Data query method and device for distributed database, equipment, storage medium and program product

By building a query plan evaluation model and push-down strategy, the query plan of distributed database is optimized, and the problem of slow data query speed in distributed databases is solved, and the data query efficiency and response time are improved.

CN120045587APending Publication Date: 2025-05-27CHINA SOUTHERN POWER GRID COMPANY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510171181.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The distributed database has problems such as slow data query speed during the query process because data is distributed on multiple nodes, resulting in low data query efficiency.

Method used

By building a query plan evaluation model, multiple query plans are evaluated and target query plans are determined based on the data transmission delay between every two work nodes in the distributed cluster, the computing power resources of each work node and the data storage layout information. The target query plan contains multiple query operators. Through a preset push-down strategy, these query operators are pushed down to the target working node where the data to be queryed is located for execution.

Benefits of technology

By optimizing query plan and push-down strategy, we can reduce data transmission delay and network load, improve data query efficiency, and reduce query response time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045587A_ABST
    Figure CN120045587A_ABST
Patent Text Reader

Abstract

The invention discloses a data query method, device and equipment for a distributed database, a storage medium and a program product, and the method comprises the steps: building a query plan evaluation model based on the data transmission time delay between every two working nodes in a distributed cluster, and the computing power resource and data storage layout information of each working node; under the condition that a query request is received, calling a query plan evaluation model, respectively evaluating a plurality of query plans generated based on the query request, and determining a target query plan from the plurality of query plans based on an evaluation result; determining at least one target working node of the to-be-queried data in the distributed cluster; and on the basis of a preset push-down strategy, pushing down each query operator contained in the target query plan to each target working node where the to-be-queried data is located, so that each target working node executes each query operator on the basis of the target query plan to obtain a data query result. By adopting the method, the data query efficiency of the distributed database can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and particularly to a method, apparatus, device, storage medium, and program product for data query in a distributed database. Background Art

[0002] With the rapid development of information technology, the amount of data has grown explosively. Traditional databases face many challenges in processing large-scale data. Thus, distributed databases have emerged. However, since the data in a distributed database is distributed across multiple nodes, there are problems such as slow data query speed in cross-node data interaction during the query process.

[0003] Therefore, how to improve the data query efficiency of a distributed database has become an urgent problem to be solved. Summary of the Invention

[0004] Embodiments of the present application provide a method, apparatus, device, storage medium, and program product for data query in a distributed database, which can improve the data query efficiency of the distributed database.

[0005] In a first aspect, an embodiment of the present application provides a method for data query in a distributed database, which is applied to a database management node in a distributed cluster where the distributed database is deployed. The method includes:

[0006] Based on the data transmission delay between every two working nodes in a plurality of working nodes included in the distributed cluster, the computing power resources of each working node, and the data storage layout information, construct a query plan evaluation model;

[0007] When a query request is received, call the query plan evaluation model to evaluate each of a plurality of query plans generated based on the query request, and based on the evaluation results, determine a target query plan from the plurality of query plans; the target query plan includes a plurality of query operators;

[0008] Determine at least one target working node in the distributed cluster where the data to be queried carried by the query request is located;

[0009] Based on a preset pushdown strategy, push down each query operator included in the target query plan to each target working node where the data to be queried is located, so that each target working node executes each query operator based on the target query plan to obtain a data query result.

[0010] In one embodiment, based on the evaluation results, a target query plan is determined from multiple query plans, including: based on the evaluation results, a query plan that meets the following conditions is determined from multiple query plans as the target query plan: the cost corresponding to the execution order of multiple query operators is the smallest; the data transmission volume corresponding to multiple target worker nodes is the smallest; the computing power resources required for the query tasks assigned to each target worker node are less than the computing power resources of the target worker node.

[0011] In one embodiment, the computing power resources of each worker node are determined by the following method: determining the hardware configuration information of each worker node among multiple worker nodes, where the hardware configuration information includes the number of central processing unit cores and the content capacity; based on a preset computing power resource evaluation strategy, performing weighted processing on each piece of information included in the hardware configuration information of each worker node, and based on the weighted information, obtaining the computing power resources of each worker node.

[0012] In one embodiment, the method further includes: real-time detecting the running state of each target worker node; in the case where a faulty target worker node is detected, allocating the query task corresponding to the faulty target worker node to any other target worker node except the faulty target worker node among at least one target worker node.

[0013] In one embodiment, the method further includes: in response to a query request, using a syntax analyzer to parse the query statement carried in the query request to obtain a target abstract syntax tree; extracting multiple query operators included in the query statement from the target abstract syntax tree; generating multiple query plans based on the multiple query operators included in the query statement.

[0014] In one embodiment, in response to a query request, parsing the query statement carried in the query request to obtain a target abstract syntax tree includes: in response to a query request, performing lexical analysis on the query statement carried in the query request, and based on the lexical analysis result, decomposing the query statement to obtain a decomposed query statement; performing syntax analysis on the decomposed query statement, and based on the syntax analysis result, constructing a target abstract syntax tree.

[0015] In a second aspect, the present application provides a data query device for a distributed database, which is applied to a database management node in a distributed cluster where a distributed database is deployed; the device includes:

[0016] A construction module, configured to construct a query plan evaluation model based on the data transmission delay between every two worker nodes included in the distributed cluster, the computing power resources of each worker node, and the data storage layout information.

[0017] An evaluation module, configured to, when receiving a query request, call a query plan evaluation model to respectively evaluate multiple query plans generated based on the query request, and determine a target query plan from the multiple query plans based on the evaluation results; the target query plan includes multiple query operators;

[0018] A determination module, configured to determine at least one target working node where the data to be queried carried by the query request is located in the distributed cluster;

[0019] A processing module, configured to, based on a preset push-down strategy, push each query operator included in the target query plan to each target working node where the data to be queried is located, so that each target working node executes each query operator based on the target query plan to obtain a data query result.

[0020] In a third aspect, the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0021] Build a query plan evaluation model based on the data transmission delay between every two working nodes included in the distributed cluster, the computing power resources of each working node, and the data storage layout information;

[0022] When receiving a query request, call the query plan evaluation model to respectively evaluate multiple query plans generated based on the query request, and determine a target query plan from the multiple query plans based on the evaluation results; the target query plan includes multiple query operators;

[0023] Determine at least one target working node where the data to be queried carried by the query request is located in the distributed cluster;

[0024] Based on a preset push-down strategy, push each query operator included in the target query plan to each target working node where the data to be queried is located, so that each target working node executes each query operator based on the target query plan to obtain a data query result.

[0025] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0026] Build a query plan evaluation model based on the data transmission delay between every two working nodes included in the distributed cluster, the computing power resources of each working node, and the data storage layout information;

[0027] In the case of receiving a query request, call a query plan evaluation model to evaluate multiple query plans generated based on the query request respectively, and determine a target query plan from the multiple query plans based on the evaluation results; the target query plan includes multiple query operators;

[0028] Determine at least one target working node in the distributed cluster where the data to be queried carried by the query request is located;

[0029] Based on a preset push-down strategy, push down each query operator included in the target query plan to each target working node where the data to be queried is located, so that each target working node executes each query operator based on the target query plan to obtain a data query result.

[0030] In a fifth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps:

[0031] Build a query plan evaluation model based on the data transmission delay between every two working nodes in the distributed cluster, the computing power resources of each working node, and the data storage layout information;

[0032] In the case of receiving a query request, call a query plan evaluation model to evaluate multiple query plans generated based on the query request respectively, and determine a target query plan from the multiple query plans based on the evaluation results; the target query plan includes multiple query operators;

[0033] Determine at least one target working node in the distributed cluster where the data to be queried carried by the query request is located;

[0034] Based on a preset push-down strategy, push down each query operator included in the target query plan to each target working node where the data to be queried is located, so that each target working node executes each query operator based on the target query plan to obtain a data query result.

[0035] The above data query method, device, equipment, storage medium and program product of the distributed database are applied to the database management node in the distributed cluster where the distributed database is deployed. The database management node can construct a query plan evaluation model based on the data transmission delay between every two working nodes included in the distributed cluster, the computing power resources of each working node, and the data storage layout information. When receiving a query request, the query plan evaluation model is called to evaluate each of the multiple query plans generated based on the query request, and based on the evaluation results, a target query plan is determined from the multiple query plans. The target query plan includes multiple query operators. At least one target working node where the data to be queried carried by the query request is located in the distributed cluster is determined. Based on a preset push-down strategy, each query operator included in the target query plan is pushed down to each target working node where the data to be queried is located, so that each target working node executes each query operator based on the target query plan to obtain a data query result. By adopting this method, since the construction of the query plan evaluation model is related to the data transmission delay between every two working nodes in the distributed cluster, the computing power resources of each working node, and the data storage layout information, therefore, based on the query plan evaluation model, each of the multiple query plans generated based on the query request is evaluated, and a target query plan with a smaller data transmission delay can be determined from the multiple query plans based on the evaluation results. In this way, the impact of data transmission delay on query performance can be reduced, thereby improving data query efficiency. After that, based on the preset push-down strategy, each query operator included in the target query plan with a smaller data transmission delay is pushed down to each target working node where the data to be queried is located. In this way, by pushing down each query operator to each target working node where the data to be queried is located, a large amount of unnecessary data transmission can be reduced, the network load can be reduced, and thereby the data query efficiency can be further improved. Brief Description of the Drawings

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0037] Figure 1 It is a schematic diagram of the application scenario of a data query method for a distributed database provided by an embodiment of the present application;

[0038] Figure 2 It is a schematic flowchart of a data query method for a distributed database provided by an embodiment of the present application;

[0039] Figure 3It is a schematic flowchart of another data query method for a distributed database provided by an embodiment of the present application;

[0040] Figure 4 It is a schematic structural diagram of a data query device for a distributed database provided by an embodiment of the present application;

[0041] Figure 5 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0042] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0043] The application scenario of the data query method for the distributed database provided by the embodiment of the present application will be introduced below.

[0044] Please refer to Figure 1 , Figure 1 It is a schematic diagram of an application scenario of a data query method for a distributed database provided by an embodiment of the present application. As Figure 1 shown, the distributed cluster 100 includes a database management node 101 and multiple working nodes ( Figure 1 In the figure, it is drawn taking multiple working nodes including a target working node 102, a target working node 103 and a working node 104 as an example).

[0045] Among them, the database management node 101 can construct a query plan evaluation model based on the data transmission delay between every two working nodes included in the distributed cluster, the computing power resources of each working node, and the data storage layout information. When receiving a query request, it calls the query plan evaluation model to evaluate multiple query plans generated based on the query request respectively, and based on the evaluation results, determines a target query plan from the multiple query plans. The target query plan contains multiple query operators. It determines at least one target working node where the data to be queried carried by the query request is located in the distributed cluster, for example, the target working node 102 and the target working node 103. Based on a preset push-down strategy, it pushes each query operator included in the target query plan to each target working node where the data to be queried is located, that is, the target working node 102 and the target working node 103, so that the target working node 102 and the target working node 103 execute each query operator based on the target query plan to obtain a data query result. By adopting this method, since the construction of the query plan evaluation model is related to the data transmission delay between every two working nodes in the distributed cluster, the computing power resources of each working node, and the data storage layout information, therefore, based on the query plan evaluation model, multiple query plans generated based on the query request are evaluated respectively, and a target query plan with a smaller data transmission delay can be determined from the multiple query plans based on the evaluation results. In this way, the impact of data transmission delay on query performance can be reduced, thereby improving data query efficiency. After that, based on the preset push-down strategy, each query operator included in the target query plan with a smaller data transmission delay is pushed to each target working node where the data to be queried is located. In this way, by pushing each query operator to each target working node where the data to be queried is located, a large amount of unnecessary data transmission can be reduced, the network load can be reduced, and thereby the data query efficiency can be further improved.

[0046] Optionally, both the database management node 101 and the multiple working nodes can be servers. Among them, the server mentioned here can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, etc.

[0047] Please refer to Figure 2 , Figure 2 is a schematic flowchart of a data query method for a distributed database provided by an embodiment of the present application. This method can be executed by a database management node (such as the above-mentioned database management node 101) in the distributed cluster. As Figure 2 shown, the data query method for this distributed database may include but is not limited to the following steps:

[0048] S201. Construct a query plan evaluation model based on the data transmission delays between every two working nodes included in a distributed cluster, the computing power resources of each working node, and the data storage layout information.

[0049] Among them, the data transmission delay is one of the important factors affecting the query performance of a distributed database and can be determined based on the time measurement method in network communication principles.

[0050] Optionally, the data transmission delay between every two working nodes can be determined by the database management node in the following way: send indication information to each working node, where the indication information is used to instruct the working node to send test data packets to other working nodes and measure the round-trip time; receive the data transmission delays between every two working nodes returned from each working node. In this way, by obtaining accurate data transmission delays, it is possible to avoid selecting node combinations with long data transmission paths and high delays when determining the target query plan, thereby reducing the data transmission waiting time.

[0051] Exemplarily, the database management system can regularly or under specific conditions (such as node status changes, network configuration adjustments, etc.), send test data packets of a specific size (such as 1KB, 1MB, etc.) from one working node (denoted as the sending node) to other working nodes, and record the sending time simultaneously. After receiving the data packets, other working nodes (denoted as receiving nodes) immediately return an acknowledgment packet, and the sending node records the receiving time when it receives the acknowledgment packet. The round-trip time (RTT) is the receiving time minus the sending time, and the average value of multiple measurements is taken as the estimated value of the network transmission delay. For example, in a distributed database containing three nodes, namely Node A, Node B, and Node C, Node A sends test data packets to Node B and Node C, and measures and calculates that the average transmission delay between Node A and Node B is 5ms, and the average transmission delay between Node A and Node C is 8ms.

[0052] Optionally, when every two working nodes perform data transmission, a custom communication protocol can be used for data transmission. Among them, the custom efficient communication protocol combines data compression and data encryption technologies. The data compression technology refers to compressing the transmitted data based on the redundancy and regularity of the data through a specific algorithm (such as Huffman coding, etc.) to reduce the data volume and the network transmission bandwidth requirements. The data encryption technology (such as symmetric encryption or asymmetric encryption algorithms) can ensure the security of the data during transmission and prevent the data from being stolen or tampered with. In this way, not only can the data transmission volume be reduced, thereby improving the data transmission efficiency, but also the security of the data transmission can be improved.

[0053] Exemplarily, assume that data needs to be transmitted between NodeA and NodeB. Then NodeA can first use a data compression algorithm to compress the data to be transmitted. For example, for a data block containing a large amount of repeated data or having a certain pattern, the Huffman coding algorithm can be used to convert it into a more compact coding form, reducing the storage space and transmission volume of the data. Then, an encryption algorithm is used to encrypt the compressed data to ensure the confidentiality of the data. After receiving the compressed and encrypted data from NodeA, NodeB can first perform a decryption operation to restore the compressed data, and then perform a decompression operation to obtain the original data.

[0054] Among them, the computing power resources of each worker node can reflect the ability of the worker node to process data. Therefore, by introducing data storage layout information when constructing the query plan evaluation model, the accuracy of the evaluation results obtained by the query plan evaluation model for the query plan can be improved.

[0055] Optionally, the computing power resources of each worker node can be determined based on the hardware performance metrics of the worker node. Among them, the hardware performance metrics can include, but are not limited to, the number of cores of the central processing unit (CPU) and the content capacity, etc. In this way, by determining the computing power resources of each worker node, it is beneficial to allocate computationally intensive tasks to nodes with strong computing capabilities, thereby improving the computing efficiency.

[0056] Among them, since the data storage layout information involves the database storage principle, different data storage layouts have different impacts on data reading and processing efficiency. Therefore, by introducing data storage layout information when constructing the query plan evaluation model, the accuracy of the evaluation results obtained by the query plan evaluation model for the query plan can be improved.

[0057] Among them, the data storage layout information includes row storage and column storage. Through the data storage layout information, it can be determined whether the data is organized on the storage medium in a row storage or column storage manner. Among them, row storage means that the data is continuously stored by row, which is suitable for reading and writing the entire row of data in a transaction processing scenario; column storage means that the data is continuously stored by column, and in scenarios such as data statistical analysis, the batch reading efficiency of column data is higher. For example, for a data warehouse application scenario, if column-based data analysis queries are frequently performed, a column storage layout may be more beneficial to improving query performance.

[0058] S202. When a query request is received, call the query plan evaluation model to evaluate each of the multiple query plans generated based on the query request, and based on the evaluation results, determine a target query plan from the multiple query plans.

[0059] Among them, the target query plan contains multiple query operators.

[0060] Since the constructed query plan evaluation model comprehensively considers various factors such as network transmission latency, node computing power resources (or computing capabilities), and data storage layout, the constructed query plan evaluation model can accurately evaluate the costs of different query plans. For example, due to the limited network bandwidth and high transmission latency between NodeA and NodeB, while NodeB has strong computing power, the cost model will tend to select an execution query plan that reduces cross-node data transmission and fully utilizes the computing power of NodeB.

[0061] S203. Determine at least one target working node in the distributed cluster where the data to be queried carried by the query request is located.

[0062] In an optional implementation manner, for the database management node to determine at least one target working node in the distributed cluster where the data to be queried carried by the query request is located, it may include: determining the identifier of the data to be queried included in the query statement, and based on the identifier of the data to be queried and the data distribution information of the distributed database, determining at least one target working node in the distributed cluster where the data to be queried carried by the query request is located; each target working node is at least one working node in the distributed cluster.

[0063] In some embodiments, the data distribution information of the distributed database may be determined by the database management node based on the metadata of the distributed database. Optionally, the database management node may read the information about the distribution of table data from the metadata storage area of the distributed database; based on the distribution information of the table data, determine the data distribution information of the distributed database.

[0064] Among them, the metadata of the database is stored in a specific data structure (such as records including the start address, end address, belonging table, and storage working node identifier of the data block, etc.). In this way, by recording key information such as the start address, end address, belonging table, and storage node identifier of the data block, it is possible to accurately determine the location where the data is located, that is, at least one target working node where the data is located, which is conducive to formulating a reasonable query execution plan according to the data distribution situation subsequently, such as pushing down the operators related to the data on a specific node to that node for execution, avoiding blind data transmission, and reducing network overhead. At the same time, it also helps with load balancing, reasonably allocating computing tasks to each node, and improving the resource utilization rate of the entire distributed database system.

[0065] Exemplarily, when the query involves the table "table1", the system can quickly obtain, by querying the metadata, on which working nodes the data blocks of "table1" are distributed, as well as the specific address ranges of each data block, so as to provide an accurate basis for data distribution for subsequent query optimization.

[0066] S204. Based on a preset push-down strategy, push each query operator included in the target query plan to each target working node where the data to be queried is located, so that each target working node executes each query operator based on the target query plan to obtain a data query result.

[0067] Optionally, the query operators may include, but are not limited to, filtering operators, joining operators, aggregation operators, projection operators, etc. Among them, the filtering operator can be used to filter out data that meets specific conditions; the joining operator can be used to associate data from different tables; the aggregation operator can be used to perform statistical calculations and grouping on data; the projection operator can be used to determine the columns included in the final returned result set.

[0068] Exemplarily, after pushing the filtering operator to NodeA and NodeB, start an independent filtering process on each node, filter the local data according to the condition "table1.value > 5", and store the filtering result in the local memory of the node. NodeB starts a joining process and performs a joining operation on the local table2 data and the filtered table1 data transmitted from NodeA according to the joining condition "table1.key = table2.key", and the joining result is also stored in the local memory. Finally, NodeB starts an aggregation process and executes the operation "SUM(column2) GROUP BY column1", combines the aggregation result with "column1" to form the final result set, and returns the result set to the client through a network communication protocol (such as a custom efficient communication protocol). In this way, by pushing the filtering operator to the target working nodes where the data is located, a large amount of unnecessary data transmission is reduced, and the network load is lowered. In addition, performing joining and aggregation operations on the small amount of transmitted filtered data can significantly improve the query execution efficiency and reduce the query response time. At the same time, reasonably using the stronger computing power of NodeB for joining and aggregation operations can avoid waste of node computing resources, thereby improving the resource utilization rate of the entire distributed database system.

[0069] In the embodiment of the present application, the database management node can construct a query plan evaluation model based on the data transmission delay between every two working nodes included in the distributed cluster, the computing power resources of each working node, and the data storage layout information. When receiving a query request, the query plan evaluation model is called to evaluate each of the multiple query plans generated based on the query request, and based on the evaluation results, a target query plan is determined from the multiple query plans. The target query plan includes multiple query operators. At least one target working node where the data to be queried carried by the query request is located in the distributed cluster is determined. Based on a preset push-down strategy, each query operator included in the target query plan is pushed down to each target working node where the data to be queried is located, so that each target working node executes each query operator based on the target query plan to obtain a data query result. By adopting this method, since the construction of the query plan evaluation model is related to the data transmission delay between every two working nodes in the distributed cluster, the computing power resources of each working node, and the data storage layout information, therefore, based on the query plan evaluation model, each of the multiple query plans generated based on the query request is evaluated, and a target query plan with a smaller data transmission delay can be determined from the multiple query plans based on the evaluation results. In this way, the impact of the data transmission delay on the query performance can be reduced, thereby improving the data query efficiency. After that, based on the preset push-down strategy, each query operator included in the target query plan with a smaller data transmission delay is pushed down to each target working node where the data to be queried is located. In this way, by pushing down each query operator to each target working node where the data to be queried is located, a large amount of unnecessary data transmission can be reduced, the network load can be reduced, and thus the data query efficiency can be further improved.

[0070] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of another data query method for a distributed database provided by an embodiment of the present application. Different from Figure 2 the data query method for the distributed database shown, Figure 3 the method shown also elaborates on how multiple query plans are generated. As Figure 3 shown, the data query method for the distributed database may include but is not limited to the following steps:

[0071] S301. Construct a query plan evaluation model based on the data transmission delay between every two working nodes included in the distributed cluster, the computing power resources of each working node, and the data storage layout information.

[0072] In an alternative embodiment, the relevant description of step S301 can refer to the description in the foregoing step S201, and will not be elaborated here.

[0073] S302. In the case of receiving a query request, in response to the query request, use a syntax analyzer to parse the query statement carried in the query request to obtain a target abstract syntax tree.

[0074] In an optional implementation manner, the database management node, in response to the query request, uses a syntax analyzer to parse the query statement carried in the query request to obtain a target abstract syntax tree, which may include: in response to the query request, perform lexical analysis on the query statement carried in the query request, and based on the lexical analysis result, decompose the query statement to obtain the decomposed query statement; perform syntax analysis on the decomposed query statement, and based on the syntax analysis result, construct a target abstract syntax tree.

[0075] Exemplarily, for a complex query statement, such as "SELECT column1, SUM (column2) FROM table1 JOIN table2 ON table1.key = table2.key WHERE table1.value > 5 GROUP BY column1", using a syntax analyzer according to predetermined syntax rules, it can be gradually decomposed into each syntax unit and a target abstract syntax tree can be constructed. In this target abstract syntax tree, each node represents a syntax structure, such as clauses like SELECT, FROM, WHERE, etc., and the leaf nodes are specific column names, table names, constant values, etc.

[0076] S303. Extract multiple query operators included in the query statement from the target abstract syntax tree, and generate multiple query plans based on the multiple query operators included in the query statement.

[0077] Optionally, the multiple query operators may include but are not limited to filtering operators, joining operators, aggregation operators, projection operators, etc.

[0078] Exemplarily, following the example in step S302, for example, the "WHERE" keyword is usually used to introduce a filtering condition, corresponding to a filtering operator; the "JOIN" keyword is used for table joining operations, corresponding to a joining operator; aggregation functions such as "SUM", "COUNT", etc. and the "GROUP BY" keyword are used for data aggregation operations, corresponding to aggregation operators; the columns specified after the "SELECT" keyword correspond to projection operators. In this way, by identifying multiple query operators in the query statement, the query intention and operation process can be clarified, providing a basis for determining the target query plan subsequently.

[0079] S304. Invoke a query plan evaluation model to evaluate each of the multiple query plans respectively, and based on the evaluation results, determine a target query plan from the multiple query plans.

[0080] Among them, the target query plan includes multiple query operators.

[0081] In an alternative implementation, the database management node determines a target query plan from multiple query plans based on the evaluation result, which may include: determining, based on the evaluation result, a query plan that meets the following conditions as the target query plan: the cost corresponding to the execution order of multiple query operators is the smallest; the data transmission volume corresponding to multiple target worker nodes is the smallest; the computing power resources required for the query tasks assigned to each target worker node are less than the computing power resources of the target worker node.

[0082] Among them, the smallest cost corresponding to the execution order of multiple query operators is conducive to reducing the costs of calculation and data transmission. For example, in a query statement containing operators such as a filtering operator (WHERE), a joining operator (JOIN), and an aggregation operator (GROUPBY), such as "SELECT COUNT (*) FROM table1 JOIN table2 ON table1.key = table2.key WHERE table1.value > 10 GROUP BY table1.column1", first perform the filtering of "table1.value > 10" at the node where the data is located, and then perform the joining and aggregation operations. Compared with joining first and then filtering, it can greatly reduce the amount of data participating in the joining and aggregation, thereby reducing the calculation and transmission costs.

[0083] Among them, the smallest data transmission volume corresponding to multiple target worker nodes is conducive to reducing data transmission latency. For example, in a distributed database, there are three nodes, NodeA, NodeB, and NodeC. The data of table1 is distributed on NodeA and NodeB, and the data of table2 is mainly on NodeC. The query involves a joining operation between table1 and table2. If the network transmission latency between NodeA and NodeC is relatively high, while the network transmission latency between NodeB and NodeC is relatively low, then select to transmit the data on NodeB to NodeC for the joining operation instead of transmitting from NodeA, which can reduce the data transmission latency.

[0084] Among them, the computing power resources required for the query tasks assigned to each target worker node are less than the computing power resources of the target worker node, which is conducive to improving the calculation efficiency. For example, for a query involving complex calculations, such as an aggregation operation involving multiple function nestings and a large amount of data, if the computing power resources of NodeA are more than those of NodeB, then the database management node can assign these complex calculation tasks to NodeA for execution, and NodeB executes relatively simple tasks, such as data filtering, etc., to improve the overall calculation efficiency.

[0085] S305. Determine at least one target working node where the data to be queried carried in the query request is located in the distributed cluster.

[0086] S306. Based on a preset push-down strategy, push down each query operator included in the target query plan to each target working node where the data to be queried is located, so that each target working node executes each query operator based on the target query plan to obtain a data query result.

[0087] In an alternative embodiment, the relevant descriptions of steps S304 to S306 can be respectively referred to the descriptions in the foregoing steps S202 to S204, and will not be elaborated here.

[0088] In the embodiment of the present application, when receiving a query request, the database management node can, in response to the query request, use a syntax analyzer to parse the query statement carried in the query request to obtain a target abstract syntax tree; based on the target abstract syntax tree, identify multiple query operators included in the query statement carried in the query request, so as to clarify the query intention and determine multiple query plans; then, based on a query plan evaluation model, evaluate each of the multiple query plans respectively, and based on the evaluation results, determine a target query plan with a smaller data transmission delay from the multiple query plans. In this way, the impact of data transmission delay on query performance can be reduced, and thus, the data query efficiency can be improved; afterwards, based on a preset push-down strategy, push down each query operator included in the target query plan with a smaller data transmission delay to each target working node where the data to be queried is located. In this way, by pushing down each query operator to each target working node where the data to be queried is located, a large amount of unnecessary data transmission can be reduced, the network load can be reduced, and thus, the data query efficiency can be further improved.

[0089] In an alternative embodiment, Figure 2 and Figure 3 in the data query method of the distributed database shown, the computing power resources of each working node can be determined by the database management node in the following manner: determine the hardware configuration information of each working node among multiple working nodes, where the hardware configuration information includes the number of central processing unit cores and the content capacity; based on a preset computing power resource evaluation strategy, perform weighted processing on each piece of information included in the hardware configuration information of each working node, and based on the weighted pieces of information, obtain the computing power resources of each working node.

[0090] Optionally, when the database management node performs weighted processing on each piece of information included in the hardware configuration information of each working node based on a preset computing power resource evaluation strategy, and obtains the computing power resources of each working node based on the weighted information, the computing power resources = number of CPU cores * main frequency coefficient + memory capacity * memory coefficient.

[0091] By adopting this implementation manner, the computing power resources of each working node can be accurately determined based on the hardware configuration information of each working node.

[0092] In an optional implementation manner, Figure 2 and Figure 3 in the data query method of the distributed database shown, the database management node can also detect the running status of each target working node in real time; in the case of detecting a faulty target working node, assign the query task corresponding to the faulty target working node to any other target working node except the faulty target working node among at least one target working node.

[0093] Optionally, the running status of each target working node may include but is not limited to CPU usage rate, memory usage rate, network connection status, etc.

[0094] By adopting this implementation manner, not only can the continuity of data query be ensured, the reliability and fault tolerance of data query be improved, but also data loss and query failure can be avoided.

[0095] In an optional implementation manner, Figure 2 and Figure 3 in the data query method of the distributed database shown, when each target working node receives a pushed-down query operator, it can execute the query operator in a multi-threaded parallel processing manner.

[0096] Optionally, when each target working node receives a pushed-down query operator and executes the query operator in a multi-threaded parallel processing manner, it may be to decompose the received query operator into multiple subtasks, and each subtask is associated with an independent thread; run multiple threads in parallel to execute multiple subtasks.

[0097] Exemplarily, when the filtering operator is pushed down to the target working node (such as NodeA) for execution, assuming NodeA has 4 CPU cores, according to the number of data blocks or the data distribution, the filtering task is divided into 4 or more subtasks (for example, divided into 4 regions according to the distribution of data blocks on the storage medium, and each region corresponds to a subtask). Each subtask is responsible for being processed by a separate thread, and these threads are executed in parallel on the CPU cores of NodeA. For example, for the filtering condition "column1>10", thread 1 is responsible for filtering the data in data block 1, thread 2 is responsible for filtering the data in data block 2, and so on. The threads are coordinated through shared memory or other synchronization mechanisms (such as semaphores) to ensure the correctness and integrity of data processing.

[0098] Adopting this implementation manner, based on the principle of multi-threaded concurrent execution in the operating system, multiple threads can share CPU resources at the same time and process different data blocks in parallel, thereby reducing the processing time and improving the local data processing speed. Compared with single-threaded processing, in a multi-core CPU environment, the processing time can be shortened to 1 / core number or even shorter of the original (considering the thread creation and scheduling overhead). For example, on a node with a 4-core CPU, using multi-threaded parallel processing for the filtering operator, the processing speed may be increased by 2-3 times, thereby accelerating the execution speed of the entire query. Especially when processing large-scale data, this gain is more obvious and can effectively improve the query performance of the distributed database system.

[0099] In an alternative implementation manner, Figure 3 In the data query method of the distributed database shown, the query statement carried in the query request may include a main query statement and a subquery statement. In this case, in response to the query request, the database management node uses a syntax analyzer to parse the query statement carried in the query request to obtain a target abstract syntax tree, which may include: in response to the query request, using the syntax analyzer to perform lexical analysis on the main query statement and the subquery statement included in the query statement respectively to obtain a first lexical analysis result corresponding to the main query statement and a second lexical analysis result corresponding to the subquery statement; based on the first lexical analysis result, decompose the main query statement to obtain the decomposed main query statement, and based on the second lexical analysis result, decompose the subquery statement to obtain the decomposed subquery statement; perform syntax analysis on the decomposed main query statement and, based on the syntax analysis result, construct a first abstract syntax tree corresponding to the main query statement, and perform syntax analysis on the decomposed subquery statement and, based on the syntax analysis result, construct a second abstract syntax tree corresponding to the subquery statement; based on the first abstract syntax tree and the second abstract syntax tree, obtain the target abstract syntax tree.

[0100] Next, taking the query statement "SELECT * FROM table3 WHERE column3 LIKE '% abc%' AND column4 IN (SELECT column5 FROM table4)" as an example, and assuming that the distributed cluster includes multiple worker nodes NodeC, NodeD, and NodeE, the overall process of the data query method for the distributed database provided by the embodiments of the present application will be described. Among them, "SELECT column5 FROM table4" is a subquery statement, which is used to obtain an intermediate result set, and then further filtering and joining operations are performed in the main query statement based on this intermediate result set.

[0101] First, the database management node can construct a query plan evaluation model based on the data transmission latency between every two of NodeC, NodeD, and NodeE, the computing power resources of each node, and the data storage layout (such as storage format, data compression method, etc.). When receiving a query request, in response to the query request, the syntax analyzer is used to perform lexical analysis on the main query statement and the subquery statement included in the query request respectively, to obtain the first lexical analysis result corresponding to the main query statement and the second lexical analysis result corresponding to the subquery statement. Based on the first lexical analysis result, the main query statement is decomposed to obtain the decomposed main query statement, and based on the second lexical analysis result, the subquery statement is decomposed to obtain the decomposed subquery statement. Syntactic analysis is performed on the decomposed main query statement, and based on the syntactic analysis result, the first abstract syntax tree corresponding to the main query statement is constructed, and syntactic analysis is performed on the decomposed subquery statement, and based on the syntactic analysis result, the second abstract syntax tree corresponding to the subquery statement is constructed. Based on the first abstract syntax tree and the second abstract syntax tree, the target abstract syntax tree is obtained.

[0102] Secondly, the database management node can extract the main query operator and the subquery operator included in the query statement from the target abstract syntax tree, and generate multiple query plans based on the main query operator and the subquery operator. For example, the LIKE and IN conditions are filtering operators in the main query operator, which are used to filter data according to specific conditions in the main query, and the subquery operator is used to implement embedding a subquery in the main query to obtain a more complex query result.

[0103] When receiving a query request, the query plan evaluation model is called to evaluate each of the multiple query plans respectively, and based on the evaluation results, the target query plan is determined from the multiple query plans. For example, the target query plan is to first execute the subquery and cache the result, then perform the filtering operation in the main query, and finally perform the join operation.

[0104] Then, the database management node can obtain the target working nodes where table3 and table4 are located from the data distribution information corresponding to the database metadata. For example, the data of table3 is distributed on NodeC and NodeD, and the data of table4 is distributed on NodeE.

[0105] After that, the database management node can, based on a preset push-down strategy, push down the sub-query operator to NodeE, push down the filtering operator in the main query operator to NodeC and NodeD, and push down the join operator in the main query operator to NodeD, so that NodeC, NodeD, and NodeE execute each query operator based on the target query plan (i.e., first execute the sub-query and cache the result, then perform the filtering operation in the main query, and finally perform the join operation) to obtain the data query result. That is, NodeE first executes the sub-query "SELECT column5 FROM table4" and stores the result in the local cache. Then, NodeC and NodeD respectively start the filtering process and filter the local table3 data according to the condition "column3 LIKE '% abc%' AND column4 IN (sub-query result)", and the filtered results are stored in the local memory. Finally, NodeD starts the join process, obtains the sub-query result from NodeE, and performs a join operation with the locally filtered result to obtain the data query result. Optionally, NodeE can also return the qualified data query result set to the database management node through the network.

[0106] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of the steps or stages in other steps or other steps.

[0107] Based on the same inventive concept, an embodiment of this application further provides a data query device for a distributed database for implementing the data query method of the distributed database involved above. The implementation solution provided by this device for solving problems is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the data query device for the distributed database provided below can refer to the limitations on the data query method for the distributed database in the foregoing, and will not be elaborated here.

[0108] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a data query device for a distributed database provided by an embodiment of this application. As Figure 4 shown, this device is applied to a database management node in a distributed cluster where a distributed database is deployed; this device may include but is not limited to:

[0109] A construction module 401, configured to construct a query plan evaluation model based on the data transmission delay between every two working nodes included in the distributed cluster, the computing power resources of each working node, and the data storage layout information;

[0110] An evaluation module 402, configured to, when receiving a query request, call the query plan evaluation model to evaluate each of multiple query plans generated based on the query request, and determine a target query plan from the multiple query plans based on the evaluation results; the target query plan includes multiple query operators;

[0111] A determination module 403, configured to determine at least one target working node where the data to be queried carried by the query request is located in the distributed cluster;

[0112] A processing module 404, configured to push each query operator included in the target query plan to each target working node where the data to be queried is located based on a preset push-down strategy, so that each target working node executes each query operator based on the target query plan to obtain a data query result.

[0113] In one embodiment, when the evaluation module 402 determines a target query plan from multiple query plans based on the evaluation results, it is specifically configured to: determine, as the target query plan, a query plan that meets the following conditions from the multiple query plans: the cost corresponding to the execution order of multiple query operators is the smallest; the data transmission volume corresponding to multiple target working nodes is the smallest; the computing power resources required for the query tasks assigned to each target working node are less than the computing power resources of this target working node.

[0114] In one embodiment, the determination module 402 is further configured to: determine, for each of the multiple working nodes, the hardware configuration information thereof, where the hardware configuration information includes the number of central processing unit cores and the content capacity; based on a preset computing power resource evaluation policy, perform weighted processing on each piece of information included in the hardware configuration information of each working node, and based on the weighted pieces of information, obtain the computing power resources of each working node.

[0115] In one embodiment, the apparatus further includes a detection module, configured to detect the running state of each target working node in real time; in the case where a faulty target working node is detected, allocate the query task corresponding to the faulty target working node to any one of the other target working nodes except the faulty target working node among at least one target working node.

[0116] In one embodiment, the apparatus further includes a parsing module, an extraction module, and a generation module. Among them, the parsing module is configured to, in response to a query request, parse the query statement carried in the query request by using a syntax analyzer to obtain a target abstract syntax tree; the extraction module is configured to extract multiple query operators included in the query statement from the target abstract syntax tree; the generation module is configured to generate multiple query plans based on the multiple query operators included in the query statement.

[0117] In one embodiment, when the parsing module is configured to, in response to a query request, parse the query statement carried in the query request to obtain a target abstract syntax tree, it is specifically configured to: in response to the query request, perform lexical analysis on the query statement carried in the query request, and based on the lexical analysis result, decompose the query statement to obtain the decomposed query statement; perform syntax analysis on the decomposed query statement, and based on the syntax analysis result, construct the target abstract syntax tree.

[0118] Each module in the above data query apparatus of the distributed database can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the terminal device in the form of hardware, or stored in the memory in the terminal device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0119] In an exemplary embodiment, the embodiment of the present application provides a computer device, which may be a server, and its internal structure diagram may be as Figure 5As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a data query method for a distributed database.

[0120] Those skilled in the art can understand that Figure 5 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0121] In an exemplary embodiment, the present application provides a computer device, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements the steps in the above-mentioned data query method for a distributed database.

[0122] In an exemplary embodiment, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in the above-mentioned data query method for a distributed database.

[0123] In an exemplary embodiment, the present application provides a computer program product, including a computer program. When the computer program is executed by the processor, it implements the steps in the above-mentioned data query method for a distributed database.

[0124] It should be noted that the data involved in the present application (including but not limited to the data transmission delay between every two working nodes, the computing power resources and data storage layout information of each working node, query requests, target query plans, etc.) are all information and data authorized by users or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0125] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0126] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.

[0127] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A data query method for a distributed database, characterized in that: Applied to a database management node in a distributed cluster where a distributed database is deployed; the method comprises: Constructing a query plan evaluation model based on the data transmission delay between each two working nodes in the multiple working nodes included in the distributed cluster, the computing resources and data storage layout information of each working node; In the case of receiving a query request, calling the query plan evaluation model, respectively evaluating multiple query plans generated based on the query request, and determining a target query plan from the multiple query plans based on the evaluation results; the target query plan includes multiple query operators; Determine at least one target working node in the distributed cluster where the to-be-queried data carried in the query request is located; Based on the preset push-down strategy, each query operator included in the target query plan is pushed down to each target working node where the data to be queried is located, so that each target working node executes each query operator based on the target query plan to obtain a data query result.

2. The method according to claim 1, characterized in that The step of determining a target query plan from the plurality of query plans based on the evaluation result includes: Based on the evaluation result, a query plan that satisfies the following conditions is determined from the multiple query plans as a target query plan: The execution order of the plurality of query operators corresponds to the minimum cost; The data transmission volume corresponding to the multiple target working nodes is the smallest; The computing resources required for the query tasks allocated to each of the target working nodes are less than the computing resources of the target working node.

3. The method according to claim 1, characterized in that The computing resources of each of the working nodes are determined in the following way: Determine the hardware configuration information of each of the plurality of working nodes, the hardware configuration information including the number of CPU cores and content capacity; Based on a preset computing power resource evaluation strategy, each item of information included in the hardware configuration information of each working node is weighted, and based on the weighted information, the computing power resources of each working node are obtained.

4. The method according to claim 1, characterized in that: The method further comprises: Detecting the operating status of each target working node in real time; In the case where a faulty target working node is detected, the query task corresponding to the faulty target working node is allocated to any other target working node among at least one of the target working nodes except the faulty target working node.

5. The method according to claim 1, characterized in that The method further comprises: In response to the query request, a syntax analyzer is used to parse the query statement carried in the query request to obtain a target abstract syntax tree; Extracting multiple query operators included in the query statement from the target abstract syntax tree; Based on the multiple query operators included in the query statement, multiple query plans are generated.

6. The method according to claim 5, characterized in that The step of parsing the query statement carried in the query request in response to the query request to obtain a target abstract syntax tree includes: In response to a query request, performing a lexical analysis on a query statement carried in the query request, and decomposing the query statement based on the lexical analysis result to obtain a decomposed query statement; The decomposed query statement is parsed, and a target abstract syntax tree is constructed based on the parsing result.

7. A data query device for a distributed database, characterized in that: The device is applied to a database management node in a distributed cluster where a distributed database is deployed; the device comprises: A construction module, used to construct a query plan evaluation model based on the data transmission delay between each two working nodes in the multiple working nodes included in the distributed cluster, the computing resources and data storage layout information of each of the working nodes; An evaluation module, configured to, upon receiving a query request, call the query plan evaluation model, respectively evaluate multiple query plans generated based on the query request, and determine a target query plan from the multiple query plans based on the evaluation results; the target query plan includes multiple query operators; A determination module, used to determine at least one target working node where the to-be-queried data carried in the query request is located in the distributed cluster; A processing module is used to push down each query operator included in the target query plan to each target working node where the data to be queried is located based on a preset push-down strategy, so that each target working node executes each query operator based on the target query plan to obtain a data query result.

8. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method according to any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Query request processing method and device, electronic equipment, medium and program product

    CN120631941A

  • Query request processing method and device, electronic equipment, medium and program product

    CN120631941B