Data processing method, data query method and device for distributed database

By de-repeat the execution plan of multiple operation statements and staged scheduling execution, the problem of low processing efficiency of multiple operation statements in distributed databases is solved, efficient concurrent execution is achieved, and computing performance is improved.

CN118550949BActive Publication Date: 2025-06-17ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411016657.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2025-06-17
Estimated Expiration
2044-07-26

AI Technical Summary

Technical Problem

How to efficiently process multiple operation statements and perform concurrent execution of distributed databases while ensuring correctness to improve computing efficiency and performance and reduce time-consuming.

Method used

By obtaining multiple operation statements, converting them into execution plans, and de-repeat the execution plans to form a sequence of sub-execution plan groups arranged in the stages of serial execution. Then, in units of each stage, the sub-execution plan group is scheduled to the corresponding multiple plan execution devices, and the corresponding sub-execution plan is executed by these devices.

Benefits of technology

It realizes that multiple operation statements are executed concurrently while ensuring correctness, improving the efficiency and performance of calculations and reducing time-consuming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118550949B_ABST
    Figure CN118550949B_ABST
Patent Text Reader

Abstract

Embodiments of this specification provide a data processing method, a data query method, and an apparatus for a distributed database. In the data processing method for the distributed database, a plurality of operation statements for the distributed database obtained are converted into corresponding execution plans, where each execution plan includes sub-execution plans arranged according to stages of serial execution; the obtained plurality of execution plans are de-duplicated and merged to obtain a sequence of groups of sub-execution plans arranged according to stages of serial execution; taking each stage as a unit, the group of sub-execution plans of this stage is scheduled to corresponding multiple plan execution devices; and each sub-execution plan in the corresponding group of sub-execution plans is executed by the multiple plan execution devices of this stage through division of labor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this specification generally relate to the field of computer technology, and particularly to data processing methods, data query methods, and devices for distributed databases. Background Art

[0002] With the rapid development of big data technology, distributed database technology has also been advancing day by day. When processing large-scale data, a distributed database often receives multiple operation statements for data processing such as addition, deletion, modification, and query. Therefore, how to efficiently respond to multiple operation statements for a distributed database to improve data processing capabilities is of great significance. Summary of the Invention

[0003] In view of the above, embodiments of this specification provide a data processing method, a data query method, and a device for a distributed database. Using this method and device, multiple operation statements can be concurrently executed on the premise of ensuring correctness, so as to improve the efficiency and performance of computing and reduce the time consumption.

[0004] According to one aspect of the embodiments of this specification, a data processing method for a distributed database is provided, including: obtaining multiple operation statements for a distributed database; converting each of the obtained operation statements into a corresponding execution plan, where each execution plan includes sub-execution plans arranged in stages of serial execution; performing deduplication and merging on the obtained multiple execution plans to obtain a sequence of sub-execution plan groups arranged in stages of serial execution; scheduling the sub-execution plan groups of each stage to corresponding multiple plan execution devices in units of each stage; and having the multiple plan execution devices of this stage execute each sub-execution plan in the corresponding sub-execution plan group respectively.

[0005] According to another aspect of the embodiments of this specification, a data query method is provided, including: obtaining multiple query statements for a distributed database; converting each of the obtained query statements into a corresponding query plan, where each query plan includes sub-query plans arranged in stages of serial execution; performing deduplication and merging on the obtained multiple query plans to obtain a sequence of sub-query plan groups arranged in stages of serial execution; scheduling the sub-query plan groups of each stage to corresponding multiple plan execution devices in units of each stage; having the multiple plan execution devices of this stage execute each sub-query plan in the corresponding sub-query plan group respectively; generating query results matching each query statement according to the execution results of the sub-query plans of the last stage, and feeding back the query results to the device that sent the corresponding query statement.

[0006] According to another aspect of the embodiments of the present specification, there is provided a data processing apparatus for a distributed database, including: an execution plan generation unit configured to obtain a plurality of operation statements for the distributed database; convert each of the obtained operation statements into a corresponding execution plan, where each execution plan includes sub-execution plans arranged according to the stages of serial execution; an execution plan merging unit configured to perform deduplication and merging on the obtained plurality of execution plans to obtain a sequence of sub-execution plan groups arranged according to the stages of serial execution; and an execution plan execution unit configured to, taking each stage as a unit, schedule the sub-execution plan group of this stage to corresponding multiple plan execution devices; and have the multiple plan execution devices of this stage execute each sub-execution plan in the corresponding sub-execution plan group in a division of labor manner.

[0007] According to yet another aspect of the embodiments of the present specification, there is provided a data query apparatus, including: a query plan generation unit configured to obtain a plurality of query statements for the distributed database; convert each of the obtained query statements into a corresponding query plan, where each query plan includes sub-query plans arranged according to the stages of serial execution; a query plan merging unit configured to perform deduplication and merging on the obtained plurality of query plans to obtain a sequence of sub-query plan groups arranged according to the stages of serial execution; a query plan execution unit configured to, taking each stage as a unit, schedule the sub-query plan group of this stage to corresponding multiple plan execution devices; and have the multiple plan execution devices of this stage execute each sub-query plan in the corresponding sub-query plan group in a division of labor manner; and a query result sending unit configured to generate query results matching each query statement according to the execution result of the sub-query plan of the last stage, and feedback the query results to the device that sent the corresponding query statement.

[0008] According to another aspect of the embodiments of the present specification, there is provided a data processing apparatus for a distributed database, including: at least one processor, and a memory coupled to the at least one processor, where the memory stores instructions that, when executed by the at least one processor, cause the at least one processor to execute the data processing method for the distributed database as described above.

[0009] According to another aspect of the embodiments of the present specification, there is provided a data query apparatus, including: at least one processor, and a memory coupled to the at least one processor, where the memory stores instructions that, when executed by the at least one processor, cause the at least one processor to execute the data query method as described above.

[0010] According to another aspect of the embodiments of the present specification, there is provided a computer-readable storage medium storing a computer program, which when executed by a processor implements the data processing method and / or data query method for a distributed database as described above.

[0011] According to another aspect of the embodiments of the present specification, there is provided a computer program product including a computer program, which is executed by a processor to implement the data processing method and / or data query method for a distributed database as described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] By referring to the following drawings, a further understanding of the essence and advantages of the content of the present specification can be achieved. In the drawings, similar components or features may have the same reference numerals.

[0013] Figure 1 An exemplary architecture of a data processing method, a data query method, and an apparatus for a distributed database according to an embodiment of the present specification is shown.

[0014] Figure 2 A flowchart of an example of a data processing method for a distributed database according to an embodiment of the present specification is shown.

[0015] Figure 3 A flowchart of an example of the deduplication and merging process of multiple execution plans according to an embodiment of the present specification is shown.

[0016] Figure 4 A schematic diagram of another example of a data processing method for a distributed database according to an embodiment of the present specification is shown.

[0017] Figure 5 A schematic diagram of an example of the execution process of a sub-execution plan according to an embodiment of the present specification is shown.

[0018] Figure 6 A flowchart of an example of a data query method according to an embodiment of the present specification is shown.

[0019] Figure 7 A block diagram of an example of a data processing apparatus for a distributed database according to an embodiment of the present specification is shown.

[0020] Figure 8 A block diagram of an example of a data query apparatus according to an embodiment of the present specification is shown.

[0021] Figure 9 A schematic diagram of an example of a data processing apparatus for a distributed database according to an embodiment of the present specification is shown.

[0022] Figure 10 A schematic diagram showing an example of a data query device according to an embodiment of the present specification. Detailed implementation manners

[0023] The subject matter described herein will be discussed below with reference to example embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein, and is not a limitation on the scope of protection, applicability, or examples set forth in the claims. The functions and arrangements of the elements discussed may be changed without departing from the scope of protection of the content of the embodiments of this specification. Each example may omit, substitute, or add various processes or components as needed. Additionally, the features described relative to some examples may be combined in other examples.

[0024] As used herein, the term "comprising" and its variants denote open terms, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" and "an embodiment" mean "at least one embodiment". The term "another embodiment" means "at least one other embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other definitions, whether explicit or implicit, may be included below. Unless explicitly specified in the context, the definition of a term is consistent throughout the specification.

[0025] The data processing method, data query method, and device for a distributed database according to an embodiment of the present specification will be described in detail below with reference to the accompanying drawings.

[0026] Figure 1 An exemplary architecture 100 of a data processing method, data query method, and device for a distributed database according to an embodiment of the present specification is shown.

[0027] As Figure 1 shown, the distributed database system 100 may include a plurality of storage nodes, such as storage nodes 10-1 to 10-4. The storage nodes 10-1 to 10-4 are distributed storage nodes, and each storage node may include a data computing engine and a data storage engine. In some examples, the data computing engine may be used to process various computing logics, such as numerical operations, attribute conversion operations, etc. In some examples, the data storage engine may perform queries or updates based on the stored data for computing use. In some examples, the stored data may be graph data.

[0028] It should be noted that Figure 1 the examples shown are merely illustrative. In other embodiments, the distributed database system 100 may include more or fewer storage nodes.

[0029] In some examples, when performing data processing, a storage node in a distributed database can act as a scheduling node for executing a data processing method for the distributed database. Other storage nodes in the distributed database can act as execution nodes for executing data processing tasks assigned by the scheduling node. In one example, the data processing tasks are in units of sub-execution plans. In one example, one or more plan execution devices can be respectively running on each execution node for executing data processing tasks relying on a data computing engine and a data storage engine.

[0030] It should be understood that Figure 1 all the network entities shown in are exemplary, and any other network entities may be involved in the architecture 100 according to specific application requirements.

[0031] Figure 2 The flowchart of a data processing method 200 for a distributed database according to an embodiment of the present specification is shown.

[0032] As Figure 2 shown, at 210, a plurality of operation statements for the distributed database are obtained.

[0033] In this embodiment, the operation statements for the distributed database may include statements for instructing operations such as data addition, data deletion, data modification, and data query on the distributed database. In some examples, the plurality of operation statements may come from the same client or from multiple different clients.

[0034] In some implementation manners, the distributed database may include a graph database. The operation statements may be operation statements for the graph database. In some examples, the operation statements may be represented in Gremlin language.

[0035] At 220, each of the obtained operation statements is converted into a corresponding execution plan.

[0036] In this embodiment, the execution plan can be used to describe the execution process and operation steps of the corresponding operation statement in the computing engine. When the execution plan is actually run, it will be divided into multiple stages, and each stage is executed serially, and the output of the previous stage will be used as the input of the next stage. In this embodiment, each execution plan may include sub-execution plans arranged according to the stages executed serially. In some examples, each stage may include at least one sub-execution plan. In one example, each execution plan may be represented as a sequence of sub-execution plans arranged according to the stages executed serially.

[0037] In some implementations, the execution plans corresponding to multiple operation statements can be obtained by converting each of the obtained operation statements using a computing engine of the same type. In some examples, the same type of computing engine can mean that the logic for converting operation statements into corresponding execution plans is the same.

[0038] At 230, duplicate elimination and merging are performed on the obtained multiple execution plans to obtain a sequence of sub-execution plan groups arranged according to the stages of serial execution.

[0039] In this embodiment, duplicate sub-execution plans among multiple execution plans can be merged. In some examples, the duplicate sub-execution plans can include, but are not limited to, at least one of the following: involving the same data set and the same computing logic, involving duplicate file reading operations. After duplicate elimination and merging, each sub-execution plan group at each stage of serial execution is composed of non-identical sub-execution plans, thereby forming a sequence of sub-execution plan groups arranged according to the stages of serial execution.

[0040] Figure 3 A flowchart showing an example of the duplicate elimination and merging process 300 of multiple execution plans according to an embodiment of this specification is shown.

[0041] As Figure 3 shown, at 310, the obtained multiple execution plans are aligned according to the stages of serial execution.

[0042] In this embodiment, each execution plan can be aligned according to the stages of serial execution. In some examples, the stages of serial execution can be numbered as stage 1, stage 2, and so on in sequence. The sub-execution plans at stage 1 in each execution plan can be grouped into a sub-execution plan group, and the sub-execution plans at stage 2 in each execution plan can be grouped into another sub-execution plan group, and so on. In some examples, since the number of stages of each execution plan is not necessarily exactly the same, virtual nodes can be added to the execution plan with fewer stages to make the number of stages of all execution plans the same.

[0043] At 320, the mergeable sub-execution plans belonging to the same stage of serial execution among the aligned multiple execution plans are merged to obtain a sequence of sub-execution plan groups arranged according to the stages of serial execution.

[0044] In this embodiment, the mergeable sub-execution plans can refer to sub-execution plans that meet the merging conditions, such as the duplicate sub-execution plans as described above.

[0045] In the above manner, a computing engine of the same type can be used to align the execution plan according to the serial execution stages, and then merge according to the aligned result, so as to improve the deduplication and merging efficiency of the execution plan on the basis of ensuring that the plan can be correctly executed.

[0046] In some implementation manners, each sub-execution plan in the sub-execution plan group sequence has a corresponding execution plan identifier. In these implementation manners, the execution plan identifier can be used to indicate the execution plan to which the sub-execution plan belongs. The execution plan identifier can be transmitted between different plan execution devices as the sub-execution plan is allocated. In some examples, the obtained execution plan can be numbered, and the number can be used as the execution plan identifier of each sub-execution plan included in the execution plan.

[0047] In some implementations, multiple plan execution devices for executing sub-execution plans in the same phase share the same state data. The sub-execution plans that can be merged may include at least one of the following: a sub-execution plan representing the same operation on the same data, a sub-execution plan representing sending messages from the same node in a distributed database to another same node. In some examples, multiple plan execution devices for executing multiple sub-execution plans in the same phase can be regarded as a whole and share the same state data loaded into the memory. In one example, the above state data may be data that needs to be loaded from a distributed database for processing operation statements, such as graph data. In one example, if 10 operation statements for the same database are obtained (such as statements related to read operations and query operations on database A), and there are 100 plan execution devices available to process these 10 operation statements. Then, the above 100 plan execution devices sharing the same state data can be regarded as a whole to jointly process the above 10 operation statements. At this time, for the whole of the above 100 plan execution devices, only one loading operation of database A needs to be completed. In one example, the distributed database may include node A and node B, and 3 plan execution devices can be respectively run on each node, such as A1, A2, A3, B1, B2, B3. If at stage 3, plan execution devices B1 and B3 respectively need the execution result 1 of the sub-execution plan in stage 2 and the execution result 2 of the sub-execution plan in stage 2, and the execution result 1 of the sub-execution plan in stage 2 and the execution result 2 of the sub-execution plan in stage 2 are respectively generated by plan execution devices A1 and A2. Then, the sub-execution plan for instructing plan execution device A1 to send the execution result 1 of the sub-execution plan in stage 2 to plan execution device B1 can be merged with the sub-execution plan for instructing plan execution device A2 to send the execution result 2 of the sub-execution plan in stage 2 to plan execution device B3. The merged sub-execution plan can be used to instruct sending the execution result 1 of the sub-execution plan in stage 2 and the execution result 2 of the sub-execution plan in stage 2 from node A to node B. Compared with multiple single-message transmissions before merging, the network overhead can be effectively reduced.

[0048] By sharing the same state data among multiple plan execution devices for executing sub-execution plans in the same phase, compared with the prior art of separately allocating computing resources for each operation statement (such as allocating 10 plan execution devices for processing each operation statement), since database A needs to be loaded once for each separate execution of an operation statement, this solution can significantly reduce the loading of duplicate data; and enables the same data (such as the vertex-edge data of graph data) to be reused among multiple plan execution devices, thereby optimizing the utilization of resources.

[0049] At 240, taking each stage as a unit, scheduling the sub - execution plan group of this stage to a corresponding plurality of plan execution devices; and having the plurality of plan execution devices of this stage execute each sub - execution plan in the corresponding sub - execution plan group by division of labor.

[0050] In this embodiment, the plurality of plan execution devices of each stage can be regarded as a whole to jointly execute each sub - execution plan in the sub - execution plan group of this stage. In some examples, the sub - execution plan group of this stage can be scheduled to the corresponding plurality of plan execution devices according to various load - balancing algorithms. Thus, each plan execution device can execute the received sub - execution plan. In some examples, the plurality of plan execution devices for executing the sub - execution plans in the sub - execution plan groups of different stages can be different, so as to support the elastic scaling of plan execution resources and provide a technical basis for adapting to dynamic scenarios with variable resources.

[0051] In some implementation manners, taking each stage as a unit, according to the nodes where the graph data involved in each sub - execution plan in the sub - execution plan group of this stage is located, each sub - execution plan can be scheduled to the plan execution device running on the corresponding node where it is located. In some examples, each node in the distributed graph database can be used to store at least part of the shards of the complete graph data. For each sub - execution plan in the sub - execution plan group of this stage, the node where the graph data targeted by this sub - execution plan is located can be determined as a candidate node, and then the target node can be determined from the candidate nodes according to the load balance between the candidate nodes, and this sub - execution plan can be scheduled to the plan execution device running on this target node. In one example, the distributed graph database includes 20 nodes, where nodes 1 - 5, nodes 6 - 10, and nodes 11 - 20 can be used to store sub - graph A, sub - graph B, and sub - graph C of the complete graph data respectively. At least part of the shards of sub - graph A, sub - graph B, or sub - graph C can be stored in each node. If the graph data involved in the sub - execution plan x belongs to sub - graph A, then this sub - execution plan x can be scheduled to the plan execution devices running on nodes 1 - 5. If the graph data involved in the sub - execution plan y belongs to sub - graph C, then this sub - execution plan y can be scheduled to the plan execution devices running on nodes 11 - 20. Through the above - mentioned method, the sub - execution plan can be scheduled to the node where the involved graph data is located, avoiding the transfer of graph data between different nodes and significantly reducing the data transfer cost.

[0052] Figure 4 FIG. shows a schematic diagram of an example of a data processing method 400 for a distributed database according to an embodiment of this specification.

[0053] As Figure 4As shown, operation statements 1 to... for a certain distributed database can be obtained first l . After that, each operation statement can be converted into a corresponding execution plan, obtaining execution plans 1 to... l . In an example, execution plan 1 can include, for example n serial execution stages, and execution plan l can include, for example m serial execution stages. Among them, each stage can be composed of at least one sub-execution plan. Next, the obtained execution plans 1 to... l can be de-duplicated and merged to obtain a sequence of sub-execution plan groups arranged according to the serial execution stages. Each stage can be regarded as a sub-execution plan group composed of several de-duplicated and merged sub-execution plans. In an example, the sequence of sub-execution plan groups can include k stages, where k can be the maximum number of stages in execution plans 1 to... l . In an example, the respective sub-execution plans of stage 1 can be scheduled to multiple plan execution devices for executing the sub-execution plans in stage 1. After the respective sub-execution plans of stage 1 are executed, the respective sub-execution plans of stage 2 can be scheduled to multiple plan execution devices for executing the sub-execution plans in stage 2. It can be understood that the multiple plan execution devices for executing the sub-execution plans in stage 1 (such as plan execution devices 1 to... x ) can be the same as or different from the multiple plan execution devices for executing the sub-execution plans in stage 2 (such as plan execution devices 1 to... y ), and similarly, can also be the same as or different from the multiple plan execution devices for executing the sub-execution plans in stage 3 (such as plan execution devices 1 to... z ). Each plan execution device can execute the respective sub-execution plans received. In some examples, the execution result of the sub-execution plan of the previous stage can be sent to the plan execution device for executing the corresponding sub-execution plan in the next stage.

[0054] Figure 5 shows a schematic diagram of an example of the execution process 500 of the sub-execution plan according to an embodiment of the present specification.

[0055] As Figure 5 shown, in a non-final stage, the respective sub-execution plans in the sub-execution plan group of this stage are executed by multiple plan execution devices of this stage in a division of labor, and the execution results with corresponding execution plan identifiers are sent to the corresponding plan execution devices in the next stage. In an example, in stage 1, the sub-execution plan i , ii can be scheduled to plan execution devices 1 and 2 running on node P respectively, and the sub-execution planiii, iv Dispatch to the plan execution device 2 running on node Q. Among them, the sub-execution plan i 、 ii, iii, iv can respectively have execution plan identifiers corresponding to execution plans 1, 2, 3, n In phase 2, the sub-execution plan v 、 vi can be dispatched to the plan execution device 1 running on node Q. Dispatch the sub-execution plan vii to the plan execution device 2 running on node Q. Among them, the sub-execution plan v 、 vii can respectively have execution plan identifiers corresponding to execution plans 1, n corresponding. The sub-execution plan vi can be merged from the sub-execution plans in phase 2 that have execution plan identifiers corresponding to execution plans 2 and 3, so that the sub-execution plan vi has execution plan identifiers corresponding to execution plans 2 and 3.

[0056] Thus, the plan execution devices 1 and 2 running on node P can respectively obtain the execution results of the sub-execution plan i 、 ii . And the plan execution devices 1 and 2 running on node P can respectively send the execution results of the sub-execution plans with corresponding execution plan identifiers i 、 ii to the plan execution device 1 running on node Q. The plan execution device 2 running on node Q can obtain the execution result of the sub-execution plan iv . And the plan execution device 2 running on node Q can send the execution result of the sub-execution plan with the corresponding execution plan identifier iii to the plan execution device 1 running on node Q. Thus, in phase 2, the plan execution device 1 running on node Q can execute the sub-execution plan i according to the execution result of the received sub-execution plan v , and execute the sub-execution plan ii 、 iii according to the execution results of the received sub-execution plans vi . The plan execution device 2 running on node Q can execute the sub-execution plan iv according to the execution result of the sub-execution plan obtained in the previous phase vii . And so on.

[0057] In the final stage, multiple plan execution devices in this stage execute each sub - execution plan in the sub - execution plan group of this stage in a division of labor, and summarize the execution results of each sub - execution plan according to the execution plan identifier to obtain the execution results of the execution plans corresponding to each operation statement. In one example, in stage k, the plan execution devices running on node P and node Q can execute the sub - execution plans scheduled for them respectively, and summarize the corresponding execution results according to the execution plan identifiers corresponding to each sub - execution plan to obtain the execution results of the execution plans corresponding to each operation statement, such as execution results 1~ n . In some examples, the above - mentioned summarizing of the corresponding execution results can be implemented through the MapReduce computing model.

[0058] In some implementation manners, the messages sent from the plan execution device running on the first node to the plan execution device running on the second node can be integrated, and then the integrated messages are sent to the second node. Among them, the above - mentioned messages can include the execution results with corresponding execution plan identifiers. The stage where the second node is located is the next stage of the stage where the first node is located. In one example, referring to the foregoing, the first node can be node P and the second node can be node Q. The execution results of the sub - execution plan with the execution plan identifier indicating execution plan 1 i and the execution results of the sub - execution plan with the execution plan identifier indicating execution plan 2 ii can be integrated, and the integrated messages are sent from node P to node Q so that the plan execution device running on node Q can execute the sub - execution plans of the next stage accordingly. Compared with the need to send single - line messages to downstream nodes multiple times before integration, the network overhead can be effectively reduced.

[0059] Using Figures 1 - 5 the data processing method for distributed databases disclosed in

[0060] Figure 6 By deduplicating and merging the execution plans corresponding to multiple operation statements, a sub - execution plan group for phased execution is obtained, and then in each stage, the current sub - execution plan group is scheduled to the corresponding multiple plan execution devices for execution, so that multiple operation statements can be executed concurrently on the premise of ensuring correctness, and duplicate calculations and file reads are eliminated as much as possible to improve the efficiency and performance of the calculation and reduce the time consumption.

[0061] As Figure 6 shown, at 610, multiple query statements for the distributed database are obtained.

[0062] In some examples, the distributed database may include a graph database, and the query statements may be represented in the Gremlin language. In some examples, when processing large-scale graph data, multiple Gremlin query statements that need to be executed simultaneously usually may involve the same data set, similar calculation logics, or repeated file reading operations, thus meeting the conditions for combined execution.

[0063] At 620, convert each of the obtained query statements into a corresponding query plan.

[0064] In this embodiment, each query plan includes a sequence of sub-query plans arranged according to the stages of serial execution.

[0065] At 630, perform deduplication and merging on the obtained multiple query plans to obtain a sequence of groups of sub-query plans arranged according to the stages of serial execution.

[0066] At 640, taking each stage as a unit, schedule the group of sub-query plans of this stage to the corresponding multiple plan execution devices; and have the multiple plan execution devices of this stage execute each sub-query plan in the corresponding group of sub-query plans in a division of labor manner.

[0067] It should be noted that the operations in the above steps 610 - 640 can refer to the corresponding descriptions of steps 210 - 240 in the foregoing embodiments, and will not be elaborated here.

[0068] At 650, generate query results matching each query statement according to the execution results of the sub-query plans in the last stage, and feed back the query results to the device that sent the corresponding query statement.

[0069] In this embodiment, in the last stage, the execution results of the sub-query plans obtained by each of the plan execution devices in this stage can be summarized to generate query results matching each query statement. In some examples, according to the execution plan identifiers corresponding to each sub-query plan in the group of sub-query plans in the last stage, summarize the execution results of the sub-query plans obtained by each of the plan execution devices to obtain query results matching each query statement. Then, the query results can be fed back to the device that sent the corresponding query statement.

[0070] Through the above method, a method of applying the data processing method for a distributed database to data query is provided, thereby realizing the concurrent execution of multiple queries, reducing the query time, and improving the efficiency and performance of calculation.

[0071] Figure 7 FIG. shows a block diagram of an example of a data processing device 700 for a distributed database according to an embodiment of the present specification. This device embodiment can be related to Figures 2 - 5The method embodiments shown correspond to this, and this apparatus can be specifically applied to various electronic devices.

[0072] As Figure 7 shown, a data processing apparatus 700 for a distributed database may include an execution plan generation unit 710, an execution plan merging unit 720, and an execution plan execution unit 730.

[0073] The execution plan generation unit 710 is configured to obtain multiple operation statements for the distributed database; and convert each of the obtained operation statements into a corresponding execution plan. Among them, each execution plan includes sub-execution plans arranged according to the stages of serial execution. The operations of the execution plan generation unit 710 can refer to the operations of 210-220 described above. Figure 2 described in the operations of 210-220.

[0074] In one example, the corresponding execution plan is obtained by converting each of the obtained operation statements using a homogeneous computing engine.

[0075] In one example, the distributed database includes a graph database, and the operation statements include operation statements for the graph database.

[0076] The execution plan merging unit 720 is configured to deduplicate and merge the obtained multiple execution plans to obtain a sequence of sub-execution plan groups arranged according to the stages of serial execution. The operations of the execution plan merging unit 720 can refer to the operations of 230 described above. Figure 2 described in the operations of 230.

[0077] In one example, the execution plan merging unit 720 is further configured to: align the obtained multiple execution plans according to the stages of serial execution; and merge the mergeable sub-execution plans belonging to the same stage of serial execution among the aligned multiple execution plans to obtain a sequence of sub-execution plan groups arranged according to the stages of serial execution. The above operations of the execution plan merging unit 720 can refer to the operations of 310-320 described above. Figure 3 described in the operations of 310-320.

[0078] In one example, each sub-execution plan in the sequence of sub-execution plan groups has a corresponding execution plan identifier.

[0079] The execution plan execution unit 730 is configured to, taking each stage as a unit, schedule the sub-execution plan group of this stage to corresponding multiple plan execution devices; and have the multiple plan execution devices of this stage execute each sub-execution plan in the corresponding sub-execution plan group by division of labor. The operations of the execution plan execution unit 730 can refer to the operations of 240 described above. Figure 2 described in the operations of 240.

[0080] In one example, the execution plan execution unit 730 is further configured to: according to the nodes where the graph data involved in each sub-execution plan in the sub-execution plan group of this phase is located, schedule each sub-execution plan to the plan execution device running on the corresponding node. The operations of the execution plan execution unit 730 can refer to the operations of some implementation manners of 240 described above Figure 2 described in the operations of some implementation manners of 240.

[0081] In one example, the execution plan execution unit 730 is further configured to: in a non-final phase, have multiple plan execution devices in this phase execute each sub-execution plan in the sub-execution plan group of this phase in a division-of-labor manner, and send the execution results with corresponding execution plan identifiers to the corresponding plan execution devices in the next phase; and in the final phase, have multiple plan execution devices in this phase execute each sub-execution plan in the sub-execution plan group of this phase in a division-of-labor manner; and summarize the execution results of each sub-execution plan according to the execution plan identifier to obtain the execution results of the execution plans corresponding to each operation statement. The operations of the execution plan execution unit 730 can refer to the operations described above Figure 5 described.

[0082] In one example, the execution plan execution unit 730 is further configured to: integrate the messages sent from the plan execution device running on the first node to the plan execution device running on the second node, where the messages include execution results with corresponding execution plan identifiers, and the phase where the second node is located is the next phase of the phase where the first node is located; and send the integrated messages to the second node. The operations of the execution plan execution unit 730 can refer to the operations of some implementation manners of the above Figure 5 described.

[0083] In one example, multiple plan execution devices for executing sub-execution plans in the same phase share the same piece of status data. The mergeable sub-execution plans include at least one of the following: a sub-execution plan representing the same operation on the same data, a sub-execution plan representing sending a message from the same node in the distributed database to another same node.

[0084] Figure 8 FIG. shows a block diagram of an example of a data query device 800 according to an embodiment of the present specification. This device embodiment can correspond to Figure 6 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0085] As Figure 8 shown, the data query device 800 may include a query plan generation unit 810, a query plan merge unit 820, a query plan execution unit 830, and a query result sending unit 840.

[0086] The query plan generation unit 810 is configured to obtain multiple query statements for a distributed database; and convert each of the obtained query statements into a corresponding query plan. Each of the query plans includes sub-query plans arranged in stages of serial execution. The operations of the query plan generation unit 810 can refer to the operations of 610-620 described above Figure 6 as described.

[0087] The query plan merging unit 820 is configured to perform duplicate removal and merging on the obtained multiple query plans to obtain a sequence of groups of sub-query plans arranged in stages of serial execution. The operations of the query plan merging unit 820 can refer to the operation of 630 described above Figure 6 as described.

[0088] The query plan execution unit 830 is configured to schedule the group of sub-query plans of each stage to corresponding multiple plan execution devices in units of each stage; and have the multiple plan execution devices of this stage divide the work to execute each sub-query plan in the corresponding group of sub-query plans. The operations of the query plan execution unit 830 can refer to the operation of 640 described above Figure 6 as described.

[0089] The query result sending unit 840 is configured to generate query results matching each query statement according to the execution results of the sub-query plans of the last stage, and feedback the query results to the device that sent the corresponding query statement. The operations of the query result sending unit 840 can refer to the operation of 650 described above Figure 6 as described.

[0090] As mentioned above Figures 1 to 8 , embodiments of a data processing method and apparatus for a distributed database, as well as a data query method and apparatus according to embodiments of this specification, have been described.

[0091] The data processing apparatus and data query apparatus for a distributed database according to embodiments of this specification can be implemented in hardware, or can be implemented using software or a combination of hardware and software. Taking software implementation as an example, as a logically meaningful apparatus, it is formed by the processor of the device where it is located reading the corresponding computer program instructions in the memory into the memory for operation. In the embodiments of this specification, the data processing apparatus and data query apparatus for a distributed database can be implemented using an electronic device, for example.

[0092] Figure 9 FIG. shows a schematic diagram of an example of a data processing apparatus 900 for a distributed database according to an embodiment of this specification.

[0093] As shown in Figure 9As shown, the data processing apparatus 900 for a distributed database may include at least one processor 910, a memory (e.g., non-volatile memory) 920, an internal memory 930, and a communication interface 940, and the at least one processor 910, the memory 920, the internal memory 930, and the communication interface 940 are connected together via a bus 950. The at least one processor 910 executes at least one computer-readable instruction stored or encoded in the memory (i.e., the elements implemented in software as described above).

[0094] In one embodiment, computer-executable instructions are stored in the memory, which when executed cause the at least one processor 910 to: obtain a plurality of operation statements for the distributed database; convert each of the obtained operation statements into a corresponding execution plan, where each execution plan includes sub-execution plans arranged in stages for serial execution; perform deduplication and merging on the obtained plurality of execution plans to obtain a sequence of groups of sub-execution plans arranged in stages for serial execution; and, taking each stage as a unit, schedule the group of sub-execution plans of that stage to corresponding multiple plan execution devices; and have the multiple plan execution devices of that stage execute the respective sub-execution plans in the corresponding group of sub-execution plans by division of labor.

[0095] It should be understood that the computer-executable instructions stored in the memory, when executed, cause the at least one processor 910 to perform the various operations and functions described above in the respective embodiments of this specification in combination with Figures 2 - 5 the descriptions.

[0096] Figure 10 FIG. shows a schematic diagram of an example of a data query apparatus 1000 according to an embodiment of this specification.

[0097] As Figure 10 shown, the data query apparatus 1000 may include at least one processor 1010, a memory (e.g., non-volatile memory) 1020, an internal memory 1030, and a communication interface 1040, and the at least one processor 1010, the memory 1020, the internal memory 1030, and the communication interface 1040 are connected together via a bus 1050. The at least one processor 1010 executes at least one computer-readable instruction stored or encoded in the memory (i.e., the elements implemented in software as described above).

[0098] In one embodiment, computer-executable instructions are stored in a memory, which when executed cause at least one processor 1010 to: obtain a plurality of query statements for a distributed database; convert each of the obtained query statements into a corresponding query plan, wherein each query plan includes sub-query plans arranged in stages for serial execution; perform deduplication and merging on the obtained plurality of query plans to obtain a sequence of groups of sub-query plans arranged in stages for serial execution; schedule the groups of sub-query plans of each stage to corresponding multiple plan execution devices in units of each stage; and have the multiple plan execution devices of this stage execute each sub-query plan in the corresponding group of sub-query plans in a division-of-labor manner; generate query results matching each query statement according to the execution results of the sub-query plans of the last stage, and feed back the query results to the device that sent the corresponding query statement.

[0099] It should be understood that the computer-executable instructions stored in the memory, when executed, cause at least one processor 1010 to perform the various operations and functions described above in the respective embodiments of this specification. Figure 6 described.

[0100] According to one embodiment, a program product such as a computer-readable medium is provided. The computer-readable medium may have instructions (i.e., the elements implemented in software as described above), which when executed by a computer cause the computer to perform the various operations and functions described above in the respective embodiments of this specification. Figures 1 - 6 described.

[0101] Specifically, a system or device equipped with a readable storage medium may be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and cause the computer or processor of the system or device to read and execute the instructions stored in the readable storage medium.

[0102] In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments, so the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of the present invention.

[0103] The computer program code required for the operations of various parts of this specification can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB, NET, and Python, conventional procedural programming languages such as C, Visual Basic 2003, Perl, COBOL 2002, PHP, and ABAP, dynamic programming languages such as Python, Ruby, and Groovy, or other programming languages. This program code can run on the user's computer, or run on the user's computer as an independent software package, or part of it runs on the user's computer and another part runs on a remote computer, or all of it runs on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any network form, such as a local area network (LAN) or a wide area network (WAN), or connected to an external computer (e.g., through the Internet), or in a cloud computing environment, or used as a service, such as software as a service (SaaS).

[0104] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer or the cloud via a communication network.

[0105] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0106] Not all steps and units in the above-mentioned various processes and system structure diagrams are necessary, and some steps or units can be ignored according to actual needs. The execution order of each step is not fixed and can be determined as needed. The device structures described in the above-mentioned various embodiments can be physical structures or logical structures, that is, some units may be implemented by the same physical entity, or some units may be implemented separately by multiple physical entities, or some components in multiple independent devices can be jointly implemented.

[0107] As used throughout this specification, the term "exemplary" means "serving as an example, instance, or illustration" and does not mean "preferred" or "advantageous" over other embodiments. For the purpose of providing an understanding of the described technology, the detailed description includes specific details. However, the technology may be practiced without these specific details. In some instances, well-known structures and devices are shown in block diagram form in order to avoid obscuring the concepts of the described embodiments.

[0108] The optional implementation manners of the embodiments of this specification have been described in detail above in conjunction with the accompanying drawings. However, the embodiments of this specification are not limited to the specific details in the above implementation manners. Within the scope of the technical concept of the embodiments of this specification, various simple modifications can be made to the technical solutions of the embodiments of this specification, and these simple modifications all fall within the protection scope of the embodiments of this specification.

[0109] The above description of the content of this specification is provided to enable any ordinary person skilled in the art to implement or use the content of this specification. For those of ordinary skill in the art, various modifications to the content of this specification are obvious, and the general principles defined herein can also be applied to other variations without departing from the protection scope of the content of this specification. Therefore, the content of this specification is not limited to the examples and designs described herein, but is consistent with the broadest scope that conforms to the principles and novel features disclosed herein.

Claims

1. A data processing method for a distributed database, wherein the distributed database includes a distributed graph database, comprising: Obtaining multiple operation statements for the distributed graph database; Convert each acquired operation statement into a corresponding execution plan, wherein each execution plan includes a plurality of sub-execution plans arranged according to serial execution stages; Aligning the obtained multiple execution plans according to the stages of serial execution, and merging the mergeable sub-execution plans belonging to the same stage of serial execution in the aligned multiple execution plans to obtain a sub-execution plan group sequence arranged according to the stages of serial execution; and Taking each stage as a unit, each sub-execution plan in the sub-execution plan group of the stage is respectively scheduled to multiple plan execution devices running on the nodes where the graph data involved in each sub-execution plan is located; and the multiple plan execution devices of the stage are divided into two parts to execute each sub-execution plan in the corresponding sub-execution plan group; wherein, the multiple plan execution devices used to execute the sub-execution plans in the same stage share the same graph data.

2. The data processing method according to claim 1, wherein: Each sub-execution plan in the sub-execution plan group sequence has a corresponding execution plan identifier. The multiple plan execution devices at this stage divide the work and execute each sub-execution plan in the corresponding sub-execution plan group, including: In a non-final stage, multiple plan execution devices in the stage divide the work to execute each sub-execution plan in the sub-execution plan group of the stage, and send the execution results with corresponding execution plan identifiers to the corresponding plan execution devices in the next stage; and In the final stage, multiple plan execution devices of this stage are divided into groups to execute each sub-execution plan in the sub-execution plan group of this stage; and the execution results of each sub-execution plan are summarized according to the execution plan identifier to obtain the execution results of the execution plan corresponding to each operation statement.

3. The data processing method according to claim 2, wherein: The step of sending the execution result with the corresponding execution plan identifier to the corresponding plan execution device of the next stage includes: Integrate messages sent from a plan execution device running on a first node to a plan execution device running on a second node, wherein the messages include execution results with corresponding execution plan identifiers, and a stage at which the second node is located is a stage next to a stage at which the first node is located; and The integrated message is sent to the second node.

4. The data processing method according to claim 1, wherein: The mergeable sub-execution plans include at least one of the following: a sub-execution plan representing a same operation on the same data, and a sub-execution plan representing sending a message from the same node to another same node in the distributed database.

5. A data query method, comprising: Get multiple query statements for a distributed graph database; Convert each query statement obtained into a corresponding query plan, wherein each query plan includes a plurality of sub-query plans arranged according to serial execution stages; Aligning the obtained multiple query plans according to the stages of serial execution, and merging the mergeable sub-query plans belonging to the same stage of serial execution in the aligned multiple query plans to obtain a sub-query plan group sequence arranged according to the stages of serial execution; Taking each stage as a unit, each sub-execution plan in the sub-query plan group of the stage is respectively scheduled to multiple plan execution devices running on the nodes where the graph data involved in each sub-execution plan is located; and the multiple plan execution devices of the stage are divided into groups to execute each sub-query plan in the corresponding sub-query plan group; wherein the multiple plan execution devices for executing the sub-execution plans in the same stage share the same state data; and A query result matching each query statement is generated according to the execution result of the sub-query plan in the final stage, and the query result is fed back to the device that sends the corresponding query statement.

6. A data processing device for a distributed database, wherein the distributed database includes a distributed graph database, comprising: An execution plan generating unit is configured to obtain a plurality of operation statements for the distributed graph database; Convert each acquired operation statement into a corresponding execution plan, wherein each execution plan includes a plurality of sub-execution plans arranged according to serial execution stages; an execution plan merging unit configured to align the obtained multiple execution plans according to the stages of serial execution, and merge the mergeable sub-execution plans belonging to the same stage of serial execution in the aligned multiple execution plans to obtain a sub-execution plan group sequence arranged according to the stages of serial execution; and The execution plan execution unit is configured to schedule each sub-execution plan in the sub-execution plan group of each stage to multiple plan execution devices running on the nodes where the graph data involved in each sub-execution plan is located, based on each stage; and the multiple plan execution devices of the stage divide the work and execute each sub-execution plan in the corresponding sub-execution plan group; wherein the multiple plan execution devices used to execute the sub-execution plans in the same stage share the same status data.

7. A data query device, comprising: A query plan generating unit, configured to obtain a plurality of query statements for a distributed graph database; Convert each query statement obtained into a corresponding query plan, wherein each query plan includes a plurality of sub-query plans arranged according to serial execution stages; A query plan merging unit is configured to align the obtained multiple query plans according to the stages of serial execution, and merge the mergeable sub-query plans belonging to the same stage of serial execution in the aligned multiple query plans to obtain a sub-query plan group sequence arranged according to the stages of serial execution; The query plan execution unit is configured to dispatch each sub-execution plan in the sub-query plan group of each stage to multiple plan execution devices running on the nodes where the graph data involved in each sub-execution plan is located, and the multiple plan execution devices of the stage divide the work to execute each sub-query plan in the corresponding sub-query plan group; wherein the multiple plan execution devices for executing the sub-execution plans in the same stage share the same state data; and The query result sending unit is configured to generate query results matching each query statement according to the execution result of the sub-query plan in the final stage, and feed back the query results to the device sending the corresponding query statement.

8. A data processing device for a distributed database, comprising: At least one processor, a memory coupled to the at least one processor, and a computer program stored in the memory, wherein the at least one processor executes the computer program to implement the data processing method for a distributed database as described in any one of claims 1 to 4.

9. A data query device, comprising: At least one processor, a memory coupled to the at least one processor, and a computer program stored in the memory, wherein the at least one processor executes the computer program to implement the data query method according to claim 5.

Citation Information

Patent Citations

  • Distributed database inquiry method, device, equipment and storage medium

    CN109117426A

  • Query statement detection method and device for distributed database, equipment and medium

    CN117992357A