Distributed database query method, related apparatus and medium

CN122838480APending Publication Date: 2026-09-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510363906.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

相关技术中,对于同一查询中的多个子任务,无论实际上逻辑串行还是并行,都串行执行,且在执行每个子任务前,向存储节点请求数据快照,根据该数据快照进行该子任务的查询,造成分布式数据库查询效率低下,网络开销大

Benefits of technology

[0101]本公开实施例中,对于分布式数据查询任务中的一条查询语句分解得到的多个子任务,分配同一查询标识。该查询标识对应于第一存储数据快照。该第一存储数据快照是在对多个子任务进行查询之前,就与查询标识对应存储,从而形成多个预存储的查询标识与多个存储数据快照的映射关系,映射关系用于在多个预存储的查询标识与多个存储数据快照之间映射,第一存储数据快照是多个存储数据快照中的一个。由于多个子任务分配的是同一查询标识,因此,针对各个子任务,基于多个预存储的查询标识与多个存储数据快照的映射关系,获取到查询标识对应的第一存储数据快照时,请求到的是同一第一存储数据快照。这样,多个子任务在执行上都是依据第一存储数据快照,都是依据各个数据的同一数据版本进行任务执行的,多个子任务中逻辑上并行的子任务就可以并行执行而不会发生错误。因为同一个查询语句的多个子任务各自的查询标识是一致的,根据同样的查询标识获取到的第一存储数据快照是一致的,所以对于同一个查询语句就可以在子任务执行之初就获取到第一存储数据快照,使各个子任务在执行时沿用获取到的第一存储数据快照,从而大大提高了分布式数据库查询效率,减小了对网络资源的占用。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838480A_ABST
    Figure CN122838480A_ABST
Patent Text Reader

Abstract

The present disclosure provides a distributed database query method, related devices and media. The method comprises: receiving, by a storage node in a distributed database, a plurality of sub-tasks from a computing node; assigning the plurality of sub-tasks with a same query identifier; for each sub-task, obtaining a first storage data snapshot corresponding to the query identifier based on a pre-stored mapping relationship, the mapping relationship being used to map between a plurality of pre-stored query identifiers and a plurality of storage data snapshots, the first storage data snapshot being one of the plurality of storage data snapshots; executing the sub-task based on the first storage data snapshot to obtain a sub-task processing result; and returning the sub-task processing result of each sub-task to the computing node of the distributed database, so that the computing node generates a processing result of the distributed data query task. The present disclosure can improve the efficiency of distributed database query and reduce the occupation of network resources. The present disclosure can be applied to various scenarios such as databases and cloud computing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of databases, and in particular to a distributed database query method, related apparatus and medium. Background Technology

[0002] Distributed databases typically consist of one or more compute nodes and one or more storage nodes. Each storage node can store a portion of the data. Multiple compute nodes can access the same storage node's data simultaneously. If one compute node is reading data from a storage node while another compute node is writing data to that storage node, a read error may occur. Therefore, it is common practice to obtain a snapshot of the storage node's data before querying it. A data snapshot represents the state of the data stored on the storage node at a specific point in time. After obtaining the data snapshot, when querying data from the storage node, the snapshot is used to determine the actual version of data that should be read, preventing read errors caused by another compute node writing data during the reading process.

[0003] Data snapshots come in two types: globally consistent data snapshots and weakly consistent data snapshots. For weakly consistent data snapshots, even queries on the same storage node may be divided into multiple subtasks. These subtasks may be logically parallel, logically sequential, or partially parallel and partially sequential. In related technologies, for multiple subtasks within the same query, regardless of whether they are actually logically sequential or parallel, they are executed sequentially. Furthermore, before executing each subtask, a data snapshot is requested from the storage node, and the query for that subtask is performed based on this snapshot. This results in low query efficiency and high network overhead in distributed databases. Summary of the Invention

[0004] This disclosure provides a distributed database query method, related apparatus, and medium, which can improve the efficiency of distributed database queries and reduce the occupation of network resources.

[0005] According to one aspect of this disclosure, a distributed database query method is provided, the distributed database comprising computing nodes and storage nodes, the method being executed by the storage nodes, the method comprising:

[0006] The computing node receives multiple subtasks, which are obtained by the computing node from the decomposition of the query statement in the distributed data query task.

[0007] Assign the same query identifier to the multiple subtasks;

[0008] For each subtask, based on a pre-stored mapping relationship, a first stored data snapshot corresponding to the query identifier is obtained. The mapping relationship is used to map multiple pre-stored query identifiers to multiple stored data snapshots, and the first stored data snapshot is one of the multiple stored data snapshots.

[0009] The subtask is executed based on the first stored data snapshot to obtain the subtask processing result;

[0010] The processing results of each subtask are returned to the computing node so that the computing node can generate the processing results of the distributed data query task.

[0011] According to one aspect of this disclosure, a distributed database query method is provided, the distributed database comprising compute nodes and storage nodes, the storage nodes comprising processing threads and a shared snapshot manager, the method being executed by the compute nodes, the method comprising:

[0012] In response to a query request from a target user, a distributed data query task is obtained, wherein the distributed data query task includes a query statement;

[0013] For the query statement in the distributed data query task, the query statement is decomposed into multiple sub-tasks;

[0014] The multiple subtasks corresponding to the query statement are sent to the processing thread of the storage node, so that the processing thread assigns the same query identifier to the multiple subtasks corresponding to the query statement, obtains the first storage data snapshot corresponding to the query identifier from the shared snapshot manager based on the query identifier, and executes the multiple subtasks based on the first storage data snapshot to obtain the subtask processing result of each subtask;

[0015] Receive the subtask processing results of each subtask of the query statement returned by the processing thread;

[0016] Based on the processing results of each subtask of the query statement, the processing result of the distributed data query task is generated.

[0017] According to one aspect of this disclosure, a distributed database query apparatus is provided, the distributed database comprising computing nodes and storage nodes, the method being executed by the storage nodes, the apparatus comprising:

[0018] A receiving unit is configured to receive multiple subtasks from the computing node, wherein the multiple subtasks are obtained by the computing node from the decomposition of the query statement in the distributed data query task;

[0019] An allocation unit is used to allocate the same query identifier to the multiple subtasks;

[0020] The acquisition unit is used to acquire, for each subtask, a first storage data snapshot corresponding to the query identifier based on a pre-stored mapping relationship, wherein the mapping relationship is used to map between multiple pre-stored query identifiers and multiple storage data snapshots, and the first storage data snapshot is one of the multiple storage data snapshots;

[0021] An execution unit is configured to execute the subtask based on the first stored data snapshot to obtain the subtask processing result;

[0022] The return unit is used to return the processing results of each subtask to the computing node so that the computing node can generate the processing results of the distributed data query task.

[0023] Optionally, the storage node includes a processing thread and a shared snapshot manager, the shared snapshot manager being used to store and maintain the mapping relationship;

[0024] The acquisition unit is used for:

[0025] Through the processing thread, for each subtask, the query identifier corresponding to the subtask is sent to the shared snapshot manager;

[0026] The shared snapshot manager returns the first stored data snapshot corresponding to the query identifier to the processing thread.

[0027] Optionally, the distributed database query device further includes a first storage unit, the first storage unit comprising:

[0028] The sending module is configured to, if the distributed data query task is a single query statement task, send the compute node identifier of the compute node, the session identifier of the single query statement task, and the query identifier of the first subtask to the shared snapshot manager through the processing thread for the first subtask among the plurality of subtasks;

[0029] The first generation module is used to generate a first storage data snapshot based on the compute node identifier and the session identifier through the shared snapshot manager, and store the first storage data snapshot in the shared snapshot manager corresponding to the query identifier;

[0030] The allocation unit is used for:

[0031] The processing thread assigns the query identifier of the first subtask to subsequent subtasks following the first subtask.

[0032] Optionally, the processing thread includes multiple processing threads corresponding to the multiple subtasks, and the multiple processing threads are parallel;

[0033] The sending module is used for:

[0034] Through the processing thread corresponding to the first subtask, the compute node identifier, the session identifier, and the query identifier of the first subtask are sent to the shared snapshot manager for the first subtask;

[0035] The allocation unit is used for:

[0036] The query identifier of the first subtask is assigned to the subsequent subtask by the processing thread corresponding to the subsequent subtask after the first subtask.

[0037] Optionally, the first generation module is used to:

[0038] The shared snapshot manager generates a first storage data snapshot based on the compute node identifier and the session identifier.

[0039] Perform a digest operation on the query identifier to obtain a digest result of the query identifier;

[0040] The first stored data snapshot is stored in the shared snapshot manager in correspondence with the summary result of the query identifier.

[0041] Optionally, returning the first stored data snapshot corresponding to the query identifier of each subtask to the processing thread through the shared snapshot manager includes:

[0042] If the shared snapshot manager stores a historical storage data snapshot corresponding to the query identifier, the historical storage data snapshot is returned to the processing thread and used as the first storage data snapshot corresponding to the query identifier.

[0043] If the shared snapshot manager stores historical storage data snapshots, but the historical storage data does not correspond to the query identifier, a first storage data snapshot corresponding to the query identifier is generated, and the generated first storage data snapshot is stored in the shared snapshot manager corresponding to the query identifier to replace the historical storage data snapshot;

[0044] If the shared snapshot manager does not store any historical storage data snapshots, a first storage data snapshot corresponding to the query identifier is generated, and the generated first storage data snapshot is stored in the shared snapshot manager corresponding to the query identifier.

[0045] Optionally, the distributed database query device further includes a counting unit, the counting unit being used for:

[0046] If the shared snapshot manager stores a historical storage data snapshot corresponding to the query identifier, it returns the historical storage data snapshot to the processing thread, uses the historical storage data snapshot as the first storage data snapshot corresponding to the query identifier, and increments the first count by 1. The first count represents the number of subtasks that are using the first storage data snapshot corresponding to the query identifier.

[0047] After returning the processing results of each subtask to the computing node, the first count is decremented by 1.

[0048] Optionally, if the distributed data query task contains multiple query statements and the isolation level of the distributed data query task is a first type of isolation level, the multiple subtasks are multiple first subtasks obtained by decomposing the first query statement among the multiple query statements.

[0049] The allocation unit is used for:

[0050] The processing thread assigns a first query identifier corresponding to the first query statement to the plurality of first subtasks;

[0051] The distributed database query device further includes a first processing unit, which is used for:

[0052] The processing thread receives multiple second subtasks derived from the decomposition of the second query statement following the first query statement from the computing node.

[0053] The processing thread assigns the first query identifier to the plurality of second subtasks;

[0054] Through the processing thread, for each of the plurality of second subtasks, the first query identifier is sent to the shared snapshot manager;

[0055] The shared snapshot manager returns the first stored data snapshot corresponding to the first query identifier to the processing thread;

[0056] The processing thread executes the second subtask based on the first stored data snapshot to obtain the processing result of the second subtask.

[0057] The processing thread returns the processing results of each second subtask to the computing node.

[0058] Optionally, the distributed database query device further includes a second storage unit, the second storage unit being used for:

[0059] In response to receiving a distributed data query task start request from the computing node, the computing node identifier of the computing node, the session identifier of the distributed data query task, and the first query identifier are sent to the shared snapshot manager;

[0060] The first storage data snapshot is generated based on the compute node identifier and the session identifier through the shared snapshot manager, and the first storage data snapshot is stored in the shared snapshot manager corresponding to the first query identifier.

[0061] Optionally, the distributed database query device further includes a first release unit, which is used to:

[0062] If the subtask processing result is the subtask processing result of multiple subtasks of the last query statement in the distributed data query task, a release request is sent to the shared snapshot manager so that the shared snapshot manager releases the stored first storage data snapshot.

[0063] Optionally, if the distributed data query task contains multiple query statements and the isolation level of the distributed data query task is the second type of isolation level, the multiple subtasks are multiple first subtasks obtained by decomposing the first query statement among the multiple query statements.

[0064] The allocation unit is used for:

[0065] The processing thread assigns a first query identifier corresponding to the first query statement to the plurality of first subtasks;

[0066] The distributed database query device further includes a second processing unit, the second processing unit being used for:

[0067] The processing thread receives multiple second subtasks derived from the decomposition of the second query statement following the first query statement from the computing node.

[0068] The processing thread assigns a second query identifier corresponding to the second query statement to the plurality of second subtasks.

[0069] Through the processing thread, for each of the plurality of second subtasks, the second query identifier is sent to the shared snapshot manager;

[0070] The shared snapshot manager returns the second storage data snapshot corresponding to the second query identifier to the processing thread;

[0071] The processing thread executes the second subtask based on the second stored data snapshot to obtain the processing result of the second subtask;

[0072] The processing thread returns the processing results of each second subtask to the computing node.

[0073] Optionally, the distributed database query device further includes a third storage unit, which is used for:

[0074] In response to receiving a distributed data query task start request from the computing node, the computing node identifier of the computing node, the session identifier of the distributed data query task, and the first query identifier are sent to the shared snapshot manager;

[0075] The first storage data snapshot is generated based on the compute node identifier and the session identifier through the shared snapshot manager, and the first storage data snapshot is stored in the shared snapshot manager corresponding to the first query identifier;

[0076] The distributed database query device further includes a third processing unit, which is used for:

[0077] The compute node identifier of the compute node, the session identifier of the distributed data query task, and the second query identifier corresponding to the second query statement are sent to the shared snapshot manager.

[0078] The shared snapshot manager generates a second storage data snapshot based on the compute node identifier and the session identifier, and stores the second storage data snapshot in the shared snapshot manager in correspondence with the second query identifier.

[0079] Optionally, the distributed database query device further includes a second release unit, the second release unit being used for:

[0080] Send a release request to the shared snapshot manager so that the shared snapshot manager releases the first or second stored data snapshot.

[0081] Optionally, the storage node includes a processing thread and a shared snapshot manager;

[0082] The distributed database query device further includes a first exception handling unit, which is used for:

[0083] In response to the detection of a user connection interruption, the connection between the storage node and the computing node is disconnected via the processing thread;

[0084] The processing thread decrements the first count corresponding to the first storage data snapshot by 1 and sends a snapshot release instruction to the shared snapshot manager, so that the shared snapshot manager moves the first storage data snapshot to the snapshot list according to the snapshot release instruction, and releases the first storage data snapshot when the first count corresponding to the first storage data snapshot is 0.

[0085] Optionally, the storage node includes a processing thread and a shared snapshot manager;

[0086] The distributed database query device further includes a second exception handling unit, which is used for:

[0087] In response to detecting an abnormal restart or shutdown of the computing node, the connection between the storage node and the computing node is disconnected via the processing thread;

[0088] The processing thread decrements the first count corresponding to the first storage data snapshot by 1 and sends a snapshot release instruction to the shared snapshot manager, so that the shared snapshot manager moves the first storage data snapshot to the snapshot list according to the snapshot release instruction, and releases the first storage data snapshot when the first count corresponding to the first storage data snapshot is 0.

[0089] According to one aspect of this disclosure, a distributed database query apparatus is provided, the distributed database comprising compute nodes and storage nodes, the storage nodes comprising processing threads and a shared snapshot manager, the apparatus being executed by the compute nodes, the apparatus comprising:

[0090] The acquisition module is used to acquire distributed data query tasks in response to query requests from target users, wherein the distributed data query tasks include query statements;

[0091] The decomposition module is used to decompose the query statement in the distributed data query task into multiple sub-tasks.

[0092] The processing module is used to send the plurality of subtasks corresponding to the query statement to the processing thread of the storage node, so that the processing thread assigns the same query identifier to the plurality of subtasks corresponding to the query statement, obtains a first storage data snapshot corresponding to the query identifier from the shared snapshot manager based on the query identifier, and executes the plurality of subtasks based on the first storage data snapshot to obtain the subtask processing result of each subtask.

[0093] A receiving module is used to receive the subtask processing results of each subtask of the query statement returned by the processing thread;

[0094] The second generation module is used to generate the processing result of the distributed data query task based on the processing results of the subtasks of each subtask of the query statement.

[0095] Optionally, the query statement may consist of multiple statements; the processing module is used for:

[0096] Identify the isolation level of the distributed data query task;

[0097] The multiple subtasks corresponding to the query statement and the isolation level are sent to the processing thread of the storage node.

[0098] According to one aspect of this disclosure, an electronic device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the distributed database query method as described above.

[0099] According to one aspect of this disclosure, a computer-readable storage medium is provided, the storage medium storing a computer program that, when executed by a processor, implements the distributed database query method described above.

[0100] According to one aspect of this disclosure, a computer program product is provided, comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the distributed database query method as described above.

[0101] In this embodiment of the disclosure, for multiple subtasks derived from a single query statement in a distributed data query task, the same query identifier is assigned. This query identifier corresponds to a first stored data snapshot. This first stored data snapshot is stored in relation to the query identifier before querying the multiple subtasks, thus forming a mapping relationship between multiple pre-stored query identifiers and multiple stored data snapshots. This mapping relationship is used to map between the multiple pre-stored query identifiers and the multiple stored data snapshots, and the first stored data snapshot is one of the multiple stored data snapshots. Since the multiple subtasks are assigned the same query identifier, when obtaining the first stored data snapshot corresponding to the query identifier for each subtask based on the mapping relationship between the multiple pre-stored query identifiers and the multiple stored data snapshots, the same first stored data snapshot is requested. In this way, the execution of multiple subtasks is based on the first stored data snapshot, and all tasks are executed based on the same data version. Logically parallel subtasks can be executed in parallel without errors. Because the query identifiers of multiple subtasks of the same query statement are consistent, and the first storage data snapshots obtained based on the same query identifier are consistent, the first storage data snapshot can be obtained at the beginning of the execution of the subtask for the same query statement. This allows each subtask to use the obtained first storage data snapshot during execution, thereby greatly improving the query efficiency of the distributed database and reducing the occupation of network resources.

[0102] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objectives and other advantages of this disclosure may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0103] The accompanying drawings are provided to further understand the technical solutions of this disclosure and constitute a part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.

[0104] Figure 1 This is an architecture diagram of the system to which the distributed database query method according to the embodiments of this disclosure is applied;

[0105] Figures 2A-2B A schematic diagram illustrating the application of the distributed database query method according to an embodiment of the present disclosure in a data query scenario is shown.

[0106] Figure 3 This is a flowchart of a distributed database query method according to an embodiment of the present disclosure;

[0107] Figure 4This is a flowchart illustrating the generation of a first storage data snapshot according to an embodiment of the present disclosure;

[0108] Figure 5 This is a schematic diagram illustrating the implementation process of storing a query identifier and a corresponding snapshot of stored data according to an embodiment of this disclosure;

[0109] Figure 6 This is a flowchart illustrating the storage of a query identifier corresponding to a stored data snapshot according to an embodiment of this disclosure;

[0110] Figure 7 This is a schematic diagram illustrating the implementation process of storing a query identifier and a corresponding snapshot of stored data according to another embodiment of this disclosure;

[0111] Figure 8 This is a schematic diagram of the overall interaction of data query when the distributed data query task is a single query statement task, according to an embodiment of the present disclosure;

[0112] Figure 9 This is a flowchart of generating a first stored data snapshot corresponding to a first query identifier according to this disclosure;

[0113] Figure 10 This is a schematic diagram of the overall interaction of data query when the distributed data query task is a multi-query statement task according to an embodiment of the present disclosure;

[0114] Figure 11A This is a flowchart illustrating the generation of a first storage data snapshot according to another embodiment of the present disclosure;

[0115] Figure 11B This is a flowchart illustrating the generation of a second storage data snapshot according to an embodiment of the present disclosure;

[0116] Figure 12 This is a schematic diagram of the overall interaction of data query when the distributed data query task is a multi-query statement task according to another embodiment of the present disclosure;

[0117] Figure 13 This is a schematic diagram illustrating the process of cleaning up a storage data snapshot of an abnormal connection according to an embodiment of the present disclosure;

[0118] Figure 14 This is a schematic diagram illustrating the process of cleaning up a storage data snapshot of an abnormal connection according to another embodiment of this disclosure;

[0119] Figure 15 This is a flowchart of a distributed database query method according to another embodiment of the present disclosure;

[0120] Figure 16This is a block diagram of a distributed database query apparatus according to an embodiment of the present disclosure;

[0121] Figure 17 This is a block diagram of a distributed database query apparatus according to another embodiment of the present disclosure;

[0122] Figure 18 This is a terminal structure diagram of a distributed database query method according to an embodiment of the present disclosure;

[0123] Figure 19 This is a server architecture diagram of a distributed database query method according to an embodiment of the present disclosure. Detailed Implementation

[0124] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this disclosure.

[0125] Data snapshots come in two types: globally consistent data snapshots and weakly consistent data snapshots. For globally consistent data snapshots, compute nodes must first obtain a globally consistent timestamp from the centralized service node, and then query data from each storage node according to that timestamp. Since each storage node generates a data snapshot with that globally consistent timestamp, the consistency of the time base of the data returned by each storage node is guaranteed. However, due to the need to request a global timestamp, real-time performance is poor. Weakly consistent data snapshots do not require obtaining a global timestamp beforehand; instead, they obtain the local timestamp from each storage node and generate corresponding data snapshots on each storage node according to their local timestamps. Because the local timestamps of each storage node are different, the time base of the data snapshots generated by each storage node may be inconsistent, resulting in lower query accuracy, but stronger real-time performance.

[0126] For weakly consistent data snapshots, even queries targeting the same storage node may be divided into multiple subtasks. These subtasks may be logically parallel, logically sequential, or partially parallel and partially sequential. In related technologies, multiple subtasks within the same query are executed sequentially, regardless of whether they are actually logically sequential or parallel. Furthermore, before executing each subtask, a data snapshot is requested from the storage node, and the query for that subtask is performed based on this snapshot. This results in low query efficiency and high network overhead in distributed databases.

[0127] The system architecture and scenarios in which this disclosure is applied are described below.

[0128] Figure 1This is a system architecture diagram of the distributed database query method applied according to embodiments of the present disclosure. It includes an object terminal 140, an Internet 130, a gateway 120, a server 110, etc.

[0129] The object terminal 140 can take various forms, including desktop computers, laptops, PDAs (personal digital assistants), tablets, mobile phones, in-vehicle terminals, home theater terminals, smart TVs, and dedicated terminals. Furthermore, it can be a single device or a collection of multiple devices. The object terminal 140 can communicate with the Internet 130 via wired or wireless means to exchange data. The object terminal 140 is used by objects to submit query requests to a distributed database to retrieve relevant data from the distributed database 110.

[0130] Server 110 refers to a computer system that can provide certain services to object terminal 140. Compared to ordinary object terminal 140, server 110 has higher requirements in terms of stability, security, and performance. Server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a portion of a high-performance computer (e.g., a virtual machine), a combination of portions of multiple high-performance computers (e.g., virtual machines), or a cloud server, etc. Server 110 contains various types of services, and the implementation of each service of server 110 is often associated with some intermediate databases or storage media. Server 110 is used to find relevant data stored locally based on the query requests submitted by the object. The distributed database includes compute nodes and storage nodes, and the compute nodes and storage nodes are independent servers. The server acting as a compute node is used to parse the query requests of the object terminal, determine the shard where the data is located, and distribute the query requests to the corresponding data nodes. The server acting as a data node is used to store and process data.

[0131] Gateway 120, also known as an internetwork connector or protocol converter, is a computer system or device that acts as a translator, enabling network interconnection at the transport layer. It bridges the gap between two systems using different communication protocols, data formats, languages, or even completely different architectures. Gateways can also provide filtering and security functions. Messages sent from target terminal 140 to server 110 are forwarded to the corresponding server via gateway 120. Messages sent from server 110 to target terminal 140 are also forwarded to the corresponding target terminal 140 via gateway 120.

[0132] The embodiments disclosed herein can be applied in various scenarios, such as Figures 2A-2B The data query scenarios shown are examples of this.

[0133] like Figure 2AThe diagram shows a simplified structural diagram of a distributed database. Specifically, the distributed database includes 3 compute nodes and 3 storage nodes. The 3 compute nodes are compute node 1, compute node 2, and compute node 3. The 3 storage nodes are storage node 1, storage node 2, and storage node 3. The data stored in storage nodes 1, 2, and 3 are different from each other. Specifically, storage node 1 contains two tables, Table 1 and Table 2. Table 1 records the number of tweets from platform A (100) and the total number of tweets from platform B (50). Table 2 records the platform IDs for platform A (xxx1) and platform B (xxx2). Storage node 2 also contains two tables, Table 1 and Table 2. Table 1 records the annual electricity consumption for zone Q (1520) and zone F (10000). Table 2 records the zone code for zone Q (xx) and zone code for zone F (xx2).

[0134] like Figure 2B As shown, an object submits a query task to compute node 1 of the distributed database: "Increase the number of tweets on platform A by 3, and sum the tweets on platforms A and B." Compute node 1 first decomposes the query task into an execution plan containing multiple subtasks, assigning the query identifier 001 to this task. The execution plan includes subtask 1 "Query the number of tweets on platform A and increment it by 3", subtask 2 "Query the number of tweets on platform B", and subtask 3 "Sum the tweets on platforms A and B". Subtasks 1 and 2 can be executed in parallel by different processing threads, while subtask 3 is executed after subtasks 1 and 2. Based on this, compute node 1 sends the query identifier 001 to storage node 1 to find the data snapshot corresponding to query identifier 001. Since no data snapshot is currently stored on this storage node, it copies the data locally to generate data snapshot 1, which corresponds to query identifier 001. Furthermore, the storage node provides Data Snapshot 1 to the processing thread executing Subtask 1 and also to the processing thread executing Subtask 2. Finally, after Subtask 1, Subtask 2, and Subtask 3 have all been completed, the compute node will know that the result of Subtask 1 is 103 tweets on Platform A, the result of Subtask 2 is 50 tweets on Platform B, and the result of Subtask 3 is a total of 153 tweets on both Platforms A and B. Therefore, it returns "the total number of tweets on Platforms A and B is 153" as the result of the query task to the user. Simultaneously, Table 1 on the storage node will be updated: the number of tweets on Platform A will change from "100" to "103". At this point, Data Snapshot 1 differs from the actual data state and is considered invalid, and will be released when the next query task is executed.

[0135] The embodiments of this disclosure are described in general below.

[0136] According to one embodiment of this disclosure, a distributed database query method is provided.

[0137] This distributed database query is often used for, for example Figures 2A-2B This disclosure describes a data query scenario. The embodiments of this disclosure employ a scheme where multiple subtasks, each decomposed from the same query statement in a distributed data query task, are assigned the same query identifier. This allows multiple subtasks to request the same stored data snapshot, improving the efficiency of distributed database queries and reducing network overhead.

[0138] The distributed database consists of compute nodes and storage nodes, with each storage node containing processing threads and a shared snapshot manager.

[0139] Compute nodes are used to access storage nodes based on user-submitted queries and to locate data corresponding to the queries in the storage nodes, so as to execute related tasks based on the retrieved data.

[0140] Storage nodes are used to store various types of data. In distributed databases, data is often divided into multiple fragments (such as shards or partitions) and stored as fragments on different storage nodes, with each storage node responsible for storing a portion of the data. Furthermore, storage nodes also provide efficient data retrieval services when compute nodes need to access data stored on them.

[0141] The processing thread is used to respond to data read and write requests from compute nodes or other clients, and to read data from the storage node or write data to the storage node according to the request content.

[0142] The shared snapshot manager maintains a consistent view of the data on the storage node, ensuring that the data snapshots used by query tasks during execution are consistent. The shared snapshot manager stores and maintains the mapping between query identifiers and data snapshots. The data snapshots maintained in the shared snapshot manager of the storage node are weakly consistent snapshots. Through the shared snapshot manager, the storage node can implement multi-version concurrency control, allowing query tasks to see the version of the data at the start of the query task, rather than the latest version. This avoids read / write conflicts and improves the concurrency performance of the storage node.

[0143] It should be noted that, in the embodiments disclosed herein, the distributed database may include multiple computing nodes and multiple storage nodes, and each storage node may include a shared snapshot manager and at least one processing thread.

[0144] like Figure 3As shown, a distributed database query method according to an embodiment of this disclosure is executed by a storage node, and the distributed database query method may include:

[0145] Step 310: Receive multiple subtasks from the computing node;

[0146] Step 320: Assign the same query identifier to multiple subtasks;

[0147] Step 330: For each subtask, based on the pre-stored mapping relationship, obtain the first stored data snapshot corresponding to the query identifier;

[0148] Step 340: Execute the subtask based on the first storage data snapshot to obtain the subtask processing result;

[0149] Step 350: Return the processing results of each subtask to the computing node so that the computing node can generate the processing results of the distributed data query task.

[0150] Steps 310-350 are described in detail below.

[0151] In step 310, multiple subtasks are received from the computing node.

[0152] Multiple subtasks are derived from the decomposition of query statements in the distributed data query task by the computing nodes.

[0153] A distributed data query task refers to a query operation submitted by a user in a distributed database environment that needs to be executed across multiple storage nodes. Because data is distributed and stored on different storage nodes in a distributed database, a distributed data query task may need to retrieve data from multiple storage nodes.

[0154] A query statement refers to an SQL statement submitted by a user to a distributed database. Query statements are used to request specific data or perform specific data operations.

[0155] In this specific implementation, a processing thread receives multiple subtasks from the compute node. Specifically, when a user submits a distributed data query task to the distributed database, the compute node decomposes the query statement into multiple subtasks. These subtasks may be executed in parallel, sequentially, or a combination of parallel and sequential execution. The compute node then submits these decomposed subtasks to the storage node in the order of execution. Based on this, the storage node's processing thread receives the multiple subtasks from the compute node.

[0156] It's important to note that when a user submits a distributed data query task to the distributed database, the compute nodes first receive the task and parse its query statement to understand its syntax and semantics. At this stage, the compute nodes extract key information such as the tables associated with the query, aggregation operations, and grouping conditions. Next, the compute nodes generate a query plan based on the distributed database's statistics (e.g., data distribution, table sizes). This plan determines how to execute the query, specifying the order of join operations, the execution method of aggregation operations, and so on. Furthermore, the compute nodes decompose the query into multiple subtasks based on the query plan. Decomposition methods include, but are not limited to, decomposition by data sharding and decomposition by operation.

[0157] When data is partitioned by data shards, and the data in the storage node is stored in shards, the compute node will generate a subtask for each shard.

[0158] Operation-based decomposition is primarily used for more complex queries. When the data referenced by a query is stored on the same storage node, the compute node may also decompose the query into multiple subtasks. For example, for a query containing filtering, aggregation, and sorting operations, the compute node might divide the query into subtask 1 corresponding to the filtering operation, subtask 2 corresponding to the aggregation operation, and subtask 3 corresponding to the sorting operation. Subtask 1, subtask 2, and subtask 3 are then executed sequentially.

[0159] For example, consider a distributed database representing the annual electricity consumption of City A. This annual electricity consumption data is stored as shards across different storage nodes, representing the annual electricity consumption of each district within City A. A user submits a query requesting statistics on the annual electricity consumption of different districts within City A. Therefore, the query in the distributed data query task is directed to the annual electricity consumption database of City A, retrieving the annual electricity consumption of different districts within City A. The computing nodes can decompose the query according to the storage nodes where the data is stored as shards by region. When City A is divided into districts A1, A2, and A3, the annual electricity consumption data for each district is stored on different storage nodes. Node 1 stores the annual electricity consumption data for district A1, node 2 stores the annual electricity consumption data for district A2, and node 3 stores the annual electricity consumption data for district A3. Based on this, the query is decomposed into three subtasks: subtask 1 calculates the annual electricity consumption data for district A1 on node 1, subtask 2 calculates the annual electricity consumption data for district A2 on node 2, and subtask 3 calculates the annual electricity consumption data for district A3 on node 3.

[0160] In step 320, the same query identifier is assigned to multiple subtasks.

[0161] The query identifier is used to uniquely identify the query statements from which multiple subtasks originate.

[0162] In the specific implementation of this embodiment, the same query identifier is assigned to multiple subtasks through a processing thread.

[0163] To save space, the specific implementation process of assigning the same query identifier to multiple subtasks through processing threads in this embodiment will be described in detail below. It will not be repeated here.

[0164] In step 330, for each subtask, a first stored data snapshot corresponding to the query identifier is obtained based on the pre-stored mapping relationship.

[0165] The mapping relationship is used to map multiple pre-stored query identifiers to multiple storage data snapshots. The mapping relationship can indicate the storage data snapshot corresponding to each of the multiple pre-stored query identifiers, and the first storage data snapshot is one of the multiple storage data snapshots.

[0166] The first storage data snapshot refers to a data snapshot maintained by the shared snapshot manager and corresponding to the query identifier.

[0167] It should be noted that the shared data snapshot manager stores and maintains the correspondence between query identifiers and data snapshots. In the shared snapshot manager, there is a one-to-one correspondence between query identifiers and data snapshots.

[0168] In this specific implementation, since the shared data snapshot manager stores and maintains the correspondence between query identifiers and data snapshots, firstly, the processing thread sends the query identifier corresponding to each subtask to the shared snapshot manager. Then, the shared snapshot manager returns the first stored data snapshot corresponding to the query identifier to the processing thread.

[0169] Specifically, when executing each subtask, the processing thread first retrieves the query identifier corresponding to the subtask, and then sends the query identifier corresponding to the subtask to the shared snapshot manager so that a data snapshot can be extracted from the shared snapshot manager based on the query identifier.

[0170] Furthermore, the shared snapshot manager maintains the mapping between query identifiers and data snapshots. Based on this, when a subtask's query identifier is sent to the shared snapshot manager via a processing thread, the shared snapshot manager searches for the corresponding data snapshot in its locally maintained mapping of query identifiers and data snapshots, and uses the found data snapshot as the first stored data snapshot. Finally, it returns the first stored data snapshot corresponding to the query identifier to the processing thread. This approach, by returning the first stored data snapshot through the shared snapshot manager, ensures that each subtask sees a consistent database state during execution. Even in a concurrent environment, modifications to data by other tasks will not affect the execution of the current subtask, thus guaranteeing task isolation and consistency. Simultaneously, the shared snapshot manager's rapid acquisition of data snapshots reduces the waiting time for subtasks in data access, improves query response speed, and ultimately enhances the query efficiency of the distributed database.

[0171] In step 340, a subtask is executed based on the first stored data snapshot to obtain the subtask processing result.

[0172] The subtask processing result is used to indicate the status of each data after the subtask is executed based on the first stored data snapshot.

[0173] In this specific implementation, a processing thread executes subtasks based on a first storage data snapshot to obtain the subtask processing results. Specifically, firstly, the processing thread extracts the data version of the data required to execute the subtask from the first storage data snapshot, and reads the data corresponding to that data version from the storage node. Next, based on the extracted data, the subtask is executed, and the data state of the data is updated. A new data version corresponding to the data is generated based on the updated data state and stored on the storage node. Finally, after the subtask is completed, the subtask processing result is obtained.

[0174] It's important to note that in multi-version storage systems, multiple versions of the same row of data that have been modified multiple times can exist in a distributed database. These versions can exist in multiple copies of the data itself, or they can be recovered through logs. For example, when data is modified from state 1 to state 2, and then from state 2 to state 3, with multiple versions coexisting, the data will be stored in three versions on the storage nodes: version 1.0 for state 1, version 2.0 for state 2, and version 3.0 for state 3. However, each version has a different query identifier. When reading the data, the query identifier recorded in the data snapshot determines which version of the data to read. Furthermore, when multiple versions of the data are saved using logs, each version is recorded as...<a,b> In the form of <state1, state2> and <state2, state3>, 'a' indicates the value before modification and 'b' indicates the value after modification,' the log for this data is then <state1, state2> and <state2, state3>. Similarly, the query identifier corresponding to the data snapshot determines which version of the data to read, and then the required version of the data is found through the log. If data before modification is needed, it can be restored from a dedicated buffer for use.

[0175] In step 350, the processing results of each subtask are returned to the computing node so that the computing node can generate the processing results of the distributed data query task.

[0176] The processing result is used to indicate the status of each data in the storage node after the query statement of the distributed data query task has been executed.

[0177] In this specific implementation, a processing thread returns the subtask processing results of each subtask to the computing node, enabling the computing node to generate the processing result of the distributed data query task. Specifically, after all subtasks corresponding to a query statement in the distributed data query task have been executed and their subtask processing results have been obtained, the processing thread first summarizes the subtask processing results of each subtask corresponding to the query statement; then, the processing thread returns the summarized subtask processing results of each subtask to the computing node. Based on this, the computing node can receive the subtask processing results of each subtask corresponding to the query statement and integrate them to obtain the processing result of the distributed data query task.

[0178] Through steps 310-350 above, in this embodiment of the disclosure, multiple subtasks derived from a single query statement in a distributed data query task are assigned the same query identifier. This query identifier corresponds to a first stored data snapshot. This first stored data snapshot is stored in a shared snapshot manager before querying the multiple subtasks, corresponding to the query identifier, thereby forming a mapping relationship between multiple pre-stored query identifiers and multiple stored data snapshots. This mapping relationship is used to map between the multiple pre-stored query identifiers and the multiple stored data snapshots. The first stored data snapshot is one of the multiple stored data snapshots. Since the multiple subtasks are assigned the same query identifier, when obtaining the first stored data snapshot corresponding to the query identifier of each subtask based on the mapping relationship between the multiple pre-stored query identifiers and the multiple stored data snapshots, the same first stored data snapshot is requested. In this way, multiple subtasks are executed based on the first storage data snapshot, and all tasks are executed according to the same data version. Logically parallel subtasks can be executed in parallel without errors because the query identifiers of multiple subtasks for the same query statement are consistent, and the first storage data snapshots obtained based on the same query identifier are consistent. Therefore, for the same query statement, the first storage data snapshot can be obtained at the beginning of the subtask execution, so that each subtask can use the obtained first storage data snapshot during execution, thereby greatly improving the query efficiency of distributed databases and reducing the occupation of network resources.

[0179] The above is a general description of steps 310-360. Steps 310, 340, and 350 have been described in detail above. The following will describe in detail the specific implementation process of steps 320 and 330.

[0180] Steps 320-330 are described in detail below.

[0181] In steps 320-330, the same query identifier is assigned to multiple subtasks;

[0182] For each subtask, based on the pre-stored mapping relationship, obtain the first stored data snapshot corresponding to the query identifier.

[0183] The storage node contains a processing thread and a shared snapshot manager, which is used to store and maintain the mapping between multiple pre-stored query identifiers and multiple stored data snapshots.

[0184] Specifically, the same query identifier is assigned to multiple subtasks through the processing thread; for each subtask, the query identifier corresponding to the subtask is sent to the shared snapshot manager through the processing thread; and the first storage data snapshot corresponding to the query identifier is returned to the processing thread through the shared snapshot manager.

[0185] Since a distributed data query task can be a single-query task (containing only one query statement) or a multi-query task (containing multiple query statements), the specific data query method will differ depending on the number of query statements included in the distributed data query task. Therefore, this disclosure will describe the data query process for a single-query task and the data query process for a multi-query task separately.

[0186] First, a detailed description will be given regarding the data query scenario where the distributed data query task in this embodiment is a single-statement task.

[0187] Please refer to Figure 4 In one embodiment, if the distributed data query task is a single query statement task, the distributed database query method further includes, but is not limited to, the following steps 410-420 before assigning the same query identifier to multiple subtasks:

[0188] Step 410: Through the processing thread, for the first subtask among multiple subtasks, send the compute node identifier of the compute node, the session identifier of the single query statement task, and the query identifier of the first subtask to the shared snapshot manager.

[0189] Step 420: Generate a first storage data snapshot based on the compute node identifier and session identifier through the shared snapshot manager, and store the first storage data snapshot in the shared snapshot manager corresponding to the query identifier.

[0190] Steps 410-420 are described in detail below.

[0191] In step 410, the compute node identifier is used to identify the compute node and distinguish it from other compute nodes in the distributed database. The session identifier of a single query task is used to indicate the network session to which the query statement in the single query task belongs, wherein different network sessions have different session identifiers.

[0192] The query identifier for the first subtask is used to identify the query statement to which the first subtask belongs.

[0193] In this specific implementation, firstly, with authorization, the processing thread retrieves the compute node identifier of the compute node and the session identifier of the network session to which the single query task belongs from the server's background logs. Next, for the first subtask, an identifier is generated based on the compute node identifier and the session identifier, and this identifier is used as the query identifier for the first subtask. Further, the processing thread, for the first subtask among multiple subtasks, sends the compute node identifier of the compute node, the session identifier of the single query task, and the query identifier of the first subtask together to the shared snapshot manager.

[0194] In step 420, when generating the first storage data snapshot based on the compute node identifier and session identifier, firstly, it is determined whether the data on each storage node in the distributed database has been synchronized, so that a data snapshot can be generated if the data on each storage node in the distributed database has been synchronized. Next, after determining that the data on each storage node in the distributed database has been synchronized, the data on the storage node where the shared snapshot manager resides is copied using the shared snapshot manager to obtain a data copy. A data snapshot is then generated based on the copied data copy, and this generated data snapshot is designated as the first storage data snapshot.

[0195] Furthermore, when storing the first storage data snapshot and the query identifier in the shared snapshot manager, the shared snapshot manager can use the query identifier as the key and the first storage data snapshot corresponding to the query identifier as the value, storing the first storage data snapshot and the query identifier in the shared snapshot manager locally in the form of key-value pairs, so as to improve the query efficiency of data snapshots.

[0196] like Figure 5The diagram shows the mapping relationship between query identifiers and data snapshots maintained by the shared snapshot manager at different points in time. Specifically, at 10:05:30, the shared snapshot manager maintains the mapping relationship between query identifier query001 and data snapshot 1. At 11:10:10, the shared snapshot manager maintains the mapping relationship between query identifier query002 and data snapshot 2. At 12:08:54, the shared snapshot manager maintains the mapping relationship between query identifier query003 and data snapshot 3. At 12:30:05, the shared snapshot manager maintains the mapping relationship between query identifier query004 and data snapshot 4. At 13:15:27, the shared snapshot manager maintains the mapping relationship between query identifier query004 and data snapshot 4. The shared snapshot manager maintains only one mapping relationship between query identifiers and data snapshots at any given time, and the mapping relationship maintained at that time is the most recent one. For example, when the time is 13:15:27, the shared snapshot manager will only record the query identifier corresponding to data snapshot 5, while the previously generated data snapshots 1 to 4 will not be saved in the shared snapshot manager. That is, the shared snapshot manager will only record the latest data snapshot and the query identifier corresponding to that data snapshot.

[0197] In this embodiment, step 320 specifically includes the following steps:

[0198] By processing threads, the query identifier of the first subtask is assigned to subsequent subtasks after the first subtask.

[0199] Specifically, by processing threads, the query identifier of the first subtask is assigned to subsequent subtasks after the first subtask, so that multiple subtasks corresponding to the same query statement have the same query identifier.

[0200] The advantage of this embodiment is that by assigning the query identifier of the first subtask to subsequent subtasks through processing threads, multiple subtasks corresponding to the same query statement can share the same query identifier. This clarifies the relationship between subsequent subtasks and the first subtask, facilitating the association and tracing of multiple subtasks corresponding to the same query statement. Furthermore, since multiple subtasks corresponding to a query statement share the same query identifier, and the query identifier corresponds one-to-one with a data snapshot, each subtask executes based on the same first stored data snapshot, improving the efficiency of distributed database queries and reducing network overhead.

[0201] Please refer to Figure 6 In one embodiment, step 420 specifically includes, but is not limited to, the following steps 610-630:

[0202] Step 610: Generate the first storage data snapshot based on the compute node identifier and session identifier using the shared snapshot manager;

[0203] Step 620: Perform a digest operation on the query identifier to obtain the digest result of the query identifier;

[0204] Step 630: Store the first storage data snapshot and the summary result of the query identifier in the shared snapshot manager.

[0205] Steps 610-630 are described in detail below.

[0206] In step 610, the specific process of generating the first storage data snapshot based on the compute node identifier and session identifier has been described in detail above. To save space, it will not be repeated here.

[0207] In step 620, the summary result is used to indicate the result of the query identifier being represented as a fixed-length string.

[0208] In this specific implementation, firstly, a preset digest function is called. Then, the preset digest function is used to perform a digest operation on the query identifier, resulting in a fixed-length random string that uniquely represents the query identifier. Therefore, this random string is used as the digest result of the query identifier.

[0209] In step 630, the shared snapshot manager can use the summary result of the query identifier as the key and the first stored data snapshot corresponding to the query identifier as the value, and store the first stored data snapshot and the summary result of the query identifier in the local shared snapshot manager in the form of key-value pairs.

[0210] like Figure 7The diagram shows the mapping between query identifier summary results and data snapshots maintained by the shared snapshot manager at different time points. Specifically, at 10:05:30, the shared snapshot manager maintains the mapping between the summary result Qeda1 of the query task with query identifier query001 and data snapshot 1. At 11:10:10, the shared snapshot manager maintains the mapping between the summary result dqaq414 of the query task with query identifier query002 and data snapshot 2. At 12:08:54, the shared snapshot manager maintains the mapping between the summary result dmqe89 of the query task with query identifier query003 and data snapshot 3. At 12:30:05, the shared snapshot manager maintains the mapping between the summary result foqnfa87 of the query task with query identifier query004 and data snapshot 4. At time 13:15:27, the shared snapshot manager maintains the mapping relationship between the summary result cmqfiqw09 of the query task with query identifier query005 and data snapshot 5.

[0211] The advantage of this embodiment is that in a distributed database, data may be distributed across multiple storage nodes and may be accessed and modified simultaneously by multiple query tasks. By generating data snapshots, a consistent data view can be provided for each query task, ensuring that the data state seen by the query task during execution is stable and unaffected by other concurrent tasks. Simultaneously, performing a digest operation (such as a hash operation) on the query identifier and storing the digest result along with the corresponding data snapshot allows for rapid verification of the uniqueness and integrity of the query identifier. When access to a data snapshot is needed, the corresponding data snapshot can be quickly located through the digest result, thereby improving the query efficiency of data snapshots.

[0212] In one embodiment, the processing thread includes multiple processing threads corresponding to multiple subtasks, and the multiple processing threads are parallel.

[0213] In this embodiment, step 410 specifically includes the following steps:

[0214] The processing thread corresponding to the first subtask sends the compute node identifier, session identifier, and query identifier of the first subtask to the shared snapshot manager.

[0215] Meanwhile, step 320 specifically includes the following steps:

[0216] The query identifier of the first subtask is assigned to the subsequent subtasks by the processing thread corresponding to the subsequent subtasks after the first subtask.

[0217] Specifically, firstly, the processing thread corresponding to the first subtask obtains the query identifier of the first subtask, the compute node identifier of the compute node, and the session identifier of the network session to which the single query statement task belongs. This compute node identifier, session identifier, and query identifier of the first subtask are then sent to the shared snapshot manager. Next, the shared snapshot manager feeds back the mapping between the query identifier of the first subtask and the first stored data snapshot to the processing thread corresponding to the first subtask. Further, the processing thread corresponding to the first subtask forwards the query identifier of the first subtask to the processing threads corresponding to subsequent subtasks. Based on this, the query identifier of the first subtask can be assigned to subsequent subtasks through the processing threads corresponding to subsequent subtasks.

[0218] For example, a distributed data query task includes sorting the students in Class A and their exam scores. The distributed database stores only a table of student IDs corresponding to the student names and a table of exam scores corresponding to those student IDs. Both the student ID table and the exam score table are stored on storage node 1. Based on this, the query statement of the distributed data query task is divided into three subtasks: subtask 1 searches the student ID table of Class A students, subtask 2 searches the exam score table corresponding to the student IDs, and subtask 3 determines the correspondence between the student names and exam scores based on the two tables found. Subtask 1 can be assigned to processing thread 1, subtask 2 to processing thread 2, and subtask 3 to processing thread 3. Processing threads 1 and 2 can execute in parallel, while processing thread 3 executes only after processing threads 1 and 2 have completed their respective subtasks. Furthermore, when the query identifier for subtask 1 is Q0x78, the query identifiers for subtasks 2 and 3 are also set to Q0x78. In this way, the data snapshots extracted by subtask 1, subtask 2 and subtask 3 are all data snapshots corresponding to the query identifier Q0x78.

[0219] The advantage of this embodiment is that by assigning the query identifier of the first subtask to subsequent subtasks, all subtasks can use the same query identifier when retrieving data snapshots from the shared snapshot manager. This allows the execution of multiple subtasks to be based on the same data snapshot, improving data consistency during task execution. Furthermore, multiple subtasks can be executed by different processing threads, achieving parallel execution of tasks in the distributed database, which can improve task execution efficiency to some extent.

[0220] In this embodiment, the specific process of returning the first stored data snapshot corresponding to the query identifier to the processing thread through the shared snapshot manager may include, but is not limited to, the following steps:

[0221] If the shared snapshot manager stores a historical storage data snapshot corresponding to the query identifier, return the historical storage data snapshot to the processing thread and use the historical storage data snapshot as the first storage data snapshot corresponding to the query identifier.

[0222] If the shared snapshot manager stores historical storage data snapshots, but the historical storage data does not correspond to the query identifier, a first storage data snapshot corresponding to the query identifier is generated and stored in the shared snapshot manager to replace the historical storage data snapshot.

[0223] If the shared snapshot manager does not store historical storage data snapshots, a first storage data snapshot corresponding to the query identifier is generated and stored in the shared snapshot manager corresponding to the query identifier.

[0224] Among them, historical storage data snapshots are the latest data snapshots maintained by the shared snapshot manager between the current time and the present time.

[0225] Specifically, when the shared snapshot manager receives a query identifier from the processing thread, it first performs a local query to check if a data snapshot is stored locally. Next, if the shared snapshot manager finds a historical data snapshot stored locally, it compares the query identifier corresponding to the historical data snapshot with the received query identifier. Furthermore, if the query identifier corresponding to the historical data snapshot matches the received query identifier, it is assumed that the shared snapshot manager stores the historical data snapshot corresponding to the query identifier. Based on this, the shared snapshot manager returns the historical data snapshot to the processing thread, using it as the first data snapshot corresponding to the query identifier.

[0226] Furthermore, if the query identifier corresponding to the historical storage data snapshot does not match the received query identifier, it is assumed that the shared snapshot manager stores a historical storage data snapshot, but the historical storage data does not correspond to the query identifier, and the historical storage data snapshot is invalid and has no reference value. Based on this, the shared snapshot manager will release the historical storage data snapshot, generate a first storage data snapshot corresponding to the received query identifier based on the session identifier and compute node identifier received along with the query identifier, and store the generated first storage data snapshot in the shared snapshot manager corresponding to the query identifier.

[0227] Furthermore, if the shared snapshot manager does not store historical storage data snapshots, it is deemed necessary to generate and store a data snapshot. Based on this, the shared snapshot manager will generate a first storage data snapshot corresponding to the query identifier based on the session identifier and compute node identifier received along with the query identifier, and store the generated first storage data snapshot corresponding to the query identifier in the shared snapshot manager.

[0228] The advantage of this embodiment is that when retrieving data snapshots from the shared snapshot manager, the lookup is based on the correspondence between the query identifier and the data snapshots stored in the shared snapshot manager. This ensures that in a multi-threaded or multi-tasking environment, different tasks access the same data snapshots on the storage nodes, improving data consistency and reliability. Furthermore, considering various scenarios such as the existence of historical data snapshots in the shared snapshot manager and whether the query identifier corresponding to an existing historical data snapshot matches the received query identifier, the embodiment directly returns the historical data snapshot if one already exists in the shared snapshot manager, avoiding the waste of time and computing resources caused by repeatedly generating data snapshots. Reusing existing historical data snapshots also saves storage space on the storage nodes to some extent. Conversely, when the query identifier of a stored historical data snapshot does not match the received query identifier, or when no historical data snapshot exists, a new data snapshot is generated promptly and stored corresponding to the query identifier, improving the flexibility of the shared snapshot manager. Furthermore, storing data snapshots and query identifiers together in a shared snapshot manager enables unified management and maintenance of data snapshots. This centralized management approach also facilitates the detection, backup, and recovery of data snapshots.

[0229] Since processing threads in a distributed database can execute in parallel, a situation may arise where multiple processing threads simultaneously occupy a data snapshot in the shared snapshot manager. Therefore, to clearly identify the reference status of data snapshots, this disclosure provides a reference counting scheme for data snapshots. This scheme can clearly indicate the reference status of data snapshots in the shared snapshot manager, reducing the risk of task execution failure due to erroneous release of data snapshots.

[0230] In this embodiment, after the shared snapshot manager stores a historical storage data snapshot corresponding to the query identifier, returns the historical storage data snapshot to the processing thread, and uses the historical storage data snapshot as the first storage data snapshot corresponding to the query identifier, the distributed database query method further includes:

[0231] Increment the first count by 1.

[0232] The first count represents the number of subtasks using the first stored data snapshot corresponding to the query identifier.

[0233] In addition, after returning the subtask processing results of each subtask to the computing node, the distributed database query method also includes:

[0234] Decrease the first count by 1.

[0235] Specifically, a first count is set for the first storage data snapshot. This first count represents the number of subtasks currently using the first storage data snapshot corresponding to the query identifier. If the shared snapshot manager stores a historical storage data snapshot corresponding to the query identifier, it returns this historical snapshot to the processing thread. After using this historical snapshot as the first storage data snapshot corresponding to the query identifier, it indicates that the first storage data snapshot is being used by a new subtask. Therefore, the first count is incremented by 1 to indicate that the number of subtasks currently using the first storage data snapshot is increasing. After the processing thread returns the subtask processing results to the compute node, it indicates that the subtask has been completed and no longer needs to use the first storage data snapshot. Therefore, the first count can be decremented by 1 to indicate that the number of subtasks currently using the first storage data snapshot is decreasing.

[0236] The advantage of this embodiment is that when the shared snapshot manager returns the first stored data snapshot to the processing thread, the first count is incremented by 1 to indicate that the subtask is using the first stored data snapshot; and after the subtask processing results of each subtask are returned to the compute node, the first count is decremented by 1 to indicate that the subtask has ended its reference to the first stored data snapshot. By performing reference counting on the data snapshots in the shared snapshot manager, the reference status of the data snapshots in the shared snapshot manager can be clearly indicated, thereby reducing the risk of task execution failure due to erroneous release of data snapshots.

[0237] like Figure 8 The diagram illustrates the interaction of a data query when the distributed data query task is a single query statement task. Specifically, when a user sends a distributed data query task containing only one query statement, the query statement in the distributed data query task is first regarded as query task 1. Then, the user session thread of the computing node decomposes query task 1 into subtask 1 and subtask 2, where subtask 1 and subtask 2 need to be executed sequentially.

[0238] Next, processing thread 1 of the storage node executes subtask 1. Upon starting subtask 1, the processing thread sends the query identifier 1 of subtask 1, the compute node identifier of the compute node, and the session identifier of the network session used by the user to send the distributed data query task to the shared snapshot manager to request a data snapshot corresponding to subtask 1. At this time, the shared snapshot manager determines a unique data snapshot 1 based on the compute node identifier and the session identifier, and returns data snapshot 1 to processing thread 1. Specifically, when determining data snapshot 1, the shared snapshot manager checks whether a historical data snapshot exists based on the session identifier, query identifier, and compute node identifier. If no historical data snapshot exists, a new data snapshot is generated based on the compute node identifier and the session identifier, and this newly generated data snapshot is used as data snapshot 1. The reference count of data snapshot 1 is incremented by 1, and data snapshot 1 is returned to processing thread 1. If a historical data snapshot exists in the shared snapshot manager, it checks whether the query identifier corresponding to that historical data snapshot is different from the received query identifier. If the query identifier corresponding to the historical data snapshot is the same as the received query identifier, then the historical data snapshot is used as data snapshot 1, the reference count of data snapshot 1 is incremented by 1, and data snapshot 1 is returned to processing thread 1. If the query identifier corresponding to the historical data snapshot is different from the received query identifier, then the historical data snapshot is released, a new data snapshot is generated based on the compute node identifier and session identifier, the newly generated data snapshot is used as data snapshot 1, the reference count of data snapshot 1 is incremented by 1, and data snapshot 1 is returned to processing thread 1.

[0239] Based on this, processing thread 1 receives data snapshot 1 returned by the shared snapshot manager and returns the mapping between query identifier 1 and data snapshot 1 to the user session thread, so that query identifier 1 can be assigned to subtask 2. Simultaneously, processing thread 1 executes subtask 1 based on data snapshot 1 and returns the processing result to the compute node. At this point, it is determined that subtask 1 has ended, and processing thread 1 decrements the reference count of data snapshot 1 by 1 and notifies the shared snapshot manager that processing thread 2 has released data snapshot 1.

[0240] Further, after subtask 1 is completed, processing thread 2 in the storage node executes subtask 2. At this point, processing thread 2 first sends query identifier 1 to the shared snapshot manager, enabling the shared snapshot manager to find the corresponding data snapshot 1 locally based on query identifier 1, increment the reference count of data snapshot 1 by 1, and return data snapshot 1 to processing thread 2. Next, processing thread 2 executes subtask 2 based on data snapshot 1 and returns the processing result to the compute node. At this point, it is determined that subtask 2 has ended, and processing thread 2 decrements the reference count of data snapshot 1 by 1 and notifies the shared snapshot manager that processing thread 2 has released data snapshot 1.

[0241] Furthermore, the user session process in the compute node will summarize the processing results of subtask 1 and subtask 2, and return the summarized processing results to the user, thereby completing the entire query task 1.

[0242] Furthermore, when a user sends the next distributed data query task, the compute node treats the query statement of that distributed query task as query task 2, decomposes query task 2 into a subtask, and assigns the subtask to the processing thread 1 of the storage node for execution. At this time, processing thread 1 sends the query identifier 2 of the subtask of query task 2, the compute node identifier of the compute node, and the session identifier of the network session used by the user to send the distributed data query task to the shared snapshot manager, in order to request a data snapshot corresponding to the subtask of query task 2 from the shared snapshot manager.

[0243] Since the execution of query task 1 changes the state of each data item in the distributed database, data snapshot 1 is essentially invalid and its data is not up-to-date. The shared snapshot manager, based on the received query identifier 2, compute node identifier, and session identifier, will find that the locally stored historical data snapshot is data snapshot 1. However, the shared snapshot manager will discover that the query identifier 1 of data snapshot 1 is different from the query identifier 2 sent during the request. When the reference count of data snapshot 1 reaches 0, the shared snapshot manager will release data snapshot 1 and generate a new data snapshot based on the compute node identifier and session identifier. This newly generated data snapshot will be used as data snapshot 2, and its reference count will be incremented by 1. Data snapshot 2 will then be returned to processing thread 1, allowing processing thread 1 to execute the subtask of query task 2 based on data snapshot 2.

[0244] The following section provides a detailed description of the data query scenario where the distributed data query task is a multi-query statement task.

[0245] In one embodiment, if the distributed data query task contains multiple query statements and the isolation level of the distributed data query task is the first type of isolation level, the multiple subtasks are multiple first subtasks obtained by decomposing the first query statement among the multiple query statements.

[0246] The first type of isolation level refers to an isolation level that requires the data snapshots used throughout the entire distributed data query task to remain consistent. For example, the repeatable read isolation level and the serializable isolation level both belong to the first type of isolation level.

[0247] Repeatable read isolation level ensures that the results of reading the same data multiple times within the same distributed data query task are consistent, even if some subtasks of the distributed data query task have been changed.

[0248] The serialization isolation level is the highest isolation level. Under the serialization isolation level, distributed data query tasks are executed serially, thereby achieving complete isolation between each distributed data query task and eliminating the problems caused by parallel execution.

[0249] It should be noted that, under the first type of isolation level, the data snapshots used by each subtask during the execution of the distributed data query task remain consistent. Therefore, in practical applications, a data snapshot can be requested from the shared snapshot manager only once and reused in subsequent processes. To support the reuse of the same data snapshot by various subtasks corresponding to different query statements in multiple query statements, this embodiment provides a scheme to fix the query identifier used by multiple query statements in a distributed data query task to the query identifier used when the data snapshot was first requested, which improves the convenience of calling data snapshots.

[0250] The first query statement refers to the first query statement among multiple query statements in a distributed data query task. The first subtask refers to the subtask obtained by decomposing the first query statement.

[0251] It should be noted that the specific process of decomposing the first query statement into multiple first subtasks is similar to the specific process of decomposing the query statement in the distributed data query task by the computing nodes in step 310 above. To save space, it will not be elaborated here.

[0252] Please refer to Figure 9 In this embodiment, prior to step 310, the distributed database query method further includes, but is not limited to, the following steps 910-920:

[0253] Step 910: In response to receiving a distributed data query task start request from the compute node, send the compute node identifier of the compute node, the session identifier of the distributed data query task, and the first query identifier to the shared snapshot manager.

[0254] Step 920: Generate a first storage data snapshot based on the compute node identifier and session identifier through the shared snapshot manager, and store the first storage data snapshot in the shared snapshot manager corresponding to the first query identifier.

[0255] Steps 910-920 are described in detail below.

[0256] In step 910, the distributed data query task start request refers to the request initiated by the compute node to the storage node when starting to execute a distributed data query task with multiple query statements. The first query identifier is used to identify the first query statement to which the multiple first subtasks belong.

[0257] In this specific implementation, when a compute node begins executing a distributed data query task with multiple query statements, it sends a distributed data query task start request to the processing thread of the storage node. Based on this, the processing thread receives the distributed data query task start request from the compute node and sends the compute node identifier, the session identifier of the distributed data query task, and the first query identifier to the shared snapshot manager. The specific process is similar to step 410 described above. The difference is that the distributed data query task in step 410 is a single query statement task, while the distributed data query task in step 910 is a multi-query statement task. For the sake of brevity, further details are omitted.

[0258] In step 920, the process of generating a first storage data snapshot based on the compute node identifier and session identifier through the shared snapshot manager, and storing the first storage data snapshot and the first query identifier in the shared snapshot manager, is similar to step 420 above. To save space, it will not be described in detail here.

[0259] The advantage of this embodiment is that when the distributed data query task is a multi-query task, the compute node generates a distributed data query task start request at the beginning of the task. This start request prompts the storage node's processing thread to send the compute node's identifier, the distributed data query task's session identifier, and the first query identifier to the shared snapshot manager before executing the subtasks corresponding to each of the multiple query statements. This enables the shared snapshot manager to generate a first storage data snapshot and store the correspondence between the first storage data snapshot and the first query identifier. In this way, each query statement in the distributed data query task can use the first query identifier as its query identifier. This ensures that the query identifier used by multiple query statements in the distributed data query task is fixed to the query identifier used when the data snapshot was first requested. By maintaining consistency in the query identifier, data snapshot consistency is achieved for the distributed data query task, improving the convenience of calling data snapshots.

[0260] In this embodiment, step 320 specifically includes the following steps:

[0261] By processing threads, first query identifiers corresponding to the first query statements are assigned to multiple first subtasks.

[0262] Specifically, the process of assigning a first query identifier to multiple first subtasks corresponding to a first query statement is similar to the process of assigning a query identifier to subsequent subtasks after the first subtask, as described above. The difference is that the first query identifier in this embodiment corresponds to the first query statement in a multi-query statement task, while the query identifier in the above embodiment corresponds to the unique query statement in a single-query statement task. For the sake of brevity, this will not be elaborated further.

[0263] In this embodiment, after returning the subtask processing results of each subtask to the computing node, the distributed database query method further includes:

[0264] By processing threads, multiple second subtasks are obtained by decomposing the second query statement after the first query statement from the computing node.

[0265] By processing threads, first query identifiers are assigned to multiple second subtasks;

[0266] By processing threads, for each of the multiple second subtasks, the first query identifier is sent to the shared snapshot manager;

[0267] The shared snapshot manager returns the first stored data snapshot corresponding to the first query identifier to the processing thread;

[0268] By using a processing thread, the second subtask is executed based on the first storage data snapshot to obtain the processing result of the second subtask;

[0269] The processing thread returns the processing results of each second subtask to the computing node.

[0270] Here, the second query statement refers to all query statements in the distributed data query task other than the first query statement (the first query statement). The second subtask refers to the subtask obtained by the computing nodes by decomposing the second query statement. The processing result of the second subtask is used to indicate the status of each data after executing the second subtask based on the first stored data snapshot.

[0271] Specifically, the implementation process of decomposing and assigning query identifiers for the second query statement in this embodiment is similar to the specific processing process for the first query statement described above. The execution of each first subtask of the first query statement and each second subtask of the second query statement depends on the first stored data snapshot. For the sake of brevity, this will not be elaborated further.

[0272] The advantage of this embodiment is that, after determining the correspondence between the first query identifier of the first query statement and the first data snapshot, the first query identifier corresponding to the first query statement is assigned to multiple first sub-tasks, so that each query statement of the distributed data query task can use the first query identifier as the query identifier. This achieves the goal of fixing the query identifier used by multiple query statements in the distributed data query task to the query identifier used when the data snapshot was first requested. By maintaining the consistency of the query identifiers of each sub-task, the consistency of data snapshots of multiple query statements in the distributed data query task is achieved, which can improve the efficiency of distributed database query and reduce network overhead.

[0273] Furthermore, in this embodiment, after returning the subtask processing results of each subtask to the computing node, the distributed database query method further includes:

[0274] If the subtask processing result is the result of multiple subtasks of the last query statement in the distributed data query task, a release request is sent to the shared snapshot manager so that the shared snapshot manager can release the first stored data snapshot.

[0275] Among them, the release request refers to the request initiated by the processing thread to release the first storage data snapshot after the processing thread has finished executing the distributed data query task.

[0276] Releasing a snapshot refers to removing it from its current state of use, so that it is no longer occupied by the current distributed data query task. Releasing a snapshot typically means releasing the current task from its reference, but the snapshot itself may still exist in the shared snapshot manager. After releasing a snapshot, the storage node can reclaim the resources associated with that snapshot (such as memory, disk space, etc.), but the snapshot may still be retained until the storage node decides to delete it.

[0277] It should be noted that releasing a data snapshot is a prerequisite for deleting a data snapshot. Only after a data snapshot has been completely released (no task references the data snapshot) will the storage node consider deleting the data snapshot.

[0278] Specifically, the storage node, through its processing thread, sequentially returns the subtask processing results of multiple subtasks for each query statement to the compute node. When the storage node, through its processing thread, sends the subtask processing results of multiple subtasks for the last query statement in the distributed data query task to the compute node, it indicates that the distributed data query task has been completed and no longer needs to use the first storage data snapshot. Based on this, the processing thread sends a release request to the shared snapshot manager. The shared snapshot manager, based on the received release request, determines that the processing thread has released its hold on the stored data snapshot, and releases the stored first storage data snapshot when it determines that no task on the storage node is holding it.

[0279] The advantages of this embodiment are that, considering the significant storage resources consumed by the first storage data snapshot stored by the shared snapshot manager, when the subtask processing result is determined to be the result of multiple subtasks of the last query statement in the distributed data query task, it is assumed that the distributed data query task no longer needs to use the first storage data snapshot. Based on the release request, the no longer needed first storage data snapshot can be released promptly to reclaim storage resources, reducing resource exhaustion caused by long-term storage resource occupation. Simultaneously, the first storage data snapshot represents the data state at a certain moment; releasing the first storage data snapshot allows subsequent queries to be based on the latest data or a new data snapshot, improving the accuracy of data queries. Furthermore, since the first storage data snapshot may be associated with locks or version control mechanisms, releasing the first storage data snapshot can remove related constraints, reduce execution conflicts between tasks, and improve the concurrent processing efficiency of the distributed database.

[0280] like Figure 10 The diagram illustrates the interaction of a distributed data query task when it is a multi-query task and the task is at the first type of isolation level. Specifically, when a user sends a distributed data query task (the query task in the diagram) containing query statement 1 and query statement 2, the compute node executes query statement 1 and query statement 2 sequentially. Before executing query statement 1, the compute node notifies the storage node's processing thread 1 to start the distributed data query task. Processing thread 1 sends the compute node identifier, the session identifier of the network session used by the user to send the distributed data query task, and the query identifier of query statement 1 to the shared snapshot manager to request a data snapshot corresponding to the query task. The specific method by which the shared snapshot manager generates the data snapshot corresponding to the query task is as described above. Figure 8The specific method for generating data snapshot 1 is basically the same. At this time, when the shared snapshot manager generates data snapshot 1 corresponding to the query task, it will return the correspondence between the query identifier of the query task and data snapshot 1 to the compute node through processing thread 1, so that query statement 2 can also use the query identifier of query statement 1 as its own query identifier for snapshot extraction.

[0281] Next, when executing query statement 1, the distributed data query task is first decomposed into subtask 1 and subtask 2, which can be executed in parallel. Then, processing thread 1 of the storage node executes subtask 1. At the start of subtask 1 execution, the processing thread sends the query identifier 1 of subtask 1, the compute node identifier of the compute node, and the session identifier of the network session used by the user to send the distributed data query task to the shared snapshot manager to request a data snapshot corresponding to subtask 1. At this time, the shared snapshot manager returns data snapshot 1 to processing thread 1. Simultaneously, processing thread 2 of the storage node executes subtask 2. Processing thread 2 first sends the query identifier to the shared snapshot manager, enabling the shared snapshot manager to find the corresponding data snapshot 1 locally based on the query identifier and return data snapshot 1 to processing thread 2. At this time, processing thread 1 executes subtask 1 based on data snapshot 1, returns the processing result to the compute node, and decrements the reference count of data snapshot 1 by 1 upon determining that subtask 1 has ended.

[0282] Furthermore, processing thread 2 will also execute subtask 2 based on data snapshot 1 and return the processing result to the compute node. Upon completion of subtask 2, processing thread 2 will decrement the reference count of data snapshot 1 by 1.

[0283] Furthermore, the user session process in the compute node will summarize the processing results of subtask 1 and subtask 2, and return the summarized processing result of query statement 1 to the user, thereby completing the execution of query statement 1 of the query task.

[0284] Further, the compute node executes query statement 2, decomposes it into a subtask 3, and assigns this subtask to processing thread 1 on the storage node for execution. At this point, processing thread 1 sends the query identifier of subtask 3 (which is the same as the query identifiers of subtasks 1 and 2 of query statement 1), the compute node's identifier, and the session identifier of the network session used by the user to send this distributed data query task to the shared snapshot manager, requesting a data snapshot corresponding to subtask 3 from the shared snapshot manager. The shared snapshot manager then returns data snapshot 1 to processing thread 1. Processing thread 1 executes subtask 3 based on data snapshot 1, returns the processing result to the compute node, and decrements the reference count of data snapshot 1 by 1 upon determining that subtask 3 has ended. Finally, the compute node returns the processing result of query statement 2 to the user, thus completing the execution of query statement 2 in the query task.

[0285] Finally, based on the processing results returned by the compute node, the user submits the query task if the processing results of query statement 1 and query statement 2 meet the requirements, thus confirming the successful execution of the query task. If the processing results of query statement 1 and / or query statement 2 do not meet the requirements, the user rolls back the query task to invalidate it, ensuring that the data in the distributed database is not affected by the query task. Therefore, when the compute node receives a user's operation to submit or roll back a query task, it notifies the shared snapshot manager to release data snapshot 1 through the processing thread. At this time, the shared snapshot manager will directly release data snapshot 1 when it detects that the reference count of data snapshot 1 is 0. If the reference count of data snapshot 1 is not 0, it will temporarily store data snapshot 1 in the snapshot list until the reference count of data snapshot 1 is 0, at which point it will be released.

[0286] In another embodiment, if the distributed data query task contains multiple query statements and the isolation level of the distributed data query task is the second type of isolation level, the multiple subtasks are multiple first subtasks obtained by decomposing the first query statement among the multiple query statements.

[0287] The second type of isolation level refers to the isolation level in a distributed data query task where, when executing multiple query statements, only data committed before the start of that query statement is read. For example, the Read Committed isolation level is a type of second-order isolation level. In this case, when the distributed data query task includes multiple query statements, each query statement has a different query identifier, but the subtasks decomposed from the query statements share the same query identifier.

[0288] Based on this, this disclosure provides a method that, when the distributed data query task is at the second type of isolation level, requires each query statement of the distributed data query task to apply for its own corresponding data snapshot, so that each query statement can normally use its own corresponding query identifier as the query identifier for applying for the data snapshot, thereby enabling each subtask of each query statement to have its corresponding data snapshot as a reference.

[0289] like Figure 11A As shown, in this embodiment, prior to step 310, the distributed database query method further includes, but is not limited to, the following steps 1110a-1120a:

[0290] Step 1110a: In response to receiving a distributed data query task start request from the compute node, send the compute node identifier of the compute node, the session identifier of the distributed data query task, and the first query identifier to the shared snapshot manager;

[0291] Step 1120a: Generate a first storage data snapshot based on the compute node identifier and session identifier through the shared snapshot manager, and store the first storage data snapshot in the shared snapshot manager corresponding to the first query identifier.

[0292] Steps 1110a-1120a are described in detail below.

[0293] The specific process of steps 1110a-1120a is similar to that of steps 910-920 above. To save space, it will not be described again.

[0294] The advantage of this embodiment is that when the distributed data query task is a multi-query statement task, the compute node generates a distributed data query task start request at the beginning of the task. This request prompts the storage node's processing thread to send the compute node's identifier, the distributed data query task's session identifier, and the first query identifier of the first query statement to the shared snapshot manager before executing the sub-tasks corresponding to each of the multi-query statements. This enables the shared snapshot manager to generate a first storage data snapshot and store the correspondence between the first storage data snapshot and the first query identifier. In this way, each first sub-task of the first query statement in the distributed data query task can use the first query identifier as its query identifier, thus fixing the query identifier used by each first sub-task of the first query statement to the same identifier. This consistency in query identifiers ensures data snapshot consistency across sub-tasks of the same query statement within the distributed data query task, improving the convenience of data snapshot retrieval and data query efficiency.

[0295] In this embodiment, step 320 specifically includes the following steps:

[0296] By processing threads, first query identifiers corresponding to the first query statements are assigned to multiple first subtasks.

[0297] In this embodiment, after the processing results of each subtask are returned to the computing node via the processing thread, the distributed database query method further includes:

[0298] By processing threads, multiple second subtasks are obtained by decomposing the second query statement after the first query statement from the computing node.

[0299] By using processing threads, second query identifiers corresponding to second query statements are assigned to multiple second subtasks.

[0300] By processing threads, for each of the multiple second subtasks, a second query identifier is sent to the shared snapshot manager;

[0301] The shared snapshot manager returns the second stored data snapshot corresponding to the second query identifier to the processing thread;

[0302] By processing the thread, the second subtask is executed based on the second storage data snapshot to obtain the processing result of the second subtask;

[0303] The processing thread returns the processing results of each second subtask to the computing node.

[0304] Specifically, the implementation process of decomposing and assigning query identifiers for the second query statement in this embodiment is similar to the specific processing process for the first query statement described above. The difference lies in that the execution of each first subtask of the first query statement depends on the correspondence between the first query identifier corresponding to the first query statement and the first stored data snapshot; while the execution of each second subtask of the second query statement depends on the correspondence between the second query identifier corresponding to the second query statement and the second stored data snapshot. For the sake of brevity, this will not be elaborated further.

[0305] In the second type of isolation level, the execution method of each query statement in a distributed data query task is basically the same as that when the distributed data query task is a single query statement task. Both methods request a data snapshot corresponding to the current query statement based on the query identifier. However, the data snapshots used for executing different subtasks of different query statements may vary slightly.

[0306] The advantage of this embodiment is that it takes into account the different isolation levels of distributed data query tasks. When the isolation level of the distributed data query task is the second type, each query statement of the distributed data query task is required to request its corresponding data snapshot from the shared snapshot manager. This ensures that each query statement normally uses its corresponding query identifier (e.g., the first query identifier of the first query statement and the second query identifier of the second query statement) as the query identifier for requesting a data snapshot from the shared snapshot manager. As a result, the processing thread has a corresponding data snapshot as a reference when executing each subtask of each query statement, thereby improving the data consistency and data accuracy of the processing thread when executing multiple subtasks of each query statement.

[0307] It should be noted that, as Figure 11B As shown, in this embodiment, before receiving multiple second subtasks derived from the decomposition of the second query statement after the first query statement from the computing node via the processing thread, the distributed database query method further includes, but is not limited to, the following steps 1110b-1120b:

[0308] Step 1110b: Send the compute node identifier of the compute node, the session identifier of the distributed data query task, and the second query identifier corresponding to the second query statement to the shared snapshot manager.

[0309] Step 1120b: Generate a second storage data snapshot based on the compute node identifier and session identifier through the shared snapshot manager, and store the second storage data snapshot in the shared snapshot manager corresponding to the second query identifier.

[0310] Steps 1110b-1120b are described in detail below.

[0311] The specific process of steps 1110b-1120b is similar to that of steps 1110a-1120a described above. The difference lies in that steps 1110a-1120a involve sending relevant data to the shared snapshot manager for the first query statement, generating a first storage data snapshot corresponding to the first query identifier; while steps 1110a-1120a involve sending relevant data to the shared snapshot manager for the second query statement, generating a second storage data snapshot corresponding to the second query identifier. The data sent, the generated data snapshots, and the constructed storage structures differ between the two steps. For the sake of brevity, these details will not be elaborated further.

[0312] The advantage of this embodiment is that when a distributed data query task is set as a multi-query statement task, before executing multiple subtasks of each query statement, the processing thread sends the compute node identifier, session identifier, and query identifier corresponding to the query statement to the shared snapshot manager. The shared snapshot manager then generates and stores the data snapshot corresponding to the query identifier of the query statement. This ensures that each subtask of each query statement can obtain the data snapshot corresponding to the query identifier when providing the query identifier to the shared snapshot manager. This approach satisfies the query requirements of distributed data query tasks that are second-type isolation level multi-query statement tasks, and improves the flexibility and universality of distributed data queries to a certain extent.

[0313] Furthermore, in this embodiment, after returning the subtask processing results of each subtask to the computing node, the distributed database query method further includes:

[0314] Send a release request to the shared snapshot manager so that the shared snapshot manager releases the first or second storage data snapshot.

[0315] Specifically, the process of releasing the first or second storage data snapshot in this embodiment is similar to the process of releasing the first storage data snapshot under the first type of isolation level described above. To save space, it will not be repeated here.

[0316] The advantage of this embodiment is that, when the distributed data query task is at the second type of isolation level, after considering that the processing results of the subtasks of each query statement are returned to the computing node, the query statement no longer needs to reference the corresponding data snapshot. The processing thread sends a release request to the shared snapshot manager, which, upon receiving the release request, checks the stored data snapshot. If it is confirmed that the data snapshot is not occupied by any task, the stored data snapshot is released, which enables timely reclamation of storage resources and improves the resource utilization of the distributed database.

[0317] like Figure 12The diagram illustrates the interaction of a distributed data query task when it is a multi-query task and the distributed data query task is at the second type of isolation level. Specifically, when a user sends a distributed data query task (the query task in the diagram) containing query statement 1 and query statement 2, the compute node will execute query statement 1 and query statement 2 sequentially. Before executing query statement 1, the compute node notifies the storage node's processing thread 1 to start the distributed data query task. Processing thread 1 sends the compute node identifier, the session identifier of the network session used by the user to send the distributed data query task, and the query identifier of query statement 1 to the shared snapshot manager to request a data snapshot corresponding to the query task. The specific method by which the shared snapshot manager generates the data snapshot corresponding to the query task is as described above. Figure 8 The specific method for generating data snapshot 1 is basically the same. At this time, when the shared snapshot manager generates data snapshot 1 corresponding to query statement 1, it will return the correspondence between the query identifier of query statement 1 and data snapshot 1 to the computing node through processing thread 1, so that each subtask in query statement 1 can use the query identifier of query statement 1 as its own query identifier to extract snapshots.

[0318] Next, when executing query statement 1, the distributed data query task is first decomposed into subtask 1 and subtask 2, which can be executed in parallel. Then, processing thread 1 of the storage node executes subtask 1. At the start of subtask 1 execution, the processing thread sends the query identifier 1 of subtask 1, the compute node identifier of the compute node, and the session identifier of the network session used by the user to send the distributed data query task to the shared snapshot manager to request a data snapshot corresponding to subtask 1. At this time, the shared snapshot manager returns data snapshot 1 to processing thread 1. Simultaneously, processing thread 2 of the storage node executes subtask 2. Processing thread 2 first sends the query identifier to the shared snapshot manager, enabling the shared snapshot manager to find the corresponding data snapshot 1 locally based on the query identifier and return data snapshot 1 to processing thread 2. At this time, processing thread 1 executes subtask 1 based on data snapshot 1, returns the processing result to the compute node, and decrements the reference count of data snapshot 1 by 1 upon determining that subtask 1 has ended.

[0319] Furthermore, processing thread 2 will also execute subtask 2 based on data snapshot 1 and return the processing result to the compute node. Upon completion of subtask 2, processing thread 2 will decrement the reference count of data snapshot 1 by 1.

[0320] Furthermore, the user session process in the compute node will summarize the processing results of subtask 1 and subtask 2, and return the summarized processing result of query statement 1 to the user, thereby completing the execution of query statement 1 of the query task.

[0321] Furthermore, the compute node executes query statement 2, decomposes it into a subtask 3, and assigns this subtask to processing thread 1 on the storage node for execution. At this point, processing thread 1 sends the query identifier of subtask 3 (which is different from the query identifiers of subtasks 1 and 2 of query statement 1), the compute node's identifier, and the session identifier of the network session used by the user to send this distributed data query task to the shared snapshot manager, requesting a data snapshot corresponding to subtask 3 from the shared snapshot manager. Since the execution of query statement 1 changes the state of data in the distributed database, data snapshot 1 is essentially invalid at this point, and its data is not up-to-date. The shared snapshot manager, based on the query identifier 2, compute node identifier, and session identifier of the received query statement 2, locates the locally stored historical data snapshot as data snapshot 1. However, the shared snapshot manager finds that the query identifier of data snapshot 1 is different from the query identifier received when executing subtask 3. When the reference count of data snapshot 1 reaches 0, the shared snapshot manager releases data snapshot 1 and generates a new data snapshot based on the compute node identifier and session identifier. This newly generated data snapshot is designated as data snapshot 2, and its reference count is incremented by 1. Data snapshot 2 is then returned to processing thread 1. Based on this, processing thread 1 executes subtask 3 based on data snapshot 2 and returns the processing result to the compute node. Upon determining that subtask 3 has ended, processing thread 1 decrements the reference count of data snapshot 2 by 1. At this point, the compute node returns the processing result of query statement 2 to the user, thus completing the execution of query statement 2 in the query task.

[0322] Finally, based on the processing results returned by the compute node, the user submits the query task if the processing results of query statement 1 and query statement 2 meet the requirements, thus confirming the successful execution of the query task. If the processing results of query statement 1 and / or query statement 2 do not meet the requirements, the user rolls back the query task to invalidate it, ensuring that the data in the distributed database is not affected by the query task. Therefore, when the compute node receives a user's operation to submit or roll back a query task, it notifies the shared snapshot manager to release data snapshot 1 and data snapshot 2 through the processing thread. At this time, the shared snapshot manager will directly release data snapshot 1 and data snapshot 2 when it detects that the reference count of data snapshot 1 / data snapshot 2 is 0. If the reference count of data snapshot 1 / data snapshot 2 is not 0, it will temporarily store data snapshot 1 / data snapshot 2 in the snapshot list until the reference count of data snapshot 1 / data snapshot 2 is 0, at which point it will be released.

[0323] The following describes a process for cleaning up a storage data snapshot of an abnormal connection according to an embodiment of this disclosure.

[0324] Because network interruptions or abnormal node shutdowns often occur during distributed data query tasks, these anomalies can prevent tasks from executing normally and lead to the prolonged erroneous occupation of certain data snapshots. Prolonged occupation or storage of data snapshots results in wasted storage node resources. Therefore, this disclosure provides a scheme for cleaning up data snapshots under abnormal conditions. This scheme can clean up data snapshots under abnormal connections, reduce resource waste in distributed databases, and improve the effective utilization of storage resources.

[0325] In one embodiment, the storage node includes a processing thread and a shared snapshot manager; after returning the subtask processing results of each subtask to the compute node, the distributed database query method further includes:

[0326] In response to the detection of a user connection interruption, the connection between the storage node and the compute node is disconnected via a processing thread;

[0327] The processing thread decrements the first count corresponding to the first storage data snapshot by 1 and sends a snapshot release command to the shared snapshot manager. The shared snapshot manager then moves the first storage data snapshot to the snapshot list according to the snapshot release command and releases the first storage data snapshot when the first count corresponding to the first storage data snapshot is 0.

[0328] The snapshot list refers to the storage space in the distributed database used to store data snapshots that need to be cleaned up.

[0329] Specifically, when a user's connection to the distributed database is interrupted, the compute nodes in the distributed database will detect the interruption. Based on this, the compute node, in response to the detected interruption, sends a user connection interruption event to the processing thread of the storage node to indicate that a user connection interruption has occurred. At this time, the storage node, through its processing thread, disconnects the connection between itself and the compute node. Then, the storage node, through its processing thread, decrements the first count corresponding to the first storage data snapshot by 1 to indicate that the distributed data query task sent by the user has ended its use of the first storage data snapshot. Simultaneously, it also sends a snapshot release command to the shared snapshot manager through its processing thread to notify that the first storage data snapshot needs to be released due to the user connection interruption. Further, when the shared snapshot manager receives the release command, considering that other tasks from other users may be using the first storage data snapshot in the storage node, it will not directly release the first data snapshot, but instead moves it to the snapshot list. After the first storage data snapshot is moved to the snapshot list, if the first count corresponding to the first storage data snapshot is not 0, it will wait for the next connection interruption notification until the first count is 0. Only then will it be considered that no task is occupying the first storage data snapshot and the first storage data snapshot can be released.

[0330] like Figure 13 The diagram illustrates the process of clearing a data snapshot when a user connection is interrupted. Specifically, after a user sends a distributed data query task, the query statement within that task is considered query task 1. The user session thread on the compute node then breaks down query task 1 into subtask 1 and subtask 2. The corresponding data snapshot for query task 1 is data snapshot 1. When subtask 1 is completed based on data snapshot 1, processing thread 1 returns the processing result to the compute node. However, during the execution of subtask 2, when processing thread 2 requests data snapshot 1 based on query identifier 1, the user session is interrupted, and the user's terminal is abnormally disconnected from the compute node of the distributed database. At this point, the compute node notifies processing threads 1 and 2 on the storage node of the connection interruption. Next, processing threads 1 and 2 notify the shared snapshot manager of the connection interruption between the storage node and the compute node, and prompt the shared snapshot manager to release data snapshot 1. Because the connection between the compute node and the storage node is interrupted, the distributed data query task submitted by the user no longer needs to reference data snapshot 1. Based on this, the shared snapshot manager retrieves the reference count of data snapshot 1. If the reference count is 0, data snapshot 1 is released directly; if the reference count is not 0, data snapshot 1 is stored in the snapshot list, meaning that the user's distributed data query task will not reuse data snapshot 1, and data snapshot 1 will be released only after the reference count is 0.

[0331] The advantage of this embodiment is that it considers abnormal situations caused by user connection interruptions. When a user connection is interrupted, the processing thread notifies the shared snapshot manager of the connection interruption event and prompts the shared snapshot manager to release the data snapshot. Furthermore, the shared snapshot manager does not directly release the data snapshot upon receiving the release instruction; instead, it cleans up the data snapshot based on its reference status. This approach not only cleans up the data snapshot promptly during abnormal connections but also reduces the risk of certain tasks failing due to direct data snapshot cleanup. In addition, this approach reduces resource waste in the distributed database and improves the effective utilization of storage resources.

[0332] In another embodiment, the storage node includes a processing thread and a shared snapshot manager; after returning the subtask processing results of each subtask to the compute node, the distributed database query method further includes:

[0333] In response to the detection of abnormal restart or abnormal shutdown of the compute node, the connection between the storage node and the compute node is disconnected through the processing thread;

[0334] The processing thread decrements the first count corresponding to the first storage data snapshot by 1 and sends a snapshot release command to the shared snapshot manager. The shared snapshot manager then moves the first storage data snapshot to the snapshot list according to the snapshot release command and releases the first storage data snapshot when the first count corresponding to the first storage data snapshot is 0.

[0335] Specifically, when a compute node abnormally restarts or shuts down, the storage node automatically interrupts the connection from that compute node and executes a connection interruption procedure. In this embodiment, the specific process of executing the connection interruption procedure to clean up the data snapshot when an abnormal restart or shutdown of a compute node is detected is similar to the specific process of cleaning up the data snapshot when a user connection interruption is detected in the previous embodiment. For the sake of brevity, it will not be described in detail again.

[0336] like Figure 14The diagram illustrates the process of clearing data snapshots when a compute node abnormally restarts or shuts down. Specifically, after a user sends a distributed data query task, the query statement within the task is considered query task 1. The user session thread on the compute node decomposes query task 1 into subtask 1 and subtask 2. The data snapshot corresponding to query task 1 is data snapshot 1. After subtask 1 is executed based on data snapshot 1, processing thread 1 returns the processing result to the compute node. While executing subtask 2, processing thread 2 requests data snapshot 1 based on query identifier 1. At this point, the compute node either shuts down or restarts. This is considered an abnormality in the compute node, and the connection between the compute node and the storage node is interrupted. Next, processing threads 1 and 2 notify the shared snapshot manager that the connection between the storage node and the compute node is interrupted, and prompt the shared snapshot manager to release data snapshot 1. Because the connection between the compute node and the storage node is interrupted, the distributed data query task submitted by the user no longer needs to reference data snapshot 1. Based on this, the shared snapshot manager retrieves the reference count of data snapshot 1. If the reference count is 0, data snapshot 1 is released directly; if the reference count is not 0, data snapshot 1 is stored in the snapshot list, meaning that the user's distributed data query task will not reuse data snapshot 1, and data snapshot 1 will be released only after the reference count is 0.

[0337] The advantage of this embodiment is that it considers abnormal situations caused by abnormal restarts or shutdowns of compute nodes. When a compute node restarts or shuts down abnormally, the processing thread notifies the shared snapshot manager of the connection interruption event and prompts the shared snapshot manager to release the data snapshot. Furthermore, the shared snapshot manager does not directly release the data snapshot upon receiving the release instruction; instead, it cleans up the data snapshot based on its reference status. This method promptly cleans up data snapshots during abnormal connections, improving the effective utilization of storage resources. In addition, this method reduces the risk of certain tasks failing due to directly cleaning up data snapshots.

[0338] The following is a detailed description of another embodiment of the distributed database query method of this disclosure.

[0339] like Figure 15 As shown, a distributed database query method according to an embodiment of this disclosure is executed by a computing node, and the distributed database query method may include:

[0340] Step 1510: In response to the target user's query request, obtain the distributed data query task;

[0341] Step 1520: For the query statement in the distributed data query task, decompose the query statement into multiple sub-tasks;

[0342] Step 1530: Send multiple subtasks corresponding to the query statement to the processing thread of the storage node so that the processing thread can assign the same query identifier to the multiple subtasks corresponding to the query statement, obtain the first storage data snapshot corresponding to the query identifier from the shared snapshot manager based on the query identifier, and execute multiple subtasks based on the first storage data snapshot to obtain the subtask processing results of each subtask.

[0343] Step 1540: Receive the subtask processing results of each subtask of the query statement returned by the processing thread.

[0344] Step 1550: Based on the subtask processing results of each subtask of the query statement, generate the processing results of the distributed data query task.

[0345] Steps 1510-1550 are described in detail below.

[0346] In step 1510, the target user refers to the user who sends a distributed data query task to the distributed database through a network session. A query request refers to a request initiated by the target user when performing certain operations that require querying and retrieving relevant data from the distributed database.

[0347] The distributed data query task includes query statements. The query statements are SQL statements.

[0348] In this specific implementation, when a target user wants to query and retrieve some data from a distributed database, the target user establishes a connection with a computing node of the distributed database and initiates a query request, specifying the data to be queried and the specific data processing strategy. Based on this, the computing node of the distributed database receives the query request and parses a distributed data query task according to the query request. This distributed data query task consists of SQL statements.

[0349] In step 1520, the specific process by which the compute node decomposes the query statement into multiple subtasks has already been described in detail in step 310 above. To save space, it will not be repeated here.

[0350] In step 1530, after the compute node and storage node establish a connection, multiple subtasks corresponding to the query statement are sent to the processing thread of the storage node. The processing thread assigns the same query identifier to each subtask, retrieves a first storage data snapshot corresponding to the query identifier from the shared snapshot manager, and executes multiple subtasks based on the first storage data snapshot. The specific process of obtaining the subtask processing results for each subtask is similar to steps 320-350 above. For brevity, it will not be elaborated further here.

[0351] In steps 1540-1550, after the storage node has executed all subtasks corresponding to the query statement, the compute node can receive the subtask processing results of each subtask corresponding to the query statement, and integrate the subtask processing results of each subtask to obtain the processing result of the distributed data query task. At this time, the compute node will feed back the processing result of the distributed data query task to the target user.

[0352] The advantage of this embodiment is that, in a distributed database, a compute node decomposes a single query statement from a received distributed data query task into multiple subtasks and assigns the same query identifier to each subtask. This query identifier corresponds to a first storage data snapshot. This first storage data snapshot is stored in a shared snapshot manager, corresponding to the query identifier, before the multiple subtasks are queried. Because multiple subtasks are assigned the same query identifier, when the processing threads on the storage node request the first storage data snapshot corresponding to the query identifier from the shared snapshot manager for each subtask, they request the same first storage data snapshot. In this way, logically parallel subtasks can be executed in parallel without errors because they are based on the same first storage data snapshot, thereby greatly improving the efficiency of distributed database queries and reducing network overhead.

[0353] In one embodiment, step 1530 specifically includes, but is not limited to, the following steps:

[0354] Identify the isolation level for distributed data query tasks;

[0355] The query statement's corresponding subtasks and isolation levels are sent to the storage node's processing thread.

[0356] Specifically, firstly, the session isolation level of the network session containing the distributed data query task can be viewed using SQL commands in the distributed database, and this session isolation level is then determined as the isolation level for the distributed data query task. The SQL command code can be represented as `SELECT @@session.tx_isolation`. Next, the compute node, through its connection with the storage node, sends the multiple subtasks corresponding to the query statement and the isolation level to the processing thread on the storage node. This allows the processing thread to determine the data snapshot for each subtask corresponding to the query statement based on the isolation level of the distributed data query task.

[0357] The advantage of this embodiment is that it takes into account the differences in isolation levels of distributed data query tasks from different network sessions. When extracting data snapshots for distributed data query tasks, taking the type of isolation level of the distributed data query task as a reference can improve the accuracy of snapshot extraction of subtasks of each query statement in the distributed data query task, thereby improving the accuracy of task execution.

[0358] The apparatus and device according to embodiments of this disclosure will now be described.

[0359] It is understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated in this embodiment, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0360] It should be noted that in various specific embodiments of this application, when processing is required based on data related to the characteristics of the target object, such as target object attribute information or a set of attribute information, the permission or consent of the target object will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require obtaining target object attribute information, separate permission or consent from the target object will be obtained through pop-ups or redirection to a confirmation page. Only after obtaining the target object's separate permission or consent will the necessary target object-related data for the normal operation of the embodiments of this application be obtained.

[0361] Figure 16 This is a schematic diagram of the structure of a distributed database query device 1600 provided in an embodiment of this disclosure. The distributed database includes computing nodes and storage nodes. The distributed database query device 1600 is executed by the storage nodes and includes:

[0362] The receiving unit 1610 is used to receive multiple subtasks from the computing node. The multiple subtasks are obtained by the computing node from the decomposition of the query statement in the distributed data query task.

[0363] Allocation unit 1620 is used to allocate the same query identifier to multiple subtasks;

[0364] The acquisition unit 1630 is used to acquire, for each subtask, a first storage data snapshot corresponding to the query identifier based on a pre-stored mapping relationship. The mapping relationship is used to map between multiple pre-stored query identifiers and multiple storage data snapshots. The first storage data snapshot is one of the multiple storage data snapshots.

[0365] Execution unit 1640 is used to execute subtasks based on the first storage data snapshot and obtain the subtask processing results;

[0366] Return unit 1650 is used to return the subtask processing results of each subtask to the computing node so that the computing node can generate the processing results of the distributed data query task.

[0367] Optionally, the storage node includes a processing thread and a shared snapshot manager, which is used to store and maintain mapping relationships;

[0368] The acquisition unit 1630 is used for:

[0369] By processing threads, for each subtask, the query identifier corresponding to the subtask is sent to the shared snapshot manager;

[0370] The shared snapshot manager returns the first stored data snapshot corresponding to the query identifier to the processing thread.

[0371] Optionally, the distributed database query apparatus 1600 further includes a first storage unit (not shown), which includes:

[0372] The sending module (not shown) is used to send the compute node identifier of the compute node, the session identifier of the single query statement task, and the query identifier of the first subtask to the shared snapshot manager through the processing thread, for the first subtask among multiple subtasks, if the distributed data query task is a single query statement task.

[0373] A first generation module (not shown) is used to generate a first storage data snapshot based on a compute node identifier and a session identifier through a shared snapshot manager, and store the first storage data snapshot in the shared snapshot manager in correspondence with a query identifier;

[0374] The allocation unit 1620 is used for:

[0375] By processing threads, the query identifier of the first subtask is assigned to subsequent subtasks after the first subtask.

[0376] Optionally, the processing thread includes multiple processing threads corresponding to multiple subtasks, and the multiple processing threads are parallel;

[0377] The first generation module (not shown) is used for:

[0378] Through the processing thread corresponding to the first subtask, the compute node identifier, session identifier, and query identifier of the first subtask are sent to the shared snapshot manager for the first subtask;

[0379] The allocation unit 1620 is used for:

[0380] The query identifier of the first subtask is assigned to the subsequent subtasks by the processing thread corresponding to the subsequent subtasks after the first subtask.

[0381] Optionally, the transmitting module (not shown) is used for:

[0382] The first storage data snapshot is generated based on the compute node identifier and session identifier using the shared snapshot manager;

[0383] Perform a digest operation on the query identifier to obtain a digest result of the query identifier;

[0384] The first storage data snapshot is stored in the shared snapshot manager, corresponding to the summary results of the query identifier.

[0385] Optionally, through a shared snapshot manager, a first stored data snapshot corresponding to the query identifier of each subtask is returned to the thread, including:

[0386] If the shared snapshot manager stores a historical storage data snapshot corresponding to the query identifier, return the historical storage data snapshot to the processing thread and use the historical storage data snapshot as the first storage data snapshot corresponding to the query identifier.

[0387] If the shared snapshot manager stores historical storage data snapshots, but the historical storage data does not correspond to the query identifier, a first storage data snapshot corresponding to the query identifier is generated and stored in the shared snapshot manager to replace the historical storage data snapshot.

[0388] If the shared snapshot manager does not store historical storage data snapshots, a first storage data snapshot corresponding to the query identifier is generated and stored in the shared snapshot manager corresponding to the query identifier.

[0389] Optionally, the distributed database query apparatus 1600 further includes a counting unit (not shown), which is used for:

[0390] If the shared snapshot manager stores a historical storage data snapshot corresponding to the query identifier, the historical storage data snapshot is returned to the processing thread. After the historical storage data snapshot is used as the first storage data snapshot corresponding to the query identifier, the first count is incremented by 1. The first count represents the number of subtasks that are using the first storage data snapshot corresponding to the query identifier.

[0391] After returning the subtask processing results of each subtask to the computing node, decrement the first count by 1.

[0392] Optionally, if the distributed data query task contains multiple query statements and the isolation level of the distributed data query task is the first type of isolation level, the multiple subtasks are multiple first subtasks obtained by decomposing the first query statement among the multiple query statements.

[0393] The allocation unit 1620 is used for:

[0394] By processing threads, first query identifiers corresponding to the first query statements are assigned to multiple first subtasks;

[0395] The distributed database query apparatus 1600 further includes a first processing unit (not shown), which is used for:

[0396] By processing threads, multiple second subtasks are obtained by decomposing the second query statement after the first query statement from the computing node.

[0397] By processing threads, first query identifiers are assigned to multiple second subtasks;

[0398] By processing threads, for each of the multiple second subtasks, the first query identifier is sent to the shared snapshot manager;

[0399] The shared snapshot manager returns the first stored data snapshot corresponding to the first query identifier to the processing thread;

[0400] By using a processing thread, the second subtask is executed based on the first storage data snapshot to obtain the processing result of the second subtask;

[0401] The processing thread returns the processing results of each second subtask to the computing node.

[0402] Optionally, the distributed database query apparatus 1600 further includes a second storage unit (not shown), which is used for:

[0403] In response to receiving a distributed data query task start request from the compute node, the compute node identifier, the session identifier of the distributed data query task, and the first query identifier are sent to the shared snapshot manager.

[0404] The first storage data snapshot is generated based on the compute node identifier and session identifier through the shared snapshot manager, and the first storage data snapshot is stored in the shared snapshot manager corresponding to the first query identifier.

[0405] Optionally, the distributed database query apparatus 1600 further includes a first release unit (not shown), which is used for:

[0406] If the subtask processing result is the result of multiple subtasks of the last query statement in the distributed data query task, a release request is sent to the shared snapshot manager so that the shared snapshot manager can release the first stored data snapshot.

[0407] Optionally, if the distributed data query task contains multiple query statements and the isolation level of the distributed data query task is the second type of isolation level, the multiple subtasks are multiple first subtasks obtained by decomposing the first query statement among the multiple query statements.

[0408] The allocation unit 1620 is used for:

[0409] By processing threads, first query identifiers corresponding to the first query statements are assigned to multiple first subtasks;

[0410] The distributed database query device also includes a second processing unit, which is used for:

[0411] By processing threads, multiple second subtasks are obtained by decomposing the second query statement after the first query statement from the computing node.

[0412] By using processing threads, second query identifiers corresponding to second query statements are assigned to multiple second subtasks.

[0413] By processing threads, for each of the multiple second subtasks, a second query identifier is sent to the shared snapshot manager;

[0414] The shared snapshot manager returns the second stored data snapshot corresponding to the second query identifier to the processing thread;

[0415] By processing the thread, the second subtask is executed based on the second storage data snapshot to obtain the processing result of the second subtask;

[0416] The processing thread returns the processing results of each second subtask to the computing node.

[0417] Optionally, the distributed database query apparatus 1600 further includes a third storage unit (not shown), which is used for:

[0418] In response to receiving a distributed data query task start request from the compute node, the compute node identifier, the session identifier of the distributed data query task, and the first query identifier are sent to the shared snapshot manager.

[0419] The first storage data snapshot is generated based on the compute node identifier and session identifier through the shared snapshot manager, and the first storage data snapshot is stored in the shared snapshot manager in correspondence with the first query identifier;

[0420] The distributed database query apparatus 1600 further includes a third processing unit (not shown), which is used for:

[0421] Send the compute node identifier of the compute node, the session identifier of the distributed data query task, and the second query identifier corresponding to the second query statement to the shared snapshot manager;

[0422] A second storage data snapshot is generated based on the compute node identifier and session identifier through the shared snapshot manager, and the second storage data snapshot is stored in the shared snapshot manager in correspondence with the second query identifier.

[0423] Optionally, the distributed database query apparatus further includes a second release unit, which is used for:

[0424] Send a release request to the shared snapshot manager so that the shared snapshot manager releases the first or second storage data snapshot.

[0425] Optionally, the storage node includes a processing thread and a shared snapshot manager; the distributed database query apparatus 1600 also includes a first exception handling unit (not shown), which is used for:

[0426] In response to the detection of a user connection interruption, the connection between the storage node and the compute node is disconnected via a processing thread;

[0427] The processing thread decrements the first count corresponding to the first storage data snapshot by 1 and sends a snapshot release command to the shared snapshot manager. The shared snapshot manager then moves the first storage data snapshot to the snapshot list according to the snapshot release command and releases the first storage data snapshot when the first count corresponding to the first storage data snapshot is 0.

[0428] Optionally, the storage node includes a processing thread and a shared snapshot manager; the distributed database query apparatus 1600 also includes a second exception handling unit (not shown), which is used for:

[0429] In response to the detection of abnormal restart or abnormal shutdown of the compute node, the connection between the storage node and the compute node is disconnected through the processing thread;

[0430] The processing thread decrements the first count corresponding to the first storage data snapshot by 1 and sends a snapshot release command to the shared snapshot manager. The shared snapshot manager then moves the first storage data snapshot to the snapshot list according to the snapshot release command and releases the first storage data snapshot when the first count corresponding to the first storage data snapshot is 0.

[0431] Figure 17 This is a schematic diagram of the structure of a distributed database query device 1700 provided in an embodiment of this disclosure. The distributed database includes compute nodes and storage nodes. The storage nodes include processing threads and a shared snapshot manager. The distributed database query device 1700 is executed by the compute nodes and includes:

[0432] The acquisition module 1710 is used to acquire distributed data query tasks in response to the query request of the target user. The distributed data query tasks include query statements.

[0433] The decomposition module 1720 is used to decompose the query statement in the distributed data query task into multiple sub-tasks.

[0434] The processing module 1730 is used to send multiple subtasks corresponding to the query statement to the processing thread of the storage node, so that the processing thread assigns the same query identifier to the multiple subtasks corresponding to the query statement, obtains the first storage data snapshot corresponding to the query identifier from the shared snapshot manager based on the query identifier, and executes multiple subtasks based on the first storage data snapshot to obtain the subtask processing results of each subtask.

[0435] The receiving module 1740 is used to receive the subtask processing results of each subtask of the query statement returned by the processing thread;

[0436] The second generation module 1750 is used to generate the processing results of the distributed data query task based on the subtask processing results of each subtask of the query statement.

[0437] Optionally, the query statement may consist of multiple statements; the processing module 1730 is used for:

[0438] Identify the isolation level for distributed data query tasks;

[0439] The query statement's corresponding subtasks and isolation levels are sent to the storage node's processing thread.

[0440] Reference Figure 18 , Figure 18To illustrate the structural block diagram of a terminal for implementing the distributed database query method of this embodiment, the terminal includes: a radio frequency (RF) circuit 1810, a memory 1815, an input unit 1830, a display unit 1840, a sensor 1850, an audio circuit 1860, a wireless fidelity (WiFi) module 1870, a processor 1880, and a power supply 1890, among other components. Those skilled in the art will understand that... Figure 18 The terminal structure shown does not constitute a limitation on mobile phones or computers and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0441] The RF circuit 1810 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 1880; in addition, it transmits uplink data to the base station.

[0442] The memory 1815 can be used to store software programs and modules. The processor 1880 executes various functional applications and data processing of the target terminal by running the software programs and modules stored in the memory 1815.

[0443] The input unit 1830 can be used to receive input numeric or character information, and to generate key signal inputs related to the settings and function control of the target terminal. Specifically, the input unit 1830 may include a touch panel 1831 and other input devices 1832.

[0444] The display unit 1840 can be used to display input or provided information, as well as various menus of the target terminal. The display unit 1840 may include a display panel 1841.

[0445] Audio circuitry 1860, speaker 1861, and microphone 1862 provide an audio interface.

[0446] In this embodiment, the processor 1880 included in the terminal can execute the distributed database query method of the previous embodiment.

[0447] The terminals disclosed in this embodiment include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, and aircraft. The embodiments of this invention can be applied to various scenarios, including but not limited to quantum computing, distributed quantum computing, and superconducting quantum computing.

[0448] Figure 19This is a partial structural block diagram of a server for implementing the distributed database query method of this disclosure. The server can vary significantly due to different configurations or performance characteristics, and may include one or more Central Processing Units (CPUs) 1922 (e.g., one or more processors) and memory 1932, and one or more storage media 1930 (e.g., one or more mass storage devices) for storing application programs 1942 or data 1944. The memory 1932 and storage media 1930 may be temporary or persistent storage. The program stored in the storage media 1930 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the CPU 1922 may be configured to communicate with the storage media 1930 and execute the series of instruction operations in the storage media 1930 on the server.

[0449] A server may also include one or more power supplies 1926, one or more wired or wireless network interfaces 1950, one or more input / output interfaces 1958, and / or one or more operating systems 1941, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0450] The central processing unit 1922 in the server can be used to execute the distributed database query method of the present disclosure embodiments.

[0451] This disclosure also provides a computer-readable storage medium for storing program code for executing the distributed database query methods of the foregoing embodiments.

[0452] This disclosure also provides a computer program product comprising a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the distributed database query method described above.

[0453] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.

[0454] In this disclosure, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0455] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0456] It should be understood that in the description of the embodiments disclosed herein, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.

[0457] In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0458] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0459] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0460] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0461] It should also be understood that the various implementation methods provided in this disclosure can be combined arbitrarily to achieve different technical effects.

[0462] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.

Claims

1. A distributed database query method, characterized in that, The distributed database includes compute nodes and storage nodes, and the method is executed by the storage nodes. The method includes: The computing node receives multiple subtasks, which are obtained by the computing node from the decomposition of the query statement in the distributed data query task. Assign the same query identifier to the multiple subtasks; For each subtask, based on a pre-stored mapping relationship, a first stored data snapshot corresponding to the query identifier is obtained. The mapping relationship is used to map multiple pre-stored query identifiers to multiple stored data snapshots, and the first stored data snapshot is one of the multiple stored data snapshots. The subtask is executed based on the first stored data snapshot to obtain the subtask processing result; The processing results of each subtask are returned to the computing node so that the computing node can generate the processing results of the distributed data query task.

2. The method according to claim 1, characterized in that, The storage node includes a processing thread and a shared snapshot manager, which is used to store and maintain the mapping relationship; For each subtask, based on a pre-stored mapping relationship, obtaining the first stored data snapshot corresponding to the query identifier includes: Through the processing thread, for each subtask, the query identifier corresponding to the subtask is sent to the shared snapshot manager; The shared snapshot manager returns the first stored data snapshot corresponding to the query identifier to the processing thread.

3. The method according to claim 2, characterized in that, If the distributed data query task is a single query statement task, then before assigning the same query identifier to the multiple subtasks, the method further includes: Through the processing thread, for the first subtask among the plurality of subtasks, the compute node identifier of the compute node, the session identifier of the single query statement task, and the query identifier of the first subtask are sent to the shared snapshot manager; The shared snapshot manager generates a first storage data snapshot based on the compute node identifier and the session identifier, and stores the first storage data snapshot in the shared snapshot manager in correspondence with the query identifier. Assigning the same query identifier to the multiple subtasks includes: assigning the query identifier of the first subtask to subsequent subtasks after the first subtask through the processing thread.

4. The method according to claim 3, characterized in that, The processing thread includes multiple processing threads corresponding to the multiple subtasks, and the multiple processing threads are parallel. The step of sending the compute node identifier of the compute node, the session identifier of the single query statement task, and the query identifier of the first subtask to the shared snapshot manager through the processing thread includes: sending the compute node identifier, the session identifier, and the query identifier of the first subtask to the shared snapshot manager through the processing thread corresponding to the first subtask; The step of assigning the query identifier of the first subtask to subsequent subtasks after the first subtask through the processing thread includes: assigning the query identifier of the first subtask to the subsequent subtasks through the processing thread corresponding to the subsequent subtasks after the first subtask.

5. The method according to any one of claims 2-4, characterized in that, The step of returning the first stored data snapshot corresponding to the query identifier to the processing thread through the shared snapshot manager includes: If the shared snapshot manager stores a historical storage data snapshot corresponding to the query identifier, the historical storage data snapshot is returned to the processing thread and used as the first storage data snapshot corresponding to the query identifier. If the shared snapshot manager stores historical storage data snapshots, but the historical storage data does not correspond to the query identifier, a first storage data snapshot corresponding to the query identifier is generated, and the generated first storage data snapshot is stored in the shared snapshot manager corresponding to the query identifier to replace the historical storage data snapshot; If the shared snapshot manager does not store any historical storage data snapshots, a first storage data snapshot corresponding to the query identifier is generated, and the generated first storage data snapshot is stored in the shared snapshot manager corresponding to the query identifier.

6. The method according to claim 5, characterized in that, After returning the historical storage data snapshot to the processing thread if the shared snapshot manager stores a historical storage data snapshot corresponding to the query identifier, and using the historical storage data snapshot as the first storage data snapshot corresponding to the query identifier, the method further includes: incrementing a first count by 1, where the first count represents the number of subtasks currently using the first storage data snapshot corresponding to the query identifier; After returning the processing results of each subtask to the computing node, the method further includes: decrementing the first count by 1.

7. The method according to any one of claims 2-6, characterized in that, If the distributed data query task contains multiple query statements, and the isolation level of the distributed data query task is the first type of isolation level, the multiple subtasks are multiple first subtasks obtained by decomposing the first query statement among the multiple query statements. Assigning the same query identifier to the multiple subtasks includes: assigning a first query identifier corresponding to the first query statement to the multiple first subtasks through the processing thread; After returning the processing results of each subtask to the computing node, the method further includes: The processing thread receives multiple second subtasks derived from the decomposition of the second query statement following the first query statement from the computing node. The processing thread assigns the first query identifier to the plurality of second subtasks; Through the processing thread, for each of the plurality of second subtasks, the first query identifier is sent to the shared snapshot manager; The shared snapshot manager returns the first stored data snapshot corresponding to the first query identifier to the processing thread; The processing thread executes the second subtask based on the first stored data snapshot to obtain the processing result of the second subtask. The processing thread returns the processing results of each second subtask to the computing node.

8. The method according to claim 7, characterized in that, Before receiving multiple subtasks from the computing node, the method further includes: In response to receiving a distributed data query task start request from the computing node, the computing node identifier of the computing node, the session identifier of the distributed data query task, and the first query identifier are sent to the shared snapshot manager; The first storage data snapshot is generated based on the compute node identifier and the session identifier through the shared snapshot manager, and the first storage data snapshot is stored in the shared snapshot manager corresponding to the first query identifier.

9. The method according to claim 7, characterized in that, After returning the processing results of each subtask to the computing node, the method further includes: If the subtask processing result is the subtask processing result of multiple subtasks of the last query statement in the distributed data query task, a release request is sent to the shared snapshot manager so that the shared snapshot manager releases the stored first storage data snapshot.

10. The method according to any one of claims 2-6, characterized in that, If the distributed data query task contains multiple query statements, and the isolation level of the distributed data query task is the second type of isolation level, the multiple subtasks are multiple first subtasks obtained by decomposing the first query statement among the multiple query statements. Assigning the same query identifier to the multiple subtasks includes: assigning a first query identifier corresponding to the first query statement to the multiple first subtasks through the processing thread; After returning the processing results of each subtask to the computing node, the method further includes: The processing thread receives multiple second subtasks derived from the decomposition of the second query statement following the first query statement from the computing node. The processing thread assigns a second query identifier corresponding to the second query statement to the plurality of second subtasks. Through the processing thread, for each of the plurality of second subtasks, the second query identifier is sent to the shared snapshot manager; The shared snapshot manager returns the second storage data snapshot corresponding to the second query identifier to the processing thread; The processing thread executes the second subtask based on the second stored data snapshot to obtain the processing result of the second subtask; The processing thread returns the processing results of each second subtask to the computing node.

11. The method according to claim 10, characterized in that, Before receiving multiple subtasks from the computing node, the method further includes: In response to receiving a distributed data query task start request from the computing node, the computing node identifier of the computing node, the session identifier of the distributed data query task, and the first query identifier are sent to the shared snapshot manager; The first storage data snapshot is generated based on the compute node identifier and the session identifier through the shared snapshot manager, and the first storage data snapshot is stored in the shared snapshot manager corresponding to the first query identifier; Before receiving multiple second subtasks derived from the decomposition of the second query statement following the first query statement from the computing node via the processing thread, the method further includes: The compute node identifier of the compute node, the session identifier of the distributed data query task, and the second query identifier corresponding to the second query statement are sent to the shared snapshot manager. The shared snapshot manager generates a second storage data snapshot based on the compute node identifier and the session identifier, and stores the second storage data snapshot in the shared snapshot manager in correspondence with the second query identifier.

12. The method according to any one of claims 1-10, characterized in that, The storage node includes a processing thread and a shared snapshot manager; After returning the processing results of each subtask to the computing node, the method further includes: In response to the detection of a user connection interruption, the connection between the storage node and the computing node is disconnected via the processing thread; The processing thread decrements the first count corresponding to the first storage data snapshot by 1 and sends a snapshot release instruction to the shared snapshot manager, so that the shared snapshot manager moves the first storage data snapshot to the snapshot list according to the snapshot release instruction, and releases the first storage data snapshot when the first count corresponding to the first storage data snapshot is 0.

13. The method according to any one of claims 1-10, characterized in that, The storage node includes a processing thread and a shared snapshot manager; After returning the processing results of each subtask to the computing node, the method further includes: In response to detecting an abnormal restart or shutdown of the computing node, the connection between the storage node and the computing node is disconnected via the processing thread; The processing thread decrements the first count corresponding to the first storage data snapshot by 1 and sends a snapshot release instruction to the shared snapshot manager, so that the shared snapshot manager moves the first storage data snapshot to the snapshot list according to the snapshot release instruction, and releases the first storage data snapshot when the first count corresponding to the first storage data snapshot is 0.

14. A distributed database query device, characterized in that, The distributed database includes computing nodes and storage nodes, and the device is executed by the storage nodes. The device includes: A receiving unit is configured to receive multiple subtasks from the computing node, wherein the multiple subtasks are obtained by the computing node from the decomposition of the query statement in the distributed data query task; An allocation unit is used to allocate the same query identifier to the multiple subtasks; The acquisition unit is used to acquire, for each subtask, a first storage data snapshot corresponding to the query identifier based on a pre-stored mapping relationship, wherein the mapping relationship is used to map between multiple pre-stored query identifiers and multiple storage data snapshots, and the first storage data snapshot is one of the multiple storage data snapshots; An execution unit is configured to execute the subtask based on the first stored data snapshot to obtain the subtask processing result; The return unit is used to return the processing results of each subtask to the computing node so that the computing node can generate the processing results of the distributed data query task.

15. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the distributed database query method according to any one of claims 1 to 13.

16. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the distributed database query method according to any one of claims 1 to 13.

17. A computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the distributed database query method according to any one of claims 1 to 13.