Query method and system of distributed database, medium and program product

By analyzing query requests in a distributed database system, identifying local execution conditions and dynamically calculating the number of execution packets, and prioritizing large data packets, the problems of resource competition and performance degradation in high concurrent queries are solved, and query performance and resource utilization efficiency are improved.

CN120371880AActive Publication Date: 2025-07-25北京镜舟科技有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510872870.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-25
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In the high concurrency query scenario, in the distributed database system, all execution units are started simultaneously, resulting in resource competition and performance degradation, especially frequent CPU cache switching, affecting query performance.

Method used

By analyzing query requests, identifying local execution conditions, dynamically compute the number of execution packets, and determining the priority sequence based on the data volume size, prioritizing large data volume packets, avoiding competition from multiple execution threads for CPU cache.

Benefits of technology

Effectively control the peak memory usage, improve query performance and resource utilization efficiency, avoid resource waste and stagnation of execution unit processing, and ensure the continuity and stability of query execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371880A_ABST
    Figure CN120371880A_ABST
Patent Text Reader

Abstract

The invention discloses a query method and system of a distributed database, a medium and a program product, and relates to the field of electrical digital data processing.The method comprises the steps that a query request is analyzed, and a calculation operation type and a data distribution key are obtained; determining an input table, and generating a local execution identifier when the input table contains a data distribution key; when a local execution identifier is detected, the number N of execution groups is calculated according to the number of system CPU cores, the minimum number of scanning lines and the number of physical groups; performing grouping calculation to generate N groups to be executed; obtaining the data volume of each to-be-executed group, and generating an execution priority sequence; calculating the to-be-executed group with the highest priority in the execution priority sequence to generate an execution result; and after the N to-be-executed groups are executed, generating a query result based on a corresponding execution result. By implementing the method and the device, the resource use efficiency of aggregation and connection operations in a distributed database system can be optimized, and the query performance is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic digital data processing, and in particular, to a query method, system, medium, and program product for a distributed database. Background Art

[0002] With the popularization of big data applications, distributed database systems need to process increasingly large-scale data query and analysis tasks. Among these tasks, aggregation operations and table join operations are the most common and resource-consuming operations. To ensure query performance and resource utilization efficiency, distributed database systems need to reasonably schedule and allocate computing resources.

[0003] In related technologies, when a distributed database system performs aggregation and join operations, it usually adopts the data redistribution (Shuffle) method to send data with the same key value to the same computing node for processing. When the data distribution meets the computing requirements, the system will choose to directly execute the calculation locally to avoid the network overhead caused by data redistribution. During the execution phase, the system will start all computing nodes to process data in parallel at the same time, and each node maintains an independent memory space for data caching and calculation.

[0004] However, in high-concurrency query scenarios, the execution method of starting all computing nodes at the same time will cause a significant increase in the peak value of system resource usage, and the parallel processing of different data sets by multiple execution threads will cause frequent switching of the CPU cache, reducing the cache hit rate and affecting query performance. Summary of the Invention

[0005] This application provides a query method, system, medium, and program product for a distributed database, which is used to optimize the resource usage efficiency of aggregation and join operations in a distributed database system and ensure query performance.

[0006] In a first aspect, this application provides a query method for a distributed database, which is applied to a data processing system. The method includes: receiving a query request and parsing it to obtain a calculation operation type and a corresponding data distribution key; determining an input table corresponding to the query request, and generating a local execution identifier when the input table contains the data distribution key; when detecting the local execution identifier, calculating an execution group number N based on the number of system CPU cores, the minimum number of scanned rows, and the number of physical groups; the N is a positive integer; performing grouped calculations according to the distribution hash values of the data in the input table to generate N pending execution groups; obtaining the data volume of each pending execution group to generate an execution priority sequence sorted in descending order of data volume; calculating the pending execution group with the highest priority in the execution priority sequence to generate an execution result; after all N pending execution groups have been executed, generating a query result based on the corresponding execution results.

[0007] In the above embodiments, the data processing system dynamically calculates the appropriate number of execution groups according to the CPU resources and data characteristics, and divides the data into multiple execution groups to be executed according to the distributed hash value, which can effectively control the peak value of memory usage, avoid the competition of multiple execution threads for the CPU cache, and at the same time ensure that the execution groups with a large amount of data can be processed preferentially, improving the overall query performance and resource utilization efficiency.

[0008] Combined with some embodiments of the first aspect, in some embodiments, the step of calculating the number of execution groups N according to the number of CPU cores in the system, the minimum number of scanned rows, and the number of physical groups when detecting the local execution identifier specifically includes: when detecting the local execution identifier, obtaining the number of CPU cores in the system and determining the corresponding single-threaded group expected value; the single-threaded group expected value is half of the number of CPU cores in the system; determining the first candidate value according to the product of the single-threaded group expected value and the preset grouping coefficient; determining the second candidate value according to the ratio of the average number of rows of the scanned nodes to the minimum number of scanned rows; determining the smaller value of the first candidate value and the second candidate value as the maximum number of groups; determining the larger value of the maximum number of groups and the single-threaded group expected value as the expected number of groups; determining the smaller value of the expected number of groups and the number of physical groups as the number of execution groups N.

[0009] In the above embodiments, the data processing system calculates the single-threaded group expected value based on the number of CPU cores, avoiding the generation of too small execution groups and preventing the performance loss caused by excessive serialization.

[0010] Combined with some embodiments of the first aspect, in some embodiments, the calculation operation types include aggregation operations and join operations; the steps of determining the input table corresponding to the query request and generating the local execution identifier when the input table contains a data distribution key specifically include: determining the input table corresponding to the query request; when the calculation operation type is a join operation, determining whether the input tables belong to the same co-distribution group; if so, generating the local execution identifier when it is determined that the input tables are the same table and the data distribution key contains the join condition column.

[0011] In the above embodiments, when processing a join operation, the data processing system first determines whether the input tables belong to the same co-distribution group and checks whether the data distribution key contains the join condition column, reducing the network transmission overhead and ensuring the accuracy of the query result at the same time.

[0012] In conjunction with some embodiments of the first aspect, in some embodiments, after receiving a query request and parsing it to obtain a calculation operation type and corresponding data distribution keys, the method further includes: determining an input table corresponding to the query request, and when the input table does not contain data distribution keys, obtaining the resource usage status of all computing nodes; calculating the available resource scores of each computing node according to the resource usage status; sorting the computing nodes based on the available resource scores, and selecting the top M computing nodes with the highest scores as candidate execution nodes; where M is a positive integer; and allocating the data records of the input table to the M candidate execution nodes.

[0013] In the above embodiments, when the input table of the data processing system does not meet the local execution condition, the system effectively balances the cluster load and improves the overall resource utilization efficiency by analyzing the resource usage status of all computing nodes.

[0014] In conjunction with some embodiments of the first aspect, in some embodiments, after all N pending execution groups have been executed and after generating a query result based on the corresponding execution results, the method further includes: collecting the execution metric information of each pending execution group; the execution metric information includes the actual execution time, resource usage, and data processing rate; and calculating a performance score for this query based on the execution metric information.

[0015] In the above embodiments, the data processing system collects the detailed execution metrics of each execution group to evaluate the execution effect.

[0016] In conjunction with some embodiments of the first aspect, in some embodiments, after the step of calculating a performance score for this query based on the execution metric information, the method further includes: when the performance score is lower than a preset performance threshold, calculating the resource utilization rate of each pending execution group based on the resource usage in the execution metric information; adjusting the preset grouping coefficient based on the resource utilization rate, and adjusting the minimum number of scanned rows based on the data processing rate in the execution metric information.

[0017] In the above embodiments, when the performance score does not meet the expectation, the system analyzes the resource utilization rate to improve the efficiency of query execution.

[0018] In conjunction with some embodiments of the first aspect, in some embodiments, after the step of obtaining the data volume of each pending execution group and generating an execution priority sequence sorted in descending order of data volume, the method further includes: obtaining the execution status information of the pending execution group; when it is determined based on the execution status information that the processing progress of the current execution group has stalled, splitting the unprocessed data of the current execution group into multiple sub-execution groups; and inserting the multiple sub-execution groups into the execution priority sequence to regenerate the execution priority sequence.

[0019] In the above embodiments, during the execution process, the data processing system monitors the status of each execution group in real time. When it is found that the processing progress of a certain execution group has stagnated, the unprocessed data will be dynamically split into multiple smaller sub-execution groups to maintain the overall execution efficiency and ensure that the query can be completed stably.

[0020] In a second aspect, an embodiment of the present application provides a data processing system, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions to cause the data processing system to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0021] In a third aspect, an embodiment of the present application provides a computer program product containing instructions. When the computer program product runs on a data processing system, it causes the data processing system to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions. When the instructions run on a data processing system, it causes the data processing system to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0023] It can be understood that the data processing system provided in the second aspect, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the method provided in the embodiments of the present application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, which will not be elaborated here.

[0024] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. Since the query request parsing and local execution judgment mechanism, as well as the grouping execution strategy based on data distribution characteristics, are adopted, the system can accurately identify the query scenarios suitable for local execution, and organize the query execution by dynamically calculating the number of execution groups and priority sorting, effectively solving the problems of resource competition and performance degradation caused by all execution units running simultaneously in the related art, and thus achieving precise control of memory usage and efficient utilization of CPU caches.

[0025] 2. Since the dynamic grouping number calculation mechanism based on multi-dimensional parameters is adopted, the system can automatically determine the most suitable number of execution groups for each query, effectively solving the problems of resource waste and performance fluctuations caused by unreasonable setting of execution parallelism in the related art, and thus achieving precise allocation of execution resources and improvement of query performance.

[0026] 3. Due to the adoption of the real-time execution status monitoring and dynamic task splitting mechanisms, the system can promptly detect and handle abnormal situations during the execution process, effectively solving the problem of overall performance degradation caused by the stagnation of the execution unit in the related art, and thus achieving the continuity and stability of query execution. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a schematic diagram of a scenario in an embodiment of the present application; Figure 2 is a schematic diagram of a scenario of a query method for a distributed database in the related art; Figure 3 is a schematic diagram of a scenario of a query method for a distributed database in an embodiment of the present application; Figure 4 is a schematic flow diagram of a query method for a distributed database in an embodiment of the present application; Figure 5 is another schematic diagram of a scenario of a query method for a distributed database in an embodiment of the present application; Figure 6 is another schematic flow diagram of a query method for a distributed database in an embodiment of the present application; Figure 7 is a schematic diagram of the structure of an entity device of a data processing system in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification of the present application, the singular forms "a", "an", "the above", "the", and "this" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations including one or more of the listed items.

[0029] Hereinafter, the terms "first" and "second" are only used for descriptive purposes, and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0030] For ease of understanding, the application scenarios of the embodiments of the present application are introduced below.

[0031] Please refer to Figure 1 , Figure 1 which is a schematic diagram of a scenario in an embodiment of the present application;Figure 1 FIG. Figure 1 shows a schematic diagram of a scenario where a distributed database performs an aggregation operation in the related art. In this execution mode, the system includes multiple parallel execution units (represented by dashed boxes in the figure), and each execution unit is provided with a scan node and an aggregation node. The scan node is responsible for reading the original data from the storage layer, while the aggregation node performs grouped statistical calculations on the data. Since the initial distribution of the data may not meet the requirements of the aggregation calculation, the system needs to perform data redistribution between the execution units (the arrow lines marked as "redistribution" in the figure). Specifically, data records with the same aggregation key value need to be transmitted to the same aggregation node for processing, which results in a large amount of network data transmission. After all the aggregation nodes complete the calculation, the results are then passed to the top-level output node to complete the entire query processing process.

[0032] In the related art, queries can be performed through parallel processing. The following describes a scenario of using the query method of the distributed database in the related art.

[0033] Please refer to Figure 2 , Figure 2 which is a schematic diagram of a scenario of the query method of the distributed database in the related art; Figure 2 FIG. Figure 2 shows a local execution scenario when the data distribution meets the calculation requirements. In this case, since the data distribution key of the table contains the columns required for the aggregation calculation, the data read by each scan node can be directly processed in the local aggregation node without data redistribution. Although this execution method avoids the network transmission overhead, there are still significant problems: all execution units will start and run simultaneously, and each aggregation node needs to allocate and maintain an independent memory space to store the intermediate calculation results. In a large-scale data processing scenario, this parallel execution method will cause the memory usage of the system to increase sharply. At the same time, multiple execution threads process different data sets in parallel, which will cause frequent switching of the CPU cache, reduce the cache hit rate, and affect the calculation performance.

[0034] However, by using the query method of the distributed database in the embodiments of the present application, the overall query efficiency is improved by preferentially processing large data volume groups. The following describes a scenario of using the query method of the distributed database in the present application.

[0035] In view of the above problems, the present application proposes an optimized execution strategy, as Figure 3 shown. Please refer to Figure 3 , Figure 3It is a schematic diagram of a scenario of the query method for a distributed database in an embodiment of the present application. The system organizes multiple execution units into a waiting queue, and each execution unit in the queue (Pending Execution Group 1 to Pending Execution Group N in the figure) contains a complete combination of scan nodes and aggregation nodes. The system determines the priority of each execution group according to the data volume of the execution group, and preferentially schedules the execution group with a larger data volume. During the specific execution process, the system activates only one execution group for calculation at a time. After the execution group completes the processing, the next execution group in the queue will be scheduled to start working. This serialized execution method ensures that only one aggregation node is performing calculations at the same time, significantly reducing the peak memory usage of the system. At the same time, since the competition of multiple execution threads for the CPU cache is avoided, the utilization efficiency of the cache is improved.

[0036] It can be seen that by adopting the query method for the distributed database in the embodiment of the present application, the problems of excessively high resource peaks and frequent cache switching caused by parallel execution in the related art can be effectively solved, thereby achieving an improvement in query performance.

[0037] For ease of understanding, the method provided in this embodiment will be described in terms of its process in combination with the above scenario. Please refer to Figure 4 , which is a schematic diagram of a process of the query method for a distributed database in an embodiment of the present application.

[0038] S401. Receive a query request and parse it to obtain the calculation operation type and the corresponding data distribution key.

[0039] Among them, the query request represents a data query instruction sent by a user or an application program to the data processing system, and includes an SQL statement or other query language forms; the calculation operation type refers to the specific calculation operation involved in the query request, such as an aggregation operation (SUM, COUNT, AVG, etc.) or a join operation (JOIN); the data distribution key represents the column or column combination specified during table creation for dispersing data storage to different computing nodes, and is specified through the "DISTRIBUTED BY HASH" statement; the parsing process refers to the process of converting the query request into an internal representation form that can be recognized and processed by the data processing system.

[0040] When a data processing system receives a database query request initiated by a user, it needs to first parse and process the query request. Specifically, the data processing system first performs lexical analysis and syntactic analysis on the input SQL statement to identify each component in the query. Then, it extracts the type of calculation operation in the query. For example, for a query like "SELECT k1, k2, sum(v FROM t1 GROUP BY k1, k2)", the system will identify that this is a GROUP BY operation with a SUM aggregation calculation. At the same time, the system will parse the table structure information involved in the query to obtain the data distribution key information specified when these tables were created. This parsing process will generate an internal data structure containing the complete query semantics, providing a basis for subsequent query optimization and execution.

[0041] During the actual execution process, query request parsing may encounter problems such as syntax errors, unknown identifiers, and data type mismatches. Especially when processing queries containing complex expressions, parsing ambiguities may occur. In response to this situation, the data processing system adopts a parsing strategy based on precedence: First, it establishes an operator precedence table to clearly define the associativity and precedence of various operators; then, during the parsing process, when encountering an expression that may be ambiguous, the system will determine the correct parsing order according to the precedence rules. For example, when parsing an expression like "SELECT a + b * c", the system will process the multiplication operation first and then the addition operation. At the same time, the system also maintains a symbol table to store the identified identifiers and their attribute information, so that it can quickly detect unknown identifiers and type mismatch problems. For the detected errors, the system will generate detailed error information, including the error location, error type, and possible correction suggestions, to help users quickly locate and solve problems.

[0042] S402. Determine the input table corresponding to the query request, and generate a local execution flag when the input table contains a data distribution key.

[0043] Among them, the input table refers to the data tables participating in the query, including the main table and the associated table; containing a data distribution key means that the distribution key column specified in the table creation statement of the input table contains the columns required for the query operation; the local execution flag is a boolean flag used to indicate whether this query can be directly executed on the local node without data redistribution; the determination process refers to the process of judging whether the query conditions match the data distribution by analyzing the metadata information of the table.

[0044] After the data processing system finishes parsing a query request, it needs to determine whether local computation optimization can be performed. Specifically, the data processing system first obtains all input table information involved in the query, including metadata such as the physical storage structure of the table and the distribution key settings. Then the system processes two main scenarios separately: for aggregation operations, it checks whether the columns in the GROUP BY clause are fully included in the distribution key of the table; for join operations, it needs to verify whether all tables participating in the join belong to the same co-distribution group and whether the columns in the join condition are included in the distribution key of the table. Only when all these conditions are met will the system generate a local execution flag indicating that this query can avoid data redistribution. This judgment process takes into account multi-table associations, subqueries, etc. in complex queries to ensure that the generated execution plan is correct and efficient.

[0045] In some embodiments, the determination of local execution conditions can be achieved in multiple ways: Optionally, the data processing system can adopt a rule-based judgment method. First, it constructs a set containing the query condition columns and the distribution key columns, and then determines whether the local execution conditions are met through the inclusion relationship of the sets. The specific process includes: extracting the GROUP BY columns or JOIN condition columns in the query to construct a condition set; extracting the distribution key columns from the table metadata to construct a distribution key set; and judging whether the condition set is a subset of the distribution key set. Optionally, the data processing system can also adopt a graph-based analysis method, representing the distribution relationship between tables as a graph structure and analyzing the compatibility of data distribution through a graph traversal algorithm. This method is particularly suitable for processing complex multi-table association queries, and the specific steps include: constructing a graph structure representing the inter-table association relationship; analyzing the connected components in the graph to determine the compatibility of data distribution; and deciding whether local execution can be performed based on the analysis results. It can be understood that other analysis methods can also be used to implement the judgment of local execution conditions, which are not limited here. In addition, the system also needs to consider various complex situations in the query, such as expression calculation, subqueries, etc., which requires more complex semantic analysis and judgment logic.

[0046] During the actual execution process, the judgment of local execution conditions may encounter problems where it is difficult to judge the matching of distribution keys due to complex expressions. For example, when the query condition contains function calls or complex expressions, simple column name matching may not be able to correctly determine whether the local execution conditions are met. In response to this situation, the data processing system adopts an expression normalization strategy: First, convert the expressions in the query condition into a standard form, which includes operations such as expanding function calls, folding constants, and simplifying expressions; then analyze the standardized expressions to determine whether they can utilize the distribution characteristics of the data. The system also maintains a function attribute registry to record the impact of different functions on the data distribution characteristics. For example, some functions may destroy the data distribution characteristics and cause local execution to be impossible. Through this fine-grained analysis, the system can more accurately judge whether a query meets the local execution conditions.

[0047] S403. When the local execution flag is detected, calculate the number of execution groups N based on the number of system CPU cores, the minimum number of scanned rows, and the number of physical partitions.

[0048] Among them, the number of system CPU cores represents the number of processor cores available for computing in the data processing system; the minimum number of scanned rows is the threshold of the minimum amount of data processed by a single execution group, used to avoid generating overly small execution groups; the number of physical partitions represents the number of data distribution buckets (BUCKETS) specified during table creation; the number of execution groups N represents the final determined number of execution groups to be executed; the calculation process refers to the process of comprehensively evaluating multiple parameters to obtain the optimal number of execution groups.

[0049] After the data processing system confirms that local execution can be performed, it needs to reasonably plan the allocation of execution resources. Specifically, the data processing system first obtains the current available number of CPU cores, and usually sets the single-thread parallel execution degree to half of the number of CPU cores to reserve resources for processing other tasks. Then the system will evaluate the actual data processing situation, including the average number of rows in the scanned nodes and the preset minimum number of scanned rows threshold, to avoid generating overly small execution groups that cause additional scheduling overhead. At the same time, the system also needs to consider the limitations at the physical storage level to ensure that the number of execution groups does not exceed the number of buckets specified during table creation to maintain the integrity of data distribution. During this process, the system will balance the parallelism and resource consumption, dynamically adjust the number of groups, and finally obtain an execution group number N that balances the parallel processing ability and system overhead.

[0050] In some embodiments, the calculation of the number of execution groups can be achieved in various ways: Optionally, the data processing system can adopt a calculation method based on resource evaluation. The steps include: calculating the expected single-thread parallelism expectedGroup = number of CPU cores / 2; calculating the initial upper limit of grouping maxGroup = min(expectedGroup * groupScale, configMaxGroup) in combination with a preset grouping coefficient; adjusting maxGroup = min(avgScanRows / minScanRows, maxGroup) according to the actual data volume; and obtaining the final number of execution groups N = min(maxGroup, physicalBuckets) considering the physical grouping limit. Optionally, the data processing system can also adopt an adaptive calculation method to dynamically adjust the grouping parameters according to historical execution data. The steps include: collecting the execution statistics of historical queries; analyzing the execution efficiency under different numbers of groups; and dynamically adjusting the grouping coefficient and the minimum scan row threshold according to the analysis results. It can be understood that other calculation methods can also be used to determine the number of execution groups, which are not limited herein.

[0051] During the actual execution process, the calculation of the number of execution groups may encounter the problem of load imbalance caused by data skew. For example, when the data is unevenly distributed among physical buckets, simply dividing the execution groups according to the number of buckets may result in some execution groups processing significantly more data than others. In response to this situation, the data processing system adopts a data-aware grouping strategy: First, sample and count the data volume in the physical buckets to evaluate the degree of data distribution uniformity; then, reorganize the original physical grouping according to the statistical results, merge the physical buckets with similar data volumes into one execution group, or split the physical bucket with too large data volume into multiple execution groups. The system also sets a data volume difference threshold to trigger regrouping when the data volume difference between execution groups exceeds the threshold. This method can effectively balance the processing load of each execution group and improve the overall execution efficiency. At the same time, the system will keep monitoring the execution situation and dynamically adjust the grouping strategy when necessary.

[0052] S404. Perform grouping calculation according to the distribution hash values of the data in the input table to generate N execution groups to be executed.

[0053] Among them, the distribution hash value represents the hash value calculated on the distribution key column of the data record and is used to determine the physical storage location of the data; the grouping calculation refers to the process of dividing the data into different execution groups according to the hash value; and the execution group to be executed represents an execution unit containing a set of relevant data records and has an independent scan node and calculation node.

[0054] After determining the number of execution groups, the data processing system needs to perform specific grouping based on the distribution characteristics of the data. Specifically, the data processing system first calculates the hash value of the values in the distribution key column using the consistent hashing algorithm to ensure that the same distribution key value can obtain the same hash value. Then, according to the number of execution groups N calculated in the previous step, the hash value space is evenly divided into N ranges, and each range corresponds to an execution group to be executed. For each data record in the input table, the system determines the execution group to which it belongs according to the hash value of its distribution key. This process needs to ensure the integrity and consistency of the data, that is, data records with the same distribution key value must be assigned to the same execution group to be executed to ensure the correctness of the calculation result. At the same time, the system also needs to create independent scan nodes and calculation nodes for each execution group to be executed to prepare for subsequent parallel processing.

[0055] S405. Obtain the data volume of each execution group to be executed and generate an execution priority sequence sorted in descending order of the data volume.

[0056] Among them, the data volume represents the total number of data records included in each execution group to be executed; the execution priority sequence refers to the queue of execution groups to be executed sorted according to the data volume size, which is used to determine the execution order; the descending order means that the execution group with a larger data volume has a higher execution priority.

[0057] After the data processing system completes data grouping, it needs to reasonably arrange the processing order of each execution group. Specifically, the data processing system first traverses all execution groups to be executed and counts the actual number of data records included in each execution group. This statistical process takes into account the size and complexity of the data records, and weight adjustment is performed for records containing large fields or complex data types. Then the system sorts all execution groups to be executed in descending order according to their data volume to generate a priority queue. This sorting strategy ensures that the execution group with a larger data volume can obtain processing resources first, which helps to improve the overall execution efficiency. At the same time, the system also records other characteristic information of each execution group, such as data distribution characteristics, estimated execution time, etc., to provide a basis for subsequent execution scheduling.

[0058] In some embodiments, the determination of the execution priority can be achieved in multiple ways: Optionally, the data processing system can adopt a multi-dimensional scoring method, and the steps include: calculating the basic data volume score; adjusting by considering the data complexity factor; weighting by combining historical execution statistics; generating a comprehensive priority score; sorting according to the score. Optionally, the data processing system can also adopt an adaptive priority strategy to dynamically adjust the priority by real-time monitoring of the system resource status. The specific steps include: monitoring the system resource usage; evaluating the resource requirements of the execution group; adjusting the priority according to the resource matching degree. It can be understood that other ways can also be adopted to determine the execution priority, which is not limited here.

[0059] During the actual execution process, the priority ranking may encounter problems with resource estimation deviation caused by complex queries. For example, when a query contains complex calculation expressions or user-defined functions, relying solely on the data volume may not accurately reflect the actual processing complexity. In response to this situation, the data processing system adopts a cost model-assisted priority calculation strategy: First, establish a basic cost model for query operations, including dimensions such as CPU calculation cost, memory usage cost, and I / O cost; then estimate the cost of various operations in the query, and calculate the comprehensive cost of the execution group in combination with the data volume; finally, determine the execution priority according to the comprehensive cost. The system will also maintain the cost statistical information of historical executions for continuously optimizing the accuracy of cost estimation. When a significant deviation is found between the actual execution cost and the estimated value, the system will trigger a dynamic adjustment of the cost model.

[0060] S406. Calculate the execution group with the highest priority in the execution priority sequence to generate an execution result.

[0061] Among them, performing the calculation means performing specific processing operations on the data in the execution group, such as aggregation calculation or join operation; the highest priority means the execution group ranked at the front in the execution priority sequence; the execution result means the intermediate result data generated after a single execution group completes the calculation.

[0062] After determining the execution priority, the data processing system starts to sequentially process each execution group. Specifically, the data processing system first obtains the execution group with the highest priority from the execution priority sequence and activates the scan node of this execution group to start data reading. The scan node will adopt a batch reading strategy to load the data into the memory. Then the system performs corresponding calculation operations according to the query type. For example, for an aggregation operation, it will construct a hash table to store grouped data and perform aggregation calculation; for a join operation, it will establish a join index and perform join matching. During the calculation process, the system will monitor the resource usage situation in real time, including memory occupancy, CPU utilization, etc., to ensure that the processing of a single execution group will not cause excessive pressure on the system. After completing the calculation, the system will store the execution result in the specified result buffer and release the relevant calculation resources.

[0063] In some embodiments, the calculation and processing of the groups to be executed can be achieved in various ways: Optionally, the data processing system can adopt a pipeline processing method, and the steps include: starting a scanning pipeline to read data in batches; activating a calculation pipeline to process the data; maintaining the calculation state to ensure data consistency; generating stage results; and merging the final execution results. Optionally, the data processing system can also adopt a vectorized execution method to improve the calculation efficiency through batch data processing. The specific steps include: organizing the data into a vector format; applying vectorized calculation operations; and optimizing the memory access mode. It can be understood that other execution methods can also be adopted to achieve the calculation and processing of the groups to be executed, which are not limited herein.

[0064] In the actual execution process, the calculation and processing may encounter the problem of memory overflow risk. Especially when processing aggregation operations or large table join operations involving a large number of groups, the intermediate results may occupy too much memory. In response to this situation, the data processing system adopts an adaptive memory management strategy: First, it evaluates the memory requirements of the execution groups and sets a reasonable upper limit for memory usage; then it monitors the memory usage situation in real time during the execution process, and triggers an overflow handling mechanism when the preset threshold is reached. The specific handling methods include: temporarily storing some intermediate results on the disk and retaining the hot data in the memory; adopting a partition processing method to split large-scale calculation tasks into multiple small batches for serial processing; dynamically adjusting the calculation parallelism to avoid processing too much data simultaneously. The system will also maintain statistical information on memory usage for optimizing the resource allocation strategy of subsequent execution groups.

[0065] S407. After all N groups to be executed are completed, generate a query result based on the corresponding execution results.

[0066] Among them, "completed" means that all N groups to be executed have completed their respective calculation and processing; "query result" means the final output data obtained by merging the execution results of all execution groups; and the generation process refers to the operation process of summarizing, merging, and post-processing multiple execution results.

[0067] After the data processing system completes the processing of all groups to be executed, it needs to perform the final result integration. Specifically, the data processing system first checks the completion status of all execution groups to ensure that there are no execution failures or timeouts. Then it selects an appropriate result merging strategy according to the query type. For example, for aggregation operations, it is necessary to merge and calculate the local aggregation results of each execution group; for join operations, it is necessary to ensure the integrity and uniqueness of the results. During the merging process, the system will perform necessary data conversion and formatting processing to ensure that the final output query result meets the format requirements specified by the user. At the same time, the system will also collect statistical information during the execution process, including processing time, resource usage, etc., for subsequent query optimization.

[0068] During the actual execution process, result merging may encounter data consistency issues. For example, when the results generated by different execution groups overlap or conflict, simple merging may lead to inaccurate results. In response to this situation, the data processing system adopts a transactional result merging strategy: First, establish a version identifier for each execution result, recording its generation time and dependencies; then check version consistency during the merging process to ensure that the merging operation is executed in the correct order. The system also implements a rollback mechanism. When an abnormal situation is found during the merging process, it can roll back to the previous consistent state and merge again. For scenarios that require deduplication, the system will maintain a set of globally unique identifiers to ensure that no duplicate data appears in the final result. At the same time, the system will also log the merging process for easy problem location and result verification.

[0069] The following provides a more specific process description of the method provided in this embodiment. Please refer to Figure 6 , which is another process schematic diagram of the query method for the distributed database in the embodiment of the present application.

[0070] S601. Receive a query request and parse it to obtain the calculation operation type and the corresponding data distribution key.

[0071] Referring to step S401, the data processing system will parse the query request.

[0072] In some embodiments, the data processing system will perform candidate planning, that is, the data processing system will determine the input table corresponding to the query request, and when the input table does not contain the data distribution key, obtain the resource usage status of all computing nodes; according to the resource usage status, calculate the available resource score of each computing node; sort the computing nodes based on the available resource score, and select the top M computing nodes with the highest scores as candidate execution nodes; allocate the data records of the input table to the M candidate execution nodes.

[0073] Among them, the resource usage status represents the current resource consumption situation of the computing node, including indicators such as CPU utilization rate and memory occupancy rate; the available resource score represents the node availability score comprehensively calculated based on multiple resource dimensions; the candidate execution node represents the computing node selected to execute the data processing task; the data record allocation represents the process of allocating the data of the input table to different execution nodes according to specific rules; the M is a positive integer representing the number of candidate nodes determined by the system according to the data scale and resource situation.

[0074] When the data processing system discovers that the input table does not contain a data distribution key, it needs to re-plan the data processing strategy. Specifically, the data processing system first obtains the real-time resource usage status of all computing nodes in the cluster through the resource monitoring component. These status data include metrics in multiple dimensions such as CPU usage rate, memory occupancy rate, disk I / O load, and network bandwidth usage. Then, the data processing system will calculate a comprehensive available resource score for each computing node based on these resource metrics. The score calculation will consider the weights and threshold constraints of different resource dimensions. Next, the data processing system will sort all computing nodes in descending order according to the available resource scores, and select the top M nodes with the highest scores as candidate execution nodes according to the data processing requirements. Finally, the data processing system will allocate the data records of the input table to these M candidate execution nodes according to the principle of data volume balance.

[0075] In some embodiments, node selection and data allocation can be implemented in various ways: Optionally, the data processing system can adopt a multi-dimensional weighted resource scoring method. After standardizing each resource metric, it calculates the comprehensive score of the node by combining the resource importance weights and sets dynamic resource usage thresholds for candidate node screening. Optionally, the data processing system can also adopt a load prediction-based allocation method. By analyzing the historical load change trend and data processing characteristics of the nodes, it predicts the resource requirements during execution, thereby achieving more reasonable data allocation. It can be understood that other methods can also be used to implement node selection and data allocation. In addition, the system also needs to consider the processing mechanisms for abnormal situations such as node failures and network delays.

[0076] During the actual execution process, problems may occur in candidate node selection where the resource evaluation does not match the data characteristics. For example, when the resource demand pattern during data processing is significantly different from the initial evaluation, node selection based on static resource status may not achieve the optimal effect. In response to this situation, the data processing system adopts an adaptive resource evaluation strategy: First, it constructs a prediction model that includes data characteristics and resource consumption patterns. This model learns the resource usage rules from historical execution data through machine learning methods. Then, according to the data characteristics of the current query, it predicts the resource consumption trend on different nodes. Finally, it dynamically adjusts the resource score calculation method of the nodes based on the prediction results. The system also implements a load balancing mechanism during execution. When a significant deviation in node load is detected, it can trigger the reallocation of data to ensure the balance of the processing load.

[0077] S602. Determine the input table corresponding to the query request, and generate a local execution identifier when the input table contains a data distribution key.

[0078] Referring to step S402, the data processing system will generate a local execution identifier.

[0079] In some embodiments, the data processing system performs a homogeneous judgment, that is, the data processing system determines the input tables corresponding to the query request; when the calculation operation type is a join operation, it determines whether the input tables belong to the same co-distribution group; if so, when it is determined that the input tables are the same table and the data distribution key contains the join condition columns, a local execution flag is generated.

[0080] Among them, the input tables corresponding to the query request represent the set of data tables participating in the query operation; the join operation represents a database operation for performing association calculations on multiple tables; the co-distribution group refers to a set of tables with the same distribution characteristics; the data distribution key represents the columns specified for data distribution when creating a table; the join condition column represents the column used for table association in the join operation; the local execution flag represents a flag that allows execution on the local node without data redistribution.

[0081] After the data processing system parses the query request, it needs to determine whether local execution optimization can be performed. Specifically, the data processing system first identifies all the input tables involved in the query and obtains the metadata information of these tables. When it is found that the query contains a join operation, the data processing system further checks whether these input tables belong to the same co-distribution group, that is, whether they adopt the same distribution strategy and distribution key. For the self-join scenario (i.e., the join operation involves the same table), the data processing system also verifies whether the distribution key of the table contains the columns used in the join condition. Only when all these conditions are met will the data processing system generate a local execution flag, indicating that the join operation can be completed locally without data redistribution.

[0082] In some embodiments, the local execution judgment of the join operation can be implemented in multiple ways: Optionally, the data processing system can adopt a static analysis method based on metadata. The specific steps include: obtaining the distribution key definition of the table from the system catalog; parsing the join condition to extract the involved columns; establishing a mapping relationship between the distribution key and the join condition columns; verifying the integrity of the mapping relationship; and checking the compatibility of the distribution strategy. Optionally, the data processing system can also adopt a dynamic analysis method based on the execution plan to verify the feasibility of local execution by simulating data distribution. The specific steps include: generating sample data distribution; analyzing the data flow path; evaluating the data movement cost; and determining the optimal execution strategy. It can be understood that other methods can also be used to implement the local execution judgment of the join operation, which is not limited here. In addition, the system also needs to consider the judgment logic in complex join scenarios, such as multi-table joins and cases where function calls are included in the conditional expressions.

[0083] For the join operation, the present application also adopts an optimized strategy of grouped execution. Please refer to Figure 5 , Figure 5It is another schematic diagram of the query method of the distributed database in the embodiment of the present application; as Figure 5 shown, in the traditional execution mode (the upper part of the figure), the system will simultaneously start multiple execution units for connection operations, and each unit contains two scan nodes and one connection node. This parallel execution mode will also bring problems of memory usage and CPU cache efficiency. In the optimized solution (the lower part of the figure), the system organizes these execution units into a waiting queue and schedules each execution unit in turn according to the pre-determined priority order. It should be noted that this optimization method is not only applicable to the self-join operation of the same table, but also applicable to the join operation between different tables belonging to the same collaborative distribution group, as long as these tables have the same data distribution characteristics.

[0084] Through this grouped execution method, the system can effectively control the peak value of resource usage and improve the computing efficiency. Specifically, when processing a query containing aggregation or join operations, the system will first parse the query request to determine whether it meets the conditions for local execution. If the conditions are met, it will calculate the appropriate number of execution groups according to factors such as the number of CPU cores of the system and the data distribution characteristics, and sort the priorities of these execution groups according to the data volume size, and finally complete the query processing through sequential scheduling.

[0085] S603. When the local execution identifier is detected, obtain the number of CPU cores of the system and determine the corresponding single-thread grouped expected value.

[0086] Among them, the number of CPU cores of the system represents the current available processor core number of the data processing system; the single-thread grouped expected value represents the ideal execution parallelism calculated based on the number of CPU cores; detecting the local execution identifier means confirming that the current query meets the local execution conditions.

[0087] After the data processing system confirms that the query can be executed locally, it first needs to calculate the appropriate execution parallelism. Specifically, the data processing system will first obtain the current available number of CPU cores. Considering system stability and the resource requirements of other tasks, the single-thread grouped expected value is usually set to half of the number of CPU cores. This setting can not only ensure the parallel efficiency of query processing, but also reserve sufficient computing resources for other tasks of the system. At the same time, the system will also consider the actual load status of the CPU and dynamically adjust this ratio when necessary to ensure the reasonable allocation of system resources.

[0088] In some embodiments, the calculation of the single-threaded grouped expected value can be achieved in various ways: Optionally, the data processing system can adopt a dynamic evaluation method, and the steps include: obtaining the system CPU usage statistics; analyzing the resource consumption patterns of historical queries; adjusting the allocation ratio according to the current system load status; calculating the final expected value; and applying the minimum and maximum value limits. Optionally, the data processing system can also adopt an adaptive calculation method to dynamically adjust the expected value by real-time monitoring of the system performance metrics. It can be understood that other methods can also be used to implement the calculation of the single-threaded grouped expected value, which is not limited herein.

[0089] During the actual execution process, the calculation of the single-threaded grouped expected value may encounter the problem of system load fluctuations. For example, when multiple resource-intensive tasks are running simultaneously in the system, simple fixed-ratio calculations may not accurately reflect the available resources. In response to this situation, the data processing system adopts a resource-aware calculation strategy: First, a monitoring mechanism for system resource usage is established to collect metrics such as CPU utilization and memory usage in real time; then, a resource scoring model is constructed based on these metrics to dynamically adjust the calculation ratio of the single-threaded grouped expected value. The system also maintains a historical statistical information of resource usage to predict the short-term resource change trend and further optimize the calculation of the expected value. When a significant change in the system load is detected, the system will trigger a recalculation of the expected value to ensure that the execution parallelism always remains within a reasonable range.

[0090] S604. Determine a first candidate value according to the product of the single-threaded grouped expected value and a preset grouping coefficient.

[0091] Wherein, the preset grouping coefficient represents a system configuration parameter used to adjust the number of groups, usually greater than 1; the first candidate value represents the preliminary grouping upper limit calculated through the single-threaded expected value and the grouping coefficient; and the product calculation represents the operation of multiplying the single-threaded grouped expected value by the preset grouping coefficient.

[0092] After obtaining the single-threaded grouped expected value, the data processing system needs to perform an amplification adjustment based on the actual bearing capacity of the system. Specifically, the data processing system first obtains the preset grouping coefficient from the configuration, which is usually determined based on the system's historical execution data and performance test results. Then, the single-threaded grouped expected value is multiplied by this coefficient to obtain the first candidate value. This calculation process takes into account the resource fluctuations and task scheduling overhead during the query execution process, providing a relatively loose upper limit value for the subsequent grouping calculation.

[0093] During the actual execution process, the use of the preset grouping coefficient may encounter problems with changes in query complexity. In response to this situation, the data processing system adopts a query-aware coefficient adjustment strategy: First, analyze the complexity characteristics of the query, including the number of tables involved, the complexity of calculation expressions, etc.; then dynamically adjust the grouping coefficient according to these characteristics, using a smaller coefficient for complex queries and a larger coefficient for simple queries.

[0094] S605. Determine the second candidate value according to the ratio of the average number of rows and the minimum number of scanned rows of the scan nodes.

[0095] Among them, the average number of rows of the scan nodes represents the average of the number of data rows in all scan nodes, which is used to reflect the overall situation of data distribution; the minimum number of scanned rows represents the minimum number of data records that a single execution group needs to process as set by the system, which is used to avoid too small execution groups; the second candidate value represents the upper limit of the number of groups calculated from the data volume dimension; the ratio represents the quotient obtained by dividing the average number of rows of the scan nodes by the minimum number of scanned rows.

[0096] When determining the number of groups, the data processing system needs to make a reasonable plan based on the actual data volume characteristics. Specifically, the data processing system first traverses all scan nodes, counts the number of data records actually contained in each node, and calculates the average value. Then the system obtains the pre-configured minimum number of scanned rows threshold, which is usually set based on the system's processing capacity and historical execution experience. Then the system divides the average number of rows of the scan nodes by the minimum number of scanned rows to obtain the second candidate value. This calculation process ensures that each execution group can obtain a sufficient amount of data for processing, avoiding problems such as excessive task scheduling overhead or reduced execution efficiency caused by overly fine grouping.

[0097] In some embodiments, the grouping calculation based on the data volume can be implemented in multiple ways: Optionally, the data processing system can adopt an adaptive calculation method, and the specific steps include: collecting data statistical information of all scan nodes; analyzing the uniformity of data distribution; dynamically adjusting the minimum number of scanned rows according to the data distribution characteristics; calculating the preliminary grouping ratio; applying upper and lower limit constraints to obtain the final second candidate value. Optionally, the data processing system can also adopt a multi-dimensional evaluation method, which, in addition to considering the number of rows, also combines the complexity and processing cost of the data for comprehensive calculation. The specific steps include: evaluating the average size of data records; analyzing the complexity of data types; calculating the processing cost weight; comprehensively calculating the grouping ratio considering multiple factors. It can be understood that other methods can also be used to implement the grouping calculation based on the data volume, which is not limited here.

[0098] S606. Determine the smaller value of the first candidate value and the second candidate value as the maximum number of groups.

[0099] Among them, the first candidate value represents the grouping upper limit of the system resource dimension calculated by the single-thread grouping expectation value and the preset grouping coefficient; the second candidate value represents the grouping upper limit of the data volume dimension calculated by the average number of rows of the scanning node and the minimum number of scanning rows; the maximum number of groups represents the upper limit of the number of groups determined after comprehensively considering the system resource limitations and data processing requirements; the smaller value means selecting the smaller one among multiple candidate values as the final result.

[0100] After obtaining the candidate values of the two dimensions, the data processing system needs to perform unified restriction processing. Specifically, the data processing system first compares the first candidate value with the second candidate value, and selects the smaller value as the maximum number of groups. This selection strategy adopts the conservative principle, which not only meets the constraints of system resources but also ensures the efficiency of data processing. If the larger value of the first candidate value and the second candidate value is used, it may cause problems such as over-allocation of system resources or too little data in a single execution group. By selecting a smaller value, the system can avoid wasting resources while ensuring execution efficiency.

[0101] In some embodiments, the determination of the maximum number of groups can be achieved in a variety of ways: Optionally, the data processing system can adopt a progressive comparison method, and the specific steps include: setting the initial maximum number of groups to the maximum value supported by the system; comparing it with the first candidate value and taking the smaller value; comparing the result of the previous step with the second candidate value and taking the smaller value; checking whether the result meets the minimum grouping requirement of the system; rounding up if necessary. Optionally, the data processing system can also adopt a multi-threshold constraint method, in addition to considering two candidate values, also introduces other system limiting factors, and the specific steps include: collecting the current resource limitation parameters of the system; calculating the upper limit of groups under various restrictions; uniformly comparing all upper limit values; selecting the final maximum number of groups. It can be understood that other methods can also be used to achieve the determination of the maximum number of groups, which is not limited here.

[0102] S607: Determine a larger value between the maximum number of groups and the expected value of single-thread grouping as the expected number of groups.

[0103] Among them, the maximum number of groups represents the upper limit of grouping obtained by comparing the first candidate value and the second candidate value; the expected value of single-threaded grouping represents the ideal execution parallelism calculated based on the number of CPU cores; the expected number of groups represents the target number of groups obtained on the basis of ensuring the minimum parallelism requirement; the larger value means selecting the larger one of the two input values as the final result, which is used to ensure that the execution parallelism is not lower than the benchmark value set by the system.

[0104] After determining the maximum number of groups, the data processing system needs to ensure that the final execution parallelism is not too low. Specifically, the data processing system first compares the maximum number of groups with the expected number of groups per single thread and selects the larger value as the expected number of groups. This selection strategy adopts a guarantee mechanism to ensure that the query execution can meet the minimum parallel processing requirements even when the data volume is small or the system resources are limited. If the maximum number of groups is directly used, it may be too small to fully utilize the parallel processing ability of the system and affect the query performance.

[0105] In some embodiments, the determination of the expected number of groups can be achieved in various ways: Optionally, the data processing system can adopt a performance evaluation-based approach. The specific steps include: constructing a performance prediction model for the query; analyzing the expected execution time under different numbers of groups; evaluating the impact degree of parallelism on performance; comprehensively considering resource consumption and performance benefits; and determining the final expected number of groups. Optionally, the data processing system can also adopt an adaptive adjustment approach, dynamically adjusting the expected number of groups by real-time monitoring of system performance metrics. The specific steps include: monitoring the resource utilization rate of the system; analyzing the execution efficiency of the query; adjusting the parallelism according to the performance feedback; and updating the expected number of groups. It can be understood that other ways can also be adopted to determine the expected number of groups, which are not limited herein.

[0106] During the actual execution process, the determination of the expected number of groups may encounter the problem of diminishing parallel benefits. For example, when the number of groups increases to a certain critical value, further increasing the parallelism may not bring obvious performance improvement but instead increase the system overhead. In response to this situation, the data processing system adopts a parallel benefit evaluation strategy: First, establish a benefit model for parallel processing, including factors such as the improvement of processing speed and the increase in resource consumption; then analyze the input-output ratio under different parallelisms through historical execution data; finally, dynamically adjust the expected number of groups according to the benefit evaluation results. The system also maintains a parallelism optimization record for recording the optimal parallelism configuration for different query types under different data scales. When a new query request is detected, the system will refer to the historical experience of similar queries to quickly determine the appropriate expected number of groups. At the same time, the system will continuously monitor the execution effect and adjust the parallelism configuration in a timely manner according to the actual situation.

[0107] S608: Determine the smaller value of the expected number of groups and the physical number of groups as the execution number of groups N.

[0108] Among them, the expected number of groups represents the target number of groups obtained after ensuring the minimum parallelism requirement; the physical number of groups represents the number of buckets (BUCKETS) specified during table creation; the execution number of groups N represents the finally determined actual number of execution groups; the smaller value means selecting the smaller one of the expected number of groups and the physical number of groups as the final result to ensure that the execution number of groups does not exceed the bucket limit of physical storage.

[0109] After obtaining the expected number of groups, the data processing system also needs to consider the constraints at the physical storage level. Specifically, the data processing system first obtains the number of buckets specified during input table creation, and this number represents the number of partitions of data at the physical storage level. Then it compares the expected number of groups with the physical number of groups and selects the smaller value as the final execution number of groups N. This selection strategy ensures the matching of the execution plan with the underlying storage structure and avoids the reduction of data access efficiency or additional data rearrangement overhead caused by the number of groups exceeding the physical number of buckets.

[0110] In some embodiments, the determination of the final execution number of groups can be achieved in various ways: Optionally, the data processing system can adopt a bucket-aware calculation method, and the specific steps include: analyzing the data distribution of physical buckets; evaluating the data skew degree between buckets; adjusting the grouping strategy according to the data distribution characteristics; calculating the actually mergable bucket combinations; determining the final execution number of groups. Optionally, the data processing system can also adopt a dynamic tuning method to optimize the determination of the number of groups by analyzing historical execution effects, and the specific steps include: collecting execution statistical information under different grouping configurations; analyzing the relationship between the number of groups and execution efficiency; establishing a grouping effect evaluation model; dynamically adjusting the grouping strategy. It can be understood that other ways can also be adopted to achieve the determination of the execution number of groups, which is not limited here.

[0111] During the actual execution process, the determination of the execution number of groups may encounter the problem of uneven physical bucket capacity. For example, when data is unevenly distributed among physical buckets, simply executing grouping according to the bucket number limit may lead to uneven processing load. In response to this situation, the data processing system adopts an optimization strategy that is aware of bucket capacity: First, statistically analyze the data volume distribution of physical buckets and calculate the actual data volume and processing cost of each bucket; then, based on the data distribution characteristics, merge adjacent small-capacity buckets or appropriately split large-capacity buckets; finally, determine the actual execution number of groups according to the optimized bucket combination. The system also maintains a bucket status monitoring mechanism to continuously record the data change trend of each bucket, and when it finds that the data distribution has changed significantly, it will trigger a re-evaluation and adjustment of the grouping strategy. In this way, the system can achieve more balanced and efficient data processing under the premise of physical storage constraints.

[0112] S609: Group and calculate according to the distributed hash values of the data in the input table to generate N groups to be executed.

[0113] Referring to step S404, the data processing system will generate groups to be executed.

[0114] S610: Obtain the data volume of each group to be executed and generate an execution priority sequence sorted in descending order according to the data volume.

[0115] Referring to step S405, the data processing system will generate a priority sequence.

[0116] S611: Obtain the execution status information of the groups to be executed.

[0117] Among them, the execution status information represents the real-time status data of each group to be executed during the running process; the group to be executed represents a set of data divided according to the distributed hash value and waiting to be processed; the obtaining process represents the operation process in which the system regularly collects and updates the execution status data; the status information includes key indicators such as processing progress, resource usage, and execution time.

[0118] During the execution process, the data processing system needs to monitor the running status of each group to be executed in real time. Specifically, the data processing system will establish a status monitoring mechanism for each group to be executed and regularly collect execution indicators from multiple dimensions. These indicators include: the number of processed data records, data processing rate, CPU usage rate, memory occupancy, I / O waiting time, etc. The system integrates this information to form a complete execution status view for subsequent execution anomaly detection and processing strategy adjustment. At the same time, the system will also record key events during the execution process, such as important time points like data loading completion and processing stage switching.

[0119] In the actual execution process, there may be a balance problem between performance overhead and monitoring accuracy in execution status monitoring. For example, overly frequent status collection may affect the actual data processing performance, while too long a collection interval may not be able to detect execution anomalies in time. In response to this situation, the data processing system adopts an intelligent sampling strategy: First, establish a status prediction model for the execution stage, and analyze the characteristic patterns of different stages based on historical execution data; then dynamically adjust the sampling frequency according to the prediction model, increasing the sampling frequency at key time points or when abnormal signs appear, and appropriately reducing the sampling frequency during the stable execution stage. The system also implements a caching mechanism for temporarily storing recent status data to avoid frequent persistence operations. When a potential execution anomaly is detected, the system will automatically increase the monitoring level and collect more detailed diagnostic information to provide a basis for subsequent problem analysis and processing.

[0120] S612. When it is determined that the processing progress of the current execution group has stalled based on the execution status information, split the unprocessed data of the current execution group into multiple sub-execution groups.

[0121] Among them, the execution status information represents the runtime status data of the execution group to be executed; the processing progress stalling means that the data processing rate is significantly lower than expected or there is no progress at all within a preset time window; the current execution group represents the set of data being processed; the unprocessed data represents the data records that have not completed the calculation; the sub-execution group represents multiple smaller execution units formed after splitting a large execution group.

[0122] After detecting an execution anomaly, the data processing system needs to adjust the task in a timely manner. Specifically, the data processing system first determines whether the current execution group has a processing progress stall based on the execution status information. The judgment criteria include: whether the data processing rate in the recent period is lower than a specific ratio (such as 50%) of the historical average, whether no new data processing has been completed for multiple consecutive sampling periods, whether the system resource usage is in an abnormal state, etc. When it is confirmed that the execution group processing has stalled, the system will immediately suspend the processing of this execution group and count the remaining unprocessed data volume. Then the system will start the task splitting mechanism, divide the unprocessed data into multiple smaller sub-execution groups, and prepare for subsequent rescheduling.

[0123] In some embodiments, the task splitting process can be implemented in multiple ways: Optionally, the data processing system can adopt an adaptive splitting strategy. The specific steps include: analyzing the cause of the stall, distinguishing whether it is caused by data characteristics or resource competition; determining the appropriate size of the sub-execution group according to the current resource status of the system; considering the data correlation for boundary division; allocating independent resource quotas for each sub-execution group; generating a new execution plan. Optionally, the data processing system can also adopt a progressive splitting method, gradually adjusting the task granularity through multiple rounds of splitting. The specific steps include: initially splitting the task into larger sub-tasks; monitoring the execution effect of the sub-tasks; performing secondary splitting if necessary; continuously optimizing the task division. It can be understood that other methods can also be used to implement task splitting, which is not limited here.

[0124] During the actual execution process, task splitting may encounter problems in handling data dependency relationships. For example, when there are complex association relationships between the data to be processed, simple equal division of the data volume may damage the integrity of data processing. In response to this situation, the data processing system adopts a semantics-aware splitting strategy: first, analyze the dependency relationships between data records to construct a data dependency graph; then, perform task partitioning according to the structural characteristics of the dependency graph to ensure that data with dependency relationships is assigned to the same sub-execution group; at the same time, consider the physical storage location of the data to minimize cross-node data access. The system also maintains a splitting effect evaluation mechanism to record the execution effects of different splitting schemes, which is used to optimize the task splitting strategy for subsequent similar scenarios. When it is found that there are still execution anomalies in the split sub-tasks, the system will trigger dynamic adjustment of the splitting strategy, including adjusting the splitting granularity, updating the dependency analysis rules, etc.

[0125] S613. Insert multiple sub-execution groups into the execution priority sequence and regenerate the execution priority sequence.

[0126] Among them, multiple sub-execution groups represent several smaller execution units obtained by splitting stagnant tasks; the execution priority sequence represents a queue of tasks to be executed sorted by data volume; the insertion operation represents the process of adding new sub-execution groups to the existing priority sequence; regeneration represents the operation of re-sorting the entire sequence after adding new sub-execution groups.

[0127] After the data processing system completes task splitting, it needs to reasonably arrange the execution order of sub-execution groups. Specifically, the data processing system first evaluates the characteristics of each sub-execution group, including information in dimensions such as data volume size, estimated execution time, and resource requirements. Then, these sub-execution groups are inserted into the existing execution priority sequence, and the insertion position is determined by the data volume of the sub-execution group. After the insertion is completed, the system will re-sort the entire priority sequence to ensure that the data volume is maintained in the order from large to small. This process needs to ensure the atomicity of the operation to avoid problems such as incorrect priority judgment or duplicate task execution during the sequence update process.

[0128] In some embodiments, the priority management of sub-execution groups can be achieved in various ways: Optionally, the data processing system can adopt a multi-dimensional priority calculation method. The specific steps include: constructing a comprehensive scoring model that includes factors such as data volume, complexity, and dependency relationships; calculating the priority scores for each sub-execution group; determining the insertion position based on the scores; handling conflicts when the priorities are similar; and maintaining the orderliness of the priority sequence. Optionally, the data processing system can also adopt a dynamic priority adjustment method to update the task priorities in real time according to the execution feedback. The specific steps include: monitoring the execution effects of the sub-execution groups; analyzing the performance; dynamically adjusting the priority weights; and reorganizing the execution sequence. It can be understood that other methods can also be used to manage the priority sequence, which is not limited herein.

[0129] S614. Calculate the highest-priority to-be-executed group in the execution priority sequence to generate an execution result.

[0130] Referring to step S406, the data processing system will generate an execution result.

[0131] S615. After all N to-be-executed groups have been executed, generate a query result based on the corresponding execution results.

[0132] Referring to step S407, the data processing system will generate a query result.

[0133] In some embodiments, the data processing system will perform a performance score. That is, the data processing system will collect the execution metric information of each to-be-executed group. The execution metric information includes the actual execution time, resource usage, and data processing rate. Based on the execution metric information, calculate the performance score of this query.

[0134] Among them, the execution metric information represents the quantitative measurement data of the running situation of the to-be-executed group. The actual execution time refers to the time interval from the start of processing to the completion of the calculation. The resource usage represents the total amount of computing resources consumed during the execution process. The data processing rate represents the number of data records processed per unit time. The performance score represents the query performance metric value calculated based on multiple execution metrics.

[0135] During the execution of a query, a data processing system needs to evaluate the execution effect in real time. Specifically, the data processing system will establish an independent performance monitoring mechanism for each group to be executed, and continuously collect key execution metrics. In terms of execution time, the system will record the task start time, data loading completion time, time distribution during the calculation phase, and the final end time. For resource usage, the system will count metrics such as CPU time, peak memory occupancy, and I / O operation volume. In data processing, the system will track and record the actual throughput of data, including the number of records processed per second, data scanning speed, calculation speed, etc. Based on this collected execution metric information, the data processing system calculates the overall performance score of this query through a preset scoring model, and this score reflects the efficiency of query execution and the rationality of resource utilization.

[0136] In some embodiments, the execution effect evaluation can be achieved in multiple ways: Optionally, the data processing system can adopt a scoring method based on historical benchmarks. By comparing with the historical execution data of similar queries, calculate the performance deviation and perform normalization processing to generate the final score. Optionally, the data processing system can also adopt a scoring method with dynamic weights. According to the characteristics of the query type and execution environment, adaptively adjust the weight ratio of different metrics in the scoring. It can be understood that other methods can also be used to achieve the execution effect evaluation. In addition, the system needs to consider the differences in scoring criteria for different scales of data and different types of queries, as well as the impact of changes in the execution environment on the scoring.

[0137] In some embodiments, the data processing system will adjust parameters based on the score. That is, when the performance score is lower than the preset performance threshold, the data processing system will calculate the resource utilization rate of each group to be executed based on the resource usage in the execution metric information; adjust the preset grouping coefficient based on the resource utilization rate, and adjust the minimum number of scanned rows based on the data processing rate in the execution metric information.

[0138] Among them, the performance score represents a comprehensive metric value of query execution quality; the preset performance threshold represents the lowest performance standard expected by the system; the resource utilization rate represents the actual usage efficiency of computing resources; the preset grouping coefficient represents an adjustment parameter used to calculate the number of execution groups; the minimum number of scanned rows represents the minimum amount of data that a single execution group should process; the resource usage represents the total amount of system resources consumed during the execution process.

[0139] After the data processing system completes the performance evaluation, it needs to optimize and adjust for insufficient performance. Specifically, the data processing system first compares the calculated performance score with a preset performance threshold, and triggers the optimization process when the score is lower than the threshold. The system will analyze the resource usage of each execution group to be executed, and calculate the actual utilization rates of key resources such as CPU and memory. By comparing the resource utilization differences between different execution groups, the system can identify whether the resource allocation is reasonable. Based on the analysis results of the resource utilization rate, the system will correspondingly adjust the preset grouping coefficient. When the resource utilization rate is generally low, the grouping coefficient is increased to improve the parallelism, and when the resource competition is serious, the grouping coefficient is decreased. At the same time, the system will dynamically adjust the minimum number of scanned rows according to the data processing rate observed during the execution process to ensure that each execution group can obtain an appropriate amount of data to be processed.

[0140] In some embodiments, parameter optimization and adjustment can be achieved in various ways: Optionally, the data processing system can adopt a progressive tuning method, adjust the parameters in small steps and observe the execution effects, and gradually find the optimal parameter configuration. Optionally, the data processing system can also adopt a model-based optimization method, establish a relationship model between parameters and performance, and predict the execution effects under different parameter values. It can be understood that other methods can also be used to achieve parameter optimization and adjustment. In addition, the system needs to consider the stability of parameter adjustment to avoid unstable execution behavior caused by frequent large-scale adjustments.

[0141] During the actual execution process, parameter adjustment may encounter the problem of conflicting optimization goals. For example, increasing the parallelism may improve the processing speed but reduce the resource utilization rate at the same time, and decreasing the minimum number of scanned rows may improve the load balancing but increase the scheduling overhead. In response to this situation, the data processing system adopts a multi-objective balance optimization strategy: First, construct a comprehensive evaluation model containing multiple performance indicators and assign weights to different optimization goals; then analyze the impact of parameter adjustment on each indicator based on historical execution data; finally, solve the multi-objective optimization problem to find a parameter configuration that can achieve a balance among various indicators. The system also implements a rollback mechanism for parameter adjustment. When the performance after adjustment fails to improve as expected, it can quickly return to the previous configuration. At the same time, the system will continuously accumulate experience in parameter adjustment, establish an optimization strategy library for different query scenarios, and improve the accuracy and efficiency of parameter adjustment.

[0142] In the embodiments of this application, due to the adoption of the dynamic grouping strategy, priority scheduling mechanism, and execution status monitoring mechanism, the execution plan can be adaptively adjusted according to system resources and data characteristics, effectively solving the problems of fierce resource competition, low cache utilization rate, and load imbalance in traditional parallel execution methods, and thus achieving a significant improvement in the query performance of the distributed database and an optimization of resource usage efficiency.

[0143] The data processing system in the embodiments of the present invention application will be described from the perspective of hardware processing. Please refer to Figure 7 FIG. 3, which is a schematic structural diagram of an entity device of the data processing system in the embodiments of the present application.

[0144] It should be noted that Figure 7 the structure of the data processing system shown is only an example, and should not bring any limitations to the functions and usage scope of the embodiments of the present invention.

[0145] As Figure 7 shown, the data processing system includes a CPU 701, which can perform various appropriate actions and processes according to the program stored in the ROM 702 or the program loaded from the storage section 708 into the RAM 703, such as executing the method described in the above embodiments. In the RAM 703, various programs and data required for system operation are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The I / O interface 705 is also connected to the bus 704.

[0146] The following components are connected to the I / O interface 705: an input section 706 including an audio input device, a button switch, etc.; an output section 707 including a liquid crystal display (LCD), an audio output device, an indicator light, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed, so that the computer program read from it can be installed into the storage section 708 as needed.

[0147] Specifically, according to the embodiments of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the CPU 701, various functions defined in the present invention are executed.

[0148] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the above-mentioned module, segment of a program, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings.

[0149] Specifically, the data processing system of this embodiment includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the query method of the distributed database provided in the above embodiment is implemented.

[0150] On the other hand, the present invention also provides a computer-readable storage medium. This storage medium may be included in the data processing system described in the above embodiment; or it may exist separately without being assembled into the data processing system. The above storage medium carries one or more computer programs. When the above one or more computer programs are executed by a processor of the data processing system, the data processing system implements the query method of the distributed database provided in the above embodiment.

Claims

1. A query method for a distributed database, characterized in that Applied to a data processing system, the method includes: Receiving a query request and parsing it to obtain a calculation operation type and corresponding data distribution keys; Determining an input table corresponding to the query request, and generating a local execution identifier when the input table contains the data distribution keys; When detecting the local execution identifier, calculating an execution grouping number N based on the number of system CPU cores, the minimum number of scanned rows, and the number of physical groupings; the N is a positive integer; Performing grouping calculations based on the distribution hash values of the data in the input table to generate N groups to be executed; Obtaining the data volume of each group to be executed, and generating an execution priority sequence sorted in descending order according to the data volume; Calculating the group to be executed with the highest priority in the execution priority sequence to generate an execution result; After all the N groups to be executed are completed, generating a query result based on the corresponding execution results.

2. The method according to claim 1, wherein The step of calculating the execution grouping number N based on the number of system CPU cores, the minimum number of scanned rows, and the number of physical groupings when detecting the local execution identifier specifically includes: When detecting the local execution identifier, obtaining the number of system CPU cores and determining a corresponding single-thread grouping expectation value; the single-thread grouping expectation value is half of the number of system CPU cores; Determining a first candidate value according to the product of the single-thread grouping expectation value and a preset grouping coefficient; Determining a second candidate value according to the ratio of the average number of rows of the scanning nodes to the minimum number of scanned rows; Determining the smaller value of the first candidate value and the second candidate value as the maximum grouping number; Determining the larger value of the maximum grouping number and the single-thread grouping expectation value as the expected grouping number; Determining the smaller value of the expected grouping number and the number of physical groupings as the execution grouping number N.

3. The method according to claim 1, characterized in that, The calculation operation type includes an aggregation operation and a join operation; The step of determining the input table corresponding to the query request and generating a local execution identifier when the input table contains the data distribution keys specifically includes: Determining the input table corresponding to the query request; When the calculation operation type is a join operation, determining whether the input tables belong to the same co-distribution group; If so, generating a local execution identifier when it is determined that the input tables are the same table and the data distribution keys include the join condition columns.

4. The method according to claim 1, wherein After the step of receiving the query request and parsing it to obtain the calculation operation type and corresponding data distribution keys, the method further includes: Determining the input table corresponding to the query request, and obtaining the resource usage status of all computing nodes when the input table does not contain the data distribution keys; Calculating the available resource scores of each computing node according to the resource usage status; Sorting the computing nodes based on the available resource scores, and selecting the M computing nodes with the highest scores as candidate execution nodes; the M is a positive integer; Allocating the data records of the input table to the M candidate execution nodes.

5. The method according to claim 1, characterized in that After the step of generating a query result based on the corresponding execution results after all the N groups to be executed are completed, the method further includes: Collect the execution metric information of each of the to-be-executed groups; the execution metric information includes the actual execution time, resource usage, and data processing rate; Calculate the performance score of the current query based on the execution metric information.

6. The method according to claim 5, characterized in that After the step of calculating the performance score of the current query based on the execution metric information, the method further includes: When the performance score is lower than a preset performance threshold, calculate the resource utilization rate of each of the to-be-executed groups based on the resource usage in the execution metric information; Adjust the preset grouping coefficient based on the resource utilization rate, and adjust the minimum scan line number based on the data processing rate in the execution metric information.

7. The method according to claim 1, wherein After the step of obtaining the data volume of each of the to-be-executed groups and generating an execution priority sequence sorted in descending order according to the data volume, the method further includes: Obtain the execution status information of the to-be-executed groups; When it is determined based on the execution status information that the processing progress of the current execution group has stalled, split the unprocessed data of the current execution group into multiple sub-execution groups; Insert the multiple sub-execution groups into the execution priority sequence and regenerate the execution priority sequence.

8. A data processing system, characterized in that, The data processing system includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the data processing system to execute the method according to any one of claims 1-7.

9. A computer-readable storage medium, comprising instructions, characterized in that, When the instruction runs on the data processing system, cause the data processing system to execute the method according to any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product runs on the data processing system, cause the data processing system to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • SQL (Structured Query Language) statement processing method and device

    CN114969101A

  • Data aggregation query method and device, computer equipment and storage medium

    CN117312412A

  • Query statement optimization method and device, equipment and storage medium

    CN119621743A

  • Characterizing Queries To Predict Execution In A Database

    US20100082599A1