A stage-aware based analytical database query scheduling method and system
Patent Information
- Application Number
- CN202610926345.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-25
AI Technical Summary
如果调度器不能识别这些阶段之间的差异,就可能将中央处理器资源分配给并行收益有限的任务,从而造成关键路径延长和系统吞吐量下降
[0024]本发明将查询级公平性调度与流水线级阶段感知调度解耦,既能够按照查询优先级分配中央处理器时间,又能够优先处理非弹性阶段任务,从而兼顾公平性和平均响应时间优化。查询级调度采用步长式比例共享算法,通过维护查询的通行值和步长,确保每个查询按照其优先级获得相应的CPU时间份额,避免出现长时间查询饥饿问题。任务级调度采用阶段感知策略,识别流水线中的弹性阶段和非弹性阶段,优先调度非弹性阶段任务,缩短查询的关键路径。
Smart Images

Figure CN122816792A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of database systems and query scheduling technology, and in particular to a phase-aware analytical database query scheduling method and system. Background Technology
[0002] With the rapid development of the big data era, online analytical processing (OASIC) systems have become the core infrastructure for enterprise data analysis and business intelligence decision-making. Modern OASIC workloads exhibit significant diversity, typically containing a large number of short-running queries and a small number of complex, long-running queries, which vary greatly in execution time, resource requirements, and response time limits. The query scheduler, as a key component determining how limited CPU cores are allocated among concurrent queries, directly impacts the system's response time, throughput, and performance predictability.
[0003] In existing technologies, block-driven parallel execution models are widely used in analytical database systems. These systems typically divide data tables into multiple data blocks and dynamically execute them using a fixed pool of worker threads. This model effectively avoids excessive thread creation and improves the utilization of multi-core processors, but it still has significant shortcomings at the scheduling level. Taking mainstream systems like DuckDB as an example, their task scheduling usually adopts a simple first-come, first-served strategy, which makes it difficult to identify the parallel characteristics of different pipelines within a query and lacks a query-level priority mechanism, resulting in a lack of flexibility in scheduling decisions.
[0004] In actual query execution, several different execution phases are typically involved. Phases such as scanning, filtering, and projection exhibit good parallelism and can be considered elastic phases; while phases such as hash join construction, aggregation and merging, and sorting and merging involve synchronization operations, lock contention, or serial bottlenecks and can be considered non-elastic phases. If the scheduler cannot recognize the differences between these phases, it may allocate CPU resources to tasks with limited parallel benefits, resulting in a longer critical path and reduced system throughput.
[0005] Furthermore, regarding the balance between scheduling fairness and response time, while simply adopting strategies such as shortest job first may reduce average response time to some extent, it can easily lead to starvation problems for long-running queries. Conversely, simply using fairness scheduling makes it difficult to utilize internal query stage information to optimize the critical path. At the same time, a fixed data block size is insufficient to accommodate differences in execution overhead across different pipelines; overly large task granularity can impair the responsiveness of short queries, while overly small task granularity increases scheduling overhead. Therefore, the industry urgently needs a database query scheduling method that can balance inter-query fairness, internal query stage optimization, and dynamic load adaptability, in order to shorten the critical path, reduce average response time under mixed online analytical processing loads, and improve system throughput while ensuring long-term fairness between queries. Summary of the Invention
[0006] The purpose of this invention is to provide a phase-aware analytical database query scheduling method and system.
[0007] To achieve the above objectives, the present invention is implemented according to the following technical solution:
[0008] This invention provides a phase-aware analytical database query scheduling method, comprising the following steps:
[0009] S1: Receive at least one analytical database query, divide the execution plan of each query into one or more pipelines, and generate corresponding scheduled tasks based on data blocks;
[0010] S2: Perform stage identification on the task to be scheduled to determine whether the pipeline to which the task belongs is a flexible stage, an inflexible stage, or an unknown stage. The stage identification process includes two levels: static analysis and dynamic monitoring.
[0011] S3: Create a resource group for each query. The resource group maintains the query priority, pass value, step size, task queue, and execution status.
[0012] S4: At the query level, select the target resource group with the smallest passage value from the active resource group based on step-size proportional shared scheduling, and update its step size according to the effective priority of the target resource group;
[0013] S5: At the task level, tasks are selected from the task queue of the target resource group based on a phase-aware strategy, where the scheduling priority of non-elastic phase tasks is higher than that of elastic phase tasks.
[0014] S6: Assign the selected tasks to worker threads for execution, and update the query status, pipeline runtime statistics, and subsequent data block granularity based on the task execution results.
[0015] Furthermore, in step S2, the static analysis layer determines the initial stage type during the query compilation phase based on the operator types included in the pipeline. Pipelines corresponding to scan operators, filter operators, projection operators, or hash join probe operators are marked as elastic stages, while pipelines corresponding to hash join construction operators, aggregation merge operators, or sort merge operators are marked as inelastic stages. The dynamic monitoring mechanism in step S2 includes event-driven detection, executed upon task completion. It determines whether the pipeline has degenerated into an inelastic stage based on whether the task execution time exceeds a preset multiple of the historical average execution time, or whether the thread waiting time exceeds a preset threshold.
[0016] The dynamic monitoring mechanism also includes periodic sampling and detection, calculating pipeline parallel efficiency and throughput changes according to a preset sampling period, and correcting the pipeline stage type when the parallel efficiency is lower than a first threshold or the current throughput is lower than a preset proportion of the historical throughput average.
[0017] Furthermore, the effective priority in step S4 is calculated based on the stage urgency, execution efficiency, waiting pain, and completion proximity of the query. The stage urgency is the proportion of non-elastic stage tasks in the pending tasks, the execution efficiency is the ratio of the current throughput to the historical average throughput, the waiting pain is the time interval since the last scheduling, and the completion proximity is the ratio of the number of completed tasks to the total number of tasks.
[0018] The overall urgency level F is calculated by weighted summation of stage urgency, execution efficiency, waiting pain level, and completion proximity. The weights of each indicator are adjusted according to the actual load characteristics, and the adjustment of short-term priority only affects the subsequent passage value growth rate.
[0019] Furthermore, in step S5, the task-level scheduling maintains a priority task queue within the resource group. The task priority is calculated by both the stage type and the waiting time. Non-elastic stage tasks are given a higher stage weight, and the waiting time serves as an auxiliary factor to gradually increase the priority of elastic tasks that have been waiting for a long time.
[0020] The task-level scheduling module maintains a local cache for worker threads. Worker threads prefetch multiple tasks from the task queue of the target resource group and write them to the local cache. Subsequently, tasks are retrieved from the local cache for execution.
[0021] Specifically, the data block granularity adaptive adjustment mechanism in step S6 dynamically adjusts the size of subsequent data blocks based on the execution time and throughput estimation of completed tasks in the same pipeline. When the average task execution time deviates from the preset target time, the size of subsequent data blocks is adjusted proportionally. When the pipeline enters the tail execution stage, the remaining data is evenly divided into a preset number of tasks.
[0022] This invention discloses a stage-aware analytical database query scheduling system, comprising: a query parsing and task generation module for receiving analytical database queries and generating tasks to be scheduled; a stage identification module for identifying the stage type of the pipeline to which the tasks to be scheduled belong; a resource group management module for managing the resource groups and their status for each query; a query-level scheduling module for performing step-by-step proportional shared scheduling; a task-level scheduling module for performing stage-aware task selection; a feedback control module for calculating the effective priority of queries; and a task granularity control module for dynamically adjusting the data block granularity.
[0023] The beneficial effects of this invention are:
[0024] This invention decouples query-level fairness scheduling from pipeline-level stage-aware scheduling, enabling the allocation of CPU time according to query priority while prioritizing tasks in inelastic stages, thus balancing fairness and average response time optimization. Query-level scheduling employs a step-size proportional sharing algorithm, maintaining a pass value and step size for each query to ensure each query receives its corresponding CPU time share according to its priority, avoiding prolonged query starvation. Task-level scheduling uses a stage-aware strategy to identify inelastic and inelastic stages in the pipeline, prioritizing tasks in inelastic stages to shorten the critical path of queries.
[0025] This invention identifies pipeline stage characteristics through static analysis and dynamic monitoring, and improves throughput, responsiveness, and performance predictability under mixed online analytical processing loads by responding to data skew, lock contention, and load changes through query-level feedback control and adaptive data block granularity control. Static analysis quickly determines the stage type using operator type information, while dynamic monitoring corrects the judgment based on actual execution conditions; the combination of both improves the accuracy of stage identification. Query-level feedback control dynamically adjusts the effective priority based on the query's running status, making scheduling decisions more flexible. Adaptive data block granularity control dynamically adjusts the data block size based on task execution time and pipeline throughput, matching task granularity with pipeline processing capacity. This avoids both excessively large task granularity leading to decreased responsiveness for short queries and excessively small task granularity leading to increased scheduling overhead. Attached Figure Description
[0026] Figure 1 This is a schematic diagram of the overall system architecture of an embodiment of the present invention;
[0027] Figure 2 This is a flowchart illustrating the query scheduling method in an embodiment of the present invention;
[0028] Figure 3 This is a flowchart illustrating the two-layer scheduling algorithm in one embodiment of the present invention;
[0029] Figure 4This is a schematic diagram of the query-level feedback control process in an embodiment of the present invention. Detailed Implementation
[0030] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions herein are used to explain the present invention, but are not intended to limit the present invention.
[0031] This embodiment provides a stage-aware analytical database query scheduling method and system, which can be deployed in DuckDB or other analytical database systems that adopt a data block-driven parallel execution model. The system includes a query parsing and task generation module, a stage identification module, a resource group management module, a query-level scheduling module, a task-level scheduling module, a feedback control module, and a task granularity control module. Its scheduling process is mainly summarized as follows:
[0032] S1. Receives a Structured Query Language (SQL) query, generates a physical execution plan, divides the execution plan into multiple pipelines, and generates data block tasks according to data partitions;
[0033] S2. Identify the task stage based on the pipeline operator type and runtime statistics, and label the task as a flexible stage, an inflexible stage, or an unknown stage.
[0034] S3 selects tasks for execution through query-level step-by-step proportional sharing scheduling and task-level stage-aware scheduling, and continuously updates scheduling parameters using feedback control and adaptive data block mechanisms.
[0035] In this embodiment, the query parsing and task generation module receives SQL queries submitted by the user, generates a physical execution plan, and divides the execution plan into multiple pipelines. Each pipeline generates one or more data block tasks according to data partitions. Each task records its query, pipeline, creation time, stage type, execution status, and scheduling-related statistics. The overall system architecture is as follows: Figure 1 As shown.
[0036] The following sections will introduce the phase identification, two-layer scheduling algorithm, query-level feedback control, and adaptive data block execution mechanism.
[0037] (1) Stage identification:
[0038] by Figure 2For example, the stage identification module first performs static analysis during the query compilation stage. Pipelines with no obvious synchronization dependencies, such as sequential scanning, filtering, and projection, are marked as elastic stages; pipelines with serial bottlenecks, such as hash join construction, final aggregation merging, and sorting merging, are marked as non-elastic stages; pipelines that cannot be determined are marked as unknown stages and corrected at runtime. Runtime dynamic monitoring includes event-driven detection and periodic sampling detection. Event-driven detection updates the exponential moving average of task execution time upon completion of each task. If the current task execution time exceeds a preset multiple of the historical average, or the thread waiting time exceeds a preset threshold, a degradation count is accumulated. Periodic sampling detection calculates the pipeline's parallel efficiency and throughput trends. If the parallel efficiency is below the first threshold or the throughput decreases significantly, the corresponding pipeline is corrected to a non-elastic stage; if the parallel efficiency recovers to above the second threshold and the throughput stabilizes, it can be restored to an elastic stage.
[0039] This embodiment sets an acknowledgment counter and a hysteresis interval in the stage identification process. On the one hand, stage correction is only performed after a preset number of consecutive triggers, which avoids frequent switching caused by occasional jitter; on the other hand, separating the degradation threshold and the recovery threshold allows the system to remain stable near the threshold. All of the above detections are lightweight statistics and do not change the query execution results.
[0040] (2) Two-level scheduling algorithm:
[0041] exist Figure 3 In this system, the two-layer scheduling algorithm consists of query-level step-size proportional shared scheduling and task-level stage-aware scheduling. Each query corresponds to a resource group, which maintains a query identifier, basic priority, pass value, step size, task queue, number of completed tasks, total number of tasks, creation time, and last scheduled time. All active resource groups are maintained by a min-heap, with the resource group with the smallest pass value at the top of the heap.
[0042] The query-level scheduling module performs step-based proportional shared scheduling. When the worker thread's local cache is empty, the resource group with the smallest pass value is retrieved from the min-heap as the target resource group. If the base query priority is p, the step size is the ratio of a preset constant to priority p; with feedback control, the step size is calculated from the effective priority. After scheduling, the pass value of the target resource group is increased by the corresponding step size, and the min-heap is readjusted.
[0043] To select specific tasks from the target resource group, this embodiment maintains a priority task queue within the resource group. Task priority is determined by both stage weight and waiting time. Inelastic tasks have higher stage weights, making them prioritized for execution; waiting time serves as an auxiliary factor, allowing elastic tasks with long waiting times to gradually increase their priority and avoid starvation.
[0044] An adaptive prefetching mechanism is used to reduce the frequency of global scheduling. Worker threads fetch multiple tasks from the target resource group in batches and write them to their local cache. Subsequent tasks are then fetched from the local cache for execution. The prefetch batch size is calculated based on the number of remaining tasks in the target resource group, the number of active worker threads, and the load factor. Under high load, the prefetch batch size is reduced to ensure fairness, while under low load, the prefetch batch size is increased to improve throughput.
[0045] Thus, the two-layer scheduling algorithm maintains long-term fairness between queries through step-size proportional shared scheduling, and prioritizes critical path tasks within queries through a stage-aware strategy, thereby decoupling fairness guarantee from performance optimization.
[0046] (3) Query-level feedback control:
[0047] Query-level feedback control process such as Figure 4 As shown, the feedback control module collects four state variables: stage urgency, execution efficiency, waiting pain level, and completion proximity. Stage urgency represents the proportion of inelastic tasks among the pending tasks; execution efficiency represents the ratio of current throughput to historical throughput; waiting pain level represents the time since the last scheduling; and completion proximity level represents the proportion of completed tasks out of the total tasks. The overall urgency F is used to adjust the effective priority after weighted calculation.
[0048] To avoid priority jitter, F can be smoothed using an exponential moving average, with an urgency cap and a priority change rate limit set. Since the query pass value remains constant, feedback control only affects the subsequent pass value growth rate, without altering the already accumulated CPU time ledger. Therefore, it can maintain long-term fairness while responding to critical phases and abnormal states in the short term.
[0049] In summary, the main technical contributions of this invention are as follows:
[0050] (1) A two-layer scheduling method combining query-level step-size proportional sharing scheduling and task-level stage-aware scheduling is proposed to address fairness and critical path optimization issues at different levels.
[0051] (2) An adaptive priority adjustment method based on multidimensional runtime feedback is proposed, which can dynamically adjust the scheduling parameters according to the query stage, execution efficiency, waiting status and completion progress.
[0052] Practice has shown that the scheduling method designed in this invention can improve the performance of hybrid online analytical processing loads without sacrificing long-term fairness among queries. In an experimental embodiment, the performance was verified on a 16-core server based on the Transaction Processing Performance Committee H benchmark (TPC-H). Compared to the native scheduler, the complete scheduling framework achieves an average performance improvement of 21.0% in single-query scenarios and a 22.3% improvement in non-elastic query performance; in multi-query concurrent scenarios, the average performance improvement reaches 16% to 18%, while the deviation between the CPU time allocation ratio and the query priority ratio is less than 5%.
[0053] The technical solutions of the present invention are not limited to the specific embodiments described above. Any technical modifications made in accordance with the technical solutions of the present invention fall within the protection scope of the present invention.
Claims
1. A phase-aware analytical database query scheduling method, characterized in that, Includes the following steps: S1: Receive at least one analytical database query, divide the execution plan of each query into one or more pipelines, and generate corresponding scheduled tasks based on data blocks; S2: Perform stage identification on the task to be scheduled to determine whether the pipeline to which the task belongs is a flexible stage, an inflexible stage, or an unknown stage. The stage identification process includes two levels: static analysis and dynamic monitoring. S3: Create a resource group for each query. The resource group maintains the query priority, pass value, step size, task queue, and execution status. S4: At the query level, select the target resource group with the smallest passage value from the active resource group based on step-size proportional shared scheduling, and update its step size according to the effective priority of the target resource group; S5: At the task level, tasks are selected from the task queue of the target resource group based on a phase-aware strategy, where the scheduling priority of non-elastic phase tasks is higher than that of elastic phase tasks. S6: Assign the selected tasks to worker threads for execution, and update the query status, pipeline runtime statistics, and subsequent data block granularity based on the task execution results.
2. The phase-aware analytical database query scheduling method according to claim 1, characterized in that: In step S2, the static analysis layer determines the initial stage type based on the operator types contained in the pipeline during the query compilation stage. The pipelines corresponding to the scan operator, filter operator, projection operator, or hash join probe operator are marked as elastic stages, while the pipelines corresponding to the hash join build operator, aggregation merge operator, or sort merge operator are marked as non-elastic stages.
3. The phase-aware analytical database query scheduling method as described in claim 2, characterized in that, The dynamic monitoring mechanism in step S2 includes event-driven detection, which is executed when the task is completed. It determines whether the pipeline has degenerated into an inelastic stage based on whether the task execution time exceeds a preset multiple of the historical average execution time or whether the thread waiting time exceeds a preset threshold.
4. The phase-aware analytical database query scheduling method as described in claim 3, characterized in that, The dynamic monitoring mechanism also includes periodic sampling and detection, calculating pipeline parallel efficiency and throughput changes according to a preset sampling period, and correcting the pipeline stage type when the parallel efficiency is lower than a first threshold or the current throughput is lower than a preset proportion of the historical throughput average.
5. The phase-aware analytical database query scheduling method as described in claim 1, characterized in that, The effective priority in step S4 is calculated based on the stage urgency, execution efficiency, waiting pain, and completion proximity of the query. The stage urgency is the proportion of non-elastic stage tasks in the pending tasks, the execution efficiency is the ratio of the current throughput to the historical average throughput, the waiting pain is the time interval since the last scheduling, and the completion proximity is the ratio of the number of completed tasks to the total number of tasks.
6. The phase-aware analytical database query scheduling method as described in claim 5, characterized in that, The overall urgency level F is calculated by weighted summation of stage urgency, execution efficiency, waiting pain level, and completion proximity. The weights of each indicator are adjusted according to the actual load characteristics, and the adjustment of short-term priority only affects the subsequent passage value growth rate.
7. The phase-aware analytical database query scheduling method as described in claim 1, characterized in that, In step S5, the task-level scheduling maintains a priority task queue within the resource group. The task priority is calculated by both the stage type and the waiting time. Non-elastic stage tasks are given a higher stage weight, and the waiting time serves as an auxiliary factor to gradually increase the priority of elastic tasks that have been waiting for a long time.
8. The phase-aware analytical database query scheduling method as described in claim 7, characterized in that, The task-level scheduling module maintains a local cache for worker threads. Worker threads prefetch multiple tasks from the task queue of the target resource group and write them to the local cache. Subsequently, tasks are retrieved from the local cache for execution.
9. The phase-aware analytical database query scheduling method as described in claim 1, characterized in that, The data block granularity adaptive adjustment mechanism in step S6 dynamically adjusts the size of subsequent data blocks based on the execution time and throughput estimation of completed tasks in the same pipeline. When the average task execution time deviates from the preset target time, the size of subsequent data blocks is adjusted proportionally. When the pipeline enters the tail execution stage, the remaining data is evenly divided into a preset number of tasks.
10. A phase-aware analytical database query scheduling system, characterized in that, include: The query parsing and task generation module is used to receive analytical database queries and generate tasks to be scheduled. The stage identification module is used to identify the stage type of the pipeline to which the task to be scheduled belongs; The resource group management module is used to manage the resource groups and their status for each query. The query-level scheduling module is used to execute step-by-step proportional shared scheduling; the task-level scheduling module is used to execute phase-aware task selection; and the feedback control module is used to calculate the effective priority of queries. The task granularity control module is used to dynamically adjust the data block granularity.