Dynamic fusing and self-healing method of task scheduling system
By using dynamic circuit breaking and self-healing methods, abnormal tasks are automatically intercepted and thread pool resources are adjusted, which solves the problems of rigid resource allocation and insufficient self-healing ability in the task scheduling system, and realizes efficient and adaptive resource management and recovery of the system.
Patent Information
- Application Number
- CN202511057170.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
AI Technical Summary
In existing task scheduling systems, thread pool resource allocation is rigid, utilization is low, and it cannot adapt to load changes. It is prone to resource exhaustion and system crashes due to abnormal tasks, lacks self-healing ability, and relies on manual intervention for recovery.
By employing dynamic circuit breaking and self-healing methods, abnormal tasks are automatically intercepted by monitoring task execution time and concurrency. Thread pool resources are adjusted in conjunction with CPU load and queue depth to achieve elastic scaling and form a closed-loop self-healing system.
Optimize resource utilization, prevent system crashes, reduce manual intervention, improve system continuity and resource utilization, and achieve low-cost adaptive recovery.
Smart Images

Figure CN120950206A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of task scheduling system technology, and in particular to a dynamic circuit breaking and self-healing method for task scheduling systems. Background Technology
[0002] In modern enterprise application systems (such as e-commerce, financial transactions, data analysis, and back-end report generation), the scheduling and execution of scheduled tasks has become an important technical means to support critical business processes. Task scheduling systems (such as Quartz, Elastic Job, and XXL-JOB) generally rely on thread pool technology to manage concurrently executed task instances, serving as the platform for these scheduled tasks. In current mainstream task scheduling system implementations, the key parameters of the thread pool, especially the core thread count and the maximum thread count, are usually statically set once during system startup or configuration and remain fixed during subsequent operation. This static thread pool configuration and management method has revealed several significant shortcomings in practice, as follows:
[0003] (1) Rigid resource allocation and low utilization: During periods of low system load (such as business downturns), the statically configured thread pool cannot automatically shrink the number of core threads. A large number of idle threads continue to occupy system resources (such as memory and CPU scheduling overhead), resulting in the unnecessary consumption of valuable computing resources. It cannot effectively serve other applications or tasks. When the system encounters a sudden high load or the execution time of some tasks unexpectedly increases (such as large data processing or slow response of external services), the static thread pool cannot detect the load change and expand thread resources in time. The sudden influx of task requests will quickly exhaust the available threads, resulting in a large number of tasks piling up in the queue and causing significant execution delays. If the queue is full, tasks may be rejected or even cause the overall performance of the scheduling system to decline or resources to be exhausted and crash (such as CPU or memory overflow), affecting the stability of core business.
[0004] (2) Lack of intelligent circuit breaking mechanism for abnormal tasks, thread resources are easily exhausted: Some scheduled tasks may have unexpected reasons such as their own logical defects (such as infinite loops), timeout or unavailability of external services they depend on, or abnormal surge in the amount of data they process, which may cause the execution time of a single task to far exceed expectations. The static thread pool lacks the ability to monitor and intervene in the execution time of tasks. Such "long tasks" will continue to occupy a thread until completion. Even if the reasonable waiting time is exceeded, when multiple similar long tasks are concurrent, it is very easy to exhaust all available threads, forming a "snowball effect" and blocking the execution of other critical tasks.
[0005] (3) Abnormal recovery relies on manual intervention and lacks self-healing ability: When the system becomes unstable or even partially fails due to the above reasons (thread resource exhaustion, task backlog, abnormal task execution, etc.), the static thread pool lacks an adaptive recovery strategy. It usually requires manual intervention by operation and maintenance personnel to intervene by analyzing logs, investigating the root cause of the problem, manually restarting the service, or adjusting the configuration (such as temporarily increasing the number of threads). This not only results in slow response speed and high management costs, but also prolongs the business interruption time, affecting service availability and user experience.
[0006] In view of the above-mentioned problems, this application proposes a dynamic circuit breaking and self-healing method for task scheduling systems. Summary of the Invention
[0007] Based on the technical problems existing in the background technology, this invention proposes a dynamic circuit breaking and self-healing method for task scheduling systems.
[0008] This invention proposes a dynamic circuit breaker and self-healing method for a task scheduling system, comprising the following steps:
[0009] S1: Perform task circuit breaking and resource protection: The specific steps are as follows:
[0010] S101: Configure monitoring parameters and complete the initialization of the monitoring system;
[0011] S102: Real-time status acquisition;
[0012] S103: Determine the circuit breaker condition and record the abnormal task identifier;
[0013] S104: Perform circuit breaker operation;
[0014] S105: Recovery condition monitoring;
[0015] S106: Task resumes execution;
[0016] S2: Dynamically adjust the thread pool based on system load: The specific steps are as follows:
[0017] S201: Flexible parameter settings;
[0018] S202: System Status Detection: Detects CPU utilization;
[0019] S203: Resource Adjustment Decisions;
[0020] S204: Thread pool change operation;
[0021] S205: Perform status tracking on the CPU after the change.
[0022] Preferably, in step S101, when configuring monitoring parameters, a periodic detection task is set, an alarm value for the number of available threads is configured, and a timeout threshold and a maximum number of concurrent tasks for a single task are determined.
[0023] Preferably, in S102, the information collected includes: (1) scanning the current state of the thread pool and calculating the number of available threads; (2) recording the start timestamp of each task; (3) counting the number of concurrent executions of different task types; wherein the number of available threads = the total number of threads - the number of active tasks.
[0024] Preferably, in S103, the circuit breaker condition is determined as follows: when the execution time of a task exceeds a threshold, or the number of parallel runs of the same task exceeds the upper limit, the circuit breaker operation is prepared to be triggered.
[0025] In step S104, when performing a circuit breaker operation, if the task meets the circuit breaker conditions, a pause command is sent to the task scheduling system so that subsequent similar tasks are no longer scheduled, the circuit-breaker task is added to the isolation list, and the occupied thread resources are released.
[0026] Preferably, in step S105, during recovery condition monitoring, the current concurrent number of each task in the isolation list is checked periodically. When the concurrent number is lower than the safe value for several consecutive periods, task scheduling is prepared to be restored.
[0027] In step S106, when a task resumes execution, if the task meets the recovery conditions, a recovery command is sent to the scheduling system to clear the corresponding record in the isolation list and allow the task of this type to re-enter the scheduling queue.
[0028] Preferably, in step S201, when setting elastic parameters, during the initialization operation in step S101, the minimum core thread count and the maximum thread count range are defined, a thread recycling strategy is configured, and a buffer task queue is established.
[0029] Preferably, in step S202, CPU utilization values are collected periodically, the depth of the task queue to be processed is detected, and the number of currently active threads is recorded. The sampling process is kept lightweight, and CPU utilization is used as an indicator. Additional monitoring indicators are added according to actual needs.
[0030] Preferably, in step S203, adjustment decisions are made by comparing the collected data with a preset threshold: when the CPU usage is too high, a reduction target value is calculated; when the CPU is idle and the queue is overflowing, a expansion target value is calculated; otherwise, the current configuration is maintained.
[0031] In step S204, when the decision value differs significantly from the current setting, the core thread count is adjusted through the management interface to control the magnitude of thread pool changes and avoid drastic fluctuations.
[0032] Preferably, in step S205, after the change, the task waiting time is monitored by observing the changes in the adjusted CPU utilization, providing a reference for subsequent decision-making.
[0033] Compared with existing technologies, the beneficial effects of this invention are:
[0034] 1. Through a dynamic scaling mechanism, thread pool resources can be elastically adjusted according to system load, avoiding resource waste and contention. In addition, by combining multiple indicators such as CPU load, queue backlog, task time / concurrency, etc., the optimization is more accurate and supports seamless integration with existing scheduling frameworks without modifying business logic, thus solving the problems of rigid resource allocation and low utilization.
[0035] 2. Through task circuit breaking and resource protection mechanisms, a dual-threshold circuit breaking algorithm based on task execution time and the number of concurrent runs of the same task automatically intercepts abnormal tasks, preventing single-point failures from spreading and causing system crashes. The self-healing capability reduces the need for manual intervention, improves system continuity, and integrates task-level circuit breaking and system-level resource adjustment into a closed-loop self-healing system. By pausing / resuming tasks, low-cost resource protection is achieved, solving the problems of thread resources being easily exhausted and lacking self-healing capabilities.
[0036] This invention, through the setting of task circuit breaking and resource protection mechanisms, can automatically intercept abnormal tasks based on a dual-threshold circuit breaking algorithm that considers both task execution duration and the number of concurrent runs of the same task. This prevents single-point failures from spreading and causing system crashes. Furthermore, it integrates task-level circuit breaking and system-level resource adjustment into a closed-loop self-healing system. By pausing / resuming tasks, it achieves low-cost resource protection. In addition, through the setting of a dynamic scaling mechanism, it enables thread pool resources to be elastically adjusted according to system load, avoiding resource waste and contention, and solving the problems of rigid resource allocation and low utilization. Attached Figure Description
[0037] Figure 1 This is a flowchart of task circuit breaking and resource protection in a dynamic circuit breaking and self-healing method for a task scheduling system proposed in this invention.
[0038] Figure 2 This is a flowchart illustrating the dynamic thread pool adjustment of system load in a dynamic circuit breaker and self-healing method for a task scheduling system proposed in this invention. Detailed Implementation
[0039] The present invention will be further explained below with reference to specific embodiments.
[0040] Example
[0041] Reference Figure 1-2 This embodiment proposes a dynamic circuit breaker and self-healing method for a task scheduling system, including the following steps:
[0042] S1: Perform task circuit breaking and resource protection: The specific steps are as follows:
[0043] S101: Monitoring parameter configuration, completes the initialization of the monitoring system, sets periodic detection tasks, configures the alarm value for the number of available threads, and determines the timeout judgment threshold and the maximum concurrency of a single task;
[0044] S102: Real-time status collection, the collected information includes: (1) scanning the current status of the thread pool and calculating the number of available threads; (2) recording the start timestamp of each task; (3) counting the number of concurrent executions of different task types; where the number of available threads = the total number of threads - the number of active tasks;
[0045] S103: Determine the circuit breaker condition and record the abnormal task identifier. The circuit breaker condition is: when the execution time of a task exceeds the threshold, or the number of parallel runs of the same task exceeds the upper limit, prepare to trigger the circuit breaker operation.
[0046] S104: Execute the circuit breaker operation. If the task meets the circuit breaker condition, send a pause command to the task scheduling system so that subsequent similar tasks will no longer be scheduled. Add the circuit breaker task to the isolation list and release the occupied thread resources.
[0047] S105: Recovery condition monitoring. By periodically checking the current concurrency of each task in the isolation list, when the concurrency is lower than the safe value for several consecutive periods, prepare to resume task scheduling.
[0048] S106: Task resumes execution. If the task meets the recovery conditions, a recovery command is sent to the scheduling system to clear the corresponding record in the isolation list and allow the task of this type to re-enter the scheduling queue.
[0049] S2: Dynamically adjust the thread pool based on system load: The specific steps are as follows:
[0050] S201: Elastic parameter settings. When setting elastic parameters, during the initialization operation in S101, define the minimum core thread count and maximum thread count range, configure the thread recycling strategy, and establish a buffer task queue.
[0051] S202: System Status Detection: Detects CPU utilization by periodically collecting CPU utilization values, detecting the depth of the task queue, recording the number of currently active threads, keeping the sampling process lightweight, using CPU utilization as the indicator, and adjusting according to actual needs to add additional monitoring indicators.
[0052] S203: Resource adjustment decision. The adjustment decision is made by comparing the collected data with the preset threshold. When the CPU is too high, the target value for shrinking is calculated; when the CPU is idle and the queue is piled up, the target value for expanding is calculated; otherwise, the current configuration is maintained.
[0053] S204: Thread pool change operation. When the decision value differs significantly from the current setting, the core thread count is adjusted through the management interface to control the magnitude of thread pool changes and avoid drastic fluctuations.
[0054] S205: After the change, the CPU status is tracked. After the change, the CPU utilization changes are observed and the task waiting time is monitored to provide a reference for subsequent decision-making.
[0055] This embodiment, through the setting of task circuit breaking and resource protection mechanisms, can automatically intercept abnormal tasks based on a dual-threshold circuit breaking algorithm that considers both task execution duration and the number of concurrent runs of the same task. This prevents single-point failures from spreading and causing system crashes. Furthermore, it integrates task-level circuit breaking and system-level resource adjustment into a closed-loop self-healing system. By pausing / resuming tasks, it achieves low-cost resource protection. In addition, through the setting of a dynamic scaling mechanism, the thread pool resources can be elastically adjusted according to the system load, avoiding resource waste and contention, and solving the problems of rigid resource allocation and low utilization.
[0056] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A dynamic circuit breaker and self-healing method for a task scheduling system, characterized in that, Includes the following steps: S1: Perform task circuit breaking and resource protection: The specific steps are as follows: S101: Configure monitoring parameters and complete the initialization of the monitoring system; S102: Real-time status acquisition; S103: Determine the circuit breaker condition and record the abnormal task identifier; S104: Perform circuit breaker operation; S105: Recovery condition monitoring; S106: Task resumes execution; S2: Dynamically adjust the thread pool based on system load: The specific steps are as follows: S201: Flexible parameter settings; S202: System Status Detection: Detects CPU utilization; S203: Resource Adjustment Decisions; S204: Thread pool change operation; S205: Perform status tracking on the CPU after the change.
2. The dynamic circuit breaker and self-healing method for a task scheduling system according to claim 1, characterized in that, In step S101, when configuring monitoring parameters, periodic detection tasks are set, alarm values for the number of available threads are configured, and timeout thresholds and maximum concurrency for a single task are determined.
3. The dynamic circuit breaker and self-healing method for a task scheduling system according to claim 1, characterized in that, In S102, the collected information includes: (1) scanning the current state of the thread pool and calculating the number of available threads; (2) recording the start timestamp of each task; (3) counting the number of concurrent executions of different task types; where the number of available threads = the total number of threads - the number of active tasks.
4. The dynamic circuit breaker and self-healing method for a task scheduling system according to claim 1, characterized in that, In S103, the circuit breaker condition is determined as follows: when the execution time of a task exceeds the threshold, or the number of parallel runs of the same task exceeds the upper limit, the circuit breaker operation is prepared to be triggered. In step S104, when performing a circuit breaker operation, if the task meets the circuit breaker conditions, a pause command is sent to the task scheduling system so that subsequent similar tasks are no longer scheduled, the circuit-breaker task is added to the isolation list, and the occupied thread resources are released.
5. The dynamic circuit breaker and self-healing method for a task scheduling system according to claim 1, characterized in that, In S105, during recovery condition monitoring, the current concurrency of each task in the isolation list is checked periodically. When the concurrency is lower than the safe value for several consecutive periods, task scheduling is prepared to be restored. In step S106, when a task resumes execution, if the task meets the recovery conditions, a recovery command is sent to the scheduling system to clear the corresponding record in the isolation list and allow the task of this type to re-enter the scheduling queue.
6. The dynamic circuit breaker and self-healing method for a task scheduling system according to claim 1, characterized in that, In S201, when setting elastic parameters, during the initialization operation in S101, the minimum core thread count and the maximum thread count range are defined, a thread recycling strategy is configured, and a buffer task queue is established.
7. The dynamic circuit breaker and self-healing method for a task scheduling system according to claim 1, characterized in that, In S202, CPU utilization values are collected periodically, the depth of the task queue to be processed is detected, and the number of currently active threads is recorded. The sampling process is kept lightweight, and CPU utilization is used as an indicator. Adjustments are made according to actual needs, and additional monitoring indicators are added.
8. The dynamic circuit breaker and self-healing method for a task scheduling system according to claim 1, characterized in that, In step S203, an adjustment decision is made by comparing the collected data with a preset threshold. When the CPU usage is too high, a reduction target value is calculated; when the CPU is idle and the queue is piling up, a expansion target value is calculated; otherwise, the current configuration is maintained. In step S204, when the decision value differs significantly from the current setting, the core thread count is adjusted through the management interface to control the magnitude of thread pool changes and avoid drastic fluctuations.
9. The dynamic circuit breaker and self-healing method for a task scheduling system according to claim 1, characterized in that, In step S205, after the change, the CPU utilization change is observed and the task waiting time is monitored to provide a reference for subsequent decision-making.