Distributed timed task scheduling method realized based on DelayQueue
Through the distributed timed task scheduling method based on DelayQueue, the problems of unbalanced task distribution and improper conflict handling in distributed systems are solved, and efficient, reliable and flexible task scheduling is achieved, which is suitable for small and medium-sized distributed scenarios.
Patent Information
- Application Number
- CN202510788722.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing distributed timed task scheduling technology has problems such as load imbalance, low resource utilization, improper task conflict handling, poor system reliability and insufficient scalability, making it difficult to meet the needs of high concurrency and large-scale scenarios.
A distributed timed task scheduling method based on DelayQueue is adopted. Through the instance heartbeat mechanism, MurmurHash algorithm and random initial delay, combined with JavaSPI mechanism and database coordination, dynamic balanced allocation of tasks, conflict avoidance and rollback mechanism are realized to ensure task status consistency and system reliability.
It achieves efficient and balanced distribution of tasks among multiple nodes, avoids repeated execution and loss of tasks, improves the system's fault tolerance and scalability, reduces system complexity and operation and maintenance costs, and supports concurrent processing of tens of thousands of tasks.
Smart Images

Figure CN120723399A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Java Web development, and in particular to a distributed timed task scheduling method implemented based on DelayQueue. Background Art
[0002] With the widespread use of distributed systems, scheduled task scheduling, as a core function to ensure the automated and intelligent operation of the system, faces many challenges. Existing technologies have obvious shortcomings, as shown below:
[0003] Limitations of traditional centralized scheduling: Early scheduled task scheduling mostly adopted a centralized architecture, relying on a single dispatch center to allocate tasks. In this model, the dispatch center became a system bottleneck. A failure would paralyze the entire dispatch system, resulting in poor reliability. Furthermore, as the scale of tasks and the number of nodes increased, the load on the centralized dispatch center increased dramatically, severely limiting scalability and making it difficult to meet the needs of high-concurrency, large-scale scenarios.
[0004] Common problems with distributed scheduling: To address the drawbacks of centralized scheduling, distributed timed task scheduling technology has emerged, but existing solutions still have flaws. Some methods use fixed allocation strategies and cannot dynamically adjust task allocation based on node load, which can easily lead to load imbalance between nodes and low resource utilization. Some solutions rely on complex distributed coordination services (such as Zookeeper) to implement task allocation and node management, increasing the complexity of the system architecture and operation and maintenance costs. In addition, in terms of task conflict handling and fault tolerance mechanisms, existing technologies often lack effective solutions. Conflicts are prone to occur when multiple nodes load tasks concurrently. When a node fails or shuts down abnormally, it is difficult to ensure consistency in task status, and there is a risk of task loss or duplication.
[0005] Deficiencies in task execution and system management: At the task execution level, traditional scheduling methods lack a flexible task processing expansion mechanism, making it difficult to quickly adapt to diverse business needs; task execution priorities and retry strategies are single, unable to meet the requirements for task execution efficiency and success rate in complex business scenarios; in terms of system management, when the scheduler is shut down, existing solutions often fail to properly handle unfinished tasks, resulting in data inconsistency or task loss, affecting system stability and reliability.
[0006] Therefore, to address the above problems, a distributed timed task scheduling method based on DelayQueue is proposed. Summary of the Invention
[0007] The purpose of the present invention is to provide a distributed timed task scheduling method based on DelayQueue to solve the problems raised in the above background technology.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] The distributed timed task scheduling method based on DelayQueue includes the following steps:
[0010] S1. Preload scheduled tasks into DelayQueue:
[0011] Load the scheduled tasks from the persistent storage through the user-defined task loader;
[0012] Use user-defined task updater to update task status and achieve multi-node task synchronization;
[0013] S2. Balanced distribution of distributed tasks:
[0014] Assign a unique identifier (UID) to each instance and periodically report heartbeats to the database to maintain instance activity.
[0015] Periodically pull the list of active instances, sort the list, and obtain the serial number B of the current instance;
[0016] Divide the task trigger time into a fixed time window T, and calculate the hash value F based on the task unique identifier and the time window;
[0017] The task belongs to the instance according to the formula F%N==B, and only the current instance adds the task to the DelayQueue, where N is the number of active instances;
[0018] S3. Task conflict avoidance and rollback mechanism:
[0019] Generate a random initial delay when the instance starts to avoid conflicts in the first task loading;
[0020] When it is detected that the task has been loaded or the instance is inactivated, the task is rolled back to the initial state;
[0021] S4. Task triggering and processing:
[0022] Pull out the expired tasks from the DelayQueue through an independent thread loop, and call the user-extended JobProcessor interface implementation class according to the task type;
[0023] Dynamically load processors of different task types based on the JavaSPI mechanism and execute tasks asynchronously according to priority;
[0024] S5. Scheduler safe shutdown:
[0025] Stop task preloading, close the thread pool, and roll back the unprocessed task status to the initial state.
[0026] As a preferred solution, the task loader and task updater in step S1 are implemented in the following manner:
[0027] The task loader filters the tasks whose status is "initialized" and whose trigger time is within the preset window from the database table;
[0028] The task updater updates the task status to "Loaded" after the task is successfully loaded, and updates it to "Success" or "Failure" after the task is executed.
[0029] As a preferred solution, the heartbeat reporting and active instance pulling cycle in step S2 is 60 seconds, and:
[0030] Heartbeat reporting is achieved by updating the last active time of the instance in the database;
[0031] The active instance list is generated by filtering the database for instances whose heartbeat times are within the last 60 seconds.
[0032] As a preferred solution, the hash algorithm in step S2 is the MurmurHash algorithm, and the time window T is one third of the scheduling period.
[0033] As a preferred solution, the random initial delay in step S3 ranges from 10 seconds to one third of the scheduling period T.
[0034] As a preferred solution, the task rollback in step S3 is implemented in the following way:
[0035] When an instance becomes inactive, other instances detect that its UID is no longer in the active list and reset the status of the tasks originally assigned to the instance to "initializing";
[0036] When the instance is shut down normally, the status of untriggered tasks in memory rolls back to "Initializing".
[0037] As a preferred solution, the JobProcessor interface in step S4 is loaded through the Java SPI mechanism, and:
[0038] Each task type corresponds to at least one JobProcessor implementation class;
[0039] When a task is triggered, it is called in the order of processor priority until the task is successfully processed or the maximum number of retries is reached.
[0040] As a preferred solution, the sorting algorithm Sort in step S2 is a lexicographic order or a numerical order, which is used to determine the position of the instance in the active list.
[0041] As a preferred solution, the safety shutdown in step S5 further includes:
[0042] Stop all preloading threads and task execution thread pools;
[0043] Persist unprocessed tasks in the DelayQueue to the database and mark them as "initialized".
[0044] The method relies on a database to achieve decentralized coordination, and the time complexity of the task allocation algorithm is O(1).
[0045] It can be seen from the technical solution provided by the present invention that the distributed timed task scheduling method based on DelayQueue provided by the present invention has the following beneficial effects:
[0046] Efficient and balanced task allocation: Based on the MurmurHash algorithm and the task allocation strategy of fixed time window division, combined with the instance heartbeat mechanism, dynamic and balanced task allocation among multiple nodes is achieved. The task allocation algorithm has a time complexity of O(1) and can quickly and evenly distribute tasks to each node, avoiding excessive load at a single point, significantly improving the overall processing capacity and resource utilization of the system.
[0047] Strong reliability and fault tolerance: Through random initial delays, task conflict detection, and rollback mechanisms, conflicts when concurrently loading tasks on multiple nodes are effectively avoided, ensuring task state consistency. When an instance becomes inactive or shuts down normally, tasks are automatically rolled back to their initial state, ensuring that tasks are not lost or repeated. Furthermore, a heartbeat mechanism monitors node status in real time, enabling automatic migration and reallocation of tasks on failed nodes, significantly improving the system's fault tolerance and availability.
[0048] Flexible and scalable architecture design: Using user-defined task loaders and updaters, as well as a JavaSPI-based task processor loading method, the system decouples task logic from the scheduling framework. Users can quickly add new task types and processing logic based on business needs without modifying the core code, greatly enhancing the system's scalability and adaptability, making it easy to handle diverse business scenarios.
[0049] Efficient and stable task execution: Independent threads trigger tasks in a loop, and combined with priority scheduling and asynchronous execution strategies, fully utilize multi-core resources and support the concurrent processing of tens of thousands of tasks. Through dynamic adjustment of learning rates, early stopping mechanisms, regularization techniques and other optimization methods, we ensure efficient and stable task processing, effectively prevent overfitting during task execution, and improve the success rate of task execution.
[0050] Safe and stable system shutdown: The scheduler's safe shutdown process ensures consistent task status during shutdown, preventing data loss or duplicate execution, and achieving a smooth shutdown by stopping task preloading, gracefully shutting down the thread pool, and rolling back unprocessed tasks. Furthermore, compensation mechanisms and timeout control in abnormal scenarios further ensure the safety and reliability of the system shutdown process.
[0051] Lightweight and low-cost: This invention relies on the database to achieve decentralized coordination, without the need to introduce complex distributed coordination services such as Zookeeper, reducing the complexity of the system architecture and operation and maintenance costs. It is especially suitable for small and medium-sized distributed scenarios. While ensuring high performance, it effectively controls the system construction and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 The figure is a flow chart of the distributed timed task scheduling method implemented based on DelayQueue in the present invention. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0054] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0055] like Figure 1 As shown, the embodiment of the present invention provides a method for distributed timed task scheduling based on DelayQueue, including the following steps:
[0056] S1. Preload scheduled tasks into DelayQueue:
[0057] Load the scheduled tasks from the persistent storage through the user-defined task loader;
[0058] Use user-defined task updater to update task status and achieve multi-node task synchronization;
[0059] S2. Balanced distribution of distributed tasks:
[0060] Assign a unique identifier (UID) to each instance and periodically report heartbeats to the database to maintain instance activity.
[0061] Periodically pull the list of active instances, sort the list, and obtain the serial number B of the current instance;
[0062] Divide the task trigger time into a fixed time window T, and calculate the hash value F based on the task unique identifier and the time window;
[0063] The task belongs to the instance according to the formula F%N==B, and only the current instance adds the task to the DelayQueue, where N is the number of active instances;
[0064] S3. Task conflict avoidance and rollback mechanism:
[0065] Generate a random initial delay when the instance starts to avoid conflicts in the first task loading;
[0066] When it is detected that the task has been loaded or the instance is inactivated, the task is rolled back to the initial state;
[0067] S4. Task triggering and processing:
[0068] Pull out the expired tasks from the DelayQueue through an independent thread loop, and call the user-extended JobProcessor interface implementation class according to the task type;
[0069] Dynamically load processors of different task types based on the JavaSPI mechanism and execute tasks asynchronously according to priority;
[0070] S5. Scheduler safe shutdown:
[0071] Stop task preloading, close the thread pool, and roll back the unprocessed task status to the initial state.
[0072] In this embodiment, the task loader and task updater in step S1 are implemented in the following manner:
[0073] The task loader filters the tasks whose status is "initialized" and whose trigger time is within the preset window from the database table;
[0074] The task updater updates the task status to "loaded" after the task is successfully loaded, and updates it to "success" or "failure" after the task is completed;
[0075] Furthermore, the purpose of preloading the scheduled task into DelayQueue in step S1 is to achieve efficient loading and state synchronization of the scheduled task through custom components, ensuring the consistency and integrity of the task in a multi-node environment. The specific steps are as follows:
[0076] Step S1-1: Task loader design and implementation:
[0077] A user-defined task loader is used to batch load tasks to be scheduled from persistent storage. The core logic includes:
[0078] Task filtering conditions: Filter tasks from the database table that are in the initialization state and whose trigger time is within the preset window. The preset window is usually set to 1 to 24 hours in the future (configurable) to avoid memory pressure caused by loading too many tasks at once.
[0079] Batch loading strategy: Use paging queries (for example, loading 1,000 records at a time) and combine database indexes to optimize query efficiency and ensure stable performance in large-scale task scenarios.
[0080] Idempotence: During the loading process, a unique database index or distributed lock mechanism is used to ensure that the same task is not loaded repeatedly by multiple nodes.
[0081] Step S1-2: Task updater implements multi-node synchronization:
[0082] The task status is maintained in real time through the user-defined task updater, achieving task synchronization between multiple nodes:
[0083] State transfer mechanism:
[0084] The state before the task is loaded is initialized, indicating that it is waiting to be scheduled;
[0085] After the task is successfully loaded into the DelayQueue, it is updated to loaded to prevent other nodes from loading it repeatedly;
[0086] After the task is completed, it is updated to success or failure according to the execution result, and the execution time, result information, etc. are recorded;
[0087] Optimistic locking implementation: A version number mechanism is used when updating task status to ensure consistency when multiple nodes are updated concurrently;
[0088] If the update fails (the number of affected rows is 0), it means that the task has been preempted by other nodes and the current node skips the task;
[0089] Step S1-3: Incremental loading and full refresh:
[0090] To balance performance and data consistency, we use the incremental loading + periodic full refresh strategy:
[0091] Incremental loading: Dynamically adjust the query range based on the time window, scanning new tasks that enter the preset window once every minute (for example, the trigger time is within the next hour);
[0092] Full refresh: A full scan is performed every hour to reload all tasks that have not expired and are in the initialization state, ensuring that unloaded tasks caused by network fluctuations or node failures are discovered in a timely manner;
[0093] Step S1-4: Task loading monitoring and exception handling:
[0094] Performance monitoring: records indicators such as task loading time, throughput, and failure rate, and triggers an alarm when the loading time exceeds a threshold (such as 5 seconds) or the failure rate remains above 5%;
[0095] Failure retry mechanism: If a task fails to load (e.g. database connection timeout), it will automatically retry three times, with an exponential backoff interval (1 second, 2 seconds, 4 seconds) each time. If it still fails, an error log will be recorded and the task will be skipped.
[0096] Memory protection: Dynamically monitor JVM memory usage during loading, and pause loading when the remaining memory is less than 20% to prevent OOM exceptions;
[0097] Step S1-5: Distributed task state consistency guarantee:
[0098] Eventual consistency: Database transactions and unique indexes ensure the eventual consistency of task status. This allows for inconsistent views between different nodes for a short period of time, but ultimately all nodes reach a consistent state.
[0099] Conflict detection: When multiple nodes attempt to load the same task simultaneously, conflicts are detected through unique constraints or distributed locks in the database to ensure that only one node successfully loads the task.
[0100] Summary of key technical points:
[0101] Efficient loading strategy: By presetting time windows and batch queries, it balances memory usage and loading efficiency, supporting millions of tasks.
[0102] State synchronization mechanism: Using database transactions and optimistic locking to ensure the consistency and atomicity of task states in a multi-node environment;
[0103] Elastic fault-tolerant design: Incremental loading combined with periodic full refresh, along with failure retry and exception handling, improves system reliability.
[0104] Performance assurance: By monitoring indicators and adaptive current limiting, we prevent resource exhaustion and ensure system stability under high concurrency.
[0105] In this embodiment, the heartbeat reporting and active instance pulling period in step S2 is 60 seconds, and:
[0106] Heartbeat reporting is achieved by updating the last active time of the instance in the database;
[0107] The list of active instances is generated by filtering the database for instances whose heartbeat times are within the last 60 seconds;
[0108] The hash algorithm in step S2 is the MurmurHash algorithm, and the time window T is one-third of the scheduling period;
[0109] Furthermore, the role of distributed task balancing in step S2 is to achieve dynamic and balanced distribution of scheduled tasks among multiple nodes through instance identification, heartbeat mechanism and hash distribution algorithm, ensuring maximum resource utilization and no repeated execution of tasks; the specific steps are as follows:
[0110] Step S2-1: Instance identification and heartbeat mechanism:
[0111] Unique Instance Identifier (UID): Each task scheduling instance generates a globally unique UID when it is started. This is used to distinguish different node identities. The UID can be generated using a UUID or a database-based auto-increment ID to ensure no duplication in a distributed environment.
[0112] Heartbeat reporting mechanism:
[0113] Cycle and Implementation: The instance updates its last active time field in the database every 60 seconds to simulate a heartbeat signal. For example, this can be achieved using the SQL statement UPDATE instance_table SET last_active_time = NOW() WHERE uid = ?.
[0114] Failure determination: The database regularly screens instances that have not had heartbeat updates in the last 60 seconds, marks them as inactive, and removes them from the list of active instances to prevent failed nodes from continuing to participate in task allocation;
[0115] Step S2-2: Active instance list maintenance and sorting:
[0116] Dynamic list generation: Each instance periodically (60 seconds) pulls the UID list of all active instances from the database, with the filtering condition being last_active_time >= NOW() - INTERVAL 60SECOND;
[0117] Instance sorting rules: Sort the active instance list in lexicographic or numerical order (such as by UID string or auto-increment ID size) to ensure that all nodes agree on the instance order, providing a unified benchmark for subsequent task allocation;
[0118] Sequence number acquisition: Each instance determines its own sequence number B (index starts at 0) in the sorted list. This sequence number is used to calculate task ownership.
[0119] Step S2-3: Task allocation time window division:
[0120] Fixed time window (T) setting: Divide the task trigger time into a fixed-length time window T, where T is one-third of the scheduling period. For example, if the scheduling period is 30 minutes, then T is 10 minutes. Tasks within the same time window are considered to be of the same type.
[0121] Window mapping logic: Tasks are assigned to corresponding time windows based on their trigger time. For example, tasks with a trigger time between 0 and 10 minutes are assigned to window 1, tasks with a trigger time between 10 and 20 minutes are assigned to window 2, and so on.
[0122] Step S2-4: Task allocation based on hash algorithm:
[0123] Hash value calculation: The MurmurHash algorithm is used to combine the task's unique identifier (such as the task ID) and the time window number to generate the hash value F. This algorithm has high efficiency and low collision characteristics, making it suitable for large-scale data scenarios.
[0124] Instance Determination Formula: The instance to which a task belongs is determined by the formula F%N==B (where N is the number of active instances, B is the sequence number of the current instance, and F is the hash value). If the equation holds, the current instance is the responsible node for the task and the task is added to the local DelayQueue. Otherwise, the task is skipped.
[0125] Load balancing effect: This algorithm ensures that tasks are evenly distributed across active nodes. In theory, the deviation in the amount of tasks undertaken by each node does not exceed ±1 / N, achieving dynamic load balancing.
[0126] Step S2-5: Dynamic adjustment and fault tolerance mechanism:
[0127] Instance addition and subtraction response: When a new instance is added or an old instance fails, the active instance list is updated, each node recalculates its own sequence number B, and redistributes tasks based on the latest list to ensure the dynamic adaptability of the allocation strategy;
[0128] Task migration: If an instance fails, other nodes will automatically recalculate the unexecuted tasks (with the status of "initialization") originally assigned to the instance after the next round of heartbeat detection, triggering secondary allocation to avoid task loss;
[0129] Summary of key technical points:
[0130] Lightweight coordination mechanism: It relies only on the database heartbeat table and active instance list, eliminating the need for complex distributed coordination services (such as Zookeeper), reducing system complexity.
[0131] Efficient allocation algorithm: MurmurHash combined with modulo operation has a time complexity of O(1), supporting the rapid allocation of tens of thousands of tasks per second.
[0132] Dynamic load balancing: Through the heartbeat mechanism and real-time list updates, tasks are automatically redistributed when nodes fail or capacity is expanded, ensuring resource utilization.
[0133] Task uniqueness: Based on the determinism of hash distribution, it ensures that the same task is processed by only one instance in a multi-node environment, avoiding repeated execution.
[0134] In this embodiment, the random initial delay in step S3 ranges from 10 seconds to one-third of the scheduling period T;
[0135] The task rollback in step S3 is achieved by:
[0136] When an instance becomes inactive, other instances detect that its UID is no longer in the active list and reset the status of the tasks originally assigned to the instance to "initializing";
[0137] When the instance is shut down normally, the status of untriggered tasks in memory rolls back to "initialization";
[0138] Furthermore, the task conflict avoidance and rollback mechanism in step S3 is used to resolve conflicts when multiple nodes are concurrently loading tasks in a distributed environment through randomized startup strategies and state rollback mechanisms, ensuring the consistency and reliability of task states. The specific steps are as follows:
[0139] Step S3-1: Random initial delay mechanism:
[0140] Delay range setting: Each instance generates a random initial delay at startup, ranging from 10 seconds to one-third of the scheduling period (T). For example, if the scheduling period T is 30 minutes, the random delay range is 10 seconds to 10 minutes.
[0141] Implementation: Use the programming language's random number generator (such as Java's Thread.sleep(Random.nextInt())) to ensure that the startup times of each instance are staggered and the time points for the first task loading are dispersed.
[0142] Purpose: Prevents all instances from initiating task loading requests at the same time, which can cause a sudden increase in database pressure or repeated task loading.
[0143] Step S3-2: Task loading conflict detection and processing:
[0144] Status check: Before loading a task, check whether the task status is still initialized through database query; if the status has changed to loaded, it means that the task has been preempted by other nodes, and the current node skips the task;
[0145] Atomicity guarantee: Use database row locks or optimistic locks (version number mechanism) to ensure the atomicity of task status updates;
[0146] If the update fails (the number of affected rows is 0), the loading task is abandoned;
[0147] Step S3-3: Task rollback when instance becomes inactive:
[0148] Detection mechanism: Each instance periodically (e.g., every 60 seconds) pulls a list of active instances and compares it with locally cached active nodes. If an instance's UID is not in the latest list, it is deemed inactive.
[0149] Task reset: Roll back the status of all tasks originally assigned to the inactive instance from loaded to initialized. This is achieved by:
[0150] Query the database for all tasks whose status is loaded and whose instance has an inactive UID.
[0151] Update the status of these tasks to initialized in batches for reallocation;
[0152] Idempotent design: Rollback operations can be executed multiple times to ensure that the system does not produce side effects during network fluctuations or retries.
[0153] Step S3-4: Task rollback during normal shutdown:
[0154] Graceful shutdown trigger: When the instance receives a shutdown signal (such as SIGTERM) or actively calls the shutdown interface, the task rollback process is triggered;
[0155] Memory task processing: traverse all untriggered tasks in the local DelayQueue, update their status to initialized in batches, and persist them to the database;
[0156] Thread pool management: Stop the task loading thread pool to ensure that no new tasks are received during the rollback period; wait for the executed tasks to complete and then close the execution thread pool to avoid task loss;
[0157] Step S3-5: Transaction guarantee for rollback operation:
[0158] Database transactions: Task status updates and rollback operations must be performed within database transactions to ensure data consistency.
[0159] Distributed transactions: If multiple databases or service calls are involved, use TCC (Try-Confirm-Cancel) or Saga mode to ensure eventual consistency;
[0160] Summary of key technical points:
[0161] Conflict prevention: Random initial delays distribute concurrent loading pressure to different time periods, reducing database peak pressure and conflict probability.
[0162] State consistency: Through atomic updates and transaction guarantees, the eventual consistency of task states across multiple nodes is ensured;
[0163] Fault-tolerance recovery: Instance inactivity detection and task rollback mechanisms ensure that tasks on failed nodes can be reassigned for execution, improving system availability.
[0164] Graceful shutdown: Task rollback design during normal shutdown avoids task loss and supports smooth system upgrades or expansions.
[0165] In this embodiment, the JobProcessor interface in step S4 is loaded through the Java SPI mechanism, and:
[0166] Each task type corresponds to at least one JobProcessor implementation class;
[0167] When a task is triggered, it is called in the order of processor priority until the task is successfully processed or the maximum number of retries is reached;
[0168] Furthermore, the purpose of task triggering and processing in step S4 is to achieve efficient triggering and asynchronous execution of scheduled tasks through independent threads, JavaSPI mechanism and priority scheduling strategy, while ensuring the flexibility and scalability of task processing; the specific steps are as follows:
[0169] Step S4-1: Independent thread loop triggering task:
[0170] Thread pool management: Create an independent task trigger thread pool and configure a fixed number of threads (for example, the number of core threads is twice the number of CPU cores) to ensure resource utilization and system stability;
[0171] Circular pull mechanism: Each thread continuously pulls expired tasks from the DelayQueue in a blocking manner; when the DelayQueue is empty or there are no expired tasks, the thread automatically blocks to avoid idling and consuming resources; when the task expires, the thread is awakened and performs subsequent processing;
[0172] Timeout processing: Set the timeout for pulling tasks (such as 5 seconds). If the task is not obtained within the timeout, recheck the thread pool status to prevent the thread from being blocked indefinitely.
[0173] Step S4-2: Processor loading based on JavaSPI mechanism:
[0174] Interface definition: defines the JobProcessor interface, including the abstract method execute(Task task), which is used to process different types of tasks. Users need to implement this interface to extend custom task logic.
[0175] Dynamic loading process:
[0176] Create a JobProcessor file in the project's META-INF / services directory. The file content is the fully qualified name of the implementation class (such as com.example.MyJobProcessor).
[0177] Use the ServiceLoader.load(JobProcessor.class) method. The Java runtime automatically scans and loads all implementation classes to build a pool of processor instances.
[0178] Multiple implementation support: The same task type can correspond to multiple JobProcessor implementation classes (for example, different strategies based on business scenarios), and the execution order is determined by the subsequent priority mechanism;
[0179] Step S4-3: Task type identification and processor matching:
[0180] Type identifier: Each task contains a task type field (such as a string identifier type="order_expire"), which is used to distinguish different business logics;
[0181] Dynamic matching: Filter matching processors from the loaded JobProcessor implementation classes based on the task type; for example, by traversing the implementation class list and calling the supports(StringtaskType) method to determine whether the current task type is supported;
[0182] Fallback strategy: If no matching handler is found, an exception is thrown or the default processing logic is executed (such as logging an error and skipping the task).
[0183] Step S4-4: Priority scheduling and asynchronous execution:
[0184] Priority definition: Assign a priority to each JobProcessor implementation class (e.g. integer type, the smaller the value, the higher the priority), specified through annotations or configuration files;
[0185] Execution strategy:
[0186] Sort the list of matching processors in ascending order of priority;
[0187] Call the execute method of the processor in sequence. If a processor successfully processes the task (such as returning a successful processing flag), the subsequent processors are skipped.
[0188] If all processors fail to execute, retries are performed according to the task configuration (e.g., the maximum number of retries is set to 3, with exponential backoff between each attempt);
[0189] Asynchronous processing: Submit task execution to the task execution thread pool and process it in asynchronous mode to avoid blocking the task triggering thread and improve system throughput;
[0190] Step S4-5: Record and monitor execution results:
[0191] Status update: After the task is completed, the task status is updated to success or failure based on the execution result, and metadata such as execution time and error information (if failed) are recorded;
[0192] Monitoring indicators: Statistics on indicators such as task execution success rate, average execution time, and failure rate. When the failure rate exceeds a threshold (such as 5%), an alarm is triggered to support system operation and maintenance analysis and optimization;
[0193] Logging: Output detailed task execution logs, including input parameters, processing process, output results, etc., to facilitate troubleshooting and auditing;
[0194] Summary of key technical points:
[0195] Efficient triggering mechanism: Independent threads are combined with DelayQueue to trigger tasks as they expire, eliminating polling overhead and reducing response latency to milliseconds.
[0196] Scalability: The JavaSPI mechanism decouples task logic and scheduling framework, allowing users to quickly expand new task types without modifying the core code;
[0197] Intelligent scheduling: Priority and retry strategies are combined to ensure that critical tasks are executed first, and the fault tolerance mechanism improves the success rate of task processing;
[0198] Asynchronous high concurrency: The thread pool asynchronous execution design fully utilizes multi-core resources and supports concurrent processing of tens of thousands of tasks.
[0199] In this embodiment, the safety shutdown in step S5 further includes:
[0200] Stop all preloading threads and task execution thread pools;
[0201] Persist unprocessed tasks in the DelayQueue to the database and mark them as "initialized".
[0202] The method relies on a database to achieve decentralized coordination, and the time complexity of the task allocation algorithm is O(1);
[0203] Furthermore, the role of the scheduler safety shutdown in step S5 is to ensure that the system does not lose tasks or generate repeated executions during the shutdown process, and to restore unfinished tasks to their initial state to achieve a smooth shutdown. The specific steps are as follows:
[0204] Step S5-1: Stop task preloading:
[0205] Loader shutdown: Send a shutdown signal to the task loader to stop pulling new tasks from the database; set the volatile flag (such as isRunning = false) to notify the loading thread to terminate the loop query logic;
[0206] Processing of remaining tasks: Wait for the currently loading batch tasks to complete to ensure that the tasks that have started loading are not interrupted; for example, use CountDownLatch to block the main thread until all loading threads confirm completion;
[0207] Step S5-2: Close the task execution thread pool:
[0208] Smooth Closure Strategy:
[0209] Call the shutdown() method of the thread pool to refuse to accept new tasks and allow submitted tasks to continue executing;
[0210] Set the waiting timeout (such as 30 seconds) and call awaitTermination() to wait for the unfinished task to be completed;
[0211] If there are still unfinished tasks after the timeout, call shutdownNow() to force a shutdown and record the list of unfinished tasks;
[0212] Interruption handling: For tasks that are forcibly interrupted, catch InterruptedException and mark the task status as pending for recovery so that it can be rolled back later;
[0213] Step S5-3: Rollback of unprocessed tasks in DelayQueue:
[0214] Task persistence: Traverse all untriggered tasks in the DelayQueue, update their status to initialized in batches, and persist them to the database; ensure the atomicity of status updates and database writes through transactions:
[0215] BEGINTRANSACTION;
[0216] UPDATE task_table
[0217] SET status = 'Initialize', next_trigger_time = ?
[0218] WHERE id IN(SELECT id FROM delay_queue_snapshot);
[0219] COMMIT;
[0220] Memory cleanup: Clear the DelayQueue memory queue to prevent repeated execution of rolled-back tasks after restart;
[0221] Step S5-4: Instance status marking and resource release:
[0222] Instance offline: Update the status of the current instance in the database to offline and delete the heartbeat record to ensure that other nodes no longer assign tasks to it;
[0223] Resource release: close external resources such as database connections and network clients to release system resources;
[0224] Graceful exit: After all cleanup operations are completed, the JVM process exits (e.g., returns status code 0) and notifies the container or operation and maintenance system of the successful shutdown.
[0225] Step S5-5: Abnormal scenario processing:
[0226] Sudden failure: If an exception occurs during the shutdown process (such as a database connection interruption), enable the compensation mechanism:
[0227] Record the current operation progress (such as the list of rolled-back task IDs);
[0228] At the next startup, the unfinished rollback operation will be continued according to the recorded progress;
[0229] Notify administrators through the alarm system for manual intervention;
[0230] Timeout control: Set a timeout threshold for each shutdown step (e.g., step S5-2 has a timeout of 30 seconds) to avoid long-term blocking that may cause the system to become unresponsive.
[0231] Summary of key technical points:
[0232] Data consistency: Transactions ensure the atomicity of task status rollback and persistence, preventing data loss or inconsistency.
[0233] Smooth transition: The thread pool is shut down smoothly to avoid direct termination that may cause tasks to be executed halfway, ensuring business continuity.
[0234] Recoverability: Recording intermediate states and compensation mechanisms, supporting breakpoint resumption in abnormal scenarios, and improving system fault tolerance;
[0235] Resource protection: Strict resource release sequence and timeout control prevent system resource leakage and ensure safe exit.
[0236] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A distributed timed task scheduling method based on DelayQueue, characterized by: The following steps are involved: S1. Preload scheduled tasks into DelayQueue: Load the scheduled tasks from the persistent storage through the user-defined task loader; Use user-defined task updater to update task status and achieve multi-node task synchronization; S2. Balanced distribution of distributed tasks: Assign a unique identifier (UID) to each instance and periodically report heartbeats to the database to maintain instance activity. Periodically pull the list of active instances, sort the list, and obtain the serial number B of the current instance; Divide the task trigger time into a fixed time window T, and calculate the hash value F based on the task unique identifier and the time window; The task belongs to the instance according to the formula F%N==B, and only the current instance adds the task to the DelayQueue, where N is the number of active instances; S3. Task conflict avoidance and rollback mechanism: Generate a random initial delay when the instance starts to avoid conflicts in the first task loading; When it is detected that the task has been loaded or the instance is inactivated, the task is rolled back to the initial state; S4. Task triggering and processing: Pull out the expired tasks from the DelayQueue through an independent thread loop, and call the user-extended JobProcessor interface implementation class according to the task type; Dynamically load processors of different task types based on the JavaSPI mechanism and execute tasks asynchronously according to priority; S5. Scheduler safe shutdown: Stop task preloading, close the thread pool, and roll back the unprocessed task status to the initial state.
2. The distributed timed task scheduling method based on DelayQueue according to claim 1 is characterized in that: The task loader and task updater in step S1 are implemented in the following manner: The task loader filters the tasks whose status is "initialization" and whose trigger time is within the preset window from the database table; The task updater updates the task status to "loaded" after the task is successfully loaded, and updates it to "success" or "failure" after the task is executed.
3. The distributed timed task scheduling method based on DelayQueue according to claim 1 is characterized in that: The heartbeat reporting and active instance pulling cycle in step S2 is 60 seconds, and: Heartbeat reporting is achieved by updating the last active time of the instance in the database; The active instance list is generated by filtering the database for instances whose heartbeat times are within the last 60 seconds.
4. The distributed timed task scheduling method based on DelayQueue according to claim 1 is characterized in that: The hash algorithm in step S2 is the MurmurHash algorithm, and the time window T is one third of the scheduling period.
5. The distributed timed task scheduling method based on DelayQueue according to claim 1 is characterized in that: The random initial delay in step S3 ranges from 10 seconds to one third of the scheduling period T.
6. The distributed timed task scheduling method based on DelayQueue according to claim 1 is characterized in that: The task rollback in step S3 is achieved by: When an instance becomes inactive, other instances detect that its UID is no longer in the active list and reset the status of the tasks originally assigned to the instance to "initializing"; When the instance is shut down normally, the status of untriggered tasks in memory rolls back to "initializing".
7. The distributed timed task scheduling method based on DelayQueue according to claim 1 is characterized in that: The JobProcessor interface in step S4 is loaded via the Java SPI mechanism, and: Each task type corresponds to at least one JobProcessor implementation class; When a task is triggered, it is called in the order of processor priority until the task is successfully processed or the maximum number of retries is reached.
8. The distributed timed task scheduling method based on DelayQueue according to claim 1 is characterized in that: The sorting algorithm Sort in step S2 is a lexicographic order or a numerical order, which is used to determine the position of the instance in the active list.
9. The distributed timed task scheduling method based on DelayQueue according to claim 1, characterized in that: The safety shutdown in step S5 further includes: Stop all preloading threads and task execution thread pools; Persist unprocessed tasks in the DelayQueue to the database and mark them as "initialized".
10. The distributed timed task scheduling method based on DelayQueue according to claim 1, characterized in that: The method relies on a database to achieve decentralized coordination, and the time complexity of the task allocation algorithm is O(1).