A distributed high-concurrency scheduling system based on decentralized job tasks
Through the combination of distributed message queues and in-memory databases, a decentralized scheduling system is built, which solves the performance bottleneck problem of xxl-job in large concurrency scenarios, and realizes an efficient and scalable scheduler cluster, improving system performance and reliability.
Patent Information
- Application Number
- CN202210669827.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-14
AI Technical Summary
The existing technology xxl-job has problems in the case of large concurrency or big data scenarios where the dependence of relational databases is high, and the scheduler does not support high-performance clusters and Ack mechanisms affect performance.
The distributed message queue and in-memory database are adopted to realize a decentralized scheduling system. Through data interactions such as heartbeat, full job loading, full job synchronization and job dispatching, a one-master-multiple-slave scheduler cluster mode is adopted. The master node regularly resets the Leader lock logo, and all nodes save job information in memory to achieve data consistency and efficient scheduling.
It realizes high-performance scheduling without the support of relational databases. The scheduler can be horizontally expanded and all nodes participate in scheduling, improving system performance and efficiency, and avoiding database bottlenecks and single point failure risks.
Smart Images

Figure CN115061814B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and in particular to a distributed high-concurrency scheduling system based on a decentralized job task. Background Art
[0002] XXL-JOB is a distributed task scheduling platform, and its core design goals are rapid development, simple learning, lightweight, and easy to expand. It has now open-sourced and been integrated into the online product lines of many companies, and can be used out of the box.
[0003] XXL-JOB adopts a centralized management solution, mainly managing various scheduled task information through MySQL. When the trigger time of the scheduled task arrives, the task information is pulled into the memory from the database, and a trigger request is initiated to the task executor. This task executor can be a bean, groovy script, python script, etc., or an external http interface.
[0004] XXL-JOB includes the following main features:
[0005] (1) Simple to use; supports CRUD operations on tasks through a web page, with simple operations, convenient deployment, and easy maintenance.
[0006] (2) Elastic scaling; once there is a new upper limit or offline of the executor machine, the tasks will be re-allocated during the next scheduling.
[0007] (3) Fault transfer; in the case of selecting the "fault transfer" task routing strategy, if a machine in the executor cluster fails, it will automatically failover to a normal executor to send a scheduling request.
[0008] (4) Consistency; the scheduling center ensures the consistency of cluster distributed scheduling through a DB lock, and each scheduling task will only be triggered once.
[0009] (5) Task failure retry; supports customizing the number of task failure retries. When a task fails, it will actively retry according to the preset number of failure retries. Among them, sharding tasks support failure retries at the sharding granularity.
[0010] (6) Blocking handling strategy; the handling strategy when the scheduling is too dense and the executor has no time to process. The strategies include: single-machine serial (default), discard subsequent scheduling, and overwrite previous scheduling.
[0011] (7) Task failure alarm; by default, it provides email-based failure alarms, and at the same time reserves an extension interface, which can easily expand alarm methods such as SMS and DingTalk.
[0012] However, this technology is mainly applied to medium and small-sized business systems with low concurrency to ensure the accuracy and reliability of tasks. For large concurrency or big data scenarios, xxl-job has the following disadvantages:
[0013] (1) It must rely on a relational database; xxl-job needs to rely on a relational database (such as mysql, postgresql), such as for job storage, and to ensure consistency through database locks, etc. Therefore, in the case of high concurrency, as the number of jobs increases and the executor nodes are horizontally scaled out, the pressure on the database will increase, so there will be a certain upper limit on the concurrency volume.
[0014] (2) The scheduler does not support high-performance clusters; the xxl-job scheduling center is deployed in a centralized manner and supports multi-node deployment. However, only one scheduling node is allowed to trigger task scheduling at the same time, and the advantages of a scheduling center cluster cannot be realized. It only simply achieves the high availability of the scheduler.
[0015] (3) The Ack mechanism affects performance; xxl-job supports tracking and monitoring the execution results of jobs, which ensures the reliability of job execution but greatly reduces its performance. Summary of the Invention
[0016] The object of the present invention is to provide a distributed high-concurrency scheduling system based on decentralized job tasks, so as to solve the foregoing problems existing in the prior art.
[0017] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0018] A distributed high-concurrency scheduling system based on decentralized job tasks, including
[0019] A distributed message queue; used to implement data interaction including heartbeat, full job loading, full job synchronization, job dispatching, and job scheduling;
[0020] Nodes; the nodes include a scheduler and an executor; the nodes regularly send heartbeats to a specified distributed message queue to inform other nodes of their survival status; when the nodes join the cluster, they send registration messages to the specified distributed message queue immediately to inform other nodes of their existence; after the scheduler is successfully registered, the main scheduler dispatches scheduling tasks and synchronizes jobs; after the executor is successfully registered, the main scheduler synchronizes all jobs, and all schedulers add information about schedulable executors.
[0021] An in-memory database; used as a cache;
[0022] The node sends data to a specified distributed message queue for other nodes to consume, thereby realizing data interaction and communication with other nodes; for important data sent, it is automatically saved in the in-memory database to improve data persistence efficiency.
[0023] Preferably, in the scheduler cluster, there is one master and multiple slaves. The election method of the master node is preemptive. It is necessary to periodically reset the Leader lock identifier of the in-memory database. The Leader lock identifier expires within a preset time. The value of the Leader lock identifier is the unique Hash identifier of the scheduler. Whichever scheduler preempts the Leader lock identifier first, that scheduler is the master node, and the remaining schedulers are slave nodes;
[0024] Preferably, in the scheduler cluster, a one-master-multiple-slaves operation mode is adopted. The master node and the slave nodes have different divisions of labor, and all nodes are working; the master scheduler needs to implement full job loading, full job synchronization, and job dispatching functions, and both the master scheduler and the slave schedulers need to implement the work scheduling of the dispatched jobs;
[0025] Full job loading is specifically as follows: Notify all executors to summarize all jobs that need to be executed, or directly query the in-memory database. Full job loading is triggered after the scheduler competes to become the master scheduler. If it has never performed full job loading itself, this function is triggered;
[0026] Full job synchronization is specifically as follows: The master node sends a synchronization message to all slave nodes, including schedulers and executors; after receiving the message, the schedulers and executors load all job data from the in-memory database to the local; this function is used to keep the local data of all nodes consistent;
[0027] Job dispatching is specifically as follows: The master scheduler needs to perform job dispatching. Through the concept of load balancing, the full amount of jobs is evenly dispatched to all schedulers, including itself; after each scheduler receives a job, it starts to schedule. Each scheduler saves the current full amount of jobs; job dispatching is to dispatch the ID of the job. Each scheduler saves the current full amount of jobs, and the scheduler is only responsible for the scheduling of the jobs it receives; Job dispatching includes three types: full amount, addition, and deletion;
[0028] Work scheduling is specifically as follows: The scheduler has been performing job scheduling calculations. This work has no role distinction, that is, both the master scheduler and the slave schedulers are performing it simultaneously. Once a certain job reaches the time point when it needs to be executed, a communication message will be sent to the executor to notify it to execute the task; Job scheduling accompanies the entire life cycle of the scheduler and never stops.
[0029] Preferably, when there are insert, update, delete, and query operations on a job, a message will be sent to the specified distributed message queue, and all nodes will consume and synchronize the insert, update, delete, and query jobs to achieve job consistency.
[0030] Preferably, the specific process of synchronously adding, deleting, modifying, and querying operations of nodes is as follows:
[0031] When a node receives a message, it will determine the operation type of the message, and the operation type will clearly identify addition, query, modification, and deletion; different actions will be triggered for different types.
[0032] Preferably, the process of node consumption is the process by which the consumer of the distributed message queue receives messages. Specifically,
[0033] S1. The consumer creates a connection to open a channel and connects to the server of the distributed message queue;
[0034] S2. Requests the server to consume messages in the corresponding distributed message queue and sets the corresponding callback function;
[0035] S3. Waits for the server to respond and deliver messages in the corresponding distributed message queue, and the consumer receives the messages;
[0036] S4. Confirms the received messages;
[0037] S5. Deletes the corresponding confirmed messages from the distributed message queue;
[0038] S6. Closes the channel;
[0039] S7. Closes the connection.
[0040] Preferably, the system adopts a weakly centralized mode, that is, after the master node fails, during the election process, the scheduling work of other slave nodes is not affected and can continuously execute the task scheduling work.
[0041] Preferably, all nodes save the full amount of jobs in a thread-safe set.
[0042] The beneficial effects of the present invention are as follows: 1. It does not depend on any relational database, thus directly removing the bottleneck restriction of the relational database on high performance. All job information and node information are saved in the in-memory database, and all node information is fully synchronized to ensure data consistency. 2. It realizes fast response, fast scheduling, triggering, etc. for a large number of jobs. The execution of jobs does not need to care whether the executor executes successfully. The key lies in scheduling and triggering; thus, high performance and high efficiency of job scheduling are achieved. 3. It realizes the horizontal scalability of the scheduler, realizes its distributed function, ensures that each scheduler is executing the scheduling task, and there is only one scheduler node working among all nodes; making the best use of the server to improve performance as much as possible. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic diagram of the operation process of the scheduling system in the embodiment of the present invention;
[0044] Figure 2 It is a schematic diagram of the registration and cancellation processes of each node in the embodiments of the present invention;
[0045] Figure 3 It is a schematic diagram of the process of producing and consuming in the distributed message queue in the embodiments of the present invention;
[0046] Figure 4 It is a schematic diagram of the job management and assignment process in the embodiments of the present invention;
[0047] Figure 5 It is a schematic diagram of the job scheduling process in the embodiments of the present invention. Detailed implementation manners
[0048] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to explain the present invention and are not used to limit the present invention.
[0049] As Figure 1 shown, in this embodiment, a distributed high-concurrency scheduling system based on decentralized job tasks is provided, including a distributed message queue, nodes, and an in-memory database; the nodes send data to a specified distributed message queue for other nodes to consume, thereby realizing data interaction and communication with other nodes; for important data sent, it is automatically saved in the in-memory database to improve the data persistence efficiency. The scheduling system of the present invention can achieve the following functions:
[0050] I. Realize message interaction relying on the distributed message queue
[0051] The distributed message queue (usually refers to message middleware, such as rabbitmq); used to realize data interaction including heartbeat, full job loading, full job synchronization, job assignment, and job scheduling;
[0052] In this scheduling system, there is a large amount of data interaction, such as heartbeat, full job loading, full job synchronization, job assignment, job scheduling, etc., which are all realized relying on the message notifications of the distributed message queue (MQ).
[0053] The distributed message queue is an open-source message broker software that can implement the Advanced Message Queuing Protocol (AMQP), and can be deployed in a cluster mode, greatly improving the message transmission efficiency and consumption efficiency.
[0054] Nodes; the nodes include a scheduler and an executor; the scheduler and the executor regularly send heartbeats to a specified distributed message queue to inform other nodes of their survival status.
[0055] AsFigure 2 As shown, when the scheduler and the executor join the cluster, they send registration messages to the specified distributed message queue immediately to inform other nodes of their existence. After the scheduler registers successfully, the main scheduler dispatches scheduling tasks and synchronizes jobs. After the executor registers successfully, the main scheduler synchronizes all jobs, and all schedulers add information about schedulable executors.
[0056] In this embodiment, when there are create, read, update, or delete operations on jobs, the master node sends messages to the specified distributed message queue, and all nodes consume and synchronize the create, read, update, or delete jobs to achieve job consistency.
[0057] 1. The specific process of node synchronization of create, read, update, or delete jobs is as follows:
[0058] When a node receives a message, it will judge the operation type of the message. The operation type will clearly identify create, read, update, or delete; different actions will be triggered for different types.
[0059] 2. The process of node consumption is the process of the consumer of the distributed message queue receiving messages. As Figure 3 shown, specifically:
[0060] S1. The consumer creates a connection (Connection), opens a channel (Channel), and connects to the server (RabbitMQ Broker) of the distributed message queue;
[0061] S2. Requests the server (Broker) to consume messages in the corresponding distributed message queue and sets the corresponding callback function;
[0062] S3. Waits for the server (Broker) to respond and deliver messages in the corresponding distributed message queue, and the consumer receives the messages;
[0063] S4. Confirms (ack, auto-acknowledgment) the received messages;
[0064] S5. Deletes the corresponding confirmed messages from the distributed message queue;
[0065] S6. Closes the channel;
[0066] S7. Closes the connection.
[0067] II. Implementing competition and data caching by relying on a distributed in-memory database
[0068] An in-memory database (such as redis) is an open-source key-value storage system, mostly used as a cache, and can be deployed in a cluster mode.
[0069] In the scheduler cluster, there is one master and multiple slaves. The election method of the master node is preemptive. It is necessary to periodically reset the Leader lock identifier of the in-memory database. The Leader lock identifier expires within a preset time. The value of the Leader lock identifier is the unique Hash identifier of the scheduler. The scheduler that preempts this Leader lock identifier first becomes the master node, and the remaining schedulers become slave nodes.
[0070] The preset time can be set according to the actual situation to better meet the actual needs. For example, it can be set to 10 seconds.
[0071] In the distributed framework, all message notifications are asynchronous, and there will be a certain probability of data asynchronization. For example, in the case of incremental jobs, after the master scheduler receives the job addition, it sends asynchronous messages to notify all nodes, and then assigns the scheduler to which the job belongs. When the scheduler executes the scheduling, if it does not receive the job with the incremental message in time, it will cause the scheduling to fail. Therefore, the scheduling system of the present invention uses an in-memory database to cache such data to prevent functional abnormalities caused by data asynchronization.
[0072] III. Achieving high performance and availability of the scheduler cluster through weak centralization
[0073] As Figure 4 and Figure 5 shown, in the scheduler cluster, a one-master-multiple-slave operation mode is adopted. The master node and the slave nodes have different divisions of labor, and all nodes are working; the master scheduler needs to implement the functions of full job loading, full job synchronization, and job assignment. Both the master scheduler and the slave schedulers need to implement the work scheduling of the assigned jobs.
[0074] 1. The specific process of full job loading is as follows: Notify all executors to summarize all jobs that need to be executed, or directly query the in-memory database. Full job loading is triggered after the scheduler competes to become the master scheduler and has never performed full job loading itself.
[0075] 2. The specific process of full job synchronization is as follows: The master node sends synchronization messages to all slave nodes, including schedulers and executors; after receiving the messages, the schedulers and executors load all job data from the in-memory database to the local; this function is used to keep the local data of all nodes consistent.
[0076] 3. The job assignment is specifically as follows: The main scheduler needs to perform job assignment (rebalance). Based on the concept of load balancing, it evenly distributes all jobs to all schedulers, including itself. After each scheduler receives a job, it starts scheduling. Each scheduler stores the current full set of jobs. Job assignment means assigning the IDs of the jobs. Each scheduler stores the current full set of jobs, and the scheduler is only responsible for scheduling the jobs it has received (SchedulerJob). Job assignment includes three types: full volume, addition, and deletion.
[0077] 4. The job scheduling is specifically as follows. The scheduler has been performing job scheduling calculations all the time. This work is not role-specific, that is, both the main scheduler and the slave schedulers are performing it simultaneously. Once a job reaches the time point when it needs to be executed, it will send a communication message to the executor to notify it to execute the task. Job scheduling accompanies the entire life cycle of the scheduler and never stops.
[0078] In this embodiment, the master-slave mode is a centralized mode. In the scheduling system of the present invention, a weakly centralized mode is adopted. That is, after the master node fails, during the election process, the scheduling work of other slave schedulers is not affected, and the scheduling work of the tasks continues to be executed all the time.
[0079] The weakly centralized mode can not only avoid the single-point failure risk of the central node in the centralized mode, but also avoid the node interaction pressure of the large nodes in the decentralized mode.
[0080] IV. Avoiding database performance bottlenecks through decentralized data distributed storage
[0081] All nodes, including schedulers and executors, store the full set of jobs in a thread-safe collection (JVM memory) and do not need to load from the database when not in use, thus saving the database I / O overhead.
[0082] By adopting the above technical solutions disclosed in the present invention, the following beneficial effects are obtained:
[0083] The present invention provides a distributed high-concurrency scheduling system based on decentralized job tasks. This system does not depend on any relational database, thus directly removing the bottleneck limitation of the relational database on high performance. All job information and node information are stored in memory, and all node information is synchronized in full to ensure data consistency. This system can achieve fast response, fast scheduling, triggering, etc. for a large number of jobs. When the job is executed, there is no need to care whether the executor executes successfully. The key lies in scheduling and triggering. In this way, high performance and high efficiency of job scheduling are achieved. This system can achieve horizontal scalability of the scheduler, realize its distributed function, ensure that each scheduler is executing scheduling tasks, and there is only one scheduler node working among all nodes; make the best use of the server to improve performance.
[0084] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A distributed high-concurrency scheduling system based on decentralized job tasks, characterized in that: include, Distributed message queue; used to implement data interaction including heartbeat, full job loading, full job synchronization, job dispatching and job scheduling; Node; Nodes include schedulers and executors; Nodes periodically send heartbeats to the designated distributed message queue to inform other nodes of their survival status; When the node joins the cluster, it immediately sends a registration message to the designated distributed message queue to inform other nodes of its existence; After the scheduler is successfully registered, the main scheduler assigns scheduling tasks and synchronizes jobs; After the executor is successfully registered, the main scheduler synchronizes all jobs, and all schedulers add information about schedulable executors; In-memory database; Used as a cache; The node sends data to the designated distributed message queue for consumption by other nodes, thereby realizing data interaction and communication with other nodes; important data sent is automatically saved in the memory database to improve data persistence efficiency; In the scheduler cluster, there is one master and multiple slaves. The election method of the master node is preemptive. The leader lock flag of the memory database needs to be reset regularly. The leader lock flag expires within a preset time. The value of the leader lock flag is the unique hash identifier of the scheduler. The scheduler that preempts the leader lock flag first is the master node, and the other schedulers are slave nodes. In the scheduler cluster, a one-master-multiple-slave operation mode is adopted. The master node and slave nodes have different divisions of labor, and all nodes are working. The master scheduler needs to realize the functions of full job loading, full job synchronization, and job dispatching. Both the master scheduler and the slave scheduler need to realize the work scheduling of the dispatched jobs. The specific steps of full job loading are: notify all executors, summarize all jobs that need to be executed, or directly query the memory database. Full job loading is triggered after the scheduler competes to become the main scheduler. If the scheduler has never executed full job loading, this function will be triggered. The specific steps of full job synchronization are as follows: the master node sends a synchronization message to all slave nodes, including the scheduler and executor. After receiving the message, the scheduler and executor load all job data from the memory database to the local. This function is used to keep the local data of all nodes consistent. Job dispatch is as follows: the main scheduler needs to perform job dispatch, and through the concept of load balancing, it evenly dispatches the full amount of jobs to all schedulers, including itself; after each scheduler receives the job, it starts to schedule it. Each scheduler saves the current full amount of jobs; job dispatch is to dispatch the job ID, each scheduler saves the current full amount of jobs, and the scheduler is only responsible for the scheduling of the jobs it receives; job dispatch includes three types: full amount, add and delete; Specifically, the scheduler is always performing the scheduling calculation of the job. This work is role-independent, that is, the master scheduler and the slave scheduler are both performing it at the same time. Once a job reaches the time point where it needs to be executed, it will send a communication message to the executor to notify it to execute the task. Job scheduling accompanies the entire life cycle of the scheduler and never stops.
2. The distributed high-concurrency scheduling system based on decentralized job tasks according to claim 1, wherein: When there are create, update, delete, or query operations on a job, a message will be sent to the specified distributed message queue, and all nodes will consume and synchronize the create, update, delete, or query operations on the job to achieve job consistency.
3. The distributed high-concurrency scheduling system based on a decentralized job task according to claim 2, characterized in that: The specific process of node synchronization of create, update, delete, or query operations is as follows: When a node receives a message, it will determine the operation type of the message. The operation type will clearly indicate create, query, update, or delete. Different actions will be triggered for different types.
4. The distributed high-concurrency scheduling system based on decentralized job tasks according to claim 3, wherein: The process of node consumption is the process by which the consumer of the distributed message queue receives the message. Specifically: S1. The consumer creates a connection and opens a channel to connect to the server of the distributed message queue. S2. Requests the server to consume the messages in the corresponding distributed message queue and sets the corresponding callback function. S3. Waits for the server to respond and deliver the messages in the corresponding distributed message queue. The consumer receives the messages. S4. Confirms the received messages. S5. Deletes the corresponding confirmed messages from the distributed message queue. S6. Closes the channel. S7. Closes the connection.
5. The distributed high-concurrency scheduling system based on a decentralized job task according to claim 1, characterized in that: The system adopts a weakly centralized mode. That is, after the master node fails, during the election process, the scheduling work of other slave nodes is not affected and can continuously execute the task scheduling work.
6. The distributed high-concurrency scheduling system based on decentralized job tasks according to claim 1, wherein: All nodes save the full amount of jobs in a thread-safe collection.
Citation Information
Patent Citations
Job scheduling method and device, computer system and computer readable storage medium
CN113032125A
Task scheduling method and device, electronic equipment and storage medium
CN113778652A