A Distributed Task Scheduling Method and System Integrating Streaming Tasks and Batch Tasks
Through the distributed scheduling system and the active reporting mechanism, the unified management of flow tasks and batch tasks is solved, efficient resource utilization and stable task operation are achieved, and the convenience of task orchestration and data processing quality are improved.
Patent Information
- Application Number
- CN202210781586.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-04
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-07-04
AI Technical Summary
The existing open source scheduling system cannot effectively manage flow tasks and batch tasks, resulting in waste of resources and cluster failures, especially in the long-term running real-time tasks that cannot be released, affecting the execution of subsequent tasks.
A distributed scheduling system is adopted, and the main master node is selected through Zookeeper to realize unified management of flow tasks and batch tasks. The task status monitoring queue and active reporting mechanism are adopted to release the task submission thread, improve parallelism, and establish task dependencies through time rules and partition detection operators.
It realizes stable and unified management of flow tasks and batch tasks, reduces resource consumption, improves task submission parallelism, prevents cluster failures, and improves task orchestration convenience and data processing quality.
Smart Images

Figure CN115168037B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data, and in particular relates to a distributed task scheduling method and system integrating stream tasks and batch tasks. Background Art
[0002] In a scheduling system, the main functions are to trigger task startup regularly, accurately, and efficiently, and to track the status until the end. In the big data task system, tasks can be roughly divided into two categories: real-time tasks (stream tasks) and offline tasks (batch tasks). Mainstream big data engines such as Spark and Flink both provide the capabilities of stream tasks and batch tasks. Offline tasks: Also known as batch tasks, they are generally completed within a limited time and hold resources for a limited time, usually not exceeding the hour level. Real-time tasks: Also known as stream tasks, they run for a long time, even continuously, and hold resources continuously.
[0003] Whether it is a stream task or a batch task, the essence is to process data. An important function of the scheduling system is to manage data processing tasks in an orderly manner. For example, in a common ETL (Extract-Transform-Load) task, data extraction must be completed first, then business processing is carried out, and finally data is exported. In existing open-source and publicly available scheduling systems, the offline scheduling system and the real-time task management system are generally managed separately, which is determined by the two task forms.
[0004] In the existing open-source mainstream offline scheduling system architectures, such as DolphinScheduler, XXL-job, Quartz, etc., the general task execution process is as follows Figure 1 , after a task is hatched and an instance is submitted, it is handed over to a thread pool for processing. The threads in the thread pool obtain the task, initialize the task, run the task, and submit the task to Yarn, and then track the task status until it ends normally or abnormally. The thread monitors the entire task status.
[0005] For real-time tasks, if the tasks run in the scheduling system for a long time, the threads held by these tasks will not be released for a long time. The more tasks there are, the more resources are occupied, and finally all resources are exhausted, resulting in subsequent tasks being unable to execute, and finally possibly causing the cluster to fail. Currently, in the existing open-source architectures and publicly available scheduling architecture systems, there is no architecture that unifies the management of offline and real-time tasks. Summary of the Invention
[0006] In fact, there is no essential difference between streaming tasks and batch tasks in terms of data processing. In view of this, the present invention proposes a distributed streaming-batch integrated scheduling system and a method based on time rules to establish a complete dependency relationship between streaming tasks and batch tasks, realizing a streaming task and batch task scheduling system with unified task management, unified resource allocation, and unified permission control.
[0007] A distributed task scheduling method for integrating streaming tasks and batch tasks disclosed in the first aspect of the present invention includes the following steps:
[0008] When the system starts, the Master node and the Worker register with Zookeeper, and the main Master is selected among the Master nodes through Zookeeper;
[0009] The main and standby Masters incubate tasks and distribute tasks. After the user creates a DAG, the Master node incubates specific DAG instances and task instances from the DAG, and then distributes the task instances to the Worker nodes in sequence according to the task dependency relationship;
[0010] After receiving the message, the Worker node classifies the tasks and manages them through the local task thread pool and the remote task thread pool;
[0011] When a remote task is submitted, the remote task is transferred to the task status monitoring queue, thereby releasing the task submission thread and improving the parallelism of task submission;
[0012] In the status check process, a timed scanning mechanism is adopted to complete the check of a large number of tasks with a small number of threads.
[0013] Further, a thread is started in the task to report the status to the Worker at regular intervals, so that the long-term task status monitoring process is transformed from a thread and the timed scanning based on this thread into a record in the queue and periodic messages.
[0014] Further, the tasks include Spark tasks or Flink tasks or Mapreduce tasks of Yarn.
[0015] Further, in the case of a Yarn cluster failure, if the Worker node does not receive task status messages for multiple cycles, a scan is performed to judge the status of the tasks.
[0016] Further, after the streaming tasks and batch tasks are uniformly managed, the batch tasks continuously detect and check the data processing time of the streaming tasks, and trigger the scheduling of the batch tasks when specific conditions are met.
[0017] Further, when the streaming task processes data based on time partitioning, the specific condition includes: the next partition has been generated and the data has not changed within multiple cycles of the current partition.
[0018] Further, when the streaming task does not process data based on time partitioning, in a customized operator, the time position of the currently processed data is periodically written into the database, and the time detection operator of the batch task determines the business processing time of Flink through this data time, triggering the execution of the batch task.
[0019] Further, the data time includes the event time (EventTime) or process time of the stream.
[0020] A distributed task scheduling system integrating streaming tasks and batch tasks disclosed in the second aspect of the present invention includes:
[0021] Registration module: When the system starts, the Master node and the Worker register with Zookeeper, and the main Master is selected among the Master nodes through Zookeeper;
[0022] Distribution module: The main and standby Masters incubate and distribute tasks. After the user creates a DAG, the Master node incubates specific DAG instances and task instances from the DAG, and then distributes the task instances to the Worker nodes in sequence according to the task dependency relationship;
[0023] Classification module: After receiving the message, the Worker node classifies the tasks and manages them through the local task thread pool and the remote task thread pool;
[0024] Thread release module: When a remote task is submitted, the remote task is transferred to the task status monitoring queue, thereby releasing the task submission thread and improving the parallelism of task submission;
[0025] Status check module: In the status check process, an active reporting mode and a failure-proof timing mechanism are adopted to complete the checks of a large number of tasks with a small number of threads.
[0026] The beneficial effects of the present invention are as follows:
[0027] Provide a set of distributed scheduling architectures. On the existing scheduling process and distributed architecture, transfer the remote tasks to the task status monitoring queue, thereby releasing the task submission thread, improving the parallelism of task submission, and ensuring the stable operation of scheduling.
[0028] Provide a processing method that supports the integration of streaming tasks and batch tasks. At the same time, through optimization, the performance of the offline tasks themselves has also been greatly improved.
[0029] On the basis of ensuring the integration of streaming tasks and batch tasks, a method for establishing dependencies between real-time tasks and offline tasks is provided to truly integrate real-time tasks and offline tasks. Description of the Drawings
[0030] Figure 1 Schematic diagram of the execution method of existing tasks;
[0031] Figure 2 Graph of the proprietary names of existing scheduling tasks;
[0032] Figure 3 Distributed architecture diagram of the existing data engine;
[0033] Figure 4 Flowchart of the existing task scheduling;
[0034] Figure 5 Polling query status mechanism for stream-batch integration of the present invention;
[0035] Figure 6 Active reporting status mechanism for stream-batch integration of the present invention;
[0036] Figure 7 Way of establishing dependencies for stream-batch of the present invention. Detailed Embodiments
[0037] The present invention will be further described below with reference to the drawings, but the present invention is not limited in any way. Any transformation or replacement made based on the teachings of the present invention falls within the protection scope of the present invention.
[0038] The present invention first introduces the existing big data distributed scheduling architecture, referring to Figure 2 ,
[0039] Operator: Refers to the abstraction of a certain same function. For example, the HTTP operator mainly completes the HTTP call function.
[0040] Task: It is to match an operator with a specific service. For example, by matching the HTTP operator with a specific URL to obtain the result of the page.
[0041] Task dependency: There is a sequential execution relationship between tasks. For example, Figure 2 in Task 2 depends on Task 1. After Task 1 is completed, Task 2 can start running.
[0042] DAG: The whole composed of one or more tasks. As shown in Figure 2 on the left, Task 1, Task 2, and Task 3 form GAG1.
[0043] Scheduling period: Tasks run periodically according to a certain rule. For example, DAG1 runs once at 1:00 am every day. It is generally provided through a Cron expression.
[0044] DAG dependency: Establish a dependency relationship between two different DAGs. For example, Figure 2 in which DAG2 depends on DAG1.
[0045] DAG instance: The instance generated each time a DAG is scheduled is called a DAG instance. Each time DAG1 is scheduled, an instance with a time scale is generated.
[0046] Task instance: When a DAG instance is generated, all tasks in the DAG are generated into a task instance. The instance brings time characteristics.
[0047] This system uses zookeeper as the distributed coordination point. When starting up, the Master and Worker will register information with Zookeeper. After successful registration, the Master selects the main Master through Zookeeper. The implementation principle is as follows: All Master nodes write a record to the main Master node of Zookeeper. The one that writes successfully is the main Master, and the one that writes fails is the standby Master. The standby Master will monitor the status of the main Master. When the main Master fails, all nodes will compete again to become the main Master. The main Master monitors all Masters and Workers. When a Master and a Worker fail, the main Master will receive a message and perform fault recovery. Normally, it is a cluster of one main Master, multiple standby Masters, and multiple Workers, as specifically Figure 3 shown.
[0048] When the cluster is normal, the responsibilities of the main Master and the standby Master are the same, which is used to incubate tasks and then distribute tasks. When the main Master node fails, the main Master will stop incubating and give priority to handling the fault situation. Communication between the Master and the Worker is through Netty.
[0049] Task distribution process: Task scheduling and distribution are to distribute tasks to appropriate machines at an appropriate time, start and run them, and then monitor the task status process. In this scheduling system, from the creation of a task to its actual execution end, it will go through several processes such as user creation, Master incubation, distribution, Worker startup, task execution, task running status tracking, task end, and exception retry.
[0050] Such as Figure 4After the user creates a DAG, the Master will accurately spawn specific DAG instances and task instances according to the user's requirements. Then, based on the task dependencies, the task instances will be distributed to the Workers in sequence. After the distribution is successful, the Master will transfer the task instances to the cache to wait for the message callback. The message sending uses Netty communication, and the message callback adopts Kafka and a very low-frequency database scan. The database scan is to prevent message loss and generate zombies.
[0051] After receiving the message, the Worker will classify the tasks and manage them using different thread pools. Local tasks that can be quickly executed, such as Python / Shell / SQL classes, are managed by the Local Task Thread Pool; big data tasks such as Flink / Spark on Yarn are submitted using the Remote Task Thread Pool. Local tasks will listen to the whole process in the thread. Since the running time is short and the task volume is small, the thread pool resources are sufficient; for Remote tasks such as Flink / Spark, after the task is submitted, the status is managed by a queue, and the threads occupied in the Remote Task Thread Pool are released.
[0052] Reference Figure 5 In the above system architecture of the present invention, when a Remote task is submitted, the task is transferred to the task status monitoring queue, thereby releasing the task submission thread and improving the parallelism of task submission. In addition, in the status check process, a timing mechanism is adopted, and the threads are shared, and a small number of threads complete the check of a large number of tasks. In this way, the streaming tasks do not occupy resources for a long time, thus achieving an architecture that integrates streaming tasks and batch tasks.
[0053] In some embodiments, the present invention reduces the consumption of scheduling resources by tasks. In the above solution, whether it is a streaming task or a batch task, resources are still consumed to poll and scan the status on Yarn; in the case of millions of tasks, the concurrency ability is still affected. Therefore, a further optimized solution, the initiative reporting status solution by the business side, is adopted. The specific solution is as follows: Reference Figure 6, whether it is a Spark task, a Flink task, or a Mapreduce task based on Yarn, a thread is started within the task itself to report the status regularly. Under normal circumstances, the Worker judges the current task status through the received task messages; in the case of a Yarn cluster failure, if the Worker does not receive the task status message for multiple cycles, it can perform a scan to determine the status of the Task. Because of the active trigger, the corresponding sending cycle can be longer and the number of times can be fewer than the regular scan cycle. Except for special cases such as the overall failure of the Yarn cluster, the status information can be accurately sent and actively reported with higher punctuality. Transforming from passive polling to active status reporting saves resource consumption.
[0054] A long-term task status monitoring process is transformed from a thread and regular scans based on that thread into a record in a queue and periodic messages. Because of the active trigger, resource consumption is greatly reduced, making it possible to integrate streaming and batch processing.
[0055] There can be dependencies between streaming and batch processing: After implementing the integration of streaming and batch processing, it is also necessary to support users in conveniently establishing reliable dependencies between streaming tasks and batch tasks.
[0056] Real-time tasks run continuously, resulting in the inability to change the status. Therefore, dependencies cannot be simply established based on the task status. The essence of a task is data processing. Real-time tasks process continuous data streams, and offline tasks process batch data. It is easy to establish dependencies between offline tasks because the time segments of each task are very clear, while the time span of real-time tasks does not seem to be clear enough on the surface. If the time segments of real-time tasks can be well divided, dependencies can be created for real-time tasks.
[0057] Reference Figure 7 , the present invention provides two ways to divide the time segments of real-time tasks:
[0058] 1. The streaming task processes data based on time partitioning.
[0059] Write data based on time partitioning, design and use a partition detection operator, and regularly check whether the partition data output by the stream has been completely processed. There are many check rules. The rule we adopt is to check two parts.
[0060] (1) The next partition has been generated;
[0061] (2) Within multiple cycles of the current partition, the data has not changed.
[0062] 2. The streaming task does not process data based on time partitioning.
[0063] This type of task cannot perceive data changes. Therefore, in the business side and customized operators, it is required to periodically write the current processing data time position into the database, such as the EventTime or ProcessTime of Flink. Design and use a time detection operator to judge the business processing time of Flink through this time and trigger the operation of offline tasks.
[0064] In summary: After the unified management of stream and batch tasks, the offline tasks can continuously detect and check the data processing time of the stream tasks, and trigger the scheduling of the offline tasks when specific conditions are met, so as to establish the dependency of the stream and batch tasks.
[0065] The present invention also provides a distributed task scheduling system integrating stream tasks and batch tasks, including:
[0066] Registration module: When the system starts, the Master node and the Worker register with Zookeeper, and the main Master is selected among the Master nodes through Zookeeper;
[0067] Distribution module: The main and standby Masters incubate and distribute tasks. After the user creates a DAG, the Master node incubates the specific DAG instance and task instance from the DAG, and then distributes the task instances to the Worker nodes in sequence according to the task dependency relationship;
[0068] Classification module: After receiving the message, the Worker node classifies the tasks and manages them through the local task thread pool and the remote task thread pool;
[0069] Thread release module: When the remote task is submitted, transfer the remote task to the task status monitoring queue, thereby releasing the task submission thread and improving the parallelism of task submission;
[0070] Status check module: In the status check process, adopt the active reporting mode and the anti-failure timing mechanism, and use a small number of threads in common to complete the check of a large number of tasks.
[0071] The beneficial effects of the present invention are as follows:
[0072] The present invention provides a set of implementation methods for distributed big data scheduling, which is the basis for stream-batch integration. By continuously optimizing the task status check mechanism, the common thread consumption and scanning mechanism are converted into a piece of data in memory and the continuously received messages, reducing the resource consumption, thereby realizing the unified and effective management of stream tasks and batch tasks.
[0073] By means of the partitioning / time detection operator, the streaming tasks and batch tasks are fully integrated, the dependency relationship between the streaming tasks and batch tasks is established, the data quality is improved, and the situation where the offline tasks are executed in advance without the execution data being ready is prevented. With the establishment of the dependency, there is no need to set a time interval between streaming and batch processing, which greatly improves the convenience of task scheduling.
[0074] As used herein, the term "preferred" is meant to be used as an example, illustration, or exemplification. Any aspect or design described herein as "preferred" is not necessarily to be construed as more advantageous than other aspects or designs. Instead, the use of the term "preferred" is intended to present concepts in a concrete manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X uses A or B" is meant to naturally include any one of the permutations. That is, if X uses A; X uses B; or X uses both A and B, then "X uses A or B" is satisfied in any of the foregoing examples.
[0075] Moreover, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art based on a reading and understanding of this specification and the drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the appended claims. In particular, with respect to the various functions performed by the above-described components (e.g., elements, etc.), the terms used to describe such components are intended to correspond to any component (unless otherwise indicated) that performs the specified function of the component (e.g., it is functionally equivalent), even if it is not structurally equivalent to the disclosed structure that performs the functions in the exemplary implementations of the present disclosure shown herein. In addition, although a particular feature of the present disclosure has been disclosed with respect to only one of several implementations, such feature may be combined with one or other features of other implementations as may be desired and advantageous for a given or particular application. Also, insofar as the terms "comprises," "has," "contains," or any variation thereof are used in the detailed description or claims, such terms are intended to include in a manner similar to the term "includes."
[0076] Each functional unit in the embodiments of the present invention may be integrated into one processing module, or each unit may exist physically alone, or multiple or more than multiple units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disk, or the like. Each of the above-mentioned devices or systems may execute the storage method in the corresponding method embodiment.
[0077] In summary, the above embodiments are one implementation manner of the present invention. However, the implementation manners of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made under the premise of departing from the spirit and principle of the present invention shall be equivalent replacement manners and are all included in the protection scope of the present invention.
Claims
1. A distributed task scheduling method integrating stream tasks and batch tasks, characterized in that The scheduling method includes the following steps: When the system starts, the Master node and the Worker register with Zookeeper, and the main Master is elected among the Master nodes through Zookeeper; The main and standby Masters incubate tasks and distribute tasks. After the user creates a DAG, the Master node incubates the DAG into specific DAG instances and task instances, and then, according to the task dependency relationship, the offline task can continuously detect and check the data processing time of the streaming task. When specific conditions are met, the scheduling of the offline task is triggered, so as to realize the establishment of dependencies between streaming and batch tasks. After the streaming task and the batch task are uniformly managed, the batch task continuously detects and checks the data processing time of the streaming task. When specific conditions are met, the scheduling of the batch task is triggered, and the task instances are distributed to the Worker nodes in turn; After receiving the message, the Worker node classifies the tasks and manages them through the local task thread pool and the remote task thread pool; When a remote task is submitted, the remote task is transferred to the task status monitoring queue, thereby releasing the task submission thread and improving the parallelism of task submission; In the status check process, an active reporting mode and a fault-proof timing mechanism are adopted, that is, a thread is started in the task to report the status to the Worker at regular intervals, and the long-term task status monitoring process is transformed into a record in the queue and periodic messages. In the case of a Yarn cluster failure, if the Worker node does not receive the task status message for multiple cycles, a scan is performed to judge the status of the task, and a small number of threads are shared to complete the inspection of a large number of tasks.
2. The distributed task scheduling method integrating stream tasks and batch tasks according to claim 1, characterized in that The tasks include Spark tasks, Flink tasks, or Mapreduce tasks of Yarn.
3. The distributed task scheduling method integrating stream tasks and batch tasks according to claim 2, characterized in that, When the streaming task processes data based on time partitioning, a partition detection operator is used to detect the data processing time of the streaming task, and the specific conditions include: the next partition has been generated and the data in the current partition has not changed for multiple cycles.
4. The distributed task scheduling method integrating stream tasks and batch tasks according to claim 3, characterized in that, When the streaming task does not process data based on time partitioning, it is required that the streaming task periodically writes the current data time position being processed into the database, and the time detection operator judges the business processing time of the stream through the data time, triggering the running of the batch task.
5. The distributed task scheduling method integrating streaming tasks and batch tasks according to claim 4, wherein The data time includes event time or processing time.
6. A distributed task scheduling system integrating stream tasks and batch tasks, characterized in that, Including: Registration module: When the system starts, the Master node and the Worker register with Zookeeper, and the main Master is elected among the Master nodes through Zookeeper; Distribution module: The main and standby Masters incubate tasks and distribute tasks. After the user creates a DAG, the Master node incubates the DAG into specific DAG instances and task instances, and then, according to the task dependency relationship, the offline task can continuously detect and check the data processing time of the streaming task. When specific conditions are met, the scheduling of the offline task is triggered, so as to realize the establishment of dependencies between streaming and batch tasks. After the streaming task and the batch task are uniformly managed, the batch task continuously detects and checks the data processing time of the streaming task. When specific conditions are met, the scheduling of the batch task is triggered, and the task instances are distributed to the Worker nodes in turn; Classification module: After receiving a message, the Worker node classifies the tasks and manages them through the local task thread pool and the remote task thread pool; Thread release module: When a remote task is submitted, the remote task is transferred to the task status monitoring queue, thereby releasing the task submission thread and improving the parallelism of task submission; Status check module: In the status check process, an active reporting mode and a fault-proof timing mechanism are adopted, that is, a thread is started in the task to report the status to the Worker at regular intervals, converting the long-term task status monitoring process into a record in the queue and periodic messages. In the case of a Yarn cluster failure, if the Worker node does not receive task status messages for multiple cycles, a scan is performed to judge the status of the tasks, and a small number of threads are shared to complete the checks of a large number of tasks.
Citation Information
Patent Citations
Task scheduling component construction method, device, storage medium and server
CN109408212A
Distributed job flow task management and scheduling system and method
CN112860405A