A task processing method and a task processing system
By combining the task processing system of distributed scheduling and execution framework, the task sharding processing and state update are realized, solving the problem of low scheduling efficiency of stand-alone task and improving system performance and reliability.
Patent Information
- Application Number
- CN202510130208.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-02-05
AI Technical Summary
In the prior art, the scheduling of stand-alone tasks is inefficient, which can easily lead to task loss, and lack efficient distributed task scheduling and abnormal task management mechanisms, which affects system performance and reliability.
A task processing system combining a distributed scheduling framework and execution framework is adopted to achieve timely update of task status and automated detection and recovery of abnormal tasks through sharding and batch processing technology, supporting flexible abortion and efficient management of tasks.
It improves task processing efficiency, ensures efficient resource utilization and reliability of the system, reduces system load, and realizes timely update of task status and automated processing of abnormal tasks.
Smart Images

Figure CN119557112B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a task processing method and a task processing system. Background Art
[0002] With the development of technology and the continuous expansion of enterprise business scale, task scheduling is often involved. Currently, task scheduling can be mainly divided into two categories, namely single-machine task scheduling and distributed task scheduling. However, due to the limitation of single-machine processing capacity, the task processing efficiency is low, and when a single machine fails, tasks are likely to be lost. Therefore, an implementation scheme for distributed task scheduling is urgently needed. Summary of the Invention
[0003] Embodiments of this application provide a task processing method and a task processing system for implementing distributed scheduling and batch processing of tasks, and improving the task processing efficiency.
[0004] To achieve the above object, the embodiments of this application adopt the following technical solutions:
[0005] In a first aspect, this application provides a task processing method, which is applied to a task processing system combining a distributed scheduling framework and a distributed execution framework. The task processing method includes:
[0006] In response to a first task being scheduled, the first task is sliced through the distributed scheduling framework to obtain multiple first sliced tasks of the first task. For each first sliced task, the first sliced task is assigned to a first target node, and the scheduling status corresponding to the first sliced task is updated through a preprocessing function.
[0007] Through the first target node, batch processing of the first sliced task is combined with the distributed execution framework to update the execution status of the first sliced task. When the execution statuses of all the first sliced tasks indicate completion, the scheduling status of the first task is set to completed through the postprocessing function of the distributed scheduling framework, where completed means success or failure.
[0008] In this application, the task processing method utilizes both the scheduling ability of the distributed scheduling framework and the execution ability (such as batch processing ability) of the distributed execution framework. Since the distributed scheduling framework and the distributed execution framework themselves can improve performance, combining the two can effectively improve performance, thereby improving the task processing efficiency. In addition, through the distributed scheduling framework and the distributed execution framework, the execution status and scheduling status of tasks and sliced tasks can be updated in a timely manner, thereby realizing full-link tracking of tasks.
[0009] In a possible design, the process of batch processing the first sliced task through the first target node includes:
[0010] Divide the first sharding task into multiple data sets. Among them, at least one of the multiple data sets includes the data in the first sharding task.
[0011] The target node processes each data set respectively. Based on this, batch processing of tasks is achieved, improving the system performance.
[0012] In a possible design, task cancellation (i.e., task abortion) can also be performed. Obtain a task cancellation request, which may include the identifier of the second task. Correspondingly, the task cancellation request is used to trigger the abortion of the second task.
[0013] After that, according to the execution status of the second task, abort the second task and update the execution status of the second task to aborted. Through the post-processing function, update the scheduling status of the second task to successful. Based on this, flexible abortion of tasks can be achieved. And since task abortion is also regarded as the task having been scheduled and completed, the scheduling status of the first task can be updated to successful instead of aborted.
[0014] The process of aborting the second task according to the execution status of the second task described above may include: when the execution status of the second task is in execution, for each second sharding task of the second task, when the execution status of the second sharding task is in execution, continue to execute the current data set corresponding to the second sharding task through the second target node where the second sharding task is located.
[0015] After the current data set is executed, stop executing the new data set corresponding to the second sharding task.
[0016] Based on this, when aborting a sharding task, if the target node executing the sharding task is currently executing the data set of the sharding task, continue to execute the current data set through the target node. After execution, stop executing the new data set of the sharding task, that is, stop executing the unexecuted data set, to achieve task abortion. And data chaos can be avoided.
[0017] In a possible design, after stopping the execution of the new data set corresponding to the second sharding task, set the execution status of the second sharding task to aborted.
[0018] Correspondingly, updating the execution status of the second task to aborted includes:
[0019] When the execution status of each sharding task of the second task is in a terminal state, update the execution status of the first task to aborted. Among them, the terminal state includes one or more of success, failure, or aborted.
[0020] Based on this, when aborting the first task, when the execution status of each shard task of the second task (i.e., the above-mentioned second shard task) is the final state, it indicates that the second task has been successfully aborted. Therefore, the execution status of the second task can be updated to aborted, which matches the actual execution situation of the second task.
[0021] In a possible design, after updating the execution status of the second task to aborted, in response to the second task being rescheduled, when the execution status of the second shard task of the second task is aborted, perform business processing on the unprocessed data set of the second shard task. Optionally, perform business processing on the unprocessed data set of the second shard task through the second target node.
[0022] When the execution status of the second shard task of the second task is successful or failed, do not perform business processing on the second shard task. Based on this, avoid the repeated execution of the second shard task.
[0023] In a possible design, it is also possible to identify abnormal tasks. For each task with an execution status of in execution, obtain the standard shard number (or planned shard number) and the actual shard number of the task, and obtain the planned execution time of the task. The actual shard number represents the number of shard tasks that have been scheduled. The task includes the above-mentioned first task.
[0024] When the standard shard number is different from the actual shard number, and the first time difference between the current time and the planned execution time is greater than the first preset time threshold, mark the task as a first abnormal task. Among them, the scheduling status of the first abnormal task is abnormal. Based on this, realize the identification of tasks with scheduling exceptions.
[0025] In a possible design, sort the tasks to be scheduled and the first abnormal tasks to obtain the sorted tasks. After that, the sorted tasks can be scheduled in sequence. Based on this, realize the rescheduling of the first abnormal tasks and ensure the reliability of the system.
[0026] Among them, optionally, the tasks to be scheduled and the first abnormal tasks can be sorted based on the order of the planned execution times of each task in the tasks to be scheduled and the first abnormal tasks. Based on this, avoid the long-term unscheduling of tasks. Of course, other sorting rules can also be used, such as if the task has a corresponding priority, it can be sorted according to the priority.
[0027] In a possible design, the above-mentioned sequential scheduling of the sorted tasks may include:
[0028] In the case where the execution status of the sharding task of the scheduled task is successful or failed, no business processing is performed on the sharding task of the scheduled task to avoid duplicate scheduling of the sharding task.
[0029] In a second aspect, the present application provides a task processing system, which combines a distributed scheduling framework and a distributed execution framework to execute the task processing method as described above.
[0030] In a third aspect, the present application provides an electronic device, which includes a memory and one or more processors; the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device is caused to execute the task processing method as described above.
[0031] In a fourth aspect, the present application provides a chip, which includes a communication interface and at least one processor:
[0032] The communication interface is used for inputting and / or outputting signaling or data;
[0033] The at least one processor is used to execute a computer program to implement the task processing method as described above.
[0034] In a fifth aspect, the present application provides a computer-readable storage medium, which includes computer instructions. When the computer instructions run on an electronic device, the electronic device is caused to execute the task processing method as described above.
[0035] In a sixth aspect, the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device is caused to execute the task processing method as described above.
[0036] Optionally, the above-mentioned electronic device may be a server.
[0037] It can be understood that the beneficial effects that can be achieved by the task processing system described in the second aspect, the electronic device described in the third aspect, the chip described in the fourth aspect, the computer-readable storage medium described in the fifth aspect, and the computer program product described in the sixth aspect provided above can refer to the beneficial effects in the first aspect and any possible design thereof, and will not be elaborated here. Description of the Drawings
[0038] Figure 1 It is a schematic diagram of a distributed task scheduling system provided by an embodiment of the present application;
[0039] Figure 2 It is a schematic diagram of a task scheduling method provided by an embodiment of the present application Figure 1 ;
[0040] Figure 3 Schematic diagram of a task scheduling provided by an embodiment of the present application;
[0041] Figure 4 Schematic of a task scheduling method provided by an embodiment of the present application Figure 2 ;
[0042] Figure 5 Schematic of a task information provided by an embodiment of the present application Figure 1 ;
[0043] Figure 6 Schematic of a task information provided by an embodiment of the present application Figure 2 ;
[0044] Figure 7 Schematic of a task information provided by an embodiment of the present application Figure 3 ;
[0045] Figure 8 Schematic of a task information provided by an embodiment of the present application Figure 4 ;
[0046] Figure 9 Schematic of a task scheduling method provided by an embodiment of the present application Figure 3 ;
[0047] Figure 10 Schematic of a task information provided by an embodiment of the present application Figure 5 ;
[0048] Figure 11 Schematic of a task information provided by an embodiment of the present application Figure 6 ;
[0049] Figure 12 Schematic of a task scheduling method provided by an embodiment of the present application Figure 4 ;
[0050] Figure 13 Schematic of a task information provided by an embodiment of the present application Figure 7 ;
[0051] Figure 14 Schematic of a task information provided by an embodiment of the present application Figure 8 ;
[0052] Figure 15 Schematic of a task information provided by an embodiment of the present application Figure 9 ;
[0053] Figure 16 Schematic of a task information provided by an embodiment of the present application Figure 10 ;
[0054] Figure 17 Schematic of a task scheduling method provided by an embodiment of the present application Figure 5 ;
[0055] Figure 18 Schematic of a task information provided by an embodiment of the present application Figure 10 One. Detailed implementation manners
[0056] For the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner. In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or a similar expression thereof refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c may be single or multiple. In the embodiments of the present application, "first", "second", "1", and "2" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise specified, the meaning of "a plurality" is two or more.
[0057] As the business scale of enterprises continues to expand, task scheduling faces challenges such as high-concurrency processing requirements, the consistency of task status, and the recovery and management of abnormal tasks. Conventional distributed task scheduling systems can adopt frameworks such as Quartz, Elastic-job, or XXL-JOB, mainly responsible for task scheduling and simple task execution. There are usually two major problems in conventional distributed task scheduling systems. One is the inefficiency of task distribution and execution. Conventional distributed task scheduling systems cannot fully optimize the task scheduling strategy, resulting in inefficient utilization of resources during task scheduling and execution. Especially in high-concurrency and large-scale task scenarios, the response speed and throughput of the system are limited. The other is the lack of an efficient batch processing framework. For tasks that need to process a large amount of data, conventional distributed task scheduling systems generally cannot support efficient task execution, resulting in long task execution times and affecting the overall performance of the system.
[0058] In addition, in conventional distributed task scheduling systems, the management of abnormal tasks usually relies on manual intervention, which is inefficient and easily leads to tasks not being processed for a long time, affecting business continuity. That is to say, conventional distributed task scheduling systems lack an automated anomaly detection and recovery mechanism, resulting in low reliability of task execution.
[0059] And in conventional distributed task scheduling systems, the task cancellation operations in the task scheduling layer and the execution layer are often independent, lacking an effective coordination mechanism. When users want to cancel tasks, they often encounter the following problems: One is the interruption or omission of task cancellation. The cancellation operation of tasks cannot guarantee that the tasks in execution can be immediately interrupted, which may lead to waste of system resources and inconsistency in task execution. The other is the task status synchronization problem. The task statuses of different nodes are inconsistent, causing confusion in task status monitoring and affecting the scheduling and status update of subsequent tasks.
[0060] Therefore, in view of the above problems, the present application proposes a distributed task scheduling solution that combines the Elastic-job and Spring batch frameworks. Elastic-job is responsible for scheduling tasks and effectively splits tasks into multiple small tasks (i.e., shard tasks) for execution through sharding technology. Spring Batch provides standardized batch processing functions during task execution, and the execution of each task shard can be independently managed to improve task execution efficiency. Moreover, from the two dimensions of scheduling and execution, the status of tasks and their shard tasks is divided into scheduling status and execution status. The status of a task depends on the status of its shard tasks, enabling full-link tracking of tasks and their shard tasks, thereby enabling flexible shard management. In addition, when applying this distributed task solution to large-scale distributed task processing scenarios, tasks can be executed in parallel through sharding, achieving efficient management and reliable execution of tasks, ensuring the efficient utilization of system resources, reducing system load, and improving processing efficiency, thereby enhancing the overall performance of the system.
[0061] In addition, based on the scheduling status and execution status of tasks and their shard tasks, automatic detection and recovery of abnormal tasks can be achieved, avoiding tasks from remaining unprocessed for a long time and ensuring the reliability of task execution.
[0062] Moreover, flexible task abortion can be achieved, ensuring the timeliness of abortion and the timely update of the status of tasks and shard tasks.
[0063] Correspondingly, as Figure 1 shown, the distributed task scheduling system corresponding to the above distributed task scheduling solution may include a task scheduling creation module, a task to-be-scheduled scanning module, a task scheduling verification module, a task scheduling module, a task execution module, and a task cancellation module.
[0064] The above task scheduling creation module is used to implement task name and configuration parameter definition, task scheduling strategy definition, task dependency definition, task configuration saving, task execution plan creation, etc.
[0065] The above task to-be-scheduled scanning module is used to implement task scanning trigger, identification of tasks to be executed, identification of abnormal tasks, task priority sorting, etc.
[0066] The above task scheduling verification module is used to implement uniqueness verification of tasks to be scheduled, verification of task execution time window, uniqueness verification of scheduled tasks, task dependency verification, etc.
[0067] The above task scheduling module is used to implement task triggering and scheduling, task distribution and sharding, registration and heartbeat monitoring, task status update, abnormal task scheduling, etc.
[0068] The above task execution module is used to implement batch task initialization, task shard uniqueness verification, batch task execution trigger, batch task shard execution, batch task status update, etc.
[0069] The above task cancellation module is used to implement execution task cancellation operations, pending cancellation status updates, scheduled scanning of tasks to be cancelled, cancellation of batch task execution, cancellation of scheduled task scheduling, task status updates, etc.
[0070] It should be noted that the above Figure 1 The modules included in the distributed task scheduling system shown are only examples. The distributed task scheduling system may also include other modules, such as a task report management module, which is used to implement task execution status tracking, execution status tracking of sharded tasks, task execution performance statistics, abnormal task identification and analysis, task dependency relationship statistics, task execution parameter recording, historical execution trend analysis, etc. This application does not limit the functional modules included in the distributed task scheduling system.
[0071] This application provides a distributed task scheduling system that supports task scheduling and sharded task execution management functions based on Elastic-job and Spring batch. Based on the pre-listener and post-listener of Elastic-job, the monitoring and update of the scheduling status of the sharded tasks of the task are realized, so as to realize the timely update of the scheduling status of the task. And, based on Spring batch, the batch processing of sharded tasks is realized, the task processing efficiency is improved, and the execution status of the sharded tasks is updated according to the execution situation of the sharded tasks, so as to realize the timely update of the execution status of the task.
[0072] The process of updating the task and the scheduling status of the sharded tasks of the task by the above pre-listener and post-listener based on Elastic-job is mainly realized through the task scheduling module. The operation process executed by the task scheduling module is mainly divided into four parts, namely task trigger and configuration registration, task shard management and scheduling, task status recording, and task status update. Exemplarily, as Figure 2 shown, the operation process (i.e., task scheduling) may include:
[0073] S201. In response to task scheduling processing, the task scheduling module determines whether to shard the task based on the configuration information of the task to be processed.
[0074] The above-mentioned task to be processed (or referred to as the first task) can be a scheduled task. The configuration information of the task can include one or more of the information such as the task name, task ID, planned execution time of the task, number of shards of the task, etc. Optionally, the number of shards of the task can be one or more. When the number of shards is one, it indicates that the task actually does not need to be sharded. In other words, this task can also be considered a single-shard task. When the number of shards is multiple, it indicates that this task is a multi-shard task, that is, there are multiple shard tasks for the task. For example, if the task is 10,000 pieces of data and the number of shards is 4, then this task is divided into 4 shard tasks, and each shard task includes 2,500 pieces of data. Of course, the number of data included in the shard tasks exemplified here is just an example, and the number of data included in each shard task can also be uneven. This application does not limit the specific number included in the shard tasks.
[0075] In the embodiment of this application, when the planned execution time of the task is reached, it indicates that the task needs to be scheduled. Then the task scheduling module can determine whether the task needs to be sharded based on the configuration information of the task. If it is sharded, it indicates that this task is a multi-shard task, and the task scheduling module can execute S202 - S208. If it is not sharded, it indicates that this task does not need to be sharded. In other words, this task can also be considered a single-shard task, and the task scheduling module can execute S209 - S214. It should be understood that the planned execution time of the task can be understood as the priority of the task. The earlier the planned execution time, the higher the priority, so that the task can be scheduled faster, avoiding affecting related services due to untimely scheduling of the task.
[0076] Exemplarily, the configuration information includes the number of shards. When the number of shards is multiple, the task scheduling module can determine that the task to be processed needs to be sharded. When the number of shards is one, the task scheduling module can determine that the task to be processed does not need to be sharded.
[0077] S202. The task scheduling module determines the target nodes from the node cluster and the shard tasks of the task corresponding to each target node based on the number of shards.
[0078] In the embodiment of this application, the task scheduling module can determine multiple target nodes from the node cluster and the shard tasks corresponding to each of the multiple target nodes based on the number of shards and in combination with a preset allocation rule. Among them, the target nodes can be docker hosts.
[0079] For example, if the preset allocation rule is the average allocation rule, the number of shards is 4, and the number of nodes is 5, then the number of target nodes can be 4. The shard task corresponding to target node 1 can be shard 0, the shard task corresponding to target node 2 can be shard 1, the shard task corresponding to target node 3 can be shard 2, and the shard task corresponding to target node 4 can be shard 3. Additionally, the preset allocation rule can also be other rules. For example, the shard tasks can be preferentially allocated to nodes with lighter loads.
[0080] S203. For each target node, the task scheduling module sends the shard task corresponding to the target node to the target node.
[0081] Continuing with the above example, as Figure 3 shown, the task scheduling module sends the data of shard 0 to target node 1, the data of shard 1 to target node 2, the data of shard 2 to target node 3, and the data of shard 3 to target node 4. Of course, the example of sending the data of the shards to the corresponding target nodes in this application is only an example under an ideal state. It is also possible that the data of the shards is not sent to the corresponding nodes but to other nodes. For example, the data of shard 0 should be sent to target node 1, but it is actually sent to target node 2. Generally speaking, the task scheduling module can send the shard tasks to the target nodes to perform business processing on the shard tasks through the target nodes. The target node can be the one corresponding to the shard task or not, and this application does not limit it.
[0082] In some embodiments, the task scheduling module can perform task sharding to obtain multiple shard tasks of the task and allocate the shard tasks. Optionally, the task scheduling module can perform operations such as task sharding and shard task allocation through a preset method, such as through a Zookeeper cluster.
[0083] S204. The task scheduling module updates the scheduling status corresponding to the shard task to running through a pre-listener.
[0084] Exemplarily, for each shard task (or referred to as the first shard task) of the task, after the pre-listener monitors that the shard task is allocated to the corresponding target node, it can update the scheduling status corresponding to the shard task to running (such as running), thereby implementing the pre-processing of the scheduling status (such as referring to the pre-processing shown above Figure 3 ). Among them, running can also be referred to as in operation or in execution. Among them, the pre-listener is an example of the pre-processing function.
[0085] In addition, optionally, in practical applications, considering that the shard tasks may not be scheduled properly as required, the pre-listener may not detect that the shard tasks are assigned, and the scheduling status corresponding to the shard tasks can still be the initial status, such as waiting to be scheduled.
[0086] S205. The task scheduling module updates the scheduling status corresponding to the task to running according to the scheduling status corresponding to the shard task.
[0087] Exemplarily, when the scheduling status corresponding to the shard task of a task is running, the pre-listener can set the scheduling status corresponding to the task to running. In addition, optionally, if the scheduling statuses corresponding to the shard tasks of a task are all waiting to be scheduled, the scheduling status corresponding to the task can still be the initial status, such as the waiting-to-be-scheduled status.
[0088] S206. The target node performs batch processing on the shard tasks.
[0089] Among them, the implementation process of S206 can be referred to below. Here, it is only stated first that after the shard task is executed, if the shard task is successfully executed, the execution status of the shard task is successful. If the shard task fails to be executed, the execution status of the shard task is failed, thereby realizing the update of the execution status (such as referring to the execution status shown above). Figure 3 shown execution status).
[0090] S207. After all the shard tasks of the task to be processed are executed, the task scheduling module updates the scheduling status corresponding to the task to success or failure through the post-listener.
[0091] In the embodiment of the present application, after the shard task is executed, the task scheduling module can set the scheduling status of the shard task to indicate completion through the post-listener. For example, if the execution status of the shard task is successful, the scheduling status of the shard task is also successful. If the execution status of the shard task is failed, the scheduling status of the shard task is also failed. Among them, the post-listener is an example of the post-processing function.
[0092] As Figure 2 shown, S207a. When the scheduling statuses of all the shard tasks of a task are all successful, it indicates that all the shard tasks have been successfully processed, and the task scheduling module updates the scheduling status corresponding to the task to success, thereby realizing the post-processing of the scheduling status (such as referring to the post-processing shown above). Figure 3 shown post-processing).
[0093] When there is a shard task of a task whose scheduling status is failed, the task scheduling module updates the scheduling status corresponding to the task to failed (see S207b).
[0094] S208. The task scheduling module deletes the registration data of the task.
[0095] In the embodiments of the present application, when the scheduling status of the task is completed (such as the successful or failed status mentioned above), indicating that all shard tasks of the task have been executed, it means that all the data of the task has been processed. The task scheduling module can determine that the task scheduling is completed and can clean up the data of the task. For example, delete the registration node data of the task through Zookeeper. Of course, other data related to the task can also be deleted to avoid unnecessary occupation of resources by task data. Or, the task scheduling module may not delete the registration data, that is, S208 is an optional step.
[0096] In some embodiments, the operations performed by the above-mentioned task scheduling module may be performed by node 1 (or device 1) of Elastic-job. The implementation processes of S201-S208 can be specifically combined with the above-mentioned task triggering and configuration registration, task shard management and scheduling, task status recording, and task status update.
[0097] For task triggering and configuration registration, when the scheduled execution time of the task arrives, the task is scheduled and triggered. Device 1 determines multiple target nodes and determines the shard tasks corresponding to each target node. Then, device 1 can start monitoring through a pre-listener. After monitoring that a shard task is assigned to a target node, device 1 can update the scheduling status of the shard task to running.
[0098] For the above-mentioned task shard management and scheduling, device 1 can record information such as the task, the ID of the task instance, the scheduling status of the task, the execution status, the scheduling status of the shard task, and the execution status. The task instance represents the task that is currently being executed. The relationship between the task and the task instance can be a one-to-N relationship.
[0099] Optionally, device 1 can record this information through a preset storage medium, and the preset storage medium can be set according to requirements. For example, the preset storage medium can be a node (such as a task management node, a sharding node, etc.), or a database, etc. The present application does not limit the preset storage medium.
[0100] In addition, the target node starts batch processing of the shard task through Spring Batch to achieve refined execution of the assigned task and realize independent execution and management of each shard task.
[0101] For the above-mentioned task status recording, when the shard task is executed and completed through Spring Batch, after device 1 monitors through a post-listener that the execution status of the shard task indicates completion, the scheduling status of the shard task is also set to completed, and this completion can be successful or failed.
[0102] Moreover, the device 1 can also update the scheduling status of the task in a timely manner through the post - listener. When the scheduling status indicators of all the shard tasks of the task indicate completion, then the scheduling status of the task is completion, and this completion can be either successful or failed.
[0103] Optionally, after the device 1 determines the scheduling status of the shard task through the post - monitor, it can update the information recorded by the relevant nodes, such as updating the scheduling status of the shard task in the above - mentioned task management node, to achieve timely dynamic update of the scheduling status of the shard task.
[0104] In addition, if the execution status or the scheduling status of the task is failed, the task can be regarded as an abnormal task, and thus the abnormal task can be rescheduled. Specifically, reference can be made to the content related to abnormal tasks in the following text.
[0105] For the above - mentioned task status update, when the device 1 determines that the scheduling status of each shard task of the task is the completion status or the failure status through the scheduling status of each shard of the task recorded by the relevant nodes, it indicates that each shard task of the task has been executed, that is, the task has been executed. Then the device 1 clears the registration node data related to the task, such as the configuration information of the task, etc. In addition, the device 1 can also delete the nodes running the shard tasks of the task in the task management node.
[0106] It should be understood that from the foregoing, storing information through nodes is only an example of storing information through a preset storage medium. Correspondingly, the device 1 clearing the registration node data related to the task is also only an example, that is to say, the device 1 can delete the relevant data of the task from the preset storage medium.
[0107] In the embodiments of the present application, the task scheduling module implements a beforeJobExecuted pre-listener and an afterJobExecuted post-listener based on the ElasticJobListener interface of the Elastic Job framework, so as to implement pre-listening and post-listening of the scheduling status of tasks and the shard tasks of the tasks, ensuring timely and accurate update of the scheduling status. Currently, the above Elastic-job is only an example of a distributed scheduling framework, and this distributed scheduling framework can also be other types of frameworks. Correspondingly, implementing the update of the pre-scheduling status of tasks and the shard tasks of the tasks through the pre-listener of Elastic-job is only an example, and it can also be achieved in other ways, such as by running code that can implement functions related to the pre-listener to update the pre-scheduling status. That is to say, the task scheduling module can update the pre-scheduling status of tasks and the shard tasks of the tasks through the pre-processing function. For example, through the pre-processing function, the scheduling status corresponding to the shard task is updated to running.
[0108] Similarly, implementing the update of the post-scheduling status of tasks and the shard tasks of the tasks through the post-listener of Elastic-job is only an example, and it can also be achieved in other ways, such as by running code that can implement functions related to the post-listener to update the post-scheduling status. That is to say, the task scheduling module can update the post-scheduling status of tasks and the shard tasks of the tasks through the post-processing function. For example, the task scheduling module updates the scheduling status corresponding to the task to success or failure through the post-processing function.
[0109] S209. The task scheduling module determines the target node.
[0110] In the embodiments of the present application, since there is no need for sharding, the task scheduling module can directly select a target node from the node cluster, and this target node is the node that executes the task to be processed.
[0111] S210. The task scheduling module assigns the task to the target node.
[0112] S211. The task scheduling module updates the scheduling status of the task to running through the pre-listener.
[0113] S212. The target node performs batch processing on the task.
[0114] S213. After the task is executed, the task scheduling module updates the scheduling status corresponding to the task to completed through the post-listener.
[0115] Wherein, completed indicates success or failure.
[0116] S214. The task scheduling module deletes the registration data of the task.
[0117] Among them, the implementation processes of S209 - S214 are similar to those of S202 - S208 above, and will not be elaborated here.
[0118] The above introduced the process of monitoring the scheduling status of the sharding tasks of a task through the pre - listener and post - listener of Elastic - job. As described above, the execution of the above - mentioned sharding tasks can be implemented through batch processing of the Spring Batch framework. The following will continue to introduce the execution process of the sharding tasks.
[0119] The execution process of the above - mentioned sharding tasks is mainly implemented through the task execution module of the target node. The operation process executed by the task execution module can be mainly divided into three parts, namely sharding task execution, sharding status update, and overall task status update. Exemplarily, as Figure 4 shown, the operation process can include:
[0120] S301. For each sharding task, after the target node receives the sharding task, the task execution module of the target node updates the execution status of the sharding task to "executing" by calling the pre - listener.
[0121] Exemplarily, after being assigned a sharding task, the task execution module of the target node calls the pre - listener to initialize the execution status of the sharding task, that is, set it to "executing". The target node deploys the Spring Batch framework. Correspondingly, the task execution module can call the beforeJob pre - listener and afterJob post - listener by calling the JobExecutionListener interface of the Spring Batch framework.
[0122] Optionally, when the execution status of a sharding task is "executing", the execution status of the task to be processed to which the sharding task belongs can be set to "executing" (as Figure 5 shown). Or, after the sharding task is assigned to the target node, the execution status of the task can be set to "executing", and this application does not limit the sequence of initializing the status of the sharding task and the task.
[0123] S302. The task execution module divides the sharding task into multiple data blocks.
[0124] S303. For each data block, the task execution module performs business processing on the data in the data block.
[0125] In the embodiments of the present application, the task execution module of the target node divides the data of the sharding task into multiple data chunks (chunks) based on the Spring Batch framework. Then, the task execution module performs corresponding business processing on the data in each data chunk through business processing code, thereby implementing batch processing of the task. Of course, similar to Elastic-job mentioned above, the Spring Batch framework is only an example of a distributed execution framework, and this distributed execution framework can also be other types of frameworks. Correspondingly, the above-mentioned division of the data of the sharding task into multiple chunks based on the Spring Batch framework is only an example of batch processing of the sharding task. Generally speaking, the task execution module can divide the sharding task into multiple data sets, and at least one of the multiple data sets includes the data in the sharding task. In addition, it should be noted that for a single-sharding task, which can be understood as a multi-sharding task with a sharding number of 1, batch processing of the single-sharding task can still be performed according to the batch processing process of the sharding task of the multi-sharding task.
[0126] S304. After all data chunks of the sharding task are processed, the task execution module updates the execution status of the sharding task to completed.
[0127] Wherein, completion indicates failure or success.
[0128] In the embodiments of the present application, after all data chunks of the sharding task are processed, it indicates that the sharding task is executed and completed. The task execution module can record the execution result of the sharding task and update the execution status of the sharding task. In the case where the execution result is a failure result, the task execution module updates the execution status of the sharding task to failed. In the case where the execution result is a success result, the task execution module updates the execution status of the sharding task to success.
[0129] Exemplarily, the task execution module can monitor the execution situation of the sharding task by calling a post listener and update the execution status of the sharding task.
[0130] S305. After all sharding tasks are executed, if the execution status of all sharding tasks is successful, the task scheduling module updates the execution status of the task to be processed to successful through a post listener.
[0131] Exemplarily, the task execution module can synchronize the execution status of the sharding task to the task scheduling module through a post listener so that the task scheduling module can update the execution status. For example, the task is a financial bill, which includes four sharding tasks: shard 0, shard 1, shard 2, and shard 3. After shard 0 and 1 are successfully executed, the execution status of shard 0 and 1 is T (see Figure 6 )). Wherein, T represents success.
[0132] After the execution status indicators of all shard tasks are completed, the task scheduling module can update the execution status of the task to completed through a post-listener, thus achieving the update of the overall execution status of the task.
[0133] Specifically, when the execution status of each shard task of the task is successful, the post-listener can set the execution status of the task to successful (such as the execution status of the task shown Figure 7 is successful). Among them, successful can also be referred to as successful execution.
[0134] S306. If the execution status of a shard task is failed, the task scheduling module calls the post-listener to update the execution status of the task to be processed to failed.
[0135] When the execution status of each shard task of the task is failed, the post-listener can set the execution status of the task to failed (such as the execution status of the task shown Figure 8 is failed, where Figure 8 the 2-F shown indicates that the execution status of shard 2 is failed). Among them, failed can also be referred to as failed execution.
[0136] In some embodiments, after the task execution is completed, whether it is successful execution or failed execution, it is the final state of the task. The task scheduling module and the task execution module can then release the resources occupied by the task, such as cache resources or other resources, to ensure that the system has the resource environment required to execute the next task.
[0137] In the embodiments of the present application, the above task scheduling module and task execution module can be on the same node, that is, on the same device, or can be distributed on different devices. When distributed on different devices, it means that the logical codes of Elastic-job and the Springbatch framework are distributed on different devices, and the business capabilities between the two are decoupled. Through interface calls, the reuse of business capabilities can be achieved, enhancing the portability of functions. And when distributed on the same device, cost can be saved.
[0138] It should be noted that, similar to the previous text, the task execution module can also update the status of the task and shard tasks through the pre-processing function and post-processing function of the distributed execution framework. For example, the status before the task execution is completed can be updated by the pre-processing function of the distributed execution framework. And the status after the task execution is completed can be updated by the post-processing function of the distributed execution framework.
[0139] In some embodiments, the distributed task scheduling system also supports a task cancellation function (or task abortion function). When a task needs to be cancelled, the task cancellation module can promptly abort the execution of the task to ensure the efficient utilization of resources. Exemplarily, the execution process of the task cancellation module mainly includes three parts, namely, triggering of the task cancellation request, updating of the task status, and checking of the task status and resource cleanup. The following will be combined with Figure 9 , and the execution process will be introduced in detail. As Figure 9 shown, the execution process may include:
[0140] S601. The task cancellation module obtains a task cancellation request. Among them, the task cancellation request includes the identifier of the task to be cancelled.
[0141] Among them, the identifier of the task to be cancelled (or the second task) represents the information that can determine the task to be cancelled, such as ID, name, etc.
[0142] Exemplarily, the above task cancellation request can be triggered by a user or a node. For example, the node is a target node. When the preset cancellation condition is met, the target node can determine that the task or shard task on the target node needs to be cancelled, and then the target node can trigger a task cancellation request.
[0143] Among them, the preset cancellation condition can be set according to requirements. For example, the preset cancellation condition includes that the used resources of the target node are greater than the preset threshold, the current load is greater than the preset load, or the target node runs abnormally (such as crashing, restarting, etc.). The present application does not limit the specific conditions included in the preset cancellation condition. The used resources may include, but are not limited to, one or more of disk occupancy rate, CPU utilization rate, and memory usage rate.
[0144] S602. The task cancellation module obtains the execution status of the task to be cancelled.
[0145] In the embodiments of the present application, the task cancellation module can use the task indicated by the task cancellation request as the task to be cancelled, that is, the task that will no longer be executed. Then, the task cancellation module can obtain the execution status of the task to be cancelled for aborting the execution of the task to be cancelled according to the execution status.
[0146] S603. When the execution status of the task to be cancelled is in execution, the task cancellation module sends a stop signal to the task execution module.
[0147] Among them, the stop signal is used to trigger the abortion of the execution of the task to be cancelled. Exemplarily, the stop signal is used to trigger the abortion of the execution of a new data block of the task to be cancelled.
[0148] Exemplarily, the task cancellation module can send a stop signal to the task execution module by calling the jobExecution.stop() or JobOperator.stop() method of Spring Batch. Of course, similar to the previous text, the Spring Batch framework is only an example of a distributed execution framework. Correspondingly, triggering the abortion of a task through a stop signal is also only an example. The task cancellation module can trigger the abortion of a task through the stop function under other frameworks. Generally speaking, the task cancellation module can send a stop request to the task execution module, and this stop request is used to trigger the abortion of the task.
[0149] S604. In response to the stop signal, for each shard task A whose execution status of the task to be cancelled is in execution, the task execution module updates the execution status of the shard task A to aborting.
[0150] In the embodiment of the present application, in response to the stop signal, for the shard task B whose execution status of the task to be cancelled is completed (success or failure), since the shard task B has been executed, the shard task B does not need to be aborted. For the shard task C whose execution status of the task to be cancelled is to be executed, since the shard task C has not started to be executed yet, therefore, the task execution module can directly set the execution status of the shard task C to aborted. And for the shard task A whose execution status of the task to be cancelled is in execution, since the data in the shard task A is currently being processed, therefore, the task execution module can set the execution status of the shard task A to aborting (i.e., stopping). Among them, the shard task A, the shard task B, and the shard task C are all shard tasks of the task to be cancelled, and they can be called the second shard tasks.
[0151] S605. After the data block of the currently executed shard task A is processed, the task execution module updates the execution status of the shard task A to aborted.
[0152] In the embodiment of the present application, in order to improve the reliability of business data execution, when it is necessary to abort the shard task A, the task execution module may not immediately abort the execution of the shard task A, but after all the data in the data block of the currently executed shard task A is executed, then abort the shard task A, so as to stop executing the remaining unexecuted data blocks. And the task execution module updates the execution status of the shard task A from aborting to aborted. Based on this, the smooth termination of the shard task A can be realized, and it can be avoided that if the remaining unexecuted data blocks of the shard task A depend on the currently executed data block and the data block of the currently executed shard task A is immediately aborted, it may cause data disorder problems when processing the remaining unexecuted data blocks next time, ensuring data consistency, and thus guaranteeing the reliability of business processing.
[0153] For example, the sharding task A includes shard 1, which includes 2000 pieces of data. Shard 1 is divided into 4 data chunks, namely chunk0, chunk1, chunk2, and chunk3, and each data chunk includes 500 pieces of data. If the target node assigned shard 1 receives a stop signal during the execution of chunk2, the task execution module of the target node can first update the execution status of shard 1 to "aborting". After that, after all 500 pieces of data in chunk2 are executed, the target node stops executing chunk2 and chunk3, thereby aborting the execution of sharding task A, and the task execution module updates the execution status of shard 1 to "aborted".
[0154] Optionally, similar to the previous text, the above task execution module can update the execution status of the sharding task A from "aborting" to "aborted" through a post listener.
[0155] In some embodiments, in response to the stop signal, the task execution module can also set the execution status of the task to be cancelled to "aborting". Optionally, the task execution module can synchronize the execution status of the sharding task A of the task to be cancelled to the task scheduling module through a post listener, and the task scheduling module updates the execution status of the task to be cancelled to "aborting".
[0156] Among them, optionally, during the abortion of the task to be cancelled, the scheduling status of the sharding task A of the task to be cancelled can also be set to "aborting". Additionally, the scheduling status of the task to be cancelled can also be set to "aborting".
[0157] It should be noted that, similar to the previous text, the above chunk is only an example of the data set. Correspondingly, after the data set processing of the currently executed sharding task A is completed, the task execution module updates the execution status of the sharding task A to "aborted".
[0158] S606. When the execution statuses of all sharding tasks A of the task to be cancelled are "aborted", the task execution module updates the execution status of the task to be cancelled to "aborted".
[0159] In the embodiments of the present application, when the scheduling statuses of all sharding tasks of the task to be cancelled are already "aborted" or "completed", it indicates that both the sharding task A and the sharding task C of the task to be cancelled have terminated. Then the task execution module can update the execution status of the task to be cancelled to "aborted" (or referred to as "cancelled", "stopped", etc.) to achieve task status update.
[0160] For example, the task to be cancelled is as described above Figure 5Task 1 in execution as shown. Task 1 includes 4 shard tasks, namely shard 0, shard 1, shard 2, and shard 3. After the task execution module successfully aborts shard 1 and shard 2, it updates the execution status of shard 1 and shard 2 to aborted (see Figure 10 in 0-S, 1-S, Figure 10 where S in indicates aborted). After that, when shard 2 and shard 3 are also successfully aborted, the execution status of shard 2 and shard 3 is also aborted (see Figure 11 in 2-S, 3-S), correspondingly, the execution status of Task 1 is also updated to aborted.
[0161] S607. The task scheduling module updates the scheduling status of the task to be cancelled to successful.
[0162] Among them, since abortion is also the final state of a task, therefore, when determining that the execution status of a task is aborted, the scheduling status of the task to be cancelled can be updated to successful, realizing the update of the task scheduling status. Exemplarily, the task scheduling module can determine that the execution status of the task is aborted through a post-processing function, such as a post-listener. Briefly speaking, the task execution module can notify the task scheduling module that the execution status of the task to be cancelled is updated to aborted.
[0163] In some embodiments, the process of aborting a task to be cancelled when the execution status of the task to be cancelled is in execution is introduced above. Additionally, when the execution status of the task to be cancelled is to be executed, it indicates that the task to be cancelled has not started execution yet. Therefore, the task cancellation module can directly cancel the task to be cancelled, and then the task cancellation module can directly update the execution status of the task to be cancelled to aborted and update the scheduling status of the task to be cancelled to successful.
[0164] In some embodiments, when the execution status of the task to be cancelled is completed, it indicates that all shard tasks of the task to be cancelled have been executed. Therefore, the task cancellation module determines that the above task to be cancelled cannot be cancelled.
[0165] S608. The task scheduling module clears the resources occupied by the task to be cancelled.
[0166] Exemplarily, when the execution status of the task to be cancelled is aborted, it indicates that the task to be cancelled is successfully aborted. That is to say, the shard tasks of the task to be cancelled have all been aborted or completed. Therefore, the task scheduling module can release the resources occupied by the task to be cancelled, such as caches or other resources, etc., to complete the task abortion operation, avoid unnecessary occupation of system resources, and ensure the success of the scheduling of the next task. Similar to the previous text, the clearing of the resources described in S608 is also an optional step.
[0167] In some embodiments, when aborting the shard task A of the task to be cancelled, a situation where the abortion fails may occur. Therefore, when the shard task A abortion fails, the task cancellation module (such as the task cancellation module through the task execution module) may set the execution status of the task to be cancelled to abortion failure. Optionally, when the task scheduling module monitors through the post - listener that the execution status of the task to be cancelled is abortion failure, the task scheduling module may set the scheduling status of the task to be cancelled to abortion failure, thus ensuring the timely and correct update of the task status. Or, the task scheduling module does not change the scheduling status.
[0168] In some embodiments, for a task with a status of aborted, the task scheduling module may re - process the task according to user requirements. For example, if a continue - processing request input by the user is received, then the task execution module may continue to execute the corresponding shard tasks according to the status of each shard task of the task (such as the scheduling status, execution status), thereby avoiding repeating the execution of the already - completed status. This process is similar to the process of re - scheduling an abnormal task described below.
[0169] In addition, the task execution module may also record the execution situation of each data block of the shard task. When it is necessary to continue executing the aborted shard task, the task execution module may directly execute the unexecuted data blocks and avoid executing the already - executed data blocks. That is to say, when it is necessary to abort the shard task, the task execution module may use the currently - executed data block as the breakpoint. When subsequently continuing to execute the shard task, the task execution module executes the data blocks after the breakpoint.
[0170] In the embodiments of the present application, essentially, the above task cancellation is actually implemented based on Elastic-job and Spring Batch. Spring Batch sends a stop signal through the jobExecution.stop() method to indicate stopping the job and the task to be cancelled, and stops executing the relevant steps (StepExecution). After that, when the task to be cancelled is a multi-shard task and the execution status of the task to be cancelled is in execution, Spring Batch sets the execution status of the shard task A of the task to be cancelled to aborted, and only executes the currently executed data block. After the current data block is executed, Spring Batch marks the shard task A as aborted. When the execution status of all shard tasks of the task to be cancelled is completed or aborted, it indicates that the task to be cancelled has been successfully aborted. Therefore, the execution status of the task to be cancelled can be updated to aborted. Correspondingly, Elastic-job can update the scheduling status of the task to be cancelled to successful through a post listener, realizing flexible cancellation of the task. In addition, by effectively combining Elastic-job and Spring Batch, the task cancellation function solves the problems of real-time and consistency of task cancellation and status update. Of course, similar to the previous text, Elastic-job and Spring Batch are only exemplary frameworks, and the present application can also implement the task cancellation function described in the present application through other frameworks. That is to say, the present application does not limit the implementation framework of the task cancellation function.
[0171] In some embodiments, the distributed task scheduling system also supports an abnormal task handling function. To implement the abnormal task handling function, the abnormal tasks can be first identified to reschedule the abnormal tasks. Among them, the identification of abnormal tasks can be implemented by an abnormal task identification module. The rescheduling of abnormal tasks can be implemented by an abnormal task scheduling module. First, the identification of abnormal tasks will be introduced below. Exemplarily, the abnormal task identification module can identify abnormal tasks during task execution. The operation process executed by the abnormal task identification module can be mainly divided into two parts, namely abnormal task detection and abnormal task marking. As Figure 12 shown, this operation process (i.e., the identification process of abnormal tasks) can include:
[0172] S401. The abnormal task identification module periodically queries the tasks whose execution status is in execution.
[0173] Among them, the tasks whose execution status is in execution represent the currently running tasks. The abnormal task identification module regularly detects all running tasks to realize timely detection of the execution situation of tasks, so as to judge whether there are abnormal tasks based on the execution situation of tasks.
[0174] In addition, the abnormal task recognition module can detect the running tasks in real time instead of at irregular intervals. This application does not limit the timing of task detection.
[0175] S402. For each task with an execution status of "executing", the abnormal task recognition module obtains the planned shard quantity and the actual shard quantity of the task, and obtains the planned start time of the task.
[0176] In the embodiments of this application, for each running task, the abnormal task recognition module can obtain the planned shard quantity and the actual shard quantity of the task. In addition, optionally, the abnormal task recognition module can determine the running tasks through the scheduling status of the tasks. For example, the running task can be a task with a scheduling status of "running" (or "being scheduled").
[0177] Among them, the planned shard quantity of the task represents the standard shard quantity of the task, that is, the total number of shard tasks that the task needs to execute. The actual shard quantity of the task represents the number of shard tasks that have been scheduled for the task, and the number of shard tasks in the current state of the task is the target state 1. For example, if the state is the execution state, the target state 1 can be "executing", "completed" (success or failure). Another example is that if the state is the scheduling state, the target state 1 can be "running", "completed" (success or failure), etc.
[0178] Among them, the planned start time of the task represents the scheduling start time of the task, and can also represent the execution start time of the task.
[0179] S403. In the case where the actual shard quantity is different from the planned shard quantity, and the time difference 1 between the current time and the planned start time is greater than the preset time threshold 1, the abnormal task recognition module marks the task as an abnormal task.
[0180] Among them, the scheduling status of the abnormal task is abnormal.
[0181] In the embodiments of this application, for each task with an execution status of "executing", the abnormal recognition module can determine whether the actual shard quantity of the task is different from the planned shard quantity (that is, determine whether the actual shard quantity of the task is less than the planned shard quantity), and determine whether the time difference between the current time and the planned start time (that is, the planned execution time) is greater than the preset time threshold 1, so as to determine whether there are shard tasks that have not been scheduled due to timeout.
[0182] When the actual number of shards is different from the planned number of shards, and the time difference 1 (or the first time difference) between the current time and the planned start time is greater than the preset time threshold 1 (or the first preset time threshold), it indicates that there are shard tasks that have timed out and not been processed in the shard tasks of this task, and the timeout is excessive. Therefore, the abnormal task recognition module determines that this task is an abnormal task and marks this task as an abnormal task, realizing the detection and marking of abnormal tasks.
[0183] Whereas when the actual number of shards is the same as the planned number of shards, or the time difference 1 between the current time and the planned start time is less than or equal to the preset time threshold 1, it indicates that there are no unprocessed shard tasks in the shard tasks of this task, or the shard tasks have not timed out excessively. Therefore, the abnormal recognition module can determine that this task is a normal task and does not need to be marked.
[0184] Optionally, the above process of marking abnormal tasks may include that when it is determined that a task is an abnormal task, the abnormal recognition module can update the scheduling status of this task to an abnormal status. Optionally, the abnormal recognition module can update the scheduling status of this task to an abnormal status by calling a post listener to indicate that this abnormal task is a scheduling abnormality.
[0185] For example, the terminal device filters out tasks with an execution status of running. This task includes tasks such as Figure 13 Task 1 as shown and Figure 14 Task 2 as shown. Figure 13 The planned execution time of Task 1 as shown is 2024-09-23 08:00:00, the number of task shards (i.e., the planned number of shards, or the standard shard data) is 4, and the execution status of the shard tasks is 0-T, 1-T, indicating that the execution status of shard 0 is successful and the execution status of shard 1 is successful. The abnormal recognition module can determine that shards 2 and 3 of Task 1 have not been executed, that is, shards 2 and 3 have not been processed. Correspondingly, the actual number of shards of Task 1 is 2. Since the actual number of shards of Task 1 is 2, which is different from the number of task shards (i.e., 4), the abnormal recognition module can determine that the actual number of shards of Task 1 is different from the number of task shards.
[0186] And the time difference 1 between the current time of the abnormal recognition module and the planned execution time of Task 1 (i.e., 2024-09-23 08:00:00). When the actual number of shards of Task 1 is different from the number of task shards and this time difference 1 is greater than the preset time threshold 1, it indicates that Task 1 is an abnormal task. Therefore, the abnormal recognition module can update the scheduling status of Task 1 to an abnormal status (see the abnormality shown in Figure 15 ).
[0187] Similarly, Figure 14The planned execution time of Task 2 shown is 2024-09-23 08:00:00, the number of task shards is 1, and Task 2 actually does not require sharding processing. The task running execution status is to be executed, indicating that Task 2 has not started processing, that is, the actual number of shards of Task 2 is 0.
[0188] And, the time difference 1 between the current time of the anomaly recognition module and the planned execution time of Task 2 (i.e., 2024-09-23 08:00:00). When the actual number of shards of Task 2 is different from the number of task shards, and this time difference 1 is greater than the preset time threshold 1, it indicates that Task 2 is an abnormal task. Therefore, the anomaly recognition module can update the scheduling status of Task 2 to the abnormal status (see Figure 16 the anomaly shown).
[0189] The abnormal tasks introduced above are mainly for tasks with scheduling anomalies. This scheduling anomaly can not only include the anomaly of tasks that should be scheduled but are not scheduled in time as described above, but may also include duplicate scheduling, that is, the anomaly that the shard tasks of the task are repeatedly scheduled. However, the duplicate scheduling anomaly can be solved through the uniqueness verification of task shards. The duplicate scheduling anomaly is not the focus of this application and is only briefly introduced here.
[0190] In some embodiments, after determining the abnormal task, a detailed log can also be generated to facilitate subsequent analysis and debugging by related devices. The log records one or more of the abnormal reasons of the abnormal task, the status of the abnormal task (i.e., the scheduling status, the execution status), and the status of the shard tasks of the abnormal task.
[0191] In some embodiments, after determining that the task is an abnormal task, the anomaly recognition module can first add the task to the abnormal task list. After that, for each abnormal task in the abnormal task list, the anomaly recognition module can update the status of the abnormal task (i.e., the scheduling status, the execution status).
[0192] The above introduced the recognition process of abnormal tasks. After determining the abnormal tasks with scheduling anomalies, the abnormal task scheduling module can also reschedule the abnormal tasks to achieve the recovery of the abnormal tasks, thereby improving the reliability of task scheduling. The operation process executed by the abnormal task scheduling module can be mainly divided into two parts, namely abnormal task scheduling and abnormal task status update. The following will be combined with Figure 17 to introduce this operation process in detail, that is, to introduce the scheduling process of abnormal tasks.
[0193] S501. The abnormal task scheduling module obtains the task to be scheduled.
[0194] Among them, the task to be scheduled refers to a task with a scheduling status of to-be-scheduled, that is, a task that has not started scheduling. The task to be scheduled may include tasks whose planned execution time is later than the current time. Of course, the task to be scheduled may also include tasks whose planned execution time is earlier than the current time but have not started scheduling yet. Optionally, the task to be scheduled can also be referred to as a task to be executed, that is, a task with an execution status of to-be-executed.
[0195] In some embodiments, the abnormal task scheduling module can look up the task to be scheduled from the task table. The task table records information of at least one task, and this information includes one or more of the planned execution time, the scheduling status of the task, or the execution status of the task. Of course, this information can also include other configuration information, such as the task identification number (identity document, id) of the task as shown above, the execution status of the task, the execution status of the sharded task, etc. In addition, it can be understood that Figure 13 the configuration parameters shown above are only examples, and the configuration parameters can be set according to requirements. For example, the configuration parameters can also include the scheduling status of the task, the scheduling status of the sharded task of the task, etc. Figure 13
[0196] S502. The abnormal task scheduling module obtains abnormal tasks with an abnormal scheduling status.
[0197] S503. Based on the order of the planned execution times of each task in the task to be scheduled and the abnormal task, the abnormal task scheduling module sorts the task to be scheduled and the abnormal task to obtain the sorted tasks.
[0198] S504. The abnormal task scheduling module sends the sorted tasks to the task scheduling module.
[0199] In the embodiments of the present application, the abnormal task scheduling module can sort the abnormal tasks (or referred to as the first abnormal tasks) and the tasks to be scheduled (that is, the tasks that need to be scheduled) in the order of the planned execution times to obtain the sorted tasks. Then, the abnormal task scheduling module sends the sorted tasks (or the information of the sorted tasks (such as identifiers such as task names, IDs, etc.)) to the task scheduling module, so that the task scheduling module can schedule each task in the sorted tasks in order, enabling the tasks with earlier planned execution times to be scheduled faster, ensuring the timeliness of task processing, and thus also ensuring that the abnormal tasks can be scheduled in a timely manner and improving the reliability of task scheduling. It should be understood that the abnormal task scheduling module can also sort the tasks in other ways, such as according to the priority of the tasks, the instructions of the user, etc.
[0200] In some embodiments, after the above abnormal task is rescheduled, the task scheduling module can update the status of the abnormal task in a timely manner (i.e., the scheduling status and the execution status). Exemplarily, the task scheduling module can update the execution status of the abnormal task according to the execution status of the shard tasks of the abnormal task. When the execution status of each shard task of the abnormal task is successful, the task scheduling module updates the execution status of the abnormal task to successful. When there is a shard task of the abnormal task whose execution status is failed, the task scheduling module updates the execution status of the abnormal task to failed. Similarly, the process of updating the scheduling status of the task is similar to the process of updating the execution status, which will not be elaborated here. In addition, the specific process of updating the status of the abnormal task can refer to the relevant content above.
[0201] Optionally, since the execution status of some shard tasks of the abnormal task may be successful, these shard tasks do not need to be rescheduled. Since the execution situation of the shard tasks is recorded when the shard tasks are executed, the task scheduling module can query the execution status of each shard task of the abnormal task. For each shard task, when the shard task is in a successful state, it indicates that the shard task has been successfully executed and does not need to be processed repeatedly. Therefore, the task scheduling module will not transparently transmit the data of the shard task to the corresponding target node. Specifically, reference can be made to the task shard uniqueness verification described above.
[0202] When the shard task is not in a successful state, it indicates that the shard task has not been successfully processed. Therefore, the task scheduling module can transparently transmit the data of the shard task to the corresponding target node, so that the target node can perform batch processing on the shard task based on the business logic code of the abnormal task to implement the execution of the shard task.
[0203] For example, as described above Figure 15 Task 1 shown is an abnormal task. After that, when Task 1 is rescheduled and executed, the task scheduling module determines that the execution statuses of shard 0 and shard 1 of Task 1 are both successful, and there is no need to reschedule shard 0 and shard 1 repeatedly, but only need to reschedule shard 2 and shard 3. When the execution statuses of shard 2 and shard 3 are successful, it indicates that all shard tasks of Task 1 have been successfully executed. Correspondingly, the task scheduling module can change the scheduling status of Task 1 from Figure 15 the abnormal status shown to Figure 18 the successful status shown.
[0204] In the embodiments of the present application, the accurate and rapid identification of abnormal tasks can be achieved through the status of tasks and their shard tasks (such as execution status, scheduling status), so that abnormal tasks can be scheduled in a timely manner, avoiding tasks from not being processed for a long time, ensuring the continuity of services, and thus ensuring the reliability of task execution. Moreover, based on the execution of the shard tasks of the abnormal tasks, the abnormal task scheduling module re-schedules the uncompleted shard tasks in a targeted manner, avoiding the repeated scheduling of shard tasks, and thus avoiding unnecessary resource waste.
[0205] In some embodiments, the abnormal task scheduling module or the abnormal task identification module can record the number of times a task is marked as an abnormal task, that is, the number of times the task is re-scheduled. When this number is greater than the preset number threshold, it indicates that the task has been re-tried and scheduled multiple times but still not successfully scheduled. Therefore, in order to avoid wasting system resources, the abnormal task scheduling module can stop re-scheduling this task, that is, abort the execution of the task and release the relevant resources.
[0206] It should be noted that the above preset number threshold can be flexibly set according to requirements. In addition, other preset values involved in the present application can also be set according to requirements.
[0207] In the embodiments of the present application, the scheduling of abnormal tasks through the sorted tasks introduced above is only an example of the retry policy. This retry policy can also be other policies (such as node failover, compensation execution, etc.) to meet the scheduling requirements in different scenarios. The abnormal task scheduling module can trigger the retry mechanism for the abnormal tasks according to the retry policy corresponding to the scenario of the abnormal tasks, realize the re-scheduling of the tasks, and attempt to recover the abnormal tasks. And when the abnormal tasks cannot be recovered through the retry mechanism, the execution of the abnormal tasks can be aborted and the resources occupied by the abnormal tasks can be released to avoid unnecessary occupation of resources.
[0208] In some embodiments, tasks may not only have the above-mentioned scheduling abnormalities, but also have normal scheduling but abnormal execution. When the execution time of the task is greater than or equal to the preset execution duration, it indicates that the task execution is too long and the task may have abnormal execution. Then the task execution module can output an abnormal prompt message to enable relevant personnel to determine whether to abort the task, so that after receiving the task cancellation request input by the user, the task is aborted. Or, perform other operations according to the abnormal handling policy corresponding to the preset execution abnormality. This abnormal handling policy indicates that when the execution time of the task is greater than or equal to the preset execution duration, an abort operation is performed. Then this other operation can be an abort operation.
[0209] For example, for a multi-slice task (i.e., a task including multiple slice tasks), if the execution time of a slice task is greater than or equal to the preset execution time 1, the abnormality identification module can determine that the multi-slice task is executed too long. The execution time of the slice task can be the time from when the execution state of the slice task is switched to being executed to the current time.
[0210] For a single-slice task, when the execution time of the single-slice task is greater than or equal to the preset execution time 1, the abnormality identification module determines that the execution of the single-slice task is too long.
[0211] In some embodiments, the distributed task scheduling method (or task processing method) provided in the present application can be applied to data processing and analysis scenarios, such as large-scale data cleaning, transformation and loading (ETL) tasks, regular data statistics and analysis tasks, such as daily and monthly report generation, etc.
[0212] Alternatively, it can be applied to bill processing scenarios, such as bill generation and settlement tasks (e.g. regular generation of water, electricity, and gas bills), financial report generation and regular settlement, etc.
[0213] Alternatively, it can be applied to inventory management scenarios, such as inventory data synchronization.
[0214] Alternatively, it can be applied to automated operation and maintenance scenarios, such as regularly checking and cleaning log files, batch updating configuration files and patches, etc.
[0215] Alternatively, it can be applied to large-scale file processing scenarios, such as distributed file transfer tasks, file compression, decompression and archiving, etc.
[0216] Alternatively, it can be applied to Internet service scenarios, such as the collection and processing of user behavior logs, large-scale content review (such as image and video review), etc.
[0217] Alternatively, it can be applied to industrial IoT scenarios, such as regular collection and analysis of equipment data, monitoring of production line status and task allocation.
[0218] Alternatively, it can be applied to intelligent recommendation and advertising delivery scenarios, such as data calculation and model update of user recommendation algorithms, optimization of advertising delivery strategies, etc.
[0219] Alternatively, it can be applied to logistics and distribution management scenarios, such as dynamic optimization of distribution routes, order batch scheduling and distribution, etc.
[0220] In general, the distributed task scheduling method provided in this application can be widely used in various scenarios that require dynamic task allocation, efficient concurrent processing, anomaly detection and recovery, especially in data-intensive and task-complex businesses, and can significantly improve task scheduling and execution efficiency and enhance system reliability and stability.
[0221] In the embodiments of the present application, the distributed scheduling system not only records the basic information about tasks (such as task name, identification number, planned execution time, etc.), but also records detailed information such as the creator, update time, sharding status, etc., which is convenient for dynamic monitoring and problem troubleshooting of tasks.
[0222] In the embodiments of the present application, the present application provides a system that combines a distributed scheduling framework and a distributed execution framework (or referred to as a task processing system). This system not only utilizes the scheduling capabilities of the distributed scheduling framework but also combines the execution capabilities of the distributed execution framework (such as batch processing capabilities). Since the distributed scheduling framework and the distributed execution framework themselves can improve performance, the combination of the two can effectively improve the performance of the system. In addition, the reliability of the system can be ensured through the identification and scheduling of abnormal tasks. The flexibility of the system can be ensured through task termination.
[0223] The above Elastic-job is only an example of a distributed scheduling framework, and Spring Batch is only an example of a distributed execution framework. This framework can also be other frameworks that can achieve the corresponding functions. For example, Elastic-job can be replaced with frameworks such as Quartz and XXL-job, and the present application does not limit the specific framework. Correspondingly, the solutions provided by the present application, such as the related functions provided by Elastic-job and Spring Batch on which the above task scheduling, task execution, task cancellation, identification of abnormal tasks, and rescheduling depend, such as the above-mentioned pre-listener, post-listener, stop signal, etc., are also only examples. The present application can implement this function through other code that can achieve this function.
[0224] It can be understood that the collection of business data involved in the present application is authorized and consented by the user. In addition, the operations performed by the above modules are only examples, and the above operations can also be performed by other modules. The present application does not limit the specific modules of the above operations.
[0225] It should be noted that the execution order of the steps introduced above is only an example, and the present application does not limit the order of execution of the above steps. This execution order can be set according to requirements. In addition, the modules involved in the present application can be deployed on the same node or on different nodes, and the present application does not limit them.
[0226] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of modules. It can be understood that the operations performed by the above modules are actually implemented by a distributed scheduling system, that is, by devices (or referred to as electronic devices, nodes, etc.). In order to implement the above functions, an electronic device includes corresponding hardware structures and / or software modules for executing each function. Combining the units and algorithm steps of each example described in the embodiments disclosed in the present application, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software.
[0227] Whether a certain function is executed in the way of hardware or computer driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described function for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiments of the present application.
[0228] Optionally, the electronic device may be a server. The number of electronic devices may be one or more. The electronic device may execute the operations performed by some or all of the above modules.
[0229] In some embodiments, the present application provides an electronic device, which includes a memory and one or more processors; the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device is caused to execute the task processing method as described above.
[0230] In some embodiments, the present application provides a chip, which includes a communication interface and at least one processor:
[0231] The communication interface is used for inputting and / or outputting signaling or data;
[0232] The at least one processor is used to execute a computer program to implement the task processing method as described above.
[0233] In some embodiments, the present application provides a computer-readable storage medium, which includes computer instructions. When the computer instructions run on an electronic device, the electronic device is caused to execute the task processing method as described above.
[0234] In some embodiments, the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device is caused to execute the task processing method as described above.
[0235] In the embodiments of the present application, the functional modules of the electronic device can be divided according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of units in the embodiments of the present application is illustrative, and is only a logical function division. There may be other division methods in actual implementation.
[0236] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0237] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0238] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it can be located in one place, or it can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0239] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0240] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0241] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.
Claims
1. A task processing method, characterized in that, Applied to a task processing system that combines a distributed scheduling framework and a distributed execution framework, the method includes: In response to the scheduling of a first task, perform sharding processing on the first task through the distributed scheduling framework to obtain multiple first sharded tasks of the first task; For each of the first sharded tasks, assign the first sharded task to a first target node, and update the scheduling status corresponding to the first sharded task through the preprocessing function; Through the first target node, combine batch processing of the first sharded task through the distributed execution framework to update the execution status of the first sharded task; When the execution statuses of all the first sharded tasks indicate completion, set the scheduling status of the first task to completion through the postprocessing function of the distributed scheduling framework, where the completion is either successful or failed; Obtain a task cancellation request; wherein the task cancellation request is used to trigger the abortion of a second task; According to the execution status of the second task, abort the second task and update the execution status of the second task to aborted; wherein the second sharded tasks of the second task are divided into multiple data sets, and when the execution status of the second sharded tasks of the second task is in execution, continue to execute the data in the currently executed data set of the second sharded task, and after all the data in the currently executed data set are executed, abort the execution of the second sharded task to stop executing the remaining unexecuted data sets of the second sharded task; Determine through the postprocessing function that the execution status of the second task is aborted, and update the scheduling status of the second task to successful; For each task with an execution status of in execution, obtain the standard shard quantity and the actual shard quantity of the task, and obtain the planned execution time of the task; wherein the actual shard quantity represents the number of sharded tasks of the task that are scheduled; When the standard shard quantity is different from the actual shard quantity, and the first time difference between the current time and the planned execution time is greater than the first preset time threshold, mark the task as a first abnormal task; wherein the scheduling status of the first abnormal task is abnormal.
2. The method according to claim 1, wherein The aborting the second task according to the execution status of the second task includes: When the execution status of the second task is in execution, for each second sharded task of the second task, if the execution status of the second sharded task is in execution, continue to execute the current data set of the second sharded task through the second target node where the second sharded task is located; wherein the current data set belongs to one of the multiple data sets obtained by dividing the second sharded task during batch processing of the second sharded task, and at least one of the multiple data sets includes the data in the second sharded task; After the current data set is executed, stop performing business processing on new data sets of the second sharded task.
3. The method according to claim 2, wherein After the stopping of performing business processing on new data sets of the second sharded task, the method further includes: Update the execution status of the second shard task to aborted; The step of updating the execution status of the second task to aborted includes: When the execution statuses of all the second shard tasks are terminal states, update the execution status of the second task to aborted; wherein the terminal states include one or more of success, failure, or aborted.
4. The method according to claim 2 or 3, characterized in that, After the execution status of the second task is updated to aborted, the method further includes: In response to the second task being rescheduled, when the execution status of the second shard task of the second task is aborted, perform business processing on the dataset of the second shard task that has not undergone business processing; When the execution status of the second shard task of the second task is success or failure, do not perform business processing on the second shard task.
5. The method according to claim 1, wherein The method further includes: Sort the tasks to be scheduled and the first abnormal tasks to obtain the sorted tasks; Schedule the sorted tasks sequentially.
6. The method according to claim 5, wherein The step of sorting the tasks to be scheduled and the first abnormal tasks to obtain the sorted tasks includes: Sort the tasks to be scheduled and the first abnormal tasks based on the chronological order of the planned execution times of each task in the tasks to be scheduled and the first abnormal tasks.
7. The method according to claim 5 or 6, characterized in that The step of scheduling the sorted tasks sequentially includes: When the execution status of the shard task of the scheduled task is success or failure, do not perform business processing on the shard task of the scheduled task.
8. A task processing system, characterized in that, The task processing system is a system that combines a distributed scheduling framework and a distributed execution framework, and the task processing system executes the task processing method according to any one of claims 1 to 7.
9. An electronic device, characterized in that, The electronic device includes a memory and one or more processors; the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device executes the task processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Includes computer instructions, when the computer instructions run on an electronic device, the electronic device executes the task processing method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the task processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Spring batch processing job webpage maintenance method and system
CN112882767A