Distributed Batch Task Scheduling Method and System

By splitting into multi-level subtasks in task scheduling and using independent JVM processes, the resource waste and performance degradation caused by resident processes are solved, and efficient and flexible task scheduling is achieved.

CN115469989BActive Publication Date: 2025-07-29IND BANK CO +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211329286.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-27
Publication Date
2025-07-29
Estimated Expiration
2042-10-27

AI Technical Summary

Technical Problem

In the existing task scheduling technology, the resident process needs to occupy memory for a long time, resulting in waste of resources. When the banking business forms are changing and the business volume grows too fast, the batch processing performance declines.

Method used

The distributed batch task scheduling method is adopted, and the task is split into subtasks at multiple levels. Each task corresponds to an independent JVM process. The scheduling service is resident, and the batch task exits after execution. The memory limit and number of task processes are dynamically adjusted, and dynamic splitting and high concurrency are supported.

Benefits of technology

Improve batch processing performance in business change and high concurrency, avoid resource crowding and JVM crashes, reduce resource waste, and achieve efficient and flexible task scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115469989B_ABST
    Figure CN115469989B_ABST
Patent Text Reader

Abstract

The present invention provides a distributed batch task scheduling method and system, including: obtaining platform front-end data; reading and registering a job schedule in the front-end data to obtain a job flow, and obtaining a task generation parent task flow table associated with the job; scanning the task flow table, obtaining executable tasks and adding optimistic locks, and calling a task execution method to execute the tasks; the platform front-end data includes: jobs, tasks, and job schedules. By adopting the method of calling a task splitting interface, the present invention has low coupling between the scheduler and the business, improves the processing performance of batch processing when the business form changes frequently and the business volume grows too fast. In addition, by using the method of splitting subtasks with the same task number as the task, the scheduling only needs to be configured to the task, solving the problem that common tasks in the current market need to be configured to the finest granularity and do not support dynamic splitting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of task scheduling, and specifically, to a distributed batch task scheduling method and system. Background Art

[0002] Task scheduling refers to the process in which the system automatically completes specific tasks at a specified moment as agreed. With task scheduling, more manpower can be liberated to be automatically executed by the system. Usually, the task scheduling program is integrated into the application. For example, the coupon service includes a scheduling program for regularly issuing coupons, and the settlement service includes a task scheduling program for regularly generating reports. Due to the adoption of a distributed architecture, a service often deploys multiple redundant instances to run our business. Running task scheduling in such a distributed system environment is called distributed task scheduling.

[0003] Patent document CN112685184A discloses a method for distributed task scheduling, a task scheduling platform, and a task executor. The task scheduling platform communicates with the task executor, and the task executor is set in a cloud server and / or an edge device. The method includes: receiving a heartbeat registration request sent by the task executor, where the heartbeat registration request includes the status information of the task executor; in response to receiving the heartbeat registration request, generating a task mailbox corresponding to the task executor; and dispatching the tasks in the task pool to the task mailbox corresponding to the task executor according to the status information of the task executor.

[0004] However, in the existing task scheduling technology, one or more tasks are for one service, and one service corresponds to one resident process. The resident process needs to occupy a certain range of memory for a long time, and more hardware resources are required during deployment, resulting in waste of resources. Moreover, the banking business forms are diverse and the business volume grows too fast, which will cause the performance of batch processing to decline. Summary of the Invention

[0005] Aiming at the defects in the prior art, the purpose of the present invention is to provide a distributed batch task scheduling method and system.

[0006] According to a distributed batch task scheduling method provided by the present invention, it includes:

[0007] Step S1: Obtain the platform front-end data;

[0008] Step S2: Read the job schedule in the front-end data and register it to obtain a job stream, and obtain a parent task stream table associated with the job;

[0009] Step S3: Scan the task stream table, obtain executable tasks and add an optimistic lock, and call the task execution method to execute the tasks;

[0010] The front-end data of the platform includes: jobs, tasks, and job schedules.

[0011] Preferably, step S2 includes:

[0012] Step S2.1: Regularly scan the job schedule to read the processable job plans;

[0013] Step S2.2: Read the task table according to the corresponding job number, obtain all tasks under the current job and add an optimistic lock; among them, the optimistic lock includes information on updating the executed date and status;

[0014] Step S2.3: Register the job and all tasks under the current job to generate a job flow table and a task flow table, where the tasks in the task flow table are parent tasks.

[0015] Preferably, step S3 includes:

[0016] Step S3.1: Regularly scan the task flow table, read the executable tasks in the task flow table and add an optimistic lock to the executable tasks, and trigger step S3.2 after the locking is successful;

[0017] Step S3.2: Check the task execution conditions corresponding to the executable tasks, and thus enter the task execution method. If the execution method is task execution, execute the task; if the execution method is to split sub-tasks, trigger step S3.3;

[0018] Step S3.3: The application defines the splitting basis according to the actual business scenario, and the platform splits the task according to the set splitting basis, splitting a parent task into multiple sub-tasks with a level of 1;

[0019] Step S3.4: Determine whether the sub-tasks at the current level still meet the splitting conditions. If not, execute the current task; if so, continue to execute the sub-task splitting method, and increment the task level corresponding to the currently split sub-tasks by 1;

[0020] Repeat triggering step S3.4 until the current sub-tasks do not meet the splitting conditions and then execute the current task.

[0021] Preferably, each task pulled up by the scheduling service corresponds to an independent JVM process;

[0022] The scheduling service and batch tasks are deployed on each server. The scheduling service runs permanently, and the batch tasks exit after execution;

[0023] Set JVM startup parameters to dynamically adjust the upper limit of memory required for the batch, and control the number of concurrent task processes by setting the upper limit of the number of task processes, where the upper limit of the number of task processes includes the single-machine upper limit and the cluster upper limit.

[0024] Preferably, the task level of the parent task is 0, where the task numbers of the subtasks and the parent task are the same;

[0025] The splitting conditions determine the business direction, and the tasks at each level correspond to different splitting conditions.

[0026] A distributed batch task scheduling system according to the present invention includes:

[0027] Module M1: Obtain platform front-end data;

[0028] Module M2: Read the job schedule in the front-end data and register it to obtain a job flow, and obtain a parent task flow table associated with the job;

[0029] Module M3: Scan the task flow table, obtain executable tasks and add an optimistic lock, and call the task execution method to execute the tasks;

[0030] The platform front-end data includes: jobs, tasks, and job schedules.

[0031] Preferably, Module M2 includes:

[0032] Module M2.1: Periodically scan the job schedule and read the processable job plans;

[0033] Module M2.2: Read the task table according to the corresponding job number, obtain all tasks under the current job and add an optimistic lock; among them, the optimistic lock includes information on updating the executed date and status;

[0034] Module M2.3: Register the job and all tasks under the current job, generate a job flow table and a task flow table, where the tasks in the task flow table are parent tasks.

[0035] Preferably, Module M3 includes:

[0036] Module M3.1: Periodically scan the task flow table, read the executable tasks in the task flow table and add an optimistic lock to the executable tasks, and trigger Module M3.2 after successful locking;

[0037] Module M3.2: Check the task execution conditions corresponding to the executable tasks, so as to enter the task execution system. If the execution system is task execution, execute the tasks; if the execution system is to split subtasks, trigger Module M3.3;

[0038] Module M3.3: The application defines the splitting basis according to the actual business scenario, and the platform splits the tasks according to the set splitting basis, and splits a parent task into multiple subtasks with a level of 1;

[0039] Module M3.4: Determine whether the subtasks at the current level still meet the splitting conditions. If not, execute the current task. If so, continue to execute the subtask splitting system, and increment the task level corresponding to the currently split subtasks by 1;

[0040] Repeat triggering Module M3.4 until the current subtasks no longer meet the splitting conditions and then execute the current task.

[0041] Preferably, each task pulled up by the scheduling service corresponds to an independent JVM process;

[0042] The scheduling service and batch tasks are deployed on each server. The scheduling service is resident, and the batch tasks exit after completion;

[0043] Set JVM startup parameters to dynamically adjust the upper memory limit required for the batch. Control the task concurrent process number by setting the upper limit of the number of task processes, where the upper limit of the number of task processes includes the single-machine upper limit and the cluster upper limit.

[0044] Preferably, the task level of the parent task is 0, and the task numbers of the subtasks and the parent task are the same;

[0045] The splitting conditions determine the business direction, and the tasks at each level correspond to different splitting conditions.

[0046] Compared with the prior art, the present invention has the following beneficial effects:

[0047] 1. By adopting the method of calling the task splitting interface, the scheduling program and the business are loosely coupled, improving the processing performance of batch processing when the business form is changeable and the business volume grows too fast. In addition, by using the method that the split subtasks and the tasks have the same task number, the scheduling only needs to be configured to the task, solving the problem that common tasks in the current market need to be configured to the finest granularity and do not support dynamic splitting.

[0048] 2. By adopting the process isolation mode, each task corresponds to an independent JVM process. Compared with the current market mode where multiple tasks of a resident process are written in one service, it avoids the defects of resource occupation between threads in the service, thread blocking caused by high concurrency, and JVM crashes caused by cumulative memory overflow of a single thread.

[0049] 3. By adopting the mode of deploying scheduling on each server and exiting after completion, the JVM exits when the task is completed. It does not need to occupy memory for a long time and does not require more hardware resources, reducing resource waste.

[0050] 4. By adopting the method of setting JVM startup parameters to dynamically adjust the upper memory limit required for a certain batch and controlling the task concurrent process number by setting the upper limit of the number of task processes, it can reasonably allocate resources between tasks. Brief Description of the Drawings

[0051] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non - restrictive embodiments with reference to the accompanying drawings:

[0052] Figure 1 It is a schematic flowchart of the present invention. Detailed Embodiments

[0053] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all fall within the protection scope of the present invention.

[0054] When the present invention performs batch processing operations, according to the characteristics of bank data, tasks are dynamically split into subtasks as needed. A task can be split at multiple levels according to actual situations. Between each task, a serial scheduling process can be formed through simple dependency relationships, or multiple scheduling processes such as parallel, aggregation, and subtask splitting can be formed through complex dependency relationships. A set of scheduling platforms and batch tasks are configured on each server, and tasks and subtasks can be scheduled and executed on any server.

[0055] Embodiment 1

[0056] A distributed batch task scheduling method provided by the present invention includes:

[0057] Step S1: Obtain platform front - end data, where the platform front - end data includes: jobs, tasks, and job schedules.

[0058] Step S2: Read the job schedule in the front - end data and register it to obtain a job stream, and obtain a parent - level task stream table associated with the job.

[0059] Specifically, step S2 includes:

[0060] Step S2.1: Regularly scan the job schedule and read the processable job plans;

[0061] Step S2.2: Read the task table according to the corresponding job number, obtain all tasks under the current job and add an optimistic lock; among them, the optimistic lock includes information on updating the executed date and status;

[0062] Step S2.3: Register the job and all tasks under the current job to generate a job stream table and a task stream table, where the tasks in the task stream table are parent - level tasks.

[0063] Further, when reading the job schedule, jobs that meet the conditions are registered. The jobs that meet the conditions include any one or more of the following: 1. Execute by day: whether the current time has reached the specified time point; 2. Execute by week: whether the current time has reached the corresponding time point of the specified cycle; 3. Specified date: whether the current time has reached the time point corresponding to the corresponding date of the current month; 4. Execute in a loop: a. whether the current time is within a time period of the loop time window, b. whether the time elapsed since the last execution end time of the job exceeds the loop interval.

[0064] Step S3: Scan the task schedule, obtain executable tasks and add optimistic locks, and call the task execution method to execute the tasks. Step S3 includes:

[0065] Step S3.1: Regularly scan the task schedule, read the executable tasks in the task schedule and add optimistic locks to the executable tasks. After the locking is successful, trigger Step S3.2;

[0066] Step S3.2: Check the task execution conditions corresponding to the executable tasks, and thus enter the task execution method. If the execution method is task execution, execute the task; if the execution method is to split sub-tasks, trigger Step S3.3;

[0067] Step S3.3: The application defines the splitting basis according to the actual business scenario. The platform splits the task according to the set splitting basis, and splits a parent task into multiple sub-tasks with a level of 1;

[0068] Step S3.4: Determine whether the sub-tasks at the current level still meet the splitting conditions. If not, execute the current task; if so, continue to execute the sub-task splitting method, and increment the task level corresponding to the currently split sub-tasks by 1;

[0069] Repeat triggering Step S3.4 until the current sub-tasks do not meet the splitting conditions and then execute the current task.

[0070] Specifically, the task level of the parent task is 0, and the task numbers of the sub-tasks and the parent task are the same; the splitting conditions determine the business direction and the tasks at each level correspond to different splitting conditions. For example, 1. Whether splitting is required; 2. Which business attribute to split by. For instance, there are 100 million pieces of data to be processed. The splitting conditions for the first-level tasks are: 1. Splitting must be performed; 2. Split by the business attribute "city where it belongs". The splitting conditions for the second-level tasks are: 1. If the data volume of "city where it belongs" is greater than 10 million, then split, otherwise do not split; when splitting, split according to the business attribute "data import date". And so on. By this scheduling method, only the task needs to be configured, solving the problem that common tasks in the current market need to be configured to the finest granularity and do not support dynamic splitting.

[0071] By adopting the method of calling the task splitting interface, the present invention realizes low coupling between the scheduler and the service. The framework provides a task splitting interface, and the application program can customize the splitting basis according to the actual business scenario, perform multi-level splitting, support parallel scheduling of tasks and subtasks, and achieve cross-machine and high-concurrency distributed batch processing. It effectively solves the problems of the ever-changing business forms, the too-fast growth of business volume, and the decline of batch processing performance in Industrial Bank. It is different from the scheduling platforms or frameworks commonly used in general data processing systems, such as Control-m.

[0072] Further, each task pulled up by the scheduling service of the present invention corresponds to an independent JVM process. By adopting the process isolation mode, each task corresponds to an independent JVM (Java Virtual Machine) process. Compared with the current market mode where multiple tasks are written in one service with a resident process, it avoids defects such as resource occupation between threads in the service, thread blocking caused by high concurrency, and JVM crashes caused by memory overflow of a single thread over time.

[0073] The scheduling service and batch tasks are deployed on each server. The scheduling service is resident, and the batch task exits after execution. Compared with the micro-service mode where one or more tasks correspond to one service and one service corresponds to one resident process, it solves the problem that the resident process needs to occupy a certain range of memory for a long time and requires more hardware resources during deployment.

[0074] Set the JVM startup parameters to dynamically adjust the upper limit of the memory required for the batch. Control the task concurrent process number by setting the upper limit of the task process number, where the upper limit of the task process number includes the single-machine upper limit and the cluster upper limit. By this method, (1) resources can be reasonably allocated between tasks, and the problem of process OOM caused by unreasonable resource allocation and data volume exceeding the critical value commonly seen in the current market can be solved. Specifically, setting the JVM startup parameters includes the following parameter items: By "-Xms" <size>"Set the initial Java heap size; via "-Xmx <size>”Set the maximum Java heap size; via "-Xmn <size>"Set the size of the young generation, which consists of the eden area + 2 survivor areas.

[0075] Example Two

[0076] The present invention also provides a distributed batch task scheduling system. Those skilled in the art can implement the distributed batch task scheduling system by executing the step process of the distributed batch task scheduling method, that is, the distributed batch task scheduling method can be understood as the preferred implementation manner of the distributed batch task scheduling system.

[0077] A distributed batch task scheduling system according to the present invention includes:

[0078] Module M1: Obtain platform front-end data, where the platform front-end data includes: jobs, tasks, and job schedules.

[0079] Module M2: Read the job schedule in the front-end data and register it to obtain a job stream, and obtain a parent task stream table associated with the job. Module M2 includes: Module M2.1: Periodically scan the job schedule and read the processable job plans; Module M2.2: Read the task table according to the corresponding job number, obtain all the tasks under the current job and add an optimistic lock; where the optimistic lock includes information for updating the executed date and status; Module M2.3: Register the job and all the tasks under the current job to generate a job stream table and a task stream table, where the tasks in the task stream table are parent tasks.

[0080] Module M3: Scan the task stream table, obtain executable tasks and add an optimistic lock, and call the task execution method to execute the tasks. Module M3 includes: Module M3.1: Periodically scan the task stream table, read the executable tasks in the task stream table and add an optimistic lock to the executable tasks. After the locking is successful, trigger Module M3.2; Module M3.2: Check the task execution conditions corresponding to the executable tasks, and thus enter the task execution system. If the execution system is task execution, execute the task; if the execution system is to split sub-tasks, trigger Module M3.3; Module M3.3: The application defines the splitting basis according to the actual business scenario, and the platform splits the task according to the set splitting basis, splitting a parent task into multiple sub-tasks with a level of 1; Module M3.4: Determine whether the sub-tasks at the current level still meet the splitting conditions. If not, execute the current task; if so, continue to execute the sub-task splitting system, and the task level corresponding to the currently split sub-tasks is incremented by 1; Repeat triggering Module M3.4 until the current sub-tasks do not meet the splitting conditions and then execute the current task.

[0081] Among them, the task level of the parent task is 0, where the task numbers of the subtasks and the parent task are the same, and the splitting condition determines the business direction, and the tasks at each level correspond to different splitting conditions.

[0082] Further, each task pulled up by the scheduling service of the present invention corresponds to an independent JVM process; the scheduling service and batch tasks are deployed on each server, the scheduling service is resident, and the batch tasks exit after execution; JVM startup parameters are set to dynamically adjust the upper limit of the memory required for the batch, and the concurrent process number of the tasks is controlled by setting the upper limit of the number of task processes, where the upper limit of the number of task processes includes the single-machine upper limit and the cluster upper limit.

[0083] Those skilled in the art know that in addition to implementing the system, device, and its various modules provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system, device, and its various modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same program. Therefore, the system, device, and its various modules provided by the present invention can be considered as a hardware component, and the modules included therein for implementing various programs can also be regarded as the structure within the hardware component; the modules for implementing various functions can also be regarded as both software programs for implementing the method and the structure within the hardware component.

[0084] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be arbitrarily combined with each other.< / size> < / size> < / size>

Claims

1. A distributed batch task scheduling method, characterized in that, It includes: Step S1: Obtain the platform front-end data; Step S2: Read the job schedule in the front-end data and register it to obtain the job flow, and obtain the task flow table of the parent tasks associated with the job; Step S3: Scan the task flow table, obtain the executable tasks and add an optimistic lock, and call the task execution method to execute the tasks; The platform front-end data includes: jobs, tasks, and job schedules; Step S2 includes: Step S2.1: Regularly scan the job schedule and read the processable job plans; Step S2.2: Read the task table according to the corresponding job number, obtain all the tasks under the current job and add an optimistic lock; among them, the optimistic lock includes information on updating the executed date and status; Step S2.3: Register the job and all the tasks under the current job to generate a job flow table and a task flow table, where the tasks in the task flow table are parent tasks; Step S3 includes: Step S3.1: Regularly scan the task flow table, read the executable tasks in the task flow table and add an optimistic lock to the executable tasks, and trigger Step S3.2 after the locking is successful; Step S3.2: Check the task execution conditions corresponding to the executable tasks, and thus enter the task execution method. If the execution method is task execution, execute the task; if the execution method is to split into sub-tasks, trigger Step S3.3; Step S3.3: The application defines the splitting basis according to the actual business scenario, and the platform splits the tasks according to the set splitting basis, and splits a parent task into multiple sub-tasks with a level of 1; Step S3.4: Determine whether the sub-tasks at the current level still meet the splitting conditions. If not, execute the current task; if so, continue to execute the sub-task splitting method, and the task level corresponding to the currently split sub-tasks is incremented by 1; Repeat triggering Step S3.4 until the current sub-tasks do not meet the splitting conditions and then execute the current task.

2. The distributed batch task scheduling method according to claim 1, wherein Each task pulled up by the scheduling service corresponds to an independent JVM process; The scheduling service and batch tasks are deployed on each server. The scheduling service is resident, and the batch tasks exit after execution; Set the JVM startup parameters to dynamically adjust the upper limit of the memory required for the batch, and control the task concurrency process number by setting the upper limit of the task process number, where the upper limit of the task process number includes the single-machine upper limit and the cluster upper limit.

3. The distributed batch task scheduling method according to claim 1, wherein The task level of the parent task is 0, where the task numbers of the sub-tasks and the parent task are the same; The splitting conditions determine the business direction and the tasks at each level correspond to different splitting conditions.

4. A distributed batch task scheduling system, characterized in that, It includes: Module M1: Obtain the platform front-end data; Module M2: Read the job schedule in the front-end data and register it to obtain the job flow, and obtain the task flow table of the parent tasks associated with the job; Module M3: Scan the task flow table, obtain the executable tasks and add an optimistic lock, and call the task execution method to execute the tasks; The platform front-end data includes: jobs, tasks, and job schedules; Module M2 includes: Module M2.1: Regularly scan the job schedule and read the processable job plans; Module M2.2: Read the task list according to the corresponding job number, obtain all tasks under the current job and add an optimistic lock; among them, the optimistic lock includes information for updating the executed date and status. Module M2.3: Register the job and all tasks under the current job, generate a job flow table and a task flow table, where the tasks in the task flow table are parent tasks. Module M3 includes: Module M3.1: Periodically scan the task flow table, read the executable tasks in the task flow table and add an optimistic lock to the executable tasks. After the locking is successful, trigger Module M3.

2. Module M3.2: Check the task execution conditions corresponding to the executable tasks, and thus enter the task execution system. If the execution system is task execution, execute the task; if the execution system is to split sub-tasks, trigger Module M3.

3. Module M3.3: The application program defines the splitting basis according to the actual business scenario. The platform splits the task according to the set splitting basis, and splits a parent task into multiple sub-tasks with a level of 1. Module M3.4: Judge whether the sub-tasks at the current level still meet the splitting conditions. If not, execute the current task; if so, continue to execute the sub-task splitting system, and increment the task level corresponding to the currently split sub-tasks by 1. Repeat triggering Module M3.4 until the current sub-tasks do not meet the splitting conditions and then execute the current task.

5. The distributed batch task scheduling system according to claim 4, characterized in that Each task pulled up by the scheduling service corresponds to an independent JVM process. The scheduling service and batch tasks are deployed on each server. The scheduling service is resident, and the batch tasks exit after execution. Set JVM startup parameters to dynamically adjust the upper limit of memory required for batches, and control the number of concurrent task processes by setting the upper limit of the number of task processes. Among them, the upper limit of the number of task processes includes the single-machine upper limit and the cluster upper limit.

6. The distributed batch task scheduling system according to claim 4, wherein The task level of the parent task is 0, where the task numbers of the sub-tasks and the parent task are the same. The splitting conditions determine the business direction, and the tasks at each level correspond to different splitting conditions.

Citation Information

Patent Citations

  • Distributed task scheduling method, task scheduling platform and task executor

    CN112685184A

  • Task scheduling method and device

    CN108563502A

  • A distributed task scheduling method and system

    CN108958920A