Task scheduling processing method and apparatus, and storage medium and electronic device
By obtaining and judging the completeness of the result set in task scheduling, the target task is allowed to run in advance and schedule data in batches, the problem of downstream tasks waiting for upstream tasks to be completed is solved, and the task scheduling and execution efficiency is improved.
Patent Information
- Application Number
- PCT/CN2024/135278
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-11-28
- Publication Date
- 2025-05-08
AI Technical Summary
In task scheduling, downstream tasks need to wait until all upstream tasks are completed before they can start running, especially when a few upstream tasks are executed slowly, this will affect the overall operation efficiency of the task.
Before executing the target task, obtain the result set information required by the target task from the metadatabase, determine the complete result set and incomplete result sets in multiple result sets, and judge whether the target task is allowed to run in advance based on the incomplete result set. Allows the target task to schedule data from the complete result set and incomplete result set to the target task when it runs in advance, and schedule incremental data from the incomplete result set to the target task during the execution of the upstream task to which the incomplete result set belongs, until the incomplete result set changes to the complete result set.
By allowing the target task to run in advance and scheduling data in batches before the incomplete result set is changed to complete, the problem of downstream tasks waiting for the upstream tasks to be completed is solved, and the task scheduling efficiency and overall execution efficiency are improved.
Smart Images

Figure CN2024135278_08052025_PF_FP_ABST
Abstract
Description
Task scheduling processing method, device, storage medium and electronic device
[0001] Related applications
[0002] This application claims priority to Chinese patent application number 2023114359150, filed on October 31, 2023, entitled “Task scheduling and processing method, device, storage medium and electronic device,” the entire text of which is hereby incorporated by reference. Technical Field
[0003] The present application relates to the field of task scheduling, and more specifically, to a task scheduling processing method, device, storage medium, and electronic device. Background Art
[0004] In the field of big data, task scheduling platforms are widely used. An efficient scheduling platform can bring huge benefits to enterprises.
[0005] In actual production, the task scheduling process is considered a DAG (Directed Acyclic Graph). Since most tasks depend on the results of upstream tasks, when scheduling tasks, downstream tasks must wait until all upstream tasks are completed before they can start running. However, when most upstream tasks of a downstream task to be executed are fast, while a few upstream tasks are slow, to ensure the data integrity of the data required for task execution, the downstream task to be executed must wait until the slowest upstream task is completed before it can be executed. This affects the timeliness of the downstream task to be executed, and thus affects the overall execution efficiency of multi-layer tasks. Summary of the Invention
[0006] The present application provides a task scheduling and processing method, device, storage medium and electronic device to solve the problem in related technologies that downstream tasks must wait until all upstream tasks are completed before they can start running, and the slow execution of a few upstream tasks will affect the overall running efficiency of the tasks.
[0007] According to one aspect of the present application, a method for scheduling tasks is provided. The method includes: before executing a target task, obtaining result set information required by the target task from a metadata database, wherein the result set information includes the completeness of multiple result sets, and the multiple result sets refer to the result sets of upstream tasks that the target task depends on, and each upstream task updates the data in the corresponding result set during operation; determining complete result sets and incomplete result sets in the multiple result sets based on the result set information, and judging whether to allow the target task to run in advance based on the incomplete result set; if the target task is allowed to run in advance, scheduling the data in the complete result set and the incomplete result set to the target task, and during the execution of the upstream task to which the incomplete result set belongs, scheduling the incremental data in the incomplete result set to the target task until the incomplete result set is changed to a complete result set, wherein the target task executes the task in batches of data during the process of receiving data.
[0008] Optionally, determining whether to allow the target task to run early based on the incomplete result sets includes: determining whether the number of incomplete result sets is less than or equal to a preset number; if the number of incomplete result sets is less than or equal to a preset number, determining whether each incomplete result set carries a preset label, wherein the preset label indicates that the target task supports executing tasks in batches based on data received in batches; if all incomplete result sets carry preset labels, determining that the target task is allowed to run early.
[0009] Optionally, each upstream task that the target task depends on includes at least one result set. When an upstream task includes more than two result sets, the result set that the target task depends on is less than or equal to all result sets included in the upstream task.
[0010] Optionally, during the execution of the upstream task to which the incomplete result set belongs, scheduling the incremental data in the incomplete result set to the target task until the incomplete result set is changed to a complete result set includes: scheduling the incremental data in the incomplete result set to the target task in batches according to preset scheduling rules, and marking the scheduled incremental data with a scheduled label after each batch of incremental data is scheduled, until the incomplete result set is changed to a complete result set, and all data in the changed result set carries the scheduled label, wherein the target task executes tasks in batches based on the data received in batches.
[0011] Optionally, scheduling the incremental data in the incomplete result set to the target task in batches according to a preset scheduling rule includes: acquiring unscheduled incremental data in the incomplete result set at preset time intervals, and scheduling the unscheduled incremental data to the target task.
[0012] Optionally, scheduling the incremental data in the incomplete result set to the target task in batches according to preset scheduling rules includes: detecting whether the unscheduled incremental data generated in the incomplete result set reaches a preset data volume, and scheduling the preset data volume of unscheduled incremental data to the target task each time the unscheduled incremental data reaches the preset data volume.
[0013] Optionally, scheduling the remaining data in the incomplete result set to the target task in batches according to preset scheduling rules includes: judging whether the unscheduled incremental data in the incomplete result set is greater than or equal to the preset data amount at every preset time interval; when the unscheduled incremental data is greater than or equal to the preset data amount, scheduling the unscheduled incremental data to the target task; when the unscheduled incremental data is less than the preset data amount, judging again after the preset time interval whether the unscheduled incremental data is greater than or equal to the preset data amount, until the unscheduled incremental data is greater than or equal to the preset data amount, and scheduling the unscheduled incremental data to the target task.
[0014] According to another aspect of the present application, a task scheduling processing device is provided. The device includes: an acquisition unit, which is used to obtain result set information required by the target task from a metadata database before executing the target task, wherein the result set information includes the completeness of multiple result sets, and the multiple result sets refer to the result sets of upstream tasks that the target task depends on, and each upstream task updates the data in the corresponding result set during operation; a judgment unit, which is used to determine the complete result set and incomplete result set in the multiple result sets based on the result set information, and judge whether to allow the target task to run in advance based on the incomplete result set; a scheduling unit, which is used to schedule the data in the complete result set and incomplete result set to the target task if the target task is allowed to run in advance, and during the execution of the upstream task to which the incomplete result set belongs, schedule the incremental data in the incomplete result set to the target task until the incomplete result set is changed to the complete result set, wherein the target task executes the task in batches of data during the process of receiving data.
[0015] According to another aspect of an embodiment of the present application, a computer storage medium is provided, which is used to store a program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute a task scheduling processing method.
[0016] According to another aspect of an embodiment of the present application, an electronic device is also provided, comprising a processor and a memory; the memory stores computer-readable instructions, and the processor is used to run the computer-readable instructions, wherein the computer-readable instructions execute a task scheduling processing method when they are run.
[0017] Through the present application, the following steps are adopted: before executing the target task, the result set information required by the target task is obtained from the metadata database, wherein the result set information includes the completeness of multiple result sets, and the multiple result sets refer to the result sets of upstream tasks that the target task depends on, and each upstream task updates the data in the corresponding result set during operation; based on the result set information, the complete result set and the incomplete result set in the multiple result sets are determined, and according to the incomplete result set, it is judged whether the target task is allowed to run in advance; if the target task is allowed to run in advance, the data of the complete result set and the incomplete result set are scheduled to the target task, and during the execution of the upstream task to which the incomplete result set belongs, the incremental data in the incomplete result set is scheduled to the target task until the incomplete result set is changed to the complete result set, wherein the target task executes the task in batches during the process of receiving the data, which solves the problem in the related technology that the downstream task needs to wait for all the upstream tasks to be executed before it can start running, and the slow execution of a few upstream tasks will affect the overall running efficiency of the task. When most upstream tasks of the target task are completed but a few are not, complete and incomplete result sets are generated. By outputting the data in the complete and incomplete result sets that the target task depends on to the downstream first, the target task can be run in advance, thereby achieving the effect of parallel execution of the target task and the upstream slow tasks corresponding to the incomplete result set, thereby improving the scheduling efficiency and overall execution efficiency of the tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0019] FIG1 is a flowchart of a method for scheduling tasks according to an embodiment of the present application;
[0020] FIG2 is a first schematic diagram of optional task scheduling according to an embodiment of the present application;
[0021] FIG3 is a second schematic diagram of optional task scheduling according to an embodiment of the present application;
[0022] FIG4 is a third schematic diagram of optional task scheduling according to an embodiment of the present application;
[0023] FIG5 is a fourth schematic diagram of optional task scheduling according to an embodiment of the present application;
[0024] FIG6 is a first schematic diagram of batch scheduling data according to an embodiment of the present application;
[0025] FIG7 is a second schematic diagram of batch scheduling data according to an embodiment of the present application;
[0026] FIG8 is a schematic diagram of task scheduling of multi-level tasks according to an embodiment of the present application;
[0027] FIG9 is a schematic diagram of a task scheduling processing device according to an embodiment of the present application;
[0028] FIG10 is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0029] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0030] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0032] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0033] For ease of description, some nouns or terms involved in the embodiments of the present application are explained below:
[0034] DAG (Directed Acyclic Graph) is a common data structure in computer science and graph theory. It consists of nodes (or vertices) and directed edges. Each edge has a direction, pointing from one node to another. This means that in a DAG, there is a definite flow of information or operations from one node to another, and there are no paths that form loops.
[0035] Concurrent computing refers to the process of multiple programs or threads simultaneously processing data or performing calculations within the same time interval. This concurrency can occur across multiple computing entities, such as multi-core processors, multi-threaded processors, and distributed computing systems.
[0036] According to an embodiment of the present application, a task scheduling method is provided.
[0037] Figure 1 is a flow chart of a method for scheduling tasks according to an embodiment of the present application. The method is performed by a task scheduler, as shown in Figure 1, and includes the following steps:
[0038] Step S102, before executing the target task, obtain the result set information required by the target task from the metadata database, wherein the result set information includes the completeness of multiple result sets. The multiple result sets refer to the result sets of upstream tasks that the target task depends on. Each upstream task updates the data in the corresponding result set during operation.
[0039] Specifically, the target task is a downstream task to be executed. The execution of the target task depends on multiple result sets. The multiple result sets are a collection of result data generated by the upstream tasks of the target task during the execution process. The target task needs to receive the result set scheduled by the task scheduler before it can be executed. The task scheduler is associated with a metadata database. The scheduler saves the result set status of all tasks (including upstream tasks and downstream tasks) and the result set information required for each task in the metadata database. After the task execution is completed, the corresponding result set status in the metadata database will be updated.
[0040] The result set status refers to whether the task to which the result set belongs has generated all the data in the result set during its execution. The result set status can be represented by the completeness of the result set. If the result set is complete, it means that all the data in the result set has been generated. If the result set is incomplete, it means that some data in the result set has not been generated.
[0041] Step S104 : determining complete result sets and incomplete result sets in the multiple result sets based on the result set information, and judging whether to allow the target task to run in advance based on the incomplete result sets.
[0042] Specifically, if the completeness of the result set is 100%, it is a complete result set, and if the completeness of the result set is less than 100%, it is an incomplete result set. If there are incomplete result sets in multiple result sets, it is determined whether to allow the target task to run in advance.
[0043] Figure 2 is a schematic diagram of optional task scheduling in an embodiment of the present application. As shown in Figure 2, the downstream task is the target task, and the upstream tasks that the target task depends on include fast task 1, fast task 1 and slow task 1. Fast task 1 and fast task 2 have both been executed, and their result sets are complete result sets, and there is at least one incomplete result set.
[0044] It should be noted that in related technologies, the target task must wait for the completion of slow task 1 before starting. If slow task 1 is in the running state, the target task is in the waiting state. If slow task 1 takes a long time to complete, it will affect the scheduling timeliness and the execution timeliness of the target task. Therefore, in certain circumstances, this embodiment allows the downstream target task to run in advance when an upstream task is not fully completed, so that the target task and the unfinished upstream task can run in parallel, thereby improving the execution efficiency of the tasks.
[0045] Whether the target task is allowed to run in advance can be determined based on the number and attributes of the incomplete result sets. Optionally, in the task scheduling processing method provided in the embodiment of the present application, determining whether the target task is allowed to run in advance based on the incomplete result sets includes: determining whether the number of incomplete result sets is less than or equal to a preset number; when the number of incomplete result sets is less than or equal to the preset number, determining whether each incomplete result set carries a preset label, wherein the preset label indicates that the target task supports executing tasks in batches based on data received in batches; when all incomplete result sets carry preset labels, determining that the target task is allowed to run in advance.
[0046] It should be noted that if there are too many incomplete result sets, it will affect the execution of the target task, and the completeness of some result sets determines whether the target task can be started. If such tasks are incomplete, the target task cannot start execution. Therefore, for result sets whose completeness does not affect the start of the target task, a preset label is added. Before allowing the target task to run in advance, it is necessary to first determine whether the number of incomplete result sets is less than or equal to the preset number, and whether all incomplete result sets carry preset labels. If any of the conditions is not met, the target task is not allowed to run in advance. If both conditions are met, the target task is allowed to run in advance. For example, the preset number can be 2. If the number of incomplete result sets is less than or equal to 2 and all incomplete result sets carry preset labels, the target task is allowed to run in advance.
[0047] Step S106, when the target task is allowed to run in advance, the data in the complete result set and the incomplete result set are scheduled to the target task, and during the execution of the upstream task to which the incomplete result set belongs, the incremental data in the incomplete result set is scheduled to the target task until the incomplete result set is changed to the complete result set, wherein the target task executes the task in batches of data during the process of receiving data.
[0048] Specifically, allowing the target task to run ahead of schedule means allowing it to run before the upstream tasks corresponding to the incomplete result set complete. As shown in Figure 2, the upstream tasks on which the target task depends include Fast Task 1, Fast Task 1, and Slow Task 1. Fast Task 1 and Fast Task 2 have both completed, while Slow Task 1 has not. Therefore, the target task is allowed to run ahead of Slow Task 1. The execution of the target task depends on the result set of the upstream tasks. Therefore, before the upstream tasks corresponding to the incomplete result set complete, the task scheduler schedules data from the complete result set and existing data from the incomplete result set to the target task. This allows the target task to execute in parallel with the incomplete upstream tasks, thereby improving overall task efficiency. During the execution of the upstream tasks to which the incomplete result set belongs, incremental data is continuously generated in the incomplete result set until it becomes a complete result set. During this process, the task scheduler continuously schedules incremental data from the incomplete result set to the target task, ensuring the integrity of the target task's data processing.
[0049] The task scheduling processing method provided in the embodiment of the present application obtains result set information required by the target task from the metadata database before executing the target task, wherein the result set information includes the completeness of multiple result sets, and the multiple result sets refer to the result sets of upstream tasks that the target task depends on, and each upstream task updates the data in the corresponding result set during operation; based on the result set information, the complete result set and the incomplete result set in the multiple result sets are determined, and it is judged whether the target task is allowed to run in advance according to the incomplete result set; if the target task is allowed to run in advance, the data of the complete result set and the incomplete result set are scheduled to the target task, and during the execution of the upstream task to which the incomplete result set belongs, the incremental data in the incomplete result set is scheduled to the target task until the incomplete result set is changed to the complete result set, wherein the target task executes the task in batches during the process of receiving data, which solves the problem in the related art that the downstream task needs to wait for all upstream tasks to be executed before it can start running, and the slow execution of a few upstream tasks will affect the overall operation efficiency of the task. When most upstream tasks of the target task are completed but a few are not, complete and incomplete result sets are generated. By outputting the data in the complete and incomplete result sets that the target task depends on to the downstream first, the target task can be run in advance, thereby achieving the effect of parallel execution of the target task and the upstream slow tasks corresponding to the incomplete result set, thereby improving the scheduling efficiency and overall execution efficiency of the tasks.
[0050] In related technologies, the target task needs to wait for all upstream tasks it depends on to be completed before it can be executed. It should be noted that when the target task is executed, it specifically depends on the result set of the upstream task. The upstream task can contain multiple result sets. According to business needs, the downstream task can only depend on part of the result set of an upstream task, rather than the entire result set. This application changes the dependency relationship during the execution of the downstream task from the task dimension to the result set dimension. Optionally, in the task scheduling processing method provided in the embodiment of the present application, each upstream task that the target task depends on includes at least one result set. When an upstream task includes more than two result sets, the result set that the target task depends on is less than or equal to all the result sets included in an upstream task.
[0051] FIG3 is a second schematic diagram of optional task scheduling according to an embodiment of the present application. As shown in FIG3 , assuming that the target task depends on n result sets, the task can be run when n-1 result sets are in a completed state. Specifically, the target task depends on result sets A and B of fast task 1, result set C of fast task 1, and result set D of slow task 1. Among these result sets, only result set D remains unfinished, at which point the target task can begin execution. That is, the target task is executed concurrently with slow task 1 to improve the overall execution efficiency of the task.
[0052] Figure 4 is a third schematic diagram of optional task scheduling in an embodiment of the present application. As shown in Figure 4, an upstream task can output multiple result sets, and the downstream task only needs to determine whether to execute based on the result set it requires, without relying on the task completion status. Specifically, fast task 1 corresponds to result set A and result set B. At this time, result set B is completed and result set A is not completed. Therefore, fast task 1 is still in progress. The target task depends on result set B, result set C and result set D. At this time, only result set D is not completed. The target task can start execution and execute concurrently with fast task 1 to improve the overall execution efficiency of the task.
[0053] In order to ensure the integrity of data processing of the target task, optionally, in the task scheduling processing method provided in the embodiment of the present application, during the execution of the upstream task to which the incomplete result set belongs, the incremental data in the incomplete result set is scheduled to the target task until the incomplete result set is changed to a complete result set, including: scheduling the incremental data in the incomplete result set to the target task in batches according to preset scheduling rules, and marking the scheduled incremental data with a scheduled label after each batch of incremental data is scheduled, until the incomplete result set is changed to a complete result set, and all data in the changed result set carries the scheduled label, wherein the target task executes tasks in batches based on the data received in batches.
[0054] Figure 5 is a fourth schematic diagram of optional task scheduling of an embodiment of the present application. As shown in Figure 5, the result sets that the target task depends on include result set A, result set B, result set C and the result set of slow task 1. Result set A, result set B and result set C are all complete result sets and static result sets. The result of slow task 1 is an incomplete result set, of which the completed part constitutes a static result set. When the target task starts execution, it first uses these static data sets to start calculations in advance. At the same time, slow task 1 is also continuously updating its own result set to obtain incremental data. In this process, the target task can use the incremental data continuously generated in slow task 1 for calculation. The task scheduler schedules the incremental data generated by slow task 1 to the target task in batches according to the preset scheduling rules, and at the same time marks the incremental data that has participated in the calculation with a scheduled label to prevent repeated calculations, thereby improving the overall task processing efficiency while ensuring the integrity of data processing.
[0055] The preset scheduling rules may be scheduling rules for periodically extracting batch data. Optionally, in the scheduling processing method for tasks provided in an embodiment of the present application, scheduling the incremental data in the incomplete result set to the target task in batches according to the preset scheduling rules includes: obtaining the unscheduled incremental data in the incomplete result set at preset time intervals, and scheduling the unscheduled incremental data to the target task.
[0056] FIG6 is a schematic diagram of batch scheduling data in an embodiment of the present application. As shown in FIG6 , the task scheduler schedules the completed incremental data of the upstream task to the downstream target task at the same time interval. The downstream target task uses the latest incremental data for calculation. After each batch of incremental data is extracted, the incremental data of the batch is marked with a "scheduled" label to prevent repeated calculation. The scheduling method of this embodiment can ensure that the downstream task can obtain data for calculation at a fixed time interval, thereby preventing the downstream task from waiting in vain.
[0057] The preset scheduling rule can be a scheduling rule for extracting batch data according to the user-defined data volume. Optionally, in the scheduling processing method for tasks provided in the embodiment of the present application, scheduling the incremental data in the incomplete result set to the target task in batches according to the preset scheduling rule includes: detecting whether the unscheduled incremental data generated in the incomplete result set reaches the preset data volume, and scheduling the preset data volume of unscheduled incremental data to the target task each time the unscheduled incremental data reaches the preset data volume.
[0058] Figure 7 is a second schematic diagram of batch scheduling data in an embodiment of the present application. As shown in Figure 7, a batch of incremental data is scheduled to the downstream target task every fixed data size, and the scheduled incremental data is marked with a scheduled label to prevent repeated calculation.
[0059] It should be noted that task startup calculations often require resource allocation, startup calculations and other operations, which will consume some resources. When repeatedly starting calculation tasks with small amounts of data, it will cause a certain amount of resource waste. This embodiment sets a preset data amount and uses a fixed size to extract incremental data to ensure that the data amount of each startup calculation reaches the batch set amount, which can effectively avoid the waste of resources caused by scheduling too small a data amount and repeatedly starting calculations.
[0060] The preset scheduling rules can also be determined comprehensively based on time and data volume. Optionally, in the scheduling processing method for tasks provided in an embodiment of the present application, scheduling the remaining data in the incomplete result set to the target task in batches according to the preset scheduling rules includes: judging whether the unscheduled incremental data in the incomplete result set is greater than or equal to the preset data volume at every preset time interval; when the unscheduled incremental data is greater than or equal to the preset data volume, scheduling the unscheduled incremental data to the target task; when the unscheduled incremental data is less than the preset data volume, judging again after the preset time interval whether the unscheduled incremental data is greater than or equal to the preset data volume, until the unscheduled incremental data is greater than or equal to the preset data volume, and scheduling the unscheduled incremental data to the target task.
[0061] Specifically, the task scheduler determines whether the incremental data completed by the upstream task has reached the preset data volume at the same time interval. If it has not reached the preset data volume, it means that the data volume is too small. In order to avoid the waste of resources caused by frequently starting the target task with a small data volume, the task scheduler can determine whether the incremental data completed by the upstream task has reached the preset data volume after the next time interval. If it has reached the preset data volume, it means that the data volume is sufficient and the data can be scheduled. It should be noted that the downstream target task can also be started when the last completed incremental data is less than the preset amount. Through this embodiment, while ensuring that the downstream task will not wait in vain for a long time, the problem of frequently starting the downstream task with a small data volume is avoided.
[0062] It should be noted that the task scheduling processing method of this embodiment can be applied to the processing of multi-level tasks. If tasks at multiple levels meet the parallel computing conditions, the computing efficiency can be further improved. Figure 8 is a schematic diagram of task scheduling of multi-level tasks in an embodiment of the present application. As shown in Figure 8, when the downstream task 1 of the upper layer only waits for the result set D, the downstream task 1 can start computing. When the downstream task 2 of the lower layer only waits for the result set H, the downstream task 2 can start computing, thereby realizing multi-level task parallel computing.
[0063] It should be noted that in this embodiment, the startup of the downstream task depends on the result set status. Assume that the time required for the x result sets required by the task is rt1, rt2, rt3...rtx. A task can have multiple result sets, so the time rt for the result set to complete will be less than or equal to the corresponding task completion time t. From this, the waiting time T_optimized of the downstream task can be calculated: T_optimized = second_largest(rt1, rt2, rt3...rtx). In related technologies, the startup of the downstream task depends on the result set status and the upstream task status. Assume that the time consumed by n upstream tasks is t1, t2, t3...tn. From this, the waiting time T_traditional of the downstream task can be calculated: T_traditional = max(t1, t2, t3...tn).
[0064] The task scheduling method of this embodiment can greatly improve computing efficiency when the generation time of a certain result set that a task depends on is significantly different from the others. When multiple levels in the entire link meet the parallel computing conditions, the computing efficiency can be further improved, thereby improving the task execution efficiency of the entire link.
[0065] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0066] The present application also provides a task scheduling processing device. It should be noted that the task scheduling processing device of the present application embodiment can be used to execute the task scheduling processing method provided in the present application embodiment. The task scheduling processing device provided in the present application embodiment is introduced below.
[0067] FIG9 is a schematic diagram of a task scheduling processing device according to an embodiment of the present application. As shown in FIG9 , the device includes: an acquisition unit 902 , a judgment unit 904 , and a scheduling unit 906 .
[0068] Specifically, the acquisition unit 902 is used to obtain the result set information required by the target task from the metadata database before executing the target task, wherein the result set information includes the completeness of multiple result sets. The multiple result sets refer to the result sets of upstream tasks that the target task depends on, and each upstream task updates the data in the corresponding result set during operation.
[0069] The judging unit 904 is configured to determine a complete result set and an incomplete result set in the multiple result sets based on the result set information, and to judge whether to allow the target task to run in advance according to the incomplete result set.
[0070] The scheduling unit 906 is used to schedule the data in the complete result set and the incomplete result set to the target task while allowing the target task to run in advance, and to schedule the incremental data in the incomplete result set to the target task during the execution of the upstream task to which the incomplete result set belongs, until the incomplete result set is changed to the complete result set, wherein the target task executes the task in batches of data during the process of receiving data.
[0071] The task scheduling processing device provided by the embodiment of the present application obtains the result set information required by the target task from the metadata database through the acquisition unit 902 before executing the target task, wherein the result set information includes the completeness of multiple result sets, and the multiple result sets refer to the result sets of the upstream tasks that the target task depends on, and each upstream task updates the data in the corresponding result set during the operation; the judgment unit 904 determines the complete result set and the incomplete result set in the multiple result sets based on the result set information, and judges whether to allow the target task to run in advance according to the incomplete result set; the scheduling unit 906 schedules the data in the complete result set and the incomplete result set to the target task when the target task is allowed to run in advance, and sends the complete result set and the incomplete result set to the target task during the execution of the upstream task to which the incomplete result set belongs. The task schedules the incremental data in the incomplete result set until the incomplete result set is changed to the complete result set. Among them, the target task executes the task in batches during the process of receiving data, which solves the problem in related technologies that downstream tasks need to wait for all upstream tasks to be completed before they can start running, and the slow execution of a few upstream tasks will affect the overall running efficiency of the task. When most of the upstream tasks of the target task are completed and a few tasks are not completed, a complete result set and an incomplete result set are generated. By outputting the data in the complete result set and the incomplete result set that the target task depends on to the downstream first, the target task can be run in advance, thereby achieving the goal of parallel operation of the target task and the upstream slow tasks corresponding to the incomplete result set, thereby improving the scheduling efficiency of the task and the overall execution efficiency of the task.
[0072] Optionally, in the task scheduling and processing device provided in an embodiment of the present application, the judgment unit 904 includes: a first judgment module, used to judge whether the number of incomplete result sets is less than or equal to a preset number; a second judgment module, used to judge whether each incomplete result set carries a preset label when the number of incomplete result sets is less than or equal to a preset number, wherein the preset label indicates that the target task supports batch execution of tasks based on data received in batches; a determination module, used to determine whether the target task is allowed to run in advance when all incomplete result sets carry preset labels.
[0073] Optionally, in the task scheduling and processing device provided in an embodiment of the present application, each upstream task that the target task depends on includes at least one result set. When an upstream task includes more than two result sets, the result set that the target task depends on is less than or equal to all result sets included in an upstream task.
[0074] Optionally, in the task scheduling processing device provided in the embodiment of the present application, the scheduling unit 906 includes: a scheduling module, which is used to schedule the incremental data in the incomplete result set to the target task in batches according to preset scheduling rules, and when each batch of incremental data is scheduled, the scheduled incremental data is marked with a scheduled label until the incomplete result set is changed to a complete result set, and all data in the changed result set carries the scheduled label, wherein the target task executes tasks in batches based on the data received in batches.
[0075] Optionally, in the task scheduling processing device provided in an embodiment of the present application, the scheduling module includes a first sub-scheduling module, which is used to obtain unscheduled incremental data in the incomplete result set at preset time intervals and schedule the unscheduled incremental data to the target task.
[0076] Optionally, in the task scheduling and processing device provided in the embodiment of the present application, the scheduling module also includes a second sub-scheduling module, which is used to detect whether the unscheduled incremental data generated in the incomplete result set reaches a preset data volume, and each time the unscheduled incremental data reaches the preset data volume, the preset data volume of unscheduled incremental data is scheduled to the target task.
[0077] Optionally, in the scheduling and processing device for tasks provided in an embodiment of the present application, the scheduling module also includes a third sub-scheduling module, which is used to determine whether the unscheduled incremental data in the incomplete result set is greater than or equal to a preset data amount at every preset time interval; when the unscheduled incremental data is greater than or equal to the preset data amount, the unscheduled incremental data is scheduled to the target task; when the unscheduled incremental data is less than the preset data amount, after the preset time interval, it is determined again whether the unscheduled incremental data is greater than or equal to the preset data amount, until the unscheduled incremental data is greater than or equal to the preset data amount, and the unscheduled incremental data is scheduled to the target task.
[0078] The scheduling processing device for the above tasks includes a processor and a memory. The above acquisition unit, judgment unit and scheduling unit are all stored in the memory as program units, and the processor executes the above program units stored in the memory to realize corresponding functions.
[0079] The processor contains a core, which retrieves the corresponding program unit from memory. One or more cores can be configured. By adjusting the core parameters, the system can address the problem in related technologies where downstream tasks must wait for all upstream tasks to complete before they can begin running. This can affect the overall efficiency of tasks if a few upstream tasks execute slowly.
[0080] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0081] An embodiment of the present application also provides a computer storage medium, which is used to store a program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute a scheduling processing method for a task: before executing a target task, obtaining result set information required by the target task from a metadata database, wherein the result set information includes the completeness of multiple result sets, and the multiple result sets refer to the result sets of upstream tasks that the target task depends on, and each upstream task updates the data in the corresponding result set during the running process; based on the result set information, a complete result set and an incomplete result set in the multiple result sets are determined, and according to the incomplete result set, it is judged whether the target task is allowed to run in advance; if the target task is allowed to run in advance, the data in the complete result set and the incomplete result set are scheduled to the target task, and during the execution of the upstream task to which the incomplete result set belongs, the incremental data in the incomplete result set is scheduled to the target task until the incomplete result set is changed to the complete result set, wherein the target task executes the task in batches of data during the process of receiving data.
[0082] The present application also provides an electronic device. FIG10 is a schematic diagram of an electronic device according to an embodiment of the present application. As shown in FIG10 , electronic device 1001 includes a processor and a memory. The memory stores computer-readable instructions, and the processor is configured to execute the computer-readable instructions. When the computer-readable instructions are executed, a method for scheduling a task is executed. The electronic device herein may be a server, a PC, a PAD, a mobile phone, or the like.
[0083] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0084] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.
[0085] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0086] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0087] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0088] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0089] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0090] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0091] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A task scheduling method, comprising: Before executing the target task, obtaining result set information required by the target task from the metadata database, wherein the result set information includes the completeness of multiple result sets, and the multiple result sets refer to the result sets of upstream tasks that the target task depends on, and each upstream task updates the data in the corresponding result set during operation; Determine a complete result set and an incomplete result set among the multiple result sets based on the result set information, and determine whether to allow the target task to run in advance according to the incomplete result set; While allowing the target task to run in advance, the data in the complete result set and the incomplete result set are scheduled to the target task, and during the execution of the upstream task to which the incomplete result set belongs, the incremental data in the incomplete result set is scheduled to the target task until the incomplete result set is changed to a complete result set, wherein the target task executes the task in batches of data during the process of receiving data.
2. The method according to claim 1, wherein: Judging whether to allow the target task to run in advance according to the incomplete result set includes: Determining whether the number of the incomplete result sets is less than or equal to a preset number; When the number of the incomplete result sets is less than or equal to the preset number, determining whether each incomplete result set carries a preset tag, wherein the preset tag indicates that the target task supports executing tasks in batches based on data received in batches; In the case that all incomplete result sets carry the preset label, it is determined to allow the target task to run in advance.
3. The method according to claim 1, wherein: Each upstream task that the target task depends on includes at least one result set. When an upstream task includes more than two result sets, the result set that the target task depends on is less than or equal to all result sets included in the upstream task.
4. The method according to claim 1, wherein: During the execution of the upstream task to which the incomplete result set belongs, scheduling incremental data in the incomplete result set to the target task until the incomplete result set is changed to a complete result set includes: The incremental data in the incomplete result set is scheduled to the target task in batches according to preset scheduling rules, and the scheduled incremental data is marked with a scheduled label every time a batch of incremental data is scheduled, until the incomplete result set is changed to a complete result set, and all data in the changed result set carries the scheduled label, wherein the target task executes tasks in batches based on the data received in batches.
5. The method according to claim 4, wherein: Scheduling the incremental data in the incomplete result set to the target task in batches according to a preset scheduling rule includes: Unscheduled incremental data in the incomplete result set is acquired at preset time intervals, and the unscheduled incremental data is scheduled to the target task.
6. The method according to claim 4, wherein: Scheduling the incremental data in the incomplete result set to the target task in batches according to a preset scheduling rule includes: Detect whether the unscheduled incremental data generated in the incomplete result set reaches a preset data amount, and schedule the preset data amount of unscheduled incremental data to the target task each time the unscheduled incremental data reaches the preset data amount.
7. The method according to claim 4, wherein: Scheduling the remaining data in the incomplete result set to the target task in batches according to a preset scheduling rule includes: Determining at preset time intervals whether the unscheduled incremental data in the incomplete result set is greater than or equal to a preset data amount; When the unscheduled incremental data is greater than or equal to the preset data amount, scheduling the unscheduled incremental data to the target task; In the case that the unscheduled incremental data is less than the preset data amount, it is determined again whether the unscheduled incremental data is greater than or equal to the preset data amount after the preset time interval, until the unscheduled incremental data is greater than or equal to the preset data amount, and the unscheduled incremental data is scheduled to the target task.
8. A task scheduling processing device, comprising: An acquisition unit is used to acquire result set information required by the target task from a metadata database before executing the target task, wherein the result set information includes the completeness of multiple result sets, and the multiple result sets refer to the result sets of upstream tasks that the target task depends on, and each upstream task updates the data in the corresponding result set during operation; a judging unit, configured to determine a complete result set and an incomplete result set among the multiple result sets based on the result set information, and to judge whether to allow the target task to run in advance according to the incomplete result set; A scheduling unit is used to schedule the data in the complete result set and the incomplete result set to the target task while allowing the target task to run in advance, and to schedule incremental data in the incomplete result set to the target task during the execution of the upstream task to which the incomplete result set belongs, until the incomplete result set is changed to a complete result set, wherein the target task executes the task in batches of data during the process of receiving data.
9. A computer storage medium, wherein: The computer storage medium is used to store a program, wherein when the program is executed, the device where the computer storage medium is located is controlled to execute the scheduling processing method for the task described in any one of claims 1 to 7.
10. An electronic device comprising a memory and a processor, wherein: The memory stores a computer program, and the processor is configured to execute the task scheduling processing method according to any one of claims 1 to 7 through the computer program.
Citation Information
Patent Citations
Timing task processing method and device, computer equipment and storage medium
CN115543568A
Processor with hardware pipeline
CN115904505A
Distributed computing scheduling system, task processing method, equipment and storage medium
CN116360993A
Task scheduling processing method and device, storage medium and electronic equipment
CN117472533A
Cited By
Power grid situation awareness edge industrial personal computer node load dynamic distribution method
CN121957920A