Data processing task scheduling method and program product
By obtaining the target resource utilization rate and upper limit value of data processing tasks, using batch task schedules to filter and dynamically orchestrate subtasks, the problem of insufficient resource utilization is solved and the efficiency and stability of data processing is improved.
Patent Information
- Application Number
- CN202510707399.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-02
AI Technical Summary
The problem of insufficient resource utilization in the prior art has led to idle resources or bottlenecks in data processing, which is difficult to meet the needs of high concurrency and large-scale data processing.
By receiving task execution requests, obtain the actual usage and usage upper limit of the target resource, use the batch task schedule to filter executable subtasks, and dynamically orchestrate the subtask execution order according to resource usage to avoid resource overload or waste.
It realizes the full and reasonable use of resources, improves the overall efficiency and throughput capabilities of data processing, and supports efficient and stable operation of large-scale data processing and complex business scenarios.
Smart Images

Figure CN120578472A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method for scheduling data processing tasks and a program product. Background Art
[0002] With the rapid development of information technology, the amount of data is growing and expanding rapidly. When processing data, the problem of insufficient resource utilization has gradually become prominent.
[0003] Some resource scheduling methods in related technologies have limitations. They simply allocate initial memory and CPU core configurations for individual data processing tasks based solely on analyzing server performance metrics. This static allocation method can lead to idle resources or bottlenecks. For example, when task demand suddenly increases and pre-allocated resources cannot meet processing requirements, processing speeds can slow or even stall. Conversely, when task demand decreases, a large number of allocated resources remain idle, resulting in significant waste. This irrational resource allocation method severely impacts the overall efficiency and throughput of data processing, making it difficult to meet the demands of high-concurrency, large-scale data processing. Summary of the Invention
[0004] The present invention provides a data processing task scheduling method and program product to solve the problem of insufficient resource utilization in the prior art.
[0005] According to one aspect of the present invention, a method for scheduling data processing tasks is provided, the method comprising:
[0006] receiving a task execution request for a data processing task, and obtaining actual usage rates of multiple target resources for executing the data processing task and upper usage limits of the target resources recorded in a resource information management table; wherein the data processing task includes multiple subtasks; and the target resources include at least central processing unit resources, memory resources, and network resources;
[0007] If the actual usage rates of the plurality of target resources do not reach the usage rate upper limit of the target resources, determining at least one executable subtask based on the task scheduling information of the plurality of subtasks recorded in the batch task scheduling table corresponding to the data processing task; wherein the task scheduling information of the plurality of subtasks is recorded in the batch task scheduling table;
[0008] A target subtask is determined from at least one of the executable subtasks, and the target subtask is executed so that the actual usage rate of at least one of the target resources is closer to and does not exceed the usage rate upper limit of the target resource.
[0009] According to another aspect of the present invention, a data processing task scheduling device is provided, the device comprising:
[0010] a target resource utilization rate and upper limit value acquisition module, configured to receive a task execution request for a data processing task, and obtain the actual utilization rates of multiple target resources for executing the data processing task and the upper limit values of the utilization rates of the target resources recorded in the resource information management table; wherein the data processing task includes multiple subtasks; and the target resources include at least central processing unit resources, memory resources, and network resources;
[0011] an executable subtask determining module, configured to determine, when the actual utilization rates of the plurality of target resources do not reach the utilization rate upper limit of the target resources, at least one executable subtask based on the task scheduling information of the plurality of subtasks recorded in the batch task scheduling table corresponding to the data processing task; wherein the task scheduling information of the plurality of subtasks is recorded in the batch task scheduling table;
[0012] The target subtask determination execution module is used to determine a target subtask from at least one of the executable subtasks and execute the target subtask so that the actual usage rate of at least one of the target resources is closer to and does not exceed the usage rate upper limit of the target resource.
[0013] According to another aspect of the present invention, an electronic device is provided, comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the scheduling method for data processing tasks described in any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the data processing task scheduling method described in any embodiment of the present invention when executed.
[0018] According to another aspect of the present invention, an embodiment of the present disclosure further provides a computer program product, including a computer program, which, when executed by a processor, implements the method for scheduling data processing tasks as described in any one of the embodiments of the present disclosure.
[0019] The technical solution of the embodiment of the present invention is as follows: first, by receiving a task execution request for a data processing task, the actual utilization rate of multiple target resources for executing the data processing task and the utilization rate upper limit value of the target resources recorded in the resource information management table are obtained; by obtaining the utilization rate and utilization rate upper limit value of the target resources involved in the data processing task, subtasks are screened and allocated according to the utilization rate upper limit value, so as to avoid problems such as resource overload or resource waste; then, in the case that the actual utilization rate of multiple target resources does not reach the utilization rate upper limit value of the target resource, the task scheduling information of multiple subtasks recorded in the batch task scheduling table corresponding to the data processing task is determined. There is one less executable subtask. Since the batch task scheduling table records the task scheduling information of multiple subtasks, the executable subtasks can be accurately screened according to the batch task scheduling table, and the order of parallel execution of multiple subtasks can be dynamically arranged according to the real-time resource usage and task demand changes, thereby improving resource utilization. Finally, by determining the target subtask from at least one of the executable subtasks and executing the target subtask, the actual utilization rate of at least one target resource is closer and does not exceed the utilization upper limit of the target resource, so that the target resource can be fully and reasonably used without exceeding the upper limit threshold, effectively avoiding the problem of excessive or insufficient resource occupation. By reasonably arranging the execution order of subtasks and resource allocation, the overall time of multi-task execution is greatly shortened, and the resource allocation is refined and automated, providing efficient and stable technical support for large-scale data processing and complex business scenarios.
[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 This is a flowchart of a method for scheduling data processing tasks according to the first embodiment of the present invention;
[0023] Figure 2 This is a flowchart of a method for scheduling data processing tasks according to a second embodiment of the present invention;
[0024] Figure 3a This is a flowchart of dynamic orchestration of data processing tasks according to a method for scheduling data processing tasks in accordance with a third embodiment of the present invention;
[0025] Figure 3b 1 is a schematic diagram of node execution of a scheduling method for data processing tasks according to a third embodiment of the present invention that can be used in an embodiment of the present invention;
[0026] Figure 3c 1 is a schematic diagram of node execution of a scheduling method for data processing tasks according to a third embodiment of the present invention that can be used in an embodiment of the present invention;
[0027] Figure 3d 1 is a schematic diagram of node execution of a scheduling method for data processing tasks according to a third embodiment of the present invention that can be used in an embodiment of the present invention;
[0028] Figure 3e 1 is a schematic diagram of node execution of a scheduling method for data processing tasks according to a third embodiment of the present invention that can be used in an embodiment of the present invention;
[0029] Figure 4 2 is a schematic structural diagram of a data processing task scheduling device provided according to a fourth embodiment of the present invention;
[0030] Figure 5 It is a structural diagram of an electronic device for implementing the method for scheduling data processing tasks according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0032] It should be noted that the terms "first," "second," "target," and the like in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this way are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.
[0033] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0034] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0035] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0036] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0037] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0038] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0039] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0040] Example 1
[0041] Figure 1 A flowchart of a method for scheduling data processing tasks is provided for the first embodiment of the present invention. This embodiment is applicable to situations where multiple tasks dynamically arrange the execution order of tasks according to resources. The method can be executed by a scheduling device for data processing tasks. The scheduling device for data processing tasks can be implemented in the form of hardware and / or software. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, PC or server. Figure 1 As shown, the method may specifically include:
[0042] S110: Receive a task execution request for a data processing task, and obtain actual usage rates of multiple target resources for executing the data processing task and usage upper limits of the target resources recorded in a resource information management table.
[0043] In embodiments of the present invention, a data processing task can be understood as a set of tasks requiring a series of operations on data. Such a data processing task can include multiple subtasks. A subtask is the smallest independently executable unit of work within a data processing task. Each subtask corresponds to a specific data processing operation; subtasks may have dependencies such as sequential execution and data transfer, or they may execute independently and in parallel. By breaking down a data processing task into multiple subtasks, resource allocation and task scheduling are facilitated.
[0044] A task execution request is an initiation instruction requesting the execution of a data processing task. The task execution request may include relevant information about the data processing task. For example, the task execution request may include a resource information management table and / or a batch task scheduling table corresponding to the data processing task. The task execution request may also include metadata carried by the request, such as task type, data volume, processing logic rules, expected execution duration, and other information.
[0045] Optionally, the task execution request can be triggered independently or in combination through one or more of the following methods: manual submission of instructions by the user, automated process timing drive, automatic triggering by specific events, periodic timing triggering, etc., to meet the startup requirements of data processing tasks in different scenarios.
[0046] The target resources are the hardware resources or system resources required to perform data processing tasks. The target resources include at least central processing unit resources, memory resources and network resources. The target resources may also include one or more resources such as input / output resources (I / O resources). Central processing unit resources are responsible for executing instructions and data processing, etc., and are the core computing resources for task operation. Memory resources are used to temporarily store data and intermediate results during program execution, etc., which affect the task operation speed and concurrency capabilities. Network resources are used to support data transmission, communication and external resource access, etc., to ensure the input and output of data required for the task and collaboration between systems. I / O resources are used to support input and output operations of data between the computing system and external devices, storage media or other systems, to ensure the read and write interaction of data required for data processing tasks and the collaborative operation between hardware and software components, and to determine the efficiency and stability of task processing data.
[0047] The actual utilization rate of the target resource can be understood as the actual utilization degree of the corresponding target resource at a certain moment or time period, usually expressed as a percentage. The actual utilization rate can be used as a real-time detection indicator to reflect the load situation of the resource. The actual utilization rate of various target resources of the data processing task may include the overall utilization rate of the central processing unit resources, the overall utilization rate of memory resources and the overall utilization rate of network resources, and / or, the respective central processing unit resource utilization rate of each subtask, the respective memory resource utilization rate of each subtask and the respective network resource utilization rate of each subtask. The utilization rate upper limit value refers to a pre-set target resource utilization threshold. The utilization rate upper limit value is used to limit the use of the corresponding resources to not exceed a safe range, so as to avoid system performance degradation, stability reduction or hardware damage due to excessive resource occupation.
[0048] The resource information management table is a core component for storing and managing basic information of target resources. The resource information management table is used to record one or more information of at least one resource, such as resource type identification, usage upper limit, configuration personnel, and configuration time.
[0049] Optionally, after receiving a task execution request for a data processing task, the actual utilization rate of various target resources of the data processing task is collected through real-time performance detection tools or system interfaces, and the recorded utilization rate upper limit value of the target resource is extracted from the resource information management table.
[0050] As an optional technical solution of an embodiment of the present invention, optionally, before obtaining the upper limit value of the utilization rate of the target resource recorded in the resource information management table, it also includes: receiving a resource information editing operation for the resource information management table through a resource information editing interface, and determining the resource information management table according to the resource information editing operation.
[0051] Optionally, before obtaining the upper limit value of the target resource usage recorded in the resource information management table, a resource information editing operation can be performed on the data processing task through a visual resource information editing interface to configure the relevant parameters of various target resources required for the data processing task. The relevant parameters of the target resource may include one or more attribute information such as resource type, upper limit value of usage, enabled status, configuration date, configuration time, configuration personnel, etc. Resource type may include one or more resources such as central processing unit resources, memory resources, and network resources. The upper limit value of usage refers to the maximum proportion of the corresponding resource that can be occupied during the execution of the data processing task.
[0052] Optionally, the upper limit of the target resource usage rate may be determined based on historical data, machine learning predictions, and other methods, including but not limited to the following implementation methods, which may be used alone or in combination:
[0053] For example, analyze the resource usage of similar data processing tasks in the past, count the average, peak and fluctuation range of the corresponding resource utilization, and reasonably set the corresponding ratio as the utilization upper limit, so as to ensure the normal operation of the task and avoid excessive resource reservation.
[0054] For example, resource caps are dynamically allocated based on the priority of data processing tasks. High-priority tasks can occupy a higher proportion of resources, while low-priority tasks are limited to a lower proportion of resources, to achieve reasonable scheduling and efficient utilization of resources.
[0055] For example, a machine learning model is used, combined with characteristics such as the input data scale and processing logic complexity of the data processing task, to predict the resource consumption trend during task execution, and then dynamically generate a usage rate cap to reduce resource waste while ensuring task completion.
[0056] Optionally, after completing the resource information editing operation, the resource information editing operation for the resource information management table may be received through the resource information editing interface, and the resource information management table corresponding to the data processing task may be generated by integrating the configuration parameters according to the resource information editing operation.
[0057] By editing the resource information management table and determining the resource information management table, the data processing task can determine the upper limit of the utilization rate according to the configured resource information management table, and can reasonably allocate resources based on the resource information management table, which not only ensures the efficient use of resources, but also avoids system performance degradation or operation failure due to excessive resource allocation, and achieves a dynamic balance between task processing and resource management.
[0058] S120. When the actual utilization rates of the plurality of target resources do not reach the utilization upper limit of the target resources, determine at least one executable subtask based on the task scheduling information of the plurality of subtasks recorded in the batch task scheduling table corresponding to the data processing task.
[0059] In an embodiment of the present invention, the batch task scheduling table can be understood as a table that records the task scheduling information of multiple subtasks in a data processing task. An executable subtask refers to a subtask that meets the preset scheduling conditions under the current resource status. The preset scheduling conditions may include conditions such as the subtask is in a pending state and has no recursive dependency on the task being executed. Task scheduling information can be understood as an information set that describes the execution rules and parameters of the subtask. Optionally, the actual utilization rate of various target resources such as central processing unit resources, memory resources and network resources is detected in real time, and dynamically compared with the utilization rate upper limit value recorded in the resource information management table. When it is detected that the actual utilization rates of various target resources are all lower than the preset utilization rate upper limit value of the target resource, the subtask screening process is triggered.
[0060] Optionally, based on the batch task scheduling table corresponding to the data processing task, the scheduling information of all subtasks related to the current data task is extracted and subtask screening is performed. Using the batch task scheduling table, when the actual utilization of multiple target resources has not reached the upper limit, executable subtasks can be screened according to preset scheduling rules to achieve reasonable resource allocation and orderly task execution.
[0061] As an optional technical solution of an embodiment of the present invention, optionally, before determining at least one executable subtask based on the task scheduling information of the multiple subtasks recorded in the batch task scheduling table corresponding to the data processing task, it also includes: parsing the data processing task to obtain the multiple subtasks corresponding to the data processing task and the task scheduling information of the multiple subtasks, and generating a batch task scheduling table corresponding to the data processing task based on the task scheduling information of the multiple subtasks.
[0062] Optionally, the identity of the initiator of the data processing task is verified to confirm whether the initiator has the authority to execute the task, and then the data processing task is parsed.
[0063] Optionally, after receiving a data processing task, the format and content of the data processing task can be verified. For example, the JSON structure can be checked to see if it is correct, if required fields are missing, and if there are circular dependencies in the task logic.
[0064] Optionally, the data processing task is parsed to split the complex data processing flow into multiple subtasks. Each subtask may include one or more task information such as execution logic information, preceding recursive dependent task information, and resource requirement information.
[0065] Optionally, while splitting the subtasks, task scheduling information for the multiple subtasks can be extracted from a batch task scheduling table corresponding to the data processing task. Optionally, after obtaining the multiple subtasks corresponding to the data processing task and the task scheduling information for the multiple subtasks, all subtasks and their task scheduling information can be aggregated to generate a batch task scheduling table.
[0066] By decomposing tasks and extracting scheduling information, complex data processing tasks are converted into multiple subtask networks, achieving refined control of resource allocation and significantly improving data processing throughput and task parallel processing capabilities.
[0067] As an optional technical solution of an embodiment of the present invention, optionally, before determining at least one executable subtask based on the task scheduling information of the multiple subtasks recorded in the batch task scheduling table corresponding to the data processing task, it also includes: receiving a scheduling information editing operation for at least one task scheduling information of the multiple subtasks recorded in the batch task scheduling table, and updating the batch task scheduling table according to the scheduling information editing operation.
[0068] Optionally, before determining the executable subtasks according to the batch task schedule, the task scheduling information can be dynamically adjusted. Specifically, an editing operation can be initiated on the scheduling information of the subtasks in the batch task schedule, such as modifying dependencies, adjusting task status, or other editing methods.
[0069] Optionally, upon receiving a scheduling information edit operation, the modified content is automatically verified for validity, for example, to check for dependency conflicts, resource over-allocation, etc. If the verification passes, the batch task schedule is immediately updated and the changed task scheduling information is synchronized.
[0070] Optionally, the updated batch task scheduling table supports version tracing and rollback functions, automatically recording the time, personnel, and modified content of each editing operation to ensure the auditability and recoverability of task scheduling information adjustments.
[0071] By editing the task scheduling information, users are allowed to modify the task scheduling strategy in real time according to actual task changes, resource conditions, etc., to achieve adaptive management of task execution and ensure that data processing tasks are always executed in the optimal way.
[0072] S130: Determine a target subtask from at least one executable subtask, and execute the target subtask so that the actual usage rate of at least one target resource is closer to and does not exceed the usage rate upper limit of the target resource.
[0073] The target subtask refers to a task that is selected from at least one executable subtask and will be actually executed.
[0074] Optionally, the target subtask is determined based on a criterion of not exceeding an upper limit value on the basis of maximizing resource utilization.
[0075] Specifically, the determination of the target subtask includes but is not limited to the following implementation methods, and each method can be used alone or in combination:
[0076] For example, when there is only a single executable subtask, the executable subtask may be directly determined as the target subtask.
[0077] Exemplarily, key target resources are selected through target resource priority or other means, the actual usage rate of the key target resource when each executable subtask is executed is obtained, and the subtasks are sorted according to the actual usage rate, and the executable subtask whose key target resource is closest to the usage upper limit after execution is selected; the key target resource can be determined by factors such as hardware resources.
[0078] For example, the impact of each executable subtask on the actual utilization rate of at least one target resource is determined, and executable subtasks that can bring the actual utilization rates of more target resources close to, but not exceeding, the corresponding utilization rate upper limit are selected. For example, the selection can be performed based on the number of target resources that can be brought close to, but not exceeding, the upper limit, and the executable subtask that can effectively utilize the largest number of target resources is selected as the target subtask.
[0079] Furthermore, the target subtask is determined and executed so that the actual utilization rate of at least one target resource is closer to and does not exceed the utilization rate upper limit of the target resource.
[0080] As an optional technical solution of an embodiment of the present invention, optionally, after executing the target subtask, it also includes: when the data processing task has been completed, analyzing the task execution log corresponding to the data processing task to obtain actual resource usage information of multiple target resources during the execution of the data processing task, and displaying the actual resource usage information.
[0081] The actual resource usage information includes the actual utilization rates of multiple target resources at the actual time of use. The task execution log is used to track, analyze, and evaluate the entire data processing task process. The task execution log at least includes the actual utilization of various target resources used during the data processing task process.
[0082] Optionally, after the data processing task is completed, the task execution log analysis process can be triggered manually or automatically.
[0083] Optionally, actual resource usage information of various target resources such as central processing unit resources, memory resources, and network resources consumed during the execution of statistical processing tasks can be collected through a distributed log collection component.
[0084] Optionally, different resource usage curve graphs may be used to visually display actual resource usage information of various target resources.
[0085] By analyzing task execution logs, we can clearly understand the actual usage of various target resources during task execution and enhance the observability of resource consumption; visualizing actual resource usage information helps evaluate resource utilization efficiency and provide reliable data basis for subsequent task scheduling, resource allocation and performance tuning.
[0086] As an optional technical solution of an embodiment of the present invention, optionally, after obtaining the actual resource usage information of the multiple target resources in the process of executing the data processing task, it also includes: generating resource configuration prompt information corresponding to the data processing task based on the actual resource usage information, and displaying the resource configuration prompt information.
[0087] The resource configuration prompt information may include first configuration prompt information for the upper limit of the usage rate of the target resource and / or second configuration prompt information for the hardware resource.
[0088] The first configuration prompt information may be a suggestion for adjusting the upper limit of the utilization rate of target resources during task execution to optimize resource scheduling strategies and improve resource utilization. The first configuration prompt information includes, but is not limited to, a suggestion for adjusting the upper limit of the utilization rate of a certain type of target resource from a current set value to a new recommended value.
[0089] The second configuration prompt information may be a suggestion for expanding or supplementing hardware resources to resolve task delays or performance degradation caused by resource bottlenecks. The second configuration prompt information may include, but is not limited to, one or more suggestions for increasing the number of nodes, expanding memory capacity, increasing network bandwidth, or adding SSD hard drives.
[0090] Optionally, based on the actual resource usage information of multiple target resources collected during task execution, one or more resource usage conditions such as the average usage rate, peak usage rate, and fluctuation trend of each type of target resource can be analyzed and calculated, and then corresponding resource configuration prompt information can be generated.
[0091] Generating configuration prompts based on actual resource usage information helps identify resource waste or bottlenecks and provides precise guidance for dynamically adjusting resource configuration to improve overall resource utilization efficiency while ensuring task performance.
[0092] The resource configuration prompt information may include first configuration prompt information for the upper limit value of the utilization rate of the target resource; the first configuration prompt information includes a recommended usage upper limit value corresponding to the upper limit value of the utilization rate. As an optional technical solution of an embodiment of the present invention, optionally, after obtaining the actual resource usage information of the plurality of target resources in the process of executing the data processing task, it further includes: displaying the first configuration prompt information for the upper limit value of the utilization rate of the target resource in the resource information management table; and / or, in response to a usage trigger operation for the recommended usage upper limit value of at least one target resource, updating the upper limit value of the utilization rate of the target resource in the resource information management table to the recommended usage upper limit value.
[0093] Optionally, after completing the analysis of the actual resource usage information, a recommended usage upper limit value corresponding to the usage rate upper limit value is generated based on historical experience data or an automated algorithm.
[0094] Optionally, first configuration prompt information of the upper limit value of the usage rate of the target resource may be displayed in the resource information management table.
[0095] Optionally, the recommended upper limit value can be viewed in the resource information management table, and the configuration update operation can be triggered by one or more methods such as clicking a confirmation button or inputting a command.
[0096] Furthermore, in response to the trigger operation, the usage upper limit value of the corresponding target resource in the resource information management table is updated to the recommended usage upper limit value.
[0097] By displaying configuration prompt information in the resource information management table, optimization suggestions based on actual usage are provided, enhancing the intelligence and foresight of resource configuration; supporting a dynamic and flexible resource adjustment mechanism, automatically updating the usage rate upper limit in response to trigger operations, achieving rapid feedback and dynamic optimization of resource configuration, and improving system flexibility and resource utilization.
[0098] As an optional technical solution of an embodiment of the present invention, optionally, the display of the actual resource usage information includes: for each of the target resources, generating a resource usage curve chart based on multiple actual usage moments of the target resource and the actual usage rate of the target resource at the actual usage moments, and displaying the resource usage curve chart.
[0099] Optionally, for each target resource, target resource usage time series data can be generated based on multiple actual usage moments of the target resource and the actual usage rate of the target resource at the actual usage moments, and the actual usage rate of target resources such as CPU resources, memory resources, network resources and I / O resources can be plotted according to the time dimension to generate and display a resource usage curve corresponding to the target resource.
[0100] Optionally, the resource usage graph can be filtered and displayed by multiple dimensions such as task and time period, making it easier to identify resource bottlenecks and inefficient periods.
[0101] Furthermore, resource usage curves of different target resources may be compared horizontally to identify imbalances in resource usage.
[0102] Optionally, the recommended resource configuration line or warning threshold mark can be superimposed on the resource usage curve graph in combination with the current total resource amount and task load to assist in resource expansion and optimization.
[0103] By generating and displaying a resource usage curve for each target resource, we can intuitively reflect the actual usage of resources at different time points, clearly understand the changing trend of resource consumption, and help identify peak and trough periods of resource usage, thereby providing a basis for more reasonable resource planning, scheduling optimization, and cost control.
[0104] The technical solution of the embodiment of the present invention is as follows: first, by receiving a task execution request for a data processing task, the actual utilization rate of multiple target resources for executing the data processing task and the utilization rate upper limit value of the target resources recorded in the resource information management table are obtained; wherein, the data processing task includes multiple subtasks; the target resources include at least central processing unit resources, memory resources and network resources; by obtaining the utilization rate and utilization rate upper limit value of the target resources involved in the data processing task, it is convenient to screen and allocate subtasks according to the utilization rate upper limit value, so as to avoid the occurrence of problems such as resource overload or resource waste; then, in the case that the actual utilization rate of multiple target resources does not reach the utilization rate upper limit value of the target resource, by performing batch tasks corresponding to the data processing task The task scheduling information of the multiple subtasks recorded in the scheduling table determines at least one executable subtask; since the task scheduling information of the multiple subtasks is recorded in the batch task scheduling table, the executable subtasks can be accurately screened according to the batch task scheduling table, and the order of parallel execution of multiple subtasks can be dynamically arranged according to the real-time resource usage and task demand changes, thereby improving resource utilization; finally, by determining the target subtask from at least one executable subtask and executing the target subtask, the actual utilization rate of at least one target resource is closer and does not exceed the utilization rate upper limit of the target resource; so that the target resource can be fully and reasonably used without exceeding the upper limit threshold, effectively avoiding the problem of excessive or insufficient resource occupation. By reasonably arranging the execution order of subtasks and resource allocation, the overall time of multi-task execution is greatly shortened, and the resource allocation is refined and automated, providing efficient and stable technical support for large-scale data processing and complex business scenarios.
[0105] Example 2
[0106] Figure 2 A flowchart of a scheduling method for data processing tasks provided in the second embodiment of the present invention further refines the implementation method of determining at least one executable subtask based on the task scheduling information of the multiple subtasks recorded in the batch task scheduling table corresponding to the data processing task. Optionally, the task scheduling information at least includes the task identification, task dependency information and task execution status of the subtask; the determination of at least one executable subtask based on the task scheduling information of the multiple subtasks recorded in the batch task scheduling table corresponding to the data processing task includes: determining at least one executable subtask based on the task identification, task dependency information and task execution status of the multiple subtasks recorded in the batch task scheduling table corresponding to the data processing task. For specific implementation methods, please refer to the description of this embodiment. Among them, the technical features that are the same or similar to those in the aforementioned embodiments are not repeated here. Figure 2As shown, the method may specifically include:
[0107] S210: Receive a task execution request for a data processing task, and obtain actual usage rates of multiple target resources for executing the data processing task and usage upper limits of the target resources recorded in a resource information management table.
[0108] S220. When the actual utilization rates of multiple target resources do not reach the utilization upper limit of the target resources, at least one executable subtask is determined based on the task identifiers, task dependency information and task execution status of multiple subtasks recorded in the batch task scheduling table corresponding to the data processing task.
[0109] The batch task scheduling table records the task scheduling information of the plurality of subtasks. The task scheduling information includes at least the task identifier, task dependency information, and task execution status of the subtask. The task identifier is used to distinguish and identify different subtasks, and can also be said to be a unique identifier for each subtask. The task dependency information can be used to describe the execution order and dependency relationship between subtasks. The task execution status can be used to record the current execution status of the subtask. For example, the task execution status may include one or more execution states such as pending execution, executing, execution suspended, and execution completed.
[0110] In some embodiments, the task scheduling information may further include scheduling time information. Optionally, the scheduling time information may specifically include a scheduling start time and / or a call end time. In an embodiment of the present invention, the scheduling time information may include scheduling time information under one or more statistical dimensions. For example, the scheduling time information may specifically include one or more information such as a scheduling date, a scheduling duration, and a scheduling time point.
[0111] Optionally, a periodic resource monitoring mechanism can be activated to collect the actual utilization rates of various target resources, such as CPU resources, memory resources, and network resources, in real time at preset time intervals. The actual utilization rates are then compared one by one with the utilization upper limits of the corresponding target resources recorded in the resource information management table. If the actual utilization rates of the various target resources for the data processing task are all lower than the utilization upper limits, a task scheduling information query operation is triggered in the batch task scheduling table to determine at least one executable subtask.
[0112] Optionally, the presence of schedulable resource space is determined when and only when the following conditions are simultaneously met: the actual utilization rate of the CPU resource is less than the upper limit of the CPU utilization rate in the resource information management table, the actual utilization rate of the memory resource is less than the upper limit of the memory resource utilization rate in the resource information management table, and the actual utilization rate of the network resource is less than the upper limit of the network resource utilization rate in the resource information management table. If the actual utilization rates of the above target resources are all less than the upper limit of the utilization rate, it indicates that there are underutilized target resources, and at least one executable subtask can be screened out based on the task scheduling information of the multiple subtasks recorded in the batch task scheduling table corresponding to the data processing task. The executable subtask at least simultaneously meets the preset scheduling conditions of being in a pending execution state and having no recursive dependency relationship with the task being executed.
[0113] Furthermore, when any of the above conditions is not met simultaneously, it is determined that there is no schedulable resource space. In this case, the operation of obtaining the actual usage rates of the multiple target resources of the data processing task can be re-executed after waiting for a preset time.
[0114] Optionally, the task scheduling information corresponding to each subtask is read from the batch task scheduling table, and the task identifier, task dependency information and task execution status of each subtask are read according to the task scheduling information, and the subtasks that are in a pending state and have no recursive dependency relationship with the currently executing subtask are filtered out as executable subtasks.
[0115] Optionally, the resource detection and task scheduling process can be continuously iterated. Once idle resources are detected, the above-mentioned screening logic for executable subtasks is immediately triggered until all subtasks of the data processing task are scheduled and executed.
[0116] Optionally, if the actual utilization of one or more target resources, such as CPU resources, memory resources, and network resources, exceeds the preset utilization limit for the corresponding target resource during the parallel execution of tasks, some non-critical executable subtasks will be automatically marked as suspended. Once the remaining subtasks have completed and sufficient target resources have been released, the suspended subtasks will be rescheduled to ensure that resource usage remains under control.
[0117] By accurately screening target subtasks, dynamic allocation of target resources can be achieved. While ensuring the safety boundaries of resource use, resource utilization can be maximized to avoid idle and wasted resources, effectively balancing task execution efficiency and system stability.
[0118] S230: Determine a target subtask from at least one of the executable subtasks, and execute the target subtask so that the actual usage rate of at least one of the target resources is closer to and does not exceed the usage rate upper limit of the target resource.
[0119] The technical solution of the embodiment of the present invention is as follows: first, by receiving a task execution request for a data processing task, the actual utilization rate of multiple target resources for executing the data processing task and the utilization upper limit value of the target resources recorded in the resource information management table are obtained; the actual utilization rate and the upper limit value are obtained to provide real-time data support for subsequent resource evaluation and task scheduling, to ensure that subsequent scheduling decisions are made within a safe range, and to realize dynamic scheduling based on actual resource conditions; then, when the actual utilization rates of multiple target resources do not reach the utilization upper limit value of the target resources, at least one executable subtask is determined according to the task identification, task dependency information and task execution status of multiple subtasks recorded in the batch task scheduling table corresponding to the data processing task; subtask screening is triggered only when resources are not fully utilized, to avoid blind scheduling. Resource overload, improve overall resource utilization, accurately identify currently executable subtasks based on the batch task scheduling table, enhance the rationality and controllability of the scheduling logic, and improve scheduling efficiency; finally, by determining the target subtask from at least one of the executable subtasks and executing the target subtask, so that the actual utilization rate of at least one of the target resources is closer and does not exceed the utilization upper limit of the target resource; select the optimal subtask and execute it through the intelligent screening mechanism to maximize resource utilization efficiency and ensure that resource utilization is increased as much as possible without breaking the resource upper limit; through real-time perception of resource status, combined with resource information management table and batch task scheduling table for analysis, and using intelligent screening mechanism to dynamically execute subtasks, the goal of efficient resource utilization while ensuring system stability is achieved, and the intelligence level of task scheduling and resource utilization are improved.
[0120] Example 3
[0121] As an optional example of an embodiment of the present invention, the method for scheduling data processing tasks in an embodiment of the present invention may specifically include:
[0122] (1) The resource information management table is used to configure and uniformly manage the utilization rate upper limit of target resources such as CPU resources, memory resources, and network resources. The resource information management table mainly includes fields such as resource type, utilization rate upper limit, whether enabled, configuration personnel, configuration date, and configuration time.
[0123]
[0124] Table 1 Resource information management table
[0125] The Resource Type field contains commonly used resources, such as CPU resources, memory resources, and network resources. The Utilization Limit field contains the usage limit of the corresponding target resource. For example, if the total memory resource is 2TB and the memory resource utilization limit is configured as 90%, the memory resource usage limit is 2TB * 90% = 1.8TB.
[0126] (2) Dynamically arrange the execution order of multiple tasks based on available target resources. A batch task scheduling table is set up to uniformly manage the scheduling of batch subtasks. The batch task scheduling table contains fields such as task number, all recursively dependent task numbers, scheduling date, scheduling time, task status, scheduling completion date, and scheduling completion time, as shown in Table 2:
[0127]
[0128] Table 2 Batch task scheduling table
[0129] The dependencies between different task nodes in Table 2 are as follows: Figure 3b shown.
[0130] When multiple tasks are executed concurrently, the system queries in real time whether the used CPU resources, memory resources, network resources and other resources are lower than the upper limit of the CPU resources, memory resources, network resources and other resources currently available for calculation. If they are lower than the upper limit of the CPU resources, memory resources, network resources and other resources available for calculation, the system queries the batch task scheduling table, selects the appropriate computing task according to the task status and starts execution. The flowchart of dynamic arrangement of data processing tasks, such as Figure 3a shown.
[0131] The specific steps for dynamic orchestration of data processing tasks are as follows:
[0132] 1. Receive a task execution request for a data processing task, and multi-task parallel scheduling begins.
[0133] 2. Real-time capture of the CPU resources, memory resources, network resources and other target resources used by the data processing task to see if they are lower than the corresponding resource usage limit in the resource information management table.
[0134] If and only if:
[0135] The actual CPU resource usage is less than the upper limit of the CPU resource usage in the resource information management table;
[0136] The actual memory resource usage is less than the upper limit of memory resource usage in the resource information management table;
[0137] The actual network resource usage rate is less than the upper limit of the network resource usage rate in the resource information management table;
[0138] If all the above conditions are met, it means that there are still underutilized resources, and step 3 is executed. Otherwise, wait for a period of time (for example, 1 minute) and continue to step 23. Determine at least one executable subtask based on the task scheduling information of multiple subtasks recorded in the batch task scheduling table corresponding to the data processing task. The following conditions must be met at the same time:
[0139] The subtask is in the pending state.
[0140] There are no recursive dependencies on the executing subtasks
[0141] After selecting and invoking a target subtask from at least one executable subtask, the batch task scheduler updates the task's schedule date, schedule time, task status, and other fields. During the parallel execution of subtasks, the resource usage monitoring program periodically records the total amount of CPU resources, memory resources, and network resources consumed by the parallel execution tasks in the execution log.
[0142] 4. Check whether all subtasks have been completed. If there are still scheduled tasks to be executed, jump to step 2; otherwise, go to step 5.
[0143] 5. Multi-task parallel scheduling ends.
[0144] An example of the dynamic orchestration process of data processing tasks is as follows:
[0145] 1) First, read the metadata information in the batch task schedule table and call only the a01 node for execution. The task node execution diagram is as follows: Figure 3c as shown (gray means it is being executed).
[0146] 2) After a01 starts executing, regular detection finds that the total amount of used CPU resources, memory resources, network resources, etc. is lower than the resource utilization limit in the resource information management table. The batch task scheduling table is queried. The b01 node has no recursive dependency relationship with the executing task a01. The b01 task is immediately called. The task node execution diagram at this time is as follows: Figure 3d shown (gray means it is being executed).
[0147] 3) After a01 and b01 start executing, regular detection finds that the total amount of used CPU resources, memory resources, network resources, etc. is lower than the resource utilization upper limit in the resource information management table. The batch task scheduling table is queried. The c01 node has no recursive dependency relationship with the executing tasks a01 and b01. The c01 task is immediately called up. The task node execution diagram at this time is as follows: Figure 3e shown (gray means it is being executed).
[0148] The same process continues until all subtasks are scheduled and executed. In extreme cases, when multiple tasks are executing in parallel and the CPU, memory, and network resource usage exceeds the upper limits of these resources, some subtasks will enter a waiting state. Only after the completed tasks have released sufficient CPU, memory, and network resources will the waiting subtasks be reactivated.
[0149] (3) Analyze the actual resource consumption logs of multi-task execution, optimize the resource utilization upper limit configuration, and fill in some resource gaps. After all tasks are completed on the day, analyze the resource usage logs of CPU resources, memory resources, and network resources consumed by the task execution to form different resource usage curves. Based on the existing resources and the amount of big data processing tasks, dynamically optimize the resource information management table. At the same time, expand the hardware resources of some resource gaps in the different resource usage curves to achieve efficient utilization of all existing resources and shorten the overall time used to process tasks as much as possible.
[0150] The technical solution of the embodiment of the present invention realizes refined and automated resource allocation by real-time statistics of various resource usage and dynamic scheduling of the order of parallel execution of multiple tasks, thereby shortening the overall execution time of multiple tasks and improving resource utilization efficiency.
[0151] Example 4
[0152] Figure 4 This is a schematic diagram of the structure of a data processing task scheduling device provided in the fourth embodiment of the present invention. This device and the data processing task scheduling method of the above embodiments belong to the same inventive concept. For details not fully described in the embodiment of the data processing task scheduling device, please refer to the embodiment of the data processing task scheduling method. Figure 4 As shown, the device includes: a target resource usage rate and upper limit value acquisition module 410, an executable subtask determination module 420 and a target subtask determination execution module 430.
[0153] Among them, the target resource utilization rate and upper limit value acquisition module 410 is used to receive a task execution request for a data processing task, obtain the actual utilization rate of multiple target resources for executing the data processing task and the utilization rate upper limit value of the target resources recorded in the resource information management table; wherein, the data processing task includes multiple subtasks; the target resources include at least central processing unit resources, memory resources and network resources; the executable subtask determination module 420 is used to determine at least one executable subtask according to the task scheduling information of multiple subtasks recorded in the batch task scheduling table corresponding to the data processing task when the actual utilization rate of multiple target resources does not reach the utilization rate upper limit value of the target resource; wherein, the task scheduling information of multiple subtasks is recorded in the batch task scheduling table; the target subtask determination execution module 430 is used to determine a target subtask from at least one executable subtask and execute the target subtask so that the actual utilization rate of at least one target resource is closer to and does not exceed the utilization rate upper limit value of the target resource.
[0154] The technical solution of the embodiment of the present invention is as follows: first, a task execution request for a data processing task is received through a target resource utilization rate and upper limit value acquisition module 410, and the actual utilization rates of a plurality of target resources for executing the data processing task and the utilization rate upper limit values of the target resources recorded in a resource information management table are acquired; wherein, the data processing task includes a plurality of subtasks; the target resources include at least central processing unit resources, memory resources and network resources; by acquiring the utilization rates and utilization rate upper limit values of the target resources involved in the data processing task, it is convenient to screen and allocate subtasks according to the utilization rate upper limit values, so as to avoid the occurrence of problems such as resource overload or resource waste; then, through an executable subtask determination module 420, when the actual utilization rates of a plurality of the target resources do not reach the utilization rate upper limit values of the target resources, the subtask is determined according to the utilization rate upper limit values related to the data processing task. At least one executable subtask is determined based on the task scheduling information of the multiple subtasks recorded in the batch task scheduling table corresponding to the task; since the task scheduling information of the multiple subtasks is recorded in the batch task scheduling table, the executable subtasks can be accurately screened according to the batch task scheduling table, and the order of parallel execution of the multiple subtasks can be dynamically arranged according to the real-time resource usage and task demand changes, thereby improving resource utilization; finally, the target subtask is determined from at least one of the executable subtasks through the target subtask determination execution module 430, and the target subtask is executed so that the actual utilization rate of at least one target resource is closer and does not exceed the utilization rate upper limit of the target resource; so that the target resource can be fully and reasonably used without exceeding the upper limit threshold, effectively avoiding the problem of excessive or insufficient resource occupation. By reasonably arranging the execution order and resource allocation of subtasks, the overall time of multi-task execution is greatly shortened, and the resource allocation is refined and automated, providing efficient and stable technical support for large-scale data processing and complex business scenarios.
[0155] Based on the above solution, the apparatus may optionally further include a resource information management table editing module. The module is configured to receive a resource information editing operation on the resource information management table through a resource information editing interface before obtaining the usage upper limit value of the target resource recorded in the resource information management table, and determine the resource information management table based on the resource information editing operation; the resource information management table is configured to record a resource type identifier, a usage upper limit value, an information editor, and an information editing time of at least one resource.
[0156] Based on the above solution, optionally, the task scheduling information includes at least the task identifier, task dependency information and task execution status of the subtask.
[0157] Based on the above solution, optionally, the executable subtask determining module 420 includes an executable subtask determining submodule, wherein the executable subtask determining submodule is configured to determine at least one executable subtask based on the task identifiers, task dependency information, and task execution status of the plurality of subtasks recorded in the batch task scheduling table corresponding to the data processing task.
[0158] Based on the above solution, the apparatus may optionally further include a batch task scheduling table generation module, wherein the batch task scheduling table generation module is configured to parse the data processing task to obtain multiple subtasks corresponding to the data processing task and task scheduling information of the multiple subtasks before determining at least one executable subtask based on the task scheduling information of the multiple subtasks recorded in the batch task scheduling table corresponding to the data processing task, and generate a batch task scheduling table corresponding to the data processing task based on the task scheduling information of the multiple subtasks.
[0159] Based on the above solution, the apparatus may optionally further include a batch task scheduling table updating module, wherein the batch task scheduling table updating module is configured to receive a scheduling information editing operation for at least one item of task scheduling information of the plurality of subtasks recorded in the batch task scheduling table corresponding to the data processing task before determining at least one executable subtask based on the task scheduling information of the plurality of subtasks recorded in the batch task scheduling table, and update the batch task scheduling table according to the scheduling information editing operation.
[0160] Based on the above solution, the apparatus may optionally further include an actual resource usage information determination module. The actual resource usage information determination module is configured to, after executing the target subtask and when the data processing task has been completed, analyze the task execution log corresponding to the data processing task to obtain actual resource usage information of the multiple target resources during the execution of the data processing task, and display the actual resource usage information; wherein the actual resource usage information includes actual usage rates of the multiple target resources at the actual usage moment.
[0161] Based on the above solution, the apparatus may optionally further include a resource configuration prompt information generation module. The resource configuration prompt information generation module is configured to, after obtaining actual resource usage information of the plurality of target resources during the execution of the data processing task, generate resource configuration prompt information corresponding to the data processing task based on the actual resource usage information, and display the resource configuration prompt information; wherein the resource configuration prompt information includes first configuration prompt information for an upper limit value of the utilization rate of the target resource and / or second configuration prompt information for hardware resources.
[0162] Based on the above solution, optionally, the resource configuration prompt information includes first configuration prompt information of a usage upper limit value of the target resource; the first configuration prompt information includes a recommended usage upper limit value corresponding to the usage upper limit value.
[0163] Based on the above solution, the apparatus may optionally further include a usage upper limit value and / or a usage upper limit value updating submodule. The usage upper limit value display submodule is configured to display first configuration prompt information of the usage upper limit value of the target resource in the resource information management table after obtaining actual resource usage information of the plurality of target resources during the execution of the data processing task; and the usage upper limit value updating submodule is configured to update the usage upper limit value of the target resource in the resource information management table to the recommended usage upper limit value in response to a usage triggering operation of the recommended usage upper limit value for at least one target resource after obtaining actual resource usage information of the plurality of target resources during the execution of the data processing task.
[0164] Based on the above solution, optionally, the actual resource usage information determination module further includes a resource usage curve graph generation submodule. The resource usage curve graph generation submodule is configured to generate a resource usage curve graph for each target resource based on multiple actual usage moments of the target resource and the actual usage rate of the target resource at the actual usage moments, and display the resource usage curve graph.
[0165] The data processing task scheduling device provided in the embodiment of the present invention can execute the data processing task scheduling method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0166] Example 5
[0167] Figure 5A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0168] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0169] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0170] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors for running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for scheduling data processing tasks.
[0171] In some embodiments, the method for scheduling data processing tasks can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for scheduling data processing tasks described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the method for scheduling data processing tasks in any other appropriate manner (for example, by means of firmware).
[0172] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0173] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0174] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0175] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0176] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0177] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0178] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication unit 19, or installed from the storage unit 18, or installed from the ROM 12. When the computer program is executed by the processor 11, the above-mentioned functions defined in the method of the embodiment of the present invention are performed.
[0179] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0180] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for scheduling data processing tasks, characterized in that: include: receiving a task execution request for a data processing task, and obtaining actual usage rates of multiple target resources for executing the data processing task and upper usage limits of the target resources recorded in a resource information management table; wherein the data processing task includes multiple subtasks; and the target resources include at least central processing unit resources, memory resources, and network resources; If the actual usage rates of the plurality of target resources do not reach the usage rate upper limit of the target resources, determining at least one executable subtask based on task scheduling information of the plurality of subtasks recorded in a batch task scheduling table corresponding to the data processing task; wherein the task scheduling information of the plurality of subtasks is recorded in the batch task scheduling table; A target subtask is determined from at least one of the executable subtasks, and the target subtask is executed so that the actual usage rate of at least one of the target resources is closer to and does not exceed the usage rate upper limit of the target resource.
2. The method for scheduling data processing tasks according to claim 1, wherein: Before obtaining the upper limit value of the utilization rate of the target resource recorded in the resource information management table, the method further includes: A resource information editing operation for a resource information management table is received through a resource information editing interface, and a resource information management table is determined based on the resource information editing operation; wherein, the resource information management table is used to record the resource type identifier, usage upper limit value, information editor and information editing time of at least one resource.
3. The method for scheduling data processing tasks according to claim 1, wherein: The task scheduling information includes at least the task identifier, task dependency information and task execution status of the subtask; The determining at least one executable subtask according to the task scheduling information of the plurality of subtasks recorded in the batch task scheduling table corresponding to the data processing task includes: At least one executable subtask is determined according to the task identifiers, task dependency information and task execution status of the plurality of subtasks recorded in the batch task scheduling table corresponding to the data processing task.
4. The method for scheduling data processing tasks according to claim 1, wherein: Before determining at least one executable subtask according to the task scheduling information of the plurality of subtasks recorded in the batch task scheduling table corresponding to the data processing task, the method further includes: The data processing task is parsed to obtain a plurality of subtasks corresponding to the data processing task and task scheduling information of the plurality of subtasks, and a batch task scheduling table corresponding to the data processing task is generated according to the task scheduling information of the plurality of subtasks.
5. The method for scheduling data processing tasks according to claim 1, wherein: Before determining at least one executable subtask according to the task scheduling information of the plurality of subtasks recorded in the batch task scheduling table corresponding to the data processing task, the method further includes: A scheduling information editing operation for at least one task scheduling information of the plurality of subtasks recorded in the batch task scheduling table is received, and the batch task scheduling table is updated according to the scheduling information editing operation.
6. The method for scheduling data processing tasks according to claim 1, wherein: After executing the target subtask, the method further includes: When the data processing task has been completed, the task execution log corresponding to the data processing task is analyzed to obtain actual resource usage information of the multiple target resources during the execution of the data processing task, and the actual resource usage information is displayed; wherein, the actual resource usage information includes the actual usage rate of multiple target resources at the actual usage time.
7. The method for scheduling data processing tasks according to claim 6, characterized in that: After obtaining actual resource usage information of the plurality of target resources during the execution of the data processing task, the method further includes: Generate resource configuration prompt information corresponding to the data processing task based on the actual resource usage information, and display the resource configuration prompt information; wherein, the resource configuration prompt information includes first configuration prompt information for the upper limit value of the utilization rate of the target resource and / or second configuration prompt information for the hardware resource.
8. The method for scheduling data processing tasks according to claim 7, characterized in that: The resource configuration prompt information includes first configuration prompt information of the usage upper limit value of the target resource; the first configuration prompt information includes a recommended usage upper limit value corresponding to the usage upper limit value; After obtaining actual resource usage information of the plurality of target resources during the execution of the data processing task, the method further includes: Displaying first configuration prompt information of the upper limit value of the utilization rate of the target resource in the resource information management table; and / or, In response to a use triggering operation for the recommended upper limit value of at least one of the target resources, the usage rate upper limit value of the target resource in the resource information management table is updated to the recommended upper limit value.
9. The method for scheduling data processing tasks according to claim 6, wherein: The displaying of the actual resource usage information includes: For each target resource, a resource usage curve graph is generated according to a plurality of actual usage moments of the target resource and the actual usage rate of the target resource at the actual usage moments, and the resource usage curve graph is displayed.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the method for scheduling data processing tasks according to any one of claims 1 to 9.