Method and device for determining task scheduling time
By automatically calculating the remaining amount of cluster resources and task scheduling coefficient, the problem of difficult to accurately determine the task scheduling time in the existing technology is solved, and the automation and intelligence level of task scheduling is improved.
Patent Information
- Application Number
- CN202510069357.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-30
AI Technical Summary
The existing big data task scheduling framework relies on manual collection of reference information during task orchestration, which increases the complexity and uncertainty of scheduling time, making it difficult to accurately determine the task scheduling time.
By obtaining the resource usage of multiple historical clusters, the total amount of current cluster resources, the resource requirements and running time of the target task, the cluster resource remaining amount and task scheduling coefficient at each time point are automatically calculated to determine the most suitable task scheduling time.
It reduces the need for manual participation and judgment, improves the automation and intelligence level of task scheduling, can determine the task scheduling time more accurately, and reduces the scheduling complexity and uncertainty.
Smart Images

Figure CN120066707A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data task scheduling, and particularly to a method and device for determining task scheduling time. Background Art
[0002] With the continuous development of big data technology, the rapid growth of data volume and the increasing complexity of data processing requirements, task scheduling and task management have become an indispensable part of big data processing systems. During the big data processing process, the dependencies between various tasks (such as data collection, data cleaning, data storage, data analysis, etc.) form a complex workflow.
[0003] Currently, there are various big data task scheduling frameworks on the market, such as DolphinScheduler, Oozie, and Azkaban, etc. These frameworks all provide task management functions, allowing task managers to define, schedule, and monitor tasks through interfaces or scripts. Among them, during the task orchestration process of big data task scheduling frameworks, task managers often collect various reference information by their own technical means to plan the task scheduling time.
[0004] However, the task orchestration work is basically carried out by task managers collecting various reference information by their own technical means to plan the task scheduling time, which easily leads to an increase in the complexity and uncertainty of task orchestration.
[0005] Therefore, how to accurately determine the task scheduling time has become a problem to be solved. Summary of the Invention
[0006] In view of this, the present invention provides a method and device for determining task scheduling time.
[0007] In a first aspect, the present invention provides a method for determining task scheduling time, the method comprising: obtaining a plurality of historical cluster resource usages, the current total cluster resources, the resource requirement and running time of a target task; determining the remaining cluster resources at each time according to the plurality of historical cluster resource usages at each time and the current total task cluster resources; determining the task scheduling coefficient at each time based on the resource requirement, running time of the target task, and the remaining cluster resources corresponding to each time; and taking the time with the largest task scheduling coefficient among the task scheduling coefficients at each time as the task scheduling time.
[0008] The method for determining the task scheduling time provided in this embodiment is a method for automatically determining the task scheduling time based on multiple historical cluster resource usages, the current total cluster resources, the resource requirements of the target task, and the running time. That is, the optimal scheduling time is calculated through an algorithm, reducing the need for manual participation and judgment, and improving the automation and intelligence levels of scheduling.
[0009] At the same time, by comprehensively considering multiple factors such as multiple historical cluster resource usages, the current total cluster resources, the resource requirements of the target task, and the running time, the remaining cluster resources and the task scheduling coefficient at each time can be determined more accurately, so as to find the most suitable scheduling time.
[0010] In a possible implementation manner, to determine the remaining cluster resources at each time according to multiple historical cluster resource usages at each time and the current total task cluster resources, it includes: detecting the task type of the target task; if the task type of the target task is a new task type, determining the remaining cluster resources at each time according to multiple historical cluster resource usages at each time and the current total task cluster resources; where determining the remaining cluster resources at each time includes: determining the first resource usage value at the target time according to multiple historical cluster resource usages at the target time and the preset weight corresponding to each historical cluster resource usage value; where the target time is any one of each time; determining the remaining cluster resources at the target time according to the difference between the current total task cluster resources and the first resource usage value at the target time.
[0011] The method for determining the task scheduling time provided in this embodiment first detects the task type of the target task, can distinguish the resource demand differences of different task types, and thus formulate a more accurate resource allocation strategy. For the new task type, this method can respond flexibly, avoiding the problem of unreasonable resource allocation caused by unknown or unconsidered task types in the traditional method.
[0012] At the same time, by comprehensively considering multiple historical cluster resource usages at each time and the preset weight corresponding to each historical resource usage value, this method can more accurately predict the first resource usage value at the target time. And by determining the remaining cluster resources according to the difference between the current total task cluster resources and the first resource usage value at the target time, this method can grasp the dynamic changes of the cluster resources in real time and avoid resource idleness or overuse.
[0013] In a possible implementation, based on the multiple historical cluster resource usages at each time and the total current task cluster resources, determine the remaining cluster resources at each time, including: detecting the task type of the target task; if the task type of the target task is a change task type, determine the remaining cluster resources at each time according to the multiple target resource usages at each time, the preset weight corresponding to each target resource usage, and the total current task cluster resources; wherein, the change task type is used to indicate that the currently existing task needs to change the task scheduling time, and determining the remaining cluster resources at each time includes: obtaining the task memory usage corresponding to each historical cluster resource usage at the target time; wherein, the target time is any one of each time; determining the multiple target resource usages at the target time according to the historical cluster resource usage and the task memory usage corresponding to the historical cluster resource usage; determining the second resource usage value at the target time according to the multiple target resource usages at the target time and the preset weight corresponding to each target resource usage; determining the remaining cluster resources at the target time according to the difference between the total current task cluster resources and the second resource usage value at the target time.
[0014] The method for determining the task scheduling time provided in this embodiment can identify and handle the requirements of the change task type, making the task scheduling more flexible. When the task needs to adjust the scheduling time, the remaining cluster resources can be re-evaluated based on the new resource requirements to ensure the smooth execution of the task.
[0015] At the same time, by obtaining the task memory usage corresponding to each historical cluster resource usage at the target time, this method can more accurately reflect the actual resource requirements of the task. This fine-grained resource assessment helps to more precisely determine the remaining cluster resources and avoid insufficient or excessive resource allocation. Moreover, introducing the preset weight to handle the multiple target resource usages takes into account the influence degree of different resource usages on the remaining cluster resources. This weight allocation method makes the resource assessment more reasonable and can more accurately reflect the actual available situation of the cluster resources.
[0016] In a possible implementation, the resource requirements of the target task include memory resource requirements and central processing unit resource requirements. Correspondingly, the remaining cluster resources include remaining memory resources and remaining central processing unit resources. Among them, determining the task scheduling coefficient for each time based on the resource requirements of the target task, the running time, and the remaining cluster resources corresponding to each time includes: determining multiple second target times corresponding to the first target time according to the running time, where the first target time is any one of each time; determining the first area corresponding to the running time according to the memory resource requirements corresponding to the first target time, the memory resource requirements corresponding to each of the multiple second target times, the first target time, and the second target times; determining the first quantity of the memory resource requirements greater than the memory resource requirements among the first target time and the multiple second target times according to the memory resource requirements corresponding to the first target time and the memory resource requirements corresponding to each of the multiple second target times; determining the memory resource matching value for each time according to the first area, the first quantity, and the memory resource requirements; determining the second area corresponding to the running time according to the remaining central processing unit resources corresponding to the first target time, the remaining central processing unit resources corresponding to each of the multiple second target times, the first target time, and the second target times; determining the second quantity of the remaining central processing unit resources greater than the central processing unit resource requirements among the first target time and the multiple second target times according to the remaining central processing unit resources corresponding to the first target time and the remaining central processing unit resources corresponding to each of the multiple second target times; determining the central processing unit resource matching value for each time according to the second area, the second quantity, and the memory resource requirements; and determining the task scheduling coefficient for each time according to the memory resource matching value and the central processing unit resource matching value for each time.
[0017] The method for determining the task scheduling time provided in this embodiment can accurately match the resource requirements of the task for memory and the central processing unit with the actual resource supply situation of the cluster. By calculating the resource matching values of memory and the central processing unit respectively, it can more accurately reflect the resource satisfaction degree of the task at different time points.
[0018] At the same time, by calculating the first area and the second area, and the quantity of the memory and central processing unit resource requirements greater than the resource remaining quantity, this method can more comprehensively evaluate the resource utilization efficiency of the task at different time points. This helps to optimize resource allocation, reduce resource waste, and improve the overall performance of the cluster. And by introducing the running time and multiple second target times, this method can dynamically evaluate the resource occupancy of the task in the time dimension. This helps to identify possible resource bottlenecks during the task execution process, so as to perform resource scheduling and optimization in advance.
[0019] In a possible implementation, determining the task scheduling coefficient for each time according to the memory resource matching value and the central processing unit resource matching value at each time includes: comparing the memory resource matching value and the central processing unit resource matching value at each time, and determining a first matching value and a second matching value from the memory resource matching value and the central processing unit resource matching value; wherein, the first matching value is greater than the second matching value; determining the task scheduling coefficient for each time according to the first matching value, the second matching value, and a preset coefficient weight.
[0020] The method for determining the task scheduling time provided in this embodiment can comprehensively consider the requirements of tasks for two core resources by comparing the memory resource matching value and the central processing unit resource matching value. This helps to ensure that the performance bottleneck will not occur due to the shortage of a certain resource during the execution of the task. Moreover, by determining the first matching value and the second matching value, this method provides a clear priority for resource allocation. The resource type represented by the first matching value (whether it is memory or the central processing unit) will be regarded as the key resource for task execution, and thus will obtain a higher priority during resource scheduling.
[0021] In a possible implementation, determining the task scheduling coefficient for each time according to the first matching value, the second matching value, and a preset coefficient weight includes:
[0022] z = h×k+(1 - h)×o; where z is the task scheduling coefficient, h is the preset coefficient weight, k is the first matching value, and o is the second matching value.
[0023] In a second aspect, the present invention provides an apparatus for determining the task scheduling time, and the apparatus includes: an acquisition module, configured to acquire a plurality of historical cluster resource usage amounts, the current total cluster resources, the resource requirements and the running time of the target task; a first determination module, configured to determine the remaining cluster resources for each time according to the plurality of historical cluster resource usage amounts and the current task cluster resources at each time; a second determination module, configured to determine the task scheduling coefficient for each time based on the resource requirements, the running time of the target task, and the remaining cluster resources corresponding to each time; a third determination module, configured to use the time with the largest task scheduling coefficient among the task scheduling coefficients for each time as the task scheduling time.
[0024] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other, and a computer instruction is stored in the memory, and the processor executes the computer instruction to execute the method for determining the task scheduling time according to the first aspect or any corresponding implementation manner thereof.
[0025] Fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method for determining task scheduling time in the above first aspect or any corresponding embodiment thereof.
[0026] Fifth aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the method for determining task scheduling time in the above first aspect or any corresponding embodiment thereof. Description of the Drawings
[0027] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0028] Figure 1 It is a flowchart of the method for determining task scheduling time according to an embodiment of the present invention;
[0029] Figure 2 It is a schematic diagram of the daily dimension available resource biaxial curve provided by an embodiment of the present invention;
[0030] Figure 3 It is a schematic diagram of the task scheduling coefficient comparison provided by an embodiment of the present invention;
[0031] Figure 4 It is a schematic diagram of the method for determining task scheduling time according to an embodiment of the present invention;
[0032] Figure 5 It is a structural block diagram of the device for determining task scheduling time according to an embodiment of the present invention;
[0033] Figure 6 It is a schematic diagram of the hardware structure of the computer device according to an embodiment of the present invention. Detailed Embodiments
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0035] Based on related technologies, at present, there are various big data task scheduling frameworks in the market, such as DolphinScheduler, Oozie, and Azkaban, etc. These frameworks all provide task management functions, allowing task managers to define, schedule, and monitor tasks through interfaces or scripts. Among them, in the process of task orchestration for big data task scheduling frameworks, task managers often collect various reference information by their own technical means to plan the task scheduling time.
[0036] However, the task orchestration work is basically for task managers to collect various reference information by their own technical means to plan the task scheduling time, which easily leads to an increase in the complexity and uncertainty of task orchestration.
[0037] Based on this, the present invention provides a method for determining the task scheduling time, which automatically determines the task scheduling time based on multiple historical cluster resource usage amounts, the current total cluster resources, the resource requirements and running time of the target task. That is, the optimal scheduling time is calculated through an algorithm, reducing the need for manual participation and judgment, and improving the automation and intelligence level of scheduling. At the same time, by comprehensively considering multiple factors such as multiple historical cluster resource usage amounts, the current total cluster resources, the resource requirements and running time of the target task, etc., it is possible to more accurately determine the remaining cluster resources and task scheduling coefficients at each time, so as to find the most suitable scheduling time.
[0038] According to an embodiment of the present invention, there is provided an embodiment of a method for determining the task scheduling time. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0039] In this embodiment, a method for determining the task scheduling time is provided, which can be used in computer devices such as computers, servers, etc. Figure 1 It is a flowchart of the method for determining the task scheduling time according to an embodiment of the present invention, as Figure 1 shown, and this process includes the following steps:
[0040] Step S101, obtain multiple historical cluster resource usage amounts, the current total cluster resources, the resource requirements and running time of the target task.
[0041] The historical cluster resource usage can represent the usage of cluster resources over a past period of time. Among them, the historical cluster resource usage can include any one or more of: data collection time, cluster CPU usage, cluster memory usage, task name, CPU resources used by the task, and memory resources used by the task. Specifically, the historical cluster resource usage can be directly extracted from a database, etc., and no specific limitation is made here, and it can be implemented by those skilled in the art as the standard.
[0042] The current total cluster resources can represent the total resources of the cluster currently in use. Among them, the current total cluster resources can be A1 or A2, and can be set manually by the user, and no specific limitation is made here.
[0043] The resource requirements of the target task can represent the cluster memory occupancy value and CPU resource occupancy value required when executing the target task, and can be set manually by the user, etc. The running time can represent the running time of the target task input by the user.
[0044] Step S102, determine the remaining cluster resources at each time according to the multiple historical cluster resource usages at each time and the current total task cluster resources.
[0045] The multiple historical cluster resource usages at each time can represent the multiple historical cluster resource usages at the same time. Among them, each time can be determined by 24 hours from 0 to 23 o'clock. Among them, each hour can have a collection time of 5 minutes, so there are a total of 288 collection points corresponding to each time, that is, multiple historical cluster resource usages can be obtained at each collection point corresponding to each time.
[0046] More specifically, the multiple historical cluster resource usages can be the historical cluster resource usages for the past 15 days, or the historical cluster resource usages for the past 16 days, etc., and no specific limitation is made here.
[0047] As an example, the method for determining the remaining cluster resources at each time according to the multiple historical cluster resource usages at each time and the current total task cluster resources can be:
[0048] Determine the remaining cluster resources at this time according to the historical cluster resource usages for the past 15 days at the collection point corresponding to any one time and the current total task cluster resources. For example: if this time is 11:25, then the historical cluster resource usages in the historical cluster resource usages for the past 15 days are all the historical cluster resource usages at 11:25, and then determine the remaining cluster resources at each time according to the historical cluster resource usages and the current total task cluster resources.
[0049] Please refer to Figure 2 ,Figure 2 It is a schematic diagram of the daily - dimension available resource biaxial curve provided according to an embodiment of the present invention.
[0050] As an example, the resource demand of the target task includes the memory resource demand and the central - processing - unit (CPU) resource demand. Correspondingly, the remaining cluster resources include the remaining memory resources and the remaining CPU resources. Among them, the daily - dimension available resource biaxial curve can be constructed based on the remaining memory resources and the remaining CPU resources. Among them, the horizontal axis is time, and the two vertical axes are the remaining memory resources and the remaining CPU resources respectively.
[0051] Step S103: Determine the task scheduling coefficient at each time based on the resource demand of the target task, the running time, and the remaining cluster resources corresponding to each time.
[0052] After determining the running time, the acquisition time period corresponding to the running time can be determined according to the running time and the above - determined acquisition interval, that is, how long the target task needs to run. After determining the acquisition time period, the remaining cluster resources at each time in the acquisition time period can be determined, and then, based on the remaining cluster resources at each time and in combination with the resource demand of the target task, the task scheduling coefficient at each time is determined.
[0053] As an example, the remaining cluster resources with positive ratios can be determined according to the ratio between the remaining cluster resources at each time and the resource demand of the target task. According to the running time and the remaining cluster resources at each time, the area on the daily - dimension available resource biaxial curve is determined, and then, based on the area on the daily - dimension available resource biaxial curve and the number of remaining cluster resources with positive ratios, the task scheduling coefficient at each time is determined.
[0054] Step S104: Use the time with the largest task scheduling coefficient among the task scheduling coefficients at each time as the task scheduling time.
[0055] As can be seen from the above, a task scheduling coefficient can be determined for one time. Among them, the task scheduling coefficient can represent the task scheduling score at this time. When the task scheduling score is higher, it is better to perform task scheduling for the target task at this time. Then, the time with the largest task scheduling coefficient among the task scheduling coefficients at each time can be used as the task scheduling time.
[0056] The method for determining the task scheduling time provided in this embodiment is a method for automatically determining the task scheduling time based on multiple historical cluster resource usage amounts, the current total cluster resources, the resource demand of the target task, and the running time. That is, the optimal scheduling time is calculated through an algorithm, reducing the need for manual participation and judgment, and improving the automation and intelligence level of scheduling.
[0057] Meanwhile, by comprehensively considering multiple factors such as the historical cluster resource usage amounts at multiple times, the current total cluster resources, the resource requirements of the target task, and the running time, the remaining cluster resources and task scheduling coefficients at each time can be determined more precisely, so as to find the most suitable scheduling time.
[0058] In a possible implementation manner, the above step S102 includes:
[0059] Step a1, detecting the task type of the target task.
[0060] The target task can represent the task to be performed. Among them, the target task can be a new task or an existing task that requires time change. After obtaining the target task input by the user, the task type of the target task can be detected to determine whether the target task belongs to a new task or an existing task that requires time change.
[0061] Step a2, if the task type of the target task is a new task type, determining the remaining cluster resources at each time according to the historical cluster resource usage amounts at multiple times and the current total cluster resources of the task; among them, determining the remaining cluster resources at each time includes: determining the first resource usage value at the target time according to the historical cluster resource usage amounts at the target time and the preset weights corresponding to each historical cluster resource usage amount; where the target time is any one of each time.
[0062] The target time is any one of each time. If the task type of the target task is a new task type, the first resource usage value at the target time can be determined according to the historical cluster resource usage amounts at the target time and the preset weights corresponding to each historical cluster resource usage amount.
[0063] As an example, each historical cluster resource usage amount corresponds to a preset weight. For example: take the cluster memory usage amount α at the target time in the recent 15 days t , α t-1 ,...α t-14 (α t is the memory usage amount yesterday, and it is recursively pushed forward one day in time sequence. If the number of newly built collection system points is not enough, then take the number that can be obtained), w t is 1, w t-1 , w t-2 , w t-3 , w t-4 are all 0.3, w t-5 , w t-6 , w t-7 , w t-8 , w t-9 , w t-10 , wt-11 , w t-12 , w t-13 , w t-14 All are 0.1, and the first resource usage value at the target time can be determined through the following formula:
[0064] where a is the first resource usage value at the target time, w is the weight, and α i is the cluster memory usage at the same time point every day in the historical days, and t is the number of historical days.
[0065] Step a3: Determine the remaining cluster resources at the target time according to the difference between the total current task cluster resources and the first resource usage value at the target time.
[0066] After determining the first resource usage values of multiple historical cluster resource usages, the total current task cluster resources can be subtracted by the first resource usage value to obtain the remaining cluster resources at the current target time. For example: The total current task cluster resources is M, and the remaining cluster resources at the target time is v: v = M - a.
[0067] It should be noted that the remaining cluster resources may include: the remaining memory resources and the remaining central processing unit resources, that is, through the above steps a1 and a3, the remaining memory resources and the remaining central processing unit resources corresponding to each time can be determined.
[0068] The method for determining the task scheduling time provided in this embodiment first detects the task type of the target task, can distinguish the resource demand differences of different task types, and thus formulate a more accurate resource allocation strategy. For new task types, this method can flexibly respond and avoid the problem of unreasonable resource allocation caused by unknown or unconsidered task types in traditional methods.
[0069] At the same time, by comprehensively considering multiple historical cluster resource usages at each time and the preset weights corresponding to each historical resource usage, this method can more accurately predict the first resource usage value at the target time. And by determining the remaining cluster resources according to the difference between the total current task cluster resources and the first resource usage value at the target time, this method can grasp the dynamic changes of the cluster resources in real time and avoid resource idleness or overuse.
[0070] In a possible implementation manner, the above step S102 includes:
[0071] Step b1: Detect the task type of the target task.
[0072] The target task can represent the task to be performed. Among them, the target task can be a new task or an existing task that requires a time change. After obtaining the target task input by the user, the task type of the target task can be detected to determine whether the target task is a new task or an existing task that requires a time change.
[0073] Step b2, if the task type of the target task is a task type change, determine the remaining cluster resources at each time according to the multiple target resource usages at each time, the preset weight corresponding to each target resource usage, and the total current task cluster resources; where the task type change is used to indicate that an existing task needs to change the task scheduling time, and determining the remaining cluster resources at each time includes: obtaining the task memory usage corresponding to each historical cluster resource usage at the target time; where the target time is any one of each time.
[0074] The target time is any one of each time. If the task type of the target task is a task type change, it is necessary to obtain the task memory usage corresponding to each historical cluster resource usage at the target time. Among them, the task memory usage can represent the actual memory usage that has not been generated. That is, if the task type of the target task is a task type change and the target task needs to be adjusted to other times for task scheduling, then the target task does not occupy memory.
[0075] As an example, in the case of a target task that needs to be changed, the memory usage ɡ of the target task that needs to be changed is correspondingly obtained t , ɡ t-1 ,...ɡ t-14 (ɡ t is the memory usage of this task yesterday. Pushing one day forward in time in sequence, if the number of newly built collection system points is insufficient, then take the number that can be obtained), and make corresponding corrections to the historical cluster resource usages.
[0076] Step b3, determine the multiple target resource usages at the target time according to the historical cluster resource usages and the task memory usages corresponding to the historical cluster resource usages.
[0077] After determining the historical cluster resource usages and the task memory usages corresponding to the historical cluster resource usages, the multiple target resource usages at the target time can be determined according to the difference between the historical cluster resource usages and the task memory usages corresponding to the historical cluster resource usages.
[0078] As an example, a t’ = a t - g t where a t’ is the target resource usage, at is the historical cluster resource usage, in g t is the task memory usage corresponding to the historical cluster resource usage.
[0079] Step b4: Determine the second resource usage value at the target time according to multiple target resource usages at the target time and the preset weight corresponding to each target resource usage.
[0080] Each historical cluster resource usage corresponds to a preset weight. For example, take the cluster memory usage α at the target time in the most recent 15 days t , α t-1 ,... α t-14 (α t is the memory usage of yesterday, and if the number of newly built collection system points is not enough when pushing one day forward in time in sequence, take the number that can be obtained), w t is 1, w t-1 , w t-2 , w t-3 , w t-4 are all 0.3, w t-5 , w t-6 , w t-7 , w t-8 , w t-9 , w t-10 , w t-11 , w t-12 , w t-13 , w t-14 are all 0.1. The first resource usage value at the target time can be determined by the following formula:
[0081] where a is the second resource usage value at the target time, w is the weight, α i is the target resource usage at the same time point every day in the historical days, and t is the number of historical days.
[0082] Step b5: Determine the remaining cluster resources at the target time according to the difference between the total current task cluster resources and the second resource usage value at the target time.
[0083] After determining the second resource usage value of multiple historical cluster resource usages, the total current task cluster resources can be subtracted by the second resource usage value to obtain the remaining cluster resources at the current target time. For example: The total current task cluster resources is M, and the remaining cluster resources at the target time is v: v = M - a.
[0084] The method for determining the task scheduling time provided in this embodiment can identify and process the need to change the task type, making the task scheduling more flexible. When the task needs to adjust the scheduling time, the remaining cluster resources can be re-evaluated based on the new resource requirements to ensure the smooth execution of the task.
[0085] Meanwhile, by obtaining the task memory usage corresponding to each historical cluster resource usage at the target time, this method can more accurately reflect the actual resource requirements of the task. This fine-grained resource assessment helps to more precisely determine the remaining cluster resources, avoiding under-allocation or over-allocation of resources. Moreover, introducing a preset weight to handle multiple target resource usages takes into account the impact degree of different resource usages on the remaining cluster resources. This way of weight allocation makes the resource assessment more reasonable and can more accurately reflect the actual available situation of the cluster resources.
[0086] In a possible implementation, the resource requirements of the target task include memory resource requirements and central processing unit resource requirements. Correspondingly, the remaining cluster resources include remaining memory resources and remaining central processing unit resources. Among them, the above step S103 includes:
[0087] Step c1, determine multiple second target times corresponding to the first target time according to the running time; where the first target time is any one of each time.
[0088] The resource requirements of the target task include memory resource requirements and central processing unit resource requirements. Correspondingly, the remaining cluster resources include remaining memory resources and remaining central processing unit resources. Specifically, the first target time can be any one of each time. As can be seen from the above, the daily dimension available resource biaxial curve consists of the time on the horizontal axis and the remaining memory resources and remaining central processing unit resources on the vertical axis. Among them, a first curve of time and remaining memory resources and a second curve of time and remaining central processing unit resources can be plotted.
[0089] According to the running time and the determined collection interval, the collection time period corresponding to the running time can be determined, that is, how long the target task needs to run. Then, taking the first target time as the starting time, and then according to how long it runs, the number of times can be determined, and then starting from the starting time, the corresponding second target times can be determined. For example: the first target time is t0, and the running time is d minutes, then the number n of the second target times after t0 on the coordinate axis is taken as: n = d÷5.
[0090] Step c2, determine the first area corresponding to the running time according to the memory resource requirements corresponding to the first target time, the memory resource requirements corresponding to each of the multiple second target times, the first target time, and the second target times.
[0091] When determining the memory resource requirement corresponding to the first target time and the memory resource requirement corresponding to each of the second target times in the second target times, the area enclosed by the length of time, the memory resource requirement corresponding to the first target time, and the memory resource requirement corresponding to each of the second target times in the multiple second target times can be determined as the first area corresponding to the running time.
[0092] As an example, the value on the memory curve corresponding to t on the coordinate axis 0 ...t n is v 0 ...v n , and the calculation method of the corresponding first area s is: (assuming that the base value of adjacent times in the formula is 1).
[0093] Where s is the first area.
[0094] Step c3, according to the memory resource requirement corresponding to the first target time and the memory resource requirement corresponding to each of the second target times in the multiple second target times, determine the first quantity of the memory resource requirement greater than the memory resource requirement at the first target time and in the multiple second target times.
[0095] The first quantity can represent the quantity of the memory resource requirement greater than the memory resource requirement. By comparing the memory resource requirement corresponding to the first target time with the memory resource requirement, and comparing the memory resource requirement corresponding to each of the second target times in the multiple second target times with the memory resource requirement, determine the quantity of the memory resource requirement greater than the memory resource requirement.
[0096] As an example, the value on the memory curve corresponding to t on the coordinate axis 0 ..t n-1 is v 0 ...v n-1 , and compare them with the expected memory resource requirement c input by the user in turn to obtain the first quantity p less than this value.
[0097] Step c4, according to the first area, the first quantity, and the memory resource requirement, determine the memory resource matching value for each time.
[0098] As an example, Where x is the memory resource matching value, c is the memory resource requirement, p is the first quantity, and s is the first area.
[0099] Step c5, according to the remaining central processing unit resources corresponding to the first target time, the remaining central processing unit resources corresponding to each of the second target times in the multiple second target times, the first target time, and the second target time, determine the second area corresponding to the running time.
[0100] Step c6: Determine the second quantity of the remaining central processing unit (CPU) resources at the first target time and each of the multiple second target times, where the remaining CPU resources are greater than the CPU resource demand, based on the remaining CPU resources corresponding to the first target time and the remaining CPU resources corresponding to each of the multiple second target times.
[0101] Step c7: Determine the CPU resource matching value for each time based on the second area, the second quantity, and the memory resource demand.
[0102] For steps c5 to c7, please refer to the above steps c1 to c4, and no further elaboration will be provided here.
[0103] It should be noted that the objects of steps c1 to c4 are the remaining memory resources, and the objects of steps c5 to c7 are the remaining CPU resources.
[0104] Step c8: Determine the task scheduling coefficient for each time based on the memory resource matching value and the CPU resource matching value for each time.
[0105] After determining the memory resource matching value and the CPU resource matching value for each time, the task scheduling coefficient for each time can be determined based on the memory resource matching value and the CPU resource matching value.
[0106] As an example, the weights of the memory resource matching value and the CPU resource matching value can be set, and then the task scheduling coefficient can be determined based on the weights.
[0107] The method for determining the task scheduling time provided in this embodiment can accurately match the resource requirements of tasks for memory and the CPU with the actual resource supply of the cluster. By calculating the resource matching values of memory and the CPU respectively, it can more accurately reflect the resource satisfaction degree of tasks at different time points.
[0108] Meanwhile, by calculating the first area and the second area, and the quantities of memory and CPU resource demands that are greater than the remaining resources, this method can more comprehensively evaluate the resource utilization efficiency of tasks at different time points. This helps to optimize resource allocation, reduce resource waste, and improve the overall performance of the cluster. Moreover, by introducing the running time and multiple second target times, this method can dynamically evaluate the resource occupancy of tasks in the time dimension. This helps to identify possible resource bottlenecks during task execution, so as to perform resource scheduling and optimization in advance.
[0109] In a possible implementation, the above step c8 includes:
[0110] Step c81: Compare the memory resource matching values and CPU resource matching values at each time, and determine a first matching value and a second matching value from the memory resource matching values and the CPU resource matching values; wherein, the first matching value is greater than the second matching value.
[0111] When determining the memory resource matching values and the CPU resource matching values, for any one of the times, the first matching value and the second matching value can be determined in the following manner.
[0112] Take the larger value among the memory resource matching value and the CPU resource matching value as the first matching value, and take the smaller value among the memory resource matching value and the CPU resource matching value as the second matching value.
[0113] As an example, define k as the smaller value of x and y, and o as the larger value of x and y.
[0114] Step c82: Determine the task scheduling coefficient at each time according to the first matching value, the second matching value, and a preset coefficient weight.
[0115] The preset coefficient weight can be set independently by a person. Among them, h can be defined as the preset coefficient weight, where 0.5 < h < 1.
[0116] In a possible implementation manner, determining the task scheduling coefficient at each time according to the first matching value, the second matching value, and the preset coefficient weight includes:
[0117] z = h × k + (1 - h) × o; where z is the task scheduling coefficient, h is the preset coefficient weight, k is the first matching value, and o is the second matching value.
[0118] Please refer to Figure 3 , Figure 3 which is a schematic diagram for comparing task scheduling coefficients provided by an embodiment of the present invention.
[0119] Figure 3 In [reference], the horizontal axis is time, and the vertical axis is the task scheduling coefficient, that is, the task scheduling score. Find the maximum value from the task scheduling coefficients, and use the time corresponding to the maximum value as the task scheduling time. In Figure 4 , the time corresponding to the maximum value can be 14:00.
[0120] The method for determining the task scheduling time provided in this embodiment can comprehensively consider the requirements of tasks for two core resources by comparing the memory resource matching value and the central processing unit resource matching value. This helps ensure that the performance bottleneck will not occur due to the shortage of a certain resource during the task execution. Moreover, by determining the first matching value and the second matching value, this method provides clear priorities for resource allocation. The resource type represented by the first matching value (whether it is memory or central processing unit) will be regarded as the key resource for task execution, and thus will obtain a higher priority during resource scheduling.
[0121] Please refer to Figure 4 , Figure 4 which is a schematic diagram of the method for determining the task scheduling time provided in the embodiment of the present invention.
[0122] Step 1: Establish a storage database for the historical resource usage data of the cluster, specifically including information items: data collection time, cluster cpu usage, cluster memory usage, task name, cpu resources used by the task, and memory resources used by the task.
[0123] Step 2: Collect data from the big data cluster in real time according to the preset data collection frequency; among them, the data at the cluster dimension includes: cpu usage, memory usage, and data collection time; the data at the task dimension includes: the task name, cpu usage, memory usage, and data collection time of each task being executed.
[0124] Step 3: Obtain the input of the user task whitelist: If there is a recently deleted task, input the task name as the whitelist task; if it is to re-plan the task execution time for the tasks arranged in the cluster, input the task name as the whitelist task.
[0125] Step 4: Generate a daily dimension available resource double-axis curve; among them, the horizontal axis is the time axis showing 24 hours from 0 to 23 o'clock, the minimum unit is 5 minutes, and there are a total of 288 points. There are two vertical axes, with memory on the left and cpu on the right. The calculation method of the points on the curve is as follows (illustrated by the calculation method of the points on the memory curve, and the calculation method of the points on the cpu curve is the same): Take the cluster memory usage amounts αt, αt-1,... αt-14 at the same time point in the recent 15 days (αt is the memory usage amount yesterday, and it is recursively pushed forward one day in time in turn. If the newly built collection system points are not enough, take the number that can be obtained), and if there is a whitelist task, correspondingly take the memory usage amounts ɡt, ɡt-1,... ɡt-14 of the whitelist task (ɡt is the memory usage amount of this task yesterday, and it is recursively pushed forward one day in time in turn. If the newly built collection system points are not enough, take the number that can be obtained), and then determine the remaining amount of memory resources and the remaining amount of central processing unit resources according to the daily dimension available resource double-axis curve.
[0126] Step 5: According to the curve generated in Step 4, the resource demand quantity input by the user, and the estimated running time, generate a score for each point on the horizontal axis. The calculation method for the score of each point is as follows:
[0127] For memory resources, traverse each time point on the horizontal axis. Taking each time point as the starting point, use the estimated duration as the base, and calculate the area covered between the memory curve and the base.
[0128] Assume that the currently selected time axis point is t0 and the estimated duration is d minutes. Then the number of points n after t0 on the coordinate axis is: n = d ÷ 5
[0129] When the time interval exceeds 24 o'clock, recalculate from the 0 o'clock side (0 o'clock and 24 o'clock are regarded as the same point). These points are represented as t0, t1... tn. The values on the memory curve corresponding to t0... tn on the coordinate axis are v0... vn. The calculation method for the corresponding area s is (assuming the base value for 5 minutes is 1 in the formula):
[0130] The values on the memory curve corresponding to t0... tn-1 on the coordinate axis are v0... vn-1. Compare them with the expected memory value c input by the user in turn, and obtain the number p of points smaller than this value.
[0131] According to the s value and p value obtained above, generate the memory resource matching value x at the t0 point:
[0132] For CPU resources, obtain the CPU resource matching value y in the same way as for memory resources.
[0133] Generate a recommended score according to the memory resource matching value x and the CPU resource matching value y. Define k as the smaller value of x and y, and o as the larger value of x and y; define h as the weight 0.5 < h < 1. The calculation method for the recommended score z is:
[0134] z = h × k + (1 - h) × o.
[0135] Step 6: Traverse the recommended scores generated for each coordinate point on the horizontal axis, and output the time corresponding to the point with the maximum score as the recommended time.
[0136] In this embodiment, a device for determining the task scheduling time is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" may be a combination of software and / or hardware that can implement a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0137] This embodiment provides a device for determining task scheduling time, as follows Figure 5 shown, including: an acquisition module 501, configured to acquire a plurality of historical cluster resource usages, the current total cluster resources, the resource requirements and running time of the target task; a first determination module 502, configured to determine the remaining cluster resources at each time according to the plurality of historical cluster resource usages at each time and the current total task cluster resources; a second determination module 503, configured to determine the task scheduling coefficient at each time based on the resource requirements, running time, and the remaining cluster resources corresponding to each time of the target task; a third determination module 504, configured to use the time with the largest task scheduling coefficient among the task scheduling coefficients at each time as the task scheduling time.
[0138] The further functional descriptions of the above various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.
[0139] The device for determining task scheduling time in this embodiment is presented in the form of functional units. Here, the functional units refer to ASIC (Application Specific Integrated Circuit) circuits, processors and memories that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0140] This embodiment of the present invention also provides a computer device having the above Figure 5 shown device for determining task scheduling time.
[0141] Please refer to Figure 6 , Figure 6 is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention, as Figure 6 shown. The computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system). Figure 6 Taking one processor 10 as an example in
[0142] The processor 10 may be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device may be a complex programmable logic device, a field-programmable gate array, a generic array logic, or any combination thereof.
[0143] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.
[0144] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may further include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0145] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memories.
[0146] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks.
[0147] The embodiments of the present invention further provide a computer-readable storage medium. The method according to the embodiments of the present invention may be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading through a network and originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium may further include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0148] A part of the present invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can call or provide the methods and / or technical solutions according to the present invention through the operations of the computer. Those skilled in the art should understand that the forms in which computer program instructions exist in a computer-readable medium include but are not limited to source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.
[0149] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for determining task scheduling time, characterized in that: The method comprises: Obtain multiple historical cluster resource usage, current cluster resource totals, target task resource requirements, and running time; Determine the remaining amount of cluster resources at each time based on multiple historical cluster resource usages at each time and the total amount of cluster resources for the current task; Determine the task scheduling coefficient for each time based on the resource demand, running time, and remaining cluster resources corresponding to each time of the target task; The time with the largest task scheduling coefficient among the task scheduling coefficients of each time is used as the task scheduling time.
2. The method for determining task scheduling time according to claim 1, characterized in that: Determining the remaining amount of cluster resources at each time according to the multiple historical cluster resource usages at each time and the total amount of cluster resources of the current task includes: Detecting a task type of the target task; If the task type of the target task is a new task type, the remaining amount of cluster resources at each time is determined according to multiple historical cluster resource usages at each time and the total amount of cluster resources of the current task; wherein determining the remaining amount of cluster resources at each time includes: Determine a first resource usage value for the target time according to multiple historical cluster resource usages at the target time and a preset weight corresponding to each historical cluster resource usage; wherein the target time is any one of the times; The remaining amount of cluster resources at the target time is determined according to the difference between the total amount of cluster resources of the current task and the first resource usage value at the target time.
3. The method for determining task scheduling time according to claim 1, characterized in that: Determining the remaining amount of cluster resources at each time according to the multiple historical cluster resource usages at each time and the total amount of cluster resources of the current task includes: Detecting a task type of the target task; If the task type of the target task is a change task type, the remaining amount of cluster resources at each time is determined according to multiple target resource usages at each time, the preset weight corresponding to each target resource usage, and the total amount of cluster resources of the current task; wherein the change task type is used to indicate that the currently existing task needs to change the task scheduling time, and determining the remaining amount of cluster resources at each time includes: Obtain the task memory usage corresponding to each historical cluster resource usage at the target time; wherein the target time is any one of the times; Determine multiple target resource usages at the target time based on historical cluster resource usages and task memory usages corresponding to the historical cluster resource usages; Determining a second resource usage value for the target time according to the multiple target resource usages for the target time and a preset weight corresponding to each target resource usage; The remaining amount of cluster resources at the target time is determined according to the difference between the total amount of cluster resources of the current task and the second resource usage value at the target time.
4. The method for determining task scheduling time according to claim 1, characterized in that: The resource requirement of the target task includes the memory resource requirement and the CPU resource requirement, and correspondingly, the cluster resource surplus includes the memory resource surplus and the CPU resource surplus; wherein, the task scheduling coefficient at each time is determined based on the resource requirement of the target task, the running time, and the cluster resource surplus corresponding to each time, including: According to the running time, determining a plurality of second target times corresponding to the first target time; wherein the first target time is any one of the times; Determine a first area corresponding to the running time according to a memory resource demand corresponding to the first target time, a memory resource demand corresponding to each of the plurality of second target times, the first target time, and the second target time; According to the memory resource demand corresponding to the first target time and the memory resource demand corresponding to each of the plurality of second target times, determining that the memory resource demand in the first target time and the plurality of second target times is greater than a first amount of the memory resource demand; Determine a memory resource matching value at each time according to the first area, the first quantity, and the memory resource demand; determining a second area corresponding to the running time according to a remaining amount of CPU resources corresponding to the first target time, a remaining amount of CPU resources corresponding to each of the plurality of second target times, the first target time, and the second target time; According to the remaining amount of the CPU resource corresponding to the first target time and the remaining amount of the CPU resource corresponding to each of the plurality of second target times, determining a second amount by which the remaining amount of the CPU resource in the first target time and the plurality of second target times is greater than the required amount of the CPU resource; Determine a CPU resource matching value at each time according to the second area, the second number, and the memory resource requirement; The task scheduling coefficient at each time is determined according to the memory resource matching value and the central processing unit resource matching value at each time.
5. The method for determining task scheduling time according to claim 4, characterized in that: Determining the task scheduling coefficient at each time according to the memory resource matching value and the central processing unit resource matching value at each time includes: Compare the memory resource matching value and the CPU resource matching value at each time, and determine a first matching value and a second matching value from the memory resource matching value and the CPU resource matching value; wherein the first matching value is greater than the second matching value; The task scheduling coefficient for each time is determined according to the first matching value, the second matching value and a preset coefficient weight.
6. The method for determining task scheduling time according to claim 5, characterized in that: The determining the task scheduling coefficient for each time according to the first matching value, the second matching value and a preset coefficient weight includes: z=h×k+(1-h)×o; wherein z is the task scheduling coefficient, h is the preset coefficient weight, k is the first matching value, and o is the second matching value.
7. A device for determining task scheduling time, characterized in that: The device comprises: The acquisition module is used to obtain multiple historical cluster resource usage, the current total cluster resource, the resource demand and running time of the target task; A first determination module is used to determine the remaining amount of cluster resources at each time according to multiple historical cluster resource usages at each time and the total amount of cluster resources of the current task; A second determination module is used to determine the task scheduling coefficient at each time based on the resource demand, running time and the remaining amount of cluster resources corresponding to each time of the target task; The third determining module is used to take the time with the largest task scheduling coefficient among the task scheduling coefficients of each time as the task scheduling time.
8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for determining task scheduling time according to any one of claims 1 to 6 by executing the computer instructions.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method for determining task scheduling time according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to cause a computer to execute the method for determining task scheduling time according to any one of claims 1 to 6.
Citation Information
Cited By
Task load balancing scheduling method and device based on big data
CN120353558A