A large-scale model computing task adaptive scheduling method
Patent Information
- Application Number
- CN202611215608.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-12
- Publication Date
- 2026-09-11
AI Technical Summary
统一调度无法识别这种差异,使得计算任务被随机分配到不匹配的节点上
通过根据计算任务的预估计算量将动态资源队列划分为多个资源层级,并为每个层级对应预设的计算资源量区间,使不同消耗特征的任务与具备相应资源能力的节点区间建立关联。在待调度任务选取执行节点时,将当前计算任务的预估计算量与各资源层级的区间进行比对,当预估计算量落入某一区间时,直接将该区间对应的资源层级确定为目标资源层级。这改变了不加区分地将任务随机分配至任意可用节点的方式。由于节点已按可用计算资源量降序排列成队列,进一步从该目标资源层级中筛选出排序最靠前的候选节点承载任务,使得高消耗任务能够优先获得资源充裕的节点,低消耗任务则匹配至资源适中的节点,减少高消耗任务被分配到资源紧缺节点或低消耗任务占用过高配置节点的错配情形。节点资源利用率在任务执行初期即被纳入分配逻辑,任务与节点的匹配不再依赖单一维度的空闲核数或剩余内存容量,而是依据可量化的消耗区间进行分层供给。在目标计算节点执行当前计算任务过程中,按固定时间间隔采集中央处理器瞬时使用率和内存瞬时占用率,并据此计算每次采集时刻的综合资源使用率。将连续多次采集的综合资源使用率分别与高负载阈值和低负载阈值进行比较。当连续多次综合资源使用率均超出高负载阈值时,该节点已处于持续过载状态,将其在动态资源队列中的排序位置向后移动,降低任务调度器从该节点分配新任务的优先级,避免新的计算任务继续涌入已过载节点,抑制过载向集群内其他节点扩散。当连续多次综合资源使用率均低于低负载阈值时,该节点负载持续较轻,将其在队列中的排序位置向前移动,增加其被选取执行新任务的概率,使得闲置资源得到更充分的利用。节点排序位置的调整发生在任务执行期间,不再需要等待任务完成或失败事件触发,负载变化的响应时延被缩短。实时负载反馈与队列位置动态更新共同作用,使资源供给方向能够依据节点实际消耗状态持续修正,缓解任务执行中后期出现的部分节点严重过载而其余节点资源大量空闲的结构性失衡。
Smart Images

Figure CN122733484A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer task scheduling technology, specifically to an adaptive scheduling method for large-scale model computing tasks. Background Technology
[0002] Large-scale model computation tasks typically refer to jobs such as deep learning model training and scientific simulation computation that require a large number of CPU cycles and memory space. Existing scheduling methods mostly adopt fair and shared scheduling strategies, forming a single resource pool of computing nodes and allocating nodes according to the order of task submission or preset weights. This approach often results in coarse resource allocation when facing load fluctuations. The matching between computing tasks and nodes relies solely on the amount of currently available resources, lacking a quantitative estimate of the actual computational consumption of the tasks. When multiple high-consumption tasks are simultaneously assigned to a few high-performance nodes, it can easily lead to node overload while the remaining nodes are underutilized. A deficiency in existing technical solutions lies in the fact that the scheduling granularity fails to incorporate the heterogeneity of tasks. Different computing tasks have significantly different CPU and memory consumption due to differences in model structure and input size. Unified scheduling cannot identify these differences, resulting in computing tasks being randomly assigned to mismatched nodes. Some tasks exceed expected execution times due to insufficient resource allocation, while others consume excessive resources for extended periods, causing a decrease in the overall throughput of the cluster. Furthermore, existing scheduling structures exhibit a lag in responding to dynamic changes in node load. Node status is typically updated only after a task is completed or fails. It cannot adjust its position based on real-time load changes during task execution, leading to resource fragmentation that gradually accumulates as the task progresses.
[0003] Solving these problems requires a two-pronged approach. First, how to stratify computational tasks based on estimated consumption and establish corresponding node supply intervals for different tiers, creating a quantitative matching relationship between tasks and nodes, rather than a simple first-come, first-served basis. Second, during task execution, how to detect real-time changes in node load and dynamically adjust resource supply according to actual consumption by moving node sorting positions in the resource queue, thereby alleviating the contradiction between localized overload and resource idleness. Summary of the Invention
[0004] This invention provides an adaptive scheduling method for large-scale model computing tasks, which aims to achieve hierarchical matching between computing tasks and computing nodes, and dynamically adjust the sorting position of nodes in the resource queue based on real-time load during task execution, so as to alleviate the problems of local overload and resource idleness.
[0005] To achieve the above objectives, the present invention provides the following technical solution: The present invention provides an adaptive scheduling method for large-scale model computing tasks, which dynamically constructs resource queues and performs hierarchical resource matching based on the estimated computational load of tasks. During task execution, the node load is continuously monitored and the queue position is adjusted, thereby achieving efficient and elastic scheduling between computing tasks and computing nodes.
[0006] The method includes: responding to changes in the resource status of multiple computing nodes in a computing cluster, generating a dynamic resource queue by arranging the computing nodes in descending order of available computing resources; dividing the dynamic resource queue into multiple resource levels based on the estimated computational load of each computing task in the task queue to be scheduled, with each resource level corresponding to a preset range of computing resources; determining the target resource level to which the current computing task belongs based on the estimated computational load, and allocating a target computing node from the target resource level to execute the current computing task; monitoring the resource utilization rate of the target computing node during the execution of the current computing task, and adjusting the position of the target computing node in the dynamic resource queue based on the resource utilization rate. In this way, the dynamic resource queue reflects the available resource status of each computing node in the cluster in real time, and matches tasks with nodes of appropriate resource levels by pre-estimating the computing load, making resource allocation more reasonable and avoiding high-resource nodes being occupied by light tasks or low-resource nodes being unable to handle heavy tasks, thereby improving overall scheduling efficiency and resource utilization. At the same time, the ranking of nodes in the queue is adjusted based on the resource utilization rate during actual operation, so that the cluster scheduling can adapt to load changes and maintain system stability and response speed.
[0007] As a preferred embodiment of the present invention, the process of generating a dynamic resource queue by arranging computing nodes in descending order of available computing resources includes: collecting the current number of idle CPU cores and the current remaining memory capacity of each computing node; calculating the available computing resources of each computing node; and sorting all computing nodes in descending order of available computing resources to generate the dynamic resource queue. By simultaneously considering the idle status of both processor and memory, the actual available capacity of the nodes can be more accurately reflected, enabling the sorting results of the dynamic resource queue to effectively guide task allocation.
[0008] As another preferred technical solution of the present invention, the process of dividing the dynamic resource queue into multiple resource levels according to the estimated computational load of each computing task in the task queue to be scheduled includes: obtaining the historical execution time and input data volume of each computing task, and calculating the estimated computational load of each computing task accordingly; statistically analyzing the estimated computational load of all computing tasks to obtain the maximum estimated computational load and the minimum estimated computational load; and equally dividing the numerical range between the maximum estimated computational load and the minimum estimated computational load into a preset number of intervals, each interval corresponding to a resource level. This method of estimating computational load based on historical task data can obtain relatively accurate resource demand indicators without knowing the internal structure of the task in advance. Furthermore, mapping the dynamic resource queue to the estimated computational load intervals as levels facilitates the classification and matching of tasks of different weights into node groups with corresponding capabilities.
[0009] As a further preferred embodiment of the present invention, when determining the target resource level, the estimated computational load of the current computing task is compared with the computational resource range corresponding to each resource level. When the estimated computational load falls within a certain computational resource range, the resource level corresponding to that range is determined as the target resource level. When allocating target computing nodes from the target resource level, all candidate computing nodes within that level are selected from the dynamic resource queue, and the candidate computing node ranked highest is selected as the target computing node. Then, the current computing task is sent to the task execution queue of the target computing node. Thus, the node with the most abundant available resources in the current level is always given priority for task execution, which helps to shorten task waiting time and execution time.
[0010] As a preferred embodiment of the present invention, monitoring the resource utilization of a target computing node includes: collecting the instantaneous CPU utilization and instantaneous memory occupancy of the target computing node at fixed time intervals during task execution, and calculating the comprehensive resource utilization at each collection time. When adjusting the position of the target computing node in the dynamic resource queue based on the resource utilization, the comprehensive resource utilization collected each time is compared with preset high load thresholds and low load thresholds; when the comprehensive resource utilization is higher than the high load threshold for a consecutive preset number of times, the ranking position of the target computing node in the dynamic resource queue is moved backward by a preset number of positions; when the comprehensive resource utilization is lower than the low load threshold for a consecutive preset number of times, its ranking position is moved forward by a preset number of positions. Through this bidirectional position adjustment mechanism, the scheduling priority of high-load nodes can be reduced, reducing the pressure on subsequent task allocation, thereby smoothing the load, while enabling low-load nodes to obtain new tasks more quickly, achieving load balancing and full utilization of resources.
[0011] As a further preferred embodiment of the present invention, after adjusting the position of the target computing node in the dynamic resource queue according to the resource utilization rate, the process further includes a migration protection process for tasks that are difficult to complete within the deadline: obtaining the remaining unexecuted computation of the current computing task, and calculating the expected remaining execution time of the current computing task in combination with the current resource utilization rate of the target computing node; when the expected remaining execution time exceeds the preset deadline of the task, selecting the standby computing node with the largest available computing resources from the dynamic resource queue, and migrating the current computing task from the target computing node to the standby computing node for continued execution. When selecting the standby computing node, computing nodes currently executing computing tasks are excluded from the dynamic resource queue, and the node with the highest sorted position is selected from the remaining standby computing node set, which is the standby computing node with the largest available computing resources. During the task migration process, the execution of the current computing task is paused on the target computing node, the current execution progress and current intermediate state data are recorded, and transmitted to the standby computing node, and execution is resumed on the standby computing node according to the recorded progress and intermediate state data. This design enables the proactive transfer of tasks to nodes with more abundant resources to continue execution when an abnormal increase in node load may cause a task to exceed its deadline, ensuring that the task is completed before the deadline and improving scheduling reliability and service quality.
[0012] The technical effects and advantages provided by the present invention in the above technical solution are as follows: By dividing the dynamic resource queue into multiple resource levels based on the estimated computational load of the task, and assigning a preset computational resource range to each level, tasks with different consumption characteristics are associated with node ranges possessing corresponding resource capabilities. When a task to be scheduled selects an execution node, the estimated computational load of the current task is compared with the range of each resource level. When the estimated computational load falls into a certain range, the resource level corresponding to that range is directly determined as the target resource level. This changes the way tasks are randomly assigned to any available node indiscriminately. Since the nodes are already arranged in descending order of available computational resources, the highest-ranked candidate nodes are further selected from the target resource level to carry the task. This ensures that high-consumption tasks can preferentially obtain nodes with sufficient resources, while low-consumption tasks are matched with nodes with moderate resources, reducing mismatches such as high-consumption tasks being assigned to nodes with scarce resources or low-consumption tasks occupying nodes with excessively high configurations. Node resource utilization is incorporated into the allocation logic at the beginning of task execution. The matching of tasks and nodes no longer depends on a single dimension of the number of idle cores or remaining memory capacity, but is based on tiered supply according to quantifiable consumption ranges. During the execution of the current computing task on the target computing node, the instantaneous CPU utilization and memory occupancy are collected at fixed time intervals, and the overall resource utilization at each collection moment is calculated accordingly. The overall resource utilization collected multiple times is compared with high-load and low-load thresholds, respectively. When the overall resource utilization exceeds the high-load threshold multiple times consecutively, the node is in a state of continuous overload. Its ranking position in the dynamic resource queue is moved backward, reducing the priority of the task scheduler in allocating new tasks to this node, preventing new computing tasks from continuing to flood into the overloaded node, and suppressing the spread of overload to other nodes in the cluster. When the overall resource utilization is below the low-load threshold multiple times consecutively, the node's load is consistently light. Its ranking position in the queue is moved forward, increasing the probability of it being selected to execute new tasks, allowing for more efficient utilization of idle resources. The adjustment of the node's ranking position occurs during task execution, eliminating the need to wait for task completion or failure events, thus shortening the response latency for load changes. The combined effect of real-time load feedback and dynamic queue position updates enables the direction of resource supply to be continuously adjusted based on the actual consumption status of nodes, alleviating the structural imbalance that occurs in the later stages of task execution, where some nodes are severely overloaded while the resources of other nodes are largely idle. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0014] Figure 1This is a flowchart of an adaptive scheduling method for large-scale model computation tasks; Figure 2 This is a flowchart of dynamic resource hierarchy partitioning and computation task scheduling; Figure 3 This is a ranking diagram of the available computing resources of computing nodes under different weighting coefficients; Figure 4 This is a diagram showing the ranking of available computing resources for computing nodes and the matching of estimated computational load for tasks. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] See Figure 1 This invention provides an adaptive scheduling method for large-scale model computing tasks, comprising: responding to changes in the resource status of multiple computing nodes in a computing cluster, generating a dynamic resource queue by arranging the computing nodes in descending order of available computing resources; dividing the dynamic resource queue into multiple resource levels according to the estimated computational load of each computing task in the task queue to be scheduled, with each resource level corresponding to a preset range of computing resources; determining the target resource level to which the current computing task belongs based on the estimated computational load, and allocating a target computing node from the target resource level to execute the current computing task; monitoring the resource utilization rate of the target computing node during the execution of the current computing task, and adjusting the position of the target computing node in the dynamic resource queue according to the resource utilization rate.
[0017] Example 1: In practice, the current number of idle CPU cores and the current remaining memory capacity of each compute node in the computing cluster are collected. A monitoring agent is deployed in the computing cluster, and this agent reports a resource status snapshot of each compute node to the cluster management node at a fixed collection period. The resource status snapshot includes the current number of idle CPU cores and the current remaining memory capacity. The current number of idle CPU cores refers to the number of physical cores on the compute node that are not allocated to any processes or threads, and the current remaining memory capacity refers to the amount of unused physical memory on the compute node, expressed in megabytes.
[0018] After collecting the current number of idle CPU cores and the current remaining memory capacity, the available computing resources for each computing node are calculated. The available computing resources are obtained by a weighted sum of the current number of idle CPU cores and the current remaining memory capacity, expressed by the following formula: ; in, This represents the amount of available computing resources for the i-th computing node in the computing cluster. This represents the number of currently idle cores in the central processing unit of the i-th computing node. Let α represent the current remaining memory capacity of the i-th computing node, and β represent the CPU weight coefficient and memory weight coefficient. The values of the CPU weight coefficient α and memory weight coefficient β are determined based on the distribution of the number of tasks of each type in the task queue to be scheduled. Each time the resource queue is updated, the scheduling system calculates the task type distribution of all computing tasks in the task queue to be scheduled, separately counting the number of computationally intensive tasks and the number of memory-intensive tasks. When the proportion of computationally intensive tasks in the task queue is greater than the proportion of memory-intensive tasks, the CPU weight coefficient α is set to 0.7 and the memory weight coefficient β is set to 0.3, biasing the calculation of available computing resources towards the number of idle CPU cores, thus prioritizing the allocation of nodes with sufficient CPU resources to computationally intensive tasks. When the proportion of memory-intensive tasks is greater than the proportion of computationally intensive tasks, the CPU weight coefficient α is set to 0.3 and the memory weight coefficient β is set to 0.7, biasing the calculation of available computing resources towards the remaining memory capacity, thus prioritizing the allocation of nodes with sufficient memory resources to memory-intensive tasks. When the number of tasks of the two types is equal or neither constitutes a significant majority, the CPU weight coefficient α and the memory weight coefficient β are both set to 0.5, ensuring that the number of idle CPU cores and the remaining memory capacity contribute equally to the calculation of available computing resources. The aforementioned ratios of 0.7 and 0.3 are based on the following principle: when the number of tasks of the two types is unevenly distributed, the resource dimension corresponding to the dominant task type receives approximately twice the weight of the other resource dimension, thereby reflecting the bias of task type towards resource demand in the ranking results. The aforementioned preset ratio can be set to 60%. The determination of task type is based on the average CPU utilization and average memory usage in the historical execution records of the computing task. When the average CPU utilization is higher than the first threshold, it is determined to be a computing-intensive task; when the average memory usage is higher than the second threshold, it is determined to be a memory-intensive task.
[0019] After obtaining the available computing resources for each computing node in the computing cluster, all computing nodes are sorted in descending order of available computing resources to generate a dynamic resource queue. The dynamic resource queue is an ordered list, where each element records the identifier of a computing node and its corresponding available computing resources. Nodes ranked higher have more available resources. When multiple computing nodes have the same amount of available computing resources, the number of currently idle CPU cores is compared, with the node having a higher number of idle CPU cores ranked higher. If the number of idle CPU cores is also the same, the remaining memory capacity is compared, with the node having a higher remaining memory capacity ranked higher. The dynamic resource queue is regenerated every time a resource status change triggers an update to ensure that scheduling decisions are based on the latest cluster resource distribution.
[0020] See Figure 3 The figure shows a curve illustrating the change in available computing resources of computing nodes in the computing cluster based on Example 1, as the node's sorting number changes. The horizontal axis represents the node's sorting number, ranging from 1 to 500; a smaller number indicates a larger amount of available computing resources. The vertical axis represents the amount of available computing resources of the node; the unit is not explicitly specified, and the value ranges from approximately 5 to 45.
[0021] The figure contains three curves, corresponding to the calculation results of available computing resources under different weight coefficient configurations in Example 1: the solid line represents the computationally intensive weight (α=0.7, β=0.3), the dashed line represents the balanced weight (α=0.5, β=0.5), and the dotted line represents the memory-intensive weight (α=0.3, β=0.7). The curves generally show a monotonically decreasing trend, indicating that the available computing resources decrease as the sorting number of the computing node increases.
[0022] Specifically, the available computing resources of the memory-intensive weight curve are generally higher than those of the balanced weight curve and the computationally intensive weight curve, especially in the lower sort numbers (approximately 1 to 300), where the maximum available computing resources are approximately 45 and the minimum is approximately 30. The balanced weight curve is in the middle position, with a maximum of approximately 40 and a minimum of approximately 25; the computationally intensive weight curve is the lowest, with a maximum of approximately 35 and a minimum of approximately 20. The three curves tend to converge towards the later sort numbers (approximately 450 to 500), indicating that on computing nodes with lower available computing resources, the difference in resource quantity under different weight configurations is relatively small.
[0023] Example 2: In practice, the historical execution time and input data volume of each computation task in the queue of tasks to be scheduled are obtained. The scheduling system maintains a task execution history database, which stores information such as the identifier of completed computation tasks, execution start and end times, execution duration, input data volume, and task type. For each computation task appearing in the queue of tasks to be scheduled, the scheduling system queries the task execution history database for the execution duration recorded during the most recent successful execution of the computation task, which is used as the historical execution duration. At the same time, it queries the data volume of input data processed in that execution, which is used as the input data volume. When a computation task has no execution record in the task execution history database, the historical execution duration is set to the default duration value corresponding to the computation task type, and the input data volume is set to the actual data volume of the data to be processed. The default duration value is calculated based on the average execution time of the task type. Specifically, the default duration for data preprocessing type computation tasks is set to 60 seconds, the default duration for model training type computation tasks is set to 3600 seconds, and the default duration for model inference type computation tasks is set to 30 seconds.
[0024] Optionally, the historical execution time is in seconds, and the input data volume is in megabytes.
[0025] The estimated computational load for each computation task is calculated based on historical execution time and input data volume. The formula for the estimated computational load is as follows: ; Among them, E j d represents the estimated computational cost of the j-th computational task in the queue of tasks to be scheduled; j S represents the historical execution time of the j-th computational task; j This represents the amount of input data for the j-th computation task; This represents the complexity coefficient corresponding to task type t of the j-th computational task, where t is the task type identifier. Estimated computational load E j The unit of measurement is seconds per megabyte (s·MB), which physically represents the computational workload required to process a unit of data volume. It is used to compare the relative computational consumption between tasks. Task types are divided into data preprocessing, model training, and model inference. The complexity coefficient corresponding to the data preprocessing type... The value is 0.8; the complexity coefficient corresponding to the model training type. The value is 1.2; the complexity coefficient corresponding to the model inference type. The value is 1.0. Complexity coefficient. The value of the complexity coefficient is determined based on the differences in computational density shown in historical execution records for different task types. During initialization, the scheduling system extracts execution records of completed computational tasks from the task execution history database, groups them by task type, and calculates the average computational density for each task type. Computational density is defined as the product of the average CPU utilization and average memory usage during task execution. Statistical results show that the average computational density of model training tasks is approximately 1.2 times that of model inference tasks, and the average computational density of data preprocessing tasks is approximately 0.8 times that of model inference tasks. Based on these statistical proportions, the complexity coefficient for model inference tasks is set to a baseline value of 1.0, the complexity coefficient for model training tasks is set to 1.2, and the complexity coefficient for data preprocessing tasks is set to 0.8. For computational tasks that cannot be classified into the above three task types, they are considered general computational types, with computational densities similar to model inference tasks; therefore, their complexity coefficient is set to 1.0. This involves converting historical execution time and input data volume into a unified conversion factor, allowing for comparison of estimated computational loads for different task types on the same scale. During operation, the scheduling system periodically recalculates the average computational density of various task types. When the average computational density of a certain task type changes by more than 10% relative to the initial statistical value, the complexity coefficient for that task type is adjusted accordingly, and the adjusted value is rounded to one decimal place.
[0026] After calculating the estimated computational cost for each computational task, the estimated computational costs of all computational tasks in the queue to be scheduled are tallied to obtain the maximum and minimum estimated computational costs. Specifically, all computational tasks in the queue to be scheduled are traversed, and the estimated computational cost of the currently traversed task is compared with the recorded maximum estimated computational cost. If the current estimated computational cost is larger, the maximum estimated computational cost is updated; simultaneously, it is compared with the recorded minimum estimated computational cost. If the current estimated computational cost is smaller, the minimum estimated computational cost is updated. After the traversal is complete, the maximum estimated computational cost is obtained. and minimum estimated computational cost If there is only one computation task in the queue of tasks to be scheduled, the maximum estimated computational load is... and minimum estimated computational cost If the same value is taken, the subsequent interval division result will be a single interval containing that value, which corresponds to the resource level to which the calculation task belongs.
[0027] The range between the maximum and minimum estimated computational load is divided into a predetermined number of intervals, each corresponding to a resource level. The predetermined number is denoted as K, which is determined by the total number of compute nodes N in the compute cluster. K is calculated by taking the square root of N and rounding up, with a minimum value of 2. For example, when the total number of compute nodes N=16, K is 4. After determining K, the width of the computation interval is... : .when and When they are equal, At this point, all computing tasks are allocated to the same resource level. At that time, the range of computational resources corresponding to the p-th resource level is: Where p is an integer from 1 to K-1; when p=K, the range of computational resources corresponding to the Kth resource level is... Resource levels with larger boundary values in the resource quantity range correspond to higher resource level numbers and are used to accommodate computationally intensive tasks. Each resource level corresponds to a resource allocation range in the dynamic resource queue, and the specific correspondence is established when the dynamic resource queue is partitioned.
[0028] Example 3: In specific implementation, please refer to Figure 2 The dynamic resource queue is divided into multiple resource levels. The total number of resource levels is the same as the number of computational resource intervals obtained by dividing the queue based on the estimated computational load of each computational task. Each resource level corresponds one-to-one with a computational resource interval. The division method is as follows: traverse each computational node in the dynamic resource queue and obtain the available computational resources for each node; for each computational node, compare its available computational resources with all computational resource intervals. When the available computational resources of a computational node fall into a certain computational resource interval, the computational node is assigned to the resource level corresponding to that interval. The lower and upper bounds of the computational resource intervals are determined by the distribution of the estimated computational load during the division process. During the comparison process, if the available computing resources are greater than or equal to the lower bound of the computing resource interval and less than the upper bound of the computing resource interval, then the available computing resources are determined to fall within that interval. For the last computing resource interval, its upper bound is the maximum estimated computing power. During the comparison, if the available computing resources are greater than or equal to the lower bound of this interval and less than or equal to the maximum estimated computing power, then it is determined to fall within it. After traversal, each resource level contains a group of computing nodes, and the relative order of computing nodes within the same resource level in the dynamic resource queue is consistent with the order before partitioning.
[0029] The estimated computational load of the current computing task is compared with the computational resource range corresponding to each resource level. The estimated computational load of the current computing task is obtained, and all the divided resource levels are traversed. For each resource level, the corresponding computational resource range is obtained, and the estimated computational load of the current computing task is compared with the lower and upper bounds of the computational resource range. If the estimated computational load of the current computing task is greater than or equal to the lower bound of the computational resource range and less than the upper bound, the resource level corresponding to that range is determined as the target resource level. If the estimated computational load of the current computing task is exactly equal to the maximum estimated computational load, and the upper bound of the last computational resource range is the maximum estimated computational load, then if the estimated computational load is greater than or equal to the lower bound of the last range and less than or equal to the maximum estimated computational load, the last resource level is determined as the target resource level. In the above comparison process, the computational resource ranges are left-closed and right-open intervals, and the last interval is a closed interval, ensuring that all possible estimated computational loads can correspond to a single resource level.
[0030] After determining the target resource level, all candidate compute nodes within the target resource level are selected from the dynamic resource queue. Specifically, a list of compute nodes contained in the target resource level is obtained, and each compute node in this list is a candidate compute node. The order of the candidate compute nodes within the target resource level is the same as their order in the dynamic resource queue.
[0031] The target compute node is selected from the candidate compute nodes that ranks highest in the dynamic resource queue. Since the dynamic resource queue is an ordered list arranged from largest to smallest available compute resources, and candidate compute nodes within the same resource level maintain this order, the candidate compute node ranking highest in the target resource level is the compute node with the largest available compute resources in that level. If there are no candidate compute nodes in the target resource level, a compute node is borrowed from a higher-ranked adjacent resource level that contains compute nodes; specifically, the target compute node is selected from the highest-ranked compute node in the dynamic resource queue within that adjacent resource level. If there are no compute nodes in all higher resource levels, a compute node is borrowed from a lower-ranked adjacent resource level that contains compute nodes.
[0032] The current computation task is assigned to the task execution queue of the target compute node. The scheduling system constructs a task assignment message, which includes the task identifier of the current computation task, the executable program path, the storage address of the input data, the estimated computational load, and the task priority information. The scheduling system sends the task assignment message to the task manager running on the target compute node via remote procedure call or message queue. After receiving the task assignment message, the task manager on the target compute node appends the task descriptors to the tail of the task execution queue located in the local memory of the target compute node in the order of receipt, waiting for the task execution thread on the target compute node to retrieve the task from the task execution queue and start execution.
[0033] See Figure 4 In the figure, the horizontal axis represents the sorting number of the computing nodes, and the vertical axis represents the available computing resources of the corresponding computing node. The units are not explicitly stated, but based on the embodiment, they should be weighted resource index values. The blue scatter dots represent the current available computing resources of each computing node in the computing cluster. The scatter dots show a decreasing trend from left to right, reflecting the characteristic of computing nodes in the dynamic resource queue being sorted from largest to smallest available resource quantity. The maximum resource quantity is approximately 42, and the minimum is approximately 5, indicating a significant difference in the distribution of computing node resources within the computing cluster.
[0034] The vertical red dashed line in the diagram represents the estimated computational load of the current task. This estimated computational load value is near the available resources of the top-ranked computing nodes, specifically around 30. As described in Example 3, the dynamic resource queue is divided into multiple resource levels, each corresponding to a preset range of computing resources. The multiple gray vertical lines in the background of the diagram delineate the boundaries of the resource levels, with the range width equally divided by the maximum and minimum estimated computational load values, ensuring that the available resources of all computing nodes are assigned to their corresponding resource levels.
[0035] The purple pentagram marks the location of the target computing node, which is approximately 170 in the computing node sorting sequence. Its available computing resources are approximately 26, falling into a resource level that matches the estimated computing load of the current computing task. This target computing node is the highest priority node among the candidate computing nodes in the dynamic resource queue, conforming to the scheduling rule in Example 3 of "selecting the computing node with the highest sorting position in the target resource level as the target computing node".
[0036] Example 4: In practice, during the execution of the current computing task on the target computing node, the instantaneous CPU utilization and instantaneous memory occupancy of the target computing node are collected at fixed time intervals. The monitoring module of the scheduling system sends collection instructions to the monitoring agent of the target computing node. The monitoring agent reads the resource usage status information provided by the operating system kernel of the target computing node at a fixed collection cycle. The instantaneous CPU utilization refers to the percentage of CPU core time that is in a non-idle state on the target computing node at the current collection time, expressed as a percentage between 0 and 100. The instantaneous memory occupancy refers to the percentage of physical memory capacity used on the target computing node at the current collection time, also expressed as a percentage between 0 and 100. The fixed time interval is set to 10 seconds. This value is based on the fact that a 10-second time interval can ensure sufficient sensitivity in tracking changes in resource usage status while avoiding additional overhead to the normal operation of the target computing node due to excessively high collection frequency.
[0037] Optionally, when the target computing node is running a Linux operating system, the instantaneous CPU utilization is calculated by reading the user-mode time, kernel-mode time, and idle time from the / proc / stat file, and the instantaneous memory utilization is calculated by reading the MemTotal and MemAvailable fields from the / proc / meminfo file.
[0038] Based on the instantaneous CPU utilization and memory occupancy, the overall resource utilization of the target computing node at each data acquisition moment is calculated. The formula for calculating the overall resource utilization is as follows: ; in, This is the resource utilization rate weighting coefficient, a real number ranging from 0 to 1. The value of is determined by . Indicates the time of data collection The overall resource utilization rate of the target computing node. For the first The timestamp of each data collection moment; Indicates the time of data collection The instantaneous utilization rate of the central processing unit of the target computing node, expressed as a percentage. Indicates the time of data collection The instantaneous memory occupancy of the target compute node, expressed as a percentage; resource utilization weighting coefficient. The value is dynamically adjusted based on the task type of the current computing task running on the target computing node: when the task type of the current computing task is a compute-intensive task, The value is 0.7; when the current computation task is a memory-intensive task, The value is 0.3; when the current computation task cannot be clearly classified as either compute-intensive or memory-intensive, The value is set to 0.5. The resource utilization weight coefficient ω is determined based on the degree of dependence of different task types on CPU and memory resources. During initialization, the scheduling system extracts the execution records of completed computational tasks from the task execution history database, groups them by task type, and calculates the ratio of average CPU utilization to average memory occupancy for all tasks within each group during execution. Statistical results show that the ratio of average CPU utilization to average memory occupancy for computationally intensive tasks is concentrated between 2.0 and 3.0, indicating that their dependence on CPU resources is more than twice that of memory resources. Therefore, the CPU weight coefficient for computationally intensive tasks is set to 0.7, so that CPU utilization contributes about 70% of the weight in the overall resource utilization. The ratio of average CPU utilization to average memory occupancy for memory-intensive tasks is concentrated between 0.3 and 0.5, indicating that their dependence on memory resources is more than twice that of CPU resources. Therefore, the CPU weight coefficient for memory-intensive tasks is set to 0.3, so that memory occupancy contributes about 70% of the weight in the overall resource utilization. When a computational task cannot be clearly categorized as either compute-intensive or memory-intensive, a ratio of average CPU utilization to average memory usage between 0.5 and 2.0 indicates that the utilization levels of the two resource types are roughly equivalent. Therefore, the CPU weighting coefficient is set to 0.5, ensuring that CPU utilization and memory usage contribute equally to the overall resource utilization. The criteria for tasks that cannot be clearly categorized are: average CPU utilization below 60% and average memory usage below 60%, or average CPU utilization above 60% and average memory usage above 60%.
[0039] After calculating the overall resource utilization rate at each data collection point, the overall resource utilization rate is compared with preset high-load and low-load thresholds. The high-load threshold is set at 85%, and the low-load threshold is set at 30%. These values (85% for high load and 30% for low load) are determined based on statistical analysis of the cluster node load distribution during the scheduling system's operation. During the initial deployment phase, the scheduling system collects overall resource utilization rate data for all computing nodes in the computing cluster under normal operating conditions. The collection period is seven consecutive days, with a collection interval of 10 seconds. All collected overall resource utilization rate data are sorted, and the 95th percentile is used as the initial reference value for the high-load threshold, and the 5th percentile as the initial reference value for the low-load threshold. Statistical results show that under typical load conditions, the 95th percentile of the overall resource utilization rate is approximately 85%, and the 5th percentile is approximately 30%. The high load threshold of 85% means that when the overall resource utilization exceeds 85%, the remaining available resources of the node are less than 15%. Continuing to allocate new tasks will significantly increase the node's response latency; therefore, the node's priority in the dynamic resource queue needs to be reduced. The low load threshold of 30% means that when the overall resource utilization is below 30%, more than 70% of the node's resources are idle; therefore, the node's priority in the dynamic resource queue needs to be increased to increase task allocation opportunities. During operation, the scheduling system recalculates the percentile distribution of overall resource utilization monthly. When the change in the 95th percentile or 5th percentile from the initial value exceeds five percentage points, the high load threshold or low load threshold is adjusted accordingly.
[0040] The scheduling system maintains a high-load counter and a low-load counter for the target computing node. Both counters are initialized to zero when the target computing node is assigned a current computing task.
[0041] After each collection and calculation of the overall resource utilization rate, the overall resource utilization rate is compared with the high load threshold and the low load threshold respectively: when the overall resource utilization rate is higher than the high load threshold, the high load counter is incremented by 1, and the low load counter is set to zero; when the overall resource utilization rate is lower than the low load threshold, the low load counter is incremented by 1, and the high load counter is set to zero; when the overall resource utilization rate is between the low load threshold and the high load threshold, both the high load counter and the low load counter are set to zero.
[0042] When the high-load counter reaches a preset number of occurrences, the target computing node is determined to be in a state of continuous high load. The preset number of occurrences is set to 3, based on the correspondence between the 10-second collection period and the duration of load fluctuations. During the initial deployment phase, the scheduling system collects time-series data on the comprehensive resource utilization of all computing nodes in the computing cluster under normal operating conditions. The collection interval is 10 seconds, and the collection period is seven consecutive days. Fluctuation analysis is performed on the collected comprehensive resource utilization time series, and the duration distribution of complete fluctuation events—from below the high-load threshold to above the high-load threshold and then back below the high-load threshold—is statistically analyzed. Statistical results show that the duration of brief load spikes caused by instantaneous task concurrency is usually within 10 seconds, i.e., no more than one collection interval; the duration of continuous load increases caused by continuous task execution is usually more than 30 seconds, i.e., spanning at least three collection intervals. Based on the above statistical patterns, the preset number of occurrences is set to 3, corresponding to a duration of 30 seconds. When the overall resource utilization rate collected for three consecutive times is higher than the high load threshold, it indicates that the load increase of the node is continuous rather than a momentary fluctuation. Using this as the trigger condition for queue adjustment can effectively filter out false triggers caused by short-term load spikes. When the overall resource utilization rate collected for three consecutive times is lower than the low load threshold, it also indicates that the low load state of the node is continuous rather than a momentary fluctuation. Using this as the trigger condition for queue advancement can avoid frequent adjustments to the node position due to short-term resource releases. In this case, the target computing node's sorting position in the dynamic resource queue is moved forward by a preset number of positions. The preset bit depth is set to 20% of the total number of compute nodes in the dynamic resource queue, rounded up. If the result of 20% is less than 1, it is rounded to 1. The 20% ratio is chosen based on the following: when a target compute node needs to move backward due to sustained high load, moving 20% of the queue length is sufficient to allow it to exit the current resource level and enter a lower resource level, significantly reducing its probability of being selected again; conversely, when a target compute node needs to move forward due to sustained low load, moving 20% of the queue length is sufficient to allow it to enter a higher resource level, increasing its probability of being selected. Too small a bit depth is insufficient to change the node's resource level, while too large a bit depth may cause nodes to frequently jump between levels, increasing the risk of scheduling oscillations. The 20% ratio strikes a balance between adjustment effectiveness and stability. For example, when a computing cluster contains 100 computing nodes, the default bit depth is 20 bits. A node with consistently high load will move 20 positions backward, dropping from the front of the queue to the middle, significantly reducing its probability of being selected for a new task. Conversely, a node with consistently low load will move 20 positions forward, moving from the back of the queue to the middle, increasing its probability of being selected for a new task. For example, when a dynamic resource queue contains 10 computing nodes, the default bit depth is 2 bits.Moving backwards refers to moving the target computing node towards the end of the dynamic resource queue, thus reducing its sorting priority.
[0043] When the low-load counter reaches a preset number of times, the target compute node is determined to be in a sustained low-load state. The preset number of times is also set to 3. In this case, the target compute node's sorting position in the dynamic resource queue is moved forward by a preset number of positions. Moving forward means moving the target compute node towards the head of the dynamic resource queue, increasing its sorting priority. The preset number of positions for moving forward is the same as the preset number of positions for moving backward.
[0044] When performing a queue position adjustment operation, if the number of bits moved backward exceeds the tail range of the dynamic resource queue, the target compute node is placed at the tail of the dynamic resource queue; if the number of bits moved forward exceeds the head range of the dynamic resource queue, the target compute node is placed at the head of the dynamic resource queue. After the queue position adjustment is completed, both the high load counter and the low load counter are reset to zero and counting starts again.
[0045] Example 5: In practice, after adjusting the target computing node's position in the dynamic resource queue based on resource utilization, the remaining unexecuted computational load of the current computing task is obtained. The task manager on the target computing node, after pausing the execution of the current computing task, reads the configuration information of the current computing task to obtain the total input data volume and the amount of data already processed. The remaining unexecuted computational load is obtained by subtracting the amount of data already processed from the total input data volume. When the amount of data already processed cannot be directly obtained from the task manager, the task manager reads the progress log file of the current computing task during processing. The progress log file records the offset value of the data volume completed for each batch processing step, line by line. The offset value of the processed data volume recorded in the last line of the progress log file is read as the amount of data already processed. The remaining unexecuted computational load is represented in bytes.
[0046] Based on the remaining unexecuted computations and the current resource utilization of the target computing node, calculate the expected remaining execution time of the current computing task. The formula for calculating the expected remaining execution time is as follows: ; in, This indicates the expected remaining execution time of the current computation task, in seconds. This indicates the remaining unexecuted computations for the current task, in bytes. The execution rate coefficient represents the number of CPU cycles required to process a unit of data, measured in cycles per byte. The value is obtained by querying the task execution history database for the average execution rate coefficient of the most recent successful execution of the current computation task. Execution rate coefficient The acquisition method is as follows: The scheduling system maintains an execution record for each computing task in the task execution history database. This record contains the total CPU cycle consumption and total data processed for each successful execution of the task. The total CPU cycle consumption is obtained by reading the user-mode cycle count recorded in the CPU performance counter on the target computing node during task execution. This counter records the total number of cycles consumed by the task process in executing instructions on the CPU. The total data processed is obtained by reading the number of bytes read and written by the input / output subsystem of the target computing node during task execution. After each task execution, the scheduling system divides the total CPU cycle consumption by the total data processed to obtain the average execution rate coefficient for that execution. When there is no execution record for the current computing task in the task execution history database, the execution rate coefficient ρ is selected from a preset default value based on the task type of the current computing task. This indicates the CPU clock speed of the target computing node, measured in Hertz, and is read from the hardware configuration information of the target computing node. This represents the load rate of the target computing node at the current moment, and is taken as the arithmetic mean of the comprehensive resource utilization rates at the three most recent consecutive data collection times. The calculation method for the comprehensive resource utilization rate is the same as that used when adjusting the position of the dynamic resource queue. If If the current node's load is close to saturation, then directly... Setting it to infinity triggers a task migration operation, preventing subsequent duration comparisons and thus avoiding division by zero errors.
[0047] The default value is determined as follows: if the current computation task is of the data preprocessing type, the execution rate coefficient is... The value is 2.5 multiplied by 10 3 If the current computation task is of the model training type, the execution rate coefficient is... The value is 8.0 multiplied by 10 4 If the current computation task is of the model inference type, the execution rate coefficient is... The value is 1.2 multiplied by 10 4 The above values are based on the fact that data preprocessing type computational tasks require a relatively low amount of computation to process a unit of data, approximately 2500 CPU cycles. Take 2.5 multiplied by 10 3 Model training computational tasks involve matrix operations and gradient calculations, requiring the highest computational load per unit byte of data, approximately 80,000 CPU cycles. Take 8.0 multiplied by 104 Model inference type computational tasks only involve forward computation, and the computational load required to process a unit of data is moderate, approximately 12,000 CPU cycles. Take 1.2 multiplied by 10 4 .
[0048] After calculating the expected remaining execution time Next, the expected remaining execution time is compared with the preset deadline of the current computation task. The preset deadline for the current computation task is specified by the user when submitting the task and is stored in the task's configuration information. The preset deadline is expressed as a relative duration from the time the task started execution, in seconds. Subtracting the actual start time of the current computation task from the current time yields the executed time, and subtracting the executed time from the preset deadline yields the remaining deadline. When the expected remaining execution time... If the remaining deadline is exceeded, a computation task migration operation will be triggered.
[0049] After triggering the computation task migration operation, the standby compute node with the largest available compute resources is selected from the dynamic resource queue. The selection process is as follows: Compute nodes currently executing computation tasks are excluded from the dynamic resource queue, resulting in a set of standby compute nodes. Specifically, the scheduling system maintains a list of active compute nodes, which records the identifiers of all compute nodes currently executing computation tasks. Each compute node in the dynamic resource queue is traversed, and for each traversed compute node, its identifier is checked to see if it exists in the list of active compute nodes. If the identifier of a compute node does not exist in the list of active compute nodes, the compute node is added to the standby compute node set. After the traversal is complete, the standby compute node with the highest sorted position in the dynamic resource queue is selected from the standby compute node set as the standby compute node with the largest available compute resources. Since the dynamic resource queue is sorted in descending order of available compute resources, the standby compute node with the highest sorted position is the standby compute node with the largest available compute resources in the standby compute node set. When the standby compute node set is empty, the standby compute node selection process is re-executed after a fixed time interval, which is set to 5 seconds.
[0050] After selecting the standby computing node with the largest available computing resources, the current computing task is migrated from the target computing node to the standby computing node for continued execution. The migration operation is as follows: The execution of the current computing task is paused on the target computing node, and the current execution progress and intermediate state data of the current computing task are recorded. The task manager on the target computing node sends a pause signal to the execution thread of the current computing task. Upon receiving the pause signal, the execution thread completes the processing of the currently being processed batch data and stops execution after the batch data processing is complete. After the execution thread stops, the task manager records the current execution progress of the current computing task, which includes the offset of the amount of data processed and the number of the completed iteration rounds. The task manager also records the current intermediate state data of the current computing task, which includes the latest values of the model parameter matrix, the optimizer state vector, and the data loader's read position pointer. The model parameter matrix is stored in the form of a multi-dimensional array, and the optimizer state vector includes internal optimizer variables such as momentum terms and gradient squared accumulation terms. For model training type computation tasks, the task manager serializes the model parameter matrix and optimizer state vector into binary files, which are then stored in a specified directory on the target compute node's local file system.
[0051] After recording the current execution progress and intermediate state data, the data is transmitted to the standby compute node. The transmission method is as follows: the task manager on the target compute node establishes a secure transmission channel with the task manager on the standby compute node. This secure transmission channel is built based on a transport layer security protocol. The task manager on the target compute node packages the current execution progress and intermediate state data into a migration data packet. The migration data packet format includes a packet header and a data payload. The packet header contains the task identifier of the current compute task, the data payload length, and a checksum. The data payload contains the serialized current execution progress and intermediate state data. The task manager on the target compute node sends the migration data packet to the task manager on the standby compute node through the secure transmission channel. Upon receiving the migration data packet, the task manager on the standby compute node performs a checksum verification on the data payload. If the checksum verification is successful, the data payload is stored in the local file system of the standby compute node.
[0052] After data transmission is complete, the current computation task resumes execution on the standby compute node based on the current execution progress and intermediate state data. The task manager on the standby compute node deserializes the current execution progress and intermediate state data from the data payload. The deserialization process includes: reading the binary data of the model parameter matrix and restoring it as a multidimensional array object; reading the optimizer state vector and restoring it as a dictionary structure object; and reading the offset of the amount of data processed and the number of the completed iteration round. After deserialization, the task manager on the standby compute node initializes a task execution thread, passes the restored model parameter matrix and optimizer state vector to the task execution thread, sets the data loader's read position pointer to the position corresponding to the offset of the amount of data processed, and sets the current iteration round to the next round after the number of the completed iteration round. After the task execution thread is started, it continues loading data and executing computation tasks from the restored execution progress. After the standby compute node successfully starts execution, the scheduling system adds the standby compute node to the list of active compute nodes and removes the target compute node from the list of active compute nodes. The scheduling system simultaneously updates the task status record of the current computing task, changing the task status to "migrated and in execution", and records the identifier of the newly executing computing node.
[0053] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. An adaptive scheduling method for large-scale model computation tasks, characterized in that, include: In response to changes in the resource status of multiple computing nodes in the computing cluster, a dynamic resource queue is generated by arranging the computing nodes in descending order of available computing resources. Based on the estimated computational load of each computational task in the task queue to be scheduled, the dynamic resource queue is divided into multiple resource levels, and each resource level corresponds to a preset range of computational resources. Based on the estimated computational load, determine the target resource level to which the current computing task belongs, and allocate target computing nodes from the target resource level to execute the current computing task; Monitor the resource utilization rate of the target computing node during the execution of the current computing task, and adjust the position of the target computing node in the dynamic resource queue according to the resource utilization rate.
2. The adaptive scheduling method for large-scale model computation tasks according to claim 1, characterized in that, The process of generating a dynamic resource queue by arranging computing nodes in descending order of available computing resources includes: Collect the current number of idle CPU cores and the current remaining memory capacity of each computing node in the computing cluster; Calculate the available computing resources for each computing node based on the current number of idle cores in the central processing unit and the current remaining memory capacity; The computing nodes in the computing cluster are sorted in descending order of available computing resources to generate the dynamic resource queue.
3. The adaptive scheduling method for large-scale model computation tasks according to claim 1, characterized in that, The dynamic resource queue is divided into multiple resource levels based on the estimated computational load of each computational task in the task queue to be scheduled, including: Obtain the historical execution time and input data volume of each computing task in the queue of tasks to be scheduled, and calculate the estimated computational load of each computing task based on the historical execution time and the input data volume; The estimated computational load of all computational tasks in the queue of tasks to be scheduled is calculated to obtain the maximum and minimum estimated computational loads. The numerical range between the maximum estimated computational load and the minimum estimated computational load is divided into a preset number of intervals, and each interval corresponds to a resource level.
4. The adaptive scheduling method for large-scale model computation tasks according to claim 1, characterized in that, The step of determining the target resource level to which the current computing task belongs based on the estimated computational load includes: The estimated computational load of the current computational task is compared with the computational resource load range corresponding to each resource level; When the estimated computational load falls within the computational resource quantity range, the resource level corresponding to the computational resource quantity range is determined as the target resource level.
5. The adaptive scheduling method for large-scale model computation tasks according to claim 1, characterized in that, The step of allocating a target computing node from the target resource level to execute the current computing task includes: Select all candidate computing nodes located within the target resource level from the dynamic resource queue; The candidate computing node that ranks first in the dynamic resource queue among the candidate computing nodes is selected as the target computing node. The current computing task is sent to the task execution queue of the target computing node.
6. The adaptive scheduling method for large-scale model computation tasks according to claim 1, characterized in that, The monitoring of resource utilization of the target computing node during the execution of the current computing task includes: During the execution of the current computing task on the target computing node, the instantaneous CPU utilization and instantaneous memory occupancy of the target computing node are collected at fixed time intervals. Based on the instantaneous utilization rate of the central processing unit and the instantaneous memory occupancy rate, the comprehensive resource utilization rate of the target computing node at each data acquisition moment is calculated.
7. The adaptive scheduling method for large-scale model computation tasks according to claim 6, characterized in that, Adjusting the position of the target computing node in the dynamic resource queue based on the resource utilization rate includes: The overall resource utilization rate at each data collection time is compared with the preset high load threshold and low load threshold; When the overall resource utilization rate is higher than the high load threshold for a consecutive preset number of times, the sorting position of the target computing node in the dynamic resource queue is moved backward by a preset number of positions. When the overall resource utilization rate is lower than the low load threshold for a consecutive preset number of times, the sorting position of the target computing node in the dynamic resource queue is moved forward by a preset number of positions.
8. The adaptive scheduling method for large-scale model computation tasks according to claim 1, characterized in that, After adjusting the position of the target computing node in the dynamic resource queue according to the resource utilization rate, the method further includes: Obtain the remaining unexecuted computation amount of the current computing task, and calculate the expected remaining execution time of the current computing task based on the remaining unexecuted computation amount and the current resource utilization rate of the target computing node; When the expected remaining execution time exceeds the preset deadline of the current computing task, the standby computing node with the largest amount of available computing resources is selected from the dynamic resource queue. The current computing task is migrated from the target computing node to the standby computing node to continue execution.
9. The adaptive scheduling method for large-scale model computation tasks according to claim 8, characterized in that, The step of selecting the standby computing node with the largest amount of available computing resources from the dynamic resource queue includes: Exclude computing nodes currently executing computing tasks from the dynamic resource queue to obtain a set of standby computing nodes; The standby computing node with the highest sorting position from the set of standby computing nodes is selected as the standby computing node with the largest amount of available computing resources.
10. The adaptive scheduling method for large-scale model computation tasks according to claim 8, characterized in that, The step of migrating the current computing task from the target computing node to the standby computing node for continued execution includes: Pause the execution of the current computing task on the target computing node, and record the current execution progress and current intermediate state data of the current computing task; The current execution progress and the current intermediate state data are transmitted to the standby computing node; The current computing task is resumed on the standby computing node based on the current execution progress and the current intermediate state data.