Large model reasoning scheduling method based on off-grid computing server

By collecting hardware information, decomposing tasks, optimizing scheduling and migration tasks in an off-grid computing power server environment, the problems of unreasonable resource allocation and unbalanced load of large model inference tasks in an off-grid environment are solved, and efficient task-resource matching and load balancing are achieved.

CN119537032BActive Publication Date: 2025-05-20BEIJING QIBU TIANXIA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510088550.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-20
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

In off-grid computing power server environment, the lack of stable network connections leads to unreasonable resource allocation, low dynamic scheduling efficiency and unbalanced load of large-scale inference tasks.

Method used

By collecting hardware information of off-grid computing power servers, building server resource vectors, and decomposing large-model inference tasks into multiple subtasks. Use optimization methods to divide tasks, dynamic task scheduling is performed based on the load state of the server, and dynamic migration of tasks ensures load balancing between servers.

Benefits of technology

It achieves efficient matching of tasks and resources, improves the overall efficiency of large-scale model inference and the utilization rate of computing power resources, avoids single point overload and resource idleness, and has good environmental adaptability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119537032B_ABST
    Figure CN119537032B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of server scheduling, and discloses a large model reasoning scheduling method based on an off-grid computing server, comprising the following steps: collecting hardware information of the off-grid computing server and constructing a server resource vector; decomposing the large model reasoning task into multiple subtasks and constructing task modeling; dividing the tasks using an optimization method based on the hardware information and task dependencies of the off-grid server; dynamically scheduling the subtasks based on the current load status of the server; dynamically migrating some tasks to other servers based on the real-time computing load of the server to ensure load balance between servers; executing the assigned tasks, monitoring the operating status of the server and the completion of the tasks, and dynamically adjusting the task allocation strategy based on the feedback data. By modeling task requirements and server resources, combined with real-time monitoring and optimized scheduling, efficient computing and maximizing resource utilization of large model reasoning in an off-grid environment are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of server scheduling, and specifically to a large model inference scheduling method based on off-grid computing power servers. Background Art

[0002] With the rapid development of artificial intelligence technology, the wide application of large models (such as Transformer, GPT, etc.) in fields such as natural language processing and computer vision has led to an exponential growth in their inference computing requirements. The inference process of large models usually involves a large number of complex computational operations, especially matrix operations, attention mechanism calculations, etc. These tasks not only require high-performance computing power resources but also pose higher requirements for task allocation and scheduling strategies. In traditional centralized computing power environments, such large-scale inference tasks are usually completed through stable network connections with cloud computing power clusters. However, as more task scenarios extend towards off-grid edge devices and distributed servers, the management of this computing demand has become more complex.

[0003] Generally speaking, under the condition of a good network environment, task scheduling can rely on centralized control and global information support, and the scheduling strategy can be adjusted dynamically in real time to achieve multi-node collaborative computing. However, in the off-grid computing power server environment, due to the lack of a stable network connection, it is difficult for computing power servers to achieve efficient cooperation. In addition, the server hardware resources in the off-grid environment usually have heterogeneity (such as different CPU, GPU performance, and memory capacity), and more refined optimization strategies are required for task allocation and scheduling. Traditional scheduling methods often do not accurately model resource requirements and are difficult to meet the diverse computing needs in the off-grid environment, resulting in a decline in computing efficiency.

[0004] Existing scheduling technologies in off-grid or weak network environments generally rely on a stable network environment to collect server resource information and task progress status in real time. When the network is unstable or missing, it is difficult for the scheduling system to obtain the task execution status in a timely manner, resulting in lag in task allocation and unreasonable scheduling. And due to the lack of accurate resource requirement modeling and dynamic feedback optimization, existing technologies usually adopt static strategies in task allocation and do not fully consider the real-time changes in server load during task execution. Problems such as resource idleness and single-point overload are inevitable, leading to low overall resource utilization.

[0005] In addition, in multi-server collaborative computing, the data communication cost between tasks is ignored or not fully optimized, which further increases the overall latency of the computing. Especially in the scenario of large model inference with complex task dependencies, the latency of data transmission may even become the main bottleneck of the system. It can be seen that the existing technologies lack an efficient scheduling mechanism that can achieve dynamic resource allocation, task optimization, and load balancing in the off-grid computing power server environment, and it is difficult to adapt to complex computing requirements and resource distribution characteristics.

[0006] Therefore, the present invention proposes a large model inference scheduling method based on an off-grid computing power server to solve the deficiencies of the existing technologies. Summary of the Invention

[0007] Aiming at the deficiencies of the existing technologies, the present invention provides a large model inference scheduling method based on an off-grid computing power server, which solves the problems of unreasonable resource allocation, low dynamic scheduling efficiency, and load imbalance of large model inference tasks in the off-grid computing power server environment due to the lack of a stable network.

[0008] To achieve the above objectives, the present invention is realized through the following technical solutions: A large model inference scheduling method based on an off-grid computing power server includes the following steps:

[0009] S1. Collect the hardware information of the off-grid computing power server and construct a server resource vector;

[0010] S2. Decompose the large model inference task into multiple subtasks and construct a task model;

[0011] S3. According to the hardware information of the off-grid server and the task dependencies, use an optimization method to perform task partitioning;

[0012] S4. Based on the current load status of the server, perform dynamic task scheduling on the subtasks;

[0013] S5. According to the real-time computing load of the server, dynamically migrate some tasks to other servers to ensure load balance between servers;

[0014] S6. Execute the assigned tasks, monitor the running status of the server and the task completion situation, and dynamically adjust the task allocation strategy according to the feedback data.

[0015] Preferably, the task model includes:

[0016] Construct a task-resource matrix, and the elements of this matrix represent the computing cost required for each task on each server, and the computing cost is calculated based on the CPU requirements, GPU requirements, and memory requirements of the task and the available resources of the server.

[0017] Preferably, the task partitioning includes:

[0018] Use the matrix decomposition method to perform low-rank decomposition on the task-resource matrix, and divide the tasks into multiple subtask sets;

[0019] Each subtask set corresponds to a server, and the computing requirements of the subtask set are matched with the available resources of the server.

[0020] Preferably, the matrix decomposition method includes using the tensor decomposition method to perform optimized decomposition on the task-resource matrix, and the goal of task division is to minimize the migration cost of tasks between servers and the imbalance of computing loads.

[0021] Preferably, the dynamic task scheduling is implemented based on the game theory optimization model, and the game theory optimization model includes the following:

[0022] Model each server as a player in the game;

[0023] The utility function of each player includes two parts: the computing cost of the subtask and the ratio of the current load of the server to its total resource capacity;

[0024] Allocate tasks according to the maximization result of the utility function, and design load balancing adjustment.

[0025] Preferably, the calculation of the utility function includes:

[0026] The computing cost of the task on the server;

[0027] The ratio of the current load of the server to the total computing power of the server;

[0028] The weight factor of task scheduling, which is used to balance the computing cost and load balancing.

[0029] Preferably, the load balancing adjustment is implemented through a dynamic load balancing model, which determines the dynamic migration strategy of tasks according to the current load of the server and the task migration cost.

[0030] Preferably, the execution of the allocated tasks includes:

[0031] Real-time monitor the running status of the server, including the utilization rates of CPU, GPU and memory;

[0032] Adjust the task allocation strategy according to the running status of the server, and preferentially allocate new tasks to the servers with lower loads;

[0033] Reallocate the unfinished tasks to ensure the fault tolerance of the system.

[0034] Preferably, the load balancing model uses a differential dynamic system to model the dynamic changes in server load. The changes in server load consist of three parts: load attenuation, spontaneous balance, and task migration. The ultimate goal is to make the loads of all servers tend to the global average load.

[0035] The present invention also provides a scheduling system for a large model inference scheduling method, including the following modules:

[0036] A task initialization module, used to collect server hardware information and the requirement information of large model tasks;

[0037] A task division module, used to optimize the division of tasks and generate a set of subtasks;

[0038] A dynamic scheduling module, used to monitor the load status of servers in real time and dynamically allocate tasks;

[0039] A load balancing module, used to dynamically migrate tasks according to the load balancing model;

[0040] An execution and feedback module, used to execute tasks and monitor the completion status of tasks in real time, and provide feedback data to optimize the scheduling strategy.

[0041] The present invention provides a large model inference scheduling method based on off-grid computing power servers. It has the following beneficial effects:

[0042] 1. By accurately modeling the server resource status and task requirements, and combining tensor decomposition and game theory to optimize the dynamic scheduling strategy, the present invention realizes the efficient matching of tasks and resources. Through dynamic load migration and task scheduling optimization, it effectively avoids the performance degradation caused by resource bottlenecks in a single server, and significantly improves the overall efficiency of large model inference and the utilization rate of computing power resources.

[0043] 2. Aiming at the problem that off-grid computing power servers cannot rely on centralized network scheduling, the present invention constructs a distributed scheduling mechanism independent of the network environment. Through localized task decomposition, division, and dynamic adjustment strategies, the system can operate efficiently in the case of network absence or instability, and has strong environmental adaptability.

[0044] 3. By monitoring the load status of servers in real time, combining the differential dynamics load balancing model and the task migration strategy, the present invention dynamically adjusts the task distribution of servers. It effectively avoids the problems of single-point overload and resource idleness, and at the same time improves the stability of task execution and the overall robustness of the system.

[0045] 4. Through the task priority scheduling, real-time feedback optimization, and dynamic concurrency adjustment mechanisms, the present invention adapts to various types of large model inference tasks. Whether it is a task with high real-time requirements or a large-scale inference task that requires high computing power support, the present invention can flexibly adjust the scheduling strategy to ensure the rapid completion of tasks and the efficient utilization of resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a flowchart of the steps of the large model inference scheduling method based on the off-grid computing power server of the present invention;

[0047] Figure 2 It is a module framework diagram of the large model inference scheduling system based on the off-grid computing power server of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0049] Please refer to the attached Figure 1 , the embodiments of the present invention provide a large model inference scheduling method based on an off-grid computing power server, including the following steps:

[0050] S1. Collect the hardware information of the off-grid computing power server and construct a server resource vector;

[0051] In the off-grid computing power server environment, to achieve efficient scheduling of large model inference tasks, the system first needs to understand the resource situation of the computing power server. This step is the basis for subsequent task decomposition, partitioning, and dynamic scheduling. By comprehensively collecting and modeling the hardware information of the server, it can provide the necessary data support for resource matching and task allocation. Generally, there is no direct network connection between off-grid servers. Therefore, the collection of hardware information and the initialization of resource status need to be completed locally. To ensure the accuracy of scheduling, the information collected in this step includes not only static hardware parameters but also real-time resource status to provide dynamic support for subsequent steps.

[0052] In this embodiment, the hardware information of the off-grid computing power server mainly includes the CPU computing power, GPU computing power, and memory capacity of each server. This information is collected through hardware monitoring tools, and the specific implementation methods include calling system-level API interfaces, parsing server hardware configuration files, or directly running test scripts. For example, the CPU computing power can be obtained by reading the logical core count and main frequency information of the system; the GPU computing power can be estimated based on the floating-point operation ability of the GPU (such as FLOPS); and the memory capacity is confirmed by reading the physical memory size of the server.

[0053] Specifically, the hardware resources of each server can be modeled as a resource vector:

[0054]

[0055] Among them, is the total CPU computing power of the server, is the total GPU computing power of the server, in floating-point operation times, is the total memory capacity of the server.

[0056] In a possible implementation, by collecting the hardware information of multiple servers in the cluster, a global resource matrix can be constructed

[0057]

[0058] Among them, each row of the matrix is the resource information of a server, and it contains a total of servers

[0059] As an option, to better adapt to the real-time requirements of task allocation, the collected hardware information can also be dynamically updated in combination with the current resource usage status of the server. Generally, the real-time resource usage information of the server includes the current CPU utilization rate, GPU utilization rate, and memory occupancy rate, which are represented by , , respectively. In this way, the available resource vector of the server can be further defined as:

[0060]

[0061] Among them, is the currently available CPU computing power of the server, is the currently available GPU computing power of the server, is the currently available memory capacity of the server

[0062] In a possible implementation, the real-time load data of the server can be obtained through a hardware monitoring tool. These real-time data can be used for subsequent dynamic task scheduling and load balancing operations.

[0063] In addition, to improve the accuracy of information collection, in this embodiment, the collection tool can sample the resources multiple times and calculate the average value. For example, the GPU computing power can be estimated by taking the average of 10 samples of the peak floating-point operation monitoring:

[0064]

[0065] where, is the GPU floating-point operation performance at the th sample. By sampling multiple times, the influence of instantaneous fluctuations can be effectively filtered, thereby obtaining more stable resource information.

[0066] In some embodiments, the resource information of the server can also be normalized in combination with the characteristics of the task requirements. Specifically, the hardware resource vector of the server can be standardized to the range of 0 to 1 to better adapt to the task allocation algorithm. For example, the normalized resource vector can be defined as:

[0067]

[0068] where, 、 、 are the maximum resource values of the CPU, GPU, and memory in the cluster respectively, is the normalized resource vector.

[0069] In another implementation, to support multiple types of tasks, the server resource information can be further extended. For example, for tasks that require storing a large amount of data, disk / performance parameters can also be added. This extension can provide more-dimensional data support for the task allocation algorithm, making the allocation strategy more flexible.

[0070] Through the above methods, the process of collecting and modeling the server hardware information is completed. The subsequent steps will use this resource information to reasonably decompose and allocate the large model inference tasks, so as to achieve the optimal utilization of resources and the efficient scheduling of tasks.

[0071] Step S2: Decompose the large model inference task into multiple subtasks and construct a task model;

[0072] The main task of this step is to decompose the complex large model inference process into multiple subtasks and construct a computational requirement model for each subtask. Through the decomposition and modeling of subtasks, the resource distribution characteristics of off-grid computing power servers can be better matched, thereby providing data support for subsequent task partitioning, dynamic scheduling, and other processes. Since the hardware information of the off-grid computing power server has been collected in step S1, this step needs to perform refined modeling and correlation description based on the resource status of the server and the logical characteristics of the task itself.

[0073] Generally, the large model inference task can be decomposed into multiple computational stages, which may involve modules such as feed-forward network calculation, attention mechanism, and non-linear activation function processing. The computational requirements of each module have different characteristics, and there are dependencies between some computational modules. For example, attention calculation needs to be executed after the feed-forward network calculation is completed. To clearly describe these logical relationships, the present invention comprehensively represents the execution characteristics of each subtask and its resource requirements by constructing a dependency graph of the task and a task requirement vector.

[0074] In some embodiments, this step also needs to consider the dynamics of tasks during the large model inference process. Since the length and shape of the model input and the intermediate results generated during the inference process may affect the computational requirements, the task modeling needs to have a certain degree of scalability and flexibility to ensure that the decomposed subtasks can adapt to different computing power nodes.

[0075] The large model inference task is first decomposed into a series of interrelated subtasks. Each subtask is modeled by a task requirement vector, and the task requirement vector represents the demand for computing power resources during the calculation process of this subtask. The specific definition is as follows:

[0076]

[0077] Among them, is the resource requirement vector of the subtask, is the demand of the task for CPU resources, usually expressed in the form of the number of cores, is the demand of the task for GPU resources, is the demand of the task for memory resources.

[0078] Specifically, the acquisition of the computational requirement vector can be determined by the following methods:

[0079] According to the structural characteristics of the large model, analyze the computational complexity of each task module. For example, the feed-forward network is usually a matrix multiplication operation, and its computational requirements are directly related to the input dimension and the size of the weight matrix;

[0080] Combined with the hardware architecture characteristics of the off-grid computing power server, further refine the resource requirements of each subtask for the CPU, GPU, and memory. For example, convolution operations may mainly rely on the GPU, while the processing of activation functions has a greater dependence on the CPU;

[0081] In some embodiments, the actual resource consumption of different task modules on the server can be measured through offline testing to obtain a more accurate task requirement vector.

[0082] To represent the dependency relationships between subtasks, the present invention further constructs a task dependency graph (DAG graph). The DAG graph consists of nodes and directed edges, where:

[0083] Each node represents a subtask; each directed edge represents the execution dependency relationship between subtasks. For example, if the output of task is the input of task , then an edge is added in the DAG graph from node to node .

[0084] As an option, the generation of the DAG graph can refer to the network structure of the large model. For example, during the inference process of the Transformer model, the attention mechanism needs to rely on the output of the input embedding layer. Therefore, the subtask node of the attention layer should depend on the embedding layer node. In this case, the structure of the DAG graph can not only reflect the order between tasks but also provide an important basis for subsequent task partitioning.

[0085] In a possible implementation manner, the task decomposition process can also be combined with the technical characteristics of dynamic task modeling. For input sequences with variable lengths (such as natural language processing tasks), the computational requirements of subtasks may change with the input length. To adapt to this characteristic, the present invention proposes a task decomposition method based on dynamic modeling:

[0086] First, according to the length of the input sequence, parametric modeling is performed on the resource requirements of each subtask. For example, for matrix multiplication operations, its GPU requirement can be expressed as:

[0087]

[0088] where , is a parameter related to the task module, determined through experiments; then, adjust the subtask partitioning strategy according to the dynamic requirements of the task. For example, for long sequence inputs, the matrix multiplication task can be further decomposed into multiple subtasks to reduce the resource occupancy of a single task.

[0089] In some embodiments, the task modeling process of the present invention further includes an adaptability analysis of task requirements. Specifically, the sub-task requirement vector is matched with the server resource vector obtained in step S1 to determine the suitable execution environment for each sub-task. The calculation formula for the matching analysis is as follows:

[0090]

[0091] Wherein, is the adaptability score of the server for the task ; , , are the CPU, GPU, and memory resource amounts of the server respectively, , , are the requirements for CPU, GPU, and memory of the task respectively.

[0092] By calculating , a reference basis can be provided for subsequent task partitioning, enabling each task to be assigned to the most suitable server for execution.

[0093] In addition, the modeling of communication requirements also needs to be considered in the task modeling process. The communication requirements are usually expressed as the amount of data dependence between tasks, defined as:

[0094]

[0095] Where is the amount of data from the output of task to task . The modeling of communication requirements can help determine the task partitioning scheme to minimize data transmission across servers. Through the above task decomposition and modeling process, the present invention lays a solid foundation for subsequent task partitioning and scheduling. The calculation requirement vectors and task dependency graphs of sub-tasks can not only accurately describe the resource requirements of tasks, but also reflect the logical relationships between tasks, thus providing support for the dynamic allocation of tasks. The above content can be further expanded in combination with specific implementation scenarios to adapt to different types of large model inference tasks.

[0096] S3. According to the hardware information of the off-grid server and the task dependency relationship, use an optimization method to perform task partitioning;

[0097] In the technical solution of the present invention, the main task of step S3 is to further divide the task into multiple sub-task sets based on the server hardware information collected in step S1 and the task model constructed in step S2. Each sub-task set consists of related sub-tasks and is assigned to a specific server for execution. The goal of task division is to achieve the best match between computing power resources and task requirements, while minimizing the communication cost across servers, thereby providing a basis for subsequent dynamic scheduling.

[0098] Generally, task division needs to consider the computing requirements of the task, the available computing power of the server, and the dependency relationships between tasks. For example, for GPU-intensive tasks, they need to be preferentially assigned to servers with high GPU computing power; for cases where there are a large number of data dependencies between tasks, they are preferably assigned to the same server to reduce the communication overhead of data transmission. As an option, this step performs optimization calculations by constructing a task-resource matrix and uses methods such as tensor decomposition to achieve task division. Specifically, the task division in this step needs to meet the following conditions:

[0099] The resource requirements of each sub-task set cannot exceed the available computing power of the target server;

[0100] Minimize the data transmission volume across servers;

[0101] Follow the dependency relationships of tasks.

[0102] To achieve the best match between tasks and server resources, first construct a task-resource matrix , which is defined as follows:

[0103]

[0104] Among them, is the computing cost of task on the server, , , are respectively the CPU requirement, GPU requirement, and memory requirement of task , , , are respectively the CPU computing power, GPU computing power, and memory capacity of the server.

[0105] In some embodiments, the task-resource matrix can also be extended to a dynamic matrix by adding the current available resource information of the server.

[0106] To optimize the task division result, the task-resource matrix can be optimized by tensor decomposition. Specifically, the task-resource matrix is decomposed into two low-rank matrices:

[0107]

[0108] Among them, is the task feature matrix, which describes the performance of the task on potential computing factors, is the server feature matrix, which describes the capabilities of the server on potential computing factors, is the rank of matrix factorization, satisfying .

[0109] Through matrix factorization, the potential matching relationship between tasks and servers can be extracted, and thus the tasks can be divided into several sub-task sets. The specific division method is as follows:

[0110]

[0111] Among them is the task set assigned to the server, and the goal of task division is to minimize the sum of all .

[0112] In order to further reduce the communication overhead of task division, the influence of task dependency relationships on the division results is also considered in this embodiment. The dependency relationships between tasks can be represented by the communication requirement matrix , and the matrix element is defined as:

[0113]

[0114] Among them is the amount of data output from task to task .

[0115] When dividing tasks, tasks with close dependency relationships are preferentially assigned to the same server. The communication optimization goal can be expressed by the following formula:

[0116]

[0117] Among them, is the communication weight coefficient, which is used to balance the computing cost and the communication cost, are the task sets assigned to the servers respectively.

[0118] To improve the flexibility of task division, the present invention can adopt a multi-stage division strategy. First, the tasks are divided into coarse-grained task groups, and the tasks within each task group have strong dependency relationships and are suitable as the same sub-task set. Then, according to the resource requirements of each task group, it is further refined to ensure that each sub-task set can be efficiently executed on the target server.

[0119] For example, in the Transformer structure of large models, coarse-grained partitioning can be carried out according to the model levels. The feed-forward network and attention mechanism of each layer are regarded as a task group. Subsequently, based on the GPU computing power of the server, the task group is further split into smaller computing units to be assigned to different servers for execution. The result of task partitioning will directly affect the subsequent scheduling efficiency. Therefore, in this embodiment, task adaptability analysis is also combined to ensure that the set of subtasks after partitioning can meet the resource conditions of the target server. The scoring function of the adaptability analysis is as follows:

[0120]

[0121] where is the adaptability score of the server for task , , , are the available resources of the server respectively, , , are the resource requirements of the task respectively.

[0122] By calculating the adaptability score, the result of task partitioning can be further optimized to ensure the maximization of resource utilization. Through the above task partitioning method, the present invention realizes the comprehensive optimization of computational cost, communication cost and task dependencies, providing a high-quality initial task distribution scheme for subsequent dynamic scheduling and load balancing. The result of task partitioning not only improves the utilization efficiency of resources, but also reduces the data transmission overhead across servers. The above method can be flexibly adjusted according to specific server configurations and task requirements.

[0123] S4. Dynamically schedule the subtasks based on the current load status of the server;

[0124] In the technical solution of the present invention, step S4 is a key link for dynamically optimizing the initial task partitioning result generated in step S3. Since the hardware resource status of off-grid computing power servers may fluctuate dynamically over time, for example, the load of some servers increases or resources become insufficient during execution, the initial task partitioning scheme may no longer be applicable. The purpose of this step is to monitor the load status of each server in real time and dynamically schedule subtasks based on the monitoring data to ensure the rationality of task allocation, thereby further improving the task execution efficiency. Generally, dynamic task scheduling needs to comprehensively consider the current load status of the server, the computing requirements of the task, and the priority of the task. On this basis, the present invention realizes dynamic task scheduling by constructing a task scheduling model and adopting a game theory optimization strategy. In this design, the server is modeled as a rational player in the game, and the goal of each player is to maximize its own utility function and achieve balanced optimization of task allocation globally.

[0125] In some embodiments, dynamic scheduling also needs to combine the execution priority of the task and allocate resources preferentially to high-priority tasks. In addition, the scheduling system needs to update the server status in real time to cope with sudden resource changes.

[0126] Dynamic task scheduling takes the load status of the server as the core scheduling basis. Specifically, the real-time load status of the server is defined as:

[0127]

[0128] where is the load status of the server, the total resource requirement of the currently running tasks is the total requirement of the running tasks on the CPU, GPU, and memory, and the total server resources is the sum of the CPU, GPU, and memory of the server.

[0129] Based on the load status of the server, the present invention designs a dynamic scheduling mechanism, and the scheduling goal is to minimize the load imbalance of all servers while meeting the resource requirements of the tasks.

[0130] Specifically,

[0131] In this embodiment, a game theory optimization model for task scheduling is constructed. The main contents of the game model include:

[0132] Definition of players: Each server is modeled as a player in the game.

[0133] Strategy space: The strategy of each player is the set of tasks allocated to itself . For each task , the player can choose whether to accept the task.

[0134] Utility function: The utility function of each server is defined as follows:

[0135]

[0136] Wherein, represents the computing cost of task on the server, which is defined by the task-resource matrix in step S3, represents the load status of the server, represents the total computing power capacity of the server, is a weight factor used to balance the optimization goals of task computing cost and load balancing.

[0137] By maximizing the utility function, each server will preferentially select tasks suitable for its own resource conditions when allocating tasks, while avoiding excessive load on itself.

[0138] In a possible implementation, the dynamic scheduling of tasks can be achieved through the following process:

[0139] First, obtain the real-time load status of the server ;

[0140] Then, calculate the adaptability score according to the resource requirements of the task and the available resources of the server , and the formula is as follows:

[0141]

[0142] Wherein, is the adaptability score of task on the server, , , are the currently available CPU, GPU, and memory resources of the server respectively, , , are the CPU, GPU, and memory requirements of task respectively.

[0143] According to the scoring results, tasks with higher adaptability are preferentially allocated to suitable servers.

[0144] When the load of some servers reaches the bottleneck (for example > 0.8), the task migration mechanism will be triggered in this embodiment to migrate some tasks on the high-load server to the low-load server. The decision basis for task migration is the migration cost and the adaptability score of the target server. The definition of the migration cost is as follows:

[0145]

[0146] Among them, is the data transfer volume of the task , , , is the task resource requirement, , , is the total server resource volume.

[0147] In some embodiments, dynamic scheduling can also be combined with the execution priority of tasks. Tasks with higher priorities will be preferentially allocated to servers with lower loads. For example, for tasks with high real-time requirements, their priorities can be set to higher values. The dynamic scheduling model combined with priorities can be expressed as:

[0148]

[0149] By introducing task priorities, the dynamic scheduling system can flexibly adjust the allocation strategy according to the importance of tasks to ensure that critical tasks are completed first.

[0150] Through the above dynamic task scheduling mechanism, the present invention effectively solves the task allocation problem of off-grid computing power servers in a load fluctuation environment. This method can not only dynamically adapt to resource changes, but also maximize the resource utilization rate of servers, while ensuring the real-time performance of task execution and the balance of the overall system.

[0151] S5. Dynamically migrate some tasks to other servers according to the real-time computing load of the servers to ensure load balance among the servers;

[0152] In the present invention, the main function of step S5 is to achieve dynamic load balancing among off-grid computing power servers. Since the dynamic scheduling mechanism in step S4 may not completely avoid load imbalance, especially during task execution, some servers may have resources approaching the bottleneck due to task accumulation, while other servers may have more idle resources. Therefore, in this step, by real-time monitoring the load status of each server and calculating the migration strategy based on the load dynamic model, some tasks on high-load servers are migrated to low-load servers to ensure that the computing resources of the entire cluster are fully utilized.

[0153] Generally, when considering load migration, the following factors need to be taken into account: the real-time load status of the server, the communication cost of task migration, and the resource adaptability of the target server. In some embodiments, the dependency relationship of tasks also needs to be considered to ensure that the migrated tasks do not affect the overall logical order of computing. In addition, to avoid the adverse impact of frequent migration operations on performance, a certain trigger threshold is set for the migration conditions in this step. For example, migration is triggered when the load status of a high-load server exceeds a preset value.

[0154] In this embodiment, the core of load migration is to monitor the load status of the server in real time and calculate the specific task migration plan based on the load dynamic model. The real-time load status of the server is defined as:

[0155]

[0156] In a possible implementation, when the load status of a certain server exceeds the set threshold (such as 0.8), the system will start the load migration mechanism to migrate some tasks on the high-load server to the low-load server.

[0157] Specifically, this embodiment realizes task migration through the following three steps: task selection, target server selection, and migration cost calculation.

[0158] Select tasks with higher computing requirements for migration from the task queue of the high-load server, and preferentially select the following types of tasks:

[0159] Tasks that run for a long time;

[0160] Tasks with high GPU resource requirements;

[0161] Tasks with weak dependency relationships with other tasks.

[0162] The specific task selection priority

[0163]

[0164] where , , are weight coefficients, set according to the specific application scenario, , , are the GPU, CPU, and memory requirements of task respectively, is the communication demand of task with other tasks.

[0165] Priority migration priority Higher-priority tasks can reduce interference with other tasks.

[0166] In a low-load server, select the target server that is most suitable for migrating tasks. The selection of the target server is based on the adaptability score , and its calculation formula is:

[0167]

[0168] Among them, represents the adaptability score of the server for task , , , respectively represent the current available CPU, GPU, and memory resources of the server.

[0169] Preferably select the server with a higher adaptability score as the target server.

[0170] The final migration goal is to minimize while maximizing .

[0171] In one possible implementation, the load balancing model can further describe the dynamic changes of the server load through differential dynamics equations. The load change of the server can be expressed as:

[0172]

[0173] Among them, is the natural decay coefficient of the server load after the task is completed, is the server to the server the speed factor of migrating tasks, is the load increment caused by new tasks.

[0174] Through this equation, the trend of server load changes can be simulated and predicted, providing a basis for task migration decisions.

[0175] In some embodiments, task migration also needs to consider the priority and execution stage of the task. For tasks with a higher priority, it is not recommended to migrate during execution, unless the load status of the server where it is located seriously exceeds the standard (for example > 0.95). In addition, tasks are more suitable for migration in the initial stage of execution (such as when they are just assigned), and migration should be avoided as much as possible when the execution is approaching completion. The timing of migration can be quantified by the following formula:

[0176]

[0177] Among them, is the priority of task migration, is the current execution time of the task, is the mid-execution time point of the task.

[0178] Through this formula, the migration priority of the task can be dynamically adjusted, thereby reducing the impact of migration on the task execution time.

[0179] As an option, this embodiment also supports batch task migration. When the load of the server exceeds the standard by a large margin, multiple tasks can be migrated simultaneously to quickly relieve the load pressure. The selection basis for batch tasks is the minimization of the total migration cost The formula is as follows:

[0180]

[0181] Among them, is the set of tasks to be migrated from the server, is the migration cost of a single task.

[0182] Through the above method, the load between servers can be quickly balanced, ensuring the efficient execution of tasks.

[0183] Through the above task migration mechanism, the present invention realizes the dynamic adjustment of the off-grid computing power server under high load, further optimizing the global utilization efficiency of computing resources. The task migration process comprehensively considers communication cost, target adaptability, and execution priority, thereby significantly improving the flexibility and robustness of the system while ensuring the task completion quality. This mechanism can be further extended and optimized according to different task scenarios to adapt to different resource distributions and computing requirements.

[0184] S6. Execute the assigned tasks, monitor the running status of the server and the task completion situation, and dynamically adjust the task allocation strategy according to the feedback data.

[0185] In the technical solution of the present invention, step S6 is the closed-loop part of the entire scheduling process. Based on the resource information collection, task modeling, task partitioning, dynamic scheduling, and load balancing implemented in steps S1 to S5, the assigned tasks finally need to be executed on the off-grid computing power server. At the same time, the execution process of the task is not static, and the running status of the server may be dynamically affected by external factors or task progress. Therefore, real-time monitoring of the task execution status and dynamically adjusting the scheduling strategy according to the monitoring results is the key link to improving the system efficiency and robustness.

[0186] In general, the monitoring of the task execution process includes two core aspects: one is the running state of the server, mainly reflected in the utilization rates of hardware resources such as CPU, GPU, and memory; the other is the completion status of the task, such as the execution progress of the current task and the queuing situation of uncompleted tasks. As an option, the feedback data can be collected in real time through the resource monitoring module and have a feedback effect on the allocation strategy of the scheduling system.

[0187] Specifically, this step realizes a closed-loop optimization method of dynamic adjustment through the monitoring and feedback mechanism to cope with abnormal situations in task execution and uneven resource distribution.

[0188] In this embodiment, the execution of the task on the off-grid computing power server is determined by the allocation plan generated by the task scheduling module. Each server executes the assigned subtasks according to the scheduling result and reports its running state and task progress to the monitoring module in real time during the execution process.

[0189] In order to comprehensively master the task execution process, a set of monitoring parameters is designed in this step to reflect the running state of the server and the completion status of the task in real time. The monitoring parameters include but are not limited to: : The CPU utilization rate of the server, expressed as the proportion of CPU resources occupied by the currently running tasks; : The GPU utilization rate of the server, expressed as the proportion of GPU computing power occupied by the currently running tasks; : The memory utilization rate of the server, expressed as the proportion of memory resources occupied by the currently running tasks; : The task Execution progress on the server, expressed as the ratio of the completed computing amount of the task to the total computing amount.

[0190] In some embodiments, the monitoring module can obtain the above parameters through built-in hardware sensors and system APIs. For example, the utilization rate of the GPU can be collected in real time through runtime tools (such as the NVIDIA-smi command), and the memory utilization rate can be obtained through the memory management interface of the operating system.

[0191] In this step, the monitoring data will be transmitted to the feedback module of the scheduling system in real time. The feedback module analyzes the running state of the server and the task completion status, and dynamically adjusts the scheduling strategy according to the analysis results.

[0192] The adjustment strategies of the feedback module include but are not limited to the following situations:

[0193] For the assigned tasks, when their execution progress lags behind the expectation (for example When it is the case of

[0194] When the hardware utilization rate of the server approaches the bottleneck (for example or ), the feedback module will trigger the task migration mechanism to re - allocate some tasks to the servers with low load.

[0195] For tasks that are still in the queue, the feedback module will optimize the resource distribution of the servers by combining the monitoring data to shorten the waiting time of the unfinished tasks.

[0196] In a possible implementation, the scheduling system can achieve balanced utilization of server resources by dynamically adjusting the task allocation ratio. Assume that the task queue in the current system , each task has a resource requirement of . To optimize the allocation ratio, the system will calculate the load weight of each server , and the formula is as follows:

[0197]

[0198] where , , are the current hardware utilization rates of the servers respectively; , , are the total resource amounts of the servers respectively.

[0199] The system preferentially allocates new tasks to the server with the lowest load weight to achieve load balancing.

[0200] As an option, the monitoring module also supports fault detection and fault - tolerance mechanisms. For example, when a server stops running due to hardware failure or abnormal status, the monitoring module will trigger an alarm and migrate all unfinished tasks to other servers. The determination conditions for fault detection can include: the resource utilization rate of the server is continuously zero (such as and ); the execution progress of the task has no change for a long time (such as remains unchanged for a period of time).

[0201] In this case, the feedback module will preferentially allocate the affected tasks to the currently optimal server and recalculate the global scheduling scheme.

[0202] In this case, the feedback module will preferentially allocate the affected tasks to the currently optimal server and recalculate the global scheduling scheme.

[0203] Through the above task execution and feedback mechanism, the present invention establishes a closed-loop optimization system. This system monitors the server status and task progress in real time during task execution, and ensures the efficient operation of the system by dynamically adjusting the allocation strategy. This design not only improves the flexibility and robustness of task execution, but also can adapt to different computing requirements in an environment with dynamically changing resources. The above content can be further expanded and optimized in combination with specific application scenarios to cope with different types of tasks and system architectures.

[0204] In summary, the present invention provides a large model inference scheduling method based on an off-grid computing power server. Through steps such as systematic resource information collection, task decomposition modeling, task partitioning, dynamic scheduling, load balancing, and feedback adjustment, an intelligent scheduling mechanism suitable for off-grid environments is constructed. By accurately modeling the server resource status and task requirements, and adopting optimization algorithms such as tensor decomposition and game theory, this method realizes the efficient matching of tasks and resources, and optimizes the task execution process through dynamic load migration and real-time feedback, improving the inference speed, resource utilization rate, and system stability. Aiming at the computing power distribution characteristics of off-grid environments, the present invention effectively solves the problems of resource imbalance and task scheduling complexity, and has broad application prospects and technical feasibility.

[0205] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A large model reasoning scheduling method based on an off-grid computing server, characterized in that: The following steps are involved: Collect the hardware information of off-grid computing servers and build server resource vectors; Decompose the large model reasoning task into multiple subtasks and build task modeling; According to the hardware information and task dependencies of the off-grid server, the task is divided using the optimization method; Dynamically schedule subtasks based on the current load status of the server; Dynamically migrate some tasks to other servers based on the real-time computing load of the server to ensure load balance between servers; Execute assigned tasks, monitor the server's operating status and task completion, and dynamically adjust task allocation strategies based on feedback data; The task modeling includes: Construct a task-resource matrix, where the elements of the matrix represent the computational cost required for each task on each server, and the computational cost is calculated based on the CPU requirement, GPU requirement, and memory requirement of the task and the available resources of the server; Use directed acyclic graphs to describe the dependencies between tasks; The task division includes: The task-resource matrix is ​​decomposed into low-rank by using matrix decomposition method, and the task is divided into multiple sub-task sets; Each subtask set corresponds to a server, and the computing requirements of the subtask set are matched with the available resources of the server; The dynamic task scheduling is implemented based on a game theory optimization model, which includes the following contents: Model each server as a player in the game; Each player’s utility function consists of two parts: the computational cost of the subtask and the ratio of the server’s current load to its total resource capacity; Assign tasks based on the maximization result of the utility function and design load balancing adjustments; The load balancing adjustment is implemented through a dynamic load balancing model, which determines the dynamic migration strategy of the task according to the current load of the server and the task migration cost; The load balancing model uses a differential dynamic system to model the dynamic changes of server loads. The changes in server loads consist of three parts: load attenuation, spontaneous balancing, and task migration. The ultimate goal is to make the loads of all servers approach the global average load.

2. The large model reasoning scheduling method based on off-grid computing server according to claim 1 is characterized in that: The matrix decomposition method includes using a tensor decomposition method to optimize the decomposition of the task-resource matrix, and the goal of task division is to minimize the migration cost of tasks between servers and the imbalance of computing load.

3. The large model reasoning scheduling method based on off-grid computing server according to claim 1 is characterized in that: The calculation of the utility function includes: The computational cost of the task on the server; The ratio of the current server load to the total server computing power; The weight factor of task scheduling is used to balance the computing cost and load balancing.

4. The large model reasoning scheduling method based on off-grid computing server according to claim 1 is characterized in that: The tasks of executing the assignment include: Real-time monitoring of server operating status, including CPU, GPU and memory utilization; Adjust the task allocation strategy according to the server's operating status, and prioritize allocating new tasks to servers with lower loads; Reallocate unfinished tasks to ensure the system's fault tolerance.

5. A scheduling system for a large model reasoning scheduling method, according to any one of claims 1-4, the large model reasoning scheduling method based on an off-grid computing server, characterized in that: Includes the following modules: Task initialization module, used to collect server hardware information and demand information of large model tasks; Task division module, used to optimize the task division and generate a set of subtasks; Dynamic scheduling module, used to monitor the server load status in real time and dynamically allocate tasks; Load balancing module, used to dynamically migrate tasks according to the load balancing model; The execution and feedback module is used to execute tasks and monitor the completion of tasks in real time, and provide feedback data to optimize scheduling strategies.

Citation Information

Patent Citations

  • Resource unified scheduling method and system for multiple types of loads

    CN118819864A

  • Self-adaptive computing power scheduling system for large model training

    CN119322682A