Multi-user computing power quota intelligent queue jumping scheduling method and system
By constructing the current computing power characteristic matrix and using neural network models for dynamic scheduling, the problems of inflexible allocation of GPU resources and inaccurate matching of task requirements in the existing technology are solved, and efficient computing power resource utilization and task execution efficiency are achieved.
Patent Information
- Application Number
- CN202510512536.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing GPU resource scheduling methods are difficult to dynamically adjust computing power allocation in a multi-user environment, resulting in high-performance GPU resources being occupied by low-demand tasks. High-priority tasks may be delayed due to insufficient resources, and there is a lack of in-depth analysis of the current computing power status of the GPU and the phased tasks requirements, resulting in high task migration costs and low resource utilization.
By obtaining the task information to be scheduled and the GPU status information, the current computing power characteristic matrix is constructed, and a neural network model integrating physical models and data-driven algorithms is used, and feature extraction and multiple iterative corrections are combined with attention mechanisms to dynamically evaluate the GPU resource allocation situation, determine whether to perform queue scheduling, and re-arrange resources for preempted tasks.
It effectively improves the utilization rate of computing power resources and task execution efficiency, ensures the accuracy of task scheduling and GPU load balancing, reduces task seizing costs, and improves fairness and adaptability in a multi-tenant environment.
Smart Images

Figure CN120029744A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computing power resource scheduling, and in particular to a method and system for intelligent queue-jumping scheduling of multi-user computing power quotas. Background Art
[0002] Existing GPU resource scheduling methods usually allocate resources based on fixed quotas or simple priority rules, which is difficult to fully adapt to the heterogeneous GPU resource management needs in a multi-user environment. In scenarios such as high-performance computing, deep learning training and reasoning, different tasks have significant differences in their requirements for GPU computing power, video memory, bandwidth and other resources. Traditional scheduling strategies often cannot dynamically adjust the allocation of computing power, resulting in high-performance GPU resources being occupied by low-demand tasks, while high-priority tasks may be delayed due to insufficient resources. In addition, when performing task preemption and queue-jump scheduling, existing methods lack in-depth analysis of the current GPU computing power status and task stage requirements, resulting in high task migration costs, low resource utilization, and may even affect the overall system throughput. At the same time, the architectural differences between heterogeneous GPU environments (such as NVIDIA A100, RTX 4090 and domestic GPUs) make task adaptation more difficult, and existing scheduling systems often cannot effectively distinguish and reasonably allocate computing power resources of different architectures. Therefore, there is an urgent need for a computing power management method that can dynamically optimize computing power allocation, support intelligent queue-jump scheduling, and is suitable for heterogeneous GPU environments to improve computing resource utilization and task execution efficiency.
[0003] For example, the Chinese patent application with publication number CN116643877A provides a computing power resource scheduling method, a computing power resource scheduling model training method and a system. After obtaining the status data of the local device in the current state, the status data is input into the computing power resource scheduling model to obtain the scheduling data in the current state, wherein the scheduling data includes the operation data of the target scheduling operation, and, based on the operation data, the target scheduling operation is executed to perform computing power resource scheduling on the local device. This solution can improve the scheduling efficiency of computing power resource scheduling.
[0004] The above existing technologies all have the problems raised by this background technology: in GPU resource scheduling in a multi-user environment, there are problems such as inflexible static quota allocation, inability to accurately match the phased computing power requirements of tasks, insufficient adaptability of heterogeneous GPUs, and lack of intelligent optimization of queue-jumping scheduling strategies. In order to solve the above problems, this application designs a multi-user computing power quota intelligent queue-jumping scheduling method and system. Summary of the invention
[0005] The technical problem to be solved by the present invention is to address the deficiencies of the prior art and provide a multi-user computing power quota intelligent queue-jumping scheduling method and system. First, the information of the tasks to be scheduled and the GPU status information are obtained, and the current computing power characteristic matrix is constructed based on hardware parameters, dynamic operation indicators and physical topology. Subsequently, a neural network model that integrates the physical model and the data-driven algorithm is used in combination with the attention mechanism for feature extraction, and an updated computing power characteristic matrix is generated through multiple rounds of iterative corrections. Based on the matrix, the GPU resource allocation is dynamically evaluated to determine whether to execute queue-jumping scheduling, and the resources of the preempted tasks are rearranged to optimize the overall computing power utilization. The present application effectively improves the efficiency and fairness of computing power allocation in a multi-tenant environment, and is suitable for complex computing task scenarios such as high-performance computing, deep learning training and reasoning.
[0006] To achieve the above object, the present invention provides the following technical solutions: A multi-user computing power quota intelligent queue-jumping scheduling method is applied to a heterogeneous GPU resource environment. The intelligent queue-jumping scheduling method includes: Obtain task information of tasks to be scheduled and status information of heterogeneous GPUs; Determine a current computing power characteristic matrix according to the state information; The task information and the current computing power characteristic matrix are input into a preset neural network model, and an updated computing power characteristic matrix is obtained through an attention mechanism and multiple rounds of iterative correction; According to the updated computing power characteristic matrix, GPU resources are allocated to the task to be scheduled.
[0007] The determining of the current computing power characteristic matrix includes: Acquire hardware parameters and dynamic operation indicators of each GPU unit according to the status information; Generate a comprehensive computing power vector based on the hardware parameters and dynamic operation indicators; Build a resource coordination matrix based on the physical topology of the GPU resource pool; The comprehensive computing power vector is imported into the resource coordination matrix to obtain the current computing power characteristic matrix.
[0008] According to the hardware parameters and dynamic operation indicators, a comprehensive computing power vector is generated, including: Generate a static computing power vector according to the hardware parameters; Calculating a dynamic attenuation coefficient according to the dynamic operation index; Generate an eco-compatibility vector based on the type of GPU; A comprehensive computing power vector is generated according to the static computing power vector, the dynamic attenuation coefficient and the ecological compatibility vector.
[0009] The neural network model is established based on physical model and data driven technology.
[0010] Taking the task information and the current computing power characteristic matrix as input, including: The task information is divided into a data loading phase, a calculation execution phase, and a result writing back phase, and a phased computing power vector is generated according to the resource requirements of each phase; The phased computing power vector is fused with the current computing power characteristic matrix to obtain an initial fusion vector that characterizes the adaptability of each phase to the current GPU idle resources.
[0011] The input to the preset neural network model includes: Load the neural network model built based on the fusion of physical model and data-driven algorithm, and set the initial parameters of the network according to the current computing power characteristic matrix; Inputting the initial fusion vector into the input layer of the neural network model, extracting features of the initial fusion vector through an attention mechanism, and generating a deep representation vector representing the coupling relationship between the real-time load of the GPU and the phased requirements of the task; The depth representation vector is processed by a multi-round iteration mechanism preset in the middle layer of the neural network model, and an intermediate result is output, wherein each iteration round corrects the output according to the dynamic constraints of the depth representation vector and the physical model, and the initial parameters of the network are adjusted by back propagation according to the correction result until the number of iterations is reached.
[0012] The updating of the computing power characteristic matrix includes: Mapping the intermediate result to the dimensional structure of the current computing power characteristic matrix to obtain a mapped preliminary computing power matrix; Conducting a feasibility test on the preliminary computing power matrix; If the test fails, the preliminary computing power matrix is partially adjusted according to the correction rules until the test passes; The initial computing power matrix that passes the inspection is used as the updated computing power characteristic matrix.
[0013] Allocating GPU resources to the task to be scheduled according to the updated computing power characteristic matrix includes: According to the GPUs matching the tasks to be scheduled in the updated computing power characteristic matrix, a GPU resource list of idle GPU resources is counted; Determine whether to perform queue-jump scheduling or preemption operation on the existing running task, if not, directly allocate idle GPU resources to the task to be scheduled; If yes, re-arrange resources for the preempted or reallocated tasks according to the updated computing power characteristic matrix; The idle GPU resources are allocated according to the computing requirements of the tasks to be scheduled, and the final allocation plan and queue-jumping scheduling results are returned to the scheduling management port to complete the running binding of the tasks to be scheduled on the GPU resources.
[0014] A multi-user computing power quota intelligent queue-jumping scheduling system, the system comprising an information collection module, an information processing module and a queue-jumping scheduling module; The information collection module is used to obtain task information of tasks to be scheduled and status information of heterogeneous GPUs; The information processing module is used to determine the current computing power characteristic matrix according to the state information, and input the task information and the current computing power characteristic matrix into a preset neural network model as input to obtain an updated computing power characteristic matrix; The queue-jumping scheduling module is used to allocate GPU resources to the tasks to be scheduled according to the updated computing power characteristic matrix.
[0015] The information processing module comprises: An information preprocessing unit, used to determine the current computing power characteristic matrix; The model processing unit is used to update the computing power characteristic matrix through the attention mechanism and multiple rounds of iterative correction.
[0016] Compared with the prior art, the present invention has the following beneficial effects: The present invention constructs a neural network model that integrates physical models and data-driven algorithms, combines task phase computing power demand analysis with attention mechanism optimization, realizes intelligent matching and dynamic queue-jumping scheduling of heterogeneous GPU resources, and can effectively improve computing power resource utilization and task execution efficiency. Through multiple rounds of iterative correction and resource rearrangement mechanisms, the accuracy of task scheduling and GPU load balancing are ensured, the cost of task preemption is reduced, and fairness and adaptability in a multi-tenant environment are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings: Figure 1 This is a flow chart of a multi-user computing power quota intelligent queue-jumping scheduling method according to Embodiment 1 of the present invention; Figure 2 This is a flow chart of the calculation of the current computing power characteristic matrix in Example 1 of the present invention; Figure 3 This is a structural diagram of a neural network model according to Embodiment 1 of the present invention; Figure 4 This is a module diagram of the multi-user computing power quota intelligent queue-jumping scheduling system in embodiment 2 of the present invention. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0019] Embodiment 1: See also Figure 1 , an embodiment provided by the present invention: a multi-user computing power quota intelligent queue-jumping scheduling method, applied to a heterogeneous GPU resource environment, for high-performance computing, AI training and other types of tasks in a multi-tenant environment, in order to make full use of heterogeneous resources such as A100, RTX series and domestic GPUs, and avoid the failure of sudden high-priority tasks to be executed in time due to fixed quotas, the following steps S1 to S4 are proposed. Through this embodiment, the problem of lack of coordination between multi-stage task requirements and GPU hardware limitations in large-scale GPU scheduling, and the difficulty in balancing multi-tenant fairness during scheduling can be solved, thereby significantly improving the overall computing power utilization and ensuring low-latency response for key tasks. The specific steps of the method are as follows: S1: Obtain task information of the task to be scheduled and status information of the heterogeneous GPU; In this embodiment, the scheduling platform collects the attributes of the tasks to be scheduled (such as task type, priority, video memory requirements, historical execution time) and real-time operating indicators on the GPU side (such as utilization, temperature, power consumption, and video memory occupancy distribution), and records the GPU hardware parameters (such as the number of CUDA cores, FP32 / FP64 peak computing power, and Tensor Core acceleration capabilities) to subsequently build a computing power feature matrix. This provides an accurate task-resource mapping basis for subsequent queue-jumping scheduling, avoiding blind queuing or resource waste due to missing information.
[0020] S2: Determine the current computing power characteristic matrix based on the status information; In this embodiment, by integrating the GPU hardware parameters and dynamic operation indicators obtained in the previous step, a comprehensive computing power vector is first generated, and then a resource coordination matrix is constructed in combination with the physical topology structure to form the current computing power characteristic matrix. Multi-dimensional information such as GPU core utilization, video memory capacity, and power consumption constraints are quantified in matrix form, which is convenient for the subsequent neural network model to directly call and make accurate queue-jumping decisions.
[0021] S3: Input the task information and the current computing power characteristic matrix into the preset neural network model to obtain an updated computing power characteristic matrix; In this embodiment, the information of the task to be scheduled and the current computing power characteristic matrix are input into the neural network model that integrates the physical model and the data-driven algorithm; the model extracts key resource features through the attention mechanism and continuously corrects the output results in multiple rounds of iterations, and finally obtains an updated computing power characteristic matrix. It can dynamically balance the requirements of high-priority tasks, GPU hardware security redundancy, and multi-tenant quota restrictions, so that queue-jumping scheduling can maximize computing power utilization while ensuring system stability.
[0022] S4: Allocate GPU resources for the tasks to be scheduled according to the updated computing power characteristic matrix; In this embodiment, available GPU resources are searched and allocated based on the resource matching degree and priority judgment output in the updated computing power characteristic matrix. When it is confirmed that queue-jumping scheduling is required, necessary resource rearrangement is performed on the low-priority or low-utilization tasks being executed, and high-priority tasks are preferentially bound to the most suitable GPU; if queue-jumping is not required, idle GPU resources are directly allocated. This effectively solves the problems of long-term occupation of resources and inability to execute sudden high-priority tasks in a timely manner caused by traditional fixed quotas, and ensures fairness and efficiency in a multi-tenant environment.
[0023] Specifically, most existing solutions are mainly based on a single or a small number of scheduling dimensions (such as priority, queuing order) for resource matching, lacking a refined evaluation of GPU physical characteristics (such as the number of Tensor Cores, video memory capacity, and degree of fragmentation), and are difficult to flexibly adapt when there are differences in software stacks between domestic GPUs and NVIDIA GPUs; once a high-priority task suddenly occurs in the system, its scheduling decision can often only rely on static rules to preempt low-priority tasks, and cannot be flexibly selected based on real-time power consumption, load, and heat dissipation safety redundancy. Even if some existing solutions have built-in dynamic monitoring mechanisms that can make preemption judgments based on GPU utilization or temperature conditions, it is difficult to subdivide the computing power bottleneck corresponding to each stage for large-scale multi-tenant environments and multi-stage tasks (such as data loading, calculation execution, and result writing back) load patterns; therefore, when there is a sudden queue interruption, there is often a situation of "preemption cost is too high" or "resource pool is still idle", resulting in a decrease in overall efficiency or long-term idleness of some GPUs. In addition, due to the differences in domestic and foreign GPU architectures, drivers, and software frameworks (TensorFlow, PyTorch, etc.), it is usually necessary to manually or based on coarse-grained rules to decide whether to deploy tasks to domestic GPUs or NVIDIA GPUs. Existing scheduling often ignores deep-level dependencies such as local compilation support and mixed-precision computing performance differences, resulting in low utilization of domestic GPUs, or some tasks running on inappropriate GPUs, dragging down performance.
[0024] In this embodiment, the computing power characteristic matrix is introduced as the core intermediate data structure, so that information of multiple dimensions (hardware parameters, dynamic operation indicators, physical topological relationships) can be quantitatively described in vectorized form and matched with the multi-stage requirements of the tasks to be scheduled. For example, A100 has higher Tensor Core acceleration performance in deep learning training scenarios, while the RTX series has advantages in general computing or graphics rendering. Domestic GPUs may have better energy consumption ratios in certain reasoning tasks. With the help of the computing power characteristic matrix, it is possible to clearly identify which type of GPU is most suitable for the computing power requirements of a certain stage during multiple rounds of iterative corrections of the neural network model. Compared with simply relying on priority queues or FIFO strategies, the computing power characteristic matrix not only considers "who arrives first, who has priority", but also explicitly measures the impact of queue-jumping behavior on the overall efficiency of the system and multi-tenant fairness. If the model predicts that preempting a low-priority task on an RTX card will only bring limited performance improvements, but may cause large-scale data writeback or power consumption surges, then the iterative update of the computing power characteristic matrix can timely correct the scheduling decision, thereby avoiding unnecessary large-scale resource migration; if the domestic GPU is idle at this time and is compatible with the demand ecosystem of the high-priority task, the matrix update results will be recommended to be executed on the domestic GPU first, thereby improving the overall utilization of the heterogeneous GPU pool.
[0025] Furthermore, the use of neural network models to integrate physical models and data-driven algorithms can continuously correct GPU quota allocation, power consumption safety redundancy, and multi-tenant queue-jumping strategies in multiple rounds of iterations, which not only solves the problem that simple priority scheduling is difficult to dynamically adapt to changes in heterogeneous GPU loads, but also effectively avoids the "hard switching" between sudden high-priority tasks and tasks being executed, which causes serious losses. With the adaptive update of the computing power characteristic matrix during the heterogeneous GPU scheduling process, not only a more refined multi-dimensional computing power evaluation and allocation is achieved, but also low-latency execution of key tasks can be guaranteed in a multi-user environment, improving overall resource utilization and fairness.
[0026] See also Figure 2 , the current computing power characteristic matrix calculation flow chart of the embodiment of the present invention, the specific steps of S2 are as follows: S2.1: Obtaining hardware parameters and dynamic operation indicators of each GPU unit according to the status information; In this embodiment, the collection of hardware parameters and dynamic operation indicators is the key to ensuring that the subsequent computing power characteristic matrix can accurately reflect the GPU load status.
[0027] Specifically, hardware parameters include the number of CUDA cores, video memory capacity, FP32 / FP64 peak computing power and TensorCore acceleration capabilities. This information is generally obtained through the driver interface or query instructions provided by the GPU; dynamic operation indicators include data such as current GPU utilization, video memory fragmentation rate, temperature and power consumption detection values, which can be read from the GPU operating environment through timed polling or event triggering. The reason for obtaining these hardware parameters and dynamic operation indicators is that different GPUs (such as A100, RTX series, domestic GPUs, etc.) have significant differences in computing power peak, video memory structure, and temperature and power consumption response. If scheduling is based only on a single parameter (such as GPU utilization), it is easy for high-priority tasks to preempt the wrong GPU or for domestic GPUs to be idle but not used. By comprehensively collecting hardware and dynamic load information, the computing power characteristic matrix generated subsequently can truly reflect the availability and potential bottlenecks of the GPU, providing more accurate basic information for multi-tenant queue scheduling.
[0028] S2.2: Generate a comprehensive computing power vector according to the hardware parameters and dynamic operation indicators; In this embodiment, the generation of the comprehensive computing power vector is to merge the hardware parameters collected in the previous step (such as the number of CUDA cores, Tensor Core acceleration capabilities, etc.) and dynamic operation indicators (GPU utilization, temperature, power consumption, memory fragmentation, etc.) into a unified data structure. When different GPUs face the same load, they may not be able to exert their theoretical peak computing power due to excessive temperature or memory fragmentation. If these dynamic factors are not considered, it will be difficult to perform accurate queue-jumping scheduling in multi-tenant scenarios.
[0029] Specifically, we first create an attribute mapping table for each GPU unit, converting hardware parameters into static computing power descriptions, such as peak computing power to characterize the computing power that can be provided under ideal conditions; then we introduce dynamic operating indicators into attenuation coefficients or weight factors to characterize possible performance degradation under the current load and ambient temperature. Through this process, the current GPU hardware limits and real-time working status can be integrated into a unified computing power indicator, thereby effectively reducing large-scale data writeback or computing delays caused by preemption decision errors.
[0030] S2.3: Construct a resource coordination matrix according to the physical topology of the GPU resource pool; In this embodiment, the resource collaboration matrix is used to characterize the topological relationship and data interaction cost between GPU units in the GPU resource pool. For multi-GPU servers or clusters, data is often transmitted between GPUs through PCIe, NVLink or internal high-speed interconnects with different bandwidth and latency characteristics; some GPUs may share caches or local NVLink channels in the same acceleration node, while others may need to communicate across nodes. When there is a need for multi-card collaboration or cross-card data dependency, simply treating all GPUs as "equivalent nodes" will lead to unnecessary communication losses or bandwidth bottlenecks. By describing the transmission bandwidth, communication latency or whether a specific cache module is shared between each pair of GPUs in the collaboration matrix, GPU combinations that are more friendly to data interaction can be given priority in subsequent queue-jumping scheduling, thereby improving the overall task execution efficiency and avoiding large-scale bandwidth occupation or communication congestion when low-priority tasks are frequently relocated to remote GPUs.
[0031] Specifically, a unique identifier is assigned to each GPU unit in the server or cluster, and the network interconnection information between it and other GPU units is collected through the underlying topology discovery module. For example, for GPUs with NVLink interconnection, query the number of NVLink interfaces, bandwidth limit, and whether there is a shared cache area; if the server uses PCIe interconnection or cross-node RDMA network, read the bandwidth and latency reference values corresponding to each PCIe channel or RDMA link, as well as whether there is a shared I / O hub. In this way, a set of connection tables that describe the interconnection status of all GPUs can be obtained.
[0032] Furthermore, the connection table is organized according to the two-dimensional structure of "GPU identifier × GPU identifier" to form an N×N initial adjacency information table, where each cell (i, j) corresponds to the type of available interconnection channel between GPU i and GPU j, the estimated bandwidth and communication delay, etc. If two GPUs are located in the same acceleration node and share a local NVLink channel, tags such as "shared cache" or "zero copy support" are added to the corresponding matrix cells to reflect the potential resource synergy advantages. If cross-node communication (such as through Infiniband or Ethernet) is detected, the specific bandwidth, delay and possible congestion points are recorded.
[0033] Furthermore, after completing the initial adjacency information aggregation, it is necessary to further weight or normalize the data interaction cost in combination with the task scenario. On the one hand, for continuous indicators such as bandwidth and latency, weights can be assigned to different numerical segments based on historical load test results to describe the difference from "high-speed interconnection" to "low-bandwidth remote"; on the other hand, for features such as shared cache or local NVLink, additional coordination coefficients can be set to reflect the significant reduction in transmission overhead when multiple cards collaborate in the same GPU node. Finally, these weighted or normalized indicators are integrated into the N×N matrix unit to obtain a resource coordination matrix that can quantify the "cooperation efficiency" or "remote communication cost" between each pair of GPUs.
[0034] As a preferred method, the resource coordination matrix can also be managed in layers or blocks according to different levels of multi-card coordination requirements: For common tasks that only require a single card to execute, only the self-loop of a single card in the matrix or the interconnection information between the card and the CPU node is focused on; for multi-card parallel tasks or distributed training scenarios, the focus is on analyzing NVLink, PCIe bandwidth, and cache sharing characteristics, and screening GPU pairs or GPU clusters with higher synergy coefficients in the matrix, thereby reducing distributed communication bottlenecks. In subsequent queue-jump scheduling, it is possible to quickly determine which GPUs are suitable for collaborative operation, which GPUs have potential congestion points or are too far apart, and avoid low-priority tasks from being randomly migrated to remote GPUs across nodes, thereby reducing bandwidth waste and latency spikes caused by queue-jumping, and further improving the flexibility and efficiency of multi-card resource allocation in a multi-tenant environment.
[0035] S2.4: Importing the comprehensive computing power vector into the resource coordination matrix to obtain a current computing power characteristic matrix; In this embodiment, in high-load or multi-tenant competition scenarios, simply emphasizing the peak performance of a certain GPU may not bring the best experience, especially when some tasks require multi-card parallelism or frequent data exchange. After matrixing, it can better understand which GPU nodes have higher collaborative efficiency.
[0036] Specifically, by combining the aforementioned comprehensive computing power vector (characterizing the available computing power, attenuation coefficient, and ecological compatibility information of a single GPU) with the resource collaboration matrix (characterizing the physical connection relationship and communication characteristics between GPUs), a more comprehensive current computing power characteristic matrix can be generated. This computing power characteristic matrix not only includes the static and dynamic computing power performance of each GPU, but also embeds the bandwidth, latency, and communication resource occupancy during cross-GPU collaboration, so as to more completely display the computing power distribution of the entire heterogeneous GPU resource pool. This enables the queue-jumping scheduling process to more accurately allocate appropriate GPUs, improve the response speed of sudden high-priority tasks, and take into account the fairness of multi-tenant resource use.
[0037] The specific steps of S2.2 are as follows: S2.2.1: Generate a static computing power vector according to the hardware parameters; In this embodiment, the hardware parameters are merged into a set of vectors according to the established field order, so that each parameter occupies a specific position in the vector and maintains a fixed dimension for subsequent computing power evaluation and comparison. Through this vectorization method, on the one hand, a unified parameter representation dimension can be maintained in the entire heterogeneous GPU resource environment, reducing the error when comparing between multi-model and multi-architecture GPUs; on the other hand, it can quickly provide a static reference value of "ideal performance" for subsequent steps, ensuring that each GPU can be accurately quantified at the performance baseline level, so that subsequent queue-jumping scheduling decisions can be evaluated under the same measurement system.
[0038] Specifically, first define a hardware description table with a fixed field order for each GPU unit to list its CUDA core number, video memory capacity, FP32 / FP64 peak computing power, and Tensor Core acceleration capabilities. Then, according to the preset order of "field 1-field 2-field 3-field 4...", convert these indicators into numerical values of the same dimension or unified reference, and arrange them in order into a set of vectors. For example, the first vector element position can be reserved for the "CUDA core number", the second position for the "video memory capacity", and so on, and the parameters of different scale levels are normalized accordingly. Ensure that during the vector construction process, the hardware parameters of different architectures (such as A100, RTX, and domestic GPUs) can be mapped to the same vector dimensions and order positions to avoid dimensional conflicts or deviations in subsequent comparisons or merges.
[0039] S2.2.2: Calculate the dynamic attenuation coefficient based on the dynamic operation index; In this embodiment, a dynamic attenuation coefficient can be calculated by performing weighted analysis and mapping on the dynamic operation indicators. The dynamic attenuation coefficient is used to characterize the difference between the available computing power of the GPU at the current moment and the static theoretical performance.
[0040] Specifically, if the temperature or power consumption is too high, the GPU may trigger automatic frequency reduction, voltage reduction and other protection mechanisms. If the fragmentation rate of video memory is too high, it may lead to insufficient actual available video memory or reduced allocation efficiency. Quantifying these unfavorable factors and incorporating them into the dynamic attenuation coefficient can make the subsequent computing power evaluation closer to the real-time working status of the GPU. By combining this coefficient with the static computing power vector, problems such as overload or efficiency reduction caused by scheduling based solely on ideal performance can be avoided, while allowing the queue-jumping scheduling process to take into account resource utilization and hardware safety redundancy.
[0041] Furthermore, first, a weight parameter is set for each dynamic operation indicator, such as utilization weight, power consumption weight, temperature weight, memory fragmentation weight, etc., and appropriate values are assigned to these weights according to business needs or actual evaluation data. Then, each dynamic operation indicator is analyzed and converted into a corresponding attenuation contribution value according to the following principles: If the current GPU utilization is high (for example, it is continuously higher than a certain threshold), it means that the running task has occupied a considerable part of the computing power. In this case, a higher "Utilization Decay Contribution Value" will be given to reflect the fact that the available computing power is partially consumed; If the temperature sensor data is close to or exceeds the safe temperature limit, it indicates that the GPU may trigger frequency reduction or cause heat dissipation pressure, making the actual available computing power lower than the static theoretical value, thus increasing the "temperature attenuation contribution value"; When the power consumption monitoring value is approaching the preset power consumption limit or heat dissipation bottleneck, the system will increase the "power consumption attenuation contribution value" accordingly to indicate that the GPU can no longer run at the peak frequency without limit; If the memory fragmentation rate is high, it means that the allocation efficiency and available memory space have decreased, which will increase the "memory fragmentation attenuation contribution value" and make additional reductions to the final available computing power.
[0042] Furthermore, after obtaining each of the above-mentioned attenuation contribution values, these contribution values will be multiplied by their corresponding weight parameters, and then they will be accumulated to obtain a "total attenuation score". In order to make the attenuation coefficient more intuitive and convenient for subsequent combination with the static computing power vector, a normalization process will be applied to the "total attenuation score" to keep the result within a reasonable range. Finally, this result is used as the dynamic attenuation coefficient, that is, if the "total attenuation score" is higher, the dynamic attenuation coefficient is closer to 0, indicating that the computing power that the GPU can provide at the current moment is significantly attenuated; if the "total attenuation score" is low, the dynamic attenuation coefficient is closer to 1, indicating that the GPU is in good condition and the static theoretical performance can be exerted to a large extent.
[0043] S2.2.3: Generate an eco-compatibility vector based on the type of GPU; In this embodiment, an ecological compatibility vector is generated based on the GPU type (such as NVIDIA GPU or domestic GPU) and its underlying driver ecology, AI framework support, compatibility of compiled libraries and other factors. The "ecological compatibility" here is not limited to the simple "whether it can run", but also includes the degree of acceleration optimization of specific frameworks or operators, stability test results at different precisions, whether it has software and hardware collaborative optimization features for certain tasks and other information. By adding these ecological compatibility factors in a vectorized form, certain GPUs can be given "specialized" evaluation weights when scheduling heterogeneous GPUs: for example, if the domestic GPU has stable performance and low energy consumption in some reasoning scenarios, a higher "reasoning adaptation score" is set for it when generating the ecological compatibility vector, so that this type of GPU is given priority when queuing scheduling. This not only fully utilizes the advantages of different GPU architectures, but also reduces the performance bottlenecks or incompatibility risks caused by blind scheduling in a multi-tenant environment.
[0044] Specifically, ecological compatibility can be divided into the following scoring dimensions: the first is the adaptability to mainstream AI frameworks (for example, whether TensorFlow and PyTorch are deeply optimized, and whether there are dedicated acceleration operators); the second is the stability of different precision modes (such as whether the throughput and accuracy during FP16 and INT8 reasoning are reliable); the third is the maintenance frequency and update compatibility of the compilation library and the underlying driver (including whether there are version conflicts, whether containerized deployment is supported, etc.); and the fourth is the special test performance for target task categories (such as reasoning, training, graphics rendering, etc.) (whether it has higher efficiency and lower error rate in the test set).
[0045] Furthermore, in terms of quantitative methods, a scoring range can be defined for each dimension, and different GPU types can be scored based on a large amount of offline test data or real operation statistics. After completing the scoring of each dimension, these scoring results are combined by weighted summation or vector splicing. The final ecological compatibility vector can not only intuitively reflect the optimization level of different GPUs in various mainstream frameworks and precision modes, but also provide a more detailed basis for software and hardware coordination for task queue-jumping decisions, ensuring that GPUs that are more suitable for current task requirements are prioritized in heterogeneous GPU scenarios, minimizing performance loss or potential failure risks caused by poor compatibility.
[0046] S2.2.4: Generate a comprehensive computing power vector based on the static computing power vector, the dynamic attenuation coefficient and the ecological compatibility vector; In this embodiment, the input parameters need to be preprocessed before the task information and the current computing power characteristic matrix are input into the neural network to ensure that the input data can accurately reflect the resource demand characteristics of the task at different computing stages and form a high degree of match with the available computing power of the GPU. A more sophisticated connection is established between the task computing process and the dynamic distribution of GPU resources, so that the scheduling system can adapt to complex heterogeneous computing environments more intelligently.
[0047] Traditional methods often regard the computing requirements of tasks as a static whole, and evaluate them only based on global computing power consumption, while ignoring the phased changes in computing power requirements during the execution of tasks. For example, in the data loading phase, the main bottleneck of the task may be video memory and PCIe bandwidth, while in the calculation execution phase, CUDA core utilization and tensor operation unit (TensorCore) load are the main influencing factors. In the result write-back phase, storage throughput and video memory recycling rate will become key parameters. If data that has not been processed in stages is directly input into the neural network, the model's prediction of the computing power requirements of the task will be too vague, which will affect the accuracy of the scheduling decision, and may even cause some GPUs to mismatch computing power resources due to insufficient consideration of their short-term bandwidth limitations.
[0048] Specifically, this application does not simply decompose the task computing requirements, but creatively improves the accuracy of task scheduling decisions by constructing a phased computing power vector and establishing a fusion mapping with the current state of the GPU. Especially in a heterogeneous GPU environment, different GPU architectures have significant differences in the degree of support for data loading, calculation execution, and result writing back. If staged processing is not performed, some GPUs may be in an inefficient state for a long time, and may even affect the throughput of the entire cluster due to the mismatch between task requirements and GPU characteristics. Through refined task feature extraction and computing power adaptation strategies, a more targeted resource scheduling solution is provided in a heterogeneous computing environment, thereby improving the intelligence level and resource utilization efficiency of the scheduling system.
[0049] The specific steps of preprocessing are as follows: The task information is divided into a data loading phase, a calculation execution phase, and a result writing back phase, and a phased computing power vector is generated according to the resource requirements of each phase; Specifically, the computing power requirements of a task are not constant throughout its life cycle, but show staged changes as the computing process progresses. For example, in deep learning tasks, the data loading stage mainly involves data preprocessing, memory allocation, and transmission from storage devices to GPU video memory, and is therefore limited by PCIe bandwidth, NVLink interconnection efficiency, and video memory throughput; the computational execution stage depends on the utilization of computing units such as CUDA cores and Tensor Cores and their peak computing power performance, while the result write-back stage mainly involves factors such as video memory release, data storage write-back, and inter-task data transmission bandwidth. Therefore, if the matching degree between tasks and GPU resources is directly measured by global computing power requirements, scheduling errors may occur due to different GPU architecture characteristics. For example, a GPU with high bandwidth but low computing power may be assigned to compute-intensive tasks, resulting in reduced resource utilization.
[0050] In this embodiment, three types of phased computing power vectors are established, and each type of vector corresponds to the key computing power requirements of the task at that stage, ensuring that subsequent scheduling decisions can more accurately evaluate the degree of adaptation of tasks to GPU resources at different stages and optimize the execution efficiency of tasks in a multi-tenant environment.
[0051] Furthermore, in the process of generating the phased computing power vector, the input data scale, calculation type, model structure and other information of the task are first analyzed to determine the memory allocation size, data transfer rate and cache reuse rate required in the data loading phase, and the information is quantified into the data loading vector. Then, the computing power requirements of the calculation execution phase are calculated based on the computational complexity, matrix operation requirements, parallel computing characteristics of the task, and the calculation execution vector is constructed based on the hardware acceleration capabilities of the GPU (such as the number of CUDA cores, Tensor Core support, etc.). Finally, in the result write-back phase, according to the amount of calculation result data that the task needs to store or return, the video memory release speed, storage I / O throughput and data consistency maintenance strategy are evaluated, and the result write-back vector is generated. The computing power requirements can be accurately described at each key stage of the task life cycle, avoiding the resource mismatch problem caused by a single computing power indicator in the traditional scheduling method, ensuring that the scheduling algorithm can fully consider the characteristics of different stages of the task when evaluating the matching degree between the task and the GPU resources, and improving the accuracy of scheduling decisions and resource utilization.
[0052] The phased computing power vector is merged with the current computing power characteristic matrix to obtain an initial fusion vector representing the adaptability of each phase to the current GPU idle resources; Specifically, the current computing power characteristic matrix is composed of information such as the GPU's hardware parameters, dynamic load status, and task execution history, which reflects the GPU's available computing power, load pressure, and resource competition at a specific point in time. However, the GPU's computing power performance is not constant, but is dynamically affected by factors such as temperature, power consumption, and task type. Therefore, it is difficult to accurately reflect its true execution capabilities based solely on static computing power indicators. In addition, the scheduling of GPU resources depends not only on computing power, but also on factors such as the degree of memory fragmentation, data transmission efficiency, and interconnection architecture. Therefore, by integrating the phased computing power vector with the current computing power characteristic matrix, the static performance, dynamic load, and phased requirements of the GPU resources can be considered simultaneously during the computing power evaluation process, so that queue-jumping scheduling can more accurately match the appropriate GPU resources.
[0053] Furthermore, when performing fusion, the current computing power characteristic matrix of the GPU is first normalized to ensure that the computing power parameters of different models of GPUs are compared on the same scale and to eliminate the deviation between different architectures. Then, the data loading vector is matched with the bandwidth resource indicators of the GPU (such as PCIe bandwidth, NVLink bandwidth, memory throughput, etc.), the calculation execution vector is matched with the core computing power resources of the GPU (such as the number of CUDA cores, Tensor Core utilization, etc.), and the result write-back vector is matched with resource indicators such as storage I / O throughput and memory release rate. Through this stage-by-stage matching process, three sets of independent fitness scores can be generated, and these scores are further weighted and summed to generate the final initial fusion vector. The fusion process can not only ensure that the task can find the most suitable GPU resources at different stages, but also dynamically adjust the resource allocation strategy during the scheduling process, so that the task can be executed on the GPU with lower load, avoiding resource bottlenecks on the high-load GPU and improving the overall scheduling efficiency.
[0054] Specifically, through the preprocessing method of this embodiment, the neural network model can be optimized based on more complete task information and GPU resource status during the training process, so that the model can more accurately predict the execution efficiency of the task and make more reasonable resource allocation based on computing power adaptability. Compared with the traditional solution based on static computing power matching or single priority scheduling, this method performs refined modeling at each stage of the task life cycle, and by integrating the GPU current computing power characteristic matrix, the scheduling process is more dynamically adaptable.
[0055] See also Figure 3 , a structural diagram of a neural network model according to an embodiment of the present invention, wherein the neural network model comprises: An input layer is used to extract features from the initial fusion vector through an attention mechanism to generate a deep representation vector that represents the coupling relationship between the GPU real-time load and the task phase requirements; An intermediate layer, used to process the depth representation vector through a multi-round iteration mechanism and output an intermediate result; The output layer is used to perform feasibility check on the intermediate results. If the check fails, the preliminary computing power matrix is locally adjusted according to the correction rules until the check passes. The preliminary computing power matrix that passes the check is output as the updated computing power characteristic matrix.
[0056] The specific steps of S3 are as follows: S3.1: Load the neural network model built based on the fusion of physical model and data-driven algorithm, and set the initial parameters of the network according to the current computing power characteristic matrix; Specifically, the loaded neural network model adopts a deep fusion architecture of physical models and data-driven algorithms. The physical model is built based on the GPU hardware characteristic equations, such as the computing power attenuation model of the CUDA core and the energy consumption formula of the domestic GPU instruction set conversion, to ensure that the model output conforms to the physical laws of the hardware. The data-driven part trains the deep temporal convolutional network (DTCN) based on historical scheduling data (such as task execution time, resource utilization, and error logs) to learn dynamic scheduling rules. When setting the initial parameters, the constraints of the physical model (such as the upper limit of video memory capacity and the temperature safety threshold) are encoded as network weights through the pre-training stage to ensure that the initial parameters are both in line with the laws of physics and have data-driven adaptability.
[0057] Furthermore, in order to make the initial parameters of the model closer to the actual operating status of the current GPU resource pool, when loading the network, the GPU available resource information in the current computing power characteristic matrix will be read first, including the GPU load, available video memory size, real-time power consumption, computing core utilization, etc., and this information will be used to initialize the input layer weights of the neural network, so that it can have a high adaptability to the current hardware environment at the beginning of the model operation.
[0058] As a preferred method, in order to improve the compatibility of the model with different GPU architectures, the standard computing power feature vectors of different GPU types will be loaded during the initialization phase, so that even if computing devices with significantly different architectures such as NVIDIA A100, RTX 4090 or domestic GPUs exist in the resource pool at the same time, the neural network can still be reasonably scheduled according to their respective characteristics. This can reduce scheduling errors during actual operation, improve the adaptability of tasks and GPU resources, and avoid task execution failures or performance waste caused by unreasonable computing power resource configuration, thereby realizing an intelligent GPU computing power allocation strategy.
[0059] S3.2: Input the initial fusion vector to the input layer of the neural network model, extract features of the initial fusion vector through the attention mechanism, and generate a deep representation vector representing the coupling relationship between the GPU real-time load and the task stage requirements; In this embodiment, since the available computing power of the GPU is affected by multiple factors such as load conditions, power consumption status, task concurrency, and the computing power requirements of the task itself also show phased characteristics, it is necessary to introduce an attention mechanism to extract features from the input data in order to generate a deep representation vector that can accurately characterize the relationship between the real-time load of the GPU and the computing power requirements of the task.
[0060] Specifically, after inputting the initial fusion vector, the neural network first normalizes the vector to ensure that the input data of different task types and different GPU models can be mapped to the same feature space to avoid training instability caused by differences in numerical scales. Subsequently, the attention mechanism automatically analyzes the degree of dependence of the current task on different GPU computing resources in different computing stages (such as data loading, calculation execution, and result writing back). For example, some tasks may be more dependent on video memory bandwidth, while other tasks may be mainly affected by the number of computing cores. In this process, the attention layer in the network calculates the weight distribution of different input features, so that the computing power features that have a greater impact on the current task are given higher weights, while the influence of features with less impact is reduced. This process can effectively avoid the resource mismatch problem that may be caused by the fixed weight allocation method used by traditional neural networks in the feature learning process, so that the final generated deep representation vector can more accurately reflect the coupling relationship between the GPU load and the task computing power requirements.
[0061] Furthermore, the initial fusion vector is first subjected to trilinear transformation to generate a query matrix (Query), a key matrix (Key), and a value matrix (Value). The query matrix is used to represent the computing power requirements of the current task, the key matrix is used to represent the computing power distribution in the current GPU resource pool, and the value matrix corresponds to the resources available for actual calculation. During the weight calculation process, different computing stages of the task will automatically learn different resource concerns. For example, in the data loading stage, the attention mechanism will give higher weights to resources such as PCIe bandwidth and NVLink interconnection, while in the calculation execution stage, it may pay more attention to computing characteristics such as CUDA core utilization and Tensor Core acceleration, while in the result write-back stage, it is more inclined to consider video memory management and storage throughput.
[0062] Furthermore, after obtaining the attention weight scores, different GPU computing power features are weighted and summed according to these weights to form a new task-computing power adaptation representation. The goal of this step is to make the neural network focus on the most critical computing power features while ignoring secondary features, thereby improving the matching degree between computing resources and task requirements. One attention head may focus mainly on the bandwidth requirements of the task, while another attention head focuses on the computing load. The final result is the weighted sum of multiple attention heads, making the final feature expression more comprehensive and avoiding information loss from a single perspective.
[0063] Furthermore, after all calculations are completed, the obtained weighted feature representation will be used as a deep representation vector and input into the subsequent layers of the neural network for further task scheduling optimization. The deep representation vector will serve as the core input for subsequent scheduling optimization, providing a higher-quality computing decision basis for the neural network model, ensuring that the scheduling strategy can be adaptively adjusted as the GPU status changes dynamically.
[0064] S3.3: Processing the depth representation vector through a multi-round iteration mechanism preset in the middle layer of the neural network model, outputting an intermediate result, wherein each iteration round corrects the output according to the dynamic constraints of the depth representation vector and the physical model, and adjusting the initial parameters of the network through back propagation according to the correction result until the number of iterations is reached; In this embodiment, in order to optimize the adaptability and robustness of the neural network in the computing power allocation decision, a multi-round iteration mechanism is adopted to gradually optimize the deep representation vector, and in each round of iteration, the scheduling scheme is corrected in combination with the dynamic constraints of the physical model. The core goal of the iteration mechanism is to enable the neural network to accurately predict the optimal GPU resource matching solution for the task under a dynamic load environment, while ensuring that the calculation results meet the actual GPU power consumption, security and task execution constraints.
[0065] Specifically, first, a multi-round cycle calculation module is preset in the middle layer of the neural network. The function of this module is to gradually optimize the output results of the neural network based on the input deep representation vector. In the first round of iterations, the network will preliminarily predict the GPU resources suitable for each task based on the computing requirements of the current task, the phased computing power requirements of the task, and the GPU computing power characteristic matrix, and store the prediction results in a temporary task allocation matrix. This task allocation matrix is used to store the number of tasks assigned to each GPU, the computing power consumption, and the computing load that the task may generate on the GPU. On this basis, the neural network will calculate the global load distribution under the current scheduling scheme, and generate a global computing power balance score based on this, which is used to measure the rationality of the current task scheduling scheme.
[0066] Furthermore, in each subsequent round of iterations, the neural network will modify the scheduling scheme based on the task allocation matrix of the previous round and the dynamic constraints in the physical model. Specifically, the physical model provides a set of key constraints, including the maximum power consumption limit of the GPU, the upper temperature limit, the memory occupancy rate, the task execution concurrency, etc. For example, if the scheduling result of the previous round causes the estimated power consumption of a GPU to exceed the safety threshold of the GPU, then in the next round of iterations, the network will automatically reduce the task load of the GPU, or adjust the tasks running on the GPU to allocate them to the GPU with lower computational load. At the same time, in order to ensure that the execution stability of other tasks will not be affected by the queue scheduling process, the physical model will also evaluate the task switching cost to avoid computing delays or excessive memory read and write overhead due to frequent task migration.
[0067] Furthermore, in each round of iteration, the neural network will not only adjust the resource allocation plan of the task, but also update the network parameters through the back-propagation algorithm to ensure that the model can gradually converge to the optimal scheduling strategy. Each round of iteration will calculate the loss function of the current scheduling plan, which comprehensively considers the following indicators: (1) the execution efficiency of the task, that is, whether the expected completion time of the task has been optimized; (2) the load balancing of the GPU, that is, whether the computing load of each GPU tends to be reasonably distributed; (3) resource utilization, that is, whether the utilization rate of computing resources is close to the optimal level; (4) the compliance with physical constraints, that is, whether the current scheduling plan violates the power consumption, temperature or video memory usage restrictions of the GPU. Based on these loss indicators, the back-propagation algorithm will calculate the gradient of each neuron and adjust the network parameters so that the output result of the next round of iteration is closer to the optimal computing scheduling strategy.
[0068] As a preferred method, in order to prevent the neural network from falling into a local optimal solution during multiple rounds of iterations, the network model will introduce a random perturbation mechanism after each round of iteration, so that the resource allocation scheme of some tasks is randomly adjusted within a certain range to explore possible better scheduling schemes. The random perturbation mechanism is similar to the simulated annealing algorithm, which can effectively prevent the neural network from converging to a suboptimal solution too early in the early iteration stage, but can search in more scheduling schemes and finally obtain the global optimal solution. At the same time, after each round of iteration, the current task allocation matrix will be compared with the matrix of the previous round of iteration, and the convergence index will be calculated. If the scheduling scheme of the current iteration has changed very little compared with the previous round, or the decrease of the loss function has fallen below the set threshold, it is determined that the current iteration has basically converged, and the iteration process is terminated in advance to improve the computational efficiency.
[0069] Specifically, in the entire iterative mechanism, each round of iteration not only optimizes the rationality of the scheduling scheme, but also continuously optimizes the parameters of the neural network, so that the model can better adapt to the dynamic changes of GPU resources. In practical applications, this method can effectively improve the intelligence of task scheduling, make task allocation more accurate, resource utilization more balanced, and ensure that task queue jumping in a multi-tenant environment will not affect the overall stability of the system. Finally, after multiple rounds of iterative optimization, the scheduling scheme output by the neural network not only has efficient computing resource utilization, but also complies with the physical constraints of the GPU, realizing a smarter and more refined GPU task scheduling strategy.
[0070] S3.4: Mapping the intermediate result to the dimensional structure of the current computing power characteristic matrix to obtain a mapped preliminary computing power matrix; In this embodiment, the intermediate results of the neural network need to be converted into an actual executable task scheduling plan, and the core of this conversion process is to map to the dimensional structure of the current computing power characteristic matrix to generate a mapped preliminary computing power matrix. The main goal of this mapping process is to ensure that the output of the neural network is not only the optimal allocation plan in theory, but also can adapt to the actual resource situation of the GPU and ensure the executableness of the task on the GPU. In order to achieve this goal, the entire mapping process is divided into several key steps, including task computing power demand conversion, GPU resource occupancy matrix construction, topology adaptation processing, constraint correction, etc., to ensure that the preliminary computing power matrix has actual execution capabilities.
[0071] Specifically, first, in the task computing power demand conversion stage, the intermediate result output by the neural network is often a vector containing task scheduling suggestions, including the computing resource categories required by the task (such as CUDA cores, Tensor Cores, video memory capacity, bandwidth, etc.) and the corresponding allocation ratios. These data need to be further converted into a standard format compatible with the GPU resource pool, so the GPU resource occupancy matrix is used for standardized representation in this embodiment. The GPU resource occupancy matrix is a multidimensional matrix, each row represents a task, and each column represents a GPU resource item (such as the number of computing cores, video memory occupancy, power consumption budget, etc.). When constructing the matrix, first search for the GPU resource items in the current computing power characteristic matrix item by item based on the computing power demand vector of the task, and fill the task demand into the corresponding column, for example: if the task requests 20GB of video memory, fill in 20GB in the "video memory occupancy" column of the matrix; if the task requires Tensor Core computing power, fill in the corresponding computing power demand value in the "Tensor Core Utilization" column. This conversion process ensures that the output of the neural network is consistent with the GPU resource characteristics, so that subsequent scheduling can be matched based on the actual hardware situation.
[0072] Furthermore, after the GPU resource occupancy matrix is constructed, topology adaptation processing is required to ensure that the task scheduling scheme can match the physical characteristics of different GPU architectures. Due to the different hardware architectures of heterogeneous GPU resources, such as the NVIDIA A100 series using NVLink high-speed interconnection, and some RTX series mainly rely on PCIe bandwidth, domestic GPUs may use self-developed communication protocols, so the interconnection topology between GPUs needs to be considered when allocating tasks to optimize data transmission efficiency. For example, if a task contains a large number of model parameter synchronization operations, and multiple computing nodes need to share data, it should be allocated to the GPU group that supports NVLink interconnection to reduce data transmission bottlenecks; for tasks that are computationally intensive but have less data interaction, they can be allocated to RTX GPUs or other independent computing units. This adaptation process is achieved by analyzing the GPU interconnection topology matrix, which stores the connection mode and bandwidth information between each GPU, and when calculating task allocation, by finding the optimal interconnection path, it ensures the efficiency of task data transmission and avoids computing bottlenecks caused by low interconnection rates between GPUs.
[0073] Furthermore, in the process of constraint correction, in order to ensure that the final generated preliminary computing power matrix is actually executable, it is also necessary to check and adjust the legality of task allocation, which mainly involves the following aspects: (1) Power budget constraint, that is, check whether the total power consumption of the current GPU exceeds the rated upper limit. For example, if the task allocation of a GPU causes the power consumption to exceed the safety threshold, it is necessary to adjust the resource occupancy of some tasks in the matrix to reduce the computing intensity or migrate to other GPUs; (2) The problem of video memory fragmentation, that is, to ensure that the video memory utilization rate is within a reasonable range after task allocation to avoid the phenomenon of task loading failure due to video memory fragmentation. For example, in the process of constructing the preliminary computing power matrix, if it is found that there is a large fragmentation in the video memory occupancy, the video memory application strategy of the task can be adjusted, such as giving priority to allocating continuous memory areas or appropriately delaying the execution of some tasks; (3) Fairness of tenant resource allocation. In a multi-tenant environment, it is necessary to check whether the tasks of each tenant comply with the preset quota strategy to prevent a single tenant from occupying high-performance GPU resources for a long time. If it is found that the resource allocation is unbalanced, the task priority in the computing power matrix is adjusted to make the computing power resource allocation between different tenants more in line with the established strategy.
[0074] S3.5: Perform a feasibility test on the preliminary computing power matrix. If the test fails, locally adjust the preliminary computing power matrix according to the correction rules until the test passes. In this embodiment, in order to ensure that the final task scheduling scheme can be actually executed, it is necessary to perform a feasibility test on the preliminary computing power matrix. The inspection process includes checking whether the GPU resource allocation exceeds the load limit, whether the task violates the tenant computing power quota, whether there is a resource competition conflict between tasks, etc. If a problem is found in the preliminary computing power matrix, the matrix will be adjusted according to the correction rules, such as reducing the computing resource allocation of some tasks, rescheduling to a GPU with a lower load, or adjusting the task priority to optimize the overall scheduling strategy. The introduction of this step can effectively avoid GPU overload or task execution failure caused by scheduling errors, and improve the stability and execution efficiency of the task allocation scheme.
[0075] S3.6: Use the verified preliminary computing power matrix as the updated computing power characteristic matrix.
[0076] The specific steps of S4 are as follows: S4.1: Counting a GPU resource list of idle GPU resources according to the GPUs matching the tasks to be scheduled in the updated computing power characteristic matrix; In this embodiment, in order to ensure that the scheduling decision can accurately match the real-time status of the current GPU resource pool, before the task is assigned, the currently available GPU resource list must be counted based on the updated computing power characteristic matrix.
[0077] Specifically, the system will first read and update the computing power parameters of each GPU in the computing power characteristic matrix, including the available ratio of computing cores (CUDA Core, Tensor Core, etc.), remaining video memory space, data throughput capacity, and parallel task occupancy, and filter based on task requirements. For example, if the task to be scheduled is a deep learning training task with high video memory requirements, the system will give priority to GPUs with sufficient remaining video memory space and strong computing power to avoid reduced task running efficiency due to insufficient video memory or limited computing power.
[0078] As a preference, the physical topology of the GPU also needs to be considered during the statistical process, such as NVLink or PCIe bandwidth limitations between GPUs, to ensure that the task can achieve the best data transmission efficiency in the distributed computing mode. By building a complete list of allocable GPU resources, the scheduling algorithm can have a more comprehensive calculation basis for subsequent resource allocation and queue-jumping decisions, thereby maximizing the utilization of GPU resources and reducing resource mismatch problems in task scheduling.
[0079] S4.2: Determine whether it is necessary to perform queue-jump scheduling or preemption operation on the existing running task. If not, directly allocate idle GPU resources to the task to be scheduled; In this embodiment, based on the statistically available GPU resource list, it is necessary to further determine whether it is necessary to execute queue-jump scheduling or preempt the resources of existing tasks. This judgment process mainly depends on multiple factors, including the priority of the task to be scheduled, the real-time load of the GPU resources, the distribution of computing requirements in the current task queue, and the computing power quota strategy of the tenant.
[0080] Specifically, if there is an idle GPU in the list of available GPU resources that meets the needs of the current task, there is no need to perform queue-jumping scheduling. The most suitable GPU can be directly assigned to the task and resources can be bound. However, in a large-scale concurrent task environment, idle GPU resources may be extremely limited, so it is necessary to further evaluate whether there are low-priority tasks occupying high-performance GPUs to determine whether preemption operations are needed. In order to avoid excessive task interruption rates due to frequent preemptions, the system will also combine historical task scheduling data to predict the calculation time of the target task and its impact on other tasks when making preemption decisions. For example, if a low-priority task still needs to run for a long time, and a high-priority task needs to be executed immediately, the system will tend to perform preemptive scheduling to ensure that the high-priority task can obtain resources first, while minimizing the impact on low-priority tasks. Through this judgment mechanism, the system can balance the urgency of the task and the availability of resources to ensure the fairness and efficient use of computing resources.
[0081] S4.3: If yes, re-arrange resources for the preempted or reallocated tasks according to the updated computing power characteristic matrix; In this embodiment, when it is determined that queue-jumping scheduling is required, it is necessary to reasonably rearrange the resources of the preempted tasks according to the updated computing power characteristic matrix to ensure that the preempted tasks will not fall into a state of complete interruption or execution failure due to resource recovery. The core of resource rearrangement is to find suitable alternative GPU resources for the preempted tasks and ensure that their computing environment is migrated as stably as possible.
[0082] Specifically, the process first checks the computing status of the preempted task, including the elapsed execution time, computing stage, memory usage, and data loading progress, to determine whether it can be safely migrated to other GPUs. If the execution progress of the preempted task is short, or its computing status allows recovery, the system will select a GPU with a lower load to migrate the task to a new computing node and synchronize its computing environment, such as using container snapshot technology (such as Checkpoint / Restore in Userspace, CRIU) to save the computing context of the task and restore the computing status on the new GPU to ensure the smoothness of the migration process. If the preempted task cannot be completely migrated, a phased preemption strategy may be adopted, that is, allowing the task to pause after executing to a certain safety checkpoint to avoid wasting computing resources due to mid-process preemption. Through such a resource rescheduling mechanism, the system can minimize the impact of preemptive scheduling on low-priority tasks, while ensuring that high-priority tasks can obtain computing resources in a timely manner, optimizing the overall GPU computing scheduling efficiency.
[0083] S4.4: Allocate idle GPU resources according to the computing requirements of the task to be scheduled, and return the final allocation plan and queue-jumping scheduling result to the scheduling management port, completing the running binding of the task to be scheduled on the GPU resources; In this embodiment, after completing the queue-jumping scheduling or directly allocating resources, the final GPU resource allocation plan needs to be confirmed, and the scheduling result is returned to the scheduling management port to complete the GPU resource binding of the task to be scheduled. The core of this step is to ensure that the scheduling of tasks on the GPU can be executed stably, and the scheduling information can be synchronized to the management system for subsequent monitoring and adjustment.
[0084] Specifically, during the resource allocation process, the system will bind appropriate computing nodes to tasks based on the task requirement information provided by the updated computing power feature matrix, and allocate key resources such as computing cores, video memory space, and data transmission bandwidth. For example, if the calculation phase of a task requires a large number of matrix operations, the system will give priority to allocating GPUs with Tensor Core acceleration capabilities to improve computing efficiency. At the same time, when binding tasks, the system will also configure appropriate execution strategies for the tasks, such as whether to allow dynamic adjustment of computing power resources and whether to allow tasks to pre-occupy more video memory space in order to maintain computing stability under high load conditions.
[0085] Preferably, the scheduling results will be synchronized to the scheduling management port, including the computing nodes of the task, GPU resource allocation, execution priority, preemption impact and other information, so as to dynamically adjust the computing resources in the subsequent task scheduling process. For example, if the computing time of a task is extended due to sudden queue-jumping scheduling, the scheduling system can automatically adjust the priority of the task to compensate for its computing time, thereby ensuring the fairness of task execution. Through this step, the system can not only complete the GPU resource binding of the task, but also ensure the transparency and traceability of task scheduling, improve the intelligence of computing power scheduling, and optimize the computing resource utilization in a multi-user environment.
[0086] Embodiment 2: See also Figure 4 ,The present invention provides an embodiment: a multi-user computing power quota intelligent queue-jumping scheduling system, the system comprising an information collection module, an information processing module and a queue-jumping scheduling module; The information collection module is used to obtain task information of tasks to be scheduled and status information of heterogeneous GPUs; The information processing module is used to determine the current computing power characteristic matrix according to the state information, and input the task information and the current computing power characteristic matrix into a preset neural network model as input to obtain an updated computing power characteristic matrix; The queue-jumping scheduling module is used to allocate GPU resources to the tasks to be scheduled according to the updated computing power characteristic matrix.
[0087] The information processing module comprises: An information preprocessing unit, used to determine the current computing power characteristic matrix; The model processing unit is used to update the computing power characteristic matrix through the attention mechanism and multiple rounds of iterative correction.
[0088] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. A multi-user computing power quota intelligent queue-jumping scheduling method, applied to a heterogeneous GPU resource environment, characterized in that: The intelligent queue-jumping scheduling method comprises: Obtain task information of tasks to be scheduled and status information of heterogeneous GPUs; Determine a current computing power characteristic matrix according to the state information; The task information and the current computing power characteristic matrix are input into a preset neural network model, and an updated computing power characteristic matrix is obtained through an attention mechanism and multiple rounds of iterative correction; According to the updated computing power characteristic matrix, GPU resources are allocated to the task to be scheduled.
2. According to claim 1, the multi-user computing power quota intelligent queue-jumping scheduling method is characterized in that: The determining of the current computing power characteristic matrix includes: Acquire hardware parameters and dynamic operation indicators of each GPU unit according to the status information; Generate a comprehensive computing power vector based on the hardware parameters and dynamic operation indicators; Build a resource coordination matrix based on the physical topology of the GPU resource pool; The comprehensive computing power vector is imported into the resource coordination matrix to obtain the current computing power characteristic matrix.
3. According to claim 2, the multi-user computing power quota intelligent queue-jumping scheduling method is characterized in that: According to the hardware parameters and dynamic operation indicators, a comprehensive computing power vector is generated, including: Generate a static computing power vector according to the hardware parameters; Calculating a dynamic attenuation coefficient according to the dynamic operation index; Generate an eco-compatibility vector based on the type of GPU; A comprehensive computing power vector is generated according to the static computing power vector, the dynamic attenuation coefficient and the ecological compatibility vector.
4. According to claim 1, the multi-user computing power quota intelligent queue-jumping scheduling method is characterized in that: The neural network model is established based on physical model and data driven technology.
5. According to claim 4, the multi-user computing power quota intelligent queue-jumping scheduling method is characterized in that: Taking the task information and the current computing power characteristic matrix as input, including: The task information is divided into a data loading phase, a calculation execution phase, and a result writing back phase, and a phased computing power vector is generated according to the resource requirements of each phase; The phased computing power vector is fused with the current computing power characteristic matrix to obtain an initial fusion vector that characterizes the adaptability of each phase to the current GPU idle resources.
6. The multi-user computing power quota intelligent queue-jumping scheduling method according to claim 5 is characterized in that: The input to the preset neural network model includes: Load the neural network model built based on the fusion of physical model and data-driven algorithm, and set the initial parameters of the network according to the current computing power characteristic matrix; Inputting the initial fusion vector into the input layer of the neural network model, extracting features of the initial fusion vector through an attention mechanism, and generating a deep representation vector representing the coupling relationship between the real-time load of the GPU and the phased requirements of the task; The depth representation vector is processed by a multi-round iteration mechanism preset in the middle layer of the neural network model, and an intermediate result is output, wherein each iteration round corrects the output according to the dynamic constraints of the depth representation vector and the physical model, and the initial parameters of the network are adjusted by back propagation according to the correction result until the number of iterations is reached.
7. The multi-user computing power quota intelligent queue-jumping scheduling method according to claim 6 is characterized in that: The updating of the computing power characteristic matrix includes: Mapping the intermediate result to the dimensional structure of the current computing power characteristic matrix to obtain a mapped preliminary computing power matrix; Conducting a feasibility test on the preliminary computing power matrix; If the test fails, the preliminary computing power matrix is partially adjusted according to the correction rules until the test passes; The initial computing power matrix that passes the inspection is used as the updated computing power characteristic matrix.
8. The multi-user computing power quota intelligent queue-jumping scheduling method according to claim 1 is characterized in that: Allocating GPU resources to the task to be scheduled according to the updated computing power characteristic matrix includes: According to the GPUs matching the tasks to be scheduled in the updated computing power characteristic matrix, a GPU resource list of idle GPU resources is counted; Determine whether to perform queue-jump scheduling or preemption operation on the existing running task, if not, directly allocate idle GPU resources to the task to be scheduled; If yes, re-arrange resources for the preempted or reallocated tasks according to the updated computing power characteristic matrix; The idle GPU resources are allocated according to the computing requirements of the tasks to be scheduled, and the final allocation plan and queue-jumping scheduling results are returned to the scheduling management port to complete the running binding of the tasks to be scheduled on the GPU resources.
9. A multi-user computing power quota intelligent queue-jumping scheduling system, used to implement the multi-user computing power quota intelligent queue-jumping scheduling method according to any one of claims 1 to 8, characterized in that: The system includes an information collection module, an information processing module and a queue-jumping scheduling module; The information collection module is used to obtain task information of tasks to be scheduled and status information of heterogeneous GPUs; The information processing module is used to determine the current computing power characteristic matrix according to the state information, and input the task information and the current computing power characteristic matrix into a preset neural network model as input to obtain an updated computing power characteristic matrix; The queue-jumping scheduling module is used to allocate GPU resources to the tasks to be scheduled according to the updated computing power characteristic matrix.
10. The multi-user computing power quota intelligent queue-jumping scheduling system according to claim 9, characterized in that: The information processing module comprises: An information preprocessing unit, used to determine the current computing power characteristic matrix; The model processing unit is used to update the computing power characteristic matrix through the attention mechanism and multiple rounds of iterative correction.
Citation Information
Patent Citations
Calculation power resource scheduling method, and training method and system of calculation power resource scheduling model
CN116643877A
Heterogeneous edge computing power network task scheduling method
CN115185650A
Task processing method and device of computing power network, electronic equipment and storage medium
CN117675823A
Large model reasoning scheduling method based on off-network computing power server
CN119537032A
Computing system and method for GPU (Graphics Processing Unit) computing power scheduling
CN119645661A
Cited By
Computing power allocation optimization method and system for multiple data processing tasks
CN120578510A
Heterogeneous computing power normalization metering method and device for intelligent computing center cloud platform
CN121255573A
Automatic aggregation scheduling method for improving GPU (Graphics Processing Unit) computing power utilization rate
CN121501524A
Self-adaptive distributed training optimization method
CN121615720A
An adaptive distributed training optimization method
CN121615720B