Multi-User Computing Power Quota Intelligent Queue-Jumping Scheduling Method and System
The method employs a neural network model to dynamically allocate GPU resources across heterogeneous environments, addressing inefficiencies in existing scheduling methods by optimizing resource utilization and task execution efficiency.
Patent Information
- Application Number
- CN202510512536.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing GPU resource scheduling methods are difficult to dynamically adjust computing power allocation in a multi-user environment, and cannot accurately match the task phased needs. The heterogeneous GPU is insufficient adaptability, resulting in low resource utilization and low task execution efficiency.
By obtaining the task information to be scheduled and GPU status information, the current computing power characteristic matrix is constructed, and the neural network model that integrates physical models and data-driven algorithms is used to perform feature extraction and multiple iterative corrections, dynamically evaluate GPU resource allocation, and queue scheduling and resource rearrangement.
It improves the utilization rate of computing power resources and task execution efficiency, ensures fairness and adaptability in a multi-tenant environment, and optimizes the intelligent matching and dynamic queue jumping scheduling of heterogeneous GPU resources.
Smart Images

Figure CN120029744B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computing power resource scheduling, and particularly to a multi-user computing power quota intelligent cut-in scheduling method and system. Background Art
[0002] Existing GPU resource scheduling methods usually allocate resources based on fixed quotas or simple priority rules, making it difficult to fully meet the heterogeneous GPU resource management requirements in a multi-user environment. In scenarios such as high-performance computing, deep learning training, and inference, different tasks have significant differences in resource requirements such as the computing power, video memory, and bandwidth of the GPU. Traditional scheduling strategies often cannot dynamically adjust the computing power allocation, resulting in high-performance GPU resources being occupied by low-demand tasks, while high-priority tasks may be delayed due to insufficient resources. In addition, when existing methods perform task preemption and cut-in scheduling, they lack in-depth analysis of the current computing power state of the GPU and the phased requirements of tasks, resulting in high task migration costs, low resource utilization, and even possible impacts on the overall throughput of the system. At the same time, the architectural differences between heterogeneous GPU environments (such as NVIDIA A100, RTX 4090, and domestic GPUs) increase the difficulty of task adaptation, and existing scheduling systems often cannot effectively distinguish and reasonably allocate the computing power resources of different architectures. Therefore, there is an urgent need for a computing power management method that can dynamically optimize computing power allocation, support intelligent cut-in scheduling, and is applicable to heterogeneous GPU environments to improve the utilization rate of computing resources and task execution efficiency.
[0003] For example, the computing power resource scheduling method, the training method and system of the computing power resource scheduling model provided by the Chinese patent application with the publication number of CN116643877A, after obtaining the state data of the local device in the current state, input the state data into the computing power resource scheduling model to obtain the scheduling data in the current state, where the scheduling data includes the operation data of the target scheduling operation, and, based on the operation data, execute the target scheduling operation to perform computing power resource scheduling on the local device; this solution can improve the scheduling efficiency of computing power resources.
[0004] The above existing technologies all have the problems proposed in this background art: there are problems in the GPU resource scheduling in a multi-user environment, such as inflexible static quota allocation, inability to accurately match the phased computing power requirements of tasks, insufficient adaptability to heterogeneous GPUs, and lack of intelligent optimization in the cut-in scheduling strategy. To solve the above problems, this application designs a multi-user computing power quota intelligent cut-in scheduling method and system. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a multi-user computing power quota intelligent queue-jumping scheduling method and system in view of the deficiencies of the prior art. First, the task information to be scheduled and the GPU status information are obtained, and the current computing power characteristic matrix is constructed based on the hardware parameters, dynamic operation metrics, and physical topology. Subsequently, a neural network model that combines a physical model and a data-driven algorithm is used, and feature extraction is performed in combination with an attention mechanism, and an updated computing power characteristic matrix is generated through multiple rounds of iterative correction. Based on this matrix, the GPU resource allocation situation is dynamically evaluated, whether to perform queue-jumping scheduling is judged, and resource rearrangement is performed on the preempted tasks to optimize the overall computing power utilization rate. This application effectively improves the computing power allocation efficiency and fairness in a multi-tenant environment and is applicable to complex computing task scenarios such as high-performance computing, deep learning training, and inference.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A multi-user computing power quota intelligent queue-jumping scheduling method is applied to a heterogeneous GPU resource environment. The intelligent queue-jumping scheduling method includes:
[0008] Obtain the task information of the task to be scheduled and the status information of the heterogeneous GPU;
[0009] Determine the current computing power characteristic matrix according to the status information;
[0010] Take the task information and the current computing power characteristic matrix as inputs and input them into a preset neural network model, and obtain an updated computing power characteristic matrix through an attention mechanism and multiple rounds of iterative correction;
[0011] Allocate GPU resources for the task to be scheduled according to the updated computing power characteristic matrix.
[0012] The determination of the current computing power characteristic matrix includes:
[0013] Obtain the hardware parameters and dynamic operation metrics of each GPU unit according to the status information;
[0014] Generate a comprehensive computing power vector according to the hardware parameters and dynamic operation metrics;
[0015] Construct a resource collaboration matrix according to the physical topology structure of the GPU resource pool;
[0016] Import the comprehensive computing power vector into the resource collaboration matrix to obtain the current computing power characteristic matrix.
[0017] Generating a comprehensive computing power vector according to the hardware parameters and dynamic operation metrics includes:
[0018] Generate a static computing power vector according to the hardware parameters;
[0019] Calculate the dynamic attenuation coefficient according to the dynamic operation index;
[0020] Generate an ecological compatibility vector according to the type of GPU;
[0021] Generate a comprehensive computing power vector according to the static computing power vector, dynamic attenuation coefficient and ecological compatibility vector.
[0022] The neural network model is established based on a physical model and data-driven technology.
[0023] Taking the task information and the current computing power characteristic matrix as inputs, includes:
[0024] Divide the task information into a data loading stage, a computing execution stage and a result writing-back stage, and generate stage-based computing power vectors for the resource requirements of each stage respectively;
[0025] Fuse the stage-based computing power vectors with the current computing power characteristic matrix to obtain an initial fusion vector representing the adaptation degree of each stage to the current GPU idle resources.
[0026] The input to a preset neural network model, includes:
[0027] Load a neural network model constructed by fusing a physical model and a data-driven algorithm, and set the network initial parameters according to the current computing power characteristic matrix;
[0028] Input the initial fusion vector into the input layer of the neural network model, and perform feature extraction on the initial fusion vector through an attention mechanism to generate a deep representation vector representing the coupling relationship between the real-time load of the GPU and the stage-based requirements of the task;
[0029] Process the deep representation vector through a multi-round iteration mechanism preset in the middle layer of the neural network model, and output an intermediate result, where each iteration round corrects the output according to the dynamic constraints of the deep representation vector and the physical model, and adjusts the network initial parameters through backpropagation according to the correction result until the number of iterations is reached.
[0030] The obtaining of the updated computing power characteristic matrix, includes:
[0031] Map the intermediate result to the dimension structure of the current computing power characteristic matrix to obtain a preliminarily mapped computing power matrix;
[0032] Conduct a feasibility test on the preliminarily mapped computing power matrix;
[0033] If the test fails, make a local adjustment to the preliminarily mapped computing power matrix according to the correction rule until the test passes;
[0034] Use the initially verified computing power matrix as the updated computing power characteristic matrix.
[0035] According to the updated computing power characteristic matrix, allocate GPU resources for the task to be scheduled, including:
[0036] Based on the GPUs in the updated computing power characteristic matrix that match the task to be scheduled, count the GPU resource list of idle GPU resources;
[0037] Determine whether to perform cut-in scheduling or preemption operations on existing running tasks. If not, directly allocate idle GPU resources for the task to be scheduled;
[0038] If so, rearrange the resources of the preempted or reallocated tasks according to the updated computing power characteristic matrix;
[0039] Allocate the idle GPU resources according to the computing requirements of the task to be scheduled, and return the final allocation plan and the cut-in scheduling result to the scheduling management port to complete the running binding of the task to be scheduled on the GPU resources.
[0040] Multi-user computing power quota intelligent cut-in scheduling system, the system includes an information collection module, an information processing module, and a cut-in scheduling module;
[0041] The information collection module is used to obtain the task information of the task to be scheduled and the status information of heterogeneous GPUs;
[0042] The information processing module is used to determine the current computing power characteristic matrix according to the status information, and use the task information and the current computing power characteristic matrix as inputs to a preset neural network model to obtain an updated computing power characteristic matrix;
[0043] The cut-in scheduling module is used to allocate GPU resources for the task to be scheduled according to the updated computing power characteristic matrix.
[0044] The information processing module includes:
[0045] An information preprocessing unit for determining the current computing power characteristic matrix;
[0046] A model processing unit for obtaining an updated computing power characteristic matrix through an attention mechanism and multi-round iterative correction.
[0047] Compared with the prior art, the beneficial effects of the present invention are:
[0048] The present invention realizes intelligent matching and dynamic queue-jumping scheduling of heterogeneous GPU resources by constructing a neural network model that integrates a physical model and a data-driven algorithm, combined with task-stage computing power demand analysis and attention mechanism optimization, which can effectively improve the utilization rate of computing power resources and task execution efficiency. Through multiple rounds of iterative correction and resource rearrangement mechanisms, the accuracy of task scheduling and GPU load balancing are ensured, the task preemption cost is reduced, and the fairness and adaptability in a multi-tenant environment are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments when read in conjunction with the accompanying drawings:
[0050] Figure 1 It is a flowchart of the intelligent queue-jumping scheduling method for multi-user computing power quotas in Embodiment 1 of the present invention;
[0051] Figure 2 It is a flowchart for calculating the current computing power characteristic matrix in Embodiment 1 of the present invention;
[0052] Figure 3 It is a structural diagram of the neural network model in Embodiment 1 of the present invention;
[0053] Figure 4 It is a module diagram of the intelligent queue-jumping scheduling system for multi-user computing power quotas in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0055] Embodiment 1:
[0056] Please refer to Figure 1 , an embodiment provided by the present invention: an intelligent queue-jumping scheduling method for multi-user computing power quotas, which is applied to a heterogeneous GPU resource environment. For multi-type tasks such as high-performance computing and AI training in a multi-tenant environment, in order to make full use of heterogeneous resources such as A100, RTX series, and domestic GPUs, and to avoid the situation that high-priority tasks cannot be executed in time due to fixed quotas, the following steps S1 to S4 are proposed. Through this embodiment, the problems of lack of coordination between multi-stage task requirements and GPU hardware limitations in large-scale GPU scheduling and difficulty in taking into account multi-tenant fairness during scheduling can be solved, thereby significantly improving the overall computing power utilization rate and ensuring low-latency response for critical tasks. The specific steps of the method are as follows:
[0057] S1: Obtain the task information of the task to be scheduled and the status information of the heterogeneous GPU;
[0058] In this embodiment, the dispatching platform collects the attributes of the tasks to be dispatched (such as task type, priority, video memory requirement, historical execution duration) and the real-time operation metrics on the GPU side (such as utilization rate, temperature, power consumption, video memory occupancy distribution), and records the hardware parameters of the GPU (such as the number of CUDA cores, FP32 / FP64 peak computing power, Tensor Core acceleration ability) to build a computing power characteristic matrix later. This provides an accurate task-resource mapping basis for subsequent cut-in dispatching, avoiding blind queuing or resource waste caused by missing information.
[0059] S2: Determine the current computing power characteristic matrix according to the status information;
[0060] In this embodiment, by integrating the GPU hardware parameters and dynamic operation metrics obtained in the previous step, a comprehensive computing power vector is first generated, and then a resource collaboration matrix is constructed in combination with the physical topology structure to form the current computing power characteristic matrix. Quantifying multi-dimensional information such as the core utilization rate, video memory capacity, and power consumption constraint of the GPU in matrix form facilitates direct invocation by subsequent neural network models and enables accurate cut-in decisions.
[0061] S3: Input the task information and the current computing power characteristic matrix into a preset neural network model to obtain an updated computing power characteristic matrix;
[0062] In this embodiment, the task information to be dispatched and the current computing power characteristic matrix are input into a neural network model that combines a physical model and a data-driven algorithm; this model extracts key resource features through an attention mechanism and continuously corrects the output results in multiple rounds of iteration, finally obtaining an updated computing power characteristic matrix. It can dynamically balance the requirements of high-priority tasks, GPU hardware safety redundancy, and multi-tenant quota restrictions, enabling cut-in dispatching to maximize computing power utilization while ensuring system stability.
[0063] S4: Allocate GPU resources for the tasks to be dispatched according to the updated computing power characteristic matrix;
[0064] In this embodiment, according to the resource matching degree and priority judgment output in the updated computing power characteristic matrix, the available GPU resources are searched and allocated. When it is confirmed that cut-in dispatching is required, necessary resource rearrangement is performed on the currently executing tasks with low priority or low utilization rate, and high-priority tasks are preferentially bound to the most suitable GPU; if cut-in is not required, idle GPU resources are directly allocated. This effectively solves problems such as resources being occupied for a long time due to traditional fixed quotas and sudden high-priority tasks being unable to be executed in a timely manner, and ensures fairness and efficiency in a multi-tenant environment.
[0065] Specifically, most existing solutions mainly perform resource matching based on a single or a small number of scheduling dimensions (such as priority, queuing order), lacking a refined assessment of the physical characteristics of GPUs (such as the number of Tensor Cores, video memory capacity, and fragmentation degree), and it is also difficult to flexibly adapt when there are software stack differences between domestic GPUs and NVIDIA GPUs. Once a high-priority task suddenly appears in the system, its scheduling decision often can only rely on static rules to preempt low-priority tasks, and cannot make flexible selections according to real-time power consumption, load, and thermal safety redundancy. Even if some existing solutions have built-in some dynamic monitoring mechanisms and can make preemption judgments based on GPU utilization or temperature conditions, for the load patterns of large-scale multi-tenant environments and multi-stage tasks (such as data loading, computing execution, result writing back), it is difficult to break down the computing power bottlenecks corresponding to each stage. Therefore, when there is a sudden queue jump, situations such as "too high preemption cost" or "the resource pool is still idle" often occur, resulting in a decrease in overall efficiency or long-term idleness of some GPUs. In addition, due to differences in GPU architectures, drivers, and software frameworks (such as TensorFlow, PyTorch, etc.) at home and abroad, it is usually necessary to manually or based on coarse-grained rules to decide whether to deploy tasks to domestic GPUs or NVIDIA GPUs. Existing scheduling often ignores deep dependencies such as local compilation support and differences in mixed-precision operation performance, resulting in low utilization of domestic GPUs, or some tasks running on inappropriate GPUs and dragging down performance.
[0066] In this embodiment, a computing power characteristic matrix is introduced as the core intermediate data structure, enabling information in multiple dimensions (hardware parameters, dynamic operation metrics, physical topology relationships) to be quantitatively described in a vectorized form and matched with the multi-stage requirements of the tasks to be scheduled. For example, the A100 has higher Tensor Core acceleration performance in the deep learning training scenario, while the RTX series has an advantage in general computing or graphics rendering. Domestic GPUs may have better energy efficiency ratios in some inference tasks. With the computing power characteristic matrix, it is possible to clearly identify which type of GPU is most suitable for the computing power requirements of a certain stage during the multi-round iterative correction process of the neural network model. Compared with simply relying on priority queues or FIFO strategies, the computing power characteristic matrix not only considers "who arrives first, who has priority", but also explicitly measures the impact of queue jumping on the overall system efficiency and multi-tenant fairness. If the model predicts that preempting a low-priority task on an RTX card will only bring limited performance improvement but may cause large-scale data writing back or a sharp increase in power consumption, then through the iterative update of the computing power characteristic matrix, the scheduling decision can be corrected in a timely manner to avoid unnecessary large-scale resource migration. If the domestic GPU is idle at this time and is compatible with the demand ecosystem of the high-priority task, the matrix update result will preferentially recommend execution on the domestic GPU, improving the overall utilization rate of the heterogeneous GPU pool.
[0067] Furthermore, by adopting a neural network model to fuse the physical model and data-driven algorithms, the GPU quota allocation, power consumption safety redundancy, and multi-tenant queue-jumping strategies can be continuously corrected in multiple rounds of iteration. This not only solves the problem that simple priority scheduling is difficult to dynamically adapt to heterogeneous GPU load changes but also effectively avoids the "hard switching" between sudden high-priority tasks and ongoing tasks, resulting in serious losses. With the adaptive update of the computing power characteristic matrix during the heterogeneous GPU scheduling process, not only more refined multi-dimensional computing power evaluation and allocation are achieved, but also low-latency execution of critical tasks can be ensured in a multi-user environment, improving the overall resource utilization rate and fairness.
[0068] Please refer to Figure 2 , the flowchart of calculating the current computing power characteristic matrix in the embodiment of the present invention. The specific steps of S2 are as follows:
[0069] S2.1: Obtain the hardware parameters and dynamic operation metrics of each GPU unit according to the status information;
[0070] In this embodiment, the collection of hardware parameters and dynamic operation metrics is the key to ensuring that the subsequent computing power characteristic matrix can accurately reflect the GPU load status.
[0071] Specifically, the hardware parameters include the number of CUDA cores, video memory capacity, FP32 / FP64 peak computing power, and TensorCore acceleration capabilities, which are generally obtained through the driver interface or query instructions provided by the GPU; the dynamic operation metrics include data such as the current GPU utilization rate, video memory fragmentation rate, temperature, and power consumption detection values, which can be read from the GPU operating environment by means of regular polling or event triggering. The reason for obtaining these hardware parameters and dynamic operation metrics is that different GPUs (such as A100, RTX series, domestic GPUs, etc.) have significant differences in computing power peaks, video memory structures, and temperature and power consumption responses. If scheduling is based only on a single parameter (such as GPU utilization rate), it is easy to have situations where high-priority tasks preempt the wrong GPU or domestic GPUs are idle but not used. By comprehensively collecting hardware and dynamic load information, the subsequent generated computing power characteristic matrix can truly reflect the availability and potential bottlenecks of the GPU, providing more accurate basic information for multi-tenant queue-jumping scheduling.
[0072] S2.2: Generate a comprehensive computing power vector according to the hardware parameters and dynamic operation metrics;
[0073] In this embodiment, the generation of the comprehensive computing power vector is to integrate the hardware parameters collected in the previous step (such as the number of CUDA cores, Tensor Core acceleration capabilities, etc.) and the dynamic operation metrics (GPU utilization rate, temperature, power consumption, video memory fragmentation degree, etc.) into a unified data structure. When different GPUs face the same load, they may not be able to exert their theoretical peak computing power due to excessive temperature or video memory fragmentation. If these dynamic factors are not considered, it will be difficult to perform accurate preemption scheduling in a multi-tenant scenario.
[0074] Specifically, first create an attribute mapping table for each GPU unit, convert the hardware parameters into static computing power descriptions, such as the peak computing power to represent the computing capabilities that can be provided in an ideal environment; then introduce attenuation coefficients or weight factors for the dynamic operation metrics to characterize the possible performance degradation under the current load and environmental temperature. Through this process, the current GPU hardware limit and the real-time working state are integrated into a unified computing power index, thereby effectively reducing the large-scale data write-back or calculation delay caused by preemption decision-making errors.
[0075] S2.3: Construct a resource collaboration matrix according to the physical topology of the GPU resource pool;
[0076] In this embodiment, the resource collaboration matrix is used to represent the topological relationship and data interaction cost between each GPU unit in the GPU resource pool. For multi-GPU servers or clusters, data is often transmitted between GPUs through PCIe, NVLink or internal high-speed interconnections with different bandwidth and latency characteristics; some GPUs may share a cache or a local NVLink channel in the same acceleration node, while others may need to communicate across nodes. When there is a need for multi-card collaboration or cross-card data dependency, simply treating all GPUs as "equivalent nodes" will lead to unnecessary communication losses or bandwidth bottlenecks. By describing the transmission bandwidth, communication latency or whether a specific cache module is shared between each pair of GPUs in the collaboration matrix, a GPU combination that is more friendly to data interaction can be preferentially selected in subsequent preemption scheduling, thereby improving the overall task execution efficiency and avoiding large-scale bandwidth occupation or communication congestion when low-priority tasks are frequently relocated to remote GPUs.
[0077] Specifically, for each GPU unit in the server or cluster, a unique identifier is assigned to it, and the network interconnection information between it and other GPU units is collected through the underlying topology discovery module. For example, for GPUs with NVLink interconnection, query the number of NVLink interfaces, the bandwidth limit, and whether there is a shared cache area; if the server uses PCIe interconnection or cross-node RDMA network, read the corresponding bandwidth, latency reference values, and whether there is a shared I / O hub for each PCIe channel or RDMA link. In this way, a connection table for describing the interconnection status of all GPUs can be obtained.
[0078] Furthermore, the connection table is sorted according to the two-dimensional structure of "GPU identifier × GPU identifier" to form an initial N×N adjacency information table. Each cell (i, j) corresponds to information such as the type of available interconnection channel, estimated bandwidth, and communication latency between GPU i and GPU j. If two GPUs are located in the same acceleration node and share the local NVLink channel, marks such as "shared Cache" or "zero-copy support" are added to the corresponding matrix cell to reflect the potential resource cooperation advantages. If cross-node communication is detected (e.g., through Infiniband or Ethernet, etc.), the specific bandwidth, latency, and possible congestion points are recorded.
[0079] Furthermore, after completing the preliminary adjacency information summary, it is necessary to further perform weighted or normalized processing on the data interaction cost in combination with the task scenario. On the one hand, for continuous metrics such as bandwidth and latency, weights can be assigned to different numerical segments according to the historical load test results to describe the difference from "high-speed interconnection" to "low-bandwidth remote"; on the other hand, for features such as shared cache or local NVLink, additional cooperation coefficients can be set to reflect the significant reduction in transmission overhead when multiple cards cooperate within the same GPU node. Finally, these weighted or normalized metrics are comprehensively filled into the N×N matrix cells to obtain a resource cooperation matrix that can quantify the "cooperation efficiency" or "remote communication cost" between each pair of GPUs.
[0080] Preferably, the resource cooperation matrix can also be managed in layers or blocks according to different levels of multi-card cooperation requirements:
[0081] For ordinary tasks that only require single - card execution, only focus on the self - loop within a single card in the matrix or the interconnection information with the CPU node; for multi - card parallel tasks or distributed training scenarios, focus on analyzing the NVLink, PCIe bandwidth, and cache sharing characteristics, and screen for GPU pairs or GPU clusters with higher cooperation coefficients in the matrix, thereby reducing the distributed communication bottleneck. During subsequent preemption scheduling, quickly determine which GPUs are suitable for cooperative operation, which GPUs have potential congestion points or are too far apart, and avoid randomly migrating low - priority tasks to remote GPUs across nodes, thereby reducing bandwidth waste and latency spikes caused by preemption, and further improving the flexibility and efficiency of multi - card resource allocation in a multi - tenant environment.
[0082] S2.4: Import the comprehensive computing power vector into the resource cooperation matrix to obtain the current computing power characteristic matrix;
[0083] In this embodiment, in a high - load or multi - tenant competition scenario, simply emphasizing the peak performance of a certain GPU may not bring the best experience. Especially when some tasks require multi - card parallelism or frequent data exchange, after matrix transformation, it is better to understand which GPU nodes have higher cooperation efficiency.
[0084] Specifically, by combining the aforementioned comprehensive computing power vector (describing the available computing power, attenuation coefficient, and ecological compatibility information of a single GPU) with the resource cooperation matrix (describing the physical connection relationship and communication characteristics between GPUs), a more comprehensive current computing power characteristic matrix can be generated. This computing power characteristic matrix not only includes the static and dynamic computing power performance of each GPU, but also embeds the bandwidth, latency, and communication resource occupancy during cross - GPU cooperation, thereby being able to more completely display the computing power distribution trend of the entire heterogeneous GPU resource pool. Thus, the preemption scheduling process can more accurately allocate appropriate GPUs, improve the response speed of sudden high - priority tasks, and take into account the fairness of multi - tenant resource usage.
[0085] The specific steps of S2.2 are as follows:
[0086] S2.2.1: Generate a static computing power vector according to the hardware parameters;
[0087] In this embodiment, the hardware parameters are merged into a group of vectors according to the established field order, so that each parameter occupies a specific position in the vector and maintains a fixed dimension, which is used for subsequent computing power evaluation and comparison. Through this vectorization method, on the one hand, it can maintain a unified parameter representation dimension in the entire heterogeneous GPU resource environment, reducing errors when comparing GPUs of multiple models and architectures; on the other hand, it can quickly provide a static reference value of "ideal performance" for subsequent steps, ensuring that each GPU can be accurately quantified at the performance baseline level, enabling subsequent preemption scheduling decisions to conduct performance evaluations under the same measurement system.
[0088] Specifically, first, a hardware description table with a fixed field order is defined for each GPU unit to list its CUDA core count, video memory capacity, FP32 / FP64 peak computing power, Tensor Core acceleration ability and other indicators. Subsequently, in the preset order of "field 1 - field 2 - field 3 - field 4...", these indicators are successively converted into values under the same dimension or unified reference, and they are arranged in order into a group of vectors. For example, the first vector element position can be reserved for the "CUDA core count", the second position for the "video memory capacity", and so on, and corresponding normalization processing is performed on parameters at different ratio levels. Ensure that during the vector construction process, the hardware parameters of different architectures (such as A100, RTX, and domestic GPUs) can be mapped to the same vector dimension and order position to avoid dimension conflicts or deviations during subsequent comparison or merging.
[0089] S2.2.2: Calculate the dynamic decay coefficient according to the dynamic operation indicators;
[0090] In this embodiment, by performing weighted analysis and mapping on the dynamic operation indicators, a dynamic decay coefficient can be calculated, which is used to characterize the difference between the available computing power of the GPU at the current moment and its static theoretical performance.
[0091] Specifically, if the temperature or power consumption is too high, the GPU may trigger protection mechanisms such as automatic frequency reduction and voltage reduction. If the video memory fragmentation rate is too high, it may lead to insufficient actual available video memory or decreased allocation efficiency. Quantifying and incorporating these adverse factors into the dynamic decay coefficient can make the subsequent computing power evaluation closer to the real-time working state of the GPU. By combining this coefficient with the static computing power vector, problems such as overload or efficiency decline caused by scheduling based solely on ideal performance can be avoided, and at the same time, the resource utilization rate and hardware safety redundancy can be taken into account during the cut-in scheduling process.
[0092] Furthermore, first, a weight parameter is set for each dynamic operation indicator, such as utilization occupancy weight, power consumption weight, temperature weight, video memory fragmentation weight, etc., and appropriate values are assigned to these weights according to business requirements or actual measurement data. Subsequently, the following principles are used to analyze each dynamic operation indicator and convert it into the corresponding decay contribution value:
[0093] If the current utilization rate of the GPU is relatively high (for example, continuously higher than a certain threshold), it means that the running tasks have occupied a considerable part of the computing power. At this time, a higher "utilization decay contribution value" will be given to reflect the fact that the available computing power has been partially consumed;
[0094] If the temperature sensor data is close to or exceeds the upper limit of the safe temperature, it indicates that the GPU may trigger downclocking or cause heat dissipation pressure, resulting in the actual available computing power being lower than the static theoretical value. Therefore, the "temperature decay contribution value" will be increased;
[0095] When the power consumption monitoring value continuously approaches the preset upper limit of power consumption or the heat dissipation bottleneck, the system will correspondingly increase the "power consumption decay contribution value" to indicate that the GPU can no longer run at the peak frequency without limit;
[0096] If the video memory fragmentation rate is high, it means that both the allocation efficiency and the available video memory space have decreased, and the "video memory fragmentation decay contribution value" will be increased, making an additional reduction to the final available computing power.
[0097] Furthermore, after obtaining each of the above decay contribution values, these contribution values will be multiplied by their corresponding weight parameters respectively, and then they will be accumulated to obtain a "total decay score". To make the decay coefficient more intuitive and facilitate subsequent combination with the static computing power vector, a normalization process will also be applied to the "total decay score" to keep the result within a reasonable range. Finally, this result is used as the dynamic decay coefficient. That is, if the "total decay score" is high, the dynamic decay coefficient is closer to 0, indicating that the computing power provided by the GPU at the current moment decays significantly; if the "total decay score" is low, the dynamic decay coefficient is closer to 1, indicating that the GPU is in good condition and the static theoretical performance can be exerted to a greater extent.
[0098] S2.2.3: Generate an ecological compatibility vector according to the type of GPU;
[0099] In this embodiment, an ecological compatibility vector is generated according to elements such as the GPU type (such as NVIDIA GPU or domestic GPU), its underlying driver ecosystem, AI framework support, and compiler library compatibility. The "ecological compatibility" here is not limited to simply "being able to run", but also includes information such as the acceleration and optimization degree for specific frameworks or operators, the stability test results at different precisions, and whether there are software and hardware co-optimization features for certain tasks. Adding these ecological compatibility elements in vector form can assign "specialized" evaluation weights to some GPUs during heterogeneous GPU scheduling: for example, if a domestic GPU has stable performance and low energy consumption in some inference scenarios, a higher "inference adaptation score" will be set for it when generating the ecological compatibility vector, so that this type of GPU will be given priority consideration during cut-in scheduling. This can not only give full play to the advantages of different GPU architectures, but also reduce the performance bottlenecks or incompatibility risks caused by blind scheduling in a multi-tenant environment.
[0100] Specifically, ecological compatibility can be divided into the following scoring dimensions: the first is the adaptability to mainstream AI frameworks (such as whether it is deeply optimized for TensorFlow and PyTorch, and whether it has dedicated acceleration operators), the second is the stability for different precision modes (such as the throughput and accuracy during FP16 and INT8 inferences are reliable), the third is the maintenance frequency and update compatibility of compilation libraries and underlying drivers (including whether there are version conflicts and whether containerized deployment is supported), and the fourth is the performance in special tests for target task categories (such as inferences, training, graphics rendering, etc.) (whether it has high efficiency and low error rate in the test set).
[0101] Furthermore, in terms of quantization methods, a scoring range can be defined for each dimension, and different GPU types can be scored based on a large amount of offline test data or real operation statistics. After completing the scoring for each dimension, these scoring results can be combined by weighted summation or vector concatenation. The finally formed ecological compatibility vector can not only intuitively reflect the optimization level of different GPUs under various mainstream frameworks and precision modes, but also provide a more refined basis for software and hardware cooperation in task queuing decisions, ensuring that the GPU more suitable for the current task requirements is preferentially scheduled in heterogeneous GPU scenarios, and minimizing the performance loss or potential failure risks caused by poor compatibility to the greatest extent.
[0102] S2.2.4: Generate a comprehensive computing power vector according to the static computing power vector, dynamic decay coefficient, and ecological compatibility vector;
[0103] In this embodiment, before inputting the task information and the current computing power characteristic matrix into the neural network, the input parameters need to be preprocessed to ensure that the input data can accurately reflect the resource demand characteristics of the task at different computing stages and highly match the available computing power status of the GPU. A more refined connection is established between the task computing process and the dynamic distribution of GPU resources, enabling the scheduling system to more intelligently adapt to complex heterogeneous computing environments.
[0104] Traditional methods often regard the computing requirements of tasks as a static whole and only evaluate based on global computing power consumption, ignoring the phased changes in computing power requirements during the execution of tasks. For example, during the data loading stage, the main bottleneck of the task may lie in video memory and PCIe bandwidth, while during the computing execution stage, the CUDA core utilization rate and the load of the Tensor Core are the main influencing factors. By the result writing back stage, the storage throughput and video memory recycling rate will become key parameters. If the data that has not been processed in stages is directly input into the neural network, it will cause the model to predict the computing power requirements of the task too vaguely, thus affecting the accuracy of scheduling decisions and even potentially causing a mismatch in computing power resources for some GPUs due to insufficient consideration of their short-term bandwidth limitations.
[0105] Specifically, this application does not simply decompose the task computing requirements. Instead, by constructing phased computing power vectors and establishing a fusion mapping with the current state of the GPU, it creatively improves the accuracy of task scheduling decisions. Especially in a heterogeneous GPU environment, there are significant differences in the support levels of different GPU architectures for data loading, computing execution, and result writing back. Without phased processing, some GPUs may be in an inefficient utilization state for a long time, and even the overall throughput of the cluster may be affected due to the mismatch between task requirements and GPU characteristics. Through refined task feature extraction and computing power adaptation strategies, a more targeted resource scheduling scheme is provided in the heterogeneous computing environment, thereby improving the intelligence level and resource utilization efficiency of the scheduling system.
[0106] The specific steps of the preprocessing are as follows:
[0107] Divide the task information into a data loading stage, a computing execution stage, and a result writing back stage, and generate phased computing power vectors for the resource requirements of each stage respectively;
[0108] Specifically, the computing power requirements of a task are not constant throughout its life cycle, but show phased changes as the computing process progresses. For example, in a deep learning task, the data loading stage mainly involves data preprocessing, memory allocation, and the transfer from the storage device to the GPU memory, so it is limited by the PCIe bandwidth, NVLink interconnection efficiency, and memory throughput capacity; the computing execution stage depends on the utilization rate of computing units such as CUDA cores and Tensor Cores and their peak computing power performance, while the result writing back stage mainly involves factors such as memory release, data storage write-back, and data transfer bandwidth between tasks. Therefore, if the global computing power requirements are directly used to measure the matching degree between the task and GPU resources, scheduling errors may occur due to different GPU architecture characteristics. For example, a GPU with high bandwidth but low computing power may be assigned to a compute-intensive task, resulting in a decrease in resource utilization.
[0109] In this embodiment, three types of phased computing power vectors are established, and each type of vector corresponds to the key computing power requirements of the task in that stage, ensuring that subsequent scheduling decisions can more accurately evaluate the adaptation degree of the task to GPU resources in different stages and optimize the execution efficiency of the task in a multi-tenant environment.
[0110] Furthermore, in the process of generating the phased computing power vector, first, parse information such as the input data scale, computing type, and model structure of the task to determine the memory allocation size, data transfer rate, and cache reuse rate required in the data loading phase, and quantify this information into a data loading vector. Then, combine the computing complexity, matrix operation requirements, parallel computing characteristics, etc. of the task to calculate the computing power requirements in the computing execution phase, and construct a computing execution vector based on the hardware acceleration capabilities of the GPU (such as the number of CUDA cores, Tensor Core support, etc.). Finally, in the result write-back phase, evaluate the video memory release speed, storage I / O throughput, and data consistency maintenance strategy according to the amount of computed result data that needs to be stored or transmitted back for the task, and generate a result write-back vector. The computing power requirements can be accurately described at each key stage of the task life cycle, avoiding the resource mismatch problem caused by a single computing power indicator in traditional scheduling methods, ensuring that the scheduling algorithm can comprehensively consider the different stage characteristics of the task when evaluating the matching degree between the task and GPU resources, and improving the accuracy of scheduling decisions and resource utilization rate.
[0111] Fuse the phased computing power vector with the current computing power characteristic matrix to obtain an initial fusion vector representing the adaptation degree of each stage to the current GPU idle resources;
[0112] Specifically, the current computing power characteristic matrix consists of information such as the hardware parameters of the GPU, dynamic load status, and task execution history, reflecting the available computing power, load pressure, and resource competition situation of the GPU at a specific time point. However, the computing power performance of the GPU is not constant but is dynamically affected by factors such as temperature, power consumption, and task type. Therefore, it is difficult to accurately reflect its true execution ability based on static computing power indicators alone. In addition, the scheduling of GPU resources depends not only on computing power but also on factors such as video memory fragmentation, data transfer efficiency, and interconnection architecture. Therefore, fusing the phased computing power vector with the current computing power characteristic matrix can simultaneously consider the static performance, dynamic load, and task phased requirements of GPU resources during the computing power evaluation process, enabling the insertion scheduling to more precisely match suitable GPU resources.
[0113] Further, during the fusion process, first, normalize the current computing power characteristic matrix of the GPU to ensure that the computing power parameters of different models of GPUs are compared on the same scale and eliminate the bias between different architectures. Then, match the data loading vector with the bandwidth resource metrics of the GPU (such as PCIe bandwidth, NVLink bandwidth, memory throughput rate, etc.), match the computing execution vector with the core computing power resources of the GPU (such as the number of CUDA cores, Tensor Core utilization rate, etc.), and match the result write-back vector with resource metrics such as storage I / O throughput and memory release rate. Through this stage-by-stage matching process, three sets of independent fitness scores can be generated, and these scores are further weighted and summed to generate the final initial fusion vector. The fusion process can not only ensure that the task can find the most suitable GPU resources at different stages but also dynamically adjust the resource allocation strategy during the scheduling process, enabling the task to be executed on a GPU with a lower load, avoiding resource bottlenecks on high-load GPUs, and improving the overall scheduling efficiency.
[0114] Specifically, through the preprocessing method of this embodiment, the neural network model can be optimized based on more complete task information and GPU resource status during the training process, enabling the model to more accurately predict the execution efficiency of the task and perform more reasonable resource allocation based on the computing power fitness. Compared with traditional solutions based on static computing power matching or single-priority scheduling, this method performs refined modeling at all stages of the task life cycle and makes the scheduling process more dynamically adaptable by fusing the current computing power characteristic matrix of the GPU.
[0115] Please refer to Figure 3 , the neural network model structure diagram of the embodiment of the present invention, the neural network model includes:
[0116] An input layer, which is used to extract features from the initial fusion vector through an attention mechanism to generate a deep representation vector representing the coupling relationship between the real-time load of the GPU and the phased requirements of the task;
[0117] An intermediate layer, which is used to process the deep representation vector through multiple rounds of iteration mechanisms and output intermediate results;
[0118] An output layer, which is used to perform a feasibility check on the intermediate result. If the check fails, locally adjust the preliminary computing power matrix according to the correction rule until the check passes, and output the preliminary computing power matrix that passes the check as the updated computing power characteristic matrix.
[0119] The specific steps of S3 are as follows:
[0120] S3.1: Load the neural network model constructed by fusing the physical model and the data-driven algorithm, and set the initial network parameters according to the current computing power characteristic matrix;
[0121] Specifically, the loaded neural network model adopts a deep fusion architecture of a physical model and a data-driven algorithm. The physical model part is constructed based on GPU hardware characteristic equations, such as the computing power attenuation model of CUDA cores and the energy consumption formula for the conversion of domestic GPU instruction sets, to ensure that the model output conforms to the hardware physical laws. The data-driven part trains a deep temporal convolutional network (DTCN) based on historical scheduling data (such as task execution time, resource utilization, error logs) to learn dynamic scheduling rules. When setting the initial parameters, the constraint conditions of the physical model (such as the upper limit of video memory capacity and the temperature safety threshold) are encoded as network weights through the pre-training stage, ensuring that the initial parameters not only conform to the physical laws but also have the adaptability of data-driven.
[0122] Furthermore, in order to make the initial parameters of the model closer to the actual operating state of the current GPU resource pool, when loading the network, the available GPU resource information in the current computing power characteristic matrix will be read first, including the GPU load situation, available video memory size, real-time power consumption, computing core utilization, etc., and this information will be used to initialize the input layer weights of the neural network, enabling it to have a high adaptability to the current hardware environment at the initial stage of model operation.
[0123] Preferably, in order to improve the compatibility of the model with different GPU architectures, the standard computing power characteristic vectors of different GPU types will also be loaded during the initialization stage, so that even if there are computing devices with significantly different architectures such as NVIDIA A100, RTX 4090, or domestic GPUs in the resource pool, the neural network can still perform reasonable scheduling according to their respective characteristics. It can reduce scheduling errors during actual operation, improve the matching degree between tasks and GPU resources, and at the same time avoid task execution failures or performance waste caused by unreasonable computing power resource allocation, thereby realizing an intelligent GPU computing power allocation strategy.
[0124] S3.2: Input the initial fusion vector into the input layer of the neural network model, and perform feature extraction on the initial fusion vector through the attention mechanism to generate a deep representation vector representing the coupling relationship between the real-time load of the GPU and the phased requirements of the task;
[0125] In this embodiment, since the available computing power of the GPU is affected by multiple factors such as the load situation, power consumption state, and task concurrency, and the computing power requirements of the task itself also show phased characteristics, it is necessary to introduce an attention mechanism to perform feature extraction on the input data to generate a deep representation vector that can accurately represent the relationship between the real-time load of the GPU and the computing power requirements of the task.
[0126] Specifically, after the initial fusion vector is input, the neural network first normalizes the vector to ensure that the input data of different task types and different GPU models can be mapped to the same feature space, avoiding training instability caused by numerical scale differences. Subsequently, the attention mechanism automatically analyzes the dependence of the current task on different GPU computing power resources at different computing stages (such as data loading, computing execution, and result writing back). For example, some tasks may be more dependent on video memory bandwidth, while others may be mainly affected by the number of computing cores. During this process, the attention layer in the network calculates the weight distribution of different input features, so that the computing power features that have a greater impact on the current task are assigned higher weights, while the influence of features with less impact is reduced. This process can effectively avoid the resource mismatch problem that may occur in the traditional neural network when using a fixed weight allocation method during feature learning, so that the finally generated deep representation vector can more accurately reflect the coupling relationship between the GPU load situation and the task computing power requirements.
[0127] Furthermore, a trilinear transformation is first performed on the initial fusion vector to generate a query matrix (Query), a key matrix (Key), and a value matrix (Value). The query matrix is used to represent the computing power requirements concerned by the current task, the key matrix is used to represent the computing power distribution in the current GPU resource pool, and the value matrix corresponds to the resources available for actual computing. During the weight calculation process, different resource focus points will be automatically learned at different computing stages of the task. For example, during the data loading stage, the attention mechanism will give higher weights to resources such as PCIe bandwidth and NVLink interconnection. During the computing execution stage, it may pay more attention to computing characteristics such as CUDA core utilization rate and Tensor Core acceleration ratio. During the result writing back stage, it is more inclined to consider video memory management and storage throughput.
[0128] Furthermore, after obtaining the attention weight scores, the different GPU computing power features are weighted and summed according to these weights to form a new task-computing power adaptation representation. The goal of this step is to let the neural network focus on the most critical computing power features while ignoring secondary features, improving the matching degree between computing resources and task requirements. One attention head may mainly focus on the bandwidth requirements of the task, while another attention head focuses on the computing load. The final result is the weighted sum of multiple attention heads, making the final feature expression more comprehensive and avoiding information loss from a single perspective.
[0129] Furthermore, after all calculations are completed, the obtained weighted feature representation will be used as the deep representation vector and input into the subsequent layers of the neural network for further task scheduling optimization. This deep representation vector will be the core input for subsequent scheduling optimization, providing a higher-quality computing decision basis for the neural network model and ensuring that the scheduling strategy can be adaptively adjusted as the dynamic changes of the GPU state.
[0130] S3.3: Process the depth representation vector through a multi-round iteration mechanism preset in the intermediate layer of the neural network model to output an intermediate result, where each iteration round corrects the output according to the dynamic constraints between the depth representation vector and the physical model, and adjusts the initial network parameters through backpropagation according to the correction result until the number of iterations is reached;
[0131] In this embodiment, in order to optimize the adaptability and robustness of the neural network in computing power allocation decision-making, a multi-round iteration mechanism is adopted to gradually optimize the depth representation vector, and the scheduling scheme is corrected in combination with the dynamic constraints of the physical model in each iteration round. The core goal of the iteration mechanism is to enable the neural network to accurately predict the optimal GPU resource matching scheme for tasks in a dynamic load environment, while ensuring that the calculation results meet the power consumption, security, and task execution constraints of the actual GPU.
[0132] Specifically, first, a multi-round loop calculation module is preset in the intermediate layer of the neural network. The role of this module is to gradually optimize the output result of the neural network based on the input depth representation vector. In the first iteration round, the network will initially predict the suitable GPU resources for each task according to the computing requirements of the current task, the phased computing power requirements of the task, and the GPU computing power characteristic matrix, and store the prediction result in a temporary task allocation matrix. This task allocation matrix is used to store the number of tasks assigned to each GPU, the computing power consumption, and the possible computing load of the task on this GPU. On this basis, the neural network will calculate the global load distribution under the current scheduling scheme and generate a global computing power balance score accordingly. This score is used to measure the rationality of the current task scheduling scheme.
[0133] Furthermore, in each subsequent iteration process, the neural network will correct the scheduling scheme based on the task allocation matrix of the previous round in combination with the dynamic constraints in the physical model. Specifically, the physical model provides a set of key constraint conditions, including the maximum power consumption limit of the GPU, the temperature upper limit, the video memory occupancy rate, the task execution concurrency, etc. For example, if the predicted power consumption of a certain GPU in the previous round of scheduling result exceeds the safety threshold of this GPU, then in the next iteration round, the network will automatically reduce the task load of this GPU, or adjust the tasks running on this GPU to assign them to GPUs with lower computing loads. At the same time, in order to ensure that the execution stability of other tasks will not be affected during the task preemption scheduling process, the physical model will also evaluate the task switching cost to avoid excessive calculation delay or video memory read / write overhead caused by frequent task migration.
[0134] Furthermore, in each iteration, the neural network not only adjusts the resource allocation scheme for tasks but also updates the network parameters through the backpropagation algorithm to ensure that the model can gradually converge to the optimal scheduling strategy. The loss function of the current scheduling scheme is calculated in each iteration. This loss function comprehensively considers the following indicators: (1) The execution efficiency of tasks, that is, whether the expected completion time of tasks has been optimized; (2) The load balancing of GPUs, that is, whether the computational loads of each GPU tend to be reasonably allocated; (3) The resource utilization rate, that is, whether the usage rate of computing power resources is close to the optimal level; (4) The compliance with physical constraints, that is, whether the current scheduling scheme violates the power consumption, temperature, or video memory usage limits of GPUs. Based on these loss indicators, the backpropagation algorithm calculates the gradient of each neuron and adjusts the network parameters to make the output result of the next iteration closer to the optimal computing power scheduling strategy.
[0135] Preferably, to prevent the neural network from falling into a local optimal solution during multiple iterations, a random perturbation mechanism is introduced after each iteration of the network model. The resource allocation scheme of some tasks is randomly adjusted within a certain range to explore possible better scheduling schemes. The random perturbation mechanism is similar to the Simulated Annealing algorithm, which can effectively prevent the neural network from prematurely converging to a suboptimal solution in the early iteration stage. Instead, it can search among more scheduling schemes and finally obtain the global optimal solution. At the same time, after each iteration ends, the current task allocation matrix is compared with the matrix of the previous iteration, and the convergence degree index is calculated. If the change range of the current iteration's scheduling scheme is extremely small compared to the previous iteration, or the decrease range of the loss function has been lower than the set threshold, it is determined that the current iteration has basically converged, and the iteration process is terminated in advance to improve the calculation efficiency.
[0136] Specifically, in the entire iteration mechanism, each iteration not only optimizes the rationality of the scheduling scheme but also continuously optimizes the parameters of the neural network, enabling the model to better adapt to the dynamic changes of GPU resources. In practical applications, this method can effectively improve the intelligence level of task scheduling, make task allocation more accurate, resource utilization more balanced, and ensure that task queuing in a multi-tenant environment does not affect the overall stability of the system. Finally, after multiple iterations of optimization, the scheduling scheme output by the neural network not only has high computing resource utilization but also complies with the physical constraints of GPUs, achieving a more intelligent and refined GPU task scheduling strategy.
[0137] S3.4: Map the intermediate result to the dimensional structure of the current computing power characteristic matrix to obtain the preliminary computing power matrix after mapping;
[0138] In this embodiment, the intermediate result of the neural network needs to be converted into an actually executable task scheduling scheme, and the core of this conversion process is to map to the dimensional structure of the current computing power characteristic matrix to generate a preliminary computing power matrix after mapping. The main objective of this mapping process is to ensure that the output of the neural network is not only the theoretically optimal allocation scheme, but also can adapt to the actual resource situation of the GPU and ensure the executability of tasks on the GPU. To achieve this goal, the entire mapping process is divided into multiple key steps, including task computing power requirement conversion, GPU resource occupancy matrix construction, topology adaptation processing, constraint condition correction, etc., to ensure that the preliminary computing power matrix has actual execution capabilities.
[0139] Specifically, first, in the task computing power requirement conversion stage, the intermediate result output by the neural network is often a vector containing task scheduling suggestions, including the types of computing resources required by the task (such as CUDA cores, Tensor Cores, video memory capacity, bandwidth, etc.) and the corresponding allocation ratios. These data need to be further converted into a standard format compatible with the GPU resource pool. Therefore, in this embodiment, a GPU resource occupancy matrix is used for standardized representation. The GPU resource occupancy matrix is a multi-dimensional matrix, where each row represents a task and each column represents a GPU resource item (such as the number of computing cores, video memory occupancy, power consumption budget, etc.). When constructing this matrix, first, according to the computing power requirement vector of the task, search for the GPU resource items in the current computing power characteristic matrix item by item, and fill in the task demand in the corresponding column. For example: if the task requests 20GB of video memory, then fill in 20GB in the "Video Memory Occupancy" column of the matrix; if the task requires Tensor Core computing power, then fill in the corresponding computing power requirement value in the "Tensor Core Utilization Rate" column. This conversion process ensures that the output of the neural network is consistent with the GPU resource characteristics, enabling subsequent scheduling to be matched based on the actual hardware situation.
[0140] Furthermore, after the GPU resource occupancy matrix is constructed, topology adaptation processing is required to ensure that the task scheduling scheme can match the physical characteristics of different GPU architectures. Due to the different hardware architectures of heterogeneous GPU resources, such as the NVIDIA A100 series using NVLink high-speed interconnection, and some RTX series mainly rely on PCIe bandwidth, domestic GPUs may use self-developed communication protocols, so the interconnection topology between GPUs needs to be considered when allocating tasks to optimize data transmission efficiency. For example, if a task contains a large number of model parameter synchronization operations, and multiple computing nodes need to share data, it should be allocated to the GPU group that supports NVLink interconnection to reduce data transmission bottlenecks; for tasks that are computationally intensive but have less data interaction, they can be allocated to RTX GPUs or other independent computing units. This adaptation process is achieved by analyzing the GPU interconnection topology matrix, which stores the connection mode and bandwidth information between each GPU, and when calculating task allocation, by finding the optimal interconnection path, it ensures the efficiency of task data transmission and avoids computing bottlenecks caused by low interconnection rates between GPUs.
[0141] Furthermore, in the process of constraint correction, in order to ensure that the final generated preliminary computing power matrix is actually executable, it is also necessary to check and adjust the legality of task allocation, which mainly involves the following aspects: (1) Power budget constraint, that is, check whether the total power consumption of the current GPU exceeds the rated upper limit. For example, if the task allocation of a GPU causes the power consumption to exceed the safety threshold, it is necessary to adjust the resource occupancy of some tasks in the matrix to reduce the computing intensity or migrate to other GPUs; (2) The problem of video memory fragmentation, that is, to ensure that the video memory utilization rate is within a reasonable range after task allocation to avoid the phenomenon of task loading failure due to video memory fragmentation. For example, in the process of constructing the preliminary computing power matrix, if it is found that there is a large fragmentation in the video memory occupancy, the video memory application strategy of the task can be adjusted, such as giving priority to allocating continuous memory areas or appropriately delaying the execution of some tasks; (3) Fairness of tenant resource allocation. In a multi-tenant environment, it is necessary to check whether the tasks of each tenant comply with the preset quota strategy to prevent a single tenant from occupying high-performance GPU resources for a long time. If it is found that the resource allocation is unbalanced, the task priority in the computing power matrix is adjusted to make the computing power resource allocation between different tenants more in line with the established strategy.
[0142] S3.5: Perform a feasibility test on the preliminary computing power matrix. If the test fails, locally adjust the preliminary computing power matrix according to the correction rules until the test passes.
[0143] In this embodiment, to ensure that the final task scheduling scheme can be actually executed, it is necessary to perform a feasibility check on the preliminary computing power matrix. This check process includes examining whether the GPU resource allocation exceeds the load limit, whether the tasks violate the tenant's computing power quota, and whether there are resource competition conflicts between tasks. If problems are found in the preliminary computing power matrix, the matrix will be adjusted according to the correction rules. For example, reducing the computing resource allocation of some tasks, rescheduling them to GPUs with lower loads, or adjusting the task priorities to optimize the overall scheduling strategy. The introduction of this step can effectively avoid GPU overload or task execution failures caused by scheduling mistakes, and improve the stability and execution efficiency of the task allocation scheme.
[0144] S3.6: Use the preliminary computing power matrix that passes the check as the updated computing power characteristic matrix.
[0145] The specific steps of S4 are as follows:
[0146] S4.1: According to the GPUs in the updated computing power characteristic matrix that match the task to be scheduled, count the GPU resource list of idle GPU resources;
[0147] In this embodiment, to ensure that the scheduling decision can accurately match the real-time state of the current GPU resource pool, before task allocation, it is necessary to count the current available GPU resource list based on the updated computing power characteristic matrix.
[0148] Specifically, first, the computing power parameters of each GPU in the updated computing power characteristic matrix will be read, including the available proportion of computing cores (such as CUDA Core, Tensor Core, etc.), the remaining video memory space, the data throughput capacity, and the occupation of parallel tasks, and screening will be performed based on task requirements. For example, if the task to be scheduled is a deep learning training task with high video memory requirements, the system will preferentially select GPUs with sufficient remaining video memory space and strong computing capabilities to avoid a decline in task running efficiency caused by insufficient video memory or limited computing capabilities.
[0149] Preferably, during the counting process, the physical topology of the GPUs, such as the NVLink or PCIe bandwidth limitations between GPUs, also needs to be considered to ensure that the tasks can obtain the best data transmission efficiency in the distributed computing mode. By constructing a complete list of allocable GPU resources, it can be ensured that the scheduling algorithm has a more comprehensive computing basis in subsequent resource allocation and cut-in decision-making, thereby maximizing the utilization rate of GPU resources and reducing resource mismatches in task scheduling.
[0150] S4.2: Determine whether it is necessary to perform cut-in scheduling or preemption operations on the existing running tasks. If not, directly allocate idle GPU resources to the task to be scheduled;
[0151] In this embodiment, based on the already - counted list of available GPU resources, it is necessary to further determine whether to perform preemptive scheduling or seize the resources of existing tasks. This determination process mainly depends on multiple factors, including the priority of the task to be scheduled, the real - time load of GPU resources, the distribution of computing requirements in the current task queue, and the computing power quota policy of the tenant.
[0152] Specifically, if there are idle GPUs in the available GPU resource list that meet the requirements of the current task, there is no need to perform preemptive scheduling. The most suitable GPU can be directly allocated to the task and resource binding can be performed. However, in a large - scale concurrent task environment, idle GPU resources may be extremely limited. Therefore, it is necessary to further evaluate whether there are low - priority tasks occupying high - performance GPUs to decide whether to perform a preemption operation. To avoid an excessive task interruption rate due to frequent preemption, when making a preemption decision, the system will also combine historical task scheduling data to predict the computing time of the target task and its impact on other tasks. For example, if a low - priority task still needs to run for a long time and a high - priority task needs to be executed immediately, the system will tend to perform preemptive scheduling to ensure that the high - priority task can obtain resources first while minimizing the impact on the low - priority task. Through this judgment mechanism, the system can balance the urgency of tasks and the availability of resources, ensuring the fairness and efficient utilization of computing power resources.
[0153] S4.3: If yes, according to the updated computing power characteristic matrix, perform resource rearrangement on the preempted or re - allocated tasks;
[0154] In this embodiment, when it is determined that preemptive scheduling needs to be performed, it is necessary to reasonably rearrange the resources of the preempted tasks according to the updated computing power characteristic matrix to ensure that the preempted tasks will not be completely interrupted or fail to execute due to resource recycling. The core of resource rearrangement is to find suitable alternative GPU resources for the preempted tasks and ensure that their computing environments are migrated as stably as possible.
[0155] Specifically, this process first checks the computing status of the preempted task, including the executed time, computing stage, memory usage, and data loading progress, to determine whether it can be safely migrated to other GPUs. If the execution progress of the preempted task is short, or its computing status allows for recovery, the system will select a GPU with lower load to migrate the task to a new computing node and synchronize its computing environment. For example, it uses container snapshot technology (such as Checkpoint / Restore in Userspace, CRIU) to save the computing context of the task and restore the computing status on the new GPU to ensure the smoothness of the migration process. If the preempted task cannot be fully migrated, a phased preemption strategy may be adopted, that is, allowing the task to pause after executing to a certain safe checkpoint to avoid waste of computing resources due to mid-course preemption. Through such a resource rearrangement mechanism, the system can minimize the impact of preemptive scheduling on low-priority tasks while ensuring that high-priority tasks can obtain computing resources in a timely manner, optimizing the overall GPU computing power scheduling efficiency.
[0156] S4.4: Allocate the idle GPU resources according to the computing requirements of the task to be scheduled, and return the final allocation plan and the result of the cut-in scheduling to the scheduling management port to complete the running binding of the task to be scheduled on the GPU resources;
[0157] In this embodiment, after completing the cut-in scheduling or directly allocating resources, it is necessary to confirm the final GPU resource allocation plan and return the scheduling result to the scheduling management port to complete the binding of the GPU resources for the task to be scheduled. The core of this step is to ensure that the scheduling of the task on the GPU can be stably executed, and the scheduling information can be synchronized to the management system for subsequent monitoring and adjustment.
[0158] Specifically, during the resource allocation process, the system will, according to the task requirement information provided by the updated computing power characteristic matrix, bind a suitable computing node for the task and allocate key resources such as computing cores, video memory space, and data transfer bandwidth. For example, if a task's computing stage requires a large amount of matrix operations, the system will preferentially allocate a GPU with Tensor Core acceleration capabilities to improve the computing efficiency. At the same time, when binding the task, the system will also configure appropriate execution strategies for the task, such as whether to allow dynamic adjustment of computing power resources and whether to allow the task to pre-occupy more video memory space to maintain computing stability under high load.
[0159] Preferably, the scheduling result will be synchronized to the scheduling management port, including information such as the computing node of the task, GPU resource allocation, execution priority, preemption impact, etc., so as to dynamically adjust computing resources during subsequent task scheduling. For example, if a task causes an extended computing time due to a sudden cut-in scheduling, the scheduling system can automatically adjust the priority of the task to compensate for its computing time, thereby ensuring the fairness of task execution. Through this step, the system can not only complete the GPU resource binding of the task, but also ensure the transparency and traceability of task scheduling, improve the intelligence of computing power scheduling, and optimize the utilization rate of computing resources in a multi-user environment.
[0160] Embodiment 2:
[0161] Please refer to Figure 4 , the present invention provides an embodiment: a multi-user computing power quota intelligent cut-in scheduling system, the system includes an information collection module, an information processing module, and a cut-in scheduling module;
[0162] The information collection module is used to obtain the task information of the task to be scheduled and the status information of heterogeneous GPUs;
[0163] The information processing module is used to determine the current computing power characteristic matrix according to the status information, and use the task information and the current computing power characteristic matrix as inputs to a preset neural network model to obtain an updated computing power characteristic matrix;
[0164] The cut-in scheduling module is used to allocate GPU resources for the task to be scheduled according to the updated computing power characteristic matrix.
[0165] The information processing module includes:
[0166] An information preprocessing unit for determining the current computing power characteristic matrix;
[0167] A model processing unit for obtaining an updated computing power characteristic matrix through an attention mechanism and multiple rounds of iterative correction.
[0168] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limitations of the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A multi-user computing power quota intelligent cut-in scheduling method, which is applied to a heterogeneous GPU resource environment, is characterized in that The intelligent queue-jumping scheduling method includes: Obtain the task information of the task to be scheduled and the status information of heterogeneous GPUs; Determine the current computing power characteristic matrix according to the status information; Take the task information and the current computing power characteristic matrix as inputs, input them into a preset neural network model, and obtain an updated computing power characteristic matrix through an attention mechanism and multiple rounds of iterative correction; Allocate GPU resources for the task to be scheduled according to the updated computing power characteristic matrix; Taking the task information and the current computing power characteristic matrix as inputs includes: Divide the task information into a data loading stage, a computing execution stage, and a result writing-back stage, and generate stage-based computing power vectors for the resource requirements of each stage respectively; Fuse the stage-based computing power vectors with the current computing power characteristic matrix to obtain an initial fusion vector representing the adaptation degree of each stage to the current GPU idle resources; The input to the preset neural network model includes: Load the neural network model and set the network initial parameters according to the current computing power characteristic matrix; Input the initial fusion vector into the input layer of the neural network model, perform feature extraction on the initial fusion vector through an attention mechanism, and generate a deep representation vector representing the coupling relationship between the real-time load of the GPU and the stage-based requirements of the task; Process the deep representation vector through a multi-round iteration mechanism preset in the middle layer of the neural network model, output intermediate results, where in each iteration round, the output is corrected according to the dynamic constraints between the deep representation vector and the physical model, and the network initial parameters are adjusted through backpropagation according to the correction results until the iteration times are reached; The obtaining of the updated computing power characteristic matrix includes: Map the intermediate results to the dimension structure of the current computing power characteristic matrix to obtain a preliminary computing power matrix after mapping; Conduct a feasibility test on the preliminary computing power matrix; If the test fails, locally adjust the preliminary computing power matrix according to the correction rules until the test passes; Take the preliminary computing power matrix that passes the test as the updated computing power characteristic matrix.
2. The multi-user computing power quota intelligent queue-jumping scheduling method according to claim 1, wherein The determination of the current computing power characteristic matrix includes: Obtain the hardware parameters and dynamic operation indicators of each GPU unit according to the status information; Generate a comprehensive computing power vector according to the hardware parameters and dynamic operation indicators; Construct a resource collaboration matrix according to the physical topology structure of the GPU resource pool; Import the comprehensive computing power vector into the resource collaboration matrix to obtain the current computing power characteristic matrix.
3. The multi-user computing power quota intelligent queue-jumping scheduling method according to claim 2, wherein Generating a comprehensive computing power vector according to the hardware parameters and dynamic operation indicators includes: Generate a static computing power vector according to the hardware parameters; Calculate the dynamic attenuation coefficient according to the dynamic operation indicators; Generate an ecological compatibility vector according to the type of GPU; Generate a comprehensive computing power vector according to the static computing power vector, the dynamic attenuation coefficient, and the ecological compatibility vector.
4. The multi-user computing power quota intelligent queue-jumping scheduling method according to claim 1, wherein The neural network model is established based on a physical model and data-driven technology.
5. The multi-user computing power quota intelligent cut-in scheduling method according to claim 1, wherein Allocating GPU resources for the task to be scheduled according to the updated computing power characteristic matrix includes: According to the GPUs in the updated computing power characteristic matrix that match the task to be scheduled, count the GPU resource list of the idle GPU resources; Determine whether to perform preemption scheduling or preemption operation on the existing running tasks. If not, directly allocate idle GPU resources to the to-be-scheduled task; If so, rearrange the resources of the preempted or reallocated tasks according to the updated computing power characteristic matrix; Allocate the idle GPU resources according to the computing requirements of the to-be-scheduled task, and return the final allocation plan and the preemption scheduling result to the scheduling management port to complete the running binding of the to-be-scheduled task on the GPU resources.
6. A multi-user computing power quota intelligent queue-jumping scheduling system for implementing the multi-user computing power quota intelligent queue-jumping scheduling method according to any one of claims 1-5, characterized in that The system includes an information collection module, an information processing module, and a preemption scheduling module; The information collection module is used to obtain the task information of the to-be-scheduled task and the status information of the heterogeneous GPU; The information processing module is used to determine the current computing power characteristic matrix according to the status information, and use the task information and the current computing power characteristic matrix as inputs to a preset neural network model to obtain an updated computing power characteristic matrix; The preemption scheduling module is used to allocate GPU resources to the to-be-scheduled task according to the updated computing power characteristic matrix.
7. The multi-user computing power quota intelligent cut-in scheduling system according to claim 6, wherein The information processing module includes: An information preprocessing unit for determining the current computing power characteristic matrix; A model processing unit for obtaining an updated computing power characteristic matrix through an attention mechanism and multi-round iterative correction.
Citation Information
Patent Citations
Calculation power resource scheduling method, and training method and system of calculation power resource scheduling model
CN116643877A
Task processing method and device of computing power network, electronic equipment and storage medium
CN117675823A
Cited By
AI server management system based on GPU computing power optimization
CN121785788A