Container-based cross-data center computing power resource scheduling method
By generating accurate resource consumption estimates through fine-grained resource reports and task profile databases from each physical GPU node in a cross-data center environment, and combining node agent programs with GPU API gateway modules, fine-grained and efficient scheduling of computing resources across data centers is achieved, solving the problems of low resource utilization and low scheduling efficiency in existing technologies.
Patent Information
- Application Number
- CN202511064500.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-07
AI Technical Summary
Existing computing resource scheduling schemes suffer from low GPU resource utilization, serious resource waste, and difficulty in achieving fine-grained management and efficient scheduling in cross-data center scenarios, especially in fine-grained tasks where resource allocation mismatch and network latency complexity lead to low scheduling efficiency.
By having each physical GPU node actively report fine-grained resource reports, the scheduler aggregates cluster resource snapshots, combines them with a task profile database to generate accurate resource consumption estimates, performs fine-grained resource allocation scoring and selection, and uses node agent programs and GPU API gateway modules for container configuration to achieve refined and efficient cross-data center computing resource scheduling.
It improves the overall utilization and scheduling efficiency of GPU resources, realizes on-demand allocation and efficient sharing, solves the problems of resource waste and low scheduling efficiency in traditional scheduling methods, and adapts to the complex challenges of heterogeneous environments across data centers.
Smart Images

Figure CN120909788A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of resource scheduling, and more particularly, to a container-based cross-data-center computing power resource scheduling method. BACKGROUND
[0002] With the rapid development of artificial intelligence and deep learning technologies, the demand for high-performance computing power resources, especially GPUs, is growing explosively. In order to meet these growing computing needs, cloud computing and distributed system architectures have become mainstream. Due to its lightweight, portability and high efficiency, container technology has shown great advantages in deploying and managing these complex applications. However, deploying GPU resources within a single data center often fails to meet the large-scale, high-concurrency computing power needs, so extending computing power resources to cross-data-center for unified management and scheduling becomes particularly important, which not only improves resource utilization, but also enhances the flexibility and reliability of the system, effectively dealing with local failures or traffic peaks, and achieving global optimization of resource allocation.
[0003] However, existing computing power resource scheduling schemes face many challenges in handling fine-grained GPU resource scheduling in cross-data-center scenarios. Traditional scheduling methods often allocate resources in units of physical machines or virtual machines, which results in low GPU resource utilization, especially for fine-grained tasks that only require partial GPU capabilities. In addition, most existing schedulers lack fine-grained perception of GPU internal resources, such as memory, stream processors (SM), and other key indicators, making it difficult to accurately match the actual needs of tasks. In a cross-data-center environment, simple resource aggregation and scheduling strategies cannot effectively deal with complex problems such as network latency, data consistency, and heterogeneous resource management, resulting in low scheduling efficiency and failing to fully exploit the potential of distributed GPU clusters, making it difficult to achieve truly on-demand allocation and efficient sharing.
[0004] Therefore, in order to achieve truly efficient GPU resource sharing, a container-based cross-data-center computing power resource scheduling scheme is desired. SUMMARY
[0005] In view of the above limitations of existing methods, according to an aspect of the present application, a container-based cross-data-center computing power resource scheduling method is provided, which includes: reporting the shareable GPU resources of each physical GPU node to a scheduler, and the scheduler aggregates the shareable GPU resource reports of the physical GPU nodes to obtain a cluster shareable GPU resource snapshot; obtaining a fine-grained GPU task to be scheduled submitted by a user; Load or generate the profiled to-be-scheduled task from a task profile database based on the to-be-scheduled fine-grained GPU task; The scheduler filters a feasible node list from the cluster-shareable GPU resource snapshot according to the profiled to-be-scheduled task; According to the profiled to-be-scheduled task, each physical GPU node in the feasible node list is scored and selected for fine-grained resource allocation to obtain a scheduling decision, which includes a target node ID. The scheduler sends the scheduling decision to a node agent program running on the target node, which is used to interact with the GPU API gateway module of the target node to configure isolation parameters for the to-be-run container. After receiving the resource limit configuration completion signal, the node agent program starts the container on the target node, and when the container is started, the GPU API call is redirected to the GPU API gateway module.
[0006] Compared with the prior art, the container-based cross-data center computing power resource scheduling method provided by the present application first requires each physical GPU node to report its shareable fine-grained GPU resource report, thereby overcoming the problem of insufficient GPU internal resource awareness of the traditional scheduler. Then, for the to-be-scheduled task submitted by the user, the estimated resource consumption data of the task is loaded or generated from the task profile database, solving the pain points of the scheduler being unable to accurately know the real demand of the task and the user's excessive application leading to resource waste. Most importantly, instead of simply filtering available nodes, the scheduler scores and selects each physical GPU node in the feasible node list for detailed resource allocation according to the profiled task demand, and determines the optimal target node by calculating the adaptation score. This intelligent matching mechanism, combined with the node agent program and the GPU API gateway module for configuring isolation parameters for the container and redirecting the GPU API call, effectively solves the problems of traditional scheduling logic being extensive, low GPU utilization, and difficulty in realizing multi-tenant fine-grained sharing, and finally realizes the fine and efficient scheduling of cross-data center computing power resources. BRIEF DESCRIPTION OF DRAWINGS
[0007] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description of embodiments of the present application, taken in conjunction with the accompanying drawings. The drawings provided in the specification and the embodiments of the present application together serve to explain the present application and, therefore, to provide further understanding of the present application, and form a part of the specification. The drawings do not limit the present application, but serve to explain the present application together with the embodiments of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0008] Figure 1 A flowchart of the container-based cross-data center computing power resource scheduling method according to the embodiments of the present application.
[0009] Figure 2 Data flow diagram for the container-based cross-data center computing resource scheduling method according to the embodiments of the present application.
[0010] Figure 3 Flowchart for step S3 in the container-based cross-data center computing resource scheduling method according to the embodiments of the present application.
[0011] Figure 4 Flowchart for step S5 in the container-based cross-data center computing resource scheduling method according to the embodiments of the present application. DETAILED DESCRIPTION
[0012] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings of the present disclosure are only for illustrative purposes and should not be construed as limiting the scope of protection of the present disclosure.
[0013] To solve the problems in the background art, the present application proposes a container-based cross-data center computing resource scheduling method. Figure 1 Flowchart for the container-based cross-data center computing resource scheduling method according to the embodiments of the present application. Figure 2 Data flow diagram for the container-based cross-data center computing resource scheduling method according to the embodiments of the present application. As shown in Figure 1 and Figure 2As shown, the container-based cross-data center computing power resource scheduling method according to the embodiment of the application comprises: S1, reporting the shareable GPU resources of each physical GPU node to a scheduler, and the scheduler aggregates the shareable GPU resource reports of the physical GPU nodes to obtain a cluster shareable GPU resource snapshot; S2, obtaining a to-be-scheduled fine-grained GPU task submitted by a user; S3, based on the to-be-scheduled fine-grained GPU task, loading or generating an imaged to-be-scheduled task from a task imaging database; S4, the scheduler filters a feasible node list from the cluster shareable GPU resource snapshot according to the imaged to-be-scheduled task; S5, according to the imaged to-be-scheduled task, performing fine-grained resource allocation scoring and selection on each physical GPU node in the feasible node list to obtain a scheduling decision, the scheduling decision comprising a target node ID; S6, the scheduler sends the scheduling decision to a node agent program running on the target node, and the node agent program is used to interact with a GPU API gateway module of the target node to configure isolation parameters for a container to be run; S7, after receiving a resource restriction configuration completion signal, the node agent program starts the container on the target node, and when the container is started, GPU API calls are redirected to the GPU API gateway module.
[0014] In step S1, the shareable GPU resource reports of each physical GPU node are reported to a scheduler, and the scheduler aggregates the shareable GPU resource reports of the physical GPU nodes to obtain a cluster shareable GPU resource snapshot. It can be understood that the existing computing power resource scheduling scheme generally lacks fine-grained perception ability of GPU internal resources, such as memory, stream processor (SM) and other key indicators, which causes the scheduler to be unable to accurately master the real available state and sharing ability of the GPU. This information loss makes the scheduling decision only stay at the coarse-grained level, and cannot realize efficient resource utilization and accurate task matching, especially in the cross-data center environment which is heterogeneous and dynamic, the problems of resource waste and low scheduling efficiency are particularly prominent. By requiring each physical GPU node to actively report the resource report containing these fine-grained information, and aggregating by the scheduler to form a global cluster shareable GPU resource snapshot, the method can provide the scheduler with a comprehensive, real-time and fine resource view, thereby overcoming the defect of insufficient perception of GPU internal resources in the traditional scheme, laying a foundation for subsequent accurate matching and efficient resource allocation based on task imaging, and further improving the overall utilization rate and scheduling efficiency of GPU resources.
[0015] The following is a specific implementation process of step S1: first, on each physical GPU node, a resource monitoring agent is deployed. The agent continuously monitors the real-time resource status of each physical GPU on the node, and periodically generates a shareable GPU resource report. The report contains detailed information about the key indicators of the GPU. Specifically, in one specific example of the present application, the shareable GPU resource report includes: GPU ID, physical GPU information, available video memory bytes, available SM percentage, available SM absolute number, current shared instance number, maximum shareable instance estimate, supported shared API list, currently active sharing strategy, and last update timestamp. Specifically: GPU ID, for example, GPU-DC1-NodeA-001, represents the first GPU on node A in data center 1, physical GPU information such as model, serial number, total video memory capacity, etc., available video memory bytes, for example, 20GB, available SM percentage, for example, 80%, available SM absolute number, for example, 100 SM, current shared instance number, for example, 2, maximum shareable instance estimate of the GPU, for example, 8, GPU support shared API list, for example, CUDA, OpenCL, and currently active sharing strategy, for example, time slice sharing, memory isolation. In addition, a last update timestamp is also attached to the report to ensure the timeliness of the information obtained by the scheduler. The resource report is periodically reported to the central scheduler through a secure network channel, such as a gRPC or HTTP / 2-based remote procedure call, every 5 seconds.
[0016] The central scheduler is internally provided with a resource snapshot management module. The module continuously receives shareable GPU resource reports from all cross-data-center physical GPU nodes. Whenever a new report is received, the resource snapshot management module will update or add to the cluster shareable GPU resource snapshot it maintains internally according to the GPU ID and last update timestamp in the report. The snapshot is a dynamically updated global view, and its structure can be a mapping table, where the key is the GPU ID and the value is the latest shareable GPU resource report content of the GPU. For example, the scheduler receives a report from GPU ID GPU-DC1-NodeA-001 on node 1 in data center A, and a report from GPU ID GPU-DC2-NodeB-001 on node 2 in data center B, and it will integrate the detailed resource status information in these reports into its cluster snapshot.
[0017] In step S2, the to-be-scheduled fine-grained GPU task submitted by the user is acquired. Accordingly, since the traditional computing power resource scheduling scheme often allocates resources in units of physical machines or virtual machines, this coarse-grained allocation method leads to low GPU resource utilization, especially for fine-grained tasks that only require part of the GPU capability. Therefore, in order to capture these specific and detailed computing power requirements from users, the to-be-scheduled fine-grained GPU task submitted by the user needs to be acquired. Only when the scheduler accurately acquires the task submitted by the user, which contains the user's preliminary fine-grained requirements for GPU resources, can a series of key processes such as subsequent task profiling, available resource matching, and final scheduling decision making be started.
[0018] The following is a specific implementation process of step S2: First, the user submits his computing task request to the scheduling service through a client interface, such as a command line tool, a web interface, or a programming interface. The request is carried in a structured data form and contains all the key information required to execute the task. These information together constitute the complete description of the to-be-scheduled fine-grained GPU task.
[0019] Specifically, a to-be-scheduled fine-grained GPU task object contains at least the following key information: an identification code for uniquely identifying the task, such as a string composed of a timestamp and a serial number; the application name or type to which the task belongs, such as image classification inference, natural language processing training, and an application profiling prompt for subsequent task profiling, such as indicating that the task is based on small batch inference of a specific model; the container image name of the specified task running environment, and the command and parameters to be executed when the container is started. In order to support fine-grained scheduling, the task information also needs to explicitly specify the specific requirements of the task for GPU resources, given in the form of a range or minimum requirement. For example, the task may declare its estimated minimum memory requirement and maximum memory requirement, such as a minimum of 768 megabytes and a maximum of 1536 megabytes, as well as an estimated minimum SM percentage requirement and a maximum SM percentage requirement, such as a minimum of 15% of the total SM of the GPU and a maximum of 30%. These fine-grained resource requirements are pre-set by the user according to the characteristics of the task, or automatically filled in by the client tool according to the application type.
[0020] In step S3, based on the to-be-scheduled fine-grained GPU task, the to-be-scheduled task that has been profiled is loaded or generated from the task profile database. It should be understood that the task request submitted by the user is often based on a rough estimate, and there may be excessive application, which causes the scheduler to be unable to accurately know the real resource consumption mode of the task, and further causes waste of GPU resources and low scheduling efficiency. By introducing the task profile database, a more accurate resource consumption estimation based on historical data or preset strategies can be generated for the to-be-scheduled task. This to-be-scheduled task that has been profiled contains the intelligent judgment of the system on the actual resource demand of the task, rather than relying only on the original request submitted by the user. This enables the scheduler to make decisions based on more real and more fine-grained resource requirements in the subsequent resource screening and allocation scoring stages, thereby realizing more efficient resource matching and utilization.
[0021] Specifically, in one specific example of the present application, Figure 3 The flowchart of step S3 in the container-based cross-data center computing resource scheduling method according to the embodiment of the present application is shown in FIG. 3. As shown in FIG. 3, in step S3, based on the to-be-scheduled fine-grained GPU task, the to-be-scheduled task that has been profiled is loaded or generated from the task profile database, including: S31, extracting an application profile prompt from the to-be-scheduled fine-grained GPU task; S32, using the application profile prompt as a query key, loading a profile matched with the query key from the task profile database; S33, generating estimated GPU resource consumption data based on the data in the profile matched with the query key; S34, appending the estimated GPU resource consumption data to the to-be-scheduled fine-grained GPU task to obtain the to-be-scheduled task that has been profiled. Figure 3
[0022] The following is a specific implementation process of step S3: first, S31 is executed. First, the structured information of the to-be-scheduled fine-grained GPU task is parsed, and a specific identifier for task profiling is identified and extracted therefrom. This identifier is called an application profile prompt, which is a string that can uniquely or approximately identify the task type, application scenario, or model feature. For example, if a task is to perform small-batch image inference of a ResNet50 model, the application profile prompt thereof can be set to ResNet50_Inference_SmallBatch. This prompt is filled in by the user or the client tool according to the task characteristics at the time of task submission, and is intended to provide a clear key for subsequent profile query.
[0023] Next, S32 is performed. After the application of the profile hint, the scheduler uses it as a query key to query the pre-built task profile database. The database stores a large amount of historical task running data and resource consumption patterns, which are statistically analyzed and summarized to form various profiles. For example, for the query key ResNet50_Inference_SmallBatch, the database can store historical statistical data of the memory, stream processor utilization, execution time, etc. of this type of task running on different GPU hardware, such as average memory consumption, 90% quantile of memory consumption, average SM utilization, etc. If there is a profile in the database that exactly matches the query key, it is loaded. In particular, in one specific example of the present application, based on the to-be-scheduled fine-grained GPU task, the profiled to-be-scheduled task is loaded or generated from the task profile database, including: if no profile matching the query key is found in the task profile database, the scheduler will use the Min value in the to-be-scheduled fine-grained GPU task to generate the profiled to-be-scheduled task. That is, if no profile matching the query key is found in the task profile database, the scheduler will use the Min value submitted by the user in the to-be-scheduled fine-grained GPU task, for example, the minimum memory of 2GB and the minimum SM percentage of 10%, as the estimated GPU resource consumption data to ensure that the task can at least obtain the minimum resources declared by it.
[0024] Then, S33 is performed. This process involves statistical analysis of historical data. For example, if the profile contains the historical distribution of memory consumption of a certain type of task, the scheduler can select a certain percentile of the distribution, for example, the 90% percentile, as the estimated memory consumption value to ensure that the task can obtain sufficient memory in most cases while avoiding excessive allocation. Similarly, for the consumption of SM, a similar method can be used to generate an estimated SM percentage based on historical SM utilization data. These estimated values are data-driven and more accurate than the rough range provided by the user initially. For example, for the ResNet50_Inference_SmallBatch task, the profile data can indicate that the 90% percentile of its memory consumption is 1.2GB and the 90% percentile of its SM utilization is 25%, which are the estimated GPU resource consumption data generated.
[0025] Finally, S34 is performed. The scheduler accesses the data structure of the original fine-grained GPU task to be scheduled. The data structure contains general information of the task, such as task identification code, container image, execution command, and user-provided initial GPU resource requirement range, such as minimum and maximum memory requirements, minimum and maximum SM percentage requirements. The scheduler adds or updates specific fields in this existing data structure to store the accurate estimated values obtained after profiling. For example, a substructure named estimatedGpuRequirements can be added, which contains fields such as estimatedMemoryBytes and estimatedSMPercent, and the specific values calculated in S33 are filled into these fields. In this way, the original fine-grained GPU task to be scheduled is converted into a profiled task to be scheduled. This new task object not only retains all the original information submitted by the user, but also embeds the GPU resource consumption prediction. For example, if the original task only specifies that the memory requirement is between 768MB and 1536MB, after profiling, the profiled task to be scheduled will explicitly contain an estimated value, such as estimated memory: 1200MB.
[0026] In step S4, the scheduler filters out a list of feasible nodes from the cluster shareable GPU resource snapshot according to the profiled task to be scheduled. Accordingly, since step S1 has provided fine-grained resource snapshots of all physical GPUs in the cluster, and step S3 has generated more accurate estimated GPU resource consumption data for the task to be scheduled, but when the cluster is large and deployed across data centers, directly performing complex adaptation scoring on all available resources will bring huge computational overhead. Therefore, the scheduler needs an efficient mechanism to quickly identify those nodes that at least meet the task requirements at the basic resource level from the vast amount of resources. Based on this, the present application filters out a list of feasible nodes from the cluster snapshot according to the estimated GPU resource consumption data of the profiled task to be scheduled, so that the scheduler can significantly reduce the range of subsequent fine-tuning and selection. This not only avoids scheduling the task to nodes that do not meet the resource conditions, thereby reducing invalid scheduling attempts and resource waste, but also improves the overall scheduling efficiency and the accuracy of resource allocation.
[0027] The following is a specific implementation process of step S4: the scheduler will check each physical GPU entry in the snapshot of cluster shareable GPU resources one by one. For each GPU in the snapshot, the scheduler will perform a series of conditional checks to determine whether the GPU can meet the basic resource requirements of the profiled task to be scheduled. The first check is the memory capacity check, in which the scheduler checks whether the available memory bytes of the GPU are greater than or at least equal to the estimated memory bytes of the profiled task to be scheduled. For example, if the task estimates 1.2 GB of memory, and a certain GPU currently only has 1 GB of available memory, then the GPU will be immediately excluded.
[0028] Next, the scheduler will perform the SM percentage check to ensure that the available SM percentage of the GPU can meet or exceed the estimated SM percentage of the profiled task to be scheduled. For example, if the task estimates 25% of SM, and a certain GPU only has 20% of available SM, then the GPU also does not meet the condition. In addition, the scheduler will also perform the shared instance capacity check to verify whether the current number of shared instances of the GPU is less than the estimated value of the maximum shareable instance number of the GPU, to ensure that the GPU still has the ability to carry new shared instances. Even if the resource amount seems sufficient, if the GPU has reached its maximum number of shared instances, it cannot allocate new tasks. Finally, the scheduler will also perform the API support check to confirm whether the GPU API required by the profiled task to be scheduled, such as CUDA, is included in the list of shared supported APIs of the GPU.
[0029] Only when a physical GPU meets all the above conditions at the same time, it is considered to be a potential candidate to carry the current task. The scheduler will add the physical node ID to which the GPU belongs to the feasible node list. In order to avoid duplication, if there are multiple GPUs on a physical node that meet the conditions, the ID of the node will also be added only once. Finally, the feasible node list will contain all the node IDs that have at least one physical GPU that meets the basic resource requirements of the task.
[0030] In step S5, based on the mapped task to be scheduled, fine-grained resource allocation scoring and selection are performed on each physical GPU node in the feasible node list to obtain a scheduling decision. The scheduling decision includes the target node ID. It should be understood that although step S4 has excluded nodes that do not meet basic resource requirements, differences still exist among the remaining feasible nodes, such as available video memory, the exact number of SMs, network latency, and current load. This step, by performing detailed resource allocation scoring on each feasible node, comprehensively considers the refined requirements of the task and the actual resource status of the node, and weighs the complex factors in a cross-datacenter environment. This scoring mechanism ensures that the scheduler can make optimal decisions, thereby achieving efficient on-demand allocation of GPU resources, maximizing overall cluster performance and resource utilization, and effectively addressing scheduling challenges in large-scale, heterogeneous environments.
[0031] Specifically, in one particular instance of this application, Figure 4 This is a flowchart of step S5 in the container-based cross-datacenter computing resource scheduling method according to an embodiment of this application. Figure 4 As shown, step S5, based on the profiled task to be scheduled, performs fine-grained resource allocation scoring and selection on each physical GPU node in the feasible node list to obtain a scheduling decision, including: S51, performing structured embedding encoding on the profiled task to be scheduled to obtain a structured embedding encoding vector for the profiled task to be scheduled; S52, performing structured embedding encoding on the shareable GPU resource reports of each physical GPU node to obtain a set of structured embedding encoding vectors for the shareable resources of physical GPU nodes; S53, calculating the fit score of the structured embedding encoding vectors for the shareable resources of physical GPU nodes in the set of structured embedding encoding vectors for the profiled task to be scheduled and the set of structured embedding encoding vectors for the shareable resources of physical GPU nodes to obtain a set of fit scores; S54, taking the ID of the physical GPU node corresponding to the largest fit score in the set of fit scores as the target node ID.
[0032] It is understandable that traditional scheduling methods often rely on simple threshold judgments or rule matching, making it difficult to capture the complex nonlinear relationships between tasks and resources, and also unable to effectively handle the fusion of different types of data (such as numerical and categorical data). Therefore, this application uses structured embedding encoding to map the refined requirements of tasks to be scheduled and the shared GPU resource reports of each physical GPU node into a high-dimensional vector space. This vectorized representation enables the scheduler to use mathematical methods to quantify the fit between tasks and resources, laying the foundation for subsequent intelligent matching and scoring based on machine learning models, thereby achieving more accurate and efficient resource allocation and overcoming the limitations of traditional methods in handling complex scheduling scenarios.
[0033] The following is a specific implementation process of steps S51 and S52: The imaged task to be scheduled contains the original information of the task and the accurate GPU resource demand obtained after imaging processing, such as the estimated video memory byte number and the estimated SM percentage, and the application type of the task. In order to convert it into a numerical vector, the scheduler will identify and process the key features in the task. For numerical features, such as the estimated video memory byte number and the estimated SM percentage, normalization processing is required. The purpose of normalization is to map these values to a standardized range, such as between 0 and 1, to eliminate the dimensional differences between different features. For example, the estimated video memory byte number can be divided by a preset maximum video memory capacity, such as the maximum video memory capacity of the GPU in the cluster, set to 48GB, and the estimated SM percentage is directly divided by 100. For categorical features, such as the application image prompt of the task, the scheduler will use one-hot encoding to convert it into a numerical vector. By concatenating all the processed numerical and categorical features, the structured embedding encoding vector of the imaged task to be scheduled is formed.
[0034] The shareable GPU resource report contains the real-time resource status of each GPU on the node. Similar to S51, the structured embedding encoding process is to convert the resource features of these physical GPU nodes into numerical vectors. The scheduler will extract key features from each shareable GPU resource report. For numerical features, such as available video memory byte number, available SM percentage, current shared instance number, and maximum shareable instance number estimate, normalization processing is also required. For example, the available video memory byte number can be divided by the total video memory capacity of the GPU, and the available SM percentage can be divided by 100. For the number of shared instances, it can be calculated by the ratio of the maximum shareable instance number estimate, resulting in a utilization rate between 0 and 1. For categorical features, such as the API list supporting sharing, such as CUDA, OpenCL, since a GPU may support multiple APIs, multi-hot encoding can be used for conversion, that is, marking 1 at the corresponding API position. For physical GPU models, one-hot encoding is used. By concatenating all the processed numerical and categorical features, the structured embedding encoding vector of the physical GPU node shareable resource is formed.
[0035] It can be understood that, after the encoding of S51 and S52, the task requirements and the node resources are uniformly represented as high-dimensional numerical vectors, which contain rich fine-grained information. In order to be able to convert these complex vector information into a single, comparable numerical value, in order to quantify the fitness or compatibility of a specific task running on a specific physical GPU node. Based on the structured embedding encoding vector of the imaged to-be-scheduled task and the structured embedding encoding vector of the physical GPU node shareable resources, the application calculates the adaptation score. That is, this score comprehensively considers the matching degree of each requirement of the task and each available resource of the node, and realizes the intelligent estimation of the resource allocation benefit. Through this quantitative evaluation, the scheduler can systematically compare all feasible nodes, so as to select the physical GPU node that can best meet the requirements of the to-be-scheduled task and maximize the resource utilization rate for the to-be-scheduled task, significantly improving the accuracy and efficiency of the scheduling decision.
[0036] More specifically, in one specific example of the present application, step S53, calculating the adaptation score of the structured embedding encoding vector of the imaged to-be-scheduled task and each physical GPU node shareable resource structured embedding encoding vector in the set of physical GPU node shareable resource structured embedding encoding vectors to obtain a set of adaptation scores, includes: S531, concatenating the structured embedding encoding vector of the imaged to-be-scheduled task and the structured embedding encoding vector of the physical GPU node shareable resources to obtain a to-be-scheduled task-physical GPU node resource joint encoding vector; S532, inputting the to-be-scheduled task-physical GPU node resource joint encoding vector into a pre-trained logistic regression model to obtain the adaptation score.
[0037] It can be understood that the task requirements and the node resources have been respectively encoded into independent, high-dimensional numerical vectors. However, to evaluate the adaptation degree between the task and the node, the model needs to consider both aspects of information and understand the interaction and matching relationship between them. Therefore, the present application fuses the features of the task and the node into a single joint encoding vector through concatenation operation, providing a complete context for the model, so that it can learn and quantify the complex matching logic between task requirements and node capabilities, thereby calculating a more accurate adaptation score.
[0038] The following is a specific implementation process of step S531: the implementation of the concatenation operation is to directly splice the two vectors in dimension. Specifically, if the structured embedding encoding vector of the imaged to-be-scheduled task is , and the structured embedding encoding vector of the physical GPU node shareable resources is , then the concatenation operation will generate a new vector, whose element order is to place all the elements of the task vector first, and then place all the elements of the node vector, that is, The output is a joint encoding vector of the to-be-scheduled task and the physical GPU node resource.
[0039] It is worth mentioning that, for the joint encoding vector of the to-be-scheduled task and the physical GPU node resource obtained by concatenating the structured embedding encoding vector of the imaged to-be-scheduled task and the structured embedding encoding vector of the physical GPU node shareable resource, when the pre-trained logistic regression model is input, since the logistic regression model performs linear combination on the input features, but the review function that determines the regression boundary finally introduces nonlinearity, therefore, it is expected to solve the nonlinearity coordination problem of the joint encoding vector of the to-be-scheduled task and the physical GPU node resource before input, so as to improve the regression performance of the logistic regression model.
[0040] Therefore, preferably, in another specific example of the present application, the structured embedding encoding vector of the imaged to-be-scheduled task and the structured embedding encoding vector of the physical GPU node shareable resource are concatenated to obtain a joint encoding vector of the to-be-scheduled task and the physical GPU node resource, comprising: The nonlinear mapping relationship between the structured embedding encoding vector of the imaged to-be-scheduled task and the structured embedding encoding vector of the physical GPU node shareable resource is established through a weight transition matrix to obtain an intra-vector nonlinear mapping matrix of the imaged to-be-scheduled task, an intra-vector nonlinear mapping matrix of the physical GPU node shareable resource, and an inter-vector nonlinear mapping relationship matrix, that is: ; wherein, is the structured embedding encoding vector of the imaged to-be-scheduled task, is the structured embedding encoding vector of the physical GPU node shareable resource, is the structured embedding nonlinear mapping encoding vector of the imaged to-be-scheduled task, is the structured embedding nonlinear mapping encoding vector of the physical GPU node shareable resource, is matrix multiplication, is the intra-vector nonlinear mapping matrix of the imaged to-be-scheduled task, is the intra-vector nonlinear mapping matrix of the physical GPU node shareable resource, is the inter-vector nonlinear mapping relationship matrix, that is, the nonlinear mapping relationship between the two is introduced at the same time of establishing the nonlinear mapping of the two themselves through the weight transition matrix.
[0041] An implicit dynamic context transition matrix between the structured embedding encoding vector of the imaged to-be-scheduled task and the structured embedding encoding vector of the physical GPU node shareable resource is calculated, that is: ; wherein, It is an implicit dynamic context transition matrix, that is, an implicit dynamic context transition matrix is introduced between the structured embedding encoding vector of the already mapped task to be scheduled and the structured embedding encoding vector of the shared resources of the physical GPU node as a prior distribution constraint.
[0042] The nonlinear mapping matrix within the vector of the mapped task to be scheduled, the nonlinear mapping matrix within the vector of the shared resources of the physical GPU nodes, and the nonlinear mapping relationship matrix between the vectors are subjected to structural association distributed alignment constraints to obtain the nonlinear structural association matrix, namely: ;in, It is subtracted based on position. It is a transpose operation. It is a non-linear structural incidence matrix, and then... and and Perform distributed alignment constraints for structural association, that is, perform distribution-based nonlinear structural association for multidimensional nonlinear representations of different orders between vectors.
[0043] Based on the cosine prior mechanism, the nonlinear structural correlation matrix and the implicit dynamic context transition matrix are similarly coupled according to structural correlation constraints to obtain the structural correlation similarity coupling matrix, namely: ;in, yes The inverse matrix, Let L be the L2 norm of the matrix. It is a structural association similarity coupling matrix, that is, based on the cosine prior mechanism, the dynamic similarity propagation of structural association constraints is carried out, that is, on the basis of association, it is further coupled with the prior to maintain the prior cooperative consistency.
[0044] Based on the structural association similarity coupling matrix, dynamic consistency aggregation is performed on the structured embedding encoding vector of the imaged task to be scheduled and the structured embedding encoding vector of the physical GPU node's shared resources to obtain the structured embedding nonlinear encoding vector of the imaged task to be scheduled and the structured embedding nonlinear encoding vector of the physical GPU node's shared resources, i.e.: ;in, It is a structured embedding of nonlinear encoded vectors for tasks that have been mapped and are awaiting scheduling. It is a structured embedded nonlinear encoded vector that can be shared by physical GPU nodes.
[0045] The imaged to-be-scheduled task structured embedding nonlinear coding vector and the physical GPU node sharable resource structured embedding nonlinear coding vector are concatenated to obtain the to-be-scheduled task-physical GPU node resource joint coding vector. In this way, by means of the nonlinear mapping relationship between vectors, the nonlinear implicitness contained in vectors in different dimensions is made explicit, and further, the nonlinear is structurally associated and aligned and dynamically propagated based on priori, so that the imaged to-be-scheduled task structured embedding coding vector and the physical GPU node sharable resource structured embedding coding vector are aggregated on a cooperative basis to obtain a nonlinear dynamic representation, thereby solving the nonlinear cooperation problem of the to-be-scheduled task-physical GPU node resource joint coding vector obtained after concatenation, and improving the regression performance of the logistic regression model. In particular, the concatenation process is the same as the above embodiment.
[0046] Correspondingly, although the features of the tasks and the nodes have been integrated, these original feature vectors cannot directly indicate the adaptation degree. Based on this, the present application inputs the to-be-scheduled task-physical GPU node resource joint coding vector into a pre-trained logistic regression model. It is worth mentioning that the logistic regression model, as a powerful prediction tool, can learn and capture the complex and nonlinear matching relationship between task requirements and node resources. Through pre-training, the model has learned from a large amount of historical scheduling data what kind of task-node combination can bring better scheduling effect. Therefore, according to these learned patterns, the matching degree between the current to-be-scheduled task and the specific physical GPU node can be intelligently evaluated, and a probability value between 0 and 1 is output, thereby providing an accurate quantitative basis for subsequent node selection and realizing fine-grained scheduling beyond simple resource matching.
[0047] The following is a specific implementation process of step S532: the scheduler inputs the to-be-scheduled task-physical GPU node resource joint coding vector into a pre-trained logistic regression model. The logistic regression model is a generalized linear model that predicts the probability of an event occurring through a linear combination, i.e., the weighted sum of the input vector plus a bias term. In this scenario, this probability is interpreted as the adaptation score between the task and the node resources. The logistic regression model receives a high-dimensional joint coding vector as the input layer, then performs a weighted sum of each dimension of the input vector through a set of learned weights, and adds a bias term. The result of this weighted sum is then converted through a Sigmoid activation function to compress the output value to between 0 and 1. This value between 0 and 1 is the final adaptation score, and the closer the score is to 1, the higher the adaptation degree of the task to the physical GPU node.
[0048] The model is pre-trained, and the weight and bias parameters inside are trained by analyzing a large amount of historical task scheduling data and task execution results. During the training process, the model learns how to predict the performance of a task running on a node according to the combination of task and node resource characteristics. For example, if historical data shows that a specific type of task performs well on a GPU with sufficient video memory but low stream processor utilization, the model will assign a higher weight to the combination of sufficient video memory and low stream processor utilization after training, so as to give a higher adaptation score when encountering similar situations. For example, if a task-physical GPU node resource joint encoding vector represents a task that requires 1.2GB of video memory and 25% of stream processor, and a GPU node with 2GB of video memory and 30% of stream processor is currently available, when this joint vector is input into the pre-trained logistic regression model, the model may output an adaptation score of 0.85, indicating that this is a highly adaptive combination.
[0049] It should be understood that although step S53 has calculated an adaptation score for each task and feasible physical GPU node combination, these scores themselves are only numerical values that measure the degree of matching. The scheduler needs a clear mechanism to select the optimal solution from these scores. Selecting the node with the highest adaptation score means that after considering the task requirements, node resource characteristics, and historical scheduling experience, the node is considered to be the most efficient physical GPU to carry the current task. This selection ensures the intelligence and optimization of the scheduling decision, avoids random or rough resource allocation, and maximizes resource utilization and task execution efficiency, which is the core embodiment of fine-grained and intelligent scheduling.
[0050] The following is a specific implementation process of step S54: After the scheduler receives the set of adaptation scores, it performs a maximum value lookup operation. The specific process is that the scheduler traverses each adaptation score in the set. During the traversal process, the scheduler maintains a record of the current maximum adaptation score and the ID of the physical GPU node corresponding to the maximum score. Initially, the maximum adaptation score can be set to a minimum value, and the target node ID can be set to empty. During the traversal process, whenever an adaptation score is encountered, the scheduler compares it with the currently recorded maximum adaptation score. If the currently traversed adaptation score is greater than the currently recorded maximum adaptation score, update the maximum adaptation score to the new score, and update the target node ID to the ID of the physical GPU node corresponding to the new score. If there are multiple physical GPU nodes with the same highest adaptation score, the scheduler will select according to pre-set rules. For example, the node with the lowest current load can be selected first, or simply the first node encountered in the set can be selected to ensure the certainty of the decision.
[0051] After traversing and comparing the entire set of adaptation scores, the ID of the physical GPU node corresponding to the maximum adaptation score recorded is determined as the target node ID for this scheduling. This target node ID represents the physical GPU node that is most suitable for carrying the profiled task to be scheduled under the current cluster resource state. For example, if the adaptation score of node A is 0.85, the adaptation score of node B is 0.78, and the adaptation score of node C is 0.91, the final target node ID determined after this step is the ID of node C.
[0052] Specifically, in one specific example of the present application, the scheduling decision further includes a target GPU ID, a virtual GPU ID, an allocated memory byte number, an allocated SM percentage, a scheduling timestamp, and a scheduler ID. It is worth mentioning that the target node ID only indicates a physical server, but a node can include multiple physical GPUs, so it is necessary to specify the target GPU ID to specify which specific GPU the task will run on. In the scenario of GPU resource sharing, the task does not exclusively occupy the entire physical GPU, but is allocated to a part of the capability, so a virtual GPU ID is needed as a unique identifier of this shared instance to facilitate subsequent resource management and isolation. The allocated memory byte number and the allocated SM percentage provide the actual amount of fine-grained resources obtained by the task, which is the direct basis for task deployment and resource billing. The scheduling timestamp and the scheduler ID provide the time point of the decision and the executor information, which are crucial for log recording, performance analysis, troubleshooting, and auditing, ensuring the transparency and traceability of the entire scheduling process.
[0053] The following is a specific implementation process of generating each value in the scheduling decision of step S5: the target GPU ID is directly derived from the decision result of step S54. In S53, the scheduler calculates an adaptation score for each physical GPU in the feasible node list. The task of S54 is to find the maximum adaptation score from these adaptation scores. The unique identifier of the physical GPU corresponding to this maximum adaptation score is the finally determined target GPU ID. For example, if GPU0 on node A obtains the highest adaptation score, the unique identifier of GPU0, such as GPU-XYZ-001, will be selected as the target GPU ID. This ID ensures that the task can be accurately deployed to the specific physical GPU that is most suitable for the cluster.
[0054] The virtual GPU ID is a logical unique identifier assigned to each scheduled task instance in a shared GPU environment. This ID does not correspond to a physical entity, but is used to distinguish different task instances running on the same physical GPU. Its generation is dynamic to ensure uniqueness. After determining the target GPU ID in S54, the scheduler generates this virtual GPU ID. A common generation method is to combine the unique identifier of the task, the current scheduling timestamp, and an incremental sequence number, or directly generate a globally unique identifier (GUID). For example, a task's virtual GPU ID may be generated as Task-ABC-202310271030_001, or a standard GUID string.
[0055] The allocated memory bytes and the allocated SM percentage are determined based on the estimated GPU resource consumption data of the profiled task to be scheduled in S3. In S53, the scheduler has determined the exact resource requirements of the task. Therefore, once the target GPU ID is determined, the scheduler directly takes the estimated memory bytes, for example, 1.2 GB, and the estimated SM percentage, for example, 25%, of the profiled task to be scheduled as the allocated memory bytes and the allocated SM percentage.
[0056] The scheduling timestamp records the exact time point when the scheduling decision is finally determined. This value is generated by the scheduler at the moment when all calculations are completed and the final scheduling result is determined, by obtaining the system time of the current server. For example, a scheduling timestamp may be a time string accurate to milliseconds, such as 2023-10-27T10:30:45.123Z.
[0057] The scheduler ID is used to identify the specific scheduler instance that executes this scheduling decision. In a distributed scheduling environment, there may be multiple scheduler instances working in parallel. Each scheduler instance is assigned a unique identifier when it starts, which can be a pre-configured string or dynamically generated. When a scheduler instance completes a scheduling decision, it will attach its own scheduler ID to the decision. For example, if the ID of the scheduler instance is Scheduler-Alpha-007, this ID will be recorded.
[0058] In step S6, the scheduler sends the scheduling decision to a node agent running on the target node, which is configured to interact with the GPU API gateway module of the target node to configure isolation parameters for the upcoming container. It is understood that the scheduler is responsible for global resource optimization and decision making, but it cannot directly configure resources and deploy containers on physical nodes. Therefore, a local agent (node agent) on the target node is needed to implement the instructions of the scheduler. This separation of responsibilities ensures that the scheduler can focus on macro resource allocation strategies, while the node agent focuses on micro resource configuration and isolation. To this end, the present application ensures that the fine-grained resource allocation decisions made by the scheduler can be accurately translated and applied to the physical GPU through the interaction of the node agent with the GPU API gateway module, providing a strict resource isolation and on-demand allocation environment for the upcoming container, thereby ensuring stable operation of the task and effective use of cluster resources.
[0059] The following is a specific implementation process of step S6: First, the scheduler obtains a scheduling decision containing the target node ID, target GPU ID, virtual GPU ID, allocated video memory bytes, allocated SM percentage, scheduling timestamp, and scheduler ID after completing step S5. The scheduler will identify the corresponding physical server according to the target node ID, and encapsulate and send the complete scheduling decision to the node agent running on the target node through a predetermined network communication protocol.
[0060] The node agent continuously listens for instructions from the scheduler. Once it receives the encapsulated scheduling decision, it immediately decapsulates and parses it to extract all the detailed parameters contained therein. These parameters are the basis for the node agent to configure resources locally.
[0061] Next, the node agent interacts with the GPU API gateway module of the target node. This GPU API gateway module is a software component running on the target node, which serves as an abstraction layer or proxy for the physical GPU hardware and its drivers. The node agent will use the target GPU ID, virtual GPU ID, allocated video memory bytes, and allocated SM percentage extracted from the scheduling decision to initiate a series of specific programming interface calls to the GPU API gateway module.
[0062] After receiving these calls, the GPU API gateway module configures the underlying physical GPU resources according to the incoming parameters, and creates an isolated resource environment for the container to be run. For example, for the allocated memory bytes, the GPU API gateway module may demarcate and reserve an exact number of bytes in the memory space of the physical GPU, and associate it with a specific virtual GPU ID, ensuring that the virtual GPU instance can only access this part of the memory. For the allocated SM percentage, if the physical GPU supports hardware virtualization technologies such as Multi-Instance GPU (MIG), the GPU API gateway module may create or configure a GPU instance with the corresponding percentage of stream processors; if it does not support hardware virtualization, it may implement a software-defined time-slice or resource scheduling strategy at the driver level to ensure that the virtual GPU instance can obtain its allocated computing power. In this way, the GPU API gateway module generates and applies a series of isolation parameters. These parameters ensure that when the container is started and attempts to access GPU resources, it can only see and use the specific memory and stream processor capabilities allocated to its virtual GPU ID, and cannot interfere with or excessively occupy the resources of other containers on the same physical GPU. These isolation parameters may include but are not limited to: creating a specific GPU device file for the container, configuring GPU resource limits for the container runtime environment, or setting resource quotas at the GPU driver level.
[0063] After successfully configuring all the isolation parameters, the node agent sends a confirmation or completion message to the scheduler, indicating that the node is ready for subsequent container startup operations.
[0064] In step S7, after receiving the resource limit configuration completion signal, the node agent starts the container on the target node. When the container is started, the GPU API calls are redirected to the GPU API gateway module. In particular, in the above steps, the node agent has interacted with the GPU API gateway module and configured strict resource isolation parameters for the container to be run. However, configuring isolation parameters alone is not enough to ensure that the container can use these limited resources as expected. Therefore, by starting the container only after the resource limit configuration is complete, the application can ensure that the container runs in a controlled and isolated environment from the beginning, avoiding errors or resource contention caused by starting when resources are not ready. At the same time, redirecting GPU API calls within the container to the GPU API gateway module is the core mechanism for implementing fine-grained resource sharing and virtualization. This allows the GPU API gateway module to intercept and manage container access to physical GPUs, enforce pre-set memory and stream processor percentage limits, and thus ensure fair allocation and efficient use of GPU resources in a multi-tenant environment.
[0065] The following is a specific implementation process of step S7: first, the node agent receives the resource limit configuration completion signal. After receiving this signal, the node agent constructs complete instructions for starting the container according to the container image, startup command, and allocated memory bytes, allocated SM percentage, and virtual GPU ID determined in the scheduling decision in the initially received to-be-scheduled fine-grained GPU task. The node agent calls the local container runtime environment, for example, Docker or Containerd, and passes these parameters.
[0066] When the container is started, the GPU API call is redirected to the GPU API gateway module. This redirection is achieved by injecting a specific configuration or library file in the runtime environment of the container. For example, the node agent may set a specific environment variable, such as LD_PRELOAD, when starting the container, which points to a shared library provided by the GPU API gateway module. This shared library acts as an intermediate layer and intercepts all standard GPU API calls issued inside the container, such as memory allocation in the CUDA programming interface, kernel function startup, etc.
[0067] When the application inside the container attempts to perform GPU operations, its GPU API calls will not directly reach the physical GPU driver, but will first be captured by this injected shared library. The shared library then forwards these calls to the GPU API gateway module running on the target node. The GPU API gateway module acts as a proxy for the physical GPU and knows the exact resource quota for each virtual GPU ID, i.e., the allocated memory bytes and the allocated SM percentage.
[0068] After receiving the GPU API call of the container, the GPU API gateway module performs permission and resource checks. For example, if the container requests memory allocation, the GPU API gateway module checks whether the requested memory amount exceeds the allocated memory bytes for its virtual GPU ID. If the container attempts to start a computing core, the GPU API gateway module manages its access to streaming processors according to the allocated SM percentage, possibly through time slicing or hardware virtualization to ensure resource isolation. Only when the request meets the preset resource limit, the GPU API gateway module forwards the request to the underlying physical GPU driver for actual operation. In this way, even if multiple shared container instances are running on the same physical GPU, effective resource isolation and performance guarantee can be achieved between them, thereby realizing fine-grained sharing and efficient utilization of GPU resources.
[0069] In summary, the container-based cross-data center computing resource scheduling method based on the embodiments of the present application is illustrated, which first requires each physical GPU node to report its shareable fine-grained GPU resource report, thereby overcoming the problem of insufficient internal resource awareness of the traditional scheduler. Then, for the user-submitted tasks to be scheduled, the estimated resource consumption data is loaded or generated from the task portrait database, solving the pain points of the scheduler unable to accurately know the real demand of the task and the user's excessive application leading to resource waste. Most importantly, the scheduler no longer simply filters available nodes, but according to the imaged task demand, each physical GPU node in the feasible node list is scored and selected for detailed resource allocation, and the optimal target node is determined by calculating the adaptation score. This intelligent matching mechanism, combined with the node agent program and the GPU API gateway module to configure the isolation parameters for the container and redirect the GPU API calls, effectively solves the problems of traditional scheduling logic being extensive, low GPU utilization, and difficulty in realizing multi-tenant fine-grained sharing, and finally realizes the fine and efficient scheduling of cross-data center computing resources.
[0070] Embodiments of the disclosure have been described above, the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes are obvious to those skilled in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles, practical applications, or improvements to the technology in the market of the embodiments, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A container-based cross-data-center computing resource scheduling method, characterized in that, The application comprises the following steps: reporting the shareable GPU resource of each physical GPU node to a scheduler, the scheduler collecting the shareable GPU resource of each physical GPU node to obtain a cluster shareable GPU resource snapshot; obtaining a to-be-scheduled fine-grained GPU task submitted by a user; loading or generating an imaged to-be-scheduled task from a task imaging database based on the to-be-scheduled fine-grained GPU task; the scheduler screening a feasible node list from the cluster shareable GPU resource snapshot based on the imaged to-be-scheduled task; performing fine-grained resource allocation scoring and selection on each physical GPU node in the feasible node list based on the imaged to-be-scheduled task to obtain a scheduling decision, the scheduling decision comprising a target node ID; the scheduler sending the scheduling decision to a node agent program running on the target node, the node agent program being used to interact with a GPU API gateway module of the target node to configure isolation parameters for a to-be-run container; after receiving a resource limit configuration completion signal, the node agent program starts a container on the target node, and when the container is started, GPU API calls are redirected to the GPU API gateway module.
2. The container-based cross-data center computing resource scheduling method according to claim 1, characterized in that, The shareable GPU resource report comprises a GPU ID, physical GPU information, available video memory byte number, available SM percentage, available SM absolute number, current shared instance number, maximum shareable instance number estimation value, API list supporting sharing, currently activated sharing strategy, and last update timestamp.
3. The container-based cross-data center computing resource scheduling method of claim 1, wherein, The method comprises the following steps: extracting an application imaging prompt from the to-be-scheduled fine-grained GPU task; loading an image matched with the query key from the task imaging database by taking the application imaging prompt as a query key; generating estimated GPU resource consumption data based on data in the image matched with the query key; attaching the estimated GPU resource consumption data to the to-be-scheduled fine-grained GPU task to obtain the imaged to-be-scheduled task.
4. The container-based cross-data center computing resource scheduling method of claim 3, wherein, If no image matched with the query key is found in the task imaging database, the scheduler generates the imaged to-be-scheduled task by taking a Min value in the to-be-scheduled fine-grained GPU task.
5. The container-based cross-datacenter compute resource scheduling method of claim 1, wherein, The method comprises the following steps: performing structured embedding coding on the imaged to-be-scheduled task to obtain an imaged to-be-scheduled task structured embedding coding vector; performing structured embedding coding on the shareable GPU resource report of each physical GPU node to obtain a set of physical GPU node shareable resource structured embedding coding vectors; calculating an adaptation score of each physical GPU node sharable resource structured embedding encoding vector in the set of the imaged to-be-scheduled task structured embedding encoding vector and the physical GPU node sharable resource structured embedding encoding vector to obtain a set of adaptation scores; taking an ID of a physical GPU node corresponding to a maximum adaptation score in the set of adaptation scores as the target node ID.
6. The container-based cross-datacenter computing resource scheduling method of claim 5, wherein, calculating an adaptation score of each physical GPU node sharable resource structured embedding encoding vector in the set of the imaged to-be-scheduled task structured embedding encoding vector and the physical GPU node sharable resource structured embedding encoding vector to obtain a set of adaptation scores, comprising: concatenating the imaged to-be-scheduled task structured embedding encoding vector and the physical GPU node sharable resource structured embedding encoding vector to obtain a to-be-scheduled task-physical GPU node resource joint encoding vector; inputting the to-be-scheduled task-physical GPU node resource joint encoding vector into a pre-trained logistic regression model to obtain the adaptation score.
7. The container-based cross-datacenter computing resource scheduling method of claim 6, wherein, concatenating the imaged to-be-scheduled task structured embedding encoding vector and the physical GPU node sharable resource structured embedding encoding vector to obtain a to-be-scheduled task-physical GPU node resource joint encoding vector, comprising: establishing a non-linear mapping relationship between the imaged to-be-scheduled task structured embedding encoding vector and the physical GPU node sharable resource structured embedding encoding vector through a weight transition matrix to obtain an imaged to-be-scheduled task vector intra-non-linear mapping matrix, a physical GPU node sharable resource vector intra-non-linear mapping matrix, and a vector inter-non-linear mapping relationship matrix; calculating an implicit dynamic context transition matrix between the imaged to-be-scheduled task structured embedding encoding vector and the physical GPU node sharable resource structured embedding encoding vector; performing structural correlation distributed alignment constraint on the imaged to-be-scheduled task vector intra-non-linear mapping matrix, the physical GPU node sharable resource vector intra-non-linear mapping matrix, and the vector inter-non-linear mapping relationship matrix to obtain a non-linear structural correlation matrix; performing structural correlation constraint similar coupling on the non-linear structural correlation matrix and the implicit dynamic context transition matrix based on a cosine prior mechanism to obtain a structural correlation similar coupling matrix; based on the structural correlation similar coupling matrix, performing dynamic consistency aggregation on the imaged to-be-scheduled task structured embedding encoding vector and the physical GPU node sharable resource structured embedding encoding vector respectively to obtain an imaged to-be-scheduled task structured embedding non-linear encoding vector and a physical GPU node sharable resource structured embedding non-linear encoding vector; performing concatenation processing on the imaged to-be-scheduled task structured embedding non-linear encoding vector and the physical GPU node sharable resource structured embedding non-linear encoding vector to obtain the to-be-scheduled task-physical GPU node resource joint encoding vector. 8.The container-based cross-data-center computing resource scheduling method of claim 1, wherein, The scheduling decision further includes a target GPU ID, a virtual GPU ID, an allocated video memory byte number, an allocated SM percentage, a scheduling timestamp, and a scheduler ID.
Citation Information
Cited By
Graphics card task resource scheduling method and system based on machine learning
CN121722574A
AI server management system based on GPU computing power optimization
CN121785788A
Model update resource scheduling method, system and product
CN121934983A
A model updating resource scheduling method, system and product
CN121934983B