Heterogeneous GPU cluster scheduling optimization method and system based on resource affinity
By building a resource affinity scheduling method in a heterogeneous GPU cluster, using job feature matching and probe testing, establishing an affinity evaluation model, and optimizing resource allocation, the problems of low resource utilization and low scheduling efficiency in a heterogeneous GPU cluster are solved, efficient resource utilization and dynamic adaptation are achieved, and overall performance and scalability are improved.
Patent Information
- Application Number
- CN202510733628.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
In heterogeneous GPU clusters, existing scheduling algorithms fail to effectively capture performance differences between heterogeneous GPUs, resulting in low resource utilization, high scheduling overhead, poor load balancing, and difficult to achieve efficient resource awareness and intelligent scheduling.
By building a heterogeneous GPU cluster scheduling method based on resource affinity, using job feature matching and probe testing, an affinity evaluation model is established, resource allocation is optimized, and precise matching and dynamic adaptation between jobs and GPU resources are achieved.
Improves the overall performance and resource utilization of heterogeneous GPU clusters, reduces scheduling complexity, enhances system scalability and scheduling efficiency, and ensures optimal performance under various workloads.
Smart Images

Figure CN120256139A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer resource scheduling, particularly to GPU resource scheduling technology in deep learning model training. Specifically, the present invention proposes an optimization method and system for heterogeneous GPU cluster scheduling based on resource affinity. Background Art
[0002] In recent years, the rapid development of deep learning technology has given rise to a huge demand for high-performance computing resources. The GPU (Graphics Processing Unit), with its powerful parallel computing capabilities, has become the core hardware for accelerating deep learning model training. However, as the complexity and scale of deep learning training tasks increase, a single type of GPU (homogeneous cluster) has become difficult to meet the performance requirements. As a computing platform that includes multiple types of GPU hardware (such as NVIDIA V100, A100, T4, etc.), heterogeneous GPU clusters can provide higher flexibility and computing efficiency for deep learning model training, but also bring challenges in aspects such as resource heterogeneity awareness, task allocation, and scheduling efficiency.
[0003] Traditional scheduling algorithms often assume that GPUs in the cluster have similar performance characteristics and do not fully consider the heterogeneity between GPUs. This leads to problems such as low resource utilization and low task execution efficiency, making it difficult to achieve the best performance in heterogeneous environments. Specific problems include:
[0004] 1) Insufficient resource awareness: Unable to effectively capture the performance differences between heterogeneous GPUs, resulting in low resource utilization.
[0005] 2) High scheduling overhead: In high-dynamic load scenarios, existing scheduling algorithms are difficult to quickly respond to changes in resource status.
[0006] 3) Poor load balancing: Training jobs do not fully match the GPU performance, which may cause some GPUs to be overloaded while others are idle.
[0007] Therefore, how to achieve efficient resource awareness and intelligent scheduling in heterogeneous GPU clusters has become a key technical problem to be solved urgently. Summary of the Invention
[0008] In view of the deficiencies of the prior art, the present invention provides an optimization method and system for heterogeneous GPU cluster scheduling based on resource affinity; The core objective of the present invention is to design an innovative heterogeneous GPU cluster resource affinity scheduling method and system. By deeply analyzing the heterogeneous characteristics of GPUs, carefully constructing a precise and efficient resource affinity scheduling algorithm, ingeniously architecting a hierarchical collaborative scheduling framework, and organically integrating a resource scheduling optimization mechanism, it realizes the highly intelligent adaptation and efficient utilization of training tasks and GPU resources, significantly improves the overall performance, resource utilization rate, and energy efficiency ratio of the cluster, and provides strong support for the training of complex models.
[0009] The technical solution of the present invention is as follows: A heterogeneous GPU cluster scheduling optimization method based on resource affinity, including: Step 1: According to the affinity history library of historical jobs, estimate the resource requirements by matching historical jobs with relatively high similarity in job characteristics. Step 2: For a new job, receive the job and conduct probe tests on each type of GPU to obtain and record the job performance data. Step 3: Use historical information or test data to establish an affinity evaluation model for each job, evaluate the affinity values of the job under different resource configurations, and provide data support for resource allocation decisions. Step 4: Enter the scheduling optimization cycle loop, construct an allocation matrix according to the objective function of the optimization problem, and allocate corresponding resources to each job.
[0010] Specifically, a heterogeneous GPU cluster scheduling optimization method based on resource affinity, including: Step S101: Record historical jobs; for a cluster with accumulated historical jobs, pre-collect historical job information of the cluster for a period of time; including: job characteristic information, execution performance data on different GPUs, and the final resource allocation plan, to form a historical job library. The recorded job characteristic information is used to match with the jobs newly submitted by users and is recorded in a numerical representation; while the execution performance data and the resource allocation plan provide a basis for the preliminary resource allocation of similar jobs; at the same time, the historical jobs are classified according to the job characteristic information and the resource allocation plan to form an affinity history library. Step S102: Receive jobs; receive the jobs newly submitted by users. Step S103: Job characteristic matching; after receiving the jobs newly submitted by users, immediately start the job characteristic matching process, and compare the characteristic information of the jobs newly submitted by users with the jobs in the affinity history library. Step S104: Preliminary resource allocation for similar jobs; After finding the job with the highest similarity, based on the execution information and resource configuration corresponding to the job with the highest similarity, preliminarily estimate the resource requirements of the current job; List the resource configuration of the job with the highest similarity as potential matching resources, and perform preliminary resource allocation for the current job; At the same time, obtain the historical affinity data of the job with the highest similarity; Step S105: Start probe testing; Step S106: Scheduling optimization loop; After determining the matching relationship between the newly submitted job and the GPU, the job enters the scheduling cycle loop to continuously optimize resource allocation.
[0011] According to the preference of the present invention, job feature matching includes: Step S1031: Extract job features; Extract feature information from the received new job. The job feature information includes: model features, data features, load characteristics, and accuracy requirements. The data features include data type and data scale; Step S1032: Traverse the historical job library; Query the recorded historical jobs in the cluster, and execute Step S1033 until the historical job most similar to the received new job is found; Step S1033: Similarity matching; Based on the vectorized representation of features, convert the job features extracted in Step S1031 into numerical representations, and combine each job feature into a vector representation; The model features use one-hot encoding to represent the model type, the data features are numerically represented, the load characteristics are classified by category encoding, and the accuracy requirements are mapped to numerical values for scalar representation; Calculate the similarity between the current job and the historical job; Use weighted similarity, weight each feature according to the importance of the feature, calculate the similarity score of each job after weighting, and find the job with the highest similarity; Step S1034: Determine whether the job with the highest similarity exists; After calculating the similarity score, set a similarity threshold to determine whether the job with the highest similarity exists; If the similarity between the current job and the historical job exceeds the similarity threshold, it is considered that the job with the highest similarity exists, otherwise it does not exist; If the job with the highest similarity exists, then load the cache and obtain the resource allocation plan of the job most similar to the current job from the historical job library as a reference for preliminary resource allocation, which means: directly give the resource allocation plan of the job most similar to the current job to the current job; If there is no similar job, then regard this job as an unknown job, start probe testing and record the actual execution efficiency of the job, achieve preliminary resource matching, and provide data support for affinity matching.
[0012] More preferably, in step S1031, for model features, obtain the L0 model according to the job submission script or the model definition of the framework, and determine the general category of the job according to the basic model architecture information; Judge its load characteristics according to the analysis job configuration information; The accuracy requirement is obtained according to the settings in the model framework.
[0013] More preferably, use weighted similarity to weight each feature according to the importance of the feature, calculate the similarity score of each job after weighting, and find the job with the highest similarity; including: For model features, calculate the cosine similarity ; For data types, calculate the cosine similarity ; For data scale, calculate the Euclidean distance ; For load characteristics, calculate the Jaccard similarity ; For accuracy requirements, calculate the Euclidean distance ; Comprehensive similarity calculation: Weight and sum the similarities of each feature according to the preset weights to obtain the total similarity score: .
[0014] More preferably, in step S105, start the probe test; including: Step S1051: Receive the unknown job; Step S1052: Conduct performance tests; Deploy the unknown job to different types of GPUs in the cluster for small-scale trial runs respectively. Multiple different combinations of running parameters are set for each GPU type during the test; in the scenario of deep learning training jobs, the parameters include batch size and gradient accumulation steps; under each parameter combination including batch size and gradient accumulation steps, strictly monitor and collect the execution performance data of the job; including job computing power requirements, video memory occupancy, data processed, data transfer time, job running time, actual throughput, GPU resource utilization rate, running power consumption.
[0015] Step S1053: Record the performance data, that is, the execution performance data; Step S1054: Determine the preliminary match; calculate the actual execution efficiency of the job according to the performance data recorded in step S1053; Actual execution efficiency The calculation formula is as follows: ; Among them, represents the job throughput, represents the normalized throughput of the job; represents the comprehensive utilization rate of the GPU; is the weight coefficient to balance the performance index; Sort the actual execution performance of unknown jobs on different GPUs from high to low, compare and screen, take the maximum execution performance as the optimal value, and take the selected optimal GPU resources as the initial resource matching scheme for unknown jobs.
[0016] According to the preference of the present invention, in step S106, the scheduling optimization loop includes: Step S1061: Initial allocation exploration; For unknown jobs, at the initial stage of scheduling, it is assumed that the throughput of unknown jobs increases linearly with the number of resources; that is: ; Among them, represents the throughput of the job under the configuration including the number of GPUs as n, the type of GPU as t, the batch size as 𝑚, and the gradient accumulation step as 𝑠; Step S1062: Throughput model estimation; Collect the actual throughput data of the job under different numbers of GPUs and types , and calculate the gradient calculation time and the gradient synchronization time according to the model fitting parameters : Among them, represents the fixed time overhead independent of the model scale during the gradient calculation; represents the proportional coefficient of the gradient calculation time increasing linearly with the model scale ; represents the fixed time overhead independent of the number of GPUs during the gradient synchronization; represents the proportional coefficient of the gradient synchronization time increasing with the number of GPUs ; Gradient calculation time: ; Gradient synchronization time: ; Calculate the time required for each iteration, including 𝑠 times of gradient calculation time and one synchronization time ; Iteration time : ; Construct a throughput model for each job; The throughput model formula is as follows: ; Among them, represents the total batch size; According to the ratio of the effective calculation time to the total iteration time, the actual utilization rate of the modeling resources : ; Step S1063: Affinity model evaluation; Define a set of resources as a resource configuration binary tuple , representing pieces ; First, by quantifying the matching degree between the floating-point operation ability of the GPU and the actual requirements of the job, obtain the computing power matching degree. The computing power matching degree includes two sub-items: basic computing power matching and architecture adaptation correction: ; Among them, represents the amount of floating-point operations required for a single iteration of the task; represents the theoretical peak computing power of the GPU; represents the architecture characteristics required by the task; represents the actual architecture of the GPU; represents the maximum architecture generation difference; As a sensitivity coefficient, represents architecture sensitivity; For the basic computing power matching item , when , , when , ; Quantify the utilization efficiency of the GPU video memory resources by the job, and obtain the video memory adaptation degree. This video memory adaptation degree covers two aspects: bandwidth utilization rate and capacity adaptation degree: ; Among them, represents the total amount of video memory operations per batch; represents the effective bandwidth of the GPU video memory; represents the processing time per batch; represents the peak video memory demand of the task; represents the available video memory after deducting the system reservation; By calculating the energy consumed per unit of computing, quantify the energy utilization efficiency, so as to obtain the energy efficiency optimization degree : ; Among them, Indicates the actual number of operands of the GPU in the job); Indicates the effective computing time of the GPU; Indicates the total energy consumption of the task; Quantify the effective production capacity under the actual workload based on the measured data of job operation and the estimated data of the throughput model, and then obtain the execution efficiency degree : ; Among them, the throughput As the core indicator to measure the execution efficiency of the job, it reflects the amount of tasks completed per unit time, and the resource utilization correction factor , measures the actual utilization rate of resources through the ratio of the effective computing time to the total iteration time; Comprehensively consider the hardware procurement or rental cost and the operating cost , evaluate the economic cost of different types and quantities of GPUs : ; For each job, construct an affinity evaluation model of the job and GPU resources in different configuration scenarios, and quantitatively evaluate the adaptability between the job tasks and GPU resources: ; Among them, Indicates the job and the resource configuration scheme 's multi-dimensional affinity value; Is the normalization process based on the Sigmoid function; As the dynamic weight coefficient; According to the evaluation data of the affinity evaluation model, construct an affinity matrix , record the affinity of different jobs under different resource configuration schemes , Indicates the job under the configuration 's affinity value; Step S1064: Global scheduling optimization; Specifically, during the implementation of the global resource scheduling, run the search algorithm according to the predetermined scheduling period to generate a series of binary allocation matrices, each matrix represents a possible resource allocation scheme, defined as the candidate allocation matrix , is the number of jobs, is the number of resource configuration schemes, and the element in the matrix is used to indicate whether the configuration is selected for the job ; For the candidate allocation matrix , the following constraint conditions should be satisfied: Allocation uniqueness: At most one 1 in each row, and each job is allocated one configuration; that is: ; Resource capacity limit: The sum of the resource configuration requirements of all allocated resources ≤ the total physical resources of the cluster; The goal of global scheduling is to achieve global optimal scheduling, that is, to obtain the optimal affinity allocation matrix , represents the basic affinity gain obtained by job ; The objective function of the optimization problem is: ; Among them, represents the basic affinity gain obtained by job ; As a fairness adjustment factor, the generalized power mean is used to achieve multi-objective balance; represents the reallocation penalty factor; Set the reallocation penalty factor , and increase the reallocation overhead based on the historical reallocation frequency of the job, which is designed as an exponential decay function : ; Among them, represents the historical reallocation times of job ; represents the time interval between the last two reallocations; represents the average running period of job ; represents the penalty intensity coefficient; Global scheduling optimization transforms the scheduling problem into an optimization problem by maximizing the sum of the affinity values of jobs in the cluster. The objective function is: ; Among them, As a fairness adjustment factor, the generalized power mean is used to achieve multi-objective balance: when , it degenerates to linear summation, pure efficiency priority, and makes resources tilt towards jobs with higher affinity; when , it approaches the geometric mean, fairness priority, and emphasizes the fairness of resource allocation; , considering max-min fairness; Step S1065: Job-level scheduling optimization; The optimization goal is to find the optimal batch size and the gradient accumulation steps : ; Step S1066: Job execution; Step S1067: Parameter fitting; During the job execution, continuously collect data Perform parameter fitting on the throughput model; According to the throughput data recorded in Step S1065, perform model parameter fitting by applying the L-BFGS-B optimization algorithm; Use the root mean square logarithmic error as the loss function : ; where n is the number of samples, is the actual measurement time, is the predicted iteration time; By minimizing this loss function, adjust the model parameters so that the predicted iteration time is closer to the actual time; Based on the optimized throughput model, provide solid data support for resource allocation decisions, thus forming a scheduling optimization loop.
[0017] A heterogeneous GPU cluster scheduling system based on resource affinity, comprising: A similar job matching module, configured to: according to the affinity history library of historical jobs, estimate resource requirements by matching historical jobs with higher similarity in job characteristics; A probe test module, configured to: for a new job, receive the job and perform probe tests on each type of GPU, obtain and record job performance data; An affinity evaluation module, configured to: use historical information or test data to establish an affinity evaluation model for each job, evaluate the affinity values of the job under different resource configurations, and provide data support for resource allocation decisions; A job scheduling module, configured to: enter the scheduling optimization cycle loop, construct an allocation matrix according to the objective function of the optimization problem, and allocate corresponding resources to each job.
[0018] According to a preferred embodiment of the present invention, the heterogeneous GPU cluster scheduling system further includes a performance monitoring module, configured to: construct a full-stack monitoring system, collect multi-dimensional metrics during job execution in real time, detect abnormal fluctuations and generate warning events.
[0019] Compared with the prior art, the beneficial effects of the present invention are: 1) Precise resource matching and data support: Through the affinity matching model between jobs and GPU resources, comprehensively considering job characteristics and GPU performance, the best match between jobs and hardware resources is achieved, making full use of the heterogeneous characteristics of the GPU cluster. The throughput model can estimate the performance data of tasks under different GPU configurations, providing more scientific and comprehensive data support for affinity-based resource allocation, effectively avoiding the blindness and inefficiency of resource allocation.
[0020] 2) Guarantee of dynamic adaptability: The periodic scheduling mechanism can adjust the resource allocation of jobs in real time according to the dynamic changes in task load, ensuring that the system always maintains the best performance under various workloads. By introducing a scheduling cycle loop, during the execution of jobs, the affinity function is evaluated in real time, and the best resource allocation matrix is generated in a timely manner to dynamically adjust the resource allocation and ensure the optimal overall performance of the cluster.
[0021] 3) Enhancement of system scalability: The hierarchical collaborative scheduling framework combines global scheduling and job-level scheduling, reducing the complexity in the scheduling process, significantly alleviating the performance bottleneck and overhead of centralized scheduling. Optimize the resource allocation between jobs at the global level and finely adjust task parameters at the job level to effectively decompose the scheduling responsibilities, improve the system scalability and scheduling efficiency, and better meet the requirements of large-scale heterogeneous GPU clusters. Brief Description of the Drawings
[0022] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation to the invention.
[0023] Figure 1 It is the overall flowchart of the heterogeneous GPU cluster scheduling optimization method based on resource affinity provided for Embodiment 2 of the present invention; Figure 2 It is the schematic diagram of the historical job matching process based on similarity provided for Embodiment 2 of the present invention; Figure 3 It is the schematic diagram of the actual efficiency probe test process of jobs provided for Embodiment 2 of the present invention; Figure 4 It is the flowchart of job periodic scheduling optimization based on a hierarchical architecture provided for Embodiment 2 of the present invention; Figure 5 It is the overall architecture diagram of the heterogeneous GPU cluster scheduling system based on resource affinity provided for Embodiment 3 of the present invention. Detailed Embodiments
[0024] The present invention will be further limited below in conjunction with the accompanying drawings and embodiments, but not limited thereto.
[0025] Embodiment 1 A heterogeneous GPU cluster scheduling optimization method based on resource affinity, including: Step 1: According to the affinity history library of historical jobs, estimate the resource requirements by matching historical jobs with relatively high similarity through job characteristics. Step 2: For a new job, receive the job and perform probe tests on each type of GPU to obtain and record job performance data. Step 3: Use historical information or test data to establish an affinity evaluation model for each job, evaluate the affinity values of the job under different resource configurations, and provide data support for resource allocation decisions. Step 4: Enter the scheduling optimization cycle loop, construct an allocation matrix according to the objective function of the optimization problem, and allocate corresponding resources to each job.
[0026] Embodiment 2 A heterogeneous GPU cluster scheduling optimization method according to Embodiment 1, characterized in that: A heterogeneous GPU cluster scheduling optimization method based on resource affinity, referring to Figure 1 , including: Step S101: Record historical jobs; for a cluster with a certain accumulation of historical jobs, pre-collect historical job information of the cluster for a period of time; various information of past jobs is detailedly recorded in the historical job library, including: job characteristic information (job characteristic information includes model characteristics, data characteristics, load characteristics, and accuracy requirements), execution performance data on different GPUs (such as throughput, execution time, resource utilization rate, and operating power consumption), and the final resource allocation scheme, forming a historical job library. The recorded job characteristic information is used to match with the jobs newly submitted by users and is recorded in the form of numerical representation; while the execution performance data and the resource allocation scheme provide a basis for the initial resource allocation of similar jobs; at the same time, the historical jobs are classified according to the job characteristic information and the resource allocation scheme to form an affinity history library; the classification is to classify jobs according to job characteristic information; the affinity history library records the resource allocation scheme of jobs so that when subsequent similar jobs are scheduled, initial resource allocation can be quickly performed.
[0027] Step S102: Receive jobs; receive the jobs newly submitted by users; this is the beginning of the actual scheduling process.
[0028] Step S103: Job characteristic matching; after receiving the jobs newly submitted by users, immediately start the job characteristic matching process, and compare the characteristic information of the jobs newly submitted by users with the jobs in the affinity history library. Step S104: Preliminary resource allocation for similar jobs; After finding the job with the highest similarity, based on the execution information and resource configuration corresponding to the job with the highest similarity, preliminarily estimate the resource requirements of the current job; List the resource configuration of the job with the highest similarity as potential matching resources, and perform preliminary resource allocation for the current job; At the same time, obtain the historical affinity data of the job with the highest similarity; As an important reference basis for subsequent accurate resource allocation and optimization.
[0029] Step S105: Start probe testing; For new jobs that cannot find sufficiently similar jobs in the affinity history library, conduct a comprehensive actual execution performance testing mechanism.
[0030] Step S106: Scheduling optimization loop; After determining the matching relationship between the newly submitted job and the GPU, the job enters the scheduling cycle loop to continuously optimize resource allocation.
[0031] As Figure 2 shown, job feature matching; includes: Step S1031: Extract job features; Extract feature information from the received new job. The job feature information includes: model features, data features, load characteristics, and precision requirements. The data features include data type and data scale (such as the number of samples, the total data volume size); Model features: represent the type of model architecture used by the job. According to the model features, its L0 model (i.e., the basic model) can be known, so that the general category of the job can be determined (such as image processing, natural language processing); Data features: represent the data type and data scale used by the job; Load characteristics: the inherent characteristics of deep learning job loads, indicating the job's requirements for computing power, video memory resources, and communication capabilities; Precision requirements: represent the precision requirements of the job in numerical calculations; Can be obtained by parsing the job configuration file.
[0032] Step S1032: Traverse the historical job library; Query the recorded historical jobs in the cluster and execute Step S1033 until the historical job most similar to the received new job is found; Step S1033: Similarity matching; Based on the vectorized representation of features, convert the job features extracted in Step S1031 into numerical representations, and combine each job feature into a vector representation; Model features are represented using one-hot encoding for model types, data features are numerically represented, load characteristics are classified by category encoding, and precision requirements are mapped to numerical values for scalar representation; The model features use one-hot encoding to represent the model type, which means: using One-Hot encoding. First, list all possible model categories, then assign a unique index to each model type, and finally create a One-Hot vector. In the vector, the dimension corresponding to the job model type takes the value of 1, and the remaining dimensions take the value of 0. For example, if there are a total of 3 model types, namely CNN, RNN, and Transformer, and the model type of a certain job is RNN, then the model feature encoding of this job is [0, 1, 0].
[0033] The numerical representation of data features means: (1) One-Hot encoding of data types: List all possible data types, assign a unique index to each data type, and perform One-Hot encoding. (2) Data scale encoding: Collect the data scale information (number of samples, total data volume) of the job, perform min-max normalization on the data scale, and map the values to the interval [0, 1].
[0034] The load characteristics are classified by category encoding, which means: using multi-label encoding. Define the load types (computation-intensive, video memory-intensive, communication-intensive, I / O-intensive), assign a unique index to each load type, and perform multi-label encoding. Set the position corresponding to the load type to 1, and the rest to 0.
[0035] The precision requirement is mapped to a numerical value for scalar representation, which means: using numerical mapping. Assign a numerical value to each precision level (FP16 → 0.5, FP32 → 1.0, FP64 → 2.0), and map the precision requirement of the job to the corresponding numerical value.
[0036] Calculate the similarity between the current job and historical jobs; use weighted similarity, weight each feature according to its importance, calculate the similarity score of each job after weighting, and find the job with the highest similarity; Step S1034: Determine whether there is a job with the highest similarity; after calculating the similarity score, set a similarity threshold. Here, the threshold setting determines the degree of definition of "similar" in the matching process. The specific threshold should be adjusted according to actual needs. High-precision matching: The threshold is set above 0.9; Medium-precision matching: The threshold is set between 0.7 and 0.9; Low-precision matching: The threshold is set between 0.5 and 0.7. It is used to determine whether there is a job with the highest similarity; if the similarity between the current job and historical jobs exceeds this similarity threshold, it is considered that there is a job with the highest similarity, otherwise there is not;
[0037] If there is a job with the highest similarity, then load the cache , obtaining the resource allocation plan of the job with the highest similarity to the current job from the historical job library as a reference for preliminary resource allocation means: directly giving the resource allocation plan of the job with the highest similarity to the current job to the current job; If there is no similar job, the job is regarded as an unknown job, and probe testing is started and the actual execution performance of the job is recorded to achieve preliminary resource matching and provide data support for affinity matching. Probe testing (refer to Figure 3 , steps S1051 - S1054), that is, running tests on the unknown job under different resource configurations, and obtaining and recording the actual running data of the job (the running data are all data required for affinity evaluation, including: job computing power demand, video memory occupancy, amount of data processed, data transmission time, job running time, actual throughput, GPU resource utilization rate, running power consumption).
[0038] In step S1031, for model features, obtain the L0 model according to the model definition in the job submission script or framework, and determine the general category of the job according to the basic model architecture information; The model feature extraction process is as follows: (1) Parse the job submission script to identify the used framework (such as TensorFlow, PyTorch) and model definition file; (2) According to the model definition file, extract the architecture information of the model and identify the L0 model (i.e., the basic model architecture); (3) Map the extracted L0 model to the predefined job category.
[0039] The L0 model refers to the basic model architecture used by the job, that is, the bottom - most structure or original architecture of the model. For example, in image processing tasks, the possible L0 models include ResNet, VGG, EfficientNet, etc.; in natural language processing tasks, the possible L0 models include BERT, GPT, Transformer, etc.
[0040] Based on the model features, the general category of the job is determined.
[0041] According to the analysis of the job configuration information, judge its load characteristics; The load characteristics refer to the degree of demand of deep - learning jobs for computing resources, including computing power, video memory resources, communication ability, and I / O ability. The method for extracting load characteristics is based on the configuration content of the job (job type, used framework, model structure), and the specific implementation steps:
[0042] (1) Parse the job submission script, configuration file, or model definition to extract relevant information; (2) Identify the key features of the job, including the type of model used, data scale, and training method; (3) Map the identified features to the corresponding load types according to predefined rules; (4) Since a job may have multiple load characteristics simultaneously, a multi-label encoding method is used for numerical representation.
[0043] Specific instructions: (1) Computation-intensive: Use complex model structures (such as deep neural networks) and a large number of mathematical operations. Judge by analyzing the depth and complexity of the model structure. Example: The job uses the ResNet-152 model for image classification, and the main calculations during training are forward propagation and backpropagation.
[0044] (2) VRAM-intensive: The model has a large number of parameters, high input data dimensions, and a large batch size. Judge by analyzing the number of model parameters and input data dimensions. Example: The job uses the GPT-3 model for text generation, and the long input sequence length results in high VRAM requirements.
[0045] (3) Communication-intensive: Adopt a distributed training method, frequently perform parameter synchronization and gradient exchange, and are sensitive to network bandwidth and latency. Judge by analyzing whether the job uses a distributed training framework. Example: The job uses the Horovod framework for multi-node distributed training and needs to frequently perform AllReduce operations.
[0046] (4) I / O-intensive: Frequently read a large number of data files, perform network communication, or access databases. Judge by analyzing the data loading method and data scale. Example: The job uses a large dataset and needs to frequently read a large-scale image dataset from disk during training, resulting in I / O becoming a bottleneck.
[0047] The accuracy requirement is obtained according to the settings in the model framework.
[0048] Extracting the numerical calculation accuracy requirement of the job is achieved by analyzing the settings in the model framework. The specific implementation process:
[0049] (1) Parse the job submission script or model definition file to identify the deep learning framework used (such as TensorFlow or PyTorch); (2) According to the specific API of the framework, identify the job accuracy strategy and determine the numerical calculation accuracy requirement of the job, such as FP16, FP32, FP64; (3) Map to a numerical representation, assign a numerical value to each accuracy level (FP16 → 0.5, FP32 → 1.0, FP64 → 2.0), and map the identified accuracy requirement to a numerical representation for subsequent similarity calculation.
[0050] Using weighted similarity, each feature is weighted according to its importance, the similarity scores of each job after weighting are calculated, and the job with the highest similarity is found; including: When calculating the similarity between the current job and historical jobs, for each job feature, after converting to a numerical representation, a suitable similarity measurement method is adopted: for one-hot encoding, cosine similarity is used for measurement, for multi-label encoding, Jaccard similarity is used for measurement, the numerical feature mapping is normalized, and the Euclidean distance is used to measure the difference.
[0051] For model features (one-hot encoding), calculate the cosine similarity ; A and B are the model feature vectors of the current job and historical jobs respectively: .
[0052] Example: Assume there are three types of model types: ResNet, BERT, and Transformer. The current job uses BERT, and the one-hot encoding is [0, 1, 0]; the historical job uses ResNet, and the one-hot encoding is [1, 0, 0]. Cosine similarity:
[0053] ; For data types (one-hot encoding), calculate the cosine similarity ; A and B are the data type vectors of the current job and historical jobs respectively: .
[0054] For data scale (numerical feature), calculate the Euclidean distance ; A and B are the data scales of the current job and historical jobs respectively, and max and min are the maximum and minimum values of the data scales of all jobs: .
[0055] Example: Assume the data scale ranges from 10GB to 100GB. The data scale of the current job is 50GB, and the data scale of the historical job is 70GB. Calculate the Euclidean distance :
[0056] ; For load characteristics (multi-label encoding), calculate the Jaccard similarity ; A and B are the sets of load characteristics of the current job and historical jobs respectively: .
[0057] Example: The current job workload characteristics are {compute-intensive, video memory-intensive}, and the historical jobs are {compute-intensive, communication-intensive}. Calculate the similarity :
[0058] ; For the accuracy requirement (numerical feature), calculate the Euclidean distance ; A and B are the accuracy requirements of the current job and the historical job respectively, and max and min are the maximum and minimum values of the data scale of all jobs: .
[0059] Comprehensive similarity calculation: Calculate the weighted sum of the similarities of each feature according to the preset weights (such as 0.3, 0.2, 0.1, 0.25, 0.15) to obtain the total similarity score: .
[0060] In step S105, as Figure 3 shown, start the probe test; including: Step S1051: Receive an unknown job; after similarity matching, if there is no similar job, it is recognized as an unknown job. Receive the unknown job from step S104 and execute the subsequent steps.
[0061] Step S1052: Conduct performance testing; Deploy the unknown job to different types of GPUs in the cluster for small-scale trial runs respectively. Multiple different combinations of running parameters are set for each GPU type during the test; in the scenario of deep learning training jobs, the parameters include batch size and gradient accumulation steps; under each parameter combination including batch size and gradient accumulation steps, strictly monitor and collect the execution performance data of the job; during the test process, it is necessary to record the multi-dimensional performance data of the actual operation of the job, including job computing power requirements, video memory occupancy, data processed, data transmission time, job running time, actual throughput, GPU resource utilization rate, and running power consumption.
[0062] The acquisition methods of each index are as follows: Job computing power requirement, which is defined as the amount of floating-point operations required for a single iteration of the job, and is obtained by using a DL framework analysis tool, such as PyTorch Profile; Video memory occupancy, which is defined as the video memory operation amount and peak demand of a single batch, and is obtained by using a training framework memory analysis tool during the execution of the job; Data processed, which is defined as the amount of data processed in the job performance test, depending on the job type, such as the number of model iterations or the number of data samples processed; Data transfer time, which is defined as the data transfer time for a single batch and is recorded using the gpustat or cudaMemcpy tool, and the average value is taken from multiple measurements; Job running time, which is defined as the completion time of the job performance test and records the time from the start of the job to the completion of a specific task volume using a timer; Actual throughput, which is defined as the amount of data that can be processed per unit time and is calculated based on the ratio of the amount of processed data to the completion time; GPU resource utilization rate, which is defined as the degree to which the GPU is actually used during the test, is obtained through a GPU performance monitoring tool, and the average value is randomly queried and calculated during the test; Operating power consumption, which is defined as the total energy consumption during the execution of the job and records the total energy consumption through a power consumption sensor.
[0063] Step S1053: Record performance data, that is, execute performance data; detailed record of the performance data obtained through testing, providing accurate data support for establishing an affinity evaluation model under different configurations for subsequent jobs, and achieving the optimal affinity matching of "job-resource configuration".
[0064] Step S1054: Determine the preliminary matching; according to the performance data recorded in step S1053, calculate the actual execution efficiency of the job; quantify the effective production capacity under the actual workload, and determine a preliminary resource matching plan for unknown jobs; Actual execution efficiency The calculation formula is as follows: ; Wherein, represents the job throughput, represents the throughput of the job after normalization; represents the comprehensive utilization rate of the GPU; is the weight coefficient to balance the efficiency index; Sort the actual execution efficiencies of unknown jobs on different GPUs from high to low, compare and screen, take the maximum execution efficiency as the optimal value, and take the selected optimal GPU resources as the initial resource matching plan for unknown jobs.
[0065] This strategy focuses on the performance of the job during actual operation, putting other limiting factors such as cost and computing power matching degree aside for the time being. Its core purpose is to allocate more excellent resources for unknown jobs as much as possible, so as to provide a scientific and reliable basis for the resource allocation of jobs in the initial stage, with the expectation of improving the execution efficiency and overall performance of the job.
[0066] As Figure 4 shown, in step S106, the scheduling optimization loop includes: Step S1061: Initial Allocation Exploration; For unknown jobs, at the initial stage of scheduling, due to the lack of sufficient historical data on the actual performance of the jobs, the "perfect scaling" assumption is adopted for resource allocation exploration, that is, it is assumed that the throughput of unknown jobs increases linearly with the number of resources; namely: ; where, represents the throughput of the job under the configuration including the number of GPUs as n, the type of GPU as t, the batch size as 𝑚, and the number of gradient accumulation steps as 𝑠; Under this assumption, the increase in default resources will bring a proportional increase in throughput, providing a reasonable basis for early resource allocation. In actual operation, at the initial stage, more GPU resources will tend to be allocated to jobs, encouraging jobs to start with a larger resource scale, ensuring that each job can fully exert its potential on the premise of expanding resources, so as to collect sufficient performance data in an environment with rich resources and provide key data support for subsequent throughput model fitting.
[0067] Step S1062: Throughput Model Estimation; Collect the actual throughput data of the job under different numbers of GPUs and types , and calculate the gradient calculation time and the gradient synchronization time according to the model fitting parameters : The set of fitting parameters is used to estimate the gradient calculation time and the gradient synchronization time . Where, represents the fixed time overhead independent of the model scale during the gradient calculation process, such as the initialization time; represents the proportionality coefficient of the gradient calculation time increasing linearly with the model scale ; represents the fixed time overhead independent of the number of GPUs during the gradient synchronization process; represents the proportionality coefficient of the gradient synchronization time increasing with the number of GPUs ;
[0068] Gradient calculation time: ; Gradient synchronization time: ; Calculate the time required for each iteration , including 𝑠 times of gradient calculation time and one synchronization time ; Iteration time : ; Build a throughput model for each job; The throughput model formula is: ; Among them, represents the total batch size; it determines the amount of tasks completed in each iteration.
[0069] Model the actual utilization rate of resources according to the ratio of the effective computing time to the total iteration time : ; By modeling multiple parameters, the throughput model can predict the performance of jobs under different numbers of GPU configurations and provide accurate data support for resource allocation.
[0070] Step S1063: Affinity model evaluation; Based on historical job information or relevant data collected by probe tests, build an affinity evaluation model for each job under different configuration conditions.
[0071] Define a set of resources as a resource configuration tuple , indicating number of , which determines the resource environment for job execution. In the deep learning model training scenario, the affinity evaluation process determines the "comprehensive affinity" between jobs and different resource configurations by considering multi-dimensional matching degrees, thereby achieving the optimal matching between jobs and GPU resources.
[0072] First, to prevent resource waste and performance bottlenecks caused by excessive or insufficient computing power, obtain the computing power matching degree by quantifying the adaptation degree between the floating-point computing power of the GPU and the actual requirements of the job. The computing power matching degree includes two sub-items: basic computing power matching and architecture adaptation correction: ; Among them, represents the amount of floating-point operations required for a single iteration of the task; represents the theoretical peak computing power of the GPU; represents the architectural features required by the task; represents the actual architecture of the GPU; represents the maximum architecture generation difference; As a sensitivity coefficient, it represents architecture sensitivity; For the basic computing power matching item , when , , to prevent the situation of excessive computing power; when , , Prevent the situation of insufficient computing power. Architecture adaptation correction items , By adapting the architecture generation mapping, when the architecture differences are large, the correction factor will significantly reduce the score.
[0073] Quantify the utilization efficiency of GPU video memory resources to obtain the video memory adaptation degree, and this video memory adaptation degree covers two aspects: bandwidth utilization rate and capacity adaptation degree: ; Among them, represents the total amount of video memory operations in a single batch; represents the effective bandwidth of the GPU video memory; represents the processing time of a single batch; represents the peak video memory demand of the task; represents the available video memory after deducting the system reservation; By calculating the energy consumed per unit of computation, quantify the energy utilization efficiency, so as to obtain the energy efficiency optimization degree : ; Among them, represents the measured number of operations of the GPU in the job (floating-point operations per second); represents the effective computing time of the GPU; represents the total energy consumption of the task; According to the measured data of job operation and the estimated data of the throughput model, quantify the effective production capacity under the actual workload, and then obtain the execution efficiency degree , reflecting the comprehensive efficiency performance under the real workload: ; Among them, the throughput as the core indicator to measure the job execution efficiency, reflects the amount of tasks completed per unit time and is a direct manifestation of efficiency; the resource utilization rate correction factor , by the ratio of the effective computing time to the total iteration time, measures the actual utilization rate of resources; Comprehensively consider the hardware procurement or leasing cost and the operation cost , evaluate the economic cost of different types and quantities of GPUs : ; Among them, the hardware cost includes the hardware purchase / lease cost and related supporting costs; the operation cost includes the energy consumption cost and the maintenance cost, etc.
[0074] For each job, systematically consider the above multi-dimensional matching elements, construct an affinity evaluation model for the job and GPU resources under different configuration scenarios, and quantitatively evaluate the adaptability between the job tasks and GPU resources: ; Among them, represents the multi-dimensional affinity value of job and resource configuration scheme ; is the normalization process based on the Sigmoid function; As a dynamic weight coefficient, it is adjusted in real time through the Bayesian optimization algorithm to adapt to the affinity evaluation under different demand scenarios; According to the data evaluated by the affinity evaluation model, construct an affinity matrix , record the affinity of different jobs under different resource configuration schemes , represents the affinity value of job under configuration ; Step S1064: Global scheduling optimization; After entering the scheduling optimization cycle loop, the global scheduling optimizes the resource allocation within the entire cluster by periodically analyzing the overall resource allocation status of the cluster.
[0075] Specifically, during the implementation of the global resource scheduling, run the search algorithm according to the predetermined scheduling cycle to generate a series of binary allocation matrices, each matrix representing a possible resource allocation scheme, defined as the candidate allocation matrix , is the number of jobs, is the number of resource configuration schemes, and the element in the matrix is used to indicate whether job has selected configuration ; For the candidate allocation matrix , the following constraint conditions should be satisfied: Allocation uniqueness: At most one 1 in each row, and each job is allocated one configuration; that is: ; Resource capacity limit: The total sum of the resource configuration requirements of all allocations ≤ the total physical resources of the cluster; Example: Assume that the cluster has 3 jobs and 3 resource configurations, and the global scheduling generates possible candidate allocation matrices: As shown in Table 1: Table 1 Table of 3 jobs and 3 resource configurations;
[0076] The goal of global scheduling is to achieve global optimal scheduling, that is, to obtain the optimal affinity allocation matrix , represents the basic affinity gain obtained by job . The objective function of the optimization problem is: ; The significance of this formula is to obtain the optimal allocation matrix . Among them, represents the basic affinity gain obtained by job ; As a fairness adjustment factor, the generalized power mean is used to achieve multi-objective balance; represents the reallocation penalty factor;
[0077] To avoid the performance degradation problem caused by frequent resource reallocation, a reallocation penalty mechanism is introduced in the global scheduling strategy. When evaluating the affinity gain of a job, a penalty is imposed on the job that requires resource reallocation. Set the reallocation penalty factor , which increases the reallocation overhead based on the historical reallocation frequency of the job and is designed as an exponential decay function :
[0078] ; The penalty factor comprehensively considers the influence of the historical reallocation times and the time interval. Among them, represents the historical reallocation times of job ; represents the time interval between the last two reallocations; represents the average running period of job ; represents the penalty intensity coefficient. The penalty factor will discount the basic affinity gain of the current job, and only when not rescheduling will cause a significant decrease in the optimal value of the scheduling objective, will rescheduling be selected.
[0079] Example: Suppose a job parameter is ; ; ; (medium penalty intensity). Then the total penalty factor:
[0080]
[0081] Suppose the basic affinity gain of the job is 0.85, and the discounted gain is: . As shown in Table 2:
[0082] Table 2 Exponential decay function Formula table;
[0083] Global scheduling optimization transforms the scheduling problem into an optimization problem by maximizing the sum of the affinity values of jobs in the cluster. The objective function is: ; Among them, As a fairness adjustment factor, the generalized power mean is used to achieve multi-objective balance: when When, it degenerates to linear summation, with pure efficiency priority, making resources tilt towards jobs with higher affinity; when , approaching the geometric mean, with fairness priority, emphasizing the fairness of resource allocation; , considering max-min fairness; Step S1065: Job-level scheduling optimization; Job-level scheduling uses lightweight agents launched together with each job to monitor and optimize the performance of individual jobs. Based on the existing resource allocation scheme, fit the throughput function of the job and dynamically adjust the batch size and gradient accumulation steps of the job to ensure that the job achieves the optimal throughput performance under the given GPU resource configuration. The optimization goal is to find the optimal batch size and gradient accumulation steps :
[0084] ; For jobs specified by some users to run with a fixed batch size, by fixing the throughput model parameter m of these jobs and allocating resources according to user requirements, ensure that users can obtain the expected resource configuration during training.
[0085] Step S1066: Job execution; After a series of resource allocation and scheduling optimizations, the job enters the formal execution state. During the execution process, the job continuously runs according to the determined resource allocation scheme, batch size, gradient accumulation steps and other parameters. And record the time required for each iteration in real time , and collect key variables related to this, including the number of GPUs n, batch size m, and gradient accumulation steps s, to provide data support for subsequent model parameter fitting and optimization steps.
[0086] Step S1067: Parameter fitting; The construction of the throughput model and the accuracy of model prediction depend on a set of parameters that can be fitted: . During the job running process, continuously collect data to fit the parameters of the throughput model;
[0087] According to the throughput data recorded in step S1065, perform model parameter fitting by using the L-BFGS-B optimization algorithm; Use the root mean square logarithmic error (RMSLE) as the loss function : ; where \(n\) is the number of samples, is the actual measurement time, and is the predicted iteration time; By minimizing this loss function, the model parameters are adjusted so that the predicted iteration time is closer to the actual time; After multiple iterations of optimization, the model can accurately reflect the actual performance in the GPU cluster and accurately predict the throughput performance of jobs under different GPU numbers, batch sizes, and gradient accumulation step configurations. Based on the optimized throughput model, solid data support is provided for resource allocation decisions, thus forming a scheduling optimization loop.
[0088] Throughput model is the core indicator for measuring the actual execution efficiency of jobs (the core content of the execution efficiency dimension in the affinity evaluation dimension). By constructing a throughput model, the actual execution efficiency of jobs under different resource configurations can be estimated, thereby providing data support for the affinity evaluation between jobs and resource configurations. Then, the affinity evaluation is performed based on the execution efficiency degree of the job and several other dimensions. According to the affinity scores of the job under different configurations, a two-dimensional affinity matrix is constructed, recording the affinity of the job under different resource configurations , where represents the job under the configuration of the affinity value.
[0089] The finally generated resource allocation plan is the optimal allocation matrix obtained according to the objective function in the "global scheduling" ; The affinity evaluation is carried out from five dimensions. Among them, the execution efficiency degree represents the actual performance of the job under different GPU resource configurations, with throughput as the core, and needs to be modeled separately (because in actual operation, the throughput does not increase in multiples with the increase of resources, that is, it cannot be perfectly scaled). The other several dimensions do not need to be modeled separately, and relevant data can be obtained according to the job configuration file and performance test, and calculated using the calculation formulas of each dimension of affinity.
[0090] Embodiment 3 A heterogeneous GPU cluster scheduling system based on resource affinity, as Figure 5 shown, includes: The Similar Job Matching Module is configured to: Based on the affinity history database of historical jobs, match historical jobs with relatively high similarity through job feature matching, and estimate resource requirements; As the initial processing unit of the system, it is responsible for matching submitted jobs with historical jobs. By extracting multi-dimensional parameters such as the model architecture, computing features, and resource requirements of the job, it performs similarity matching with the historical job database. When the matching degree reaches the preset threshold, it loads the cache , calls the historical resource configuration plan to perform initial resource allocation for the job, and at the same time obtains the historical affinity data of similar jobs, which serves as an important reference basis for subsequent precise resource allocation and optimization; If it is a new type of job, it is marked as an unknown job, triggering the probe test module to start performance testing. This module ensures that the matching strategy is adaptively optimized according to the cluster load status through a dynamic weight adjustment mechanism.
[0091] The Probe Test Module is configured to: For new jobs, receive the jobs and perform probe tests on each type of GPU, obtain and record job performance data; For unknown jobs or jobs with failed historical matches, generate and collect the efficiency data of the actual operation of the jobs through performance testing. In a resource isolation environment, the module conducts multi-parameter combination tests on unknown jobs, covering configurations such as different GPU types, batch sizes, and gradient accumulation steps, and collects key performance indicators such as throughput, video memory occupancy, and computing power utilization. The test data quickly locates the optimal parameter combination through an intelligent sampling algorithm, and generates a configuration table containing the actual efficiency data of each GPU type, providing benchmark support for subsequent resource allocation. The probe test module adopts a lightweight design to ensure the minimum impact on cluster resources.
[0092] The Affinity Evaluation Module is configured to: Use historical information or test data to establish an affinity evaluation model for each job, evaluate the affinity values of the job under different resource configurations, and provide data support for resource allocation decisions; Based on historical data or test results, construct an affinity evaluation system to evaluate the affinity values of the job under different resource configurations. By outputting the evaluation results to the job scheduling module, it provides data support for the scheduler to generate resource allocation strategies. At the same time, this module uses the feedback data of the performance monitoring module to continuously perform parameter fitting to optimize the model and improve the model accuracy.
[0093] The job scheduling module is configured to: enter the scheduling optimization cycle loop, construct an allocation matrix according to the objective function of the optimization problem, and allocate corresponding resources to each job. Generate a scheduling plan according to the fitness report of the affinity evaluation module. The module adopts a hierarchical architecture design, including two subsystems of global scheduling and job-level scheduling that operate in coordination. The global scheduler operates with a periodic scheduling strategy, uses a multi-objective optimization algorithm to generate an optimal resource allocation plan at the cluster level, and at the same time introduces a dynamic penalty factor to suppress resource oscillations and improve scheduling stability. The job-level scheduler, as the execution terminal, is responsible for performing fine-grained resource adjustment and converting the global policy into specific resource configurations. By dynamically adjusting the batch size and gradient accumulation steps of the job, local optimization of a single job is performed to ensure that the job achieves optimal performance under the given GPU resources.
[0094] Embodiment 4 A heterogeneous GPU cluster scheduling system according to Embodiment 3, wherein: The heterogeneous GPU cluster scheduling system further includes a performance monitoring module, which is configured to: construct a full-stack monitoring system, collect multi-dimensional metrics during job operation in real time, detect abnormal fluctuations and generate alarm events. Use performance monitoring to synchronously feedback runtime data to the throughput prediction module to calibrate the prediction model and form a closed-loop control flow of "execution - monitoring - optimization". At the same time, identify performance deviations through an anomaly detection algorithm and feedback resource conflict events to the job scheduling module to trigger resource rescheduling or parameter calibration.
[0095] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A heterogeneous GPU cluster scheduling optimization method based on resource affinity, characterized in that Including: Step 1: According to the affinity history library of historical jobs, estimate the resource requirements by matching historical jobs with relatively high similarity through job characteristics; Step 2: For new jobs, receive the jobs and conduct probe tests on each type of GPU, and obtain and record job performance data; Step 3: Use historical information or test data to establish an affinity evaluation model for each job, evaluate the affinity values of jobs under different resource configurations, and provide data support for resource allocation decisions; Step 4: Enter the scheduling optimization cycle loop, construct an allocation matrix according to the objective function of the optimization problem, and allocate corresponding resources to each job.
2. The optimization method for heterogeneous GPU cluster scheduling based on resource affinity according to claim 1, wherein Specifically including: Step S101: Record historical jobs; for clusters with accumulated historical jobs, pre-collect historical job information of the cluster for a period of time; including: job characteristic information, execution performance data on different GPUs, and the final resource allocation plan, to form a historical job library; The recorded job characteristic information is used to match with the jobs newly submitted by users and is recorded in the form of numerical representation; while the execution performance data and the resource allocation plan provide a basis for the preliminary resource allocation of similar jobs; at the same time, the historical jobs are classified according to the job characteristic information and the resource allocation plan to form an affinity history library; Step S102: Receive jobs; receive the jobs newly submitted by users; Step S103: Job characteristic matching; after receiving the jobs newly submitted by users, immediately start the job characteristic matching process, and compare the characteristic information of the jobs newly submitted by users with the jobs in the affinity history library; Step S104: Preliminary resource allocation for similar jobs; after finding the job with the highest similarity, preliminarily estimate the resource requirements of the current job according to the execution information and resource configuration corresponding to the job with the highest similarity; list the resource configuration of the job with the highest similarity as the potential matching resources, and conduct preliminary resource allocation for the current job; at the same time, obtain the historical affinity data of the job with the highest similarity; Step S105: Start probe tests; Step S106: Scheduling optimization cycle; after determining the matching relationship between the newly submitted job and the GPU, the job enters the scheduling cycle loop to continuously optimize resource allocation.
3. An optimization method for heterogeneous GPU cluster scheduling based on resource affinity according to claim 2, characterized in that Job characteristic matching; including: Step S1031: Extract job characteristics; extract characteristic information from the received new job, and the job characteristic information includes: model characteristics, data characteristics, load characteristics, and accuracy requirements, and the data characteristics include data type and data scale; Step S1032: Traverse the historical job library; query the recorded historical jobs in the cluster, and execute Step S1033 until the historical job most similar to the received new job is found; Step S1033: Similarity matching; Based on the vectorized representation of characteristics, convert the job characteristics extracted in Step S1031 into numerical representation, and merge each job characteristic into a vector representation; the model characteristics use one-hot encoding to represent the model type, the data characteristics are numerically represented, the load characteristics are classified by category encoding, and the accuracy requirements are mapped to numerical values for scalar representation; Calculate the similarity between the current job and historical jobs; use weighted similarity to weight each feature according to its importance, calculate the similarity scores of each job after weighting, and find the job with the highest similarity. Step S1034: Determine whether there exists a job with the highest similarity; after calculating the similarity scores, set a similarity threshold to determine whether there exists a job with the highest similarity; if the similarity between the current job and historical jobs exceeds this similarity threshold, it is considered that there exists a job with the highest similarity, otherwise not. If there is a job with the highest similarity, load the cache , obtaining the resource allocation plan of the job with the highest similarity to the current job from the historical job library as a reference for preliminary resource allocation means: directly giving the resource allocation plan of the job with the highest similarity to the current job to the current job; If there are no similar jobs, recognize this job as an unknown job, start the probe test and record the actual execution performance of the job, achieve preliminary resource matching, and provide data support for affinity matching.
4. An optimization method for heterogeneous GPU cluster scheduling based on resource affinity according to claim 3, characterized in that, In step S1031, for model features, obtain the L0 model according to the job submission script or the model definition of the framework, and determine the general category of the job based on the basic model architecture information. Judge its load characteristics according to the analyzed job configuration information. The accuracy requirement is obtained according to the settings in the model framework.
5. The heterogeneous GPU cluster scheduling optimization method based on resource affinity according to claim 3, characterized in that Use weighted similarity to weight each feature according to its importance, calculate the similarity scores of each job after weighting, and find the job with the highest similarity; including: For model features, calculate the cosine similarity ; Calculate the cosine similarity for data types ; For the data scale, calculate the Euclidean distance ; For the load characteristics, calculate the Jaccard similarity ; For the accuracy requirement, calculate the Euclidean distance ; Comprehensive similarity calculation: The similarities of each feature are weighted and summed according to the preset weights to obtain the total similarity score: .
6. The heterogeneous GPU cluster scheduling optimization method based on resource affinity according to claim 2, wherein, In step S105, start the probe test. Including: Step S1051: Receive the unknown job. Step S1052: Conduct performance tests. Deploy the unknown job to different types of GPUs in the cluster for small-scale trial runs respectively, and set multiple different combinations of running parameters for each GPU type during the test; in the scenario of deep learning training jobs, the parameters include batch size and gradient accumulation steps; under each combination of parameters including batch size and gradient accumulation steps, strictly monitor and collect the execution performance data of the job; including job computing power requirements, video memory occupancy, data processed, data transfer time, job running time, actual throughput, GPU resource utilization rate, running power consumption. Step S1053: Record the performance data, i.e., the execution performance data. Step S1054: Determine the preliminary matching; according to the performance data recorded in step S1053, calculate the actual execution performance of the job. Actual execution performance The calculation formula is as follows: ; Among them, represents the job throughput, represents the normalized throughput of the job; represents the comprehensive utilization rate of the GPU; is the weight coefficient to balance the performance index; Sort the actual execution performances of the unknown job on different GPUs from high to low, compare and screen, take the maximum execution performance as the optimal value, and take the selected optimal GPU resources as the initial resource matching scheme for the unknown job.
7. A method for optimizing the scheduling of heterogeneous GPU clusters based on resource affinity according to any one of claims 2-6, characterized in that, In step S106, the scheduling optimization loop; including: Step S1061: Initial allocation exploration. For unknown jobs, at the initial stage of scheduling, assume that the throughput of the unknown job increases linearly with the number of resources; that is: ; Among them, represents the throughput of the job under the configuration with the number of GPUs being n, the type of GPU being t, the batch size being 𝑚, and the gradient accumulation step number being 𝑠; Step S1062: Throughput model estimation. Collect operation throughput data under different numbers of GPUs and types and calculate the gradient calculation time according to the model fitting parameters and the gradient synchronization time : where represents the fixed time overhead independent of the model scale during the gradient calculation; represents the proportionality coefficient of the linear growth of the gradient calculation time with the model scale ; represents the fixed time overhead independent of the number of GPUs during the gradient synchronization; represents the proportionality coefficient of the growth of the gradient synchronization time with the number of GPUs ; Gradient calculation time: ; Gradient synchronization time: ; Calculate the time required for each iteration, including the time for 𝑠 gradient calculations and one synchronization time ; ; Iteration time : ; Build a throughput model for each job. The throughput model formula is as follows: ; Among them, represents the total batch size; The actual utilization rate of the modeling resources based on the ratio of the effective computing time to the total iteration time : ; Step S1063: Affinity model evaluation. Define a set of resources as a resource configuration binary tuple , representing individual ; First, by quantifying the adaptation degree between the floating-point computing power of the GPU and the actual requirements of the job, the computing power matching degree is obtained. The computing power matching degree includes two sub-items: basic computing power matching and architecture adaptation correction ; Among them, represents the amount of floating-point operations required for a single iteration of the task; represents the theoretical peak computing power of the GPU; represents the architectural features required by the task; represents the actual architecture of the GPU; represents the maximum value of the architecture generation gap; As a sensitivity coefficient, it represents architectural sensitivity; For the basic computing power matching item When then When then ; Quantify the utilization efficiency of the GPU video memory resources to obtain the video memory adaptation degree, where the video memory adaptation degree covers two aspects: bandwidth utilization rate and capacity adaptation degree ; Among them, represents the total amount of video memory operations in a single batch; represents the effective bandwidth of the GPU video memory; represents the processing time in a single batch; represents the peak video memory requirement of the task; represents the available video memory after deducting the system reservation; Quantify the energy utilization efficiency by calculating the energy consumed per unit of computational quantity, so as to obtain the energy efficiency optimization degree : ; Among them, represents the actual number of operands measured by the GPU in the job); represents the effective computing time of the GPU; represents the total energy consumption of the task; Quantify the effective production capacity under the actual workload according to the measured data of job operation and the estimated data of the throughput model, and then obtain the execution efficiency : ; Among them, throughput As the core indicator to measure the execution efficiency of jobs, it reflects the amount of tasks completed per unit time, and the resource utilization correction factor , by the ratio of the effective computing time to the total iteration time, measures the actual utilization rate of resources; Comprehensively consider the hardware procurement or rental costs and operating costs , and evaluate the economic cost-effectiveness of different types and quantities of GPUs : ; For each job, build an affinity evaluation model between the job and GPU resources under different configuration scenarios, and quantitatively evaluate the suitability between the job tasks and GPU resources: ; Among them, represents the job and the multi-dimensional affinity value of the resource allocation scheme ; is the normalization process based on the Sigmoid function; serves as the dynamic weight coefficient; Evaluate data according to the affinity evaluation model and construct an affinity matrix , record the affinity of different jobs under different resource allocation schemes , denote the job under the configuration the affinity value; Step S1064: Global scheduling optimization. Specifically, during the implementation of global resource scheduling, a search algorithm is run according to a predetermined scheduling period to generate a series of binary allocation matrices, each matrix representing a possible resource allocation scheme, defined as a candidate allocation matrix , where is the number of jobs, is the number of resource configuration schemes, and the element in the matrix is used to indicate whether job has selected configuration ; for the candidate allocation matrix , the following constraint conditions should be met: Allocation uniqueness: at most one 1 per row, and each job is allocated one configuration; that is: ; Resource capacity limit: The total sum of the resource configuration requirements of all allocated resources ≤ the total physical resources of the cluster. The goal of global scheduling is to achieve global optimal scheduling, that is, to obtain the optimal affinity allocation matrix , denotes the basic affinity gain obtained by job ; the objective function of the optimization problem is as follows: ; Among them, represents the basic affinity gain obtained by the job ; As a fairness adjustment factor, the generalized power mean is used to achieve multi-objective balance; represents the reallocation penalty factor; Set the reassignment penalty factor , which increases the reassignment overhead based on the historical reassignment frequency of the job and is designed as an exponential decay function : ; Among them, represents the historical reassignment times of the job ; represents the time interval between the last two reassignments; represents the job 's average running period; represents the penalty intensity coefficient; Global scheduling optimization transforms the scheduling problem into an optimization problem by maximizing the sum of the affinity values of jobs in the cluster. The objective function is as follows: ; Among them, As a fairness adjustment factor, the generalized power mean is used to achieve multi-objective balance: when it degenerates into linear summation, with pure efficiency priority, tilting resources towards jobs with higher affinity; when , it approaches the geometric mean, with fairness priority, emphasizing the fairness of resource allocation; , the max-min fairness is considered; Step S1065: Job-level scheduling optimization; the optimization goal is to find the optimal batch size and the number of gradient accumulation steps : ; Step S1066: Job execution; Step S1067: Parameter fitting; Continuously collect data during job execution Perform parameter fitting on the throughput model; According to the throughput data recorded in step S1065, model parameter fitting is performed by applying the L-BFGS-B optimization algorithm; Use the root mean square logarithmic error as the loss function : ; where n is the number of samples, is the actual measurement time, is the predicted iteration time; By minimizing this loss function, the model parameters are adjusted so that the predicted iteration time is closer to the actual time; Based on the optimized throughput model, solid data support is provided for resource allocation decisions, thus forming a scheduling optimization loop.
8. A heterogeneous GPU cluster scheduling system based on resource affinity, characterized in that Including: A similar job matching module, configured to: estimate resource requirements by matching historical jobs with relatively high similarity in job characteristics according to the affinity history library of historical jobs; A probe test module, configured to: for a new job, receive the job and perform probe tests on each type of GPU, and obtain and record job performance data; An affinity evaluation module, configured to: use historical information or test data to establish an affinity evaluation model for each job, evaluate the affinity values of the job under different resource configurations, and provide data support for resource allocation decisions; A job scheduling module, configured to: enter the scheduling optimization cycle loop, construct an allocation matrix according to the objective function of the optimization problem, and allocate corresponding resources to each job.
9. The heterogeneous GPU cluster scheduling system based on resource affinity according to claim 8, characterized in that The heterogeneous GPU cluster scheduling system further includes a performance monitoring module, configured to: build a full-stack monitoring system, collect multi-dimensional metrics during job runtime in real time, detect abnormal fluctuations, and generate alarm events.
Citation Information
Patent Citations
Computing network, computing force measurement method, scheduling device and related products
CN115373836A
Deep learning training task scheduling system facing GPU cluster
CN115904666A
Resource scheduling method and system for training tasks of deep recommendation system
CN117492997A
Resource scheduling method and device based on immune algorithm, equipment and medium
CN118733215A
Load balancing method and device for GPU (Graphics Processing Unit) resources and computer equipment
CN119201468A
Cited By
Model training method and device based on heterogeneous GPU cluster and storage medium
CN120508395A
GPU cluster scheduling strategy optimization system based on deep reinforcement learning
CN120821575A
Campus resource collaborative management method and system based on big data
CN121212752A
Private large model-oriented AI computing power scheduling method and system
CN122387633A
An AI computing power scheduling method and system for a private large model
CN122387633B