Load balancing and energy consumption optimization method for computer heterogeneous computing power server cluster

CN122195630BActive Publication Date: 2026-09-25YANCHENG JIKAITONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610059831.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-16
Publication Date
2026-09-25
Estimated Expiration
2046-01-16

AI Technical Summary

Technical Problem

[0004]针对上述问题,本发明提出计算机异构算力服务器集群的负载均衡与能耗优化方法,该计算机异构算力服务器集群的负载均衡与能耗优化方法采用融合计算复杂度、数据并行度等多维度特征的任务分类模型,突破传统粗放式任务分配的局限,实现任务类型与CPU/GPU算力单元优势的精准匹配,通过加权欧氏距离算法计算任务与算力单元的特征相似度,提升分类匹配的科学性与精准度,从源头规避因任务与算力单元不匹配导致的负载失衡问题,经实验验证,任务分类准确率提升35%以上,为集群负载均衡奠定坚实基础

Benefits of technology

[0028]1、本发明采用融合计算复杂度、数据并行度等多维度特征的任务分类模型,突破传统粗放式任务分配的局限,实现任务类型与CPU/GPU算力单元优势的精准匹配,通过加权欧氏距离算法计算任务与算力单元的特征相似度,提升分类匹配的科学性与精准度,从源头规避因任务与算力单元不匹配导致的负载失衡问题,经实验验证,任务分类准确率提升35%以上,为集群负载均衡奠定坚实基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122195630B_ABST
    Figure CN122195630B_ABST
Patent Text Reader

Abstract

The application provides a load balancing and energy consumption optimization method of a computer heterogeneous computing power server cluster, relates to the technical field of computing and cloud computing, and comprises the following steps: S1: obtaining multi-dimensional characteristic parameters of a to-be-scheduled task in the heterogeneous computing power server cluster, and constructing a task characteristic set; S2: constructing a multi-dimensional task classification model based on the task characteristic set, and realizing accurate matching of the to-be-scheduled task and a CPU / GPU computing power unit in the cluster through the classification model; the task classification model fusing multi-dimensional characteristics such as computing complexity and data parallelism breaks through the limitation of traditional extensive task allocation, realizes accurate matching of the task type and the advantage of the CPU / GPU computing power unit, calculates the characteristic similarity of the task and the computing power unit through a weighted Euclidean distance algorithm, improves the scientificity and accuracy of the classification matching, and avoids the load imbalance problem caused by the mismatch between the task and the computing power unit from the source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computing and cloud computing technology, and in particular to a method for load balancing and energy consumption optimization of heterogeneous computing server clusters. Background Technology

[0002] With the development of heterogeneous computing technology, heterogeneous computing server clusters have been widely used in high-performance computing scenarios such as artificial intelligence and big data processing. In existing technologies, load balancing and energy consumption optimization for heterogeneous computing clusters mainly employ traditional task allocation and scheduling methods, including static scheduling based on load thresholds, queue scheduling based on task priorities, and energy-saving scheduling based on energy consumption models. Static scheduling allocates tasks to computing units with lower loads by setting a preset load threshold; queue scheduling prioritizes tasks based on their priority; and energy-saving scheduling reduces energy consumption by shutting down idle nodes or adjusting the operating frequency of computing units. Simultaneously, existing technologies also attempt to introduce simple task characteristics for preliminary classification to achieve basic matching between tasks and computing units.

[0003] However, existing technologies employ a coarse-grained task allocation model, considering only single or a small number of task characteristics without fully integrating multi-dimensional features such as computational complexity and data parallelism. This results in low matching accuracy between tasks and CPU / GPU computing units, easily leading to load imbalance and reduced overall cluster performance. Furthermore, they lack a dynamic trade-off mechanism between performance and energy consumption. Most scheduling algorithms are single-objective optimizations, prioritizing either performance or energy consumption, failing to adaptively adjust to the real-time load status of the cluster. Under high load, energy consumption optimization leads to performance degradation, while under low load, excessive pursuit of performance results in energy waste. Simultaneously, scheduling strategies have poor adaptability; existing methods are mostly rule-based scheduling, failing to fully utilize historical scheduling data and real-time feedback for continuous optimization. This makes them ill-suited to complex and ever-changing computing environments, unable to achieve long-term stable improvement in load balancing and energy consumption optimization. Moreover, existing energy consumption scheduling algorithms do not effectively integrate real-time electricity prices and carbon emission intensity signals, failing to meet the development needs of green computing. Therefore, this invention proposes a load balancing and energy consumption optimization method for heterogeneous computing server clusters to address the problems existing in the prior art. Summary of the Invention

[0004] To address the aforementioned issues, this invention proposes a load balancing and energy consumption optimization method for heterogeneous computing server clusters. This method employs a task classification model that integrates multi-dimensional features such as computational complexity and data parallelism, breaking through the limitations of traditional extensive task allocation. It achieves precise matching between task types and the advantages of CPU / GPU computing units. By calculating the feature similarity between tasks and computing units using a weighted Euclidean distance algorithm, the scientific nature and accuracy of classification matching are improved. This avoids load imbalance caused by mismatch between tasks and computing units from the source. Experimental results show that the task classification accuracy is improved by more than 35%, laying a solid foundation for cluster load balancing.

[0005] To achieve the objectives of this invention, the invention is implemented through the following technical solution: a method for load balancing and energy consumption optimization of heterogeneous computing server clusters, comprising the following steps:

[0006] S1: Obtain multi-dimensional feature parameters of tasks to be scheduled in the heterogeneous computing server cluster and construct a task feature set;

[0007] S2: Construct a multi-dimensional task classification model based on the task feature set, and achieve accurate matching between the task to be scheduled and the CPU / GPU computing power unit in the cluster through the classification model;

[0008] S3: Construct a dynamic trade-off scheduling function with dual performance and energy consumption objectives, and dynamically adjust the weight coefficients of performance and energy consumption objectives according to the real-time load status of the cluster;

[0009] S4: Integrate the cluster's real-time power consumption model with external electricity price signals or carbon emission intensity signals, substitute them into the dual-objective dynamic trade-off scheduling function, and generate an initial scheduling scheme;

[0010] S5: Construct a deep reinforcement learning model, train the model using historical scheduling data of the cluster, continuously optimize the model parameters through real-time scheduling feedback, and generate an adaptive scheduling strategy.

[0011] S6: Adjust the initial scheduling scheme based on the adaptive scheduling strategy, execute the final scheduling, and monitor the cluster load and energy consumption status in real time.

[0012] Further improvements are made in the following: In S1, the multi-dimensional feature parameters include computational complexity, data parallelism, task priority, data throughput, and memory requirements. The computational complexity is represented by the product of the number of instruction sets required by the task and the number of operation cycles, and the data parallelism is represented by the number of subtasks that can be divided and executed in parallel.

[0013] A further improvement is made in S2, where the multi-dimensional task classification model achieves accurate matching between tasks and computing units through task feature similarity matching. The similarity calculation uses a weighted Euclidean distance algorithm, with the specific formula as follows:

[0014] ,

[0015] Where Sim(x,y) represents the similarity between the feature vector x of the task to be scheduled and the feature vector y of the computing unit adaptation; n represents the number of feature dimensions; ω i Let the weight coefficient of the i-th feature dimension satisfy the following condition: Furthermore, the sum of the weight coefficients corresponding to the computational complexity and the data parallelism is not less than 0.6; x i y represents the normalized value of the task to be scheduled in the i-th feature dimension; i This represents the normalized value of the adaptation threshold for the computing power unit in the i-th feature dimension.

[0016] A further improvement lies in: the weight coefficient ω of the feature dimension. i The analytic hierarchy process (AHP) is used to determine the importance of features. Specifically, a feature importance judgment matrix is ​​constructed, the eigenvector corresponding to the largest eigenvalue of the judgment matrix is ​​calculated, and the weight coefficients of each feature dimension are obtained after normalization.

[0017] A further improvement lies in the following: In S3, the expression for the performance-energy consumption dual-objective dynamic trade-off scheduling function is:

[0018] ,

[0019] Where f(α,P,E) is the objective value of the bi-objective scheduling optimization; α is the dynamic weight coefficient, with a value range of [0,1]; P is the task completion performance index, in seconds. -1 E represents the number of tasks completed per unit of time; E is the energy consumption optimization index, in W. -1 α represents the number of tasks completed per unit of energy consumption; when the cluster load rate is ≥70%, α is 0.7-1.0, prioritizing performance; when the cluster load rate is ≤30%, α is 0-0.3, prioritizing energy consumption optimization.

[0020] A further improvement is that the cluster load rate is calculated by the average real-time load of all computing units in the cluster. The real-time load is represented by a weighted sum of the CPU / GPU utilization, memory usage, and task queue length of the computing units.

[0021] Further improvements are made in the following: In S4, the real-time power consumption model of the cluster includes static power consumption, dynamic power consumption and power consumption inflection point. Static power consumption is the basic power consumption of the computing unit in the idle state. Dynamic power consumption is positively correlated with the computing load of the computing unit. The power consumption inflection point is the point of sudden change in dynamic power consumption with load. When the load exceeds the inflection point value, the growth rate of dynamic power consumption increases by more than 20%.

[0022] A further improvement is made in S5, where the reward function of the deep reinforcement learning model is:

[0023] ,

[0024] Where R(t) is the scheduling reward value at time t; β, γ, and δ are the weighting coefficients of performance reward, energy consumption reward, and load imbalance penalty, respectively, satisfying β+γ+δ=1; R P (t) represents the performance reward value at time t, which is positively correlated with the task completion efficiency; R E (t) represents the energy reward value at time t, which is negatively correlated with the energy consumption per unit task; R L (t) represents the load imbalance penalty value at time t, which is positively correlated with the load standard deviation of each computing unit in the cluster.

[0025] Further improvements are made in that the deep reinforcement learning model adopts the DQN (DeepQ-Network) architecture. The model input is the real-time load status of the cluster, task characteristics and current energy consumption data, and the output is the scheduling action space, including the target computing power unit for task allocation, the adjustment of task execution order and the start and stop of idle nodes. The model parameters are updated through the temporal difference learning algorithm.

[0026] Further improvements are made in S6, where the real-time monitored metrics include CPU / GPU utilization, task completion time, total cluster energy consumption, load balancing, and carbon emissions per task for each computing unit. When the load balancing is found to be below a preset threshold or the energy consumption per task is found to be above a preset threshold, the retraining of the deep reinforcement learning model and the updating of the scheduling strategy are triggered.

[0027] The beneficial effects of this invention are as follows:

[0028] 1. This invention adopts a task classification model that integrates multi-dimensional features such as computational complexity and data parallelism, breaking through the limitations of traditional extensive task allocation. It achieves accurate matching between task types and the advantages of CPU / GPU computing power units. By calculating the feature similarity between tasks and computing power units through a weighted Euclidean distance algorithm, it improves the scientificity and accuracy of classification matching, and avoids the load imbalance problem caused by the mismatch between tasks and computing power units from the source. Experimental verification shows that the task classification accuracy is improved by more than 35%, laying a solid foundation for cluster load balancing.

[0029] 2. This invention designs a dynamic weighted performance-energy consumption dual-objective optimization function to achieve an adaptive dynamic trade-off between performance and energy consumption. The weight coefficients are adjusted according to the real-time load rate of the cluster. Under high load, performance is prioritized to reduce task completion time, while under low load, energy consumption is prioritized to reduce energy waste. At the same time, a real-time power consumption model of the cluster is integrated, including static power consumption, dynamic power consumption, and power consumption inflection point. It also responds to real-time electricity price or carbon emission intensity signals, so that the scheduling decision takes into account performance, energy consumption, and green environmental protection requirements, effectively improving the cluster's energy efficiency ratio and reducing operating costs.

[0030] 3. This invention constructs an adaptive scheduling strategy based on deep reinforcement learning, trains a model using historical scheduling data, and continuously optimizes model parameters and scheduling strategies through real-time scheduling feedback. This strategy can adapt to complex and ever-changing computing environments, dynamically adjust task allocation schemes, execution order, and idle node status, and continuously improve load balancing efficiency and energy consumption optimization. At the same time, by setting a multi-dimensional reward function, it guides the model to learn the optimal scheduling strategy, ensures the long-term stable operation of the cluster, and improves the intelligence level of scheduling. Attached Figure Description

[0031] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0032] To enhance understanding of the present invention, the present invention will be further described in detail below with reference to embodiments. These embodiments are only used to explain the present invention and do not constitute a limitation on the scope of protection of the present invention.

[0033] Example 1

[0034] according to Figure 1 As shown, this embodiment proposes a load balancing and energy consumption optimization method for heterogeneous computing server clusters, including the following steps:

[0035] S1: Obtain multi-dimensional feature parameters of tasks to be scheduled in a heterogeneous computing server cluster and construct a task feature set. The multi-dimensional feature parameters include computational complexity, data parallelism, task priority, data throughput, and memory requirements. Computational complexity is represented by the product of the number of instructions required by the task and the computation cycle. Data parallelism is represented by the number of subtasks that can be split and executed in parallel. Comprehensively capture the core attributes of the task to avoid the task perception bias caused by a single feature description and provide complete data support for subsequent accurate scheduling. At the same time, through standardized feature representation methods, improve the comparability of features of different types of tasks and reduce the difficulty of building subsequent classification models.

[0036] S2: A multi-dimensional task classification model is constructed based on the task feature set. This model enables precise matching between tasks to be scheduled and CPU / GPU computing units in the cluster. The multi-dimensional task classification model achieves precise matching between tasks and computing units through task feature similarity matching. The similarity calculation uses a weighted Euclidean distance algorithm, and the specific formula is as follows:

[0037] ,

[0038] Where Sim(x,y) represents the similarity between the feature vector x of the task to be scheduled and the feature vector y of the computing unit adaptation; n represents the number of feature dimensions; ω i Let the weight coefficient of the i-th feature dimension satisfy the following condition: Furthermore, the sum of the weight coefficients corresponding to the computational complexity and the data parallelism is not less than 0.6; x i y represents the normalized value of the task to be scheduled in the i-th feature dimension; i This represents the normalized adaptation threshold value of the computing power unit in the i-th feature dimension. The weight coefficient ω of the feature dimension... i The analytic hierarchy process (AHP) is used to determine the following: a feature importance judgment matrix is ​​constructed, the eigenvector corresponding to the largest eigenvalue of the judgment matrix is ​​calculated, and the weight coefficients of each feature dimension are obtained after normalization. The hardware advantages of CPU / GPU computing units are fully utilized to improve task execution efficiency and reduce the waste of computing resources from the source. At the same time, precise matching is used to reduce the load difference between different computing units, laying a solid foundation for overall load balancing.

[0039] S3: Construct a dynamic trade-off scheduling function that balances performance and energy consumption goals, dynamically adjusting the weighting coefficients of these goals based on the real-time cluster load. The expression for the performance-energy consumption dual-goal dynamic trade-off scheduling function is as follows:

[0040] ,

[0041] Where f(α,P,E) is the objective value of the bi-objective scheduling optimization; α is the dynamic weight coefficient, with a value range of [0,1]; P is the task completion performance index, in seconds. -1 E represents the number of tasks completed per unit of time; E is the energy consumption optimization index, in W. -1The α value represents the number of tasks completed per unit of energy consumption. When the cluster load rate is ≥70%, α is 0.7-1.0, prioritizing performance; when the cluster load rate is ≤30%, α is 0-0.3, prioritizing energy consumption optimization. The cluster load rate is calculated by the average real-time load of all computing units in the cluster. The real-time load is represented by a weighted sum of the CPU / GPU utilization, memory usage, and task queue length of the computing units. This achieves adaptive matching between performance and energy consumption, avoiding extreme problems caused by single-objective optimization and improving the flexibility of cluster operation. At the same time, the dynamic weight adjustment mechanism can accurately match changes in cluster load, ensuring fast task response speed under high load and energy saving effect under low load.

[0042] S4: Integrates the cluster's real-time power consumption model with external electricity price signals or carbon emission intensity signals, and substitutes them into a dual-objective dynamic trade-off scheduling function to generate an initial scheduling scheme. The cluster's real-time power consumption model includes static power consumption, dynamic power consumption, and power consumption inflection point. Static power consumption is the basic power consumption of the computing unit in the idle state, dynamic power consumption is positively correlated with the computing load of the computing unit, and the power consumption inflection point is the point of abrupt change in dynamic power consumption with load. When the load exceeds the inflection point value, the rate of increase in dynamic power consumption increases by more than 20%. By integrating real-time power consumption characteristics, the scheduling scheme is made more in line with the hardware energy consumption law, improving the accuracy of energy consumption optimization. At the same time, it responds to external electricity price and carbon emission signals, taking into account both operating cost control and green computing requirements, thus expanding the application value of the scheduling scheme.

[0043] S5: Construct a deep reinforcement learning model, train the model using historical cluster scheduling data, continuously optimize model parameters through real-time scheduling feedback, and generate an adaptive scheduling strategy; the reward function of the deep reinforcement learning model is:

[0044] ,

[0045] Where R(t) is the scheduling reward value at time t; β, γ, and δ are the weighting coefficients of performance reward, energy consumption reward, and load imbalance penalty, respectively, satisfying β+γ+δ=1; R P (t) represents the performance reward value at time t, which is positively correlated with the task completion efficiency; R E (t) represents the energy reward value at time t, which is negatively correlated with the energy consumption per unit task; R L(t) represents the load imbalance penalty value at time t, which is positively correlated with the standard deviation of the load of each computing unit in the cluster. The deep reinforcement learning model adopts the DQN (DeepQ-Network) architecture. The model input includes the real-time load status of the cluster, task characteristics, and current energy consumption data. The output is the scheduling action space, including task allocation to target computing units, adjustment of task execution order, and start / stop of idle nodes. The model parameters are updated through a temporal difference learning algorithm. Historical data is used to accumulate scheduling experience, improve the initial adaptability of the strategy, and reduce trial and error costs. At the same time, the real-time feedback optimization mechanism enables the scheduling strategy to dynamically adapt to changes in the computing environment, ensuring the load balancing and energy consumption optimization effect in long-term operation.

[0046] S6: Adjust the initial scheduling scheme based on an adaptive scheduling strategy, execute the final scheduling, and monitor the cluster load and energy consumption status in real time. Real-time monitored metrics include CPU / GPU utilization of each computing unit, task completion time, total cluster energy consumption, load balancing, and carbon emissions per task. When the load balancing falls below a preset threshold or the energy consumption per task exceeds a preset threshold, the deep reinforcement learning model is retrained and the scheduling strategy is updated. This secondary adjustment optimizes the initial scheme, further improving scheduling accuracy and reducing potential load imbalances and energy waste. Simultaneously, the real-time monitoring and retraining trigger mechanism form a closed-loop optimization, ensuring the scheduling strategy always adapts to dynamic cluster changes and guaranteeing long-term stable and efficient system operation.

[0047] Example 2

[0048] according to Figure 1 As shown, this embodiment proposes a load balancing and energy consumption optimization method for heterogeneous computing server clusters, including the following steps:

[0049] Task Feature Extraction: 1000 tasks to be scheduled (including image recognition, natural language processing, data mining, and other types) were selected from an AI training cluster. Five core features were extracted for each task: computational complexity (…). (Unit: MIPS, i.e., million instructions per second), data parallelism ( That is, the number of subtasks that a task can be broken down into, and task priority ( Levels 1-5, with Level 1 being the highest), data throughput ( (Unit: GB / s), Memory Requirements ( (Unit: GB). Normalize each feature value, mapping it to the [0,1] interval, and construct the task feature vector x=( ).

[0050] Determining the adaptation features of computing units: For the CPU computing units (Intel Xeon Gold 6330) and GPU computing units (NVIDIA A100) in the cluster, the optimal adaptation feature thresholds for different task types were tested. After normalization, the CPU adaptation feature vector yCPU=(0.6,0.3,0.4,0.5,0.7) and the GPU adaptation feature vector yGPU=(0.9,0.8,0.6,0.7,0.5) were obtained.

[0051] Weight coefficient determination: A feature importance judgment matrix was constructed using the analytic hierarchy process (AHP). Five experts in heterogeneous computing were invited to score the importance of the five features. The calculated weight coefficients ω = (0.3, 0.3, 0.2, 0.1, 0.1) satisfy the following conditions: (Computational complexity and data parallelism are the core characteristics).

[0052] Similarity Calculation and Matching: A weighted Euclidean distance algorithm is used to calculate the similarity between each task's feature vector and yCPU and yGPU, and tasks are assigned to computing units with lower similarity (higher matching degree). Experimental results show that the task classification accuracy of this embodiment reaches 92%, which is 35% higher than the traditional single feature classification method (accuracy 57%), and the load balancing is improved by 40%.

[0053] Example 3

[0054] according to Figure 1 As shown, this embodiment proposes a load balancing and energy consumption optimization method for heterogeneous computing server clusters, including the following steps:

[0055] Cluster load monitoring: Real-time collection of CPU and GPU computing unit utilization, memory usage, and task queue length in the cluster. A weighted summation formula is used to calculate the load value of each computing unit, and the average value is taken to obtain the real-time cluster load rate. Load rate threshold settings: High load threshold 70%, low load threshold 30%.

[0056] Dual-objective optimization function parameter settings: The performance index P is the number of tasks completed per unit time (s). -1 The energy consumption index E is expressed as the number of tasks completed per unit of energy consumption (W). -1 The dynamic weighting coefficient α is set according to the load rate: when the load rate is ≥70% (high load), α=0.8; when 30% < load rate <70% (medium load), α=0.5; when the load rate is ≤30% (low load), α=0.2.

[0057] Real-time power consumption model integration: By collecting the static and dynamic power consumption of the CPU and GPU through a power meter, the static power consumption of the CPU is determined to be 80W, and the dynamic power consumption is positively correlated with the utilization rate (for every 10% increase in utilization rate, the dynamic power consumption increases by 15W). The power consumption inflection point is 80% utilization rate (after which the dynamic power consumption growth rate increases by 25%). The static power consumption of the GPU is 100W, and the dynamic power consumption is positively correlated with the utilization rate (for every 10% increase in utilization rate, the dynamic power consumption increases by 20W). The power consumption inflection point is 75% utilization rate (after which the dynamic power consumption growth rate increases by 22%).

[0058] Signal Response and Scheduling Execution: The algorithm integrates real-time grid price signals (peak price 1.2 yuan / kWh, off-peak price 0.6 yuan / kWh). During peak price periods, α is further reduced to 0.1 under low load conditions, prioritizing the shutdown of idle nodes. During off-peak price periods, α is maintained at 0.8 under high load conditions, ensuring performance while reducing energy consumption costs. Experiments show that this algorithm reduces task completion time by 28% under high load and energy consumption by 32% under low load conditions.

[0059] Example 4

[0060] according to Figure 1 As shown, this embodiment proposes a load balancing and energy consumption optimization method for heterogeneous computing server clusters, including the following steps:

[0061] Reinforcement learning model construction: A deep reinforcement learning model is constructed using the DQN architecture. The input layer is a 12-dimensional vector (containing cluster load rate, utilization rate of each computing unit, average value of task feature vector, real-time energy consumption, electricity price signal, etc.), the hidden layer is a 2-layer fully connected layer (with 64 and 32 neurons respectively), and the output layer is a 6-dimensional action space (corresponding to 6 scheduling actions: allocate to CPU, allocate to GPU, adjust task priority, split task, shut down idle CPU nodes, and shut down idle GPU nodes).

[0062] Reward function parameter settings: Set the reward function weight coefficients β=0.4, γ=0.4, and δ=0.2, satisfying β+γ+δ=1. Where R... P (t) = 1 - (actual task completion time / preset task completion time), R E (t) = 1 - (actual unit task energy consumption / preset unit task energy consumption), R L (t) = (cluster load standard deviation / preset load standard deviation threshold).

[0063] Model Training and Optimization: Historical scheduling data from the cluster over the past 6 months (including 100,000 scheduling records) was collected as the training set. The model was trained using a temporal difference learning algorithm with a learning rate of 0.001 and 10,000 training iterations. During training, real-time scheduling data was used for validation every 100 iterations to adjust the model parameters.

[0064] Adaptive scheduling execution: Real-time cluster status data is collected and input into the trained model to obtain and execute the optimal scheduling action; scheduling feedback data (task completion status, energy consumption data, and load data) is collected hourly to fine-tune the model parameters. Experiments show that this strategy keeps the cluster load balance stable above 0.9 over the long term, and reduces energy consumption per task by 38% compared to traditional methods.

[0065] Validation data:

[0066] The load balancing and energy consumption optimization method for a heterogeneous computing server cluster of the present invention has been applied to a cluster containing 20 CPU nodes (Intel Xeon Gold 6330) and 10 GPU nodes (NVIDIA). A three-month experimental verification was conducted on a heterogeneous cluster of A100. Compared with existing mainstream scheduling methods (traditional load threshold scheduling, single-objective performance scheduling, and single-objective energy consumption scheduling), the core performance indicators are improved as follows: Task classification accuracy: 92%, while the average of existing methods is 57%, an improvement of more than 35%; Load balancing: The present invention is consistently stable at 0.92, while the average of existing methods is 0.75, an improvement of 22.7%; High load (load rate ≥ 70%) task completion time: The present invention is 28% shorter than the average of existing methods; Low load (load rate ≤ 30%) unit task energy consumption: The present invention is 32% lower than the average of existing methods; Long-term average unit task energy consumption: The present invention is 38% lower than the average of existing methods; Overall cluster energy efficiency ratio (number of tasks completed per unit of energy consumption): The present invention is 61.3% higher than the average of existing methods; Energy cost in response to electricity price / carbon emission signals: Energy cost during peak electricity price periods is reduced by 25%, and carbon emission intensity is reduced by 30%.

[0067] Experimental results show that the present invention can effectively solve the problems of load imbalance and excessive energy consumption in heterogeneous computing power clusters, significantly improve the energy efficiency ratio while ensuring cluster performance, adapt to complex and ever-changing computing environments, and has good practicality and promotion value.

[0068] This method for load balancing and energy consumption optimization in a heterogeneous computing server cluster employs a task classification model that integrates multi-dimensional features such as computational complexity and data parallelism. It overcomes the limitations of traditional extensive task allocation, achieving precise matching between task types and the advantages of CPU / GPU computing units. By calculating the feature similarity between tasks and computing units using a weighted Euclidean distance algorithm, it improves the scientific rigor and accuracy of classification matching, preventing load imbalance caused by task-computing unit mismatch from the outset. Experiments have verified that the task classification accuracy is improved by more than 35%, laying a solid foundation for cluster load balancing. Furthermore, this invention designs a dynamic weighted performance-energy consumption dual-objective optimization function, achieving an adaptive dynamic trade-off between performance and energy consumption. The weight coefficients are adjusted according to the real-time cluster load rate, prioritizing performance to reduce task completion time under high load and prioritizing energy consumption to reduce energy waste under low load. Simultaneously, it integrates a real-time cluster power consumption model, including static power consumption, dynamic power consumption, and power consumption inflection points, and responds to real-time electricity prices or carbon emission intensity signals. This ensures that scheduling decisions consider performance, energy consumption, and environmental protection needs, effectively improving the cluster's energy efficiency ratio and reducing operating costs. Meanwhile, this invention constructs an adaptive scheduling strategy based on deep reinforcement learning, trains the model using historical scheduling data, and continuously optimizes the model parameters and scheduling strategy through real-time scheduling feedback. This strategy can adapt to complex and ever-changing computing environments, dynamically adjust task allocation schemes, execution order, and idle node status, and continuously improve load balancing efficiency and energy consumption optimization. At the same time, by setting a multi-dimensional reward function, it guides the model to learn the optimal scheduling strategy, ensures the long-term stable operation of the cluster, and improves the intelligence level of scheduling.

[0069] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for load balancing and energy consumption optimization of heterogeneous computing server clusters, characterized in that, Includes the following steps: S1: Obtain multi-dimensional feature parameters of tasks to be scheduled in the heterogeneous computing server cluster and construct a task feature set; S2: A multi-dimensional task classification model is constructed based on the task feature set. This model enables precise matching between tasks to be scheduled and CPU / GPU computing units in the cluster. The multi-dimensional task classification model achieves precise matching between tasks and computing units through task feature similarity matching. The similarity calculation uses a weighted Euclidean distance algorithm, and the specific formula is as follows: , Where Sim(x,y) represents the similarity between the feature vector x of the task to be scheduled and the feature vector y of the computing unit adaptation; n represents the number of feature dimensions; ω i Let the weight coefficient of the i-th feature dimension satisfy the following condition: Furthermore, the sum of the weight coefficients corresponding to the computational complexity and the data parallelism is not less than 0.6; x i y represents the normalized value of the task to be scheduled in the i-th feature dimension; i This represents the normalized value of the adaptation threshold for the computing power unit in the i-th feature dimension; S3: Construct a dynamic trade-off scheduling function that balances performance and energy consumption goals, dynamically adjusting the weighting coefficients of these goals based on the real-time cluster load. The expression for the performance-energy consumption dual-goal dynamic trade-off scheduling function is as follows: , Where f(α,P,E) is the objective value of the bi-objective scheduling optimization; α is the dynamic weight coefficient, with a value range of [0,1]; P is the task completion performance index, in seconds. -1 E represents the number of tasks completed per unit of time; E is the energy consumption optimization index, in W. -1 α represents the number of tasks completed per unit of energy consumption; when the cluster load rate is ≥70%, α is 0.7-1.0, prioritizing performance; when the cluster load rate is ≤30%, α is 0-0.3, prioritizing energy consumption optimization. S4: Integrate the cluster's real-time power consumption model with external electricity price signals or carbon emission intensity signals, substitute them into the dual-objective dynamic trade-off scheduling function, and generate an initial scheduling scheme; S5: Construct a deep reinforcement learning model, train the model using historical cluster scheduling data, continuously optimize model parameters through real-time scheduling feedback, and generate an adaptive scheduling strategy; the reward function of the deep reinforcement learning model is: , Where R(t) is the scheduling reward value at time t; β, γ, and δ are the weighting coefficients of performance reward, energy consumption reward, and load imbalance penalty, respectively, satisfying β+γ+δ=1; R P (t) represents the performance reward value at time t, which is positively correlated with the task completion efficiency; R E (t) represents the energy reward value at time t, which is negatively correlated with the energy consumption per unit task; R L (t) represents the load imbalance penalty value at time t, which is positively correlated with the load standard deviation of each computing unit in the cluster; S6: Adjust the initial scheduling scheme based on the adaptive scheduling strategy, execute the final scheduling, and monitor the cluster load and energy consumption status in real time.

2. The load balancing and energy consumption optimization method for heterogeneous computing server clusters according to claim 1, characterized in that: In S1, the multi-dimensional feature parameters include computational complexity, data parallelism, task priority, data throughput, and memory requirements. The computational complexity is represented by the product of the number of instruction sets required by the task and the number of operation cycles, and the data parallelism is represented by the number of subtasks that can be divided and executed in parallel.

3. The load balancing and energy consumption optimization method for heterogeneous computing server clusters according to claim 1, characterized in that: The weight coefficient ω of the feature dimension i The analytic hierarchy process (AHP) is used to determine the importance of features. Specifically, a feature importance judgment matrix is ​​constructed, the eigenvector corresponding to the largest eigenvalue of the judgment matrix is ​​calculated, and the weight coefficients of each feature dimension are obtained after normalization.

4. The load balancing and energy consumption optimization method for heterogeneous computing server clusters according to claim 1, characterized in that: The cluster load rate is calculated by the average real-time load of all computing units in the cluster. The real-time load is represented by a weighted sum of the CPU / GPU utilization, memory usage, and task queue length of the computing units.

5. The load balancing and energy consumption optimization method for heterogeneous computing server clusters according to claim 1, characterized in that: In S4, the cluster real-time power consumption model includes static power consumption, dynamic power consumption, and power consumption inflection point. Static power consumption is the basic power consumption of the computing unit in the idle state. Dynamic power consumption is positively correlated with the computing load of the computing unit. The power consumption inflection point is the point of sudden change in dynamic power consumption as the load changes. When the load exceeds the inflection point value, the dynamic power consumption growth rate increases by more than 20%.

6. The load balancing and energy consumption optimization method for heterogeneous computing server clusters according to claim 1, characterized in that: The deep reinforcement learning model adopts the DQN (DeepQ-Network) architecture. The model input is the real-time load status of the cluster, task characteristics and current energy consumption data, and the output is the scheduling action space, including task allocation target computing power units, task execution order adjustment and idle node start and stop. The model parameters are updated through the temporal difference learning algorithm.

7. The load balancing and energy consumption optimization method for heterogeneous computing server clusters according to claim 1, characterized in that: In S6, the real-time monitored metrics include CPU / GPU utilization of each computing unit, task completion time, total cluster energy consumption, load balancing, and carbon emissions per task. When the load balancing is found to be below a preset threshold or the energy consumption per task is found to be above a preset threshold, the retraining of the deep reinforcement learning model and the updating of the scheduling strategy are triggered.

Citation Information

Patent Citations

  • Robot grabbing stability detection method based on multi-modal feature decoupling technology

    CN116945234A

  • Distributed computing-based intelligent design method and system for electric power engineering cloud resources

    CN120494574A