A multi-cluster computing algorithm cooperative scheduling method and system based on karmada

CN122816780APending Publication Date: 2026-09-25JINGNENG DIGITAL IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610614710.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]本发明提供一种基于Karmada的多集群算电协同调度方法及系统,以解决现有技术中缺乏对数据中心能效PUE的量化调度机制、无法精准评估工作负载跨集群迁移的经济性、以及重调度决策缺乏防抖设计导致系统稳定性差的问题

Benefits of technology

[0046](1)本发明通过将PUE值作为乘性因子嵌入成本子评分计算公式,确立了PUE对IT运行成本的线性放大作用。相比于现有技术仅依据电价判断或仅将PUE视为静态约束,本发明能够准确识别低电价但高PUE集群的总成本劣势,避免调度决策偏差,确保工作负载始终被调度至综合运营成本最低的集群。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816780A_ABST
    Figure CN122816780A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of algorithm and electricity cooperation, and particularly relates to a multi-cluster algorithm and electricity cooperation scheduling method and system based on Karmada. The method comprises the following steps: S1: collecting data, periodically collecting real-time energy efficiency data from each member cluster through the energy data collector deployed on the multi-cluster management platform control plane; S2: calculating cost score, when receiving a workload scheduling request, the multi-dimensional comprehensive score engine embeds the PUE value as a multiplicative factor in the cost sub-score calculation of each candidate member cluster to correct the influence of the real-time electricity price on the actual IT device energy consumption cost; S3: executing scheduling according to the score, the multi-cluster management platform control plane generates a propagation strategy according to the comprehensive scheduling score result containing at least the cost sub-score. The problems that the existing technology cannot accurately evaluate the economy of workload cross-cluster migration and the lack of anti-jitter design of rescheduling decision leads to poor system stability are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computing and power coordination technology, and in particular to a multi-cluster computing and power coordination scheduling method and system based on Karmada. Background Technology

[0002] As enterprises deepen their digital transformation, more and more business applications are being deployed on Kubernetes-based containerized platforms. To meet the needs of cross-regional deployment, disaster recovery failover, and avoiding lock-in to a single cloud vendor, multi-cluster management platforms, represented by Karmada (Kubernetes Armada, a cloud-native multi-cloud multi-cluster container orchestration platform), are widely used. These platforms manage multiple member clusters through a unified control plane, distributing workloads to different data centers for execution.

[0003] Meanwhile, against the backdrop of "dual carbon" targets and electricity market reforms, the energy consumption costs and carbon emissions of data centers have attracted significant attention. The total energy consumption of a data center consists of the energy consumption of IT equipment and the energy consumption of infrastructure such as cooling and power supply. The efficiency of the latter is typically measured by PUE (Power Usage Effectiveness). Due to differences in data center location, climate conditions, and cooling architecture across different regions, PUE values ​​vary, resulting in different actual electricity costs for the same computing task across different clusters.

[0004] Current common multi-cluster scheduling schemes primarily make decisions based on resource dimensions such as CPU availability, memory availability, and network latency of each member cluster. These schemes rarely incorporate energy-side factors such as real-time electricity prices, PUE (Power Usage Effectiveness) levels, and the proportion of renewable energy in the data centers where the clusters reside into their scheduling scoring systems. When the electricity prices or energy efficiency of each data center change dynamically, scheduling decisions are unlikely to automatically favor clusters with lower overall energy costs, potentially leading to higher actual operating costs or insufficient utilization of green electricity.

[0005] For workloads already in operation, some scheduling systems can support rescheduling when resources are scarce, but they typically lack quantitative assessment of the migration process itself. Cross-cluster migration often involves overhead such as data network transmission, brief service interruptions, and context environment reconstruction. If scheduling decisions do not calculate these overheads, the overall efficiency may be affected because the migration benefits outweigh the costs. Furthermore, if the rescheduling trigger conditions are overly sensitive to instantaneous fluctuations in energy prices, it may lead to frequent workload migrations, which is detrimental to system stability. Summary of the Invention

[0006] This invention provides a multi-cluster computing power collaborative scheduling method and system based on Karmada to solve the problems in the prior art, such as the lack of a quantitative scheduling mechanism for data center power efficiency (PUE), the inability to accurately assess the economics of workload migration across clusters, and the lack of anti-jitter design in rescheduling decisions leading to poor system stability.

[0007] This invention is achieved through the following technical solution, providing a multi-cluster computing power cooperative scheduling method based on Karmada, comprising the following steps:

[0008] S1: Data collection: Real-time energy efficiency data is periodically collected from each member cluster by an energy data collector deployed on the control plane of the multi-cluster management platform. The energy efficiency data includes at least the power utilization efficiency (PUE) value of the data center where each member cluster is located and the real-time electricity price of the corresponding region.

[0009] S2: Calculate the cost score. When a workload scheduling request is received, the multi-dimensional comprehensive scoring engine uses the PUE value as a multiplicative factor and embeds it into the cost sub-score calculation of each candidate member cluster to correct the impact of the real-time electricity price on the actual IT equipment energy consumption cost, thereby obtaining a cost sub-score that reflects the differences in infrastructure energy efficiency.

[0010] S3: Based on the score, the multi-cluster management platform control plane generates a propagation strategy based on the comprehensive scheduling score result that includes at least the cost sub-score, and schedules the workload to the target member cluster with the best overall performance.

[0011] Specifically, the cost sub-score in step S2 is calculated based on the actual cost, and the formula for calculating the actual cost is as follows:

[0012] ActualCost(c i ) = Price i × PUE i × P_IT i

[0013] Where ActualCost is the actual cost, c i For member clusters, Price i For member cluster c i Real-time electricity price in the area, PUE i For member cluster c i The real-time energy utilization efficiency value, P_IT i For member cluster c i Estimated power consumption of IT equipment;

[0014] The cost sub-score is inversely proportional to the ActualCost.

[0015] Specifically, this also includes the following rescheduling processes:

[0016] S4: Trigger condition monitoring, the rescheduling monitor continuously monitors the energy efficiency status changes of the member clusters where the deployed workloads are located. When the preset energy efficiency deterioration trigger condition is met, the rescheduling evaluation process is triggered.

[0017] S5: Quantitative assessment of migration costs. In the rescheduling assessment process, a migration cost model is introduced to quantitatively assess the migration overhead of migrating workloads from the source cluster to the target cluster.

[0018] S6: Net benefit decision and migration execution. The actual migration operation is only executed when the scheduling score gain of the target cluster is greater than the sum of the normalized migration cost and the preset anti-jitter threshold.

[0019] Specifically, the migration cost model described in step S5 includes data transmission cost, service interruption cost, and state reconstruction cost. The total migration cost is calculated using the following formula:

[0020] C migration = C transfer + C downtime + C rebuild

[0021] Among them, C migration For the total migration cost, C transfer For data transmission cost, for C downtime Service interruption cost, C rebuild State reconstruction cost;

[0022] Data transmission cost C transfer The calculation formula is:

[0023] C transfer = DataSize / Bandwidth×UnitCost

[0024] DataSize refers to the total amount of data that needs to be transferred across clusters during the migration process, including at least container images, persistent volume data, configuration mappings, and encrypted credentials.

[0025] Bandwidth refers to the available network bandwidth between the source cluster and the target cluster;

[0026] UnitCost refers to the cost of renting or using a unit of network bandwidth per unit of time.

[0027] Service interruption cost C downtime The calculation formula is:

[0028] C downtime = DowntimeDuration × PenaltyRate

[0029] DowntimeDuration refers to the duration during which the service cannot provide normal access to the outside world during the migration process;

[0030] PenaltyRate refers to the service level agreement penalty factor for workload.

[0031] State reconstruction cost C rebuild The calculation formula is:

[0032] C rebuild = StateSize × RebuildFactor

[0033] StateSize refers to the amount of state data that the application maintains at runtime in memory, local temporary storage, or cache.

[0034] RebuildFactor refers to the state reconstruction complexity coefficient, which reflects the resource consumption of an application rebuilding its runtime context from scratch.

[0035] Specifically, the penalty coefficient in the service interruption cost is set differently according to the workload type, including: the lowest penalty coefficient for stateless workloads; the highest penalty coefficient for stateful workloads; and zero penalty coefficient for batch processing tasks.

[0036] Specifically, the preset energy efficiency deterioration triggering conditions mentioned in S4 include at least one of the following: the real-time electricity price of the currently operating cluster fluctuates more than a preset threshold compared to the benchmark electricity price; the real-time PUE value of the currently operating cluster increases more than a preset threshold compared to the historical average benchmark value; the green electricity ratio of the currently operating cluster decreases more than a preset threshold compared to the dispatch benchmark.

[0037] Specifically, the rescheduling evaluation process described in S4 is subject to a cooling-off period mechanism: a preset cooling-off period is enforced after each rescheduling decision, during which no new migration process is initiated to suppress the back-and-forth migration oscillations caused by energy data fluctuations. Emergency rescheduling triggered by abnormal cluster health or energy price deviations exceeding extreme circuit breaker thresholds is not subject to the cooling-off period mechanism.

[0038] Specifically, the comprehensive scheduling score in S3 also includes a low-carbon sub-score, which is calculated based on a weighted combination of the real-time green electricity ratio of the power grid where the member cluster is located and the future predicted green electricity ratio. This sub-score is used to support time-dimensional delayed scheduling, spatial-dimensional cross-cluster scheduling, and joint scheduling of both.

[0039] Specifically, when the workload is a stateful workload and the migration cost exceeds the preset red line, a local scaling strategy is adopted instead of full migration. The local scaling strategy prioritizes scaling up the clusters that are geographically adjacent or have the lowest network latency, and establishes an asynchronous / synchronous data replication link between the source cluster and the target cluster at the application layer or storage layer. After the data status is caught up and the traffic is smoothly switched, the source cluster is scaled down.

[0040] It also includes a multi-cluster computing power collaborative scheduling system based on Karmada, comprising an energy-aware extension module deployed in the Karmada control plane. This energy-aware extension module is integrated in a plug-in manner, without intruding into the native Karmada kernel, and is used to execute the Karmada-based multi-cluster computing power collaborative scheduling method described above. The energy-aware extension module includes:

[0041] The energy data acquisition unit serves as a data link between the system and the member cluster layer. Its input end connects to the PUE sensors, electricity price interfaces, and green electricity monitoring units of each cluster in the member cluster layer. It periodically collects real-time electricity prices, PUE energy efficiency indicators, green electricity ratio data of the power grid, cluster resource utilization rate, and health status data of each member cluster. Its output end synchronizes the standardized data to the multi-dimensional comprehensive scoring engine and rescheduling monitor.

[0042] The multi-dimensional comprehensive scoring engine receives all the data reported by the energy data collector, completes cluster pre-filtering through hard constraints, calculates the comprehensive score of each candidate cluster based on a preset four-dimensional weighted scoring model of cost, low carbon, resources and performance, and outputs the scoring results to the policy controller and Karmada scheduling execution component.

[0043] The policy controller serves as a bridge between energy sensing capabilities and the native scheduling system of Karmada. It receives the cluster scoring results from the multi-dimensional comprehensive scoring engine, transforms the computing and power collaborative decision-making logic into native scheduling rules that Karmada can recognize, and injects the optimal cluster selection, replica allocation, and differentiated configuration strategies into the Karmada Scheduler by dynamically generating and updating propagation and coverage strategies, thereby achieving lossless injection of energy sensing capabilities into the native scheduling system.

[0044] The rescheduling monitor serves as the core unit for system closed-loop optimization. It continuously receives real-time data from the energy data acquisition unit, monitors the clusters carrying the running workloads, and determines whether the electricity price, PUE, green electricity ratio, cluster load, and health status indicators meet the preset rescheduling trigger conditions. After triggering the rescheduling assessment, it performs a quantitative verification of scheduling benefits and migration costs using a migration cost model. Once the verification is successful, it issues a rescheduling instruction to the Karmada scheduling execution component and the policy controller, driving the scheduling policy update and workload cross-cluster migration.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] (1) This invention establishes the linear amplification effect of PUE on IT operating costs by embedding the PUE value as a multiplicative factor into the cost sub-scoring calculation formula. Compared with the prior art that judges only based on electricity price or only regards PUE as a static constraint, this invention can accurately identify the total cost disadvantage of clusters with low electricity price but high PUE, avoid scheduling decision bias, and ensure that the workload is always scheduled to the cluster with the lowest overall operating cost.

[0047] (2) By introducing a three-component migration cost model consisting of data transmission cost, service interruption cost and state reconstruction cost, and combining it with the net benefit decision rule, this invention upgrades cross-cluster rescheduling from qualitative judgment to quantitative calculation, ensuring that the benefits of each migration operation are greater than the costs, and effectively reducing the damage to business continuity caused by invalid migration.

[0048] (3) Energy efficiency triggering conditions covering multiple dimensions such as electricity price fluctuations, PUE deterioration, green electricity ratio decline, and excessive load are set, and a cooling period is enforced after each rescheduling, constraining the migration frequency from both time and revenue dimensions. This design effectively avoids the repeated migration of workload caused by instantaneous fluctuations in energy data while dynamically capturing the optimization window, thus balancing the two major goals of energy efficiency optimization and system stability. Attached Figure Description

[0049] Figure 1 This is an architecture diagram of the multi-cluster computing and power coordination scheduling system based on Karmada according to the present invention. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0051] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature; in the description of this application, unless otherwise stated, "multiple" means two or more.

[0052] This invention provides a multi-cluster computing power collaborative scheduling method and system based on Karmada. For ease of understanding, the core concepts and overall architecture involved in this invention will first be explained.

[0053] Karmada (Kubernetes Armada) is an open-source, cloud-native multi-cluster container orchestration platform. Its core design goal is to allow users to manage applications distributed across multiple Kubernetes clusters as if they were a single Kubernetes cluster. The Karmada architecture consists of a control plane and multiple member clusters. The control plane includes core components such as the API Server (Application Programming Interface Server), Scheduler, and Control Manager. The member clusters are the Kubernetes clusters that actually run the business applications. Users submit application deployment resource objects and PropagationPolicies to the Karmada control plane. The Karmada Scheduler selects target member clusters for the application based on the policy, and the controller then distributes the application resources to the corresponding member clusters for execution.

[0054] This invention, while fully retaining Karmada's native multi-cluster scheduling capabilities, adds an energy sensing extension module in a plug-in format, opening up a collaborative link between computing power scheduling and power data, and realizing cross-cluster intelligent scheduling that takes into account business service quality, power cost control and low-carbon emission reduction goals.

[0055] like Figure 1 As shown, the multi-cluster computing and power collaborative scheduling system based on Karmada provided in this embodiment of the invention adopts a three-layer architecture, which consists of the user / application layer, the central cluster Karmada control plane layer, and the member cluster layer from top to bottom.

[0056] The user / application layer serves as the unified access point for the system, comprising three core access units: the kubectl operation unit, the CI / CD tool unit, and the workload submission unit. It receives manual operation commands from operations and maintenance personnel, deployment triggers from automated delivery toolchains, and submissions of all types of cloud-native workloads, including stateless workloads (Deployments), stateful workloads (StatefulSets), and batch processing tasks (Jobs / CronJobs). After standardizing and encapsulating the requests, it distributes them to the Karmada API Server.

[0057] The central cluster Karmada control plane layer is the core of the system's scheduling decision-making, consisting of two parts: the native Karmada scheduling component and the newly added energy sensing extension module.

[0058] Karmada's native scheduling components include: Karmada API Server, which serves as the unified access gateway for the control plane, receiving all requests from the user / application layer and completing authentication, authorization, and distribution; Karmada Control Manager, which serves as the core execution unit for scheduling decisions and coordinates the lifecycle management of the entire scheduling process; and Karmada Scheduler, which has built-in PropagationPolicy configuration capabilities and executes specific multi-cluster workload distribution actions.

[0059] The new energy sensing extension module is integrated as a plug-in, without intruding on the native Karmada kernel, and includes four sub-modules:

[0060] The energy data acquisition unit serves as a data link between the system and the member cluster layer. Its input end connects to the PUE sensors, electricity price interfaces, and green electricity monitoring units of each cluster in the member cluster layer. It periodically collects real-time electricity prices, PUE energy efficiency indicators, green electricity ratio data of the power grid, cluster resource utilization rate, and health status data of each member cluster. Its output end synchronizes the standardized data to the multi-dimensional comprehensive scoring engine and re-dispatch monitor.

[0061] The multi-dimensional comprehensive scoring engine receives all the data reported by the energy data acquisition unit, first completes the cluster pre-filtering through hard constraints, and then calculates the comprehensive score of each candidate cluster based on the preset four-dimensional weighted scoring model of cost, low carbon, resources and performance, and outputs the scoring results to the policy controller and Karmada Control Manager.

[0062] The policy controller, acting as a bridge between energy sensing capabilities and the native scheduling system of Karmada, receives cluster scoring results from the multi-dimensional comprehensive scoring engine, transforms the computing and power collaborative decision-making logic into native scheduling rules that Karmada can recognize, and injects the optimal cluster selection, replica allocation, and differentiated configuration strategies into the Karmada Scheduler by dynamically generating and updating PropagationPolicy and OverridePolicy.

[0063] The rescheduling monitor, as the core unit of the system closed-loop optimization, continuously receives real-time data from the energy data acquisition unit and monitors the running workload cluster 24 / 7. It determines whether each energy indicator meets the preset rescheduling trigger conditions. After triggering the rescheduling assessment, it completes the quantitative verification of "scheduling benefits - migration costs" in combination with the migration cost model. After the verification is passed, it issues a rescheduling instruction to the Karmada Control Manager and the policy controller.

[0064] The member cluster layer is the actual operating environment of the workload, consisting of multiple geographically distributed member clusters (cluster A, cluster B, cluster C, etc.), which can support heterogeneous cluster access across regions, clouds, and data centers. Each member cluster is equipped with a PUE sensor, electricity price interface, and green electricity monitoring unit, and undertakes the dual responsibility of reporting energy data to the energy data acquisition unit and issuing scheduling policies to the execution center cluster.

[0065] This embodiment also provides a multi-cluster computing power collaborative scheduling method based on Karmada, which can be applied to the above system. This method includes two stages: an initial scheduling process and a rescheduling process.

[0066] The initial scheduling process completes the entire process of a workload going from submission to deployment to the optimal cluster, including:

[0067] S1: Data collection: Real-time energy efficiency data is periodically collected from each member cluster by an energy data collector deployed on the control plane of the multi-cluster management platform. The energy efficiency data includes at least the power utilization efficiency (PUE) value of the data center where each member cluster is located and the real-time electricity price of the corresponding region.

[0068] The energy data acquisition unit periodically collects real-time energy efficiency data from each member cluster. Specifically, it obtains the PUE value of each data center in real time by connecting to the DCIM (Data Center Infrastructure Management) interface of the data center where each member cluster is located, or by deploying PUE sensors at key power nodes; and obtains the real-time electricity price of the region where each member cluster is located by connecting to the external electricity market interface. At the same time, the energy data acquisition unit also collects operational data such as resource utilization (CPU utilization, memory utilization) and health status of each member cluster, and obtains the real-time green electricity ratio of the power grid where each cluster is located through the green electricity monitoring unit.

[0069] In this embodiment, the PUE value is not a static value; it dynamically changes with seasonal changes, variations in cooling load due to diurnal temperature differences, and fluctuations in the overall load rate of IT equipment. The energy data acquisition unit can capture these fluctuations and update the energy status view of each cluster at fixed intervals (e.g., every minute) to ensure the real-time performance and accuracy of subsequent scheduling decisions.

[0070] S2: Calculate the cost score. When a workload scheduling request is received, the multi-dimensional comprehensive scoring engine uses the PUE value as a multiplicative factor and embeds it into the cost sub-score calculation of each candidate member cluster to correct the impact of the real-time electricity price on the actual IT equipment energy consumption cost, thereby obtaining a cost sub-score that reflects the differences in infrastructure energy efficiency.

[0071] Specifically, when a user submits a workload scheduling request to the Karmada API Server via kubectl or CI / CD tools, the Karmada API Server routes the request to the Karmada Control Manager, which then triggers the multidimensional comprehensive scoring engine to perform cluster score calculation.

[0072] The multi-dimensional comprehensive scoring engine first performs cluster pre-filtering (hard constraint filtering): it checks whether the IT resources (CPU, memory, storage) of each member cluster meet the workload's request values, whether the cluster meets the business-defined SLA constraints, and the current health status of the cluster. Only clusters that pass the pre-filtering are considered valid candidate clusters.

[0073] The multi-dimensional comprehensive scoring engine performs weighted scoring and ranking on the pre-filtered candidate clusters. This embodiment uses a four-dimensional weighted scoring model, and the comprehensive scoring formula is as follows:

[0074] Score(c i ) = w cost × S cost (c i ) + w carbon × Scarbon (c i ) + w resource × S resource (c i ) + w perf × S perf (c i )

[0075] Score(c i S is the overall score. cost (c i S is the cost sub-score. carbon (c i ) represents the low-carbon fraction score, S resource (c i S represents the resource sub-score. perf (c i ) represents the performance sub-score, w cost As the weight for the cost dimension, w carbon As the weight for the low-carbon dimension, w resource For the weights of the resource dimension, w perf Weights for performance dimension

[0076] Each sub-score is normalized to the [0, 1] interval; a higher score indicates better performance in that dimension. The cost sub-score in this embodiment embeds the PUE value as a multiplicative factor in cost calculation, reflecting the linear amplification effect of infrastructure efficiency on IT operating costs. Specifically, the actual cost of each candidate cluster is first calculated:

[0077] ActualCost(c i ) = Price i ×PUE i ×P_IT i

[0078] Where ActualCost is the actual cost, c i For member clusters, Price i For member cluster c i Real-time electricity price in the area, PUE i For member cluster c i The real-time energy utilization efficiency value, P_IT i For member cluster c i Estimated power consumption of IT equipment;

[0079] The cost sub-score is inversely proportional to the Actual Cost, that is, the lower the Actual Cost, the lower the S. cost (c i The higher the value, the better.

[0080] This embodiment provides a set of numerical examples. Assume clusters A and B are two candidate clusters, and the estimated IT equipment power consumption P_IT for the workload is 1 kWh:

[0081] Cluster A: Real-time electricity price A = 0.5 yuan / kWh, PUE A = 1.2. Calculated using the multiplicative mechanism: ActualCost(A) = 0.5 × 1.2 × 1 = 0.6 yuan / kWh.

[0082] Cluster B: Real-time electricity price B = 0.4 yuan / kWh (better than A in terms of electricity price alone), PUE B = 1.8. Calculation: ActualCost(B) = 0.4 × 1.8 × 1 = 0.72 yuan / kWh.

[0083] Although cluster B has an advantage in base electricity price, its poor data center energy efficiency (high PUE) results in a higher total energy cost per unit of IT power consumed compared to cluster A. This embodiment successfully identifies cluster A as the more cost-effective scheduling target by using the PUE multiplicative factor, with the corresponding cost sub-score S_cost(A) > S_cost(B).

[0084] The calculation methods for other sub-scores are as follows: Low-carbon sub-score S carbon (c i Based on the real-time green electricity ratio assessment of the power grid where the cluster is located, the higher the green electricity ratio, the higher the score; resource sub-score S resource (c i A comprehensive evaluation of the cluster's CPU, memory utilization, and remaining capacity aims to balance the load across clusters; performance sub-score S perf (c i It measures network latency, geographical distance, and historical health metrics between clusters.

[0085] This embodiment can use three preset weighting strategies for users to choose from according to their business needs: cost-first strategy (w cost =0.5, w carbon =0.1, w resource =0.25, w perf =0.15), suitable for businesses sensitive to operating costs; low-carbon priority strategy (w cost =0.15, w carbon =0.5, w resource =0.2, w perf =0.15), suitable for businesses with strict carbon emission targets; Balanced strategy (w cost=0.3, w carbon =0.3, w resource =0.25, w perf =0.15), taking into account cost, low carbon and resource balance.

[0086] S3: Based on the score, the multi-cluster management platform control plane generates a propagation strategy based on the comprehensive scheduling score result that includes at least the cost sub-score, and schedules the workload to the target member cluster with the best overall performance.

[0087] Specifically, the multi-dimensional comprehensive scoring engine synchronizes the comprehensive scoring results of each candidate cluster to the policy controller and Karmada Control Manager. Based on the scoring results, the policy controller transforms the decision logic of computing and power coordination into native scheduling rules that Karmada can recognize, dynamically generates PropagationPolicy, clarifies the target cluster and replica weight allocation of the workload, and generates OverridePolicy to handle configuration differences between clusters (such as differential injection of environment variables, storage classes, etc.).

[0088] Karmada Control Manager receives policy instructions and drives Karmada Scheduler to execute specific scheduling actions, accurately distributing workloads to the target member clusters with the best overall cost. For example, in the numerical example above, cluster A has a higher overall cost score, so the workload will be scheduled to cluster A.

[0089] While workloads are running on the target cluster, the system does not stop optimizing; instead, it dynamically manages existing workloads throughout their entire lifecycle through a rescheduling process. This rescheduling process, led by a rescheduling monitor, strikes a balance between finding better cluster opportunities and maintaining system stability. The rescheduling process includes:

[0090] S4: Trigger condition monitoring, the rescheduling monitor continuously monitors the energy efficiency status changes of the member clusters where the deployed workloads are located. When the preset energy efficiency deterioration trigger condition is met, the rescheduling evaluation process is triggered.

[0091] The rescheduling monitor continuously receives real-time data from the energy data acquisition unit and polls to monitor the energy efficiency status of the member clusters containing the deployed workloads. This embodiment pre-defines the following five energy efficiency degradation trigger conditions; when any one of these conditions is met, the rescheduling evaluation process is triggered:

[0092] (1) Electricity price fluctuation trigger: The real-time electricity price of the currently running cluster fluctuates more than the preset threshold θ compared with the benchmark electricity price at the time of the initial scheduling of the workload. price (Default value 20%), i.e., |Price current - Pricebaseline | / Price baseline ≥ 20%; of which, Price current Price is the real-time electricity price for the currently running cluster. baseline This is the base electricity price for the initial scheduling of the workload.

[0093] (2) PUE deterioration trigger: The real-time PUE value of the currently running cluster increases by more than the preset threshold θ compared with its historical average baseline value. pue (Default value 10%), i.e. (PUE) current - PUE baseline ) / PUE baseline ≥ 10%; of which PUE current This represents the real-time PUE value of the currently running cluster. baseline This is the historical average baseline PUE for this cluster.

[0094] (3) Green electricity ratio decrease trigger: The green electricity ratio of the currently running cluster decreases by more than the preset threshold θ compared with the scheduling baseline. green (Default value 15%), i.e., GreenRatio baseline - GreenRatio current ≥ 15%; GreenRatio baseline GreenRatio serves as the baseline for green electricity allocation during dispatch. current This represents the current real-time percentage of green electricity.

[0095] (4) High cluster load trigger: The CPU or memory utilization of the currently running cluster exceeds the preset proportion θ of the planned threshold. load (Default value 30%) or higher, Load current (c src Load planned (c src ≥30%; Load current (c src ) represents the current actual load of the source cluster. Load current (c src ) represents the planned load threshold for the source cluster.

[0096] (5) Cluster health anomaly trigger: When the health of the currently running cluster decreases due to node failure, network partition, etc., an emergency rescheduling is immediately triggered, without any threshold or cooldown period restrictions.

[0097] All of the above thresholds are user-configurable parameters, and in this embodiment, the above values ​​are the system default values.

[0098] S5: Quantitative assessment of migration costs. After the rescheduling assessment process is triggered, the multi-dimensional comprehensive scoring engine re-evaluates all candidate clusters. Simultaneously, a migration cost model is introduced to quantitatively assess the workload shifting from the source cluster c. src Migrate to target cluster c dst Migration overhead.

[0099] The migration cost model includes data transfer costs, service interruption costs, and state reconstruction costs. The total migration cost is calculated using the following formula:

[0100] C migration = C transfer + C downtime + C rebuild

[0101] Among them, C migration For the total migration cost, C transfer For data transmission costs, C downtime For service interruption costs, C rebuild Cost of state reconstruction;

[0102] Data transmission cost C transfer The calculation formula is:

[0103] C transfer = DataSize / Bandwidth×UnitCost

[0104] DataSize refers to the total amount of data that needs to be transferred across clusters during the migration process, including at least container images, persistent volume data, configuration mappings, and encrypted credentials.

[0105] Bandwidth refers to the available network bandwidth between the source cluster and the target cluster;

[0106] UnitCost refers to the cost of renting or using a unit of network bandwidth per unit of time.

[0107] Service interruption cost C downtime The calculation formula is:

[0108] C downtime = DowntimeDuration × PenaltyRate

[0109] DowntimeDuration refers to the duration during which the service cannot be accessed normally.

[0110] PenaltyRate refers to the service level agreement penalty factor for workload.

[0111] PenaltyRate is set differently based on workload type: Stateless workloads (Deployment) support rolling updates, and can achieve zero or very short downtime during migration by starting a new replica and then terminating the old replica, resulting in the lowest penalty coefficient; Stateful workloads (StatefulSet) involve disk mount switching and data consistency checks, resulting in longer downtime and the highest penalty coefficient; Batch processing tasks (Job / CronJob) have discrete execution and can wait for the current execution cycle to end before migration, so the penalty coefficient is set to zero.

[0112] State reconstruction cost C rebuild The calculation formula is:

[0113] C rebuild = StateSize × RebuildFactor

[0114] StateSize refers to the amount of state data that the application maintains at runtime in memory, local temporary storage, or cache.

[0115] RebuildFactor refers to the state reconstruction complexity coefficient, which reflects the resource consumption of an application rebuilding its runtime context from scratch.

[0116] S6: Net benefit decision and migration execution. The actual migration operation is only executed when the scheduling score gain of the target cluster is greater than the sum of the normalized migration cost and the preset anti-jitter threshold.

[0117] To ensure the economic rationality of each migration decision, this embodiment introduces an "effective score." The effective score represents the net competitiveness of the target cluster after deducting migration losses.

[0118] EffectiveScore(c dst = Score(c dst -α×NormCost

[0119] Where NormCost = C migration / C max For the normalized migration cost, C max The maximum tolerable migration cost set for the system; α is the migration sensitivity weighting coefficient; EffectiveScore(c dst Score(c) represents the effective score of the target cluster. dst ) represents the original overall score of the target cluster, and NormCost represents the normalized migration cost.

[0120] Migration gain is defined as the difference between the target cluster score and the source cluster score:

[0121] Benefit = Score(c dst ) - Score(c src )

[0122] Where Benefit is the migration benefit, and Score(c src Score(c) represents the original score of the target cluster. src () represents the original score of the source cluster.

[0123] Ultimately, a migration plan will be generated and the actual migration operation will be executed only if the following criteria are met:

[0124] Benefit > NormCost + Threshold min

[0125] Among them, Threshold min The preset minimum revenue threshold (anti-jitter threshold) is designed to prevent workloads from frequently migrating between clusters due to minor fluctuations in electricity prices or resource ratings, thus playing a role in resisting shocks and protecting system stability.

[0126] Once the benefit assessment is approved, the rescheduling monitor sends a rescheduling command to the Karmada Control Manager and the policy controller. The policy controller updates the target cluster configuration and replica distribution in the PropagationPolicy, and the Karmada Control Manager drives the Karmada Scheduler to perform the actual cross-cluster migration, completing the switch of the workload from the source cluster to the target cluster.

[0127] Cooldown mechanism

[0128] To prevent workloads from migrating back and forth between two clusters due to fluctuations in energy data (oscillation phenomenon), this embodiment employs a cooling-off period mechanism. After each rescheduling decision (regardless of whether an actual migration is ultimately performed), the system enforces a cooling-off period T. cooldown (Default value 5 to 15 minutes, configurable). During the cooldown period, the rescheduling monitor continues to collect data and compare thresholds, but will not initiate a new migration process. Only when t_ is met... now > t_ last_reschedule + T cooldown Only after this time will the system allow the next migration to proceed. After the cooldown period ends, update t_ last_reschedule The timestamp records the rescheduling decision log for this round, marking the start of the next monitoring cycle. Specifically, t_ now t_ is the current time. last_reschedule The time when the last rescheduling occurred.

[0129] The only exception is when the cluster health is abnormal or the energy price deviates beyond the extreme circuit breaker threshold (Category 5 above). This trigger will bypass the cooling-off period limit and perform a forced migration to ensure business continuity.

[0130] Green electricity forecasting and dispatch

[0131] Preferably, the comprehensive scheduling score in step S3 may also include a low-carbon sub-score based on green electricity prediction to support more refined spatiotemporal joint scheduling.

[0132] The multi-dimensional comprehensive scoring engine obtains green electricity prediction data of the power grid where each cluster is located through external interfaces, including the current real-time green electricity ratio. current (c i ) and future time period T horizon The projected green ratio within the country predicted (c i, t+T horizon It should be noted that the system of this invention does not directly participate in the predictive modeling of the power system, nor does it assume predictive responsibilities related to machine learning or deep learning. Instead, it acts as a data consumer, acquiring external predictive data through a standardized interface. External data sources include, but are not limited to, power grid dispatching systems, third-party new energy prediction platforms, or meteorological and energy monitoring centers.

[0133] Low carbon fraction score S carbon (c i A weighted combination of real-time and predicted values ​​is used:

[0134] S carbon (c i ) = α × GreenRatio current (c i )+β×GreenRatio predicted (c i, t+T horizon )

[0135] Here, α and β are weighting coefficients, satisfying α+β= 1. The ratio of α to β can be flexibly adjusted according to the degree of dependence on real-time data and predicted data.

[0136] Based on the support of low-carbon sub-scores, this embodiment supports three green electricity sensing and scheduling modes:

[0137] (1) Time-based scheduling (delayed scheduling): Suitable for batch processing tasks with low time sensitivity (such as big data processing jobs and periodic CronJobs). When the system detects that the weighted average green electricity ratio in the future prediction period is significantly higher than the current ratio (i.e., GreenRatio), it can be used for these tasks. predicted - GreenRatio current> δ_ green ,δ_ green If a preset green electricity revenue threshold is set, the task will be placed in a waiting state until the predicted high green electricity window arrives before it is started, thereby avoiding the peak of high carbon intensity electricity load.

[0138] (2) Spatial Dimension Scheduling (Cross-Cluster Scheduling): Applicable to general workloads with certain requirements for execution time but no strict restrictions on geographical location. The scheduler compares the current green electricity ratio of different regional clusters in real time and routes the load to the cluster with the highest current green electricity ratio. For example, when the western cluster is at its peak photovoltaic output while the eastern cluster still mainly relies on thermal power, the system will automatically prioritize the western cluster for deployment, realizing "moving with green".

[0139] (3) Joint time-space scheduling: Suitable for highly elastic workloads that have both execution time flexibility and cross-regional distribution capabilities. The system comprehensively evaluates the distribution differences in the spatial dimension and the predicted evolution in the temporal dimension, calculates the optimal (cluster, time) combination, and minimizes the carbon footprint.

[0140] When the external prediction interface becomes inaccessible or the returned data does not meet quality standards, the system automatically enters degradation mode: GreenRatio is discontinued. predicted The variable degenerates into one that depends only on GreenRatio. current The real-time green electricity dispatch mode is used; if real-time data is also unavailable, it further degenerates into a conventional dispatch mode based on traditional PUE values ​​or remaining resources, ensuring that the basic functions of computing power dispatch are not rendered ineffective due to the interruption of external energy forecast data.

[0141] Special handling strategies for stateful loads

[0142] StatefulSets, due to their inherent state persistence characteristics, face challenges such as PVC binding restrictions and ordering requirements when migrating across clusters. For this scenario, when the migration cost C calculated in step S5... migration If the preset red line is exceeded, or if the underlying storage facility does not support cross-cluster volume replication, this embodiment adopts the "nearest expansion" strategy instead of full migration.

[0143] Specifically, the scheduler no longer forces a full migration, but instead prioritizes scaling out clusters that are geographically adjacent or have the lowest network latency. After the target cluster successfully creates Pods and is ready, traffic is smoothly switched through a front-end traffic distributor (such as Service Mesh or Global Ingress). Scaling in is then performed once the traffic in the source cluster has returned to zero, thus achieving a near-migration effect while ensuring data security and business continuity.

[0144] This invention extends the Karmada control plane through plug-in architecture, adding four core modules: an energy data collector, a multi-dimensional comprehensive scoring engine, a policy controller, and a rescheduling monitor. This constructs a complete closed-loop chain of "energy data collection - multi-dimensional quantitative scoring - policy transformation and execution - dynamic rescheduling optimization," achieving the goal of computing-coordinated scheduling that effectively reduces operating costs and increases the proportion of green electricity consumption while ensuring business SLA.

[0145] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. This application is not limited to the exact structures described above and illustrated in the accompanying drawings, and it should not be considered that the specific implementation of this application is limited to these descriptions. For those skilled in the art, various changes and modifications made without departing from the concept of this application should be considered to fall within the protection scope of this application.

Claims

1. A multi-cluster computing power cooperative scheduling method based on Karmada, characterized in that, Includes the following steps: S1: Data collection: Real-time energy efficiency data is periodically collected from each member cluster by an energy data collector deployed on the control plane of the multi-cluster management platform. The energy efficiency data includes at least the power utilization efficiency (PUE) value of the data center where each member cluster is located and the real-time electricity price of the corresponding region. S2: Calculate the cost score. When a workload scheduling request is received, the multi-dimensional comprehensive scoring engine uses the PUE value as a multiplicative factor and embeds it into the cost sub-score calculation of each candidate member cluster to correct the impact of the real-time electricity price on the actual IT equipment energy consumption cost, thereby obtaining a cost sub-score that reflects the differences in infrastructure energy efficiency. S3: Based on the score, the multi-cluster management platform control plane generates a propagation strategy based on the comprehensive scheduling score result that includes at least the cost sub-score, and schedules the workload to the target member cluster with the best overall performance.

2. The multi-cluster computing power cooperative scheduling method based on Karmada according to claim 1, characterized in that, The cost sub-score in step S2 is calculated based on the actual cost, and the formula for calculating the actual cost is as follows: ActualCost(c i ) = Price i × PUE i × P_IT i Where ActualCost is the actual cost, c i For member clusters, Price i For member cluster c i Real-time electricity price in the area, PUE i For member cluster c i The real-time energy utilization efficiency value, P_IT i For member cluster c i Estimated power consumption of IT equipment; The cost sub-score is inversely proportional to the ActualCost.

3. The multi-cluster computing power cooperative scheduling method based on Karmada according to claim 1, characterized in that, It also includes the following rescheduling processes: S4: Trigger condition monitoring, the rescheduling monitor continuously monitors the energy efficiency status changes of the member clusters where the deployed workloads are located. When the preset energy efficiency deterioration trigger condition is met, the rescheduling evaluation process is triggered. S5: Quantitative assessment of migration costs. In the rescheduling assessment process, a migration cost model is introduced to quantitatively assess the migration overhead of migrating workloads from the source cluster to the target cluster. S6: Net benefit decision and migration execution. The actual migration operation is only executed when the scheduling score gain of the target cluster is greater than the sum of the normalized migration cost and the preset anti-jitter threshold.

4. The multi-cluster computing power cooperative scheduling method based on Karmada according to claim 3, characterized in that, The migration cost model described in step S5 includes data transmission cost, service interruption cost, and state reconstruction cost. The total migration cost is calculated using the following formula: C migration = C transfer + C downtime + C rebuild Among them, C migration For the total migration cost, C transfer For data transmission costs, C downtime For service interruption costs, C rebuild Cost of state reconstruction; Data transmission cost C transfer The calculation formula is: C transfer = DataSize / Bandwidth×UnitCost DataSize refers to the total amount of data that needs to be transferred across clusters during the migration process, including at least container images, persistent volume data, configuration mappings, and encrypted credentials. Bandwidth refers to the available network bandwidth between the source cluster and the target cluster; UnitCost refers to the cost of renting or using a unit of network bandwidth per unit of time. Service interruption cost C downtime The calculation formula is: C downtime = DowntimeDuration × PenaltyRate DowntimeDuration refers to the duration during which the service cannot provide normal access to the outside world during the migration process; PenaltyRate refers to the service level agreement penalty factor for workload. State reconstruction cost C rebuild The calculation formula is: C rebuild = StateSize×RebuildFactor StateSize refers to the amount of state data that the application maintains at runtime in memory, local temporary storage, or cache. RebuildFactor refers to the state reconstruction complexity coefficient, which reflects the resource consumption of an application rebuilding its runtime context from scratch.

5. The multi-cluster computing power cooperative scheduling method based on Karmada according to claim 4, characterized in that: The penalty coefficient in the service interruption cost is set differently based on the workload type, including: Stateless loads have the lowest penalty coefficient; Stateful loads have the highest penalty coefficient; The penalty coefficient for batch processing tasks is zero.

6. The multi-cluster computing power cooperative scheduling method based on Karmada according to claim 3, characterized in that, The preset energy efficiency degradation triggering conditions described in S4 include at least one of the following: The real-time electricity price of the currently running cluster fluctuates more than a preset threshold compared to the benchmark electricity price; The real-time PUE value of the currently running cluster has increased by more than a preset threshold compared to the historical average baseline value; The green electricity ratio of the currently running cluster has decreased by more than the preset threshold compared to the scheduling baseline.

7. The multi-cluster computing power cooperative scheduling method based on Karmada according to claim 3, characterized in that, The rescheduling evaluation process described in S4 is subject to a cooling-off period mechanism: a preset cooling-off period is enforced after each rescheduling decision, during which no new migration process is initiated to suppress the back-and-forth migration oscillations caused by energy data fluctuations. Emergency rescheduling triggered by abnormal cluster health or energy price deviations exceeding extreme circuit breaker thresholds is not subject to the cooling-off period mechanism.

8. The multi-cluster computing power cooperative scheduling method based on Karmada according to claim 1, characterized in that: The comprehensive scheduling score in S3 also includes a low-carbon sub-score, which is calculated based on a weighted combination of the real-time green electricity ratio of the power grid where the member cluster is located and the future predicted green electricity ratio. It is used to support time-dimensional delayed scheduling, spatial-dimensional cross-cluster scheduling, and joint scheduling of the two.

9. The multi-cluster computing power cooperative scheduling method based on Karmada according to claim 4, characterized in that, When the workload is a stateful workload and the migration cost exceeds the preset red line, a local expansion strategy is adopted instead of full migration. The local expansion strategy prioritizes expansion in the clusters that are geographically adjacent or have the lowest network latency, and establishes an asynchronous / synchronous data replication link between the source cluster and the target cluster at the application layer or storage layer. After the data status is caught up and the traffic is smoothly switched, the source cluster is scaled down.

10. A multi-cluster computing power collaborative scheduling system based on Karmada, characterized in that, This includes an energy-aware extension module deployed in the Karmada control plane. The energy-aware extension module is integrated in a plug-in manner, without intruding into the native Karmada kernel, and is used to execute the Karmada-based multi-cluster computing power cooperative scheduling method as described in any one of claims 1 to 9. The energy-aware extension module includes: The energy data acquisition unit serves as a data link between the system and the member cluster layer. Its input end connects to the PUE sensors, electricity price interfaces, and green electricity monitoring units of each cluster in the member cluster layer. It periodically collects real-time electricity prices, PUE energy efficiency indicators, green electricity ratio data of the power grid, cluster resource utilization rate, and health status data of each member cluster. Its output end synchronizes the standardized data to the multi-dimensional comprehensive scoring engine and rescheduling monitor. The multi-dimensional comprehensive scoring engine receives all the data reported by the energy data collector, completes cluster pre-filtering through hard constraints, calculates the comprehensive score of each candidate cluster based on a preset four-dimensional weighted scoring model of cost, low carbon, resources and performance, and outputs the scoring results to the policy controller and Karmada scheduling execution component. The policy controller serves as a bridge between energy sensing capabilities and the native scheduling system of Karmada. It receives the cluster scoring results from the multi-dimensional comprehensive scoring engine, transforms the computing and power collaborative decision-making logic into native scheduling rules that Karmada can recognize, and injects the optimal cluster selection, replica allocation, and differentiated configuration strategies into the Karmada Scheduler by dynamically generating and updating propagation and coverage strategies, thereby achieving lossless injection of energy sensing capabilities into the native scheduling system. The rescheduling monitor serves as the core unit for system closed-loop optimization. It continuously receives real-time data from the energy data acquisition unit, monitors the clusters carrying the running workloads, and determines whether the electricity price, PUE, green electricity ratio, cluster load, and health status indicators meet the preset rescheduling trigger conditions. After triggering the rescheduling assessment, it performs a quantitative verification of scheduling benefits and migration costs using a migration cost model. Once the verification is successful, it issues a rescheduling instruction to the Karmada scheduling execution component and the policy controller, driving the scheduling policy update and workload cross-cluster migration.