Energy efficiency optimization oriented intelligent computing center computing and resource collaborative scheduling method, system, device and medium

By constructing a multi-level resource and energy efficiency perception model and a full-stack energy efficiency assessment model for the intelligent computing center, and combining virtual digital twins and hybrid optimization algorithms, the collaborative scheduling of computing tasks and infrastructure resources of the intelligent computing center was realized, solving the fragmented and static problems of energy efficiency management and improving the system's energy efficiency and adaptability.

CN122111658APending Publication Date: 2026-05-29LIAOYANG POWER SUPPLY COMPANY OF STATE GRID LIAONING ELECTRIC POWER SUPPLY +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIAOYANG POWER SUPPLY COMPANY OF STATE GRID LIAONING ELECTRIC POWER SUPPLY
Filing Date
2026-02-05
Publication Date
2026-05-29

Smart Images

  • Figure CN122111658A_ABST
    Figure CN122111658A_ABST
Patent Text Reader

Abstract

The application discloses a method and system for computing and resource collaborative scheduling of a wisdom calculation center facing energy efficiency optimization, belongs to the technical field of data centers and high-performance computing, and comprises the following steps: establishing a multi-level resource and energy efficiency perception model of the wisdom calculation center to acquire real-time perception data; constructing a full-stack energy efficiency evaluation and prediction model based on the real-time perception data; solving an optimal collaborative scheduling strategy through a joint optimization engine based on the full-stack energy efficiency evaluation and prediction model, and outputting a prediction value according to the collaborative scheduling strategy; executing the collaborative scheduling strategy, monitoring actual energy consumption and task performance data after execution of the collaborative scheduling strategy, taking the difference between the monitoring result and the prediction value as a feedback signal, and inputting the feedback signal into the full-stack energy efficiency evaluation and prediction model to realize online self-adaptation and closed-loop optimization of the model. The application realizes optimization of the overall energy efficiency ratio of the wisdom calculation center by constructing a unified energy efficiency model and intelligent collaborative decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data center and high-performance computing technology, and in particular relates to a method, system, device and medium for intelligent computing center computing and resource collaborative scheduling for energy efficiency optimization, which is used for intelligent scheduling of computing tasks and infrastructure resources such as cooling and power supply. Background Technology

[0002] With the rapid development of artificial intelligence technology, especially the explosive growth in demand for large-scale deep learning model training and inference, intelligent computing centers, as key computing infrastructure, have reached unprecedented levels in scale, computing density, and energy consumption. Energy costs for intelligent computing centers now constitute a significant portion of their total lifecycle cost of ownership, and have brought severe carbon emission pressure. Against this backdrop, improving energy efficiency is no longer simply a matter of energy conservation, but has become a core issue concerning operational economics, environmental sustainability, and industry competitiveness. Traditional extensive energy management methods are no longer suitable for the high-intensity, heterogeneous, and dynamic computing load characteristics of intelligent computing centers, making refined energy efficiency optimization imperative.

[0003] Currently, the industry has explored various approaches to energy efficiency management in intelligent computing centers, but mainstream methods still have significant limitations. First, computing task scheduling systems and physical infrastructure management systems are typically deployed independently, operating in silos. The computing scheduler prioritizes job completion time and resource utilization, lacking awareness of the underlying hardware power consumption and the resulting cooling and power distribution needs. The infrastructure system passively responds to heat load and power changes caused by computing loads, lacking foresight regarding upstream computing load characteristics, leading to overall response lag and energy efficiency losses. Second, optimization dimensions are relatively singular and static. Existing optimization measures are mostly localized, such as adjusting chiller setpoints, increasing supply air temperature to reduce cooling energy consumption, or using heterogeneous hardware hybrid deployment. These methods lack synergistic consideration of dynamic factors. Third, the scheduling granularity is too coarse. Most scheduling systems use servers or virtual machines as the smallest resource unit, ignoring the energy efficiency interactions caused by different tasks being deployed together on the same server, as well as the differences in energy efficiency characteristics of tasks at different stages of operation.

[0004] The aforementioned fragmented, static, and coarse-grained management model results in a significant blind spot for optimizing the energy efficiency of intelligent computing centers. Oversupply or improper allocation of computing resources not only directly increases the energy consumption of IT equipment but also triggers unnecessary cooling and power distribution losses. Furthermore, infrastructure systems, unable to obtain accurate load forecasts, often adopt conservative operating strategies, further increasing non-IT energy consumption. This disconnect between computing and energy management loops makes it difficult for the actual operating power consumption (PUE) of intelligent computing centers to consistently approach the theoretical optimum. Therefore, a revolutionary method is urgently needed that can penetrate traditional management boundaries and achieve intelligent collaborative scheduling of computing load and infrastructure resources to achieve optimal system-level global energy efficiency. Summary of the Invention

[0005] In view of the shortcomings of the above or existing technologies, this invention proposes a method and system for collaborative scheduling of computing and resources in intelligent computing centers for energy efficiency optimization. This breaks down the barriers to resource scheduling of computing and infrastructure, and optimizes the overall energy efficiency ratio of intelligent computing centers by constructing a unified energy efficiency model and intelligent collaborative decision-making.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a method for collaborative scheduling of computing and resources in intelligent computing centers for energy efficiency optimization, comprising:

[0008] Establish a multi-level resource and energy efficiency perception model for the intelligent computing center to obtain real-time perception data;

[0009] Based on the real-time sensing data, a full-stack energy efficiency assessment and prediction model is constructed.

[0010] The predicted value is output based on the full-stack energy efficiency assessment and prediction model;

[0011] The collaborative scheduling strategy enables the model to achieve online adaptation and closed-loop optimization.

[0012] As a further preferred embodiment of the present invention, the multi-level resource and energy efficiency perception model of the intelligent computing center specifically includes:

[0013] The computing resource layer is used to collect real-time data on the load rate, hardware operating frequency, temperature, and real-time power consumption of computing nodes.

[0014] The task feature layer is used to extract the computing power requirements, memory and communication characteristics, priority and deadline constraints of the computing tasks to be scheduled.

[0015] The infrastructure resource layer is used to collect real-time data on the operating efficiency of the cooling and power distribution systems, ambient temperature and humidity, and external energy signals.

[0016] As a further preferred embodiment of the present invention, a full-stack energy efficiency assessment and prediction model is constructed based on the real-time sensing data; specifically including:

[0017] The joint optimization objective function aims to minimize the perceived effective energy consumption of the overall system ownership cost. The real-time perceived data is used as the input to the full-stack energy efficiency assessment and prediction model, and the output is a quantitative prediction of the overall energy efficiency performance in the future.

[0018] The joint optimization objective function incorporates a performance penalty term for computational tasks as a constraint or a weighted cost term.

[0019] By dynamically learning historical data through machine learning methods, we can quantitatively predict the total energy efficiency of the system under different combinations of resource allocation schemes and infrastructure control strategies.

[0020] The evaluation targets for the overall energy efficiency of the system include computational energy consumption, cooling energy consumption, power supply and distribution losses, and the cost of breach of service quality.

[0021] As a further preferred embodiment of the present invention, the construction of the full-stack energy efficiency assessment and prediction model further includes: constructing a virtual digital twin that is synchronized or quasi-synchronized with the physical intelligent computing center. The virtual digital twin is used to perform low-cost simulation, energy efficiency pre-assessment and risk assessment of candidate scheduling strategies generated by the joint optimization engine before they are put into execution in the physical system.

[0022] Among them, the virtual digital twin has thermodynamic simulation capabilities. Based on the predicted computing load distribution and equipment power consumption, it generates a three-dimensional dynamic temperature field and airflow organization simulation diagram inside the computer room, and optimizes the setpoint adjustment scheme of the cooling system based on this simulation diagram.

[0023] As a further preferred embodiment of the present invention, the step of outputting predicted values ​​based on the full-stack energy efficiency assessment and prediction model specifically includes: solving the optimal collaborative scheduling strategy through a joint optimization engine based on the full-stack energy efficiency assessment and prediction model, and outputting predicted values ​​according to the collaborative scheduling strategy.

[0024] As a further preferred embodiment of the present invention, the step of solving for the optimal cooperative scheduling strategy through a joint optimization engine and outputting a predicted value based on the cooperative scheduling strategy specifically includes:

[0025] In response to the arrival of new computing tasks or changes in system state, a collaborative optimization decision-making process is triggered. Based on a full-stack energy efficiency assessment and prediction model, the optimal collaborative scheduling strategy is solved through a joint optimization engine while satisfying task constraints. The collaborative scheduling strategy simultaneously outputs the placement scheme of computing tasks on heterogeneous hardware, the dynamic frequency adjustment command of computing devices, and the setpoint adjustment scheme of cooling and power supply systems.

[0026] The collaborative scheduling strategy explicitly responds to external dynamic signals; when the real-time electricity price is detected to be entering a peak period, the joint optimization engine dynamically adjusts the weights of the joint optimization objective function to generate a strategy to reduce real-time electricity costs.

[0027] As a further preferred embodiment of the present invention, before solving for the optimal cooperative scheduling strategy through a joint optimization engine and outputting the predicted value according to the cooperative scheduling strategy, the method further includes energy efficiency-aware task preprocessing; the energy efficiency-aware task preprocessing includes:

[0028] The queue of computing tasks waiting to be scheduled is subjected to feature analysis to identify multiple independent tasks with complementary computing features. These independent tasks are then intelligently combined into a composite task package, which is submitted as a whole to the collaborative optimization engine for unified resource allocation and scheduling to improve the overall resource utilization and energy efficiency of the cluster.

[0029] Specifically, the complementarity of the computing features refers to: identifying and combining a first task with high memory bandwidth requirements with a second task with high utilization of computing cores; when the first and second tasks share the same computing node, utilizing the heterogeneous hardware resources of that node to reduce resource idle time.

[0030] As a further preferred embodiment of the present invention, the online adaptation and closed-loop optimization of the model by executing the collaborative scheduling strategy specifically includes: executing the collaborative scheduling strategy, monitoring the actual energy consumption and task performance data after the execution of the collaborative scheduling strategy, using the difference between the monitoring results and the predicted values ​​as a feedback signal, and inputting it into the full-stack energy efficiency assessment and prediction model to realize the online adaptation and closed-loop optimization of the model.

[0031] As a further preferred embodiment of the present invention, the closed-loop optimization specifically includes: establishing a model performance monitor, and when the average deviation between the actual energy consumption data and the predicted value is continuously exceeded by a preset threshold, automatically triggering the incremental learning or retraining process of the energy efficiency assessment and prediction model, and updating the model parameters using the latest collected data to maintain prediction accuracy.

[0032] Secondly, this invention provides a smart computing center computing and resource collaborative scheduling system for energy efficiency optimization, comprising:

[0033] The perception model building unit is used to establish a multi-level resource and energy efficiency perception model for the intelligent computing center and to acquire real-time perception data.

[0034] The assessment and prediction model building unit constructs a full-stack energy efficiency assessment and prediction model based on the real-time sensing data.

[0035] The collaborative scheduling strategy unit outputs predicted values ​​based on the full-stack energy efficiency assessment and prediction model.

[0036] The monitoring and feedback optimization unit is used to execute the collaborative scheduling strategy to achieve online adaptation and closed-loop optimization of the model.

[0037] Thirdly, an electronic device includes: a processor and a memory; the memory is used to store a computer program, which, when executed by the processor, causes the electronic device to perform the energy-efficient intelligent computing center computing and resource collaborative scheduling method described in the first aspect.

[0038] Fourthly, a computer-readable storage medium includes: a computer program or instructions; when the computer program or instructions are executed on a computer, the computer causes the computer to perform the energy-efficiency-optimized intelligent computing center computing and resource collaborative scheduling method described in the first aspect.

[0039] The beneficial effects of this invention are as follows:

[0040] 1. This invention achieves a breakthrough from local static optimization to global dynamic collaboration. By incorporating computing task scheduling, hardware resource regulation, and infrastructure management into a unified joint optimization framework, and establishing a full-link digital twin model, it can accurately quantify the chain-like energy efficiency impact caused by any scheduling decision, thereby finding the globally optimal or near-optimal solution at the system level. This collaborative scheduling mechanism can effectively break down resource silos, solve energy efficiency losses caused by inter-system target conflicts or response delays, and achieve a significant approach of the PUE value to the theoretical limit.

[0041] 2. This invention endows the intelligent computing center with high adaptability and multi-objective balancing capabilities in complex and dynamic environments. By utilizing machine learning models for real-time learning and prediction, the system can proactively adapt to drastic fluctuations in task load, changes in the external environment, and the heterogeneity of hardware devices. It can not only achieve dynamic resource adjustments at the second to minute level, but also intelligently weigh and compromise among multiple sometimes conflicting objectives, ensuring the deadlines of core tasks while reducing operating costs, thus achieving a balance between economy and reliability.

[0042] 3. This invention uses digital twin technology to simulate, verify, and rehearse scheduling strategies in virtual space, reducing the risks and costs of energy efficiency optimization trial and error in production systems and ensuring the stability and security of data center operations. At the same time, the closed-loop feedback learning mechanism enables the system energy efficiency model to continuously evolve and adapt to new hardware architectures, task types, and operational strategies, ensuring the long-term viability of the technology. Attached Figure Description

[0043] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a flowchart illustrating the intelligent computing center computing and resource collaborative scheduling method for energy efficiency optimization provided by the present invention.

[0045] Figure 2 This invention provides a flowchart of the energy efficiency sensing task preprocessing process.

[0046] Figure 3 The flowchart of the intelligent computing center computing and resource collaborative scheduling system for energy efficiency optimization provided by the present invention is shown. Detailed Implementation

[0047] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0048] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0049] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "comprising" or "including," and similar terms as used in this disclosure, mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, but do not exclude other elements or objects. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0050] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that mutually excludes other embodiments. It should be noted that the embodiments of the present invention can be applied to any applicable scenario.

[0051] In this embodiment of the invention, "instruction" can include direct and indirect instructions, as well as explicit and implicit instructions. The information indicated by a certain piece of information is called the information to be instructed. In specific implementation, there are many ways to instruct the information to be instructed, such as, but not limited to, directly instructing the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly instruct the information to be instructed by instructing other information, where there is a correlation between the other information and the information to be instructed. It can also instruct only a part of the information to be instructed, while the other parts are known or pre-agreed upon. For example, the instruction of specific information can be achieved by using a pre-agreed (e.g., protocol-defined) arrangement of various pieces of information, thereby reducing instruction overhead to some extent. Simultaneously, common parts of various pieces of information can be identified and uniformly indicated to reduce the instruction overhead caused by individually indicating the same information.

[0052] Furthermore, the specific indication method can also be any existing indication method, such as, but not limited to, the above-mentioned indication methods and their various combinations. Specific details of various indication methods can be found in existing technologies, and will not be elaborated upon here. As described above, for example, when multiple pieces of information of the same type need to be indicated, the indication methods for different pieces of information may differ. In specific implementation, the required indication method can be selected according to specific needs. This embodiment of the invention does not limit the selected indication method; therefore, the indication methods involved in this embodiment of the invention should be understood to cover various methods that enable the party to be indicated to obtain the information to be indicated.

[0053] It should be understood that the information to be indicated can be sent as a whole or divided into multiple sub-information messages sent separately, and the sending period and / or timing of these sub-information messages can be the same or different. The specific sending method is not limited in this embodiment of the invention. The sending period and / or timing of these sub-information messages can be predefined, for example, according to a protocol, or configured by the sending device by sending configuration information to the receiving device.

[0054] "Predefined" or "pre-configured" can be achieved by pre-saving corresponding codes, tables, or other means that can be used to indicate relevant information in the device. This embodiment of the invention does not limit the specific implementation method. "Saving" can refer to saving in one or more memories. These memories can be separate installations or integrated into the encoder, decoder, processor, or electronic device. Alternatively, some memories can be separately installed, while others are integrated into the decoder, processor, or electronic device. The type of memory can be any form of storage medium, and this embodiment of the invention does not limit this.

[0055] In the embodiments of this invention, "protocol" may refer to a protocol family in the field of communication, a standard protocol with a similar protocol family frame structure, or a related protocol applied to a future intelligent computing center computing and resource collaborative scheduling method system for energy efficiency optimization. The embodiments of this invention do not specifically limit this.

[0056] In this embodiment of the invention, descriptions such as "when," "under the circumstances," "if," and "if" all refer to the device making corresponding processing under certain objective circumstances, and are not limited to a specific time. They do not require the device to make a judgment action during implementation, nor do they imply any other limitations.

[0057] See Figure 1 This invention provides a method for collaborative scheduling of computing and resources in intelligent computing centers for energy efficiency optimization, comprising:

[0058] Step 101: Establish a multi-level resource and energy efficiency perception model for the intelligent computing center and obtain real-time perception data;

[0059] In step 101, the multi-level resource and energy efficiency perception model of the intelligent computing center specifically includes:

[0060] The computing resource layer is used to collect real-time data on the load rate, hardware operating frequency, temperature, and real-time power consumption of computing nodes.

[0061] By deploying agent programs on each computing node, individual and cluster-level metrics of CPUs, GPUs, and dedicated AI accelerator cards are collected in real time, including but not limited to core utilization, operating frequency and voltage, cache and memory utilization at all levels, bandwidth utilization and latency of interconnect networks, chip junction temperature, and real-time power consumption obtained through smart cabinet PDUs or onboard sensors.

[0062] The task feature layer is used to extract the computing power requirements, memory and communication characteristics, priority and deadline constraints of the computing tasks to be scheduled.

[0063] At the time of task submission or through runtime analysis, extract its static and dynamic characteristics. Static characteristics include task type, declared resource requirements, software framework and dependencies; dynamic characteristics include computational feature profiles derived from historical or real-time analysis, such as computing power density, memory access patterns, communication topology and data volume, task phase division, and tolerance level for hardware errors; at the same time, the business priority and service level agreement constraints associated with the task, including the expected completion deadline.

[0064] The infrastructure resource layer is used to collect real-time data on the operating efficiency of the cooling and power distribution systems, ambient temperature and humidity, and external energy signals.

[0065] It integrates data streams from building automation systems, power distribution management systems, and environmental sensors to monitor the load rate, energy efficiency ratio, valve opening degree, and set point of the cooling subsystem in real time; monitors the load, conversion efficiency, and harmonics of the power supply and distribution electronic system; and collects temperature and humidity data at multiple points in the computer room, cabinet intake / return air temperature, and outdoor meteorological parameters; at the same time, it accesses external dynamic signals such as time-of-use electricity pricing, grid carbon intensity factor, and renewable energy output forecasts.

[0066] Step 102: Based on the real-time sensing data, construct a full-stack energy efficiency assessment and prediction model;

[0067] In step 102, based on the real-time sensing data, a full-stack energy efficiency assessment and prediction model is constructed; specifically including:

[0068] The joint optimization objective function aims to minimize the perceived effective energy consumption of the overall system ownership cost. The real-time perceived data is used as the input to the full-stack energy efficiency assessment and prediction model, and the output is a quantitative prediction of the overall energy efficiency performance in the future.

[0069] The joint optimization objective function incorporates a performance penalty term for computational tasks as a constraint or a weighted cost term.

[0070] The system dynamically learns from historical data using machine learning methods to quantitatively predict the overall energy efficiency of the system under different resource allocation schemes and infrastructure control strategies. It utilizes a combination of supervised learning and reinforcement learning. In the supervised learning part, the model is trained using historical operating data to predict the execution time and power consumption curves of specific tasks under specific hardware configurations, as well as the optimal operating point and energy consumption of the infrastructure subsystem under specific thermal loads and environmental conditions. In the reinforcement learning part, the long-term strategy value is explored in a simulation environment to optimize the dynamic weight coefficients.

[0071] The overall energy efficiency assessment targets of the system include computational energy consumption, cooling energy consumption, power supply and distribution losses, and the cost of service quality breach. The cost of service quality breach refers to the service quality breach cost incurred due to violations of the task SLA. This is modeled using dynamic weighting coefficients, where the breach cost can be monetized for business priorities, achieving an automatic trade-off between energy efficiency and performance. The integrated modeling using dynamically adjustable weighting coefficients incorporates economic costs, energy consumption, and service quality into a unified decision-making framework for comprehensive consideration.

[0072] This invention constructs and runs a unified full-stack energy efficiency assessment and prediction model. It employs a machine learning model trained on historical data to rapidly infer and predict performance and power consumption under specific task-hardware mapping relationships, as well as the optimal operating point of infrastructure under specific thermal loads and external environments. By defining a perceptual joint optimization objective function, the originally independent performance and energy consumption objectives are unified into a quantifiable mathematical model for trade-offs. This enables scheduling decisions to accurately balance task completion time and energy costs, avoiding energy waste caused by pursuing a single indicator, and achieving an optimal combination of economic benefits and technological efficiency.

[0073] A set of machine learning models trained on large-scale historical and real-time data is employed to establish accurate predictive inference capabilities. This set of models includes sub-models specifically designed to predict the execution time and dynamic power consumption curves of specific tasks under target hardware configurations, and sub-models capable of inferring the optimal operating setpoint and energy consumption of infrastructure based on predicted thermal load and environmental parameters. This enables rapid quantitative evaluation of the overall system energy efficiency under different scheduling strategies. Finally, the machine learning model set incorporates online learning and dynamic adaptation mechanisms. By continuously comparing the deviation between predicted values ​​and actual execution results, it automatically triggers incremental learning or parameter fine-tuning, ensuring that the model can adapt to hardware aging, load pattern changes, and external environmental changes over the long term, maintaining high predictive accuracy and thus supporting the continuous effectiveness of the closed-loop optimization system.

[0074] In step 102, constructing a full-stack energy efficiency assessment and prediction model further includes: constructing a virtual digital twin synchronized or quasi-synchronized with the physical computing center. This virtual digital twin is used to perform low-cost simulation, energy efficiency pre-assessment, and risk assessment of candidate scheduling strategies generated by the joint optimization engine before they are implemented in the physical system. This invention, by introducing digital twin technology to construct a virtual simulation environment, enables complex collaborative scheduling strategies to be fully validated before being implemented in the real system. This reduces the risk of system performance fluctuations or cooling failures due to inappropriate strategies, providing a safe testing ground for exploring more aggressive energy efficiency optimization strategies, thereby uncovering deeper energy-saving potential.

[0075] The virtual digital twin possesses thermodynamic simulation capabilities. Based on predicted computational load distribution and equipment power consumption, it generates a three-dimensional dynamic temperature field and airflow organization simulation diagram inside the computer room, and optimizes the setpoint adjustment scheme of the cooling system based on this simulation diagram. By integrating thermodynamic simulation capabilities into the digital twin, scheduling decisions not only consider power consumption but also enable accurate prediction and proactive management. By optimizing task placement to reduce local hotspots and formulating differentiated cooling strategies accordingly, the ineffective energy consumption of the cooling system can be reduced.

[0076] The virtual digital twin continuously receives and integrates real-time sensing data from step S1 through a data interface, thereby dynamically mapping the real state of the physical entity and forming a comprehensive simulation environment that includes a virtual image of computing resources, a task execution logic model, and a refined physical model of infrastructure.

[0077] The core function of a virtual digital twin is to perform a low-cost, high-speed simulation of the entire lifecycle of any candidate scheduling strategy generated by the collaborative optimization engine before it is sent to the physical system for execution. The simulation process accurately simulates the execution of computing tasks on virtual hardware, the dynamic power consumption and heat distribution generated, and the response and energy consumption of the infrastructure control system.

[0078] Based on simulation results, the system conducts multi-dimensional pre-evaluation and risk assessment: in terms of energy efficiency pre-evaluation, it accurately calculates the predicted total energy efficiency index under the strategy; in terms of risk assessment, it focuses on analyzing whether potential local hotspots exceed the safety threshold, whether the power supply link is overloaded, and whether critical tasks may violate the SLA due to resource competition. Only strategies that are verified as efficient and safe through twin verification will be finally approved for implementation, thereby achieving zero-risk energy efficiency optimization exploration for the physical system.

[0079] The virtual digital twin further integrates high-precision thermodynamics and computational fluid dynamics simulation capabilities. This simulation capability takes the predicted computing load distribution and fine-grained device power consumption data as input, and dynamically simulates the heat generation, transfer and diffusion processes in the data center from the chip level, server level to the rack level and even the entire data center level by solving the heat transfer and flow control equations in three-dimensional space.

[0080] The simulation engine can generate and visualize three-dimensional dynamic temperature field cloud maps and airflow organization vector maps inside the computer room in real time. The temperature field cloud map can not only identify potential overheated areas, but also show the spatiotemporal evolution of the temperature gradient; the airflow organization map clearly depicts the paths of hot and cold air, mixing conditions, and possible inefficiencies such as short-circuit circulation, insufficient airflow, or bypass.

[0081] Based on this high-fidelity thermal environment simulation, the system can implement data-driven, precise cooling optimization. Specifically, it can dynamically optimize the setpoint adjustment scheme of the cooling system according to the simulated temperature distribution and airflow efficiency. For example, it can selectively adjust the air supply temperature and air volume of different areas to achieve a shift from unified cooling for the entire computer room to differentiated cooling based on the needs of racks or clusters; or optimize the opening distribution of chilled water valves to precisely match cooling capacity with local heat loads. This essentially upgrades the cooling strategy from a rough response based on empirical rules to predictive, precise control based on physical simulation, thereby minimizing unnecessary energy consumption of the cooling system while ensuring the safe temperature of the equipment.

[0082] Step 103: Output predicted values ​​based on the full-stack energy efficiency assessment and prediction model; specifically including:

[0083] Based on the full-stack energy efficiency assessment and prediction model, the optimal collaborative scheduling strategy is solved through a joint optimization engine, and the predicted value is output according to the collaborative scheduling strategy.

[0084] In step 103, based on the full-stack energy efficiency assessment and prediction model, the optimal cooperative scheduling strategy is solved through a joint optimization engine, and the predicted value is output according to the cooperative scheduling strategy; specifically including:

[0085] In response to the arrival of new computing tasks or changes in system state, a collaborative optimization decision-making process is triggered. Based on a full-stack energy efficiency assessment and prediction model, the optimal collaborative scheduling strategy is solved through a joint optimization engine while satisfying task constraints. The collaborative scheduling strategy simultaneously outputs the placement scheme of computing tasks on heterogeneous hardware, the dynamic frequency adjustment command of computing devices, and the setpoint adjustment scheme of cooling and power supply systems.

[0086] The collaborative scheduling strategy process explicitly responds to external dynamic signals; when the real-time electricity price is detected to be entering a peak period, the joint optimization engine dynamically adjusts the weights of the joint optimization objective function to generate a strategy to reduce real-time electricity costs.

[0087] In this embodiment of the invention, strategies for reducing immediate electricity costs include: scheduling non-urgent tasks that can be delayed to be executed during off-peak electricity prices, migrating tasks to computing units with higher overall energy efficiency, or proactively reducing the operating frequency of some non-critical hardware to reduce power consumption.

[0088] Based on the goal of reconstruction, the engine prioritizes searching and generating scheduling strategies in the solution space that can effectively reduce immediate power expenditure. This is reflected in a multi-layered strategy package, which mainly includes:

[0089] Time-shifted scheduling: Deep analysis of the task queue identifies non-urgent tasks that are time-flexible and insensitive to latency. Combined with the task dependency graph, the start time of these tasks is scheduled to be during off-peak hours or normal periods when electricity prices are lower.

[0090] Space energy efficiency migration: While ensuring network and data locality, assess the migration of currently running or soon-to-be-started tasks from current high-power computing units to other computing pools within the data center with higher overall energy efficiency, or to edge sites or cloud availability zones located in low-carbon electricity price areas.

[0091] Dynamic adjustment of hardware performance and power consumption: The workload is broken down and analyzed into compute-intensive and memory-intensive tasks. For computing tasks on non-critical paths or hardware in a low-utilization state, the operating frequency and voltage are proactively and safely reduced, or some servers are put into a low-power sleep state while ensuring business continuity.

[0092] Multi-system collaborative execution: The generated strategy will simultaneously consider the linkage between computing and infrastructure.

[0093] This mechanism not only directly reduces operating electricity costs, but also transforms the intelligent computing center from a simple power load into a flexible resource with demand-side response capabilities. At the same time, it indirectly optimizes the carbon footprint by guiding the load to align with green electricity periods.

[0094] This invention enables intelligent computing centers to transform from passive power consumers to active energy consumers by explicitly responding to external signals such as real-time electricity prices. It can intelligently shift computing loads from peak electricity consumption periods to off-peak periods or migrate them to high-energy-efficiency infrastructure, directly reducing electricity costs and enhancing the ability of data centers to participate in grid demand-side response as flexible loads, thereby improving operational economics.

[0095] The joint optimization engine employs a hybrid optimization algorithm for solving the problem. This algorithm uses metaheuristics to perform global exploration in a large-scale solution space to generate a high-quality initial policy population, and combines it with reinforcement learning to fine-tune the policy locally and evaluate long-term benefits. This invention uses a hybrid optimization algorithm combining metaheuristics and reinforcement learning to effectively solve the problem of solving high-dimensional, nonlinear cooperative scheduling problems. The metaheuristic algorithm ensures the breadth of the search in the vast solution space, avoiding getting trapped in local optima; reinforcement learning, through the evaluation of long-term benefits, enables the policy to learn and evolve. The combination of the two ensures the timeliness, quality, and adaptability of the decision-making process.

[0096] In this embodiment of the invention, based on a full-stack energy efficiency assessment and prediction model, a collaborative optimization engine makes real-time decisions to generate a cross-domain collaborative scheduling strategy. Decisions are triggered by events such as new task submissions, the arrival of periodic scheduling windows, or significant changes in system state. The scheduling problem is constructed as a mixed-integer nonlinear programming problem that satisfies all task resource requirements and SLA constraints. The evaluation function is the objective, and a hybrid optimization algorithm is used for efficient solution. This hybrid optimization algorithm can combine metaheuristic algorithms for global solution search and embed gradient-based methods or reinforcement learning agents for local fine-tuning. The solution result is an executable collaborative scheduling strategy, which is a composite of three instructions: 1) Computational scheduling instructions: specifying the specific servers, accelerator cards, and container quotas for each task; 2) Hardware control instructions: setting specific operating frequency-voltage pairs for each computing unit to implement dynamic energy efficiency management; 3) Infrastructure control instructions: issuing setpoints such as chilled water temperature and fan speed to the cooling system and proposing load allocation suggestions to the power management system.

[0097] To address high-dimensional, nonlinear, and complexly constrained collaborative scheduling problems, the joint optimization engine employs a deeply integrated hybrid optimization algorithm architecture. Through a hierarchical and collaborative working mechanism, it organically integrates the global exploration advantages of metaheuristic algorithms with the sequential decision-making and long-term value assessment capabilities of reinforcement learning algorithms. The specific implementation is as follows:

[0098] First, the metaheuristic algorithm layer, acting as a global explorer, is responsible for efficient searching within a vast solution space comprised of variables such as task placement, hardware frequency tuning, and infrastructure settings. By maintaining a diverse policy population, the metaheuristic algorithm layer extensively explores potential optimal regions using operations such as crossover and mutation, generating a set of high-quality initial policy solutions that are balanced across multiple objectives, including energy efficiency, cost, and performance. This effectively prevents the solution process from prematurely falling into local optima, providing a high-quality starting point for subsequent fine-tuning.

[0099] Secondly, the reinforcement learning algorithm layer, acting as a local tuner and long-term evaluator, receives high-quality policies from the metaheuristic algorithm layer as initial policies or starting points for exploration. Building upon this, the reinforcement learning agent continuously interacts with the digital twin simulation environment, making small, sequential adjustments to the policy. By receiving multi-dimensional reward signals from the simulation environment, it not only learns how to fine-tune individual policies but, more importantly, assesses the long-term impact of current decisions on the future system state, thereby learning to formulate forward-looking scheduling policies.

[0100] Finally, the two algorithms are tightly coupled through a collaborative feedback loop. The metaheuristic algorithm periodically abstracts the high-value policy patterns it discovers into empirical knowledge and injects it into the state space or policy network of the reinforcement learning agent to guide its more efficient exploration. Simultaneously, new optimal regions or value assessment information discovered by reinforcement learning in local search can also be fed back to the metaheuristic algorithm to dynamically adjust its population evolution direction. This collaborative mechanism ensures that the joint optimization engine can both quickly locate promising regions in a large-scale solution space and perform deep, refined optimization of the policy while considering long-term gains, thereby stably and efficiently solving for high-quality collaborative scheduling policies.

[0101] See Figure 2 In this embodiment of the invention, before solving for the optimal cooperative scheduling strategy using a joint optimization engine based on a full-stack energy efficiency assessment and prediction model, and before outputting the predicted value according to the cooperative scheduling strategy, energy efficiency-aware task preprocessing is also included; the energy efficiency-aware task preprocessing includes:

[0102] The queue of computing tasks waiting to be scheduled is subjected to feature analysis to identify multiple independent tasks with complementary computing features. These independent tasks are then intelligently combined into a composite task package, which is submitted as a whole to the collaborative optimization engine for unified resource allocation and scheduling, thereby improving the overall resource utilization and energy efficiency of the cluster.

[0103] This invention adds an energy efficiency-aware task preprocessing step, intelligently packaging tasks before scheduling and performing pre-emptive resource integration and optimization. This improves resource matching efficiency from a macro-level perspective of the task queue, thereby increasing the average utilization of computing resources such as servers. The specific process is as follows:

[0104] The system performs fine-grained feature analysis and profile construction on the queue of computing tasks waiting to be scheduled. Based on the task features extracted in step 101, the system constructs a multi-dimensional computing resource requirement profile for each task. This profile not only includes the demand for explicit resources such as CPU, GPU, and memory, but also characterizes the strength of its dependence on specific hardware resources, its usage pattern, and its timing characteristics during runtime.

[0105] The intelligent task combination analysis engine is activated. Based on a predefined resource complementarity model and cooperative scheduling affinity rules, the engine performs pairwise or multiple matching analyses on tasks in the queue. Its goal is to identify independent tasks that exhibit spatiotemporal complementarity in resource requirements. Typical complementarity patterns include combining a memory-bandwidth-intensive task with a computationally-intensive task; or combining a collective task requiring frequent communication with multiple tasks with extremely low communication requirements.

[0106] For the identified highly complementary task groups, the system intelligently bundles and encapsulates them into a unified, atomic composite task package. This composite task package will be given a unified resource request specification and scheduling priority, and submitted as a single scheduling unit to the joint optimization engine in step 103.

[0107] This invention effectively simulates resource demand mapping, enabling composite task packages to be scheduled as a whole to a single computing node or a tightly coupled group of nodes for execution in subsequent micro-resource allocation. It maximizes the concurrent utilization of heterogeneous computing resources within a single server, reduces core idleness, memory idleness, or high-speed interconnect link idleness caused by resource type mismatch or fragmentation, fundamentally reducing the cluster's base power consumption during resource idleness and significantly improving the effective computing output per unit of energy consumption. It is particularly suitable for scenarios handling massive numbers of small or heterogeneous tasks.

[0108] Specifically, the complementarity of computational features refers to identifying and combining a first task with high memory bandwidth requirements with a second task that maximizes the utilization of computing cores. When the first and second tasks share the same computing node, the heterogeneous hardware resources of that node are utilized to reduce resource idle time. This invention combines and packages memory bandwidth-intensive and computationally intensive tasks, achieving efficient utilization of the internal resources of modern heterogeneous computing nodes. This refined resource allocation results in a smoother utilization curve for individual hardware components, maximizing hardware energy efficiency and avoiding the waste of other resources being idle due to bottlenecks in certain types of resources.

[0109] Computational feature complementarity is based on a multi-dimensional and dynamic matching principle. Through intelligent task combination, it maximizes the concurrent utilization of various heterogeneous resources within a single physical computing node, thereby reducing fixed power consumption and improving overall energy efficiency.

[0110] In this embodiment of the invention, the external energy signals acquired in real-time include at least time-of-use electricity price signals, regional power grid carbon emission intensity signals, and real-time renewable energy output forecast data. When the joint optimization engine makes decisions, it quantifies carbon emission intensity into carbon costs and transforms renewable energy forecasts into green incentive factors, thereby setting the minimization of carbon footprint or the maximization of green electricity consumption as a core optimization objective alongside economic benefits, or incorporating it as a dynamic constraint into the decision-making model. This mechanism enables the engine to proactively generate and execute green scheduling strategies, such as initiating batch computing during peak green electricity output periods and reducing non-emergency loads during periods of high carbon and electricity prices. This drives intelligent computing centers to shift from passive energy consumption to proactive participation in energy system balance, becoming flexible computing assets that support dual-carbon goals. This invention incorporates carbon emission intensity and renewable energy forecasts into the optimization objectives, elevating this method beyond simple energy saving to green scheduling. It guides computing loads to tilt towards low-carbon electricity periods or regions in time or space, proactively promoting the consumption of intermittent green electricity such as wind and solar power, and directly reducing the carbon footprint of the computing industry.

[0111] Step 104, the implementation of the collaborative scheduling strategy to realize the online adaptation and closed-loop optimization of the model, specifically includes: implementing the collaborative scheduling strategy, monitoring the actual energy consumption and task performance data after the implementation of the collaborative scheduling strategy, using the difference between the monitoring results and the predicted values ​​as a feedback signal, and inputting it into the full-stack energy efficiency assessment and prediction model to realize the online adaptation and closed-loop optimization of the model.

[0112] Specifically, the closed-loop optimization includes: establishing a model performance monitor; when the average deviation between the actual energy consumption data and the predicted value continues to exceed a preset threshold, automatically triggering the incremental learning or retraining process of the energy efficiency assessment and prediction model, and updating the model parameters using the latest collected data to maintain prediction accuracy.

[0113] Through standardized interfaces or adapters, the composite policy instruction package generated in step 103 is synchronously distributed to the computing cluster scheduler, hardware management platform, and infrastructure management system to ensure that each system performs corresponding operations within the coordinated time window, achieving cross-domain action synchronization. During the policy execution cycle, the actual computing task progress, power consumption of each subsystem, and key environmental indicators are collected at high frequency. The actual measured total energy efficiency index is compared and analyzed with the predicted value of the full-stack energy efficiency assessment and prediction model at the time of decision-making. The prediction deviation is calculated, and this deviation data, together with the corresponding system status and decision scheme, forms a new training sample. This sample is incrementally fed back to the machine learning prediction engine in step 102 to drive the online fine-tuning of model parameters or trigger periodic retraining, forming a continuously self-improving closed-loop optimization system.

[0114] This invention establishes a model self-updating mechanism based on deviation monitoring, which ensures the long-term effectiveness of the method. As hardware ages, task types change, or infrastructure is upgraded, the system can automatically detect the degradation of the model's predictive ability and trigger learning, so that the energy efficiency optimization strategy always remains consistent with the real state of the physical system. This achieves adaptive optimization throughout the system's entire life cycle and ensures the sustainability of the technical solution.

[0115] This invention achieves system-level optimization of energy efficiency management in intelligent computing centers by establishing a closed-loop system of full-stack perception and collaborative scheduling from computing tasks and hardware resources to infrastructure. It breaks down the barriers of independent operation of traditional subsystems and can dynamically allocate resources from a global perspective. This significantly reduces the overall energy consumption of the system while ensuring the performance of computing tasks, thus achieving a dual reduction in operating costs and carbon emissions.

[0116] This invention achieves a breakthrough from local static optimization to global dynamic collaboration. By incorporating computing task scheduling, hardware resource regulation, and infrastructure management into a unified joint optimization framework and establishing a full-link digital twin model, it can accurately quantify the chain-like energy efficiency impact caused by any scheduling decision, thereby finding the globally optimal or near-optimal solution at the system level. This collaborative scheduling mechanism can effectively break down resource silos, solve energy efficiency losses caused by inter-system target conflicts or response delays, and significantly approach the theoretical limit of PUE.

[0117] This invention endows intelligent computing centers with high adaptability and multi-objective balancing capabilities in complex and dynamic environments. By utilizing machine learning models for real-time learning and prediction, the system can proactively adapt to drastic fluctuations in task load, changes in the external environment, and the heterogeneity of hardware devices. It can not only achieve dynamic resource adjustments at the second to minute level, but also intelligently weigh and compromise among multiple sometimes conflicting objectives, ensuring the deadlines of core tasks while reducing operating costs, thus achieving a balance between economy and reliability.

[0118] This invention uses digital twin technology to simulate, verify, and rehearse scheduling strategies in virtual space, reducing the risks and costs of energy efficiency optimization trial and error in production systems and ensuring the stability and security of data center operations. At the same time, the closed-loop feedback learning mechanism enables the system's energy efficiency model to continuously evolve and adapt to new hardware architectures, task types, and operational strategies, ensuring the long-term viability of the technology.

[0119] This invention achieves real-time global insight into system status by constructing a multi-layered perception model covering computing tasks, hardware resources, and infrastructure. Based on the perceived data, it utilizes machine learning and digital twin technologies to establish a unified full-stack energy efficiency assessment and prediction model, which can accurately quantify the energy efficiency impact of different scheduling strategies. Through a joint decision engine integrating hybrid optimization algorithms, it dynamically responds to task demands, real-time electricity prices, and carbon emission signals, simultaneously generating coordinated scheduling instructions for computing task placement, hardware frequency adjustment, and infrastructure control. This enables a multi-objective intelligent trade-off between energy consumption, cost, performance, and carbon emissions from a global perspective. This invention also introduces energy efficiency-aware task preprocessing and intelligent packaging steps, and continuously optimizes the model through a closed-loop learning mechanism after execution, thus forming a self-improving systematic solution. Ultimately, this significantly improves the overall energy efficiency and sustainable operation capabilities of intelligent computing centers.

[0120] Example 2:

[0121] See Figure 3 This invention provides a smart computing center computing and resource collaborative scheduling system for energy efficiency optimization, comprising:

[0122] The perception model construction unit 201 is used to establish a multi-level resource and energy efficiency perception model for the intelligent computing center and to acquire real-time perception data.

[0123] The assessment and prediction model building unit 202 constructs a full-stack energy efficiency assessment and prediction model based on the real-time sensing data.

[0124] The collaborative scheduling strategy unit 203 outputs predicted values ​​based on the full-stack energy efficiency assessment and prediction model.

[0125] The monitoring feedback optimization unit 204 is used to execute the collaborative scheduling strategy to achieve online adaptive and closed-loop optimization of the model.

[0126] The various variations and specific examples of the energy efficiency optimization-oriented intelligent computing center computing and resource collaborative scheduling method in the foregoing embodiments are also applicable to the power system computing power and power collaborative scheduling system of the regional intelligent computing center in this embodiment. Through the foregoing detailed description of the energy efficiency optimization-oriented intelligent computing center computing and resource collaborative scheduling method, those skilled in the art can clearly understand the power system computing power and power collaborative scheduling system of the regional intelligent computing center in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here.

[0127] Example 3:

[0128] This invention proposes an electronic device, which can be a network device, or a chip (system) or other component or assembly that can be disposed in a network device. The electronic device may include a processor. Optionally, the electronic device may also include a memory and / or a transceiver. The processor is coupled to the memory and transceiver, for example, by means of a communication bus connection.

[0129] The following is a detailed introduction to the various components of the electronic device:

[0130] In this context, the processor is the control center of the electronic device. It can be a single processor or a collective term for multiple processing elements. For example, a processor can be one or more central processing units (CPUs), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).

[0131] Alternatively, the processor can perform various functions of the electronic device by running or executing software programs stored in memory and calling data stored in memory, such as executing the above-mentioned intelligent computing center computing and resource collaborative scheduling method for energy efficiency optimization.

[0132] In a specific implementation, as one example, the processor may include one or more CPUs, such as CPU0 and CPU1.

[0133] In a specific implementation, as one example, the electronic device may also include multiple processors. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, a processor may refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).

[0134] The memory is used to store the software program that executes the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.

[0135] Optionally, the memory can be read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory can be integrated with the processor or exist independently and coupled to the processor through an interface circuit of an electronic device; the embodiments of the present invention do not specifically limit this.

[0136] A transceiver is used for communication with other electronic devices. For example, if the electronic device is a terminal, the transceiver can be used to communicate with a network device or with another terminal device. Similarly, if the electronic device is a network device, the transceiver can be used to communicate with a terminal or with another network device.

[0137] Optionally, the transceiver may include a receiver and a transmitter. The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.

[0138] Optionally, the transceiver can be integrated with the processor or exist independently and coupled to the processor through the interface circuit of the electronic device. This embodiment of the invention does not specifically limit this.

[0139] It is understood that the structure of an electronic device does not constitute a limitation on the electronic device. An actual electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0140] Furthermore, the technical effects of the electronic devices can be referenced from the technical effects of the energy-efficiency-optimized intelligent computing center computing and resource collaborative scheduling method described in the above method embodiments, and will not be repeated here.

[0141] It should be understood that the processor in the embodiments of the present invention can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0142] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DRRAM).

[0143] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0144] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0145] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0146] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0147] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0148] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0149] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0150] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0151] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the invention to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A method for collaborative scheduling of computing and resources in intelligent computing centers for energy efficiency optimization, characterized in that, include: Establish a multi-level resource and energy efficiency perception model for the intelligent computing center to obtain real-time perception data; Based on the real-time sensing data, a full-stack energy efficiency assessment and prediction model is constructed. The predicted value is output based on the full-stack energy efficiency assessment and prediction model; The collaborative scheduling strategy enables the model to achieve online adaptation and closed-loop optimization.

2. The intelligent computing center computing and resource collaborative scheduling method for energy efficiency optimization according to claim 1, characterized in that, The multi-layered resource and energy efficiency perception model of the intelligent computing center specifically includes: The computing resource layer is used to collect real-time data on the load rate, hardware operating frequency, temperature, and real-time power consumption of computing nodes. The task feature layer is used to extract the computing power requirements, memory and communication characteristics, priority and deadline constraints of the computing tasks to be scheduled. The infrastructure resource layer is used to collect real-time data on the operating efficiency of the cooling and power distribution systems, ambient temperature and humidity, and external energy signals.

3. The intelligent computing center computing and resource collaborative scheduling method for energy efficiency optimization according to claim 1, characterized in that, Based on the real-time sensing data, a full-stack energy efficiency assessment and prediction model is constructed; specifically including: The joint optimization objective function aims to minimize the perceived effective energy consumption of the overall system ownership cost. The real-time perceived data is used as the input to the full-stack energy efficiency assessment and prediction model, and the output is a quantitative prediction of the overall energy efficiency performance in the future. The joint optimization objective function incorporates a performance penalty term for computational tasks as a constraint or a weighted cost term. By dynamically learning historical data through machine learning methods, we can quantitatively predict the total energy efficiency of the system under different combinations of resource allocation schemes and infrastructure control strategies. The evaluation targets for the overall energy efficiency of the system include computational energy consumption, cooling energy consumption, power supply and distribution losses, and the cost of breach of service quality.

4. The intelligent computing center computing and resource collaborative scheduling method for energy efficiency optimization according to claim 3, characterized in that, The construction of the full-stack energy efficiency assessment and prediction model also includes: constructing a virtual digital twin that is synchronized or quasi-synchronized with the physical intelligent computing center. The virtual digital twin is used to perform low-cost simulation, energy efficiency pre-assessment, and risk assessment of candidate scheduling strategies generated by the joint optimization engine before they are put into execution in the physical system. Among them, the virtual digital twin has thermodynamic simulation capabilities. Based on the predicted computing load distribution and equipment power consumption, it generates a three-dimensional dynamic temperature field and airflow organization simulation diagram inside the computer room, and optimizes the setpoint adjustment scheme of the cooling system based on this simulation diagram.

5. The intelligent computing center computing and resource collaborative scheduling method for energy efficiency optimization according to claim 1, characterized in that, The predicted values ​​output by the full-stack energy efficiency assessment and prediction model specifically include: Based on the full-stack energy efficiency assessment and prediction model, the optimal collaborative scheduling strategy is solved through a joint optimization engine, and the predicted value is output according to the collaborative scheduling strategy.

6. The intelligent computing center computing and resource collaborative scheduling method for energy efficiency optimization according to claim 5, characterized in that, The process of solving for the optimal cooperative scheduling strategy using a joint optimization engine and outputting predicted values ​​based on the cooperative scheduling strategy specifically includes: In response to the arrival of new computing tasks or changes in system state, a collaborative optimization decision-making process is triggered. Based on a full-stack energy efficiency assessment and prediction model, the optimal collaborative scheduling strategy is solved through a joint optimization engine while satisfying task constraints. The collaborative scheduling strategy simultaneously outputs the placement scheme of computing tasks on heterogeneous hardware, the dynamic frequency adjustment command of computing devices, and the setpoint adjustment scheme of cooling and power supply systems. The collaborative scheduling strategy explicitly responds to external dynamic signals; when the real-time electricity price is detected to be entering a peak period, the joint optimization engine dynamically adjusts the weights of the joint optimization objective function to generate a strategy to reduce real-time electricity costs.

7. The intelligent computing center computing and resource collaborative scheduling method for energy efficiency optimization according to claim 6, characterized in that, Before solving for the optimal cooperative scheduling strategy using a joint optimization engine and outputting predicted values ​​based on the cooperative scheduling strategy, the process also includes energy efficiency-aware task preprocessing; the energy efficiency-aware task preprocessing includes: The queue of computing tasks waiting to be scheduled is subjected to feature analysis to identify multiple independent tasks with complementary computing features. These independent tasks are then intelligently combined into a composite task package, which is submitted as a whole to the collaborative optimization engine for unified resource allocation and scheduling to improve the overall resource utilization and energy efficiency of the cluster. Specifically, the complementarity of the computing features refers to: identifying and combining a first task with high memory bandwidth requirements with a second task with high utilization of computing cores; when the first and second tasks share the same computing node, utilizing the heterogeneous hardware resources of that node to reduce resource idle time.

8. The intelligent computing center computing and resource collaborative scheduling method for energy efficiency optimization according to claim 1, characterized in that, The aforementioned collaborative scheduling strategy enables the model to achieve online adaptation and closed-loop optimization, specifically including: The system executes a collaborative scheduling strategy, monitors the actual energy consumption and task performance data after the strategy is executed, and uses the difference between the monitoring results and the predicted values ​​as a feedback signal to input into the full-stack energy efficiency assessment and prediction model to achieve online adaptation and closed-loop optimization of the model.

9. The intelligent computing center computing and resource collaborative scheduling method for energy efficiency optimization according to claim 8, characterized in that, The closed-loop optimization specifically includes: establishing a model performance monitor, which automatically triggers the incremental learning or retraining process of the energy efficiency assessment and prediction model when the average deviation between the actual energy consumption data and the predicted value continues to exceed a preset threshold, and updates the model parameters using the latest collected data.

10. A smart computing center computing and resource collaborative scheduling system for energy efficiency optimization, characterized in that, The method for collaborative scheduling of computing and resources in intelligent computing centers, oriented towards energy efficiency optimization, as described in any one of claims 1-9, includes: The perception model building unit is used to establish a multi-level resource and energy efficiency perception model for the intelligent computing center and to acquire real-time perception data. The assessment and prediction model building unit constructs a full-stack energy efficiency assessment and prediction model based on the real-time sensing data. The collaborative scheduling strategy unit outputs predicted values ​​based on the full-stack energy efficiency assessment and prediction model. The monitoring and feedback optimization unit is used to execute the collaborative scheduling strategy to achieve online adaptation and closed-loop optimization of the model.

11. An electronic device, characterized in that, include: Processor and memory; The memory is used to store a computer program, which, when executed by the processor, causes the electronic device to perform the energy-efficient intelligent computing center computing and resource collaborative scheduling method as described in any one of claims 1-9.

12. A computer-readable storage medium, comprising: Computer programs or instructions; When the computer program or instructions are run on the computer, the computer performs the intelligent computing center computing and resource collaborative scheduling method for energy efficiency optimization as described in any one of claims 1-9.