Data center multi-objective scheduling device and method fusing digital twin and meta-learning

CN122763597APending Publication Date: 2026-09-15CHINA MOBILE GRP GANSU CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610100493.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-26
Publication Date
2026-09-15

Smart Images

  • Figure CN122763597A_ABST
    Figure CN122763597A_ABST
Patent Text Reader

Abstract

The application provides a data center multi-objective scheduling device and method fusing digital twinning and meta-learning, the device comprising: a multi-modal data acquisition module for acquiring multi-source heterogeneous data; a data preprocessing and feature engineering module for processing data and calculating health scores of power equipment; a prediction module for predicting power generation power sequences and predicting load power sequences; a digital twinning module for constructing a high-fidelity virtual model and providing a simulation environment for online optimization of a scheduling strategy; and a multi-objective optimization module adopting a meta-reinforcement learning intelligent agent for online fine-tuning and simulation verification in the digital twinning module based on real-time measurement values, predicted data and health scores of power equipment, determining an optimal power scheduling plan by solving a problem of a joint optimization target. The application solves the technical problems of the prior art, such as insufficient adaptability in a dynamic uncertain environment, single optimization target and failure to consider long-term health states of power equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of data center optimization scheduling, and in particular to a data center multi-objective scheduling device and method that integrates digital twins and meta-learning. Background Technology

[0002] As the cornerstone of the digital economy, data centers are experiencing ever-increasing computing power, but their energy consumption is also rising sharply. In response to low-carbon development goals, data centers are gradually introducing a high proportion of renewable energy. However, the inherent intermittency and volatility of renewable energy have posed a serious challenge to the stable operation and economic dispatch of the power system. At the same time, the deepening of market mechanisms such as time-of-use pricing and carbon emission trading are placing higher demands on the economic efficiency and market adaptability of data center energy management strategies.

[0003] Faced with the above challenges, existing scheduling technologies mainly fall into two categories, but both have significant limitations. One category is traditional optimization methods based on precise mathematical models, such as mixed-integer linear programming. While effective when operating conditions are fixed, the static modeling paradigm upon which these methods rely struggles to adapt to the dynamic uncertainties of market prices, load fluctuations, and renewable energy output. Furthermore, they typically only optimize short-term operating costs, neglecting the physical losses of critical equipment like energy storage during cyclical use, which is detrimental to long-term economic viability. The other category is data-driven reinforcement learning methods. Although capable of handling complex dynamic decisions, standard algorithms usually require massive amounts of trial and error to converge, resulting in low training efficiency and limited generalization ability to unseen operating scenarios. This fails to meet the basic requirements of data centers for high security and high timeliness in scheduling commands.

[0004] Therefore, existing technologies lack adaptability to dynamic and uncertain environments and have a singular optimization objective, making it difficult to ensure both short-term economic efficiency and long-term equipment health. This results in significant bottlenecks in the actual effectiveness of scheduling strategies and the long-term security of the system. Therefore, there is an urgent need in this field for a new scheduling paradigm that can quickly adapt to changing conditions and coordinate multi-dimensional objectives to support the safe, economical, low-carbon, and sustainable operation of data centers. Summary of the Invention

[0005] The purpose of this invention is to provide a data center multi-objective scheduling device and method that integrates digital twins and meta-learning, which solves the technical problems of insufficient adaptability of existing data center scheduling technology in dynamic and uncertain environments, single optimization objective, and failure to comprehensively consider the long-term health status of key power equipment.

[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: In a first aspect, the present invention provides a data center multi-objective scheduling device integrating digital twins and meta-learning. This device is applied to the integrated energy system of the data center, comprising a local energy system connected to the public power grid, and the local energy system including a renewable energy system and an energy storage system. The device includes: The multimodal data acquisition module is configured to monitor and acquire multi-source heterogeneous data from the data center, the energy system, the public power grid, and power equipment in real time. The data preprocessing and feature engineering module, connected to the multimodal data acquisition module, is configured to process the acquired multi-source heterogeneous data and dynamically calculate the health score of the power equipment. The prediction module, which communicates with the data preprocessing and feature engineering module, is configured to generate a predicted power generation sequence of the renewable energy system and a predicted load power sequence for future scheduling cycles based on historical data and external meteorological information. The digital twin module is configured to construct a high-fidelity virtual model of the data center's source-grid-load-storage energy system based on the physical parameters of power equipment and the grid topology, and to provide a secure and isolated simulation environment for online optimization of scheduling strategies; The multi-objective optimization module, as the decision-making core of the device, adopts a meta-reinforcement learning agent at its core. It is configured to perform online fine-tuning and simulation verification in the digital twin module based on real-time measurement values, future prediction data, and health scores of dynamic power equipment. By solving the problem of joint optimization objectives, the optimal power scheduling plan is determined. The control and scheduling module communicates with the multi-objective optimization module and is configured to generate control commands based on the optimal power scheduling plan and issue them to the execution units in the energy system.

[0007] Furthermore, the meta-reinforcement learning agent in the multi-objective optimization module is configured to execute the following mechanism: The offline meta-training loop mechanism is configured to train on a series of tasks with different operating scenarios sampled from historical data, in order to learn and solve for a set of initial parameters of a meta-learning model that can quickly adapt to new tasks. The online meta-adaptive loop mechanism is configured to, when faced with a new scheduling task defined by the latest prediction data, rapidly and finely adjust the initial parameters of the meta-learning model in the simulation environment provided by the digital twin module using interactive data related to the current new task, thereby generating a specialized scheduling strategy for the current scheduling task.

[0008] Furthermore, the multi-objective optimization module is configured to solve the problem of joint optimization objective by minimizing the joint cost function. The problem is that the joint cost function is defined as follows: ; in, For scheduling strategy, To schedule discrete time steps in the time domain, express The economic cost of time The dynamic weighting coefficients are, in order, economic costs, environmental costs, and power equipment degradation costs. express The environmental cost at any given moment, and express The cost of electrical equipment degradation at any given time.

[0009] Furthermore, the process of minimizing the joint cost function is subject to one or more constraints, including power balance constraints, energy storage system operation constraints, and grid interaction power constraints.

[0010] Furthermore, the health score of the power equipment is dynamically calculated based on multimodal physical sensor data related to the operational health of the power equipment and the real-time load change rate of the data center, specifically including: A basic health assessment model is trained based on historical data from multimodal physical sensors to obtain a raw health score reflecting the internal state of power equipment. Calculate an external operating stress factor based on the real-time load change rate of the data center; The original health score is dynamically combined with the external operational stress factors to obtain the final dynamic health score.

[0011] Furthermore, the digital twin module is configured to provide a physical, high-fidelity simulation environment that can accurately model the electrical topology of the data center and the dynamic response characteristics of its energy assets, thereby providing accurate simulation data for the online meta-adaptive cycle.

[0012] Secondly, the present invention also provides a multi-objective scheduling method for data centers that integrates digital twins and meta-learning, the method comprising the following steps: Step 301: Monitor and collect multi-source heterogeneous data from the data center, the energy system, the public power grid, and power equipment in real time; Step 302: Process the collected multi-source heterogeneous data and dynamically calculate the health score reflecting the current pressure status of the power equipment; Step 303: Based on historical data and external meteorological information, generate a load power prediction sequence and a renewable energy power generation prediction sequence for the future scheduling time domain; Step 304: Combine the real-time running data obtained in step 301, the health score dynamically calculated in step 302, and the prediction sequence generated in step 303 to construct a high-dimensional system state vector that can comprehensively characterize the current and expected future state of the system. Step 305: In a digital twin environment that is precisely mapped to the physical system, launch a reinforcement learning agent that has been pre-trained offline. Step 306: The agent takes the system state vector constructed in step 304 as input and performs an online meta-adaptation process. By simulating interaction in the digital twin environment, it quickly fine-tunes its policy network to adapt to the current scheduling task. Step 307: Solve the Markov decision process of the joint optimization objective through the fine-tuned agent, and output the optimal scheduling action for the current decision cycle. Step 308: Through the control and scheduling module, the optimal scheduling action output in step 307 is converted into control commands for specific power equipment and sent to the field execution unit. Step 309: At the start of the next decision cycle, return to step 301 to form a closed-loop scheduling mechanism that continuously senses, predicts, optimizes, and executes safely in the rolling time domain.

[0013] Furthermore, in terms of algorithm implementation, the agent employs the Model Independent Meta-Learning (MAML) algorithm to construct a two-layer learning framework of offline meta-training and online meta-adaptation; and in the policy update of the framework, the Proximal Policy Optimization (PPO) algorithm is used to perform gradient updates to improve training stability and convergence efficiency.

[0014] Furthermore, the online decision optimization steps in steps 305 to 307 involve modeling the scheduling problem as a Markov Decision Process (MDP), which includes: The state space, defined by the high-dimensional system state vector constructed in step 304, includes real-time measurements, future prediction sequences, and health scores of dynamic power equipment. Action space is the set of scheduling decisions for controllable energy assets within a data center, including the charging and discharging power setpoints of energy storage systems and the power interaction setpoints with the power grid. The reward function is used to evaluate the immediate benefits of performing a specific action in a specific state and is designed to guide the agent to learn to achieve the joint optimization objective.

[0015] Furthermore, the reward function in the Markov decision process At each time step It is precisely defined as: , The first term is the negative of the joint cost function, designed to incentivize the agent to minimize the cost; the second term... As a penalty, it is used to punish behaviors that violate system operating constraints, so as to ensure that the policies learned by the agent strictly adhere to physical and operational constraints.

[0016] Compared with the prior art, the present invention has at least the following beneficial effects: (1) This invention empowers the scheduling system with the ability to “learn” by deploying an intelligent agent based on meta-reinforcement learning. In the offline stage, the intelligent agent uses massive historical data to master meta-knowledge across scenarios. In the online stage, it uses a digital twin environment and only needs a small amount of current data to quickly fine-tune and generate a scheduling strategy that is highly adaptable to the current working conditions. This significantly improves the response speed and decision robustness under uncertain conditions such as renewable energy fluctuations, load changes and electricity price fluctuations, and solves the problem of poor adaptability of traditional methods.

[0017] (2) This invention innovatively constructs a multi-objective optimization framework that integrates economic costs, carbon emission costs and power equipment degradation costs, and achieves intelligent trade-offs through a dynamic weighting mechanism. In particular, it introduces a power equipment health score based on multimodal sensor data and load change rate dynamic calculation, quantifies the long-term losses of power equipment and incorporates them into real-time optimization, and realizes the extension of the goal from short-term operating economy to full life cycle asset health management, ensuring the stability and economy of the system's long-term operation.

[0018] (3) This invention, by combining high-fidelity digital twin technology, provides a safe and isolated simulation environment for the online adaptation and policy verification of meta-reinforcement learning agents. This not only avoids the risks that online learning may bring to the physical system, but also allows for full verification of the optimization strategy before the instruction is issued, thereby enhancing the reliability and security of the entire scheduling system.

[0019] (4) In terms of algorithm implementation, this invention adopts model-independent meta-learning to construct a two-layer learning framework and uses the proximal policy optimization algorithm for policy update. The PPO algorithm limits the update step size by truncating the surrogate objective function, which effectively ensures the stability and convergence efficiency of the training process in complex multi-objective scheduling problems, overcomes the defects of unstable training and slow convergence of traditional reinforcement learning, and improves the practicality and feasibility of the method. Attached Figure Description

[0020] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the overall structure of the multi-target dynamic scheduling device provided in this embodiment; Figure 2 This is an overall flowchart of the multi-objective dynamic scheduling method provided in this embodiment; Figure 3 This is a schematic diagram of the meta-reinforcement learning dual-loop optimization framework provided in this embodiment. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0024] This embodiment provides a data center multi-objective scheduling device that integrates digital twins and meta-learning. Please refer to... Figure 1 As shown, it is applied to the integrated energy system of source-grid-load-storage within the data center. The energy system includes a local energy system connected to the public power grid, wherein the local energy system further comprises a renewable energy system and an energy storage system. The device includes: The multimodal data acquisition module is configured to collect load data of the data center, power generation data of the renewable energy system, operating status data of the energy storage system, time-of-use electricity price and carbon emission factor information of the public power grid, and multimodal physical sensing data related to the operational health of key power equipment in real time and at high frequency. The data preprocessing and feature engineering module, connected to the multimodal data acquisition module, is configured to process the acquired multi-source heterogeneous data, including cleaning, normalization and time-series alignment, and extract key state features for decision-making. The state features include: dynamically calculating a health score reflecting the current pressure state of the power equipment based on the multimodal physical sensing data and load change rate. The prediction module, which communicates with the data preprocessing and feature engineering module, is configured to generate a predicted power generation sequence of the renewable energy system and a predicted load power sequence for future scheduling cycles based on historical data and external meteorological information. The digital twin module is configured to construct a high-fidelity virtual model of the data center's source-grid-load-storage energy system / physical system based on the physical parameters of power equipment and the grid topology, and to provide a secure and isolated simulation environment for online optimization of scheduling strategies; The multi-objective optimization module, as the decision-making core of the device, adopts a meta-reinforcement learning agent, which is configured to perform online fine-tuning and simulation verification based on real-time measurement values, future prediction data, and dynamic power equipment health scores in the digital twin module. By solving a problem with total electricity cost, total carbon emissions, and the impact of power equipment health as joint optimization objectives, an optimal power scheduling plan is determined. The (safety) control and scheduling module communicates with the multi-objective optimization module and is configured to generate control commands conforming to the industrial communication protocol according to the optimal power scheduling plan, and issue them to the execution unit in the energy system through an encrypted channel.

[0025] In this embodiment, the meta-reinforcement learning agent in the multi-objective optimization module is designed as a multi-dimensional learning structure, including: An offline meta-training loop mechanism is configured to train on a series of tasks with different operating scenarios sampled from historical data. Its goal is to learn and solve for a set of initial parameters of a meta-learning model that can quickly adapt to new tasks. An online meta-adaptive loop mechanism is configured to, when faced with a new scheduling task defined by the latest prediction data, rapidly and finely adjust the initial parameters of the meta-learning model in the simulation environment provided by the digital twin module using only a small amount of interaction data related to the current new task, thereby tailoring a highly targeted specialized scheduling strategy for the current scheduling task.

[0026] In this embodiment, the multi-objective optimization module is configured to solve for a function that minimizes a joint cost function. For the objective problem, the joint cost function is defined as follows: ,in: To schedule discrete time steps in the time domain, For scheduling strategy, express The economic cost at any given time mainly accounts for the system's expenditure on purchasing electricity from the grid. express The environmental cost at any given moment is quantified based on the purchased electricity volume and the grid's real-time carbon emission intensity (or factor). express The critical power equipment degradation cost at any given moment is quantified based on the changes in the health status of the energy storage system due to its charging and discharging operations, and is used to assess the operational losses of core components. These are dynamic weighting coefficients that correspond to economic costs, environmental costs, and power equipment degradation costs, respectively, and their values ​​are adjusted based on the real-time assessment of the system's operating status.

[0027] In this embodiment, the process of minimizing the joint cost function is subject to one or more of the following constraints: Power balance constraints: ; Energy storage system operating constraints: , , ; Power grid interaction constraints: ,in: express The interaction power between the time-of-use system and the power grid, where the power purchased is positive. express The power generation capacity of renewable energy at all times, express Total load power of the data center at any given time These represent the charging and discharging power of the energy storage system at time t, respectively. These represent the upper and lower limits of the state of charge of the energy storage system, i.e., the maximum and minimum allowable values. These are the maximum charging and discharging power of the energy storage system, respectively. These are the minimum and maximum limits for the power exchanged with the power grid, respectively.

[0028] In this embodiment, the data processing and feature engineering module calculates the health score of the power equipment according to the following steps: First, a basic health assessment model is trained based on historical data to obtain an original health score reflecting the internal state of the power equipment; second, based on the real-time load change rate... Calculate an external operating stress factor; finally, dynamically combine the original health score with the external operating stress factor to obtain the final dynamic health score.

[0029] In this embodiment, the (security) control scheduling module is configured to use a virtual private network (VPN) tunnel based on the IPSec protocol to achieve secure encrypted communication between the device and external networks, and to use the AES-256 encryption algorithm to encrypt the transmitted control command data packets to ensure the confidentiality, integrity and anti-tampering capability of the scheduling commands during the transmission to the field execution unit.

[0030] In this embodiment, the digital twin module is configured to provide a physical, high-fidelity simulation environment that can accurately model the electrical topology of the data center and the dynamic response characteristics of its energy assets, thereby providing accurate simulation data for the online meta-adaptive cycle.

[0031] This embodiment also provides a multi-objective scheduling method for data centers that integrates digital twins and meta-learning. Please refer to [link / reference]. Figure 2 As shown, the method includes the following steps: Step 301: Real-time acquisition of data center load data, renewable energy power generation data, energy storage system operation status data, power grid information, and sensor data related to the health of key power equipment through the multimodal data acquisition module; Specifically, the data collection in step 301 includes: acquiring real-time load power, photovoltaic / wind power output, energy storage system SOC, grid time-of-use electricity price, carbon emission factor, and health-related data such as energy storage battery temperature and cycle count through sensors and communication interfaces deployed throughout the data center.

[0032] Step 302: The collected data is denoised, missing values ​​are filled and normalized by the data processing and feature engineering module. Combined with the basic health model trained based on historical data and the real-time load change rate, a health score reflecting the current pressure status of the power equipment is dynamically calculated. Step 303: Invoke the prediction module to generate a load power prediction sequence and a renewable energy power generation prediction sequence for the future scheduling time domain based on historical data and external meteorological information; Specifically, in step 303, load and power generation forecasting includes: using a time series forecasting model, such as a Long Short-Term Memory Network (LSTM), inputting historical data, and outputting load and renewable energy power generation forecast curves for the next 24 hours or longer.

[0033] Step 304: Combine the real-time running data obtained in step 301, the health score dynamically calculated in step 302, and the prediction sequence generated in step 303 to construct a high-dimensional system state vector that can comprehensively characterize the current and expected future state of the system. Specifically, in step 304, the system state vector construction includes: concatenating real-time measurement values ​​(such as current SOC, electricity price), dynamic health scores, and future prediction sequences into a high-dimensional vector. This vector fully describes the current state and future trends of the system and serves as the input to the reinforcement learning agent.

[0034] Step 305: In a digital twin environment that is precisely mapped to the physical system, launch a reinforcement learning agent that has been pre-trained offline. Specifically, in step 305, starting the meta-reinforcement learning agent includes: loading an agent model that has been pre-trained offline in the digital twin environment; the model has learned general knowledge to deal with various scheduling tasks, and its network parameters are in a good initialization state.

[0035] Step 306: The agent takes the system state vector constructed in step 304 as input and performs an online meta-adaptation process. By performing a small number of simulated interactions in the digital twin environment, it quickly fine-tunes its policy network to adapt to the current scheduling task. For details, see Figure 3 In step 306, regarding online meta-adaptation: this step corresponds to the inner loop of the MAML algorithm; the agent, starting from the state vector of step 304, interacts with the simulation system several times in the digital twin environment, collecting a small amount of empirical data relevant to the current specific task. Using this data, the model parameters are adjusted through one or more steps of gradient descent. Fine-tuning is performed to obtain adaptive parameters. This process aims to quickly specialize the strategy to suit the current task.

[0036] Step 307: Through a finely tuned agent, solve a Markov decision process with total electricity cost, total carbon emissions and power equipment health as joint optimization objectives, and output the optimal scheduling action for the current decision cycle. Specifically, in step 307, the optimal action solution includes: using the policy network fine-tuned in step 306. The optimal scheduling action for the current state is determined. This process is modeled as a Markov decision process, with its state space defined by the vector in step 304 and its action space encompassing multiple decision variables, including energy storage charging and discharging power. The agent's goal is to maximize the reward function defined in the given process. The long-term cumulative expectation. The design of the reward function guides the agent to minimize the joint costs of economic, environmental, and power equipment degradation while satisfying all physical constraints.

[0037] Step 308: The optimal scheduling action output in step 307 is converted into control commands for specific power equipment through the (safety) control and scheduling module, and then sent to the field execution unit through an encrypted channel. Specifically, in step 308, the secure issuance of instructions includes converting the abstract action (such as the energy storage power setting value) output in step 307 into control instructions conforming to industry-standard communication protocols. According to claim 7, the instructions are securely sent to execution units such as energy storage converters and smart meters within the data center using IPSec VPN and AES-256 encryption technology.

[0038] Step 309: At the start of the next decision cycle, return to step 301 to form a closed-loop scheduling mechanism that continuously senses, predicts, optimizes, and executes safely in the rolling time domain.

[0039] Specifically, in step 309, the rolling optimization includes: after a scheduling cycle ends, the system returns to step 301, obtains the latest system state, and repeats the entire process. This rolling execution method forms a closed-loop control, enabling the system to continuously adapt to environmental changes.

[0040] In this embodiment, see Figure 3 The training of the meta-reinforcement learning agent in this invention is divided into two stages: Offline meta-training phase: The goal of this phase is to learn and obtain meta-initial model parameters that can generalize across tasks. The training process samples a large number of diverse scheduling tasks from the historical database. For each sampled task, a fast adaptation process is executed within an inner loop to obtain task-related parameters. Then, evaluate it on a new validation dataset. The performance of the sampled tasks is then assessed. Finally, the gradient of the meta-objective function is calculated based on the sum of the validation performance of all sampled tasks, and the meta-initial parameters are updated. .

[0041] Online deployment phase: When the system encounters an unseen scheduled task, the inner loop will be activated to process the initial parameters obtained during offline training. Make quick fine adjustments.

[0042] Whether in the inner or outer loop, the policy update uses the Proximal Policy Optimization (PPO) algorithm.

[0043] In this embodiment, the meta-reinforcement learning agent employs a Model-Independent Meta-Learning (MAML) algorithm to construct a two-layer learning framework of offline meta-training and online meta-adaptation. Furthermore, in the policy update of this framework, a Proximal Policy Optimization (PPO) algorithm is used to perform gradient updates, thereby improving training stability and convergence efficiency. Specifically, this implementation includes an outer loop for offline meta-training and an inner loop for online meta-adaptation, and gradient updates are performed by maximizing a truncated surrogate objective function.

[0044] In this embodiment, the online decision optimization steps in steps 305 to 307 involve modeling the scheduling problem as a Markov Decision Process (MDP), which is defined as follows: The state space is defined by the high-dimensional system state vector constructed in step 304, which includes real-time measurements, future prediction sequences, and health scores of dynamic power equipment. The action space is a set of scheduling decisions for controllable energy assets within a data center, including the charging and discharging power settings of energy storage systems and the power settings for interaction with the power grid. The reward function, which evaluates the immediate benefits of performing a specific action in a given state, is designed to guide the agent to learn and achieve the joint optimization objective.

[0045] In this embodiment, the reward function in the Markov decision process At each time step It is precisely defined as: The first term is the negative of the joint cost function, designed to incentivize the agent to minimize the cost; the second term... As a penalty term, it is used to punish behaviors that violate system operating constraints, and it is precisely defined as follows: ,in, The deviation of the power balance constraint has a value of ; All of these are preset, sufficiently large positive penalty coefficients (the definitions of the other coefficients are described above) to ensure that the policies learned by the agent strictly adhere to the key physical and operational constraints defined above.

[0046] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data center multi-objective scheduling device fusing digital twinning and meta-learning, characterized in that, This device is applied to the integrated energy system of the data center, which includes a local energy system connected to the public power grid. The local energy system includes a renewable energy system and an energy storage system. The device comprises: The multimodal data acquisition module is configured to monitor and acquire multi-source heterogeneous data from the data center, the energy system, the public power grid, and power equipment in real time. The data preprocessing and feature engineering module is configured to process the collected multi-source heterogeneous data and dynamically calculate the health score of the power equipment. The prediction module, which communicates with the data preprocessing and feature engineering module, is configured to generate a predicted power generation sequence of the renewable energy system and a predicted load power sequence for future scheduling cycles based on historical data and external meteorological information. The digital twin module is configured to construct a high-fidelity virtual model of the data center's source-grid-load-storage energy system based on the physical parameters of power equipment and the grid topology, and to provide a secure and isolated simulation environment for online optimization of scheduling strategies; The multi-objective optimization module, as the decision-making core of the device, adopts a meta-reinforcement learning agent at its core. It is configured to perform online fine-tuning and simulation verification in the digital twin module based on real-time measurement values, future prediction data, and health scores of dynamic power equipment. By solving the problem of joint optimization objectives, the optimal power scheduling plan is determined. The control and scheduling module communicates with the multi-objective optimization module and is configured to generate control commands based on the optimal power scheduling plan and issue them to the execution units in the energy system.

2. The data center multi-objective scheduling apparatus fusing digital twin and meta-learning according to claim 1, wherein, The meta-reinforcement learning agent in the multi-objective optimization module is configured to execute the following mechanism: The offline meta-training loop mechanism is configured to train on a series of tasks with different operating scenarios sampled from historical data, in order to learn and solve for a set of initial parameters of a meta-learning model that can quickly adapt to new tasks. The online meta-adaptive loop mechanism is configured to, when faced with a new scheduling task defined by the latest prediction data, rapidly and finely adjust the initial parameters of the meta-learning model in the simulation environment provided by the digital twin module using interactive data related to the current new task, thereby generating a specialized scheduling strategy for the current scheduling task.

3. The data center multi-objective scheduling device integrating digital twins and meta-learning according to claim 1, characterized in that, The multi-objective optimization module solves a problem of jointly optimizing objectives, configured to solve a problem of minimizing a joint cost function defined as: ; wherein, is a dispatch policy, is a discrete time step within the dispatch time domain, represents an economic cost at time instant, are dynamic weight coefficients for the economic cost, the environmental cost, and the power equipment degradation cost, respectively, represents an environmental cost at time instant, and represents a power equipment degradation cost at time instant.

4. The data center multi-objective scheduling device integrating digital twins and meta-learning according to claim 3, characterized in that, The process of minimizing the joint cost function is subject to one or more constraints, including power balance constraints, energy storage system operation constraints, and grid interaction power constraints.

5. The data center multi-objective scheduling device integrating digital twins and meta-learning according to claim 1, characterized in that, The health score of the power equipment is dynamically calculated based on multimodal physical sensor data related to the operational health of the power equipment and the real-time load change rate of the data center, specifically including: A basic health assessment model is trained based on historical data from multimodal physical sensors to obtain a raw health score reflecting the internal state of power equipment. Calculate an external operating stress factor based on the real-time load change rate of the data center; The original health score is dynamically combined with the external operational stress factors to obtain the final dynamic health score.

6. The data center multi-objective scheduling device integrating digital twins and meta-learning according to claim 2, characterized in that, The digital twin module is configured to provide a physical, high-fidelity simulation environment that can accurately model the electrical topology of the data center and the dynamic response characteristics of its energy assets, thereby providing accurate simulation data for the online meta-adaptive cycle.

7. A multi-objective scheduling method for data centers integrating digital twins and meta-learning, characterized in that, The method includes the following steps: Step 301: Monitor and collect multi-source heterogeneous data from the data center, the energy system, the public power grid, and power equipment in real time; Step 302: Process the collected multi-source heterogeneous data and dynamically calculate the health score reflecting the current pressure status of the power equipment; Step 303: Based on historical data and external meteorological information, generate a load power prediction sequence and a renewable energy power generation prediction sequence for the future scheduling time domain; Step 304: Combine the real-time running data obtained in step 301, the health score dynamically calculated in step 302, and the prediction sequence generated in step 303 to construct a high-dimensional system state vector that can comprehensively characterize the current and expected future state of the system. Step 305: In a digital twin environment that is precisely mapped to the physical system, launch a reinforcement learning agent that has been pre-trained offline. Step 306: The agent takes the system state vector constructed in step 304 as input and performs an online meta-adaptation process. By simulating interaction in the digital twin environment, it quickly fine-tunes its policy network to adapt to the current scheduling task. Step 307: Solve the Markov decision process of the joint optimization objective through the fine-tuned agent, and output the optimal scheduling action for the current decision cycle. Step 308: Through the control and scheduling module, the optimal scheduling action output in step 307 is converted into control commands for specific power equipment and sent to the field execution unit. Step 309: At the start of the next decision cycle, return to step 301 to form a closed-loop scheduling mechanism that continuously senses, predicts, optimizes, and executes safely in the rolling time domain.

8. The data center multi-objective scheduling method integrating digital twins and meta-learning according to claim 7, characterized in that, In terms of algorithm implementation, the intelligent agent adopts the Model Independent Meta-Learning (MAML) algorithm to construct a two-layer learning framework of offline meta-training and online meta-adaptation; Furthermore, in the policy update of the aforementioned framework, the proximal policy optimization (PPO) algorithm is used to perform gradient updates to improve training stability and convergence efficiency.

9. The data center multi-objective scheduling method integrating digital twins and meta-learning according to claim 7, characterized in that, The online decision optimization steps in steps 305 to 307 involve modeling the scheduling problem as a Markov Decision Process (MDP), which includes: The state space, defined by the high-dimensional system state vector constructed in step 304, includes real-time measurements, future prediction sequences, and health scores of dynamic power equipment. Action space is the set of scheduling decisions for controllable energy assets within a data center, including the charging and discharging power setpoints of energy storage systems and the power interaction setpoints with the power grid. The reward function is used to evaluate the immediate benefits of performing a specific action in a specific state and is designed to guide the agent to learn to achieve the joint optimization objective.

10. The data center multi-objective scheduling method integrating digital twins and meta-learning according to claim 7 or 9, characterized in that, The reward function in the Markov decision process At each time step It is precisely defined as: , The first term is the negative of the joint cost function, designed to incentivize the agent to minimize the cost; the second term... As a penalty, it is used to punish behaviors that violate system operating constraints, so as to ensure that the policies learned by the agent strictly adhere to physical and operational constraints.