Comprehensive energy optimization scheduling method and device, storage medium and computer equipment

Through layered decoupling control and intelligent decision-making algorithms, the problems of second-level fluctuations and multivariate coupling conflicts in the integrated energy scheduling system are solved, and efficient energy scheduling and optimization are achieved.

CN120509675APending Publication Date: 2025-08-19GUANGDONG POWER GRID CORP ZHAOQING POWER SUPPLY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510691073.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing integrated energy scheduling system cannot adapt to second-level fluctuations and multivariate coupling conflicts, resulting in poor real-time and accuracy of scheduling instructions.

Method used

By collecting multi-source heterogeneous energy data in real time, using hierarchical decoupling control, using event-triggered reinforcement learning model to generate second-level instructions, incremental model prediction control generates minute-level instructions, and security verification is performed based on KKT condition constraints to generate optimization scheduling instructions.

Benefits of technology

Real-time scheduling at millisecond level is realized, real-time and accuracy of energy scheduling is improved, the actual operation characteristics of the equipment are ensured, the calculation accuracy and response capabilities are balanced, and multi-objective collaborative optimization is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509675A_ABST
    Figure CN120509675A_ABST
Patent Text Reader

Abstract

According to the comprehensive energy optimization scheduling method and device, the storage medium and the computer equipment provided by the invention, during energy scheduling, multi-source heterogeneous energy data in a comprehensive energy system are collected in real time and are divided into the fast variable data and the slow variable data, so that hierarchical decoupling control is realized; wherein the reinforcement learning model can be triggered through an event to generate a second-level instruction corresponding to the fast variable data, and millisecond-level dynamic response is realized; in addition, a minute-level instruction corresponding to slow variable data can be generated through incremental model predictive control, an equipment physical equation is embedded in the process as a constraint condition, the actual operation characteristics of load equipment in an optimization result can be ensured, and the calculation precision and the real-time response capability are balanced. And finally, security verification processing can be carried out on the generated instruction based on KKT condition constraints to obtain an optimized scheduling instruction so as to further improve the accuracy of the instruction, and the optimized scheduling instruction is utilized to carry out energy optimized scheduling so as to realize millisecond-level real-time scheduling and multi-target collaborative optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of energy scheduling technology, and in particular to a comprehensive energy optimization scheduling method, device, storage medium and computer equipment. Background Art

[0002] With the global energy structure shifting and the continued advancement of electricity market reforms, integrated energy systems, as the key link connecting all types of energy production and consumption, are undergoing a profound transformation from centralized dispatch to distributed, coordinated dispatch. This shift aims to more efficiently utilize energy resources and promote the large-scale integration and consumption of renewable energy.

[0003] However, the inherent characteristics of distributed energy resources, namely the high volatility and uncertainty of their output, such as second-level variations in light intensity and minute-level load changes, make it difficult for traditional static dispatch models to respond to dynamic disturbances in real time. Furthermore, with the continued opening and deepening of the electricity market, high-frequency trading is becoming increasingly prevalent, and spot electricity prices are becoming the norm, placing higher demands on the real-time decision-making capabilities of dispatch systems. In summary, existing integrated energy dispatch systems suffer from issues such as an update frequency that cannot adapt to second-level fluctuations and command oscillation caused by multi-variable coupling conflicts, resulting in poor real-time and accuracy of dispatch commands. Summary of the Invention

[0004] The purpose of this application is to solve at least one of the above-mentioned technical defects, especially the technical defects in the existing technology that the integrated energy scheduling system cannot adapt to the second-level fluctuations in update frequency, multi-variable coupling conflicts lead to instruction oscillations, etc., resulting in poor real-time and accuracy of scheduling instructions.

[0005] This application provides a comprehensive energy optimization scheduling method, which includes:

[0006] Real-time collection of multi-source heterogeneous energy data in an integrated energy system; the multi-source heterogeneous energy data includes fast variable data and slow variable data;

[0007] Determine the state variable fluctuation of the integrated energy system according to the fast variable data, and when the state variable fluctuation exceeds a preset threshold, generate a second-level instruction corresponding to the fast variable data through an event-triggered reinforcement learning model;

[0008] Generating minute-level instructions corresponding to the slow variable data through incremental model predictive control; the incremental model predictive control embeds device physical equations as constraints;

[0009] After the target instruction is generated, the target instruction is security-checked based on the KKT condition constraint to obtain the optimized scheduling instruction, and the optimized scheduling instruction is used to perform energy optimization scheduling on the integrated energy system; wherein, the target instruction includes second-level instructions and minute-level instructions.

[0010] Optionally, determining the state variable fluctuation amount of the integrated energy system according to the fast variable data includes:

[0011] Determine the state variable at the previous moment and the state variable at the current moment according to the fast variable data;

[0012] The state variable at the current moment is subtracted from the state variable at the previous moment to obtain the state variable fluctuation amount of the integrated energy system at the current moment.

[0013] Optionally, the generating of second-level instructions corresponding to the fast variable data by the event-triggered reinforcement learning model includes:

[0014] The fast variable data is input into the event-triggered reinforcement learning model, so that the event-triggered reinforcement learning model uses a double-buffered experience pool to update the control strategy and outputs a second-level instruction.

[0015] Optionally, the control strategy update formula of the event-triggered reinforcement learning model includes:

[0016]

[0017]

[0018] Where, represents the policy parameters; represents the learning rate; represents the action value function, which is used to predict the long-term benefits of performing action a in state s; Indicates policy parameters gradient; represents the sample set collected from the double-buffered experience pool; Indicates the real-time cache sampling ratio; Indicates real-time cache; Indicates long-term caching.

[0019] Optionally, the generating minute-level instructions corresponding to the slow variable data through incremental model predictive control includes:

[0020] Construct an objective function and embed the device physics equations as constraints to form an incremental model;

[0021] The slow variable data is input into the incremental model, so that the incremental model performs incremental adjustment based on the optimization result of the previous cycle of the slow variable data, and outputs the minute-level instructions of the current cycle.

[0022] Optionally, performing security verification processing on the target instruction based on the KKT condition constraint to obtain the optimized scheduling instruction includes:

[0023] Performing a security check on the target instruction using KKT conditional constraints to obtain a check result;

[0024] When the verification result is failed, a correction algorithm is used to correct the target instruction to obtain an optimized scheduling instruction;

[0025] When the verification result is passed, the target instruction is used as the final optimized scheduling instruction.

[0026] Optionally, the calculation expression of the correction algorithm includes:

[0027]

[0028] Where, Indicates optimized scheduling instructions; Indicates the original second-level instruction or minute-level instruction; Indicates correction compensation; Represents the orthographic projection operator.

[0029] This application also provides a comprehensive energy optimization scheduling device, including:

[0030] A data acquisition module is used to collect multi-source heterogeneous energy data in the integrated energy system in real time; the multi-source heterogeneous energy data includes fast variable data and slow variable data;

[0031] a fast variable prediction module, configured to determine the state variable fluctuation of the integrated energy system based on the fast variable data, and, when the state variable fluctuation exceeds a preset threshold, generate a second-level instruction corresponding to the fast variable data through an event-triggered reinforcement learning model;

[0032] A slow variable prediction module, configured to generate minute-level instructions corresponding to the slow variable data through incremental model predictive control; the incremental model predictive control embeds device physical equations as constraints;

[0033] The instruction optimization module is used to perform security verification on the target instruction based on the KKT condition constraint after the target instruction is generated, obtain the optimized scheduling instruction, and use the optimized scheduling instruction to perform energy optimization scheduling on the integrated energy system; wherein, the target instruction includes second-level instructions and minute-level instructions.

[0034] The present application also provides a storage medium, which stores computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the comprehensive energy optimization scheduling method as described in any of the above embodiments.

[0035] The present application also provides a computer device, comprising: one or more processors, and a memory;

[0036] The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the comprehensive energy optimization scheduling method as described in any one of the above embodiments are executed.

[0037] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0038] The integrated energy optimization scheduling method, device, storage medium and computer equipment provided by the present application can collect multi-source heterogeneous energy data in the integrated energy system in real time during energy scheduling; the multi-source heterogeneous energy data here includes fast variable data and slow variable data, so that hierarchical decoupling control can be achieved to avoid multi-variable coupling conflicts; wherein, for fast variable data, the state variable fluctuation of the integrated energy system can be determined based on the fast variable data, and then when the state variable fluctuation exceeds the preset threshold, the event-triggered reinforcement learning model is used to generate a second-level instruction corresponding to the fast variable data, thereby achieving millisecond-level dynamic response and improving the real-time performance of energy scheduling; for slow variable data, the incremental model predictive control can be used to generate a minute-level instruction corresponding to the slow variable data. The incremental model predictive control here embeds the device physical equation as a constraint condition, so as to ensure the actual operating characteristics of the load equipment in the optimization result, balance the calculation accuracy and real-time response capability. In addition, the generated second-level instruction or minute-level instruction can be security-checked based on the KKT condition constraint to obtain the optimized scheduling instruction to further improve the accuracy of the instruction. Finally, the optimized scheduling instruction can be used to perform energy optimization scheduling on the integrated energy system, achieving millisecond-level real-time scheduling and multi-objective collaborative optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0040] Figure 1 A flow chart of a comprehensive energy optimization scheduling method provided in an embodiment of the present application;

[0041] Figure 2 A flowchart of an instruction security verification process provided in an embodiment of the present application;

[0042] Figure 3 A schematic diagram of the structure of a comprehensive energy optimization and scheduling device provided in an embodiment of the present application;

[0043] Figure 4 A schematic diagram of the internal structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0044] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0045] Due to the inherent characteristics of distributed energy resources, namely the high volatility and uncertainty of their output, such as second-level variations in light intensity and minute-level load fluctuations, traditional static dispatch models struggle to respond to dynamic disturbances in real time. Furthermore, with the continued opening and deepening of the electricity market, high-frequency trading is becoming increasingly prevalent, and second-level fluctuations in spot electricity prices have become the norm, placing higher demands on the real-time decision-making capabilities of dispatch systems. In summary, existing integrated energy dispatch systems suffer from issues such as an update frequency that cannot adapt to second-level fluctuations and command oscillation caused by multi-variable coupling conflicts, resulting in poor real-time and accuracy of dispatch instructions.

[0046] Based on this, this application proposes the following technical solutions, please refer to the following for details:

[0047] In one embodiment, Figure 1 As shown, Figure 1 This is a flow chart of a comprehensive energy optimization scheduling method provided in an embodiment of the present application; this application also provides a comprehensive energy optimization scheduling method, which specifically includes the following:

[0048] S110: Real-time collection of multi-source heterogeneous energy data in the integrated energy system; the multi-source heterogeneous energy data includes fast variable data and slow variable data.

[0049] In this step, during energy scheduling, computer equipment can collect multi-source heterogeneous energy data in the integrated energy system in real time; the multi-source heterogeneous energy data here can also be divided into fast variable data and slow variable data according to the update frequency, so as to realize hierarchical decoupling control and avoid multi-variable coupling conflicts.

[0050] Among them, fast variable data refers to variable data with a high update frequency and rapid changes over time. It involves energy system parameters that are dynamically adjusted in a short period of time, such as new energy output PRE(t), electricity price p(t), energy storage SOCS(t), grid operation status (voltage V(t), power flow F(t)), etc.; while slow variable data refers to variable data with a low update frequency and relatively slow changes. It involves long-term planning and strategic decision-making, such as equipment efficiency depreciation and carbon emission quotas.

[0051] Specifically, computer equipment can collect multi-source heterogeneous energy data from integrated energy systems in real time through sensors, smart meters, data acquisition and monitoring systems, and other devices. Because different types of data have different change frequencies and response characteristics, in order to improve the efficiency and stability of energy scheduling, this application can classify data into fast-changing data and slow-changing data based on its update frequency.

[0052] It is understandable that through the division of fast variable data and slow variable data, computer equipment can adopt a hierarchical decoupling control method, handing over the processing of fast variable data to the real-time scheduling layer to quickly respond to dynamic changes in the system, while the slow variable data is handed over to the long-term optimization layer for trend prediction and strategy planning, so as to avoid coupling conflicts between fast and slow variables, so that short-term decisions will not be overly restricted by long-term parameter changes, and at the same time, long-term optimization strategies will not be frequently adjusted due to short-term fluctuations, thereby achieving efficient system energy scheduling.

[0053] S120: Determine the state variable fluctuation of the integrated energy system based on the fast variable data, and when the state variable fluctuation exceeds a preset threshold, generate a second-level instruction corresponding to the fast variable data through an event-triggered reinforcement learning model.

[0054] In this step, after the fast variable data is collected in step S110, the computer device can determine the state variable fluctuation of the integrated energy system based on the fast variable data, and when the state variable fluctuation exceeds the preset threshold, the event-triggered reinforcement learning (ET-DRL) model generates second-level instructions corresponding to the fast variable data, thereby achieving millisecond-level dynamic response and improving the real-time performance of energy scheduling.

[0055] The event-triggered reinforcement learning model refers to an intelligent decision-making approach that combines event-triggered control (ETC) and deep reinforcement learning (DRL). It aims to reduce computational burden and communication overhead while maintaining efficient decision-making capabilities. Specifically, it introduces an event-triggered mechanism based on reinforcement learning, enabling the agent to make decisions only when key states change, rather than continuously controlling at fixed intervals. Therefore, in this application, the model can trigger calculations only when state changes exceed a threshold, thereby reducing computational complexity and improving system stability.

[0056] Specifically, to quantify changes in system states, the computer can continuously calculate the fluctuation of the state variables based on fast variable data—that is, the degree of change in the fast variable between adjacent time steps. Then, based on preset scheduling strategies and stability requirements, it can set a preset threshold to measure whether the fluctuation of the state variables exceeds an acceptable range. When the computer detects that the fluctuation of the state variables exceeds the preset threshold, it can immediately trigger an event-triggered reinforcement learning model, enter intelligent decision-making mode, and generate second-level instructions corresponding to the fast variable data.

[0057] Furthermore, the preset threshold can be adaptively adjusted based on historical volatility, or can be set based on the actual system acceptance, such as when the electricity price jump Δp>5% or when the new energy output sudden change ΔPRE>10%Prated. The adaptive dynamic threshold calculation formula can be expressed as follows:

[0058]

[0059] Where, Indicates the preset threshold; It represents the threshold coefficient, which can be set according to the acceptable fluctuation of the project; Indicates the standard deviation of the historical state variable fluctuation.

[0060] The triggering conditions are as follows:

[0061]

[0062] Where, Indicates the judgment result; represents the fluctuation of the state variable at time t; Indicates the preset threshold.

[0063] Understandably, because the event-triggered mechanism employed by the event-triggered reinforcement learning model reduces unnecessary computational burden and only makes decisions when key states change, the system is able to achieve millisecond-level responses without being constrained by the excessive computational overhead of traditional fixed-time-step reinforcement learning. Furthermore, post-execution feedback data can be fed back to the event-triggered reinforcement learning model in real time for optimization and training, enabling the model to adapt more quickly to similar events in the future, improving the intelligence and responsiveness of energy scheduling.

[0064] S130: Generating minute-level instructions corresponding to slow variable data through incremental model predictive control; the incremental model predictive control embeds device physical equations as constraints.

[0065] In this step, after the slow variable data is collected in step S110, the computer device can generate minute-level instructions corresponding to the slow variable data through incremental model predictive control (iMPC). The iMPC here embeds the device physical equations as constraints, thereby ensuring that the optimization result can reflect the actual operating characteristics of the load device and balance calculation accuracy and real-time response capabilities.

[0066] Incremental model predictive control (MPC) is an improved model predictive control (MPC) method that improves computational efficiency and real-time performance by optimizing the increments of the control variables—the amount of change between adjacent time steps—while ensuring that the control strategy complies with the physical characteristics of the equipment. Because incremental model predictive control only optimizes the increments of the control variables, its computational complexity is lower than that of traditional model predictive control, thus ensuring optimization accuracy while improving real-time computing capabilities.

[0067] Furthermore, device physics equations are mathematical expressions that describe the operating characteristics, behavioral constraints, and state evolution of physical devices. These equations are typically derived from physical laws, empirical models, or experimental data. They characterize the relationship between a device's input, output, and internal state, and are applied as constraints in optimization scheduling, control algorithms, or simulation analysis to ensure that optimization calculations meet the device's actual operational capabilities and avoid engineering-infeasible results.

[0068] Specifically, the computer device can combine the device's physical equations and use incremental model predictive control to generate minute-level instructions corresponding to the slow variable data, ensuring that the optimization results meet the device's actual operating capabilities. During this process, the computer device can use a rolling optimization mechanism to obtain the latest slow variable data at regular intervals and combine historical data to predict future system states. For each optimization cycle, incremental model predictive control calculates the incremental value of the current control variable instead of directly calculating the absolute value of the control variable. Finally, after completing the incremental optimization calculation, minute-level instructions can be generated.

[0069] For example, in energy storage optimization and scheduling, incremental model predictive control can calculate the change in energy storage charging and discharging power ΔP, rather than directly solving for the charging and discharging power P. This ensures that energy storage equipment does not experience drastic power fluctuations and thus avoids impacts on the power grid. In load management, incremental model predictive control can smoothly adjust user-side load demand by adjusting the rate of change of load response ΔL, preventing large-scale load fluctuations from affecting the stability of the power system.

[0070] It is understandable that because incremental model predictive control uses an incremental optimization method, its computational complexity is significantly reduced compared to traditional model predictive control. Therefore, it can improve real-time performance while maintaining optimization accuracy, enabling computer equipment to solve optimization problems in a shorter timeframe and ensuring minute-level dispatch response. Furthermore, incremental model predictive control can adapt to the dynamic changes of integrated energy systems. By continuously adjusting the optimization model's prediction parameters through rolling updates of historical data, the optimization strategy consistently aligns with the system's actual operating conditions.

[0071] S140: After the target instruction is generated, the target instruction is security-checked based on the KKT condition constraint to obtain the optimized scheduling instruction, and the optimized scheduling instruction is used to perform energy optimization scheduling on the integrated energy system; wherein, the target instruction includes second-level instructions and minute-level instructions.

[0072] In this step, after generating the dispatch instructions in steps S120 and S130, the computer device can perform security verification on the target instructions based on KKT (Karush-Kuhn-Tucker) constraints to obtain optimized dispatch instructions, further improving the accuracy of the instructions. Finally, the computer device can use the optimized dispatch instructions to optimize energy scheduling for the integrated energy system, achieving millisecond-level real-time scheduling and multi-objective collaborative optimization.

[0073] Among them, the types of target instructions include second-level instructions and minute-level instructions. In addition, the KKT condition is a necessary condition in mathematical optimization and is applicable to nonlinear optimization problems with constraints. In the scheduling optimization problem of the integrated energy system of this application, the KKT condition can be used as a constraint to optimize the solution process, ensure that the optimized scheduling instructions meet the physical constraints and safe operation requirements of the equipment, avoid instructions exceeding the equipment's tolerance range, and thus improve the accuracy and reliability of the scheduling strategy.

[0074] Specifically, after generating a scheduling instruction, the computer device can perform a security check on the target instruction to ensure that it meets the system's operating constraints. To this end, the computer device can constrain the target instruction based on the KKT conditional constraints to ensure that the resulting optimized scheduling instruction can achieve the optimal result while meeting the constraints.

[0075] For example, during the optimization process, the computer can verify the feasibility and optimality of the target instructions based on KKT constraints, ensuring that the optimized instructions do not deviate from the actual operating boundaries of the equipment and do not lead to unreasonable resource allocation. For example, when the charge and discharge power of the energy storage system exceeds the physical limits of the equipment, the KKT constraints can force the optimization results to meet the equipment's operating constraints.

[0076] Finally, the optimized dispatch instructions obtained through KKT conditional constraint optimization can be applied by computer equipment to the integrated energy system to achieve precise control of links such as renewable energy power generation, energy storage system charging and discharging, and power grid flow dispatching. For example, the control of energy storage charging and discharging power and adjustable load start and stop corresponding to fast variable data, as well as the control of load rate adjustment and carbon emission quota adjustment corresponding to slow variable data, can be used to dynamically adjust the optimization strategy by real-time monitoring of fast and slow variables to ensure that the final dispatch instructions can adapt to the complex and changing energy system environment.

[0077] Taking the second-level instructions generated based on fast variable data as an example, the energy storage charging and discharging power and adjustable load start and stop instructions can be expressed as follows:

[0078]

[0079]

[0080] Where, represents the charging and discharging power of the energy storage system at time t; Represents the continuous actions output by the Actor network; Indicates the start and stop command of the adjustable load; Represents the action value predicted by the Critic network, where Represents a sign function, which is used to convert the Q value into a discrete action, such as 1 for turning on the adjustable load and -1 for turning off the adjustable load.

[0081] It should be noted that the Actor network and the Critic network are the two sub-networks that constitute the core of the event-triggered reinforcement learning model; among them, the Actor network is mainly responsible for outputting action strategies, and the Critic network is mainly responsible for evaluating the value of the actions taken by the Actor.

[0082] In the above embodiment, during energy scheduling, multi-source heterogeneous energy data in the integrated energy system can be collected in real time; the multi-source heterogeneous energy data here includes fast variable data and slow variable data, thereby realizing hierarchical decoupling control and avoiding multi-variable coupling conflicts; wherein, for fast variable data, the state variable fluctuation of the integrated energy system can be determined based on the fast variable data, and then when the state variable fluctuation exceeds the preset threshold, the event-triggered reinforcement learning model is used to generate second-level instructions corresponding to the fast variable data, thereby achieving millisecond-level dynamic response and improving the real-time performance of energy scheduling; for slow variable data, minute-level instructions corresponding to the slow variable data can be generated through incremental model predictive control. Here, the incremental model predictive control embeds the device physical equation as a constraint condition, thereby ensuring that the optimization result reflects the actual operating characteristics of the load equipment and balancing calculation accuracy and real-time response capability. In addition, the generated second-level instructions or minute-level instructions can be security-verified based on the KKT condition constraint to obtain optimized scheduling instructions to further improve the accuracy of the instructions. Finally, the optimized scheduling instructions can be used to perform energy optimization scheduling on the integrated energy system, achieving millisecond-level real-time scheduling and multi-objective collaborative optimization.

[0083] In one embodiment, the process of determining the state variable fluctuation amount of the integrated energy system based on the fast variable data in step S120 may include:

[0084] S121: Determine the state variables at the previous moment and the state variables at the current moment according to the fast variable data.

[0085] S122: Subtract the state variable at the previous moment from the state variable at the current moment to obtain the state variable fluctuation of the integrated energy system at the current moment.

[0086] In this embodiment, after collecting fast variable data, the computer device can determine the state variables at the previous moment and the state variables at the current moment based on the fast variable data, and then subtract the state variables at the previous moment from the state variables at the current moment to obtain the state variable fluctuation of the integrated energy system at the current moment. This fluctuation can then be used as an important reference parameter for energy scheduling optimization to guide subsequent scheduling strategy adjustments and achieve precise control and real-time optimization of the integrated energy system.

[0087] Specifically, the computer device can first parse and preprocess the fast-variable data to extract key energy scheduling information, determine the current operating status of the integrated energy system, and calculate the current state variables based on this operating status. Furthermore, the computer device can obtain the previous state variables by using the previous operating status. Subsequently, the computer device can compare the current state variables with the previous state variables and calculate the change between them, thereby obtaining the fluctuation of the state variables of the integrated energy system at the current moment.

[0088] More specifically, the calculation formula for the state variable fluctuation is as follows:

[0089]

[0090] Where, represents the fluctuation of the state variable at time t; Represents the state variables of the i-th state variable at time t, including renewable energy output PRE(t), electricity price p(t), energy storage SOCS(t), grid operation status (voltage V(t), power flow F(t)), etc.; represents the state variable of the i-th state variable at time t-1.

[0091] In one embodiment, the process of generating second-level instructions corresponding to fast variable data by triggering the reinforcement learning model through events in step S120 may include:

[0092] S123: Inputting the fast variable data into the event-triggered reinforcement learning model, so that the event-triggered reinforcement learning model uses the double-buffered experience pool to update the control strategy and outputs a second-level instruction.

[0093] In this embodiment, when it is monitored that the fluctuation of the state variable exceeds a preset threshold, the computer device can input the fast variable data into the event-triggered reinforcement learning model, so that the event-triggered reinforcement learning model uses a double-buffered experience pool to update the control strategy and output second-level instructions.

[0094] It's important to note that the dual-buffered experience pool is divided into a real-time cache and a long-term cache. The real-time cache stores the most recent 10 seconds of experience data, has a small capacity, and can be used for rapid policy fine-tuning. The long-term cache stores historical data, has a large capacity, and is used for offline deep learning training. Therefore, using the dual-buffered experience pool, event-triggered reinforcement learning models can reduce model complexity and optimize computational efficiency.

[0095] Specifically, after a triggering event triggers the reinforcement learning model, the computer device can input fast variable data into the reinforcement learning model. The model uses a deep reinforcement learning algorithm to evaluate the current system state and predict the optimal control strategy. During this process, the event-triggered reinforcement learning model can adopt a double-buffered experience pool mechanism to achieve efficient control strategy updates by storing and managing past interaction experience. In the double-buffered experience pool, a portion of the experience is used for short-term updates to quickly adapt to current state fluctuations, while the other portion of the experience is used for long-term optimization to enhance the model's global generalization capabilities. Based on this, the event-triggered reinforcement learning model can perceive the state of the environment and model future trends based on historical data and the current state to determine the optimal second-level scheduling instructions.

[0096] In one embodiment, the control strategy update formula of the event-triggered reinforcement learning model in step S123 may include:

[0097]

[0098]

[0099] Where, represents the policy parameters; represents the learning rate; represents the action value function, which is used to predict the long-term benefits of performing action a in state s; Indicates policy parameters gradient; represents the sample set collected from the double-buffered experience pool; Indicates the real-time cache sampling ratio; Indicates real-time cache; Indicates long-term caching.

[0100] In this embodiment, through the control strategy update formula, the event-triggered reinforcement learning model can dynamically adjust the control strategy based on data characteristics at different time scales to achieve efficient and stable energy scheduling. Strategy parameters can be continuously adjusted during the reinforcement learning process to adapt to the state changes of the integrated energy system, while balancing short-term rapid response and long-term stable optimization. This enables computer equipment to efficiently manage renewable energy output, energy storage system charging and discharging, and grid power flow control, improving scheduling real-time performance, enhancing strategy adaptability, and reducing computing costs.

[0101] In one embodiment, the process of generating minute-level instructions corresponding to slow variable data through incremental model predictive control in step S130 may include:

[0102] S131: Construct the objective function and embed the device physical equations as constraints to form an incremental model.

[0103] S132: Input the slow variable data into the incremental model, so that the incremental model performs incremental adjustment based on the optimization result of the previous cycle of the slow variable data, and outputs the minute-level instructions of the current cycle.

[0104] In this embodiment, after initially acquiring slow variable data, the computer device can construct an objective function and embed the device's physical equations as constraints to form an incremental model. The slow variable data is then input into the incremental model, causing the incremental model to make incremental adjustments based on the optimization results of the previous cycle of the slow variable data, and output minute-level instructions for the current cycle. Compared to traditional model predictive control, this application can significantly reduce the amount of computation and adapt to minute-level updates.

[0105] In a specific implementation, taking the dual-objective optimization function of economy and carbon emissions as an example, the objective function can be constructed as follows:

[0106]

[0107] Where, represents the objective function; represents the increment of the control variable; represents the economic cost at time t; represents the carbon emission cost at time t.

[0108] Then, operating constraints can be established. Here, the application can embed the device physical equations (efficiency model) as constraints to ensure that the optimization results reflect the actual operating characteristics of the load device. Taking the energy storage charge and discharge efficiency constraint as an example, the operating constraint can be expressed as follows:

[0109]

[0110] Where, Indicates the efficiency of the energy storage system; Indicates rated power; Indicates the current attenuation coefficient; Represents the current of the energy storage system.

[0111] Taking the gas turbine efficiency constraint as an example, the operation constraint can be expressed as follows:

[0112]

[0113] Where, represents the efficiency of the gas turbine; Indicates load rate; and Respectively represent the minimum and maximum load rates allowed for the unit; 、 and Both represent fitting coefficients.

[0114] Furthermore, the process of incremental adjustment of the incremental model based on the optimization results of the previous cycle of slow variable data can be expressed as follows:

[0115]

[0116] Where, represents the increment of the control variable; Represents the Hessian matrix, that is, the second-order derivative of the objective function, which is used to reflect the curvature; Represents the gradient vector, that is, the first-order derivative of the objective function, which is used to reflect the slope direction.

[0117] Finally, the incremental model can output minute-level instructions for slow variable objects. , such as the load rate adjustment of a new energy unit, the carbon emission quota adjustment, etc.

[0118] In one embodiment, Figure 2 As shown, Figure 2 A flowchart of an instruction security verification process provided in an embodiment of the present application; the process of performing security verification on the target instruction based on the KKT condition constraint in step S140 to obtain the optimized scheduling instruction may include:

[0119] S141: Use KKT conditional constraints to perform security verification on the target instruction and obtain the verification result.

[0120] S142: When the verification result is failure, a correction algorithm is used to correct the target instruction to obtain an optimized scheduling instruction.

[0121] S143: When the verification result is passed, the target instruction is used as the final optimized scheduling instruction.

[0122] In this embodiment, the computer device can use KKT conditional constraints to perform a safety check on the target instruction. If the check fails, it means that the target instruction exceeds the device's tolerance range. Therefore, the computer device needs to use a correction algorithm to correct the target instruction and obtain an optimized scheduling instruction. If the check passes, it means that the instruction meets the device's tolerance range and can be directly used as the final optimized scheduling instruction, thereby ensuring the efficient and safe operation of the integrated energy system on a time scale of seconds or minutes.

[0123] Specifically, computer equipment uses safety constraints to perform real-time verification on generated instructions and retracts out-of-limit instructions along the constraint gradient to ensure power grid safety. The expression of the KKT condition constraint can be shown as follows:

[0124]

[0125] Where, represents the objective function; represents an inequality constraint; represents the Lagrange multiplier.

[0126] In one embodiment, the calculation expression of the correction algorithm in step S142 may include:

[0127]

[0128] Where, Indicates optimized scheduling instructions; Represents the original target instruction; Indicates correction compensation; Represents the orthographic projection operator.

[0129] In this embodiment, by introducing a correction compensation term and combining it with a forward projection operator, the correction algorithm can optimize second- or minute-level instructions while ensuring that the instructions meet the equipment's operating constraints. This allows the original target instruction to be optimized to the greatest extent possible while satisfying the KKT conditional constraints. This prevents scheduling instructions from exceeding the equipment's capacity, reduces energy scheduling deviations caused by instruction failure, and improves system stability. Furthermore, this correction algorithm can reduce the number of scheduling calculation iterations, accelerate optimization convergence, and ensure that the energy system can still achieve real-time scheduling in high-frequency dynamic environments, ultimately improving the economic efficiency, reliability, and control accuracy of the integrated energy system.

[0130] The following describes the comprehensive energy optimization scheduling device provided in the embodiment of the present application. The comprehensive energy optimization scheduling device described below and the comprehensive energy optimization scheduling method described above can be referenced to each other.

[0131] In one embodiment, Figure 3 As shown, Figure 3 This is a schematic diagram of the structure of a comprehensive energy optimization and scheduling device provided in an embodiment of the present application. The present application also provides a comprehensive energy optimization and scheduling device, including a data acquisition module 210, a fast variable prediction module 220, a slow variable prediction module 230, and an instruction optimization module 240, specifically including the following:

[0132] The data acquisition module 210 is used to collect multi-source heterogeneous energy data in the integrated energy system in real time; the multi-source heterogeneous energy data includes fast variable data and slow variable data.

[0133] The fast variable prediction module 220 is used to determine the state variable fluctuation of the integrated energy system based on the fast variable data, and when the state variable fluctuation exceeds a preset threshold, the event triggers the reinforcement learning model to generate second-level instructions corresponding to the fast variable data.

[0134] The slow variable prediction module 230 is used to generate minute-level instructions corresponding to the slow variable data through incremental model predictive control; the incremental model predictive control embeds the device physical equations as constraints.

[0135] The instruction optimization module 240 is used to perform security verification on the target instructions based on KKT condition constraints after generating the target instructions, obtain the optimized scheduling instructions, and use the optimized scheduling instructions to perform energy optimization scheduling on the integrated energy system; wherein, the target instructions include second-level instructions and minute-level instructions.

[0136] In the above embodiment, during energy scheduling, multi-source heterogeneous energy data in the integrated energy system can be collected in real time; the multi-source heterogeneous energy data here includes fast variable data and slow variable data, thereby realizing hierarchical decoupling control and avoiding multi-variable coupling conflicts; wherein, for fast variable data, the state variable fluctuation of the integrated energy system can be determined based on the fast variable data, and then when the state variable fluctuation exceeds the preset threshold, the event-triggered reinforcement learning model is used to generate second-level instructions corresponding to the fast variable data, thereby achieving millisecond-level dynamic response and improving the real-time performance of energy scheduling; for slow variable data, minute-level instructions corresponding to the slow variable data can be generated through incremental model predictive control. Here, the incremental model predictive control embeds the device physical equation as a constraint condition, thereby ensuring that the optimization result reflects the actual operating characteristics of the load equipment and balancing calculation accuracy and real-time response capability. In addition, the generated second-level instructions or minute-level instructions can be security-verified based on the KKT condition constraint to obtain optimized scheduling instructions to further improve the accuracy of the instructions. Finally, the optimized scheduling instructions can be used to perform energy optimization scheduling on the integrated energy system, achieving millisecond-level real-time scheduling and multi-objective collaborative optimization.

[0137] In one embodiment, the fast variable prediction module 220 may include:

[0138] The variable determination submodule is used to determine the state variables at the previous moment and the state variables at the current moment based on the fast variable data.

[0139] The difference calculation submodule is used to subtract the state variable at the previous moment from the state variable at the current moment to obtain the state variable fluctuation of the integrated energy system at the current moment.

[0140] In one embodiment, the fast variable prediction module 220 may further include:

[0141] The model prediction submodule is used to input fast variable data into the event-triggered reinforcement learning model, so that the event-triggered reinforcement learning model uses a double-buffered experience pool to update the control strategy and output second-level instructions.

[0142] In one embodiment, the model prediction submodule may include:

[0143]

[0144]

[0145] Where, represents the policy parameters; represents the learning rate; represents the action value function, which is used to predict the long-term benefits of performing action a in state s; Indicates policy parameters gradient; represents the sample set collected from the double-buffered experience pool; Indicates the real-time cache sampling ratio; Indicates real-time cache; Indicates long-term caching.

[0146] In one embodiment, the slow variable prediction module 230 may include:

[0147] The model building submodule is used to build the objective function and embed the device physical equations as constraints to form an incremental model.

[0148] The incremental adjustment submodule is used to input slow variable data into the incremental model so that the incremental model performs incremental adjustment based on the optimization results of the previous cycle of the slow variable data, and outputs the minute-level instructions for the current cycle.

[0149] In one embodiment, the instruction optimization module 240 may include:

[0150] The safety check submodule is used to perform safety check on the target instruction using KKT condition constraints to obtain the check result.

[0151] The instruction correction submodule is used to use a correction algorithm to correct the target instruction when the verification result is failed, so as to obtain an optimized scheduling instruction.

[0152] The instruction determination submodule is used to use the target instruction as the final optimized scheduling instruction when the verification result is passed.

[0153] In one embodiment, the calculation expression of the correction algorithm in the instruction correction submodule may include:

[0154]

[0155] Where, Indicates optimized scheduling instructions; Indicates the original second-level instruction or minute-level instruction; Indicates correction compensation; Represents the orthographic projection operator.

[0156] In one embodiment, the present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the comprehensive energy optimization scheduling method as described in any of the above embodiments.

[0157] In one embodiment, the present application also provides a computer device having computer-readable instructions stored therein. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the comprehensive energy optimization scheduling method as described in any one of the above embodiments.

[0158] Schematically, as Figure 4 As shown, Figure 4 This is a schematic diagram of the internal structure of a computer device provided in an embodiment of the present application. The computer device 300 can be provided as a server. Figure 4 Computer device 300 includes a processing component 302, which further includes one or more processors, and memory resources represented by memory 301 for storing instructions executable by processing component 302, such as application programs. The application programs stored in memory 301 may include one or more modules, each corresponding to a set of instructions. Furthermore, processing component 302 is configured to execute the instructions to perform the integrated energy optimization scheduling method according to any of the above-described embodiments.

[0159] The computer device 300 may further include a power supply component 303 configured to perform power management of the computer device 300, a wired or wireless network interface 304 configured to connect the computer device 300 to a network, and an input / output (I / O) interface 305. The computer device 300 may operate based on an operating system stored in the memory 301, such as Windows Server™, Mac OS X™, Unix™, Linux™, Free BSD™, or the like.

[0160] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0161] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0162] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referenced to each other.

[0163] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A comprehensive energy optimization scheduling method, characterized in that: The method comprises: Real-time collection of multi-source heterogeneous energy data in an integrated energy system; the multi-source heterogeneous energy data includes fast variable data and slow variable data; Determine the state variable fluctuation of the integrated energy system according to the fast variable data, and when the state variable fluctuation exceeds a preset threshold, generate a second-level instruction corresponding to the fast variable data through an event-triggered reinforcement learning model; Generating minute-level instructions corresponding to the slow variable data through incremental model predictive control; the incremental model predictive control embeds device physical equations as constraints; After the target instruction is generated, the target instruction is security-checked based on the KKT condition constraint to obtain the optimized scheduling instruction, and the optimized scheduling instruction is used to perform energy optimization scheduling on the integrated energy system; wherein, the target instruction includes second-level instructions and minute-level instructions.

2. The comprehensive energy optimization scheduling method according to claim 1 is characterized in that: Determining the state variable fluctuation amount of the integrated energy system according to the fast variable data includes: Determine the state variable at the previous moment and the state variable at the current moment according to the fast variable data; The state variable at the current moment is subtracted from the state variable at the previous moment to obtain the state variable fluctuation amount of the integrated energy system at the current moment.

3. The comprehensive energy optimization scheduling method according to claim 1, characterized in that: The event-triggered reinforcement learning model generates second-level instructions corresponding to the fast variable data, including: The fast variable data is input into the event-triggered reinforcement learning model, so that the event-triggered reinforcement learning model uses a double-buffered experience pool to update the control strategy and outputs a second-level instruction.

4. The comprehensive energy optimization scheduling method according to claim 3 is characterized in that: The control strategy update formula of the event-triggered reinforcement learning model includes: Where, represents the policy parameters; represents the learning rate; represents the action value function, which is used to predict the long-term benefits of performing action a in state s; Indicates policy parameters gradient; represents the sample set collected from the double-buffered experience pool; Indicates the real-time cache sampling ratio; Indicates real-time cache; Indicates long-term caching.

5. The comprehensive energy optimization scheduling method according to claim 1 is characterized in that: Generating minute-level instructions corresponding to the slow variable data through incremental model predictive control includes: Construct an objective function and embed the device physics equations as constraints to form an incremental model; The slow variable data is input into the incremental model, so that the incremental model performs incremental adjustment based on the optimization result of the previous cycle of the slow variable data, and outputs the minute-level instructions of the current cycle.

6. The comprehensive energy optimization scheduling method according to claim 1, characterized in that: The security verification process of the target instruction based on the KKT condition constraint to obtain the optimized scheduling instruction includes: Performing a security check on the target instruction using KKT conditional constraints to obtain a check result; When the verification result is failed, a correction algorithm is used to correct the target instruction to obtain an optimized scheduling instruction; When the verification result is passed, the target instruction is used as the final optimized scheduling instruction.

7. The comprehensive energy optimization scheduling method according to claim 6, characterized in that: The calculation expression of the correction algorithm includes: Where, Indicates optimized scheduling instructions; Indicates the original second-level instruction or minute-level instruction; Indicates correction compensation; Represents the orthographic projection operator.

8. A comprehensive energy optimization scheduling device, characterized in that: include: A data acquisition module is used to collect multi-source heterogeneous energy data in the integrated energy system in real time; the multi-source heterogeneous energy data includes fast variable data and slow variable data; a fast variable prediction module, configured to determine the state variable fluctuation of the integrated energy system based on the fast variable data, and, when the state variable fluctuation exceeds a preset threshold, generate a second-level instruction corresponding to the fast variable data through an event-triggered reinforcement learning model; A slow variable prediction module, configured to generate minute-level instructions corresponding to the slow variable data through incremental model predictive control; the incremental model predictive control embeds device physical equations as constraints; The instruction optimization module is used to perform security verification on the target instruction based on the KKT condition constraint after the target instruction is generated, obtain the optimized scheduling instruction, and use the optimized scheduling instruction to perform energy optimization scheduling on the integrated energy system; wherein, the target instruction includes second-level instructions and minute-level instructions.

9. A storage medium, characterized in that: The storage medium stores computer-readable instructions, which, when executed by one or more processors, enable the one or more processors to execute the steps of the integrated energy optimization scheduling method according to any one of claims 1 to 7.

10. A computer device, characterized in that: include: one or more processors, and memory; The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the integrated energy optimization scheduling method according to any one of claims 1 to 7 are executed.