Comprehensive energy system low-carbon optimization scheduling method based on deep reinforcement learning
By constructing a comprehensive energy system model and improving the diffusion neural network through deep reinforcement learning, multi-scenario source-load uncertainty data is generated. By using SAC agents for transfer learning, the shortcomings of action optimization in the low-carbon optimization scheduling of the comprehensive energy system are solved, and an efficient and robust low-carbon scheduling strategy is realized to reduce energy consumption and carbon emissions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-03-27
AI Technical Summary
In the low-carbon optimization scheduling of integrated energy systems, existing technologies lack continuity and real-time performance in action optimization, resulting in insufficient adaptability of the system to environmental changes, limited optimization effect, and insufficient precision in scheduling strategy optimization.
A comprehensive energy system model is constructed using a deep reinforcement learning approach. An improved diffusion neural network model is built using multi-dimensional operational datasets and historical load data of various types to generate source-load uncertainty data for multiple scenarios. The SAC agent is then used for transfer learning to output the optimal scheduling strategy, which is then combined with energy storage charging and discharging and renewable energy output for coordinated optimization.
Significantly reduce overall energy consumption and energy purchase costs, reduce the use of fossil fuels and carbon emissions, improve the optimization accuracy and robustness of dispatch strategies, reduce the risk of dispatch mismatch caused by forecast deviations, meet energy supply demand and security constraints, and achieve a balance between economic efficiency and low carbon emissions.
Smart Images

Figure CN121745696A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of energy system operation and optimization, and particularly relates to a low-carbon optimization scheduling method for a comprehensive energy system based on deep reinforcement learning. BACKGROUND
[0002] Under the dual urgent requirements of global energy crisis and climate deterioration, it is an urgent strategic task to establish a clean and efficient energy paradigm. Distributed renewable energy represented by wind turbines and photovoltaics is rapidly expanding, providing a key path for sustainable energy transformation. In this context, the comprehensive energy system deeply couples power, heating and cooling and other multi-energy flows and plays a dynamic complementary role, fundamentally reshaping the traditional energy operation mode and showing unique system-level solution advantages. However, the inherent intermittency and volatility of renewable energy, combined with the random characteristics of load demand, bring unprecedented challenges to the safe, reliable and efficient operation of the comprehensive energy system.
[0003] A low-carbon optimization scheduling method for a comprehensive energy system based on action adjustment reinforcement learning is disclosed in Chinese Patent Publication No. CN120046784A. The method first establishes a carbon emission flow model for the comprehensive energy system based on the EH coupling model, with the minimum system operation cost and carbon trading cost as the target. Reinforcement learning is designed to use the known state quantity of the system at time t as the state space, and the output of each device at time t as the action space. The reward function includes cost rewards and penalties for violating system constraints. In the early exploration stage, the strategy network is used to output actions for the current system state, and the network parameters are updated. When the training test reaches the preset threshold, enter the action space adjustment stage, adjust the action output by the strategy network at time t according to the constraint conditions, and introduce the regularization term of the action offset into the loss function of the strategy network. The updated network is used to optimize the low-carbon scheduling of the comprehensive energy system, and the output of each device is output. However, the above-mentioned comparative document only adjusts the action space when the preset threshold is reached in the action optimization, and the adjustment strategy lacks continuity and real-time performance, the system's ability to adapt to environmental changes is insufficient and the optimization effect is limited, resulting in insufficient optimization precision of the scheduling strategy. Therefore, it is very necessary to provide a low-carbon optimization scheduling method for a comprehensive energy system based on deep reinforcement learning to improve the optimization precision of the scheduling strategy. SUMMARY
[0004] Therefore, the present application provides a low-carbon optimization scheduling method for a comprehensive energy system based on deep reinforcement learning.
[0005] The present application provides a low-carbon optimization scheduling method for a comprehensive energy system based on deep reinforcement learning, which comprises: Collecting a multi-dimensional operation data set of a monitoring object group, and constructing a comprehensive energy system model based on the multi-dimensional operation data set; minimize the total operation cost of the system corresponding to the integrated energy system model as an objective, build a deep reinforcement learning scheduling model, and obtain an initial scheduling strategy based on the operation constraints of the integrated energy system model and the deep reinforcement learning scheduling model; Based on the multi-dimensional operation data set and multi-class load historical data, an improved diffusion neural network model is constructed, and the improved diffusion neural network model generates multi-scenario source-load uncertainty data; According to the initial scheduling strategy and the multi-scenario source-load uncertainty data, an SAC agent is constructed, a pre-training strategy is used to pre-train the SAC agent, the pre-trained SAC agent is deployed in the integrated energy system model, and an optimal scheduling strategy is output to the monitoring object group.
[0006] On the basis of the above technical solutions, preferably, the integrated energy system model is constructed based on the multi-dimensional operation data set, specifically including: The monitoring object group is determined, and a multi-dimensional operation data set generated during the operation of the monitoring object group is collected; The multi-dimensional operation data set is pre-processed to form a standardized time series input matrix, and an integrated energy system model is constructed based on the standardized time series input matrix.
[0007] On the basis of the above technical solutions, preferably, the monitoring object group includes a combined heat and power unit, an electric boiler, a gas boiler, an electric refrigerator, an absorption refrigerator, a battery energy storage system, a photovoltaic power generation device, a wind power generation device, an electric load, a heat load, a cold load, and an external power grid and gas network interface. The multi-dimensional operation data set includes the input power of each energy device, the output power of each energy device, the energy conversion efficiency, the charge and discharge power and state of charge of the energy storage system, the output of the renewable energy, the multi-class load demand, the external environment temperature, the energy price signal, the carbon emission price signal and the green certificate trading price signal.
[0008] Further preferably, the total operation cost includes carbon emission cost, green certificate trading cost, battery energy storage system degradation cost, natural gas procurement cost, and main grid power procurement cost, wherein, Based on the carbon emission quota of the integrated energy system model, the allocation quota of the combined heat and power unit and the gas boiler, the power generation reference quota of the combined heat and power unit and the gas boiler, and the heat supply reference quota of the combined heat and power unit and the gas boiler, the carbon emission cost is calculated using a step-type carbon trading price function; Based on the number of certificates required for compliance, the green certificate quota coefficient allocated by the regulatory agency, and the number of green certificates obtained by the integrated energy system model, the green certificate trading cost is calculated; a battery energy storage model is constructed according to the cycle life of the battery energy storage system in the integrated energy system model and the degradation cost of the battery energy storage system, and the degradation cost of the battery energy storage system is converted into a unit energy throughput from a capital investment cost in the battery energy storage model to obtain the degradation cost of the battery energy storage system; Based on the purchase and sale price settlement of the main grid power, the cost of purchasing natural gas from the gas pipeline network, and the supply amount of natural gas in the integrated energy system model, the main grid power purchase cost is calculated.
[0009] Further preferably, the step of constructing an improved diffusion neural network model based on the multi-dimensional operation data set and multi-class load historical data comprises: The historical source load time series data in the multi-dimensional operation data set is converted into a standard Gaussian distribution by gradually adding Gaussian noise to construct a forward diffusion process, and the standard Gaussian noise is gradually restored to form multi-scenario source load data to construct a backward diffusion process; The improved diffusion neural network model is trained based on the forward diffusion process and the reverse denoising process, and the improved diffusion neural network model generates multi-scenario source load uncertainty data.
[0010] Further preferably, the pre-training of the SAC agent with a transfer learning strategy specifically comprises: The multi-scenario source load uncertainty data generated from the improved diffusion neural network model is classified according to the characteristics of the scenarios to construct a typical scenario library; The typical scenarios are divided into different training environments according to complexity and uncertainty, and the source load time series data, state transition function and reward calculation mechanism of the corresponding scenarios are encapsulated in each training environment to pre-train the SAC agent.
[0011] Further preferably, the operation constraints include combined heat and power equipment constraints, electric boiler equipment constraints, gas boiler equipment constraints, electric refrigeration equipment constraints, absorption refrigeration equipment constraints, battery energy storage system constraints, and main grid interaction constraints.
[0012] In a second aspect of the present application, an integrated energy system low-carbon optimization scheduling system based on deep reinforcement learning is provided, which comprises a data acquisition module, a model construction module and a policy optimization module, wherein, The data acquisition module is used to acquire a multi-dimensional operation data set of a monitoring object group, and to construct an integrated energy system model based on the multi-dimensional operation data set; The model construction module is configured to construct a deep reinforcement learning scheduling model with the minimum total operation cost of the integrated energy system model as the target, and obtain an initial scheduling strategy based on the operation constraint condition of the integrated energy system model and the deep reinforcement learning scheduling model, construct an improved diffusion neural network model based on the multi-dimensional operation data set and multi-class load historical data, and enable the improved diffusion neural network model to generate multi-scenario source-load uncertainty data. The strategy optimization module is configured to construct an SAC agent according to the initial scheduling strategy and the multi-scenario source-load uncertainty data, pre-train the SAC agent with a transfer learning strategy, deploy the pre-trained SAC agent to the integrated energy system model, and output an optimal scheduling strategy to the monitoring object group.
[0013] In a third aspect of the present application, an electronic device is provided, which includes a processor, a memory, a user interface and a network interface, the memory is configured to store instructions, the user interface and the network interface are configured to communicate with other devices, and the processor is configured to execute the instructions stored in the memory.
[0014] In a fourth aspect of the present application, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps of the deep reinforcement learning-based integrated energy system low-carbon optimization scheduling method.
[0015] The deep reinforcement learning-based integrated energy system low-carbon optimization scheduling method provided by the present application has the following beneficial effects compared with the prior art: (1) With the minimum total operation cost as the target, the energy storage charging and discharging, the electric-thermal-gas coupled device and the renewable energy output are optimized in combination with reinforcement learning, the start-stop loss and the curtailment of wind and light are reduced, the integrated energy consumption and the energy purchase cost are significantly reduced, and through the preferential calling and flexible scheduling of low-carbon resources, the use of fossil energy and the carbon emission intensity are reduced on the premise of meeting the energy supply demand and safety constraints, the unity of economy and low carbon is realized, the improved diffusion neural network generates multi-scenario source-load uncertainty data, the training covers extreme and rare working conditions, the robustness and generalization ability of the strategy under the fluctuation of renewable output and the mutation of load are improved, the scheduling mismatch risk caused by prediction deviation is reduced, the optimization accuracy of the scheduling strategy is effectively improved, the operation constraints of the integrated energy system are explicitly introduced in the reinforcement learning strategy search process, the obtained strategy naturally meets the engineering feasibility and safety boundary, and the out-of-bound operation is avoided.
[0016] (2) By forward diffusion, the complex historical source load distribution is mapped to the standard Gaussian, and the full distribution is learned by reverse denoising, which can retain the non-Gaussian characteristics and time correlation and cross-variable correlation, avoid the deviation caused by only point prediction or simple noise model, sample from high noise state and gradually denoise, can synthesize tail and extreme scenarios, improve the coverage of rare events such as wind and light sudden drop and peak load, help to configure backup and safety margin, reduce the risk of loss of load and wind and light abandonment, by adjusting the diffusion step, noise intensity or conditional information, diversified scenarios of different severity and different time scale can be generated, reduce the mode collapse phenomenon, and be more stable and reliable than traditional generation methods. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0018] Figure 1 The flowchart of the low-carbon optimal scheduling method of the comprehensive energy system based on deep reinforcement learning provided by the present application is shown. Figure 2 The architecture diagram of the comprehensive energy system model provided by the present application is shown. Figure 3 The architecture diagram of the improved diffusion neural network model provided by the present application is shown. Figure 4 The optimization control architecture diagram of the comprehensive energy system model provided by the present application is shown. Figure 5 The structure diagram of the low-carbon optimal scheduling system of the comprehensive energy system provided by the present application is shown. Figure 6 The structure diagram of the electronic device provided by the present application is shown.
[0019] Explanation of reference numerals: 1, low-carbon optimal scheduling system of comprehensive energy system; 11, data acquisition module; 12, model construction module; 13, strategy optimization module; 2, electronic device; 21, processor; 22, communication bus; 23, user interface; 24, network interface; 25, memory. DETAILED DESCRIPTION
[0020] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0021] This invention discloses a low-carbon optimization scheduling method for integrated energy systems based on deep reinforcement learning, referencing... Figure 1 The steps of this method include S1 to S4.
[0022] Step S1: Collect a multi-dimensional operational dataset of the monitored object group and construct a comprehensive energy system model based on the multi-dimensional operational dataset.
[0023] This step also includes steps S11 to S12.
[0024] Step S11: Determine the monitoring target group and collect the multidimensional operation dataset generated during the operation of the monitoring target group. The monitoring target group includes cogeneration units, electric boilers, gas boilers, electric chillers, absorption chillers, battery energy storage systems, photovoltaic power generation devices, wind power generation devices, power load, heat load, cold load, and the interface between the external power grid and the gas pipeline network. The multidimensional operation dataset includes the input power of each energy device, the output power of each energy device, the energy conversion efficiency, the charging and discharging power and state of charge of the energy storage system, the output of renewable energy, various load demands, external ambient temperature, energy price signals, carbon emission price signals, and green certificate trading price signals.
[0025] This step first clarifies the physical and functional boundaries of the integrated energy system, identifying all energy equipment and load types within the system. Based on energy type and function, the equipment is categorized into power generation equipment, heating equipment, cooling equipment, energy storage equipment, loads, and external interfaces. A detailed inventory of each type of equipment is conducted, recording basic information such as equipment model, rated parameters, installation location, and operating status. The energy flow, material flow, and information flow networks are analyzed to identify key nodes and equipment in the system, including all 12 listed monitoring categories. Based on the system topology and operating characteristics, the specific measurement point locations for each monitoring object are determined to ensure data representativeness and completeness.
[0026] Please see Figure 2 , Figure 2The schematic diagram of the integrated energy system architecture includes an energy input end, an energy conversion device, a battery energy storage system, a load end, and a market mechanism. The energy input end includes a main grid, a fan, a photovoltaic device, and a pipe network. The energy conversion device includes an electric boiler, a battery energy storage system, a heat pump cogeneration device, an absorption refrigeration device, and a gas boiler. The load end includes an electrical load, a cold load, and a heat load. The market mechanism includes a carbon market that supports the purchase and sale of carbon quotas and a green certificate market that supports the purchase and sale of green power certificates. Figure 2 The various energy flows are represented by different colored lines. Blue lines represent the flow of electrical energy, red lines represent the flow of thermal energy, gray lines represent the flow of natural gas, and light blue lines represent the flow of cold energy.
[0027] The multi-dimensional operation data set collection method can collect temperature and other external condition data by installing power meters, flow meters, thermometers, and other sensing devices at the input and output ends of various energy devices, or installing state-of-charge monitoring devices in energy storage systems. Existing metering devices such as electricity meters, heat meters, and flow meters are connected. The electricity market system is connected to obtain electricity price signals, the gas price information system is connected to obtain gas prices, and the carbon trading platform and green certificate trading system are connected to obtain relevant price signals.
[0028] In step S12, the multi-dimensional operation data set is preprocessed to form a standardized time series input matrix, and based on the standardized time series input matrix, an integrated energy system model is constructed.
[0029] In this step, the key link of converting the original multi-dimensional operation data into a standardized model for analysis is to first clean the collected multi-dimensional data, then align and resample the time series to unify the data time scale, then normalize the data and unify the units to achieve data standardization, and finally create a feature engineering by combining time feature extraction and derived variable creation to build a standardized time series input matrix with time points as rows and device parameters as columns. Based on this matrix, the integrated energy system model is constructed by dividing the system into power, heat, cold, and energy storage subsystems, selecting appropriate model types for different parts, applying machine learning or deep learning algorithms to establish data relationships and integrate physical constraints, ensuring accuracy through parameter calibration and model verification, and finally integrating the subsystem models and performing global optimization to form an integrated model that can accurately represent the dynamic characteristics of the energy system.
[0030] In step S2, a deep reinforcement learning scheduling model is constructed with the goal of minimizing the total operating cost of the integrated energy system model corresponding to the system, and based on the operating constraints of the integrated energy system model and the deep reinforcement learning scheduling model, an initial scheduling strategy is obtained.
[0031] In this step, the total system operation cost includes carbon emission cost, green certificate transaction cost, battery energy storage system degradation cost, natural gas procurement cost and main grid power procurement cost, wherein, Based on the carbon emission quota of the integrated energy system model, the allocation quota of the combined heat and power and gas boiler, the power generation reference quota of the combined heat and power and gas boiler, and the heat supply reference quota of the combined heat and power and gas boiler, the carbon emission cost is calculated by using the step-type carbon trading price function; Based on the number of green certificates required for compliance, the green certificate quota coefficient allocated by the regulatory agency, and the number of green certificates obtained by the integrated energy system model, the green certificate transaction cost is calculated; According to the cycle life of the battery energy storage system in the integrated energy system model and the degradation cost of the battery energy storage system, a battery energy storage model is constructed, and in the battery energy storage model, the capital investment cost is converted into the degradation cost per unit of energy throughput to obtain the degradation cost of the battery energy storage system; Based on the purchase and sale electricity price settlement of the main grid power, the cost of purchasing natural gas from the gas pipeline network, and the supply amount of natural gas from the integrated energy system model, the main grid power procurement cost is calculated.
[0032] In one example, the objective function of the deep reinforcement learning model is constructed to minimize the total system operation cost, and the specific expression is:
[0033] Wherein, represents the total system operation cost, represents the carbon emission cost, represents the green certificate transaction cost, represents the battery energy storage system degradation cost, represents the natural gas procurement cost, represents the main grid power procurement cost.
[0034] The carbon emission cost can be represented as:
[0035]
[0036] Wherein, represents the carbon emission quota allocated to the integrated energy system, and respectively represent the carbon emission quota allocated to the combined heat and power equipment and the gas boiler, and respectively represent the reference carbon emission quota for power generation and heat supply of the combined heat and power equipment, represents the reference carbon emission quota for heat supply of the gas boiler, and respectively represent the power generation and heat supply of the combined heat and power equipment, represents the heat supply of the gas boiler, represents the carbon trading cost of the integrated energy system, represents the step-type carbon trading price, represents the traded carbon quota amount, Q out represents the actual carbon emission amount of the integrated energy system, d represents the step of the step-type carbon trading price, c represents the carbon trading benchmark price, k 1 represents the carbon trading incremental coefficient, k 2 represents the carbon trading compensation coefficient.
[0037] The green certificate trading cost can be represented as:
[0038]
[0039]
[0040] wherein, G q represents the number of green certificates required to be held for compliance, β represents the green certificate quota coefficient allocated by the regulatory agency, represents the system load power at the t moment, G r represents the number of green certificates obtained by the integrated energy system, represents the photovoltaic system power generation at the t moment, represents the wind turbine power generation at the t moment, C GCT represents the green certificate trading cost, G GCT represents the green certificate quota deviation, Lambda GCT represents the green certificate trading price.
[0041] Battery energy storage system cost:
[0042]
[0043]
[0044] wherein, T cycle represents the cycle life of the battery energy storage system, represents the degradation cost of the battery energy storage system at the t moment, the battery energy storage system at the t-th time point, t the power of the battery energy storage system at the t-th time point, C E the capital investment cost of the battery energy storage system, SOC t the battery state of charge at the t-th time point, t DOD the depth of discharge,
[0045] gas and electricity purchase cost:
[0046]
[0047] wherein, the amount of natural gas supplied to the system at the t-th time point, t the unit price of natural gas at the t-th time point, T represents the total number of scheduling periods, t the electricity purchase price of the integrated energy system at the t-th time point, when t the income obtained by the integrated energy system by selling electricity to the main grid, when the expenditure of the integrated energy system for purchasing electricity from the main grid, Δ T the unit time period length, m
[0048] Further, the operation constraints include combined heat and power equipment constraints, electric boiler equipment constraints, gas boiler equipment constraints, electric refrigeration equipment constraints, absorption refrigeration equipment constraints, battery energy storage system constraints, and main grid interaction constraints.
[0049] The combined heat and power equipment constraints are represented as:
[0050] wherein, the power output of the combined heat and power equipment at the t-th time point, t the heat output of the combined heat and power equipment at the t-th time point, t the natural gas consumption at the t-th time point, t the low calorific value of natural gas; Epsilon CHP the heat-to-power ratio of the combined heat and power equipment, Pmax, cogen(t) represents the upper limit of the electric power output of the cogeneration plant, Pmin, cogen(t) represents the lower limit of the thermal power output of the cogeneration plant, Pmax, cogen(t) represents the upper limit of the thermal power output of the cogeneration plant, Pdown, cogen(t) represents the ramp-down constraint of the cogeneration plant, Pup, cogen(t) represents the ramp-up constraint of the cogeneration plant.
[0051] The electric boiler plant constraints are represented as:
[0052] wherein, Qb(t) represents the thermal output of the electric boiler at time t; t Pb(t) represents the electric power consumption of the electric boiler at time t, t EB ηb represents the energy conversion efficiency of the electric boiler; Epsilon Pmax, b(t) represents the upper limit of the thermal output of the electric boiler, Pmin, b(t) represents the lower limit of the thermal output of the electric boiler. The gas boiler plant constraints are represented as:
[0053]
[0054] wherein, Qg(t) represents the thermal output of the gas boiler at time t, t Pgas(t) represents the natural gas consumption power of the gas boiler at time t, t GB ηg represents the energy conversion efficiency of the gas boiler, Eta Pmax, g(t) represents the upper limit of the thermal output of the gas boiler, Pmin, g(t) represents the lower limit of the thermal output of the gas boiler. The electric refrigeration plant constraints are represented as:
[0055]
[0056] wherein, Qr(t) represents the refrigeration output of the electric refrigerator at time t, t Pfr(t) represents the electric power consumption of the electric refrigerator at time t, t EC ηr represents the energy conversion efficiency of the electric refrigerator, Eta Pmax, r(t) represents the upper limit of the refrigeration output of the electric refrigerator, Pmin, r(t) represents the lower limit of the refrigeration output of the electric refrigerator. The absorption refrigeration plant constraints are represented as:
[0057]
[0058] where, represents the cooling output of the absorption refrigeration device at the t time, represents the heat input of the absorption refrigeration device at the t time, Eta AC represents the energy conversion efficiency of the absorption refrigeration device, represents the upper limit of the cooling output of the absorption refrigeration device, represents the lower limit of the cooling output of the absorption refrigeration device.
[0059] The battery energy storage system constraints are represented as:
[0060]
[0061] where, represents the battery energy storage system capacity at the t time, represents the charging efficiency of the battery energy storage system at the t time, represents the discharging efficiency of the battery energy storage system at the t time, represents the upper limit of the battery energy storage system capacity, represents the lower limit of the battery energy storage system capacity, SOC init represents the initial battery state of charge, SOC end represents the final battery state of charge, Eta BESS,ch represents the initial state of charge of the battery energy storage system, Eta BESS,dis represents the final state of charge of the battery energy storage system, represents the state of charge of the battery energy storage system at the t time, represents the state of discharge of the battery energy storage system at the t time, represents the upper limit of the battery energy storage system charging power at the t time, represents the upper limit of the battery energy storage system discharging power at the t time.
[0062] The main grid interaction constraints are represented as:
[0063] where, represents the maximum power for the main grid power trade. represents the minimum power of the main grid power transaction.
[0064] The power balance constraint conditions that the integrated energy system optimization scheduling model needs to meet are:
[0065]
[0066]
[0067]
[0068]
[0069] wherein, represents the electrical load at the t time, represents the thermal load at the t time, represents the cold load at the t time, represents the natural gas consumption of the gas boiler at the t time, represents the natural gas consumption of the combined heat and power equipment at the t time, represents the amount of natural gas purchased from the gas pipeline network, represents the maximum gas purchase power of the integrated energy system model, represents the minimum gas purchase power of the integrated energy system model.
[0070] Step S3, based on the multi-dimensional operation data set and the multi-class load historical data, an improved diffusion neural network model is constructed, and the improved diffusion neural network model generates multi-scenario source-load uncertainty data.
[0071] In this step, steps S31-S32 are also included.
[0072] Step S31, the historical source-load time series data in the multi-dimensional operation data set is converted into a standard Gaussian distribution by gradually adding Gaussian noise, to construct a forward diffusion process, and the standard Gaussian noise is gradually restored to form multi-scenario source-load data, to construct a backward diffusion process.
[0073] In this step, please refer to Figure 3 , a forward diffusion process is constructed, and noise is gradually added to the historical renewable energy and load data to establish a probability distribution:
[0074]
[0075]
[0076]
[0077]
[0078] Where T represents the total number of diffusion steps, N Indicates a standard Gaussian distribution; β t Indicates the first t A predefined constant for the step, with a value range of (0,1), is used to control the noise level and satisfies the following conditions: This indicates that the noise gradually increases over time; q ( x t | x t-1 ) represents a conditional Gaussian distribution with a mean of . The variance is ; x t Indicates the first t The diffusion data of the step, x 0 represents the original data. q ( x 1:T | x 0) represents the joint probability distribution of the forward process, that is, the complete transformation process from the original data to the final noisy data. I This represents the distribution range of standard Gaussian noise. Indicates the parameter α t Sum.
[0079] A reverse denoising process is constructed, using a trained improved diffusion model to gradually restore the noisy data, generating current multi-scenario data:
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087] in,q x t x t q x t-1 x t-1 q x 0 q x t x t-1 t t q x t-1 x t t t q x t x 0 t q x t-1 x t x 0 t t p θ x 0:T c c p θ x t-1 x t c c t t Mu θ x t ,t ∣ c ) represents the posterior Gaussian distribution mean under the current state, diffusion step t and condition information c, Σ θ ( x t , t ∣ c ) represents the neural network parameterized covariance function, represents the noise term, and c represents the given condition state, represents the noise injected at t step.
[0088] Step S32, training the improved diffusion neural network model based on the forward diffusion process and the backward denoising process, and making the improved diffusion neural network model generate multi-scenario source load uncertainty data.
[0089] In this embodiment, the forward diffusion maps the complex historical source load distribution to the standard Gaussian, and the backward denoising learns its full distribution, which can retain the non-Gaussian characteristics and time correlation and cross-variable correlation, avoid the deviation caused by only point prediction or simple noise model, sample from high-noise state and gradually denoise, can synthesize tail and extreme scenarios, improve the coverage of rare events such as wind and light sudden drop and peak load, help to configure reserve and safety margin, reduce the risk of loss of load and wind and light curtailment, and through adjusting the diffusion step, noise intensity or condition information, diversified scenarios of different severity and different time scale can be generated, reducing the mode collapse phenomenon, and being more stable and reliable than traditional generation methods.
[0090] Step S4, constructing a SAC agent according to the initial scheduling strategy and the multi-scenario source load uncertainty data, pre-training the SAC agent by using the transfer learning strategy, deploying the pre-trained SAC agent on the integrated energy system model, and outputting the optimal scheduling strategy to the monitored object group.
[0091] In this step, please refer to Figure 4 , construct a SAC agent, pre-train it by using historical operation data set, and the cumulative return of the random strategy can be represented as:
[0092] Wherein, π* represents a policy function that can maximize the expected cumulative return, s t represents the environment state of the system at time step t . a t represents the action taken by the agent at time step t . r(s t ,at ) denotes a time step t performing an action a t the obtained immediate reward value, denotes a discount factor, denotes a policy entropy term, denotes a temperature parameter for balancing the trade-off between expected reward and entropy, Rho π denotes a policy π the resulting state distribution and state-action distribution, E (st,at)~ρπ [ ] denotes the average over all possible state-action trajectories under the policy π guided by argmax, which denotes the maximization operation.
[0093] The Bellman value function can be expressed as:
[0094]
[0095] where, Q ( s t , a t ) denotes a state-action value function, denotes a soft state value function, p denotes a state transition probability, s t+1 ~ p denotes the sampling of the next state according to the transition probability p , π ( s t , a t ) denotes a policy function at the current state, Es t+1 ~ p [ ] denotes the average over all possible next states based on the state transition probability p , r(s t ,a t ) denotes a time step t performing an action a t the obtained immediate reward value.
[0096] The SAC agent update can be expressed as:
[0097]
[0098]
[0099]
[0100]
[0101]
[0102] where, π new denotes the updated new policy function, D KL denotes the KL divergence, r t denotes the time step t performs an action a t the obtained immediate reward value, Π denotes the feasible policy space, Z old ( s t ) is usually represented by a standard Gaussian distribution for normalizing the distribution, Q old ( s t , ·) denotes the old state-action value function, denotes the target soft Q network, J Q ( Theta ) denotes the loss function of the Q network, Q θ ( s t , a t ) denotes the parameterized Q value network, ( s t , a t ) denotes the target Q value, Theta and denote the parameters of the value network and its corresponding target network, respectively, denotes the soft update smoothing factor, Phi denotes the parameters of the policy network, π ϕ ( a t | s t ) denotes the parameterized policy function, i.e., selecting action s t under state a tThe probability, J π ( Phi () represents the loss function of the policy network. E at~π [ ]: Regarding strategy π The expected distribution of actions, E st~D,at~πϕ [ ]: Expectations regarding sampling state from the replay buffer and sampling action from the current policy.
[0103] Strategy selection can be expressed as:
[0104] in, This represents the output policy network. a t Indicates the system at time step t The actions taken by the intelligent agent This represents the mean of the output strategy. This represents the variance of the output strategy. Epsilon t ⊙ represents random noise, and ⊙ represents the Hadamard product symbol.
[0105] This step also includes steps S41 to S42.
[0106] Step S41: From the multi-scenario source-load uncertainty data generated by the improved diffusion neural network model, classify the multi-scenario source-load uncertainty data according to the scenario characteristics, and construct a typical scenario library.
[0107] In this step, the generated data is first analyzed for its scene characteristics, extracting key feature parameters such as the fluctuation characteristics of photovoltaic and wind power output, the time-varying patterns of various loads, quantitative indicators of uncertainty, and periodic patterns across multiple time scales. Then, classification criteria are established based on multiple dimensions, including renewable energy output level, load demand intensity, uncertainty, operating condition type, and time attributes. Machine learning algorithms such as K-means clustering, hierarchical clustering, and density clustering are used to automatically classify the multidimensional feature space and optimize it with expert knowledge. Next, the most representative typical scenes are extracted from each category by analyzing the representativeness of each cluster, calculating the distance between scenes, applying principal component analysis for dimensionality reduction, and using information gain criteria. Finally, a standardized typical scene library is constructed, assigning a unique identifier and descriptive label to each scene, establishing a multidimensional index system and metadata database, and designing a dynamic update mechanism to ensure that the scene library can comprehensively cover the main operating conditions that the system may encounter, providing a diverse and targeted training environment for the pre-training of the SAC agent.
[0108] Step S42, typical scenarios are divided into different training environments according to complexity and uncertainty level, and the source-load time series data, state transition function and reward calculation mechanism of the corresponding scenario are encapsulated in each training environment to pre-train the SAC agent.
[0109] In this step, the complexity and uncertainty level of each typical scenario are evaluated by calculating the source-load combination dimension, constraint complexity, decision variable coupling relationship, and prediction error variance, random disturbance amplitude, and extreme event impact, and then the scenarios are divided into three levels of training environments: simple, medium and complex, and a progressive training sequence is designed. Then, in each training environment, the source-load time series data interface, state transition function based on physical constraints, reward calculation mechanism integrating multiple objectives, and constraint violation penalty strategy of the corresponding scenario are encapsulated. Finally, a phased pre-training strategy is adopted to let the SAC agent gradually learn the basic control law, multi-objective balancing ability and extreme working condition coping ability from simple environments. Through the optimization of hyperparameters, the design of transfer learning mechanism and the establishment of cross-environment experience replay mechanism, the agent can obtain sufficient training and accumulate rich control experience in different complexity environments.
[0110] In this embodiment, the energy storage charging and discharging, electric-thermal-gas coupled equipment and renewable energy output are optimized collaboratively to minimize the total system operating cost, reduce start-stop loss and curtailment of wind and light, significantly reduce comprehensive energy consumption and energy purchase cost, and reduce the use of fossil energy and carbon emission intensity while meeting energy demand and safety constraints. The unification of economy and low carbon is achieved, and the diffusion neural network is improved to generate multi-scenario source-load uncertainty data, so that the training covers extreme and rare working conditions, and the robustness and generalization ability of the strategy under renewable power fluctuation and load mutation are improved. The risk of scheduling mismatch caused by prediction deviation is reduced, thereby effectively improving the optimization accuracy of the scheduling strategy. The operating constraints of the comprehensive energy system are explicitly introduced into the reinforcement learning strategy search process, so that the resulting strategy naturally meets the engineering feasibility and safety boundary, and avoids out-of-bound operations.
[0111] In one example, referring to Tables 1 and 2, for the uncertainty of renewable energy and load, compared with other methods, the improved diffusion network proposed in this embodiment performs better in scenario generation, with renewable energy coverage rate improved by 2.74% and load coverage rate improved by 2.07% compared with the Monte Carlo method. Compared with the SAC and PSO algorithms, the proposed method reduces the overall cost by 4.36% and 8.26%, respectively, and reduces the carbon emission related cost by 4.71% and 7.48%, respectively, fully verifying the advantages of the proposed method in economy and low carbon.
[0112] Table 1 Comparison of different scenario generation methods
[0113] Table 2 Cost comparison of different optimization scheduling methods
[0114] Based on the above method, the embodiment of the application discloses a comprehensive energy system low-carbon optimization scheduling system based on deep reinforcement learning, which refers to Figure 5 The comprehensive energy system low-carbon optimization scheduling system 1 comprises a data acquisition module 11, a model construction module 12 and a strategy optimization module 13, wherein The data acquisition module 11 is used for acquiring a multi-dimensional operation data set of a monitored object group, and constructing a comprehensive energy system model based on the multi-dimensional operation data set; The model construction module 12 is used for constructing a deep reinforcement learning scheduling model with the minimum system total operation cost corresponding to the comprehensive energy system model as the target, and obtaining an initial scheduling strategy based on the operation constraint condition of the comprehensive energy system model and the deep reinforcement learning scheduling model; constructing an improved diffusion neural network model based on the multi-dimensional operation data set and multi-class load historical data, and enabling the improved diffusion neural network model to generate multi-scenario source-load uncertainty data; The strategy optimization module 13 is used for constructing an SAC agent according to the initial scheduling strategy and the multi-scenario source-load uncertainty data, pre-training the SAC agent with a transfer learning strategy, deploying the pre-trained SAC agent to the comprehensive energy system model, and outputting an optimal scheduling strategy to the monitored object group.
[0115] In one example, the data acquisition module 11 is used for determining a monitored object group, and acquiring a multi-dimensional operation data set generated in the operation process of the monitored object group; performing data preprocessing on the multi-dimensional operation data set to form a standardized time sequence input matrix, and constructing a comprehensive energy system model based on the standardized time sequence input matrix.
[0116] In one example, the monitored object group comprises a combined heat and power unit, an electric boiler, a gas boiler, an electric refrigerator, an absorption refrigerator, a battery energy storage system, a photovoltaic power generation device, a wind power generation device, an electric load, a heat load, a cold load and an interface between an external power grid and a gas pipeline network, and the multi-dimensional operation data set comprises input power of each energy device, output power of each energy device, energy conversion efficiency, charge and discharge power and state of charge of the energy storage system, output of the renewable energy, multi-class load demand, external environment temperature, energy price signal, carbon emission price signal and green certificate transaction price signal.
[0117] In one example, the system total operation cost comprises carbon emission cost, green certificate transaction cost, battery energy storage system degradation cost, natural gas procurement cost and main grid power procurement cost, wherein The carbon emission quota based on the comprehensive energy system model, the allocation quota of the combined heat and power and the gas boiler, the power generation reference quota of the combined heat and power and the gas boiler, and the heat supply reference quota of the combined heat and power and the gas boiler, and the carbon emission cost are calculated by using the step-type carbon trading price function; The green certificate trading cost is calculated based on the number of certificates required for compliance, the green certificate quota coefficient allocated by the regulatory agency, and the number of green certificates obtained by the comprehensive energy system model. The battery energy storage model is constructed according to the cycle life of the battery energy storage system in the comprehensive energy system model and the degradation cost of the battery energy storage system, and the degradation cost of the battery energy storage system is obtained by converting the capital investment cost into the degradation cost per unit of energy throughput in the battery energy storage model. The main grid power purchase cost is calculated based on the purchase and sale price settlement of the main grid power, the cost of purchasing natural gas from the gas pipeline network, and the supply amount of natural gas obtained by the comprehensive energy system model.
[0118] In one example, the model construction module 12 is configured to convert the historical source load time series data in the multi-dimensional operation data set into a standard Gaussian distribution by gradually adding Gaussian noise to construct a forward diffusion process, and gradually restore the standard Gaussian noise to form multi-scenario source load data to construct a backward diffusion process; train the improved diffusion neural network model based on the forward diffusion process and the reverse denoising process, and make the improved diffusion neural network model generate multi-scenario source load uncertainty data.
[0119] In one example, the strategy optimization module 13 is configured to classify the multi-scenario source load uncertainty data according to the scene characteristics to construct a typical scene library, divide the typical scenes into different training environments according to the complexity and uncertainty degree, and encapsulate the source load time series data, state transition function and reward calculation mechanism of the corresponding scenes in each training environment to pre-train the SAC agent.
[0120] In one example, the operation constraints include combined heat and power equipment constraints, electric boiler equipment constraints, gas boiler equipment constraints, electric refrigeration equipment constraints, absorption refrigeration equipment constraints, battery energy storage system constraints, and main grid interaction constraints.
[0121] See Figure 6 A structural schematic diagram of an electronic device is provided for the embodiments of the present application. As shown in Figure 6 The electronic device 2 can include at least one processor 21, at least one network interface 24, a user interface 23, a memory 25, and at least one communication bus 22.
[0122] The communication bus 22 is used to realize the connection and communication between the components.
[0123] The user interface 23 can include a display, a camera, and optionally a standard wired interface and a wireless interface.
[0124] The network interface 24 can optionally include a standard wired interface and a wireless interface (e.g., a WI-FI interface).
[0125] The processor 21 can include one or more processing cores. The processor 21 connects various parts of the server through various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 25, and calling data stored in the memory 25. Optionally, the processor 21 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 21 can be integrated with a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU is mainly used to process an operating system, a user interface, and an application program. The GPU is used to render and draw the content to be displayed on the display. The modem is used to process wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 21, but can be implemented by a separate chip.
[0126] The memory 25 can include a random access memory (RAM) and a read-only memory (ROM). Optionally, the memory 25 includes a non-transitory computer-readable storage medium. The memory 25 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 25 can include a program storage area and a data storage area. The program storage area can store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The data storage area can store data involved in the above-mentioned various method embodiments, etc. The memory 25 can optionally be at least one storage device located away from the above-mentioned processor 21. For example, Figure 6As shown, the memory 25 as a computer storage medium can include an operating system, a network communication module, a user interface module, and an application program of the low-carbon optimization scheduling method of the integrated energy system based on deep reinforcement learning.
[0127] In Figure 6 As shown in the electronic device 2, the user interface 23 is mainly used to provide an interface for user input and obtain data input by the user; and the processor 21 can be used to call the application program of the low-carbon optimization scheduling method of the integrated energy system based on deep reinforcement learning stored in the memory 25, and when executed by one or more processors, make the electronic device execute one or more methods in the above embodiments.
[0128] A computer-readable storage medium stores instructions. When executed by one or more processors, the computer performs one or more methods as described in the above embodiments.
[0129] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all described as a series of action combinations, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other order or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0130] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0131] In the several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic. The division of the units is only a logical function division. There can be another division manner for actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different parts can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical or other forms.
[0132] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, some or all of the units can be selected to achieve the purpose of the embodiment scheme.
[0133] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.
[0134] When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable memory. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a memory, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned memory includes: a U disk, a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0135] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A low-carbon optimization scheduling method for integrated energy systems based on deep reinforcement learning, characterized in that, The method includes: Collect multidimensional operational datasets of the monitored object group, and construct a comprehensive energy system model based on the multidimensional operational datasets; With the goal of minimizing the total operating cost of the integrated energy system model, a deep reinforcement learning scheduling model is constructed, and an initial scheduling strategy is obtained based on the operating constraints of the integrated energy system model and the deep reinforcement learning scheduling model. Based on the multidimensional operational dataset and various types of historical load data, an improved diffusion neural network model is constructed, and the improved diffusion neural network model generates source-load uncertainty data for multiple scenarios. Based on the initial scheduling strategy and the multi-scenario source-load uncertainty data, a SAC agent is constructed. The SAC agent is pre-trained using a transfer learning strategy. The pre-trained SAC agent is deployed in the integrated energy system model and outputs the optimal scheduling strategy to the monitoring object group.
2. The low-carbon optimization scheduling method for a comprehensive energy system based on deep reinforcement learning as described in claim 1, characterized in that, The construction of the integrated energy system model based on the multidimensional operational dataset specifically includes: Identify the monitoring object group and collect the multidimensional operational dataset generated during the operation of the monitoring object group; The multidimensional operational dataset is preprocessed to form a standardized time-series input matrix, and a comprehensive energy system model is constructed based on the standardized time-series input matrix.
3. The low-carbon optimization scheduling method for integrated energy systems based on deep reinforcement learning as described in claim 2, wherein the monitoring object group includes cogeneration units, electric boilers, gas boilers, electric chillers, absorption chillers, battery energy storage systems, photovoltaic power generation devices, wind power generation devices, power loads, heat loads, cold loads, and interfaces between the external power grid and the gas pipeline network; the multi-dimensional operation dataset includes the input power of each energy device, the output power of each energy device, the energy conversion efficiency, the charging and discharging power and state of charge of the energy storage system, the output of renewable energy, various load demands, external ambient temperature, energy price signals, carbon emission price signals, and green certificate trading price signals.
4. The low-carbon optimization scheduling method for a comprehensive energy system based on deep reinforcement learning as described in claim 1, characterized in that, The total operating cost of the system includes carbon emission costs, green certificate trading costs, battery energy storage system degradation costs, natural gas purchase costs, and mains grid power purchase costs. The carbon emission cost is calculated using a tiered carbon trading price function based on the carbon emission quotas, allocation quotas for cogeneration and gas-fired boilers, power generation benchmark quotas for cogeneration and gas-fired boilers, and heating benchmark quotas for cogeneration and gas-fired boilers, all based on the integrated energy system model. The transaction cost of the green certificates is calculated based on the number of certificates required for compliance, the green certificate quota coefficient allocated by the regulatory agency, and the number of green certificates obtained from the integrated energy system model. A battery energy storage model is constructed based on the cycle life and degradation cost of the battery energy storage system in the integrated energy system model. In the battery energy storage model, the degradation cost per unit energy throughput is converted from the capital investment cost to obtain the degradation cost of the battery energy storage system. The main grid power purchase cost is calculated based on the settlement of electricity purchase and sale prices from the main grid, the cost of purchasing natural gas from the gas pipeline network, and the natural gas supply volume of the integrated energy system model.
5. The low-carbon optimization scheduling method for a comprehensive energy system based on deep reinforcement learning as described in claim 1, characterized in that, The steps for constructing the improved diffusion neural network model based on the multidimensional operational dataset and various types of historical load data include: The historical source-load time-series data in the multidimensional running dataset is transformed into a standard Gaussian distribution by gradually adding Gaussian noise to construct the forward diffusion process, and the standard Gaussian noise is gradually restored to form multi-scenario source-load data to construct the backward diffusion process. The improved diffusion neural network model is trained based on the forward diffusion process and the reverse denoising process, and the improved diffusion neural network model generates source-load uncertainty data for multiple scenarios.
6. The low-carbon optimization scheduling method for a comprehensive energy system based on deep reinforcement learning as described in claim 1, characterized in that, The pre-training of the SAC agent using a transfer learning strategy specifically includes: From the multi-scenario source-load uncertainty data generated by the improved diffusion neural network model, the multi-scenario source-load uncertainty data is classified according to scenario characteristics to construct a typical scenario library; Typical scenarios are divided into different training environments according to their complexity and uncertainty. In each training environment, the source load time series data, state transition function and reward calculation mechanism of the corresponding scenario are encapsulated to pre-train the SAC agent.
7. The low-carbon optimization scheduling method for a comprehensive energy system based on deep reinforcement learning as described in claim 1, characterized in that, The operational constraints include constraints on combined heat and power (CHP) equipment, electric boiler equipment, gas boiler equipment, electric refrigeration equipment, absorption refrigeration equipment, battery energy storage system, and grid interaction constraints.
8. A low-carbon optimization scheduling system for integrated energy systems based on deep reinforcement learning, characterized in that, The integrated energy system low-carbon optimization scheduling system (1) includes a data acquisition module (11), a model building module (12), and a strategy optimization module (13), wherein, The data acquisition module (11) is used to collect a multi-dimensional operational dataset of the monitored object group and to construct a comprehensive energy system model based on the multi-dimensional operational dataset; The model building module (12) is used to construct a deep reinforcement learning scheduling model with the goal of minimizing the total operating cost of the system corresponding to the integrated energy system model, and to obtain an initial scheduling strategy based on the operating constraints of the integrated energy system model and the deep reinforcement learning scheduling model. Based on the multidimensional operating dataset and multi-type load historical data, an improved diffusion neural network model is constructed, and the improved diffusion neural network model generates multi-scenario source-load uncertainty data. The strategy optimization module (13) is used to construct an SAC agent based on the initial scheduling strategy and the multi-scenario source-load uncertainty data, pre-train the SAC agent with a transfer learning strategy, deploy the pre-trained SAC agent on the integrated energy system model, and output the optimal scheduling strategy to the monitoring object group.
9. An electronic device, characterized in that, The device includes a processor (21), a memory (25), a user interface (23), and a network interface (24). The memory (25) is used to store instructions. The user interface (23) and the network interface (24) are used to communicate with other devices. The processor (21) is used to execute the instructions stored in the memory (25) to cause the electronic device (2) to perform the method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Integrated energy system low-carbon optimization scheduling method based on action adjustment reinforcement learning
CN120046784A
Cited By
A method for intelligent scheduling of an offshore energy island system
CN122288332A