Drl-based low-carbon scheduling method and system for shaguo desert power-hydrogen-carbon-methanol system
By constructing an electric-hydrogen-carbon-methanol multi-energy coupled system model and DRL optimization scheduling, the problems of difficult new energy consumption and high carbon emissions in the desert and Gobi areas have been solved, achieving efficient and low-carbon new energy consumption and economic optimization.
Patent Information
- Application Number
- CN202511634084.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-11-10
AI Technical Summary
The new energy base in the desert area is located in a remote area with a weak power grid structure and a lack of flexible adjustment resources, which leads to difficulties in the consumption of new energy, high carbon emissions and poor economic efficiency. Existing scheduling methods have limitations such as complex modeling, strong conservatism or high computational cost when dealing with high-dimensional uncertainty problems.
A multi-energy coupled system model of electricity-hydrogen-carbon-methanol is constructed, and a flexible carbon capture mechanism and a carbon resource utilization path for green hydrogen to synthesize methanol are introduced. The orthogonal attention mechanism based on DRL and the double-delay deep deterministic strategy gradient algorithm (OA-TD3) are used for optimization scheduling. Through multi-energy coupling and coordinated regulation, the on-site consumption capacity and economy of new energy are improved.
It has significantly improved the local consumption capacity of new energy and the economic efficiency of system operation, enhanced the robustness and practical feasibility of dispatch strategies, achieved coordination between energy supply and demand and multi-energy complementarity, and reduced carbon emissions and operating costs.
Smart Images

Figure CN121119630B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of new energy power system optimal dispatching, and particularly relates to a low-carbon dispatching method and system for a desert-gobi-desert electricity-hydrogen-carbon-methanol system based on DRL (Deep Reinforcement Learning). BACKGROUND
[0002] Promoting the construction of large-scale wind and photovoltaic bases in desert, gobi and desert areas is an important measure for China to achieve the "double carbon" goal. However, new energy bases in desert, gobi and desert areas are located in remote areas, and the power grid structure is weak, with a lack of flexible adjustment resources. In addition, the intermittency and volatility of wind power and photovoltaic power output lead to multiple challenges such as difficulty in new energy consumption, high carbon emissions and poor economic efficiency in system operation.
[0003] At present, the dispatching research on new energy bases in desert, gobi and desert areas mostly focuses on the "wind-solar-thermal storage" bundling and external transmission mode, and insufficient attention is paid to local collaborative consumption and multi-energy coupling utilization. Although some research has introduced electric-hydrogen collaboration and carbon capture technologies, the flexible operation characteristics of the carbon capture system are not fully considered in the dispatching model, and the resource utilization path of green hydrogen and carbon dioxide to synthesize methanol is not effectively utilized, resulting in that the system regulation capacity and economic benefits cannot be fully utilized. In addition, traditional optimization methods such as stochastic programming and robust optimization have limitations such as complex modeling, strong conservatism or high computational cost when dealing with high-dimensional uncertainty problems. Therefore, how to effectively utilize the electric-hydrogen-carbon-methanol multi-energy coupling and collaborative control means, break through the limitations of a lack of flexible resources and complex operation constraints in new energy bases in desert, gobi and desert areas, and realize high-proportion consumption of new energy and low-carbon economic collaborative optimization of the system on the premise of ensuring stable operation of the power grid has become an important problem to be solved in the field of power system dispatching. SUMMARY
[0004] In view of the above problems, the purpose of the present application is to provide a low-carbon dispatching method and system for a desert-gobi-desert electricity-hydrogen-carbon-methanol system based on DRL, which improves the local consumption capacity of new energy through multi-energy coupling and flexible carbon capture mechanism, and realizes economic and low-carbon collaborative optimization of the system by using an improved reinforcement learning algorithm. The technical solution is as follows:
[0005] A low-carbon dispatching method for a desert-gobi-desert electricity-hydrogen-carbon-methanol system based on DRL, comprising the following steps:
[0006] Step 1: building an electric-hydrogen-carbon-methanol multi-energy coupling system model, specifically including a flexible carbon capture model, an electrolytic tank and a hydrogen buffer tank model, a methanol synthesis model and an electrochemical energy storage model;
[0007] Step 2: Construct a target function with the minimum total system operation cost, including the fuel cost of the thermal power unit, the loss cost of the carbon capture solvent, the carbon sequestration cost, the hydrogen production cost by water electrolysis, the energy storage usage cost, the hydrogen storage cost, the methanol revenue, the wind and light abandonment cost, and the carbon emission cost;
[0008] Step 3: Establish system operation constraints, including wind and light output constraints, hydrogen storage balance constraints, carbon capture system constraints, hydrogen production by water electrolysis constraints, and electrochemical energy storage device constraints;
[0009] Step 4: Model the system scheduling problem as a Markov decision process, with the system scheduling center as the agent and the electric-hydrogen-carbon-methanol system as the environment, defining the state space, action space, and reward function;
[0010] Step 5: Solve the Markov decision process using a double-delay deep deterministic policy gradient algorithm based on the orthogonal attention mechanism to obtain the optimal scheduling strategy.
[0011] A DRL-based low-carbon scheduling system for the Shaguohe electric-hydrogen-carbon-methanol system, comprising a system modeling module, a target and constraint module, an algorithm framework module, and an optimized scheduling module;
[0012] The system modeling module is used to construct an electric-hydrogen-carbon-methanol multi-energy coupling system model containing wind power generation, photovoltaic power generation, carbon capture thermal power units, hydrogen production by water electrolysis, methanol synthesis, electrochemical energy storage, and hydrogen buffer tanks;
[0013] The target and constraint module is used to set the minimum total system operation cost as the objective function, and to set the system operation constraints, wherein the total operation cost includes the fuel cost of the thermal power unit, the loss cost of the carbon capture solvent, the carbon sequestration cost, the hydrogen production cost by water electrolysis, the energy storage usage cost, the hydrogen storage cost, the methanol revenue, the wind and light abandonment cost, and the carbon emission cost; the constraints include wind and light output constraints, hydrogen storage balance constraints, carbon capture system constraints, hydrogen production by water electrolysis constraints, and electrochemical energy storage device constraints;
[0014] The algorithm framework module is used to establish a double-delay deep deterministic policy gradient algorithm framework embedded with an orthogonal attention mechanism based on a Markov decision process, with the system scheduling center as the agent and the electric-hydrogen-carbon-methanol system as the environment, defining the state space, action space, and reward function;
[0015] The optimized scheduling module is used to optimize the scheduling of the electric-hydrogen-carbon-methanol multi-energy coupling system model based on the double-delay deep deterministic policy gradient algorithm framework embedded with the orthogonal attention mechanism, outputting the optimal output strategy of each device and achieving the collaborative optimization of system economy and low carbon.
[0016] The beneficial effects of the present application are:
[0017] 1) The present application fully considers the strong volatility of wind power and photovoltaic output in the Shagexuan region and the complexity of system operation constraints. By constructing an electricity-hydrogen-carbon-methanol multi-energy coupling system model, introducing a flexible carbon capture unit with a storage tank and a carbon resourceization path of green hydrogen synthesis methanol, the energy supply and demand on both sides are coordinated and multi-energy is complementary, which significantly improves the local consumption capacity of new energy and the economic efficiency of system operation;
[0018] 2) The algorithm of the present application embeds an orthogonal attention module in the critic network. By recalibrating the features, the accuracy of the value function estimation is improved, and the operating constraints are integrated into the reward function in the form of a penalty term, guiding the agent to consider system safety and optimization objectives during decision-making, thereby enhancing the robustness and practical feasibility of the scheduling strategy;
[0019] 3) The present application uses Markov decision process to establish an orthogonal attention mechanism-based double-delay deep deterministic policy gradient (Orthogonal Attention-Twin Delayed Deep Deterministic policy gradient algorithm, OA-TD3) algorithm framework. Through orthogonal constraints, feature decoupling and redundancy suppression are realized, the convergence speed and stability of the algorithm in high-dimensional uncertain environments are improved, and the optimal scheduling strategy can be generated automatically according to the system state, realizing the collaborative optimization of economic efficiency and low carbon. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 The flowchart of the scheduling method of the present application.
[0021] Figure 2 The orthogonal attention mechanism-based double-delay deep deterministic policy gradient (OA-TD3) algorithm of the present application.
[0022] Figure 3 The training effect comparison chart of the scheduling control model in the present application.
[0023] Figure 4 The scheduling result diagram after applying the present application. DETAILED DESCRIPTION
[0024] The present application will be further described in detail below in conjunction with the drawings and specific embodiments.
[0025] Referring to Figure 1 and Figure 2 , the present application provides a deep reinforcement learning-based low-carbon scheduling method for Shagexuan electricity-hydrogen-carbon-methanol systems, which specifically includes the following steps:
[0026] Step S1: Construct an electro-hydrogen-carbon-methanol multi-energy coupled system model, which includes a flexible carbon capture model, an electrolyzer and hydrogen buffer tank model, a methanol synthesis model, and an electrochemical energy storage model.
[0027] Step S1.1: Flexible carbon capture model;
[0028] Specifically, the flexible carbon capture model is modeled as follows:
[0029] ;
[0030] In the formula, For carbon capture units in Total power generation during the period; This refers to the net output power of the carbon capture unit; Fixed energy consumption for carbon capture units; For carbon capture units Energy consumption during different time periods. Overall. The high energy consumption and high cost in the capture process mainly stem from the energy supply required in the regeneration stage. That is, the energy consumption of desorption in the operation is much higher than the energy consumption of absorption. Therefore, in this invention, the operation energy consumption only considers the energy consumption of desorption.
[0031] The overall carbon emission level of the unit is related to the different power generation outputs of the carbon capture unit, as shown in the following formula:
[0032] ;
[0033] In the formula, For thermal power units Total generation of time period quantity; For the unit of power generation output quantity.
[0034] Here, we distinguish between the carbon dioxide being processed by the absorption tower and the carbon dioxide being processed by the regeneration tower. The expression is:
[0035] ;
[0036] ;
[0037] In the formula, For absorption tower The time period is being processed quantity; This is the ratio of the amount of flue gas diverted into the carbon capture system to the total amount of flue gas generated on the power generation side. For carbon capture and absorption efficiency; For the collection unit Energy consumption; For regeneration tower The time period is being processed quantity.
[0038] Due to the presence of the storage tank, the regeneration tower processes... quantity Satisfy the following formula:
[0039] ;
[0040] In the formula, for Periodic flow into the rich liquid tank The amount, A negative value indicates that the solution flows from the rich liquid tank to the regeneration tower, while a positive value indicates that the solution flows from the absorption tower into the rich liquid tank.
[0041] When present in solution as a compound, the relationship between the compounds and the alcoholamines is as follows:
[0042] ;
[0043] In the formula, and The first The amount of solution flowing into the lean fluid storage device during the time period and the first The amount of solution flowing into the rich liquid storage tank during a given time period; a positive value indicates inflow, and a negative value indicates outflow. and These represent ethanolamine (MEA) solution and molar mass; The regeneration resolution coefficient; This represents the mass percentage concentration of ethanolamine. The density of the ethanolamine solution; and The first The storage capacity of the low-lying fluid storage tank and the storage capacity of the high-lying fluid storage tank during the time period.
[0044] Step S1.2: Model of electrolyzer and hydrogen buffer tank;
[0045] The hydrogen production capacity of an alkaline electrolyzer is directly proportional to the input power of the electrolyzer, as shown in the simplified expression below:
[0046] ;
[0047] ;
[0048] In the formula, express The molar flow rate of hydrogen produced during a given period; Indicates the hydrogen production efficiency of the electrolyzer; This indicates the current input power of the electrolytic cell; This indicates the calorific value of hydrogen. express The quality of hydrogen produced during a given period; This indicates the molar mass of hydrogen.
[0049] A hydrogen buffer tank is also installed, through which green hydrogen enters the methanol synthesis unit. This buffer tank serves as a buffer and provides some storage. The hydrogen buffer model can be represented as follows:
[0050] ;
[0051] ;
[0052] ;
[0053] In the formula, Indicates in Hydrogen storage capacity in the hydrogen storage tank during a given period; express The amount of hydrogen entering the storage tank during a given period is the amount of green hydrogen produced. express The outflow rate of hydrogen storage tank during a given period; express The amount of hydrogen consumed at any given moment.
[0054] Step S1.3: Methanol synthesis model;
[0055] The mass relationships between methanol and its compounds are as follows:
[0056] ;
[0057] In the formula, and These represent the mass ratios of hydrogen and carbon dioxide consumed to produce one unit of methanol. express The amount of methanol produced at any given time; express The amount of carbon dioxide consumed at any given moment; express The power consumed in the production of methanol at any given time; This represents the power consumed to produce one unit mass of methanol.
[0058] Step S1.4: Electrochemical energy storage model;
[0059] Wind and solar power systems come with built-in energy storage devices. Considering the electrochemical energy storage inherent in wind and solar power systems as a single energy storage system, its mathematical state model can be written as:
[0060] ;
[0061] In the formula, For electrochemical energy storage Capacity of a time period; , These represent the battery charging and discharging efficiencies, respectively. , For energy storage batteries Charging and discharging power during a given time period; This represents a time period within the scheduling cycle.
[0062] Step S2: The objective function is to minimize the total operating cost of the system. The total operating cost includes the fuel cost of thermal power units, the cost of carbon capture solvent loss, the cost of carbon sequestration, the cost of hydrogen production by water electrolysis, the cost of energy storage, the cost of hydrogen storage, the revenue from methanol, the cost of wind and solar curtailment, and the cost of carbon emissions.
[0063] The objective of the model proposed in this invention is to effectively adjust the output and status of various devices in the system while meeting load and transmission demands, thereby reducing wind and solar power curtailment and carbon dioxide emission costs, and simultaneously minimizing the overall economic cost of the system. Objective function Defined as follows:
[0064] ;
[0065] In the formula, For system operating costs; Costs associated with wind and solar power curtailment; For the system's carbon emission costs; For the entire scheduling cycle.
[0066] Step S2.1: System operating cost;
[0067] Operating costs include fuel costs for thermal power units. Solvent loss costs during carbon capture process Carbon sequestration costs Cost of hydrogen production by water electrolysis Hydrogen storage cost Energy storage usage costs and methanol revenue The model is as follows:
[0068] ;
[0069] ;
[0070] ;
[0071] ;
[0072] ;
[0073] ;
[0074] ;
[0075] ;
[0076] In the formula, , and This is the power generation cost coefficient for thermal power units; This is the cost coefficient for ethanolamine solvent; This is the solvent operating loss coefficient; Absorbed by ethanolamine The amount; as a unit Cost coefficients for transportation and storage; Consumption for methanol production The amount; This is the cost coefficient for hydrogen production via water electrolysis; The hydrogen storage cost coefficient for hydrogen buffer tanks; This refers to the energy storage charging and discharging cost coefficient. This refers to the unit price at which methanol is sold. This represents the unit cost of methanol production.
[0077] Step S2.2: System curtailment costs;
[0078] The curtailment of wind and solar power is transformed into part of the economic cost through a penalty factor, which is defined as follows:
[0079] ;
[0080] ;
[0081] ;
[0082] In the formula, and These are the costs of wind curtailment and solar curtailment, respectively. and These are respectively the wind curtailment penalty factor and the solar curtailment penalty factor; and Each is the original power output of the wind and solar power system; and Each contributes to the actual scenery.
[0083] Step S2.3: System carbon emission cost;
[0084] The main entities included in this carbon quota management are carbon capture power plants, whose carbon quotas are allocated based on the amount of electricity they supply. It can be defined as follows:
[0085] ;
[0086] In the formula, This is the benchmark value for carbon emissions per unit of electricity used by coal-fired power units.
[0087] Specifically, the cost of carbon emissions can be defined as:
[0088] ;
[0089] In the formula, The unit carbon price; for The time system discharges outward The total amount.
[0090] Step S3: Establish system operation constraints, including power balance constraints, equipment operation safety constraints, carbon capture system constraints, electrolysis hydrogen production constraints, energy storage operation constraints, and hydrogen storage balance constraints.
[0091] Specifically, the constraint modeling involved in this invention is as follows:
[0092] Step S3.1: System wind and solar power output constraints;
[0093] To meet the requirements for green hydrogen production, thermal power units do not participate in hydrogen synthesis; hydrogen production power is provided solely by wind, solar, and energy storage. For better utilization and absorption of wind and solar power, their output participates in both hydrogen production and carbon capture and regeneration, as shown in the following formula:
[0094] ;
[0095] In the formula, To contribute to wind and solar power participation in carbon capture and recycling; To contribute wind and solar power to the load.
[0096] Step S3.2: Hydrogen storage balance constraints;
[0097] The balance between the inflow and outflow of hydrogen storage tanks within a scheduling cycle is defined by the following formula:
[0098] ;
[0099] Step S3.3: Constraints on the carbon capture system;
[0100] The output range and ramping constraints of the carbon capture unit are defined as follows:
[0101] ;
[0102] In the formula, and These represent the upper and lower limits of the carbon capture unit's output, respectively. This is the upper limit of the ramping power of the carbon capture unit.
[0103] The carbon capture storage tank has volume constraints, and the storage capacity of the solution storage must be kept consistent at the beginning and end of the scheduling day to prepare for optimized operation on the next scheduling day, as expressed by the following formula:
[0104] ;
[0105] In the formula, , These are the minimum storage volumes for the lean and rich liquid tanks, respectively. , These are the maximum liquid storage volumes of the lean and rich liquid tanks, respectively. , These represent the volumes of the lean and rich liquid tanks at the start of the scheduling cycle. , These represent the volumes of the lean and rich liquid tanks at the end of the scheduling cycle.
[0106] Step S3.4: Constraints on hydrogen production and storage via water electrolysis;
[0107] The electrolysis power of the electrolyzer is constrained by equipment limitations, and the hydrogen buffer tank also has its corresponding constraints to ensure that the thermal power plant does not produce hydrogen. The model is as follows:
[0108] ;
[0109] In the formula, and These are the upper and lower limits of the electrolytic cell power, respectively; and These are the upper and lower limits of the hydrogen storage tank capacity, respectively. and These represent the hydrogen storage capacity at the beginning and end of the scheduling cycle, respectively.
[0110] Step S3.5: Constraint of electrochemical energy storage device;
[0111] ;
[0112] In the formula, , These are the upper and lower limits of energy storage capacity, respectively. and A Boolean variable, representing Charging and discharging cannot be performed simultaneously during a given period. , These are the upper limits for charging power and discharging power, respectively.
[0113] Step S4: Model the system scheduling problem as a Markov decision process, defining the state space, action space, and reward function.
[0114] Step S4.1: State space;
[0115] During the training process of the electricity-hydrogen-carbon-methanol system, the information provided by the environment to the agent mainly includes the original output of the wind turbine and photovoltaic system, fluctuating load, the status of the carbon capture device's storage tank, energy storage status, and scheduling time, as shown in the following formula:
[0116] ;
[0117] In the formula, express The load status of the entire system during a given time period.
[0118] Step S4.2: Motion space;
[0119] After observing the system's state space information, the agent selects an action in the action space based on the policy function. This action is then input into the model as the system's power value. The action space includes actual wind and solar power output, net output power of the carbon capture unit, energy consumption of the carbon capture unit, wind and solar power output participating in carbon capture and regeneration, hydrogen production output from the electrolyzer, split ratio, and energy storage output, and is modeled as follows:
[0120] ;
[0121] In the formula, for The output of energy storage during different time periods.
[0122] Step S4.3: Reward function;
[0123] The reward function represents the agent's performance in... Take action in a state The immediate reward obtained is consistent with the system's objective function. Considering that system constraints must also be satisfied, to accelerate training and improve scheduling control effectiveness, a penalty term for exceeding system constraints is added to the above objective function, including a penalty for power imbalance. Penalties for exceeding the climbing limit Penalties for exceeding capacity limits for each device in the system .
[0124] ;
[0125] In the formula, , , and These are the weighting factors that balance these terms. In the system optimization process, the linear weighting strategy not only does not interfere with the actual cost, but also effectively accelerates model training and further improves the overall system performance by reasonably selecting a set of weight parameters.
[0126] Step S5: The Markov decision process is solved using the dual-delay deep deterministic policy gradient (OA-TD3) algorithm based on orthogonal attention mechanism to obtain the optimal scheduling policy. A schematic diagram of the algorithm's solution is shown below. Figure 2 As shown.
[0127] (1) Orthogonal attention model:
[0128] Specifically, the Orthogonal Attention (OA) model improves the discriminativeness and stability of feature representations by introducing orthogonal constraints during the attention generation process and learning a set of mutually decoupled, non-redundant feature bases. Unlike traditional channel attention mechanisms, OA does not rely solely on global statistics for weighting; instead, it enhances the structured representation capability of features by constraining the geometric properties of the projection matrix, ensuring parameter efficiency. The modeling is as follows:
[0129] Orthogonal attention module input feature vector It is expressed as follows:
[0130] ;
[0131] In the formula, Represents the set of real numbers. This represents the dimension of the input feature vector. This means that the feature vector... It is a A vector with n real components.
[0132] The orthogonal attention module generates attention weights and calibrates the input features through a dual "dimensionality reduction-dimensionality increase" mapping process. First, the weight matrix... The input features are projected onto a low-dimensional latent space, as follows:
[0133] ;
[0134] ;
[0135] In the formula, The compressed feature vector; This represents the dimension of the compressed feature vector; It is an activation function that sets the part of the input that is less than 0 to 0, while leaving the part that is greater than or equal to 0 unchanged. for transpose; It is an identity matrix; orthogonality constraints guarantee The row vectors are approximately orthogonal, thus reducing redundancy between features.
[0136] Subsequently, another set of approximately orthogonal matrices Map it back to the original dimension:
[0137] ;
[0138] ;
[0139] In the formula, This is the attention vector, representing the modeling of the importance of each channel of the input feature; This represents the Sigmoid activation function; for The transpose of .
[0140] Finally, element-wise weighting is applied to the input channels to achieve adaptive recalibration of features, modeled as follows:
[0141] ;
[0142] In the formula, This represents the recalibrated features. This indicates element-wise multiplication.
[0143] Strict orthogonality conditions are difficult to maintain during training. Therefore, this invention employs soft orthogonality constraints, introducing the following regularization term into the total loss function. :
[0144] ;
[0145] In the formula, This is the regularization coefficient, used to balance the strength of the error loss between the model prediction and the true value and the orthogonal regularization loss; Let Frobenius norm be denoted. This regularization term guides the behavior by penalizing the deviation between the projection matrix and the orthogonal matrix. and Approximating orthogonality improves the robustness of attention weight generation.
[0146] (2) Algorithm solution:
[0147] In traditional deep reinforcement learning algorithms, the critic network typically uses fully connected layers to process features, which easily introduces redundancy and noise, leading to unstable target value estimation and large training fluctuations. This invention proposes embedding an orthogonal attention mechanism into the Critic network within the TD3 algorithm framework. Specifically, the OA model applies orthogonal constraints, forcing the network to learn a set of orthogonal bases to recalibrate features, ensuring decoupling of input features, effectively reducing feature redundancy, and thus improving the accuracy of value function estimation and the overall performance of the algorithm.
[0148] Specifically, with a daily scheduling cycle and an hourly unit set as the time step for the daily scheduling task, the solution process is as follows:
[0149] First, the network parameters are initialized. The OA-TD3 algorithm includes a policy network, a value network consisting of two orthogonal attention networks, and their corresponding target policy network and target value network. The network parameters are as follows: Policy network parameters Parameters of Value Network 1 Parameters of Value Network 2 Target policy network parameters Parameters of Target Value Network 1 Parameters of Target Value Network 2 Initialize the above network parameters, initialize the experience storage buffer, and initialize the value network weight matrix. And set the orthogonal regularization coefficient. .
[0150] Secondly, in the electric-hydrogen-carbon-methanol system environment, the agent acts according to the current strategy. With exploring noise selection action The environment returns to the next state. and rewards and will Store in the experience replay buffer.
[0151] Then, random sampling is performed from the experience playback buffer. Batch data The next action is generated through the target policy network. ,in Smooth the noise to the target; It is a cutoff function that limits the noise value to the interval middle, This indicates that the mean is 0 and the standard deviation is 0. The normal distribution; These are the upper and lower limits for truncation.
[0152] The target value is calculated using the target value network, as shown in the following formula:
[0153] ;
[0154] In the formula, for Target value for the time period; In the state Take action below The reward value obtained; Discount factor; The values are 1 and 2, corresponding to network 1 and network 2 respectively; For the first Parameters of a target value network; For the first A target value network in state Take action below The target value.
[0155] Minimize the total loss function of the value network To update the network parameters, the model is as follows:
[0156] ;
[0157] ;
[0158] ;
[0159] In the formula, N is the total number of batches; , and They represent the first The status, actions, and target values of each batch; This is the orthogonal attention constraint term, i.e., the regularization term of the soft orthogonal constraint, used to ensure that the feature base learned by the OA module maintains approximately orthogonality; For the first A target value network in state Take action below Target value; For the first Parameters of a value network; For value network hyperparameters; This represents the gradient of the total loss function. This is the differential symbol.
[0160] Following this, the policy network parameters are updated every few steps according to the delayed update mechanism. The policy is improved by maximizing the target value network's evaluation of the action, as modeled below:
[0161] ;
[0162] ;
[0163] in, To determine the expected reward that can be obtained by taking the corresponding action distribution. For expected return The gradient; Indicates the state Calculate the gradient of the policy network below; Representing state With action The gradient of value network 1 is then calculated. These are the hyperparameters of the policy network.
[0164] The target network parameters are updated incrementally using a soft update strategy, as modeled below:
[0165] ;
[0166] In the formula, This is the soft update coefficient, used to ensure network stability through a soft update strategy.
[0167] Finally, through continuous interaction, sampling, and training, the system schedule target converges or the number of training rounds is reached.
[0168] Figure 3 and Figure 4 The results of running the low-carbon scheduling method for the Shago wasteland electricity-hydrogen-carbon-methanol system based on deep reinforcement learning proposed in this invention are presented.
[0169] The algorithm training of this invention is implemented based on the PyTorch framework, using Python 3.8 as the programming language, and the hardware platform configuration is an Intel Core i7-13650HX processor and an NVIDIA GeForce RTX 4060 graphics card. Regarding the algorithm hyperparameter settings, the maximum number of training epochs for OA-TD3 is set to 3000, the soft update coefficient is 0.005, the batch size is 256, and the experience replay buffer capacity is 10000. The learning rates for the actor network and the critic network are set to 0.003 and 0.0003, respectively, to ensure stable convergence during the policy evaluation and policy improvement processes.
[0170] To verify the effectiveness and performance of the proposed OA-TD3 algorithm, the TD3 algorithm and the DDPG (deep deterministic policy gradient) algorithm were selected for comparison. All algorithms were trained on the same electric-hydrogen-carbon-methanol multi-energy system optimization scheduling model and compared under the same training environment and hyperparameter configuration. The change curve of the reward function during training is shown in the figure. Figure 3 As shown.
[0171] In terms of convergence speed, the DDPG algorithm exhibits the fastest initial convergence rate, stabilizing around 1000 training rounds. However, its final convergence value is significantly lower than the other two algorithms, indicating relatively poor performance. The TD3 algorithm converges slightly slower than DDPG, reaching a stable state after approximately 2000 training rounds; however, its convergence value is significantly improved, demonstrating superior policy optimization capabilities. The OA-TD3 algorithm shows the most rapid increase in its reward function during the early stages of training, stabilizing around 1500 rounds. Its convergence speed is significantly better than the TD3 algorithm, and its final reward value is also higher, indicating that it can learn a more effective scheduling policy. Regarding convergence stability, while the DDPG algorithm converges quickly, it exhibits large fluctuations after reaching a steady state, resulting in poor stability. The TD3 algorithm shows significantly improved stability compared to DDPG, while the OA-TD3 algorithm further improves convergence stability, performing best among the three.
[0172] The TD3 algorithm optimizes the DDPG algorithm by introducing a dual-Q network and adding a delayed policy update mechanism and a target policy smoothing mechanism. Therefore, it significantly outperforms DDPG in terms of convergence speed and stability. The OA-TD3 algorithm, building upon TD3, introduces an orthogonal attention mechanism into its Critic network. By applying orthogonal constraints, it effectively decouples features, reduces feature redundancy, and recalibrates different features, enabling the Critic network to more accurately evaluate the value of state-action pairs. This accelerates convergence while guiding the policy to a better-performing local optimum. Simultaneously, the orthogonal regularization term, as a powerful inductive bias, constrains the model's optimization space, reduces random fluctuations during training, and decreases sensitivity to hyperparameter perturbations, resulting in better stability.
[0173] Based on the above optimal parameters, the output of each unit during the algorithm's solution process is as follows: Figure 4As shown in the figure, wind and solar power output exhibits typical intermittent and fluctuating characteristics. Solar power output is mainly concentrated between 9:00 and 18:00, with a sustained output above 3 million kilowatts from 11:00 to 16:00. Wind power output is more prominent at night and in the early morning, especially maintaining a high level from 1:00 to 3:00 AM and from 20:00 to 24:00 PM. Electrolyzers utilize wind and solar power to produce hydrogen. Hydrogen production capacity increases significantly during periods of high wind and solar output, maintaining above 2.5 million kilowatts during 1:00 to 3:00 AM, 11:00 to 16:00 AM, and 23:00 to 24:00 PM, effectively absorbing wind and solar power output. Conversely, when wind and solar output is low, the electrolyzers maintain a lower operating level to ensure stable load delivery. During peak wind and solar output periods, surplus electricity is used for carbon capture in thermal power units and for charging energy storage devices. If there is still a surplus after load transmission and methanol synthesis, wind and solar power are curtailed, which often occurs during periods of high wind and solar output. Thermal power units, as a flexible system adjustment resource, significantly increase net output during periods of insufficient renewable energy output (7:00–10:00 and 18:00–20:00) to ensure power balance. During periods of significant renewable energy shortage, such as 8:00–9:00, energy storage devices discharge to supplement the power deficit. During peak wind and solar output periods, thermal power units reduce output to avoid over-generation and increased carbon emissions. The methanol synthesis unit operates synchronously with the supply of green hydrogen and captured carbon dioxide, maintaining high output during peak wind and solar output periods and suspending operation during periods of insufficient wind and solar output, such as 8:00–9:00. The results fully demonstrate the effectiveness of the proposed model framework.
[0174] As demonstrated by the above embodiments, this invention proposes a low-carbon scheduling method for the electricity-hydrogen-carbon-methanol system in the desert region based on deep reinforcement learning. This method constructs a multi-energy coupled system model of electricity-hydrogen-carbon-methanol, introduces a flexible carbon capture mechanism and carbon resource utilization path, and achieves coordination between energy supply and demand and efficient absorption of new energy sources. Algorithmically, an OA-TD3 algorithm integrating orthogonal attention mechanism and TD3 algorithm is proposed. Feature decoupling and redundancy suppression improve the algorithm's convergence speed and stability, adapting to highly uncertain scheduling environments. This design effectively improves the system's new energy absorption capacity, operational economy, and low-carbon performance, while enhancing the algorithm's robustness and convergence performance in complex scheduling scenarios, providing an efficient and stable optimized scheduling solution for green energy bases in the desert region.
Claims
1. A low-carbon scheduling method for the Shago wasteland electricity-hydrogen-carbon-methanol system based on DRL, characterized in that, Includes the following steps: Step 1: Construct a multi-energy coupled system model of electricity-hydrogen-carbon-methanol, specifically including a flexible carbon capture model, an electrolyzer and hydrogen buffer tank model, a methanol synthesis model, and an electrochemical energy storage model; Step 2: Construct an objective function to minimize the total system operating cost, which includes the fuel cost of thermal power units, the cost of carbon capture solvent loss, the cost of carbon sequestration, the cost of hydrogen production by water electrolysis, the cost of energy storage, the cost of hydrogen storage, the revenue from methanol, the cost of wind and solar curtailment, and the cost of carbon emissions. Step 3: Establish system operation constraints, including wind and solar power output constraints, hydrogen storage balance constraints, carbon capture system constraints, water electrolysis hydrogen production and storage constraints, and electrochemical energy storage device constraints. Step 4: Model the system scheduling problem as a Markov decision process, in which the system scheduling center is the agent, the electric-hydrogen-carbon-methanol system is the environment, and the state space, action space and reward function are defined. Step 5: Solve the Markov decision process using a dual-delay deep deterministic policy gradient algorithm based on orthogonal attention mechanism to obtain the optimal scheduling policy; In step 1, the modeling of each part is as follows: 1) Flexible carbon capture model: ; In the formula, For carbon capture units in Total power generation during the period; This refers to the net output power of the carbon capture unit; Fixed energy consumption for carbon capture units; For carbon capture units Energy consumption during different time periods; The overall carbon emission level of the unit is related to the different power generation outputs of the carbon capture unit, as shown in the following formula: ; In the formula, For thermal power units Total generation of time period quantity; For the unit of power generation output quantity; The expression that distinguishes between the carbon dioxide being processed by the absorption tower and the carbon dioxide being processed by the regeneration tower is: ; ; In the formula, For absorption tower The time period is being processed quantity; This is the ratio of the amount of flue gas diverted into the carbon capture system to the total amount of flue gas generated on the power generation side. For carbon capture and absorption efficiency; For the collection unit Energy consumption; For regeneration tower The time period is being processed quantity; Due to the presence of the storage tank, the regeneration tower processes... quantity Satisfy the following formula: ; In the formula, for Periodic flow into the rich liquid tank The amount, A negative value indicates that the solution flows from the rich liquid tank to the regeneration tower, while a positive value indicates that the solution flows from the absorption tower into the rich liquid tank. When present in solution as a compound, the following relationship holds: ; In the formula, and The first The amount of solution flowing into the lean fluid storage tank during the time period and the first The amount of solution flowing into the rich liquid storage tank during a given time period; a positive value indicates inflow, and a negative value indicates outflow. and Representing ethanolamine solution and molar mass; The regeneration resolution coefficient; This represents the mass percentage concentration of ethanolamine. The density of the ethanolamine solution; and The first The storage capacity of the low-lying fluid storage tank and the storage capacity of the high-lying fluid storage tank during the time period; 2) Model of electrolyzer and hydrogen buffer tank: The hydrogen production capacity of an alkaline electrolyzer is directly proportional to the input power of the electrolyzer, as shown below: ; ; In the formula, express The molar flow rate of hydrogen produced during a given period; This indicates the hydrogen production efficiency of the electrolyzer; This indicates the current input power of the electrolytic cell; This indicates the calorific value of hydrogen. express The quality of hydrogen produced during a given period; Indicates the molar mass of hydrogen; A hydrogen buffer tank is also installed, through which green hydrogen enters the methanol synthesis unit, serving as a buffer and storage mechanism; the hydrogen buffer model is represented as follows: ; ; ; In the formula, Indicates in Hydrogen storage capacity in the hydrogen storage tank during a given period; express The amount of hydrogen entering the storage tank during a given period is the amount of green hydrogen produced. express The outflow rate of hydrogen storage tank during a given period; express The amount of hydrogen consumed at any given moment; 3) Methanol synthesis model: The mass relationships between methanol and its compounds are as follows: ; In the formula, , These represent the mass ratios of hydrogen and carbon dioxide consumed to produce one unit of methanol, respectively. express The amount of methanol produced at any given time; express The amount of carbon dioxide consumed at any given moment; express The power consumed in the production of methanol at any given time; This represents the power consumed to produce one unit mass of methanol. 4) Electrochemical energy storage model: Treating the electrochemical energy storage inherent in wind and solar power as a whole, its mathematical state model can be written as follows: ; In the formula, For electrochemical energy storage Capacity of a time period; , These represent the battery charging and discharging efficiencies, respectively. , For energy storage batteries Charging and discharging power during a given time period; This represents a time period within the scheduling cycle.
2. The low-carbon scheduling method for the Shago wasteland electricity-hydrogen-carbon-methanol system based on DRL according to claim 1, characterized in that, In step 2, the objective function As shown in the following formula: ; In the formula, For system operating costs; Costs associated with wind and solar power curtailment; For the system's carbon emission costs; For the entire scheduling cycle; 1) System operating costs: The system operating costs include the fuel costs of thermal power units. Carbon capture solvent loss cost Carbon sequestration costs Cost of hydrogen production by water electrolysis Hydrogen storage cost Energy storage usage costs and methanol revenue The model is as follows: ; ; ; ; ; ; ; ; In the formula, , and This is the power generation cost coefficient for thermal power units; This is the cost coefficient for ethanolamine solvent; This is the solvent operating loss coefficient; Absorbed by ethanolamine The amount; as a unit Cost coefficients for transportation and storage; Consumption for methanol production The amount; This is the cost coefficient for hydrogen production via water electrolysis; The hydrogen storage cost coefficient for hydrogen buffer tanks; This refers to the energy storage charging and discharging cost coefficient. This refers to the unit price at which methanol is sold. This represents the unit cost of methanol production. 2) Costs of wind and solar power curtailment: The curtailment of wind and solar power is transformed into part of the economic cost through a penalty factor, which is defined as follows: ; ; ; In the formula, and These are the costs of wind curtailment and solar curtailment, respectively. and These are respectively the wind curtailment penalty factor and the solar curtailment penalty factor; and Each is the original power output of the wind and solar power system; and Each contributes to the actual scenery; 3) Carbon emission costs: The entities included in carbon quota management are carbon capture power plants, whose carbon quotas are allocated based on the amount of electricity they supply. Defined as follows: ; In the formula, The benchmark value for carbon emissions per unit of electricity used by coal-fired power units; The carbon emission cost is defined as: ; In the formula, The unit carbon price; for The time system discharges outward The total amount.
3. The low-carbon scheduling method for the Shago wasteland electricity-hydrogen-carbon-methanol system based on DRL according to claim 2, characterized in that, In step 3, the constraint condition is as follows: 1) Constraints on wind and solar power output: To meet the requirements for green hydrogen production, thermal power units do not participate in hydrogen synthesis; hydrogen production power is provided solely by wind, solar, and energy storage. Wind and solar power output participates in hydrogen production as well as carbon capture and regeneration, as shown in the following formula: ; In the formula, To contribute to wind and solar power participation in carbon capture and recycling; For wind and solar power output to participate in load transmission; 2) Hydrogen storage balance constraints: The balance between the inflow and outflow of hydrogen storage tanks within a scheduling cycle is defined by the following formula: ; 3) Constraints of carbon capture systems: The output range and ramping constraints of the carbon capture unit are defined as follows: ; In the formula, and These represent the upper and lower limits of the carbon capture unit's output, respectively. This is the upper limit of the ramping power of the carbon capture unit; The carbon capture storage tank has volume constraints, and the storage capacity of the solution storage must be kept consistent at the beginning and end of the scheduling day to prepare for optimized operation on the next scheduling day, as expressed by the following formula: ; In the formula, , These are the minimum storage volumes for the lean and rich liquid tanks, respectively. , These are the maximum liquid storage volumes of the lean and rich liquid tanks, respectively. , These represent the volumes of the lean and rich liquid tanks at the start of the scheduling cycle. , These are the volumes of the lean and rich liquid tanks at the end of the scheduling cycle, respectively. 4) Constraints on hydrogen production and storage via water electrolysis: The electrolysis power output range of the electrolyzer is constrained, and the capacity of the hydrogen buffer tank is constrained, as shown in the following model: ; In the formula, and These are the upper and lower limits of the electrolytic cell power, respectively; and These are the upper and lower limits of the hydrogen storage tank capacity, respectively. and These represent the hydrogen storage capacity at the beginning and end of the scheduling cycle, respectively. 5) Constraints of electrochemical energy storage devices: ; In the formula, , These are the upper and lower limits of energy storage capacity, respectively. and These are Boolean variables, representing respectively Charging and discharging cannot be performed simultaneously during a given period. , These are the upper limits for charging power and discharging power, respectively.
4. The low-carbon scheduling method for the Shago wasteland electricity-hydrogen-carbon-methanol system based on DRL according to claim 3, characterized in that, In step 4, the state space, action space, and reward function of the Markov decision process are represented as follows; 1) State space: During the training process of the electricity-hydrogen-carbon-methanol system, the information provided by the environment to the agent includes the original output of the wind turbine and photovoltaic system, fluctuating load, the status of the carbon capture device's storage tank, energy storage status, and scheduling time, as shown in the following formula: ; In the formula, express The load status of the entire system during the specified time period; 2) Motion space: The action space includes actual wind and solar power output, net output power of carbon capture units, energy consumption of carbon capture units, wind and solar power output participating in carbon capture and regeneration, hydrogen production output from electrolyzers, split ratio, and energy storage output, modeled as follows: ; In the formula, for The output of energy storage during different time periods; 3) Reward function: In the objective function Based on this, a penalty term exceeding system constraints is added, including a penalty for power imbalance. Penalties for exceeding the climbing limit Penalties for exceeding capacity limits for each device in the system , means as follows: ; In the formula, , , and It is a weighting factor that balances the various items.
5. The low-carbon scheduling method for the Shago wasteland electricity-hydrogen-carbon-methanol system based on DRL according to claim 4, characterized in that, In step 5, the orthogonal attention mechanism is represented as follows; Orthogonal attention module input feature vector It is expressed as follows: ; In the formula, Represents the set of real numbers. Indicates the dimension of the input feature vector; Attention weights are generated through a dual mapping process and calibrated against the input features; firstly, through the weight matrix... The input features are projected onto a low-dimensional latent space, as follows: ; ; In the formula, The compressed feature vector; This represents the dimension of the compressed feature vector; It is an activation function used to set the part of the input that is less than 0 to 0, while leaving the part that is greater than or equal to 0 unchanged; for Transpose of; It is the identity matrix; Then, another set of approximately orthogonal matrices Mapping back to the original dimension: ; ; In the formula, This is the attention vector, representing the modeling of the importance of each channel of the input feature; This represents the Sigmoid activation function; for Transpose of; Finally, element-wise weighting is applied to the input channels to achieve adaptive recalibration of features, modeled as follows: ; In the formula, This represents the recalibrated features. This represents element-wise multiplication; By employing soft orthogonal constraints, the following regularization term is introduced into the total loss function. : ; in, This is the regularization coefficient, used to balance the strength of the error loss between the model prediction and the true value and the orthogonal regularization loss; This represents the Frobenius norm.
6. The low-carbon scheduling method for the Shago wasteland electricity-hydrogen-carbon-methanol system based on DRL according to claim 5, characterized in that, In step 5, the solution process is as follows: Step 5.1: Network parameter initialization; The dual-delay deep deterministic policy gradient algorithm based on orthogonal attention mechanism includes a policy network, a value network composed of two orthogonal attention mechanisms, and their corresponding target policy network and target value network; the network parameters are as follows: policy network parameters Parameters of Value Network 1 Parameters of Value Network 2 Target policy network parameters Parameters of Target Value Network 1 Parameters of Target Value Network 2 Initialize the above network parameters, initialize the experience storage buffer, and initialize the value network weight matrix. And set the orthogonal regularization coefficient. ; Step 5.2: Environmental interaction and data sampling; In an electric-hydrogen-carbon-methanol system environment, the agent acts according to the current strategy. With exploring noise selection action The environment returns to the next state. and rewards and will Store in the experience replay buffer; Step 5.3: Batch sampling and value network update; Random sampling from the experience playback buffer Batch data The next action is generated through the target policy network. ,in Smooth the noise to the target; It is a cutoff function that limits the noise value to the interval middle, This indicates that the mean is 0 and the standard deviation is 0. The normal distribution; These are the upper and lower limits for truncation; The target value is calculated using the target value network, as shown in the following formula: ; In the formula, for Target value for the time period; In the state Take action below The reward value obtained; Discount factor; The values are 1 and 2, corresponding to network 1 and network 2 respectively; For the first Parameters of a target value network; For the first A target value network in state Take action below The target value; Minimize the total loss function of the value network To update the network parameters, the model is as follows: ; ; ; In the formula, N is the total number of batches; , and They represent the first The status, actions, and target values of each batch; For the first A target value network in state Take action below The target value; For the first Parameters of a value network; For value network hyperparameters; This represents the gradient of the total loss function. The differential symbol; Step 5.4: Policy network update; Following a delayed update mechanism, the policy network parameters are updated every set number of steps. The policy is improved by maximizing the value network's evaluation of the action. The model is as follows: ; ; in, To determine the expected reward that can be obtained by taking the corresponding action distribution. For expected return The gradient; Indicates the state Calculate the gradient of the policy network below; Representing state With action The gradient of value network 1 is then calculated. For policy network hyperparameters; Step 5.5: Soft update the target network; Use a soft update strategy to gradually update the target network parameters: ; in, This is the soft update coefficient, used to ensure network stability through a soft update strategy; Step 5.6: Iterate training until convergence; The system continuously interacts, samples, and trains until the system scheduling target converges or the set number of training rounds is reached.
7. A low-carbon scheduling system for the Shago wasteland electricity-hydrogen-carbon-methanol system based on DRL, implemented based on the method of claim 1, characterized in that, include: The system includes a modeling module, an objective and constraint module, an algorithm framework module, and an optimization and scheduling module. The system modeling module is used to construct an electric-hydrogen-carbon-methanol multi-energy coupled system model that includes wind power generation, photovoltaic power generation, carbon capture thermal power units, water electrolysis hydrogen production, methanol synthesis, electrochemical energy storage and hydrogen buffer tanks. The objective and constraint module is used to minimize the total system operating cost as the objective function and to set system operating constraints. The total operating cost includes the fuel cost of thermal power units, the cost of carbon capture solvent loss, the cost of carbon sequestration, the cost of hydrogen production by water electrolysis, the cost of energy storage, the cost of hydrogen storage, the revenue from methanol, the cost of wind and solar curtailment, and the cost of carbon emissions. The constraints include constraints on wind and solar power output, hydrogen storage balance, carbon capture system, hydrogen production and storage by water electrolysis, and electrochemical energy storage devices. The algorithm framework module is used to establish a dual-delay deep deterministic policy gradient algorithm framework based on Markov decision process and embedding orthogonal attention mechanism. The Markov decision process takes the system scheduling center as the agent and the electric-hydrogen-carbon-methanol system as the environment, and defines the state space, action space and reward function. The optimization scheduling module is used to optimize the scheduling of the electric-hydrogen-carbon-methanol multi-energy coupled system model based on the orthogonal attention mechanism and the dual-delay deep deterministic strategy gradient algorithm framework, and outputs the optimal output strategy of each device to achieve coordinated optimization of system economy and low carbon emissions.
Citation Information
Patent Citations
Hydrogen-containing comprehensive energy system low-carbon economic dispatching method based on near-end strategy optimization algorithm
CN120525223A
KR20250101169A