A green certificate and ccer coupled power dispatch decision optimization method and device
By combining LSTM and BP neural networks with multi-agent reinforcement learning, a multi-dimensional information fusion system for the electricity-carbon-certificate market is constructed. This solves the problems of dynamic characteristics and strategic game effects in multi-market coupled power dispatch, and realizes intelligent collaborative decision-making and market-based trading in power dispatch.
Patent Information
- Application Number
- CN202511332146.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-09-18
AI Technical Summary
Existing technologies are ill-suited to the dynamic nature of power dispatch decisions in multi-market coupled systems such as the electricity market, carbon market, and green certificate market. They are unable to capture price fluctuations and the strategic game effects among market participants in real time, leading to dispatch decision failures and the inability to quantify market risks.
By employing an LSTM prediction model and a BP neural network combined with multi-agent reinforcement learning, a multi-dimensional information fusion system for the electricity-carbon-certificate market is constructed. Through predicting price trends and optimizing decision-making, optimal power dispatch for power generation companies is achieved.
It enables intelligent collaborative decision-making in power dispatch, improves forecast accuracy and market adaptability, supports flexible decision-making under high-frequency fluctuations, and promotes the market-based trading of renewable energy consumption and carbon emission reduction.
Smart Images

Figure CN120822806B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of new energy power dispatching, and in particular to a green certificate and CCER coupled power dispatching decision optimization method and device. BACKGROUND
[0002] The prior art has made some progress in the power dispatching decision related to the electricity market, carbon market and green certificate market. However, the prior art has obvious deficiencies in dealing with the power dispatching decision of the multi-market coupling of the electricity market, carbon market and green certificate market.
[0003] Firstly, the existing dispatching decision model is mostly a static optimization model, which is difficult to adapt to the dynamic characteristics of multi-market linkage, cannot capture the complex dynamic coupling relationship between the fluctuation of electricity spot price, the change of carbon quota supply and demand and the liquidity of the green certificate market in real time, and the dispatching decision will fail when the market mutates. Secondly, the existing technology generally ignores the strategy game effect between market subjects. The decision-making behavior of multi-subjects such as power generation enterprises, power selling companies and carbon asset management institutions based on their own maximum benefit will influence each other through market clearing mechanism, forming a complex Nash equilibrium, and the current model cannot quantify the market risk brought by the interaction of multiple subjects. In addition, although artificial intelligence such as deep learning has made progress in single market prediction, there are still technical limitations in multi-market coordination and multi-subject modeling. SUMMARY
[0004] Therefore, the embodiment of the present application provides a green certificate and CCER coupled power dispatching decision optimization method and device, aiming to realize the optimal power dispatching decision of power plant in the multi-market and green certificate-CCER coupling scene.
[0005] To achieve the above-mentioned purpose, according to the first aspect of the embodiment of the present application, a green certificate and CCER coupled power dispatching decision optimization method is provided, comprising:
[0006] obtaining historical data of a plurality of price influencing factors of the electricity-carbon-certificate market, dispatching information data of a plurality of power plant, and coupling relationship between green certificate and CCER;
[0007] According to the historical data of each of the price influencing factors of the electricity-carbon-certificate market, a pre-trained LSTM prediction model is used to predict the prediction data of each of the price influencing factors in the future T time periods;
[0008] According to the historical data and the prediction data of each of the price influencing factors, a pre-trained BP prediction model is used to predict the prediction data of various prices of the electricity-carbon-certificate market in the future T time periods, the prices including thermal power generation price, green power generation price, CCER price, carbon emission quota price and green certificate price;
[0009] Using the predicted price data of various electricity-carbon-certificate markets, the dispatch information data of power generation companies, and the coupling relationship between green certificates and CCERs, the state space, action space, and reward function of power generation companies in the electricity-carbon-certificate market are constructed, and the optimal power dispatch decision of power generation companies is obtained by iterative optimization using a deep Q-network algorithm.
[0010] Furthermore, the method also includes a step of pre-training the LSTM prediction model for each of the price influencing factors, specifically:
[0011] The first training sample set is constructed by obtaining historical data of the price influencing factors;
[0012] Initialize the network parameters in the LSTM neural network;
[0013] The input data of each sample in the first training sample set is input into the LSTM neural network to obtain the predicted value, and then compared with the true value of the sample and the loss function value is calculated.
[0014] Gradient descent is used to update the network parameters, and iterative training is performed.
[0015] When the loss function is less than a first preset value, the iteration stops, the network parameters of the LSTM neural network are determined, and the trained LSTM prediction model is obtained.
[0016] Furthermore, the loss function value of the LSTM network is calculated according to the following formula:
[0017] ;
[0018] ;
[0019] ;
[0020] in, This represents the loss function value during the s-th iteration of training. Indicates the training of the s-th iteration. The orientation factor of each sample Indicates the training of the s-th iteration. The nth sample pair Influence factors of each network parameter; It is a constant. , For the s-th iteration, train the th... Predicted values for each sample For the first The true value corresponding to each sample; This indicates that the s-th iteration of training the LSTM neural network represents the... Network parameters, It is the curvature parameter trained in the s-th iteration. sinh represents the hyperbolic sine function, and sign represents the sign function. This represents the inverse hyperbolic tangent function.
[0021] Furthermore, the network parameters in the LSTM network include the weights and biases of its input gate, forget gate, cell unit, and output gate;
[0022] The network parameters are updated using gradient descent, including:
[0023] ;
[0024] ;
[0025] in, For learning rate, This indicates that the LSTM network is trained in the (s+1)th iteration. Network parameters.
[0026] Furthermore,
[0027] ,
[0028] .
[0029] Furthermore, the method also includes a step of pre-training the BP prediction model, specifically:
[0030] Obtain historical data of the price and its corresponding price influencing factors, and construct a second training sample set;
[0031] Initialize the network parameters in the BP neural network;
[0032] The input data of each sample in the second training sample set is input into the BP neural network to obtain the predicted value, and then compared with the true value of the sample and the loss function value is calculated.
[0033] Gradient descent is used to update the network parameters, and iterative training is performed.
[0034] When the loss function is less than the second preset value, the iteration stops, the network parameters in the BP neural network are determined, and the trained BP prediction model is obtained.
[0035] Furthermore, for each training iteration, the loss function value of the BP neural network is calculated according to the following formula:
[0036] ,
[0037] ,
[0038] ,
[0039] in, Represents the total loss function. , The first loss function is respectively and the first loss function Adaptive weights, For the second training sample set The true value of each sample For the first One predicted value, This represents the number of samples in the second training sample set.
[0040] Furthermore, after each iteration of training the BP neural network, it is updated according to the following formula. , :
[0041] ,
[0042] ,
[0043] in, The first loss function at the t-th iteration Adaptive weights of the function The first loss function at the t-th iteration The function value, The second loss function at the t-th iteration The function value; The first loss function at the (t+1)th iteration Adaptive weight values, The second loss function at iteration t+1. The adaptive weight value.
[0044] Furthermore, the dispatch information data of the power generation companies includes green electricity generation, load, green certificate demand, CCER demand, carbon quota demand, and risk preference coefficient; the green electricity includes photovoltaic power generation and wind power generation; the coupling relationship between green certificates and CCERs is that green certificates and CCERs are exchanged proportionally; the state space, action space, and reward function of each power generation company in the electricity-carbon-certificate market for multi-agent reinforcement learning are independent;
[0045] The state space of each of the aforementioned power generation companies is represented as follows:
[0046] ,
[0047] in, Let the state space be... For the price of photovoltaic power generation, For the price of wind power generation, For load, and These are wind power generation and photovoltaic power generation, respectively. For the price of green certificates, To meet the demand for green certificates, For CCER price, For CCER demand, This refers to the carbon allowance requirement. The price of carbon allowances. This is the risk preference coefficient; the subscript 't' indicates that it is related to the time period 't'.
[0048] The action space is represented as follows:
[0049] ,
[0050] in, For the action space, These refer to the volume of electricity traded. These represent the number of green certificates traded. These represent the trading quantities of CCER. The exchange rate for green certificates to CCERs;
[0051] The reward function is a combination of rewards from multiple sub-tasks of the power generation company, expressed as follows: ,in, It is the first The weight of each subtask It is the first Rewards for each sub-task This refers to the status of the power generation company. This refers to the actions of the power generation company. .
[0052] Further, the price influencing factors of the thermal power generation price include crude oil futures price, natural gas futures price, national coal price, thermal power installed capacity and thermal power generation amount; the price influencing factors of the carbon quota price include carbon trading amount, monthly carbon futures price and thermal power generation price; the price influencing factors of the CCER price include carbon quota price; the price influencing factors of the photovoltaic power generation price include photovoltaic power generation installed capacity, photovoltaic power generation amount and sunshine hours; the price influencing factors of the wind power generation price include wind power generation installed capacity, wind power generation amount and wind speed; and the price influencing factors of the green certificate price include renewable energy consumption weight.
[0053] According to a second aspect of the embodiment of the present application, a green certificate and CCER coupled power dispatch decision optimization device is provided, comprising:
[0054] a data acquisition module configured to acquire historical data of a plurality of price influencing factors of an electricity-carbon-certificate market, dispatch information data of a plurality of power generation companies, and a coupling relationship between a green certificate and a CCER;
[0055] a second prediction module configured to, according to the historical data of each of the price influencing factors of the electricity-carbon-certificate market, predict prediction data of each of the price influencing factors in a future T time period by using a pre-trained LSTM prediction model;
[0056] a second prediction module configured to, according to the historical data and the prediction data of each of the price influencing factors, predict prediction data of various prices of the electricity-carbon-certificate market in the future T time period by using a pre-trained BP prediction model, the prices including thermal power generation price, green power generation price, CCER price, carbon emission quota price and green certificate price;
[0057] a dispatch optimization module configured to, by using the prediction data of various prices of the electricity-carbon-certificate market, the dispatch information data of the power generation companies, and the coupling relationship between the green certificate and the CCER, construct a state space, an action space and a reward function for multi-agent reinforcement learning of the power generation companies in the electricity-carbon-certificate market, and obtain an optimal power dispatch decision of the power generation companies by using a deep Q network algorithm for iterative optimization.
[0058] The embodiment of the present application has at least one of the following advantages or beneficial effects:
[0059] The embodiment of the present application realizes intelligentized collaborative decision of power dispatch in a deep integration scenario of green certificate, CCER market and electricity market by constructing a power dispatch decision optimization system based on combination of an LSTM prediction model, a BP neural network and multi-agent reinforcement learning.
[0060] The embodiment of the present application can comprehensively consider complex factors affecting transaction decision-making by collecting and processing various historical data variables such as electricity price, green certificate price, CCER price, carbon price, power generation, and power consumption, and using an LSTM model for multivariate time series prediction, thereby improving the comprehensiveness and accuracy of power dispatch decision-making. In addition, by combining a BP neural network to further fit the future market price trend, the accuracy of the prediction result is improved, thereby providing more comprehensive and detailed input support for subsequent power dispatch strategies.
[0061] The embodiment of the present application can iteratively optimize and update the model based on continuous reception of new data, and respond to dynamic changes in the external environment in real time. The embodiment of the present application can support small-granularity power dispatch prediction and strategy adjustment, meet the flexible decision-making needs in high-frequency fluctuating markets, and improve the adaptability of power dispatch subjects to market changes.
[0062] The embodiment of the present application uses a multi-agent reinforcement learning method to autonomously optimize transaction behavior. Each type of power plant (such as thermal power enterprises, green power generators, power users, and power grid companies) can learn and evolve individualized power dispatch strategies based on its independent state space, action space, and reward function. Through the training and optimization of the experience pool D and the deep Q network, the power dispatch decision-making is transformed from experience-driven to data-driven and strategy-adaptive, significantly improving the game ability and collaboration level of multi-agent in the coupled market environment.
[0063] The embodiment of the present application breaks down the information silos between traditional electricity markets and carbon trading markets, integrates green certificates and CCER into a unified power dispatch decision-making system, and realizes market linkage and collaborative optimization. The embodiment of the present application not only improves the allocation efficiency of electricity and carbon resources, but also promotes the marketization of renewable energy consumption and carbon emission reduction, serves the green energy development goal, and has good social benefits and application prospects.
[0064] The embodiment of the present application technically realizes multi-source data fusion prediction, intelligent transaction strategy optimization, and deep collaboration of green certificates and CCER with the electricity market, and constructs an intelligent, dynamic, and multi-agent decision support system that can adapt to the development trend of future energy markets. It has obvious advantages in improving prediction accuracy, enhancing strategy flexibility and market adaptability, and promoting innovation of green energy trading mechanisms, and has significant practical value and promotion potential.
[0065] The technical solutions in the present application can be combined with each other to realize more preferred combination solutions. Other features and advantages of the present application will be described in the subsequent description, and some advantages will become apparent from the description, or will be understood by those skilled in the art through implementation of the present application. The objects and other advantages of the present application can be realized and obtained through the contents particularly pointed out in the description and the drawings. BRIEF DESCRIPTION OF DRAWINGS
[0066] The drawings are only for the purpose of illustrating specific embodiments and are not considered as limiting the present application, and in the whole drawings, the same reference signs represent the same parts;
[0067] Figure 1 is a schematic diagram of the main process of a green certificate and CCER coupled power dispatch decision optimization method according to an embodiment of the present application;
[0068] Figure 2 is a network structure schematic diagram of an LSTM prediction model according to an embodiment of the present application;
[0069] Figure 3 is a network structure schematic diagram of a BP prediction model according to an embodiment of the present application;
[0070] Figure 4 is a composition schematic diagram of a green certificate and CCER coupled power dispatch decision optimization device according to an embodiment of the present application. DETAILED DESCRIPTION
[0071] The preferred embodiments of the present application will be specifically described below in combination with the drawings, wherein the drawings form a part of the present application and are used to explain the principles of the embodiments of the present application, and are not used to limit the scope of the present application.
[0072] Embodiment one
[0073] In the electricity-carbon-green certificate collaborative trading scenario, the electricity spot price fluctuation, the carbon quota supply and demand change and the green certificate market liquidity influence each other, forming a complex dynamic coupling relationship. The traditional single market optimization model such as the power market clearing algorithm or the simple linear linkage carbon price-power price transmission coefficient is difficult to capture this multi-dimensional dynamic correlation in real time, resulting in that when the market mutates, the carbon price fluctuates sharply or the green certificate policy adjustment decision fails. The strategy game effect between market subjects is generally ignored in the prior art-the decision behavior of multiple subjects such as power generation enterprises, power selling companies and carbon asset management institutions based on their own maximum benefit will influence each other through the market clearing mechanism, forming a complex Nash equilibrium, and the current static model cannot quantify the market risk brought by the interaction of multiple subjects.
[0074] Figure 1 is a schematic diagram of the main process of a green certificate and CCER coupled power dispatch decision optimization method according to an embodiment of the present application. As shown inFigure 1 As shown, the green certificate and CCER coupled power dispatch decision optimization method in this embodiment includes the following steps S101 to S104.
[0075] Step S101, obtaining historical data of a plurality of price influencing factors of an electricity-carbon-certificate market, dispatch information data of a plurality of power plant operators, and coupling relationship between green certificates and CCERs;
[0076] Step S102, according to the historical data of each of the price influencing factors of the electricity-carbon-certificate market, using a pre-trained LSTM prediction model to predict prediction data of each of the price influencing factors in the future T time periods;
[0077] Step S103, according to the historical data and the prediction data of each of the price influencing factors, using a pre-trained BP prediction model to predict prediction data of various prices of the electricity-carbon-certificate market in the future T time periods, the prices including thermal power generation price, green power generation price, CCER price, carbon emission quota price and green certificate price;
[0078] Step S104, using the prediction data of various prices of the electricity-carbon-certificate market, the dispatch information data of the power plant operators, and the coupling relationship between the green certificates and the CCERs to construct a state space, an action space and a reward function for multi-agent reinforcement learning of the power plant operators in the electricity-carbon-certificate market, and using a deep Q network algorithm to iteratively optimize to obtain an optimal power dispatch decision of the power plant operators.
[0079] It can be understood that in this embodiment and some embodiments of the present application, the objects involved in the power dispatch of the power plant operators include five types of resources, i.e. thermal power generation, green power generation (including centralized photovoltaic power generation and land-based wind power generation), CCER, carbon emission quota and green certificate. Each type of resource has its own price, and there is one or more price influencing factors affecting a price. Specifically, as shown in Table 1, the price influencing factors of the thermal power generation price include crude oil futures price, natural gas futures price, national coal price, thermal power generation capacity and thermal power generation amount; the price influencing factors of the carbon quota price include carbon trading volume, monthly carbon futures price and thermal power generation price; the price influencing factors of the CCER price include carbon quota price; the price influencing factors of the photovoltaic power generation price include photovoltaic power generation capacity, photovoltaic power generation amount and sunshine hours; the price influencing factors of the wind power generation price include wind power generation capacity, wind power generation amount and wind speed; and the price influencing factors of the green certificate price include renewable energy consumption weight.
[0080] Table 1
[0081]
[0082] Understandably, in this embodiment and some embodiments of the present invention, for a certain price in the electricity-carbon-certificate market, an LSTM prediction model is first used to predict the predicted data of each price influencing factor in the future T time period. Then, the historical data and predicted data of each price influencing factor are input into a BP prediction model to predict the predicted data of the various prices in the future T time period.
[0083] Furthermore, in this embodiment and some embodiments of the present invention, the pre-training step of the LSTM prediction model specifically includes steps S201 to S205.
[0084] Step S201: Obtain historical data of the price influencing factors to construct the first training sample set;
[0085] Step S202: Initialize the network parameters in the LSTM neural network;
[0086] Step S203: Input the input data of each sample in the first training sample set into the LSTM neural network to obtain the predicted value, compare it with the true value of the sample and calculate the loss function value;
[0087] Step S204: Use gradient descent to update the network parameters and perform iterative training;
[0088] Step S205: When the loss function is less than the first preset value, stop the iteration, determine the network parameters of the LSTM neural network, and obtain the trained LSTM prediction model; otherwise, return to step S203 to continue the iteration.
[0089] Specifically, in this embodiment and some embodiments of the present invention, historical data for each influencing factor in Table 1 above is collected on a monthly basis, and the historical data should generally include at least 120 months. A first training sample set is used to train the LSTM network, with 60% of the samples in this set serving as the training set and 40% as the test set. Each price influencing factor corresponds to one LSTM prediction model.
[0090] Specifically, such as Figure 2 As shown, an LSTM neural network is constructed. The LSTM neural network includes three gates: a forget gate, an input gate, and an output gate, as well as a cell unit.
[0091] (1) The forgetting gate is as follows:
[0092] ,
[0093] in, It is the output of the forget gate. It is the weight matrix of the output data of the forget gate. It is the bias vector of the output data of the forget gate. is the hidden state of the previous time, is the current input.
[0094] (2) The input gate determines which new information should be added to the cell state, the formula contains two parts: the activation value of the input gate, as follows:
[0095] ,
[0096] where, is the activation value, is the weight matrix of the activation value, is the bias vector of the activation value.
[0097] (3) The new state candidate value of the cell unit, as follows:
[0098] ,
[0099] where, is the new state candidate value, is the weight matrix of the new state candidate value, is the bias vector of the new state candidate value.
[0100] The updated value of the cell unit state, as follows:
[0101] ,
[0102] where, is the updated value of the cell state.
[0103] (4) The output gate decides which part of the cell state to output, as follows:
[0104] ,
[0105] where, is the output gate, is the weight matrix of the cell output state, is the bias vector of the cell output state.
[0106] The final output value, as follows:
[0107] ,
[0108] where, is the output value of the LSTM model.
[0109] Specifically, in this embodiment and some embodiments of the present invention, for the price influencing factors related to photovoltaic and wind power prices, in order to consider the locational influence relationship between wind turbines or solar panels and observation stations and reflect geographical differences, the forget gate weight coefficient is spatially weighted. The forget gate determines which information should be discarded. It is the first Latitude and longitude coordinates of a wind turbine or solar panel It is the first The latitude and longitude coordinates of a solar radiation or wind speed observation station It is the attenuation coefficient, for example, set to 20km.
[0110] ,
[0111] ,
[0112] in, The Hadamard product represents the element-wise multiplication of corresponding positions in a matrix.
[0113] Understandably, the network parameters in an LSTM neural network include the weights of its input gate, forget gate, cell unit, and output gate. and bias Specifically, in this embodiment and some embodiments of the present invention, the weights and biases are extracted from a standard normal distribution, with the weights set to 0.01 and the biases set to 0 during initialization.
[0114] For ease of description below, all weights that need to be updated are set. and bias Corresponding to The loss function of the LSTM network described in this embodiment and some embodiments of the present invention introduces the direction sensitivity coefficients of various network parameters. and curvature parameters , The loss function of the LSTM network is as follows:
[0115] ;
[0116] ;
[0117] ;
[0118] in, This represents the loss function value during the s-th iteration of training. Indicates the training of the s-th iteration. The orientation factor of each sample Indicates the training of the s-th iteration. The nth sample pair an influence factor of a network parameter; is a constant, , is a predicted value corresponding to the historical data in the first training sample set in the s-th iteration of training the LSTM network, is a true value corresponding to the historical data in the first training sample set in the s-th iteration of training the LSTM network; is a predicted value corresponding to the historical data in the first training sample set in the s-th iteration of training the LSTM network, is a true value corresponding to the historical data in the first training sample set in the s-th iteration of training the LSTM network; represents the s-th iteration of training the LSTM network, is a network parameter, is a curvature parameter in the s-th iteration of training, , is a direction-sensitive coefficient, and sinh represents a hyperbolic sine function, and sign represents a sign function, represents an inverse hyperbolic tangent function. is degenerated into a symmetric loss when a direction-sensitive penalty, i.e., a penalty for overestimation or underestimation, is achieved.
[0119] Specifically, in the embodiment and some embodiments of the application, the gradient of the loss function of the LSTM network includes:
[0120] ,
[0121] .
[0122] The gradient descent method is adopted to update the network parameters, including:
[0123] ;
[0124] ;
[0125] wherein, is a learning rate, which is set to 0.001, represents the s+1-th iteration of training the LSTM network, is a network parameter. When the iteration is stopped, the weight matrix and bias vector values of the input gate, the forget gate and the output gate at this time are determined, and a trained LSTM network is obtained.
[0126] Alternatively, in some other embodiments of the application, the determination that the loss function is less than the first preset value is achieved by determining whether the gradient of the curvature parameter is less than a second preset value (for example, set to 0.0001), and then the iteration is stopped, and the values at this time are obtained. .
[0127] It can be understood that, in the embodiment and some embodiments of the application, the time sequence corresponding to one historical data in the first training sample set is Input, get a predicted value of output . Each time series The predicted value of the LSTM training model is compared with the corresponding true value to construct a loss value function. Compared with the traditional MSE mean square error gradient descent method, the embodiment of the application first introduces a direction-sensitive coefficient , which realizes the punishment of overestimation or underestimation of the prediction result. Secondly, the function and are used to realize the stability smoothing processing of noise and outliers, avoid gradient explosion, and accelerate the critical point convergence. Finally, the curvature c of adaptive update is introduced in the loss function, which reflects the sensitivity to errors of different degrees. Since the price index has both noise and regularity, small errors can be allowed, but medium deviations should be avoided. Therefore, when is small, the sensitivity of the central region (small error) is reduced, the loss value grows slowly, the saturation speed of the edge region (large error) is slow, and the punishment degree will continue to grow for a longer distance as the error increases. The model is not so strict for small errors, but the punishment for medium and large errors is relatively more “linear” and persistent. Using the loss function of the embodiment of the application can make the prediction value of the future change of the price influencing factor more accurate and scientific.
[0128] Specifically, in the embodiment and some embodiments of the application, the trained LSTM model is used to predict the predicted value of the price influencing factor in the future T time periods. Taking the price influencing factor of the thermal power price, i.e., the crude oil futures price, as an example, assuming that the historical data of the price influencing factor is , the historical data is input into the trained LSTM model, the model outputs the first predicted value ; adds to the original data sequence , and removes the earliest input value to keep the length of the input sequence unchanged, to obtain the input of the second prediction LSTM model, the model outputs the second predicted value , and the same way is used to obtain the input ;…; the input of the tth prediction model , the model outputs the tth predicted value …, and finally the predicted value of the price influencing factor in the future T time periods is obtained . By analogy, the above steps are repeated for other price influencing factors of the thermal power price, and then the predicted value of each price influencing factor in the future T time periods is obtained.
[0129] Further, in the embodiment and some embodiments of the application, the pre-training of the BP prediction model comprises steps S301 to S305.
[0130] In step S301, historical data of the price and the corresponding price influencing factors are obtained to construct a second training sample set.
[0131] In step S302, network parameters in the BP neural network are initialized.
[0132] In step S303, input data of each sample in the second training sample set is input into the BP neural network to obtain a predicted value, which is compared with the true value of the sample to calculate a loss function value.
[0133] In step S304, gradient descent method is used to update the network parameters for iterative training.
[0134] In step S305, when the loss function is less than a second preset value, the iteration is stopped, the network parameters in the BP neural network are determined, and the trained BP prediction model is obtained.
[0135] Specifically, in the embodiment and some embodiments of the application, the BP neural network is trained by using the second training sample set to obtain the future variation trend of the thermal power generation price, the green power generation price, the CCER price, and the carbon quota price. Specifically, for example, the predicted values of all influencing factors of the thermal power generation price in the future T time periods are a set , the predicted values of all influencing factors of the green power generation price in the future T time periods are a set , the predicted values of all influencing factors of the CCER price in the future T time periods are a set , the predicted values of all influencing factors of the carbon emission quota price in the future T time periods are a set , and the predicted values of all influencing factors of the green certificate price in the future T time periods are a set When training the BP prediction model for predicting the thermal power generation price, the predicted values of all influencing factors of the thermal power generation price in the future T time periods are combined to form a set as the input of the second training sample set. In the embodiment and some embodiments of the application, the first 70% is used as the training set, and the last 30% is used as the test set.
[0136] Specifically, in the embodiment and some embodiments of the application, as shown in Figure 3 , the BP neural network comprises an input layer, a hidden layer, and an output layer. The hidden layer activation function is tan-sigmoid, as follows:
[0137] .
[0138] The activation function output range is [−1, 1], and has good gradient characteristics near the origin, which helps to accelerate the training process, and the output layer activation function is a purelin linear mapping, as follows:
[0139] .
[0140] The weights and biases of the BP neural network are extracted from the standard normal distribution In this embodiment and some embodiments of the present application, the weights are initialized to 0.01 and the biases are set to 0.
[0141] Specifically, in this embodiment and some embodiments of the present application, in the forward propagation process, the input data is transmitted through the network layer by layer until it reaches the output layer. The specific steps are as follows:
[0142] Input layer to hidden layer, as follows:
[0143] ,
[0144] ,
[0145] wherein, represents the input, is the weight matrix from the input layer to the hidden layer, is the bias vector of the hidden layer, is the net input of the hidden layer, is the activation output of the hidden layer.
[0146] Hidden layer to output layer, as follows:
[0147] ,
[0148] wherein, is the weight matrix from the hidden layer to the output layer, is the bias vector of the output layer, is the net input of the output layer, which is also the final output.
[0149] Specifically, in this embodiment and some embodiments of the present application, the loss function value of the BP neural network is calculated according to the following formula each time the training is performed:
[0150] ,
[0151] ,
[0152] ,
[0153] wherein, Represents the total loss function. , The first loss function is respectively and the first loss function Adaptive weights, For the second training sample set The true value of each sample For the first The predicted value, i.e. the th predicted value. Input in each sample The corresponding actual value and predicted value, The number of samples in the second training sample set. Adaptive weights for the two loss functions. , Initially, all values were set to 0.5. , This reflects the importance of the two loss terms in the current batch of data.
[0154] Furthermore, after each iteration of training the BP neural network, it is updated according to the following formula. , :
[0155] ,
[0156] ,
[0157] in, The first loss function at the t-th iteration Adaptive weights of the function The first loss function at the t-th iteration The function value, The second loss function at the t-th iteration The function value; The first loss function at the (t+1)th iteration Adaptive weight values, The second loss function at iteration t+1. The adaptive weight values are determined. After each iteration, the total loss value for that iteration is calculated. .
[0158] During each iteration of training the BP neural network, the weights and biases are updated as follows:
[0159] ,
[0160] ,
[0161] ,
[0162] ,
[0163] ,
[0164] .
[0165] wherein, is a dynamic learning rate, is an initial learning rate, , are the weight matrices of the input layer to the hidden layer before and after updating, respectively, , are the weight matrices of the hidden layer to the output layer before and after updating, respectively, is a bias vector of the hidden layer, is a bias vector of the output layer. is a hyperparameter for controlling the decay of the learning rate, and is initially set to 0.01. Exemplarily, the initial learning rate is set to 0.001.
[0166] Specifically, in the embodiment and some embodiments of the application, when the total loss value is less than a second preset value, the iteration is ended. Exemplarily, the second preset value is set to 0.01. When the iteration is ended, the weight matrix and the bias vector of the input layer to the hidden layer at this time, and the weight matrix and the bias vector of the hidden layer to the output layer are determined.
[0167] At this time, the trained BP prediction model can be used to predict the future trend value of the product thermal power generation price. Similarly, the , , , are sequentially used as the second sample set, and the corresponding BP prediction model is obtained by training to obtain the future trend value of the green electricity generation price, the CCER price, the carbon emission quota price, and the green certificate price.
[0168] It can be understood that, in the embodiment and some embodiments of the application, compared with the traditional MSE mean square error gradient descent method, during the training process of the BP neural network using the second training sample set, firstly, a segmented loss function is introduced. When the error is small, the square penalty is sensitive to small errors, and can quickly converge. When the error is large, the linear penalty has less effect on outliers, avoiding that the model is dominated by abnormal values. Secondly, a dynamic learning rate is introduced, and the learning rate value is dynamically updated according to each gradient descent, to speed up the learning convergence process.
[0169] Further, in the embodiment and some embodiments of the application, the LSTM prediction model and the BP prediction model can be nested and combined together, and the historical data of all price influencing factors of a certain price are input into the corresponding nested neural network to predict the future trend of the price.
[0170] Specifically, in the embodiment and some embodiments of the application, the scheduling information data of the power plant includes green power generation, load, green certificate demand, CCER demand, carbon quota demand, and risk preference coefficient; the green power includes photovoltaic power generation and wind power generation; the scheduling information data of the power plant is obtained by internal calculation of the power plant and is directly input into multi-agent reinforcement learning.
[0171] Specifically, in the embodiment and some embodiments of the application, green power can be exchanged for CCER, and green power can be exchanged for green certificates. The coupling relationship between green certificates and CCER is that green certificates and CCER are exchanged in proportion. The exchange ratio can be set as the power carbon dioxide emission factor of the region where the power plant is located. For example, if the power plant is located in the North China region, the average power carbon dioxide emission factor of the North China region in 2021 is 0.7120, so 1MWh of green power = 1 green certificate = 0.7120 CCER = 0.7120 carbon emission quota.
[0172] Specifically, in the embodiment and some embodiments of the application, the state space, the action space and the reward function of each power plant in the multi-agent reinforcement learning in the electricity-carbon-certificate market are independent.
[0173] Specifically, in the embodiment and some embodiments of the application, the state space of each power plant is represented as:
[0174] ,
[0175] wherein, is the state space, indicating the state of the power plant in the electricity-carbon-certificate market, is the photovoltaic power generation price, is the wind power generation price, is the load, and are the wind power generation and the photovoltaic power generation, respectively, is the green certificate price, is the green certificate demand, is the CCER price, is the CCER demand, is the carbon quota demand, is the carbon quota price, is the risk preference coefficient; the subscript t indicates the time period t.
[0176] Specifically, in the embodiment and some embodiments of the application, the action space is represented as:
[0177] ,
[0178] in, For the action space, These refer to the volume of electricity traded. These represent the number of green certificates traded. These represent the trading quantities of CCER. The exchange rate for green certificates to CCERs;
[0179] Specifically, in this embodiment and some embodiments of the present invention, the reward function is a combination of multiple sub-task rewards of the power generation manufacturer, expressed as follows: ,in, It is the first The weight of each sub-task reflects the level of importance that power generation companies attach to it. Set an initial value, which is in the range [0,1]. It is the first Rewards for each sub-task This refers to the status of the power generation company. This refers to the actions of the power generation company. In this embodiment of the invention, the task weights are either preset values or dynamically changed during the decision-making process.
[0180] It is understandable that in this embodiment and some embodiments of the present invention, the power generation manufacturer's status during time period t... Take action Receive a reward , According to the status Take action The status of the power generation companies mentioned later is , This will generate a set of experiences from each interaction between power generators and the electricity-carbon-certificate market. The experience data is stored in the experience replay pool. When the experience data stored in the experience replay pool meets the requirements for updating the DQN network parameters, a set of experiences is randomly sampled from the experience replay pool. The sampled experience is used to calculate the loss function and update the network parameters.
[0181] Specifically, in this embodiment and some embodiments of the present invention, a deep Q-network algorithm is used for iterative optimization to obtain the optimal power dispatch decision of the power generation manufacturer. Each interaction between the power generation manufacturer and the electricity-carbon-certificate market generates a set of experience... The experience data is stored in the experience replay pool. When the experience data stored in the experience replay pool meets the requirements for updating the DQN network parameters, a set of experiences is randomly sampled from the experience replay pool. The sampled experiences are used to calculate the loss function and update the network parameters. Specifically, this includes steps S401 to S407.
[0182] Step S401, initialize the parameters of the Q network in the DQN network and the parameters of the target Q network , and let ; set the training round ;
[0183] Step S402, set the current step number ,
[0184] Step S403, according to the state of the power plant merchant , select an action and its corresponding Q value using the Q network in a greedy strategy , execute the action to obtain a reward , the state of the power plant merchant becomes , and store the experience in the experience pool; represents the corresponding Q value obtained by selecting an action according to the state and using the Q network with parameters ;
[0185] Step S404, calculate the target Q value according to the following formula :
[0186] ,
[0187] wherein represents the corresponding Q value obtained by selecting an action according to the state and using the target Q network with parameters ; is a discount factor, representing the weight of future rewards, used to quantify the consideration of the next time market state when the power plant merchant makes decisions.
[0188] Step S405, determine whether the number of experiences stored in the experience pool meets or exceeds the pre-set number threshold , in response to the determination result being met, randomly select experiences, update the parameters of the Q value network according to the experiences and the loss function , specifically including:
[0189] The loss function is
[0190] ,
[0191] The parameters are updated using the gradient descent method ,
[0192] ;in, The learning rate;
[0193] Step S406, Q network each After the next update The parameters of the Q network Copy to the target Q network, i.e., each Q network After the next update, ;
[0194] Step S406: Determine if the loss function is less than the third threshold. If it is not less than the third threshold, then... If the condition is not met, return to step S403; otherwise, proceed to step S407.
[0195] Step S407: Let If not satisfied , If the maximum number of training epochs is reached, return to step S402; otherwise, end the training. Parameters are obtained after training. The optimal value The set of optimal decision values for time period t corresponding to the value. .
[0196] Understandably, by repeating steps S401 and S407, the optimal decision for each time period can be obtained in turn, thereby enabling power generation companies to make decisions that take into account the impact of the future electricity-carbon-certificate market environment on the current dispatch decision.
[0197] Understandably, in this embodiment and some embodiments of the present invention, the reward for time period t... Similarly, the reward corresponding to the subtask function can also be used. We get the weighted sum, that is This will generate a set of experiences from each interaction between power generators and the electricity-carbon-certificate market. The experience data is stored in the experience replay pool. When the experience data stored in the experience replay pool meets the requirements for updating the DQN network parameters, a set of experiences is randomly sampled from the experience replay pool. The sampled experience is used to calculate the loss function and update the network parameters and the weights of the subtask rewards.
[0198] Specifically, in this embodiment and some embodiments of the present invention, the reward function of the power manufacturer is a combination of rewards for its dispatch sub-tasks in the electricity market, carbon market, and green certificate market, specifically... Sub-task rewards Represents revenue from the electricity market, sub-task rewards. This represents the cost of purchasing carbon allowances in the carbon market or the revenue from selling carbon allowances; sub-task rewards. This indicates the cost of purchasing green certificates in the green certificate market or the revenue from selling green certificates. the reward corresponding to the subtask, i.e., the reward of the time period t the reward corresponding to the three subtask functions , , the weighted sum, i.e., The deep Q network algorithm is used in these embodiments to iteratively optimize the optimal power dispatching decision of the power plant supplier, which specifically includes steps S501 to S508.
[0199] Step S501, initialize the parameters of the Q network in the DQN network , the parameters of the target Q network , and the weights of the subtask rewards , for example , is the number of subtasks, and let ; set the training round ;
[0200] Step S502, set the current step number ;
[0201] Step S503, according to the state of the power plant supplier , the Q network is used to select the action and the corresponding Q value and the Q value of each subtask , for example, the Q value of the dispatching subtask of the above-mentioned power market, carbon market, and green certificate market , , ; execute the action to obtain the reward and the subtask reward , the state of the power plant supplier becomes , and the experience is stored in the experience pool; , represents the total task and subtask corresponding Q value obtained by selecting the action according to the state and using the Q network with the parameters ;
[0202] Step S504, calculate the total task and subtask target Q value according to the following formula and :
[0203] ,
[0204] ,
[0205] wherein, , Indicates based on state And use parameters as Target Q network selects action The corresponding Q values for the total task and subtasks are obtained; It is a discount factor, representing the weight of future rewards, used to quantify the considerations of power generation companies regarding the market state at the next time point when making decisions.
[0206] Step S505: Determine whether the number of experiences stored in the experience pool is greater than or equal to the pre-approval threshold. If the judgment result is satisfied, then a random selection is made. One experience, ,according to The group describes the use of experience and loss functions to update the parameters of the Q-value network. ,
[0207] The total loss function is:
[0208] ,
[0209] Update parameters using gradient descent ,
[0210] ;in, The learning rate;
[0211] Furthermore, in this embodiment and some embodiments of the present invention, the weight of the subtask reward is updated using the following formula:
[0212] The total error is:
[0213] ,
[0214] No. The error of each subtask is:
[0215] ,
[0216] No. The weights of each subtask are updated as follows:
[0217] ,
[0218] Step S506, Q network each After the next update The parameters of the Q network Copy to the target Q network, i.e., each Q network After the next update, ;
[0219] Step S506: judging whether the loss function is less than a third threshold value, if not, then and returning to step S503, otherwise entering step S507;
[0220] Step S507: setting , if not satisfying , the maximum training round, then returning to step S502, otherwise ending.
[0221] After ending, the parameters are obtained as the optimal values and , and the optimal decision value set under the t period corresponding to the values of the optimal values are obtained, and the optimal decisions under each time period are obtained in sequence by repeating step S501 and step S507, so that the power plant decision considering the influence of the future electricity-carbon-certificate market environment state on the current scheduling decision is realized.
[0222] Embodiment Two
[0223] Figure 4 is a component schematic diagram of a green certificate and CCER coupled power scheduling decision optimization device according to an embodiment of the present application. As Figure 4 indicated, the green certificate and CCER coupled power scheduling decision optimization device according to an embodiment of the present application includes a data acquisition module, a first prediction module, a second prediction module, and a scheduling optimization module.
[0224] The data acquisition module is configured to acquire historical data of a plurality of price influence factors of an electricity-carbon-certificate market, scheduling information data of a plurality of power plant operators, and a coupling relationship between green certificates and CCERs.
[0225] The second prediction module is configured to, according to the historical data of each of the price influence factors of the electricity-carbon-certificate market, adopt a pre-trained LSTM prediction model to predict prediction data of each of the price influence factors in a future T time period.
[0226] The second prediction module is configured to, according to the historical data and the prediction data of each of the price influence factors, adopt a pre-trained BP prediction model to predict prediction data of various prices of the electricity-carbon-certificate market in the future T time period, the prices including a thermal power generation price, a green power generation price, a CCER price, a carbon emission quota price, and a green certificate price.
[0227] The scheduling optimization module is configured to use the predicted data of various prices of the electricity-carbon-certificate market, the scheduling information data of the power plant, and the coupling relationship between the green certificate and the CCER, to construct a state space, an action space, and a reward function for multi-agent reinforcement learning of the power plant in the electricity-carbon-certificate market, and to use a deep Q network algorithm to iteratively optimize to obtain an optimal power scheduling decision of the power plant.
[0228] The embodiment of the present application realizes intelligent collaborative decision-making of power scheduling in a deep integration scenario of the green certificate market, the CCER market, and the electricity market by constructing a power scheduling decision optimization system based on the combination of the LSTM prediction model, the BP neural network, and the multi-agent reinforcement learning.
[0229] The embodiment of the present application can comprehensively consider complex factors affecting transaction decisions and improve the comprehensiveness and precision of power scheduling decisions by collecting and processing various historical data variables such as electricity prices, green certificate prices, CCER prices, carbon prices, power generation, and power consumption for multi-dimensional information fusion, and using an LSTM model for multivariate time series prediction. In addition, by combining a BP neural network to further fit the future market price trend, the accuracy of the prediction results is improved, thereby providing more comprehensive and detailed input support for subsequent power scheduling strategies.
[0230] The embodiment of the present application can iteratively optimize and update the model based on continuous reception of new data, and respond to dynamic changes in the external environment in real time. The embodiment of the present application can support small-granularity power scheduling prediction and strategy adjustment, meet the flexible decision-making needs in a high-frequency fluctuating market, and improve the adaptability of the power scheduling subject to market changes.
[0231] The embodiment of the present application uses a multi-agent reinforcement learning method to autonomously optimize transaction behavior, and each type of power plant (such as a thermal power enterprise, a green power generator, a power user, a power grid company, etc.) can learn and evolve individualized power scheduling strategies based on its independent state space, action space, and reward function. Through the training and optimization of the experience pool D and the deep Q network, the power scheduling decision is changed from experience-driven to data-driven and strategy-adaptive, significantly improving the game ability and collaboration level of multiple subjects in a coupled market environment.
[0232] The embodiment of the present application breaks down the information silos between the traditional electricity market and the carbon trading market, integrates the green certificate and the CCER into a unified power scheduling decision system, and realizes market linkage and collaborative optimization. The embodiment of the present application not only improves the allocation efficiency of power and carbon resources, but also promotes the marketization of renewable energy consumption and carbon emission reduction, serves the green energy development goal, has good social benefits and application and promotion prospects.
[0233] The embodiment of the present application realizes multi-source data fusion prediction, intelligent transaction strategy optimization and deep cooperation of green certificates, CCER and the electricity market, and constructs an intelligent, dynamic and multi-agent decision support system that can adapt to the development trend of future energy markets, has obvious advantages in improving prediction accuracy, enhancing strategy flexibility and market adaptability, promoting green energy transaction mechanism innovation, and has obvious practical value and popularization potential.
[0234] Those skilled in the art can understand that all or part of the processes of the above-mentioned embodiment methods can be completed by a computer program instructing relevant hardware, and the program can be stored in a computer readable storage medium. The computer readable storage medium is a disk, an optical disk, a read-only memory or a random access memory, etc.
[0235] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. A green certificate and CCER coupled power dispatch decision optimization method, characterized in that, The method comprises the following steps: obtaining historical data of multiple price influencing factors of an electricity-carbon-certificate market, scheduling information data of multiple power generation companies, and coupling relationship between green certificates and CCERs; using a pre-trained LSTM prediction model to predict prediction data of the price influencing factors respectively in the future T time periods according to historical data of the price influencing factors respectively of the electricity-carbon-certificate market; using a pre-trained BP prediction model to predict prediction data of various prices of the electricity-carbon-certificate market in the future T time periods according to the historical data and the prediction data of the price influencing factors respectively, the prices including thermal power generation price, green power generation price, CCER price, carbon emission quota price and green certificate price; using the prediction data of various prices of the electricity-carbon-certificate market, the scheduling information data of the power generation companies, and the coupling relationship between green certificates and CCERs to construct state space, action space and reward function of multi-agent reinforcement learning of the power generation companies in the electricity-carbon-certificate market, and using deep Q network algorithm to iteratively optimize to obtain optimal power scheduling decision of the power generation companies; the scheduling information data of the power generation companies includes green power generation amount, load amount, green certificate demand amount, CCER demand amount, carbon quota demand amount and risk preference coefficient; the green power includes photovoltaic power generation and wind power generation; the coupling relationship between green certificates and CCERs is that green certificates and CCERs are exchanged in proportion; the state space of each power generation company is represented as: ; wherein, is the state space, is the photovoltaic power generation price, is the wind power generation price, is the load amount, and are the wind power generation amount and the photovoltaic power generation amount, respectively, is the green certificate price, is the green certificate demand amount, is the CCER price, is the CCER demand amount, is the carbon quota demand amount, is the carbon quota price, is the risk preference coefficient; the subscript t indicates a relationship with a time period t. the action space is represented as: , wherein, is the action space, is the electricity trade volume, is the green certificate trade volume, is the CCER trade volume, is the green certificate to CCER ratio; The reward function of the power plant is the combination of the rewards of its dispatch sub-tasks in the electricity market, the carbon market and the green certificate market, specifically , the sub-task reward represents the income of the electricity market, the sub-task reward represents the cost of purchasing carbon quota in the carbon market or the income of selling carbon quota, represents the cost of purchasing green certificate in the green certificate market or the income of selling green certificate, is the weight corresponding to the sub-task reward, is the state of the power plant, is the action of the power plant, .
2. The method of claim 1, wherein the method further comprises a step of pre-training the LSTM prediction model of the price influencing factors, specifically: obtaining historical data of the price influencing factors to construct a first training sample set; initializing network parameters in the LSTM neural network; inputting input data of each sample in the first training sample set into the LSTM neural network to obtain prediction values, comparing the prediction values with true values of the samples, and calculating loss function values; using gradient descent method to update the network parameters and performing iterative training; when the loss function is less than a first preset value, stopping iteration, determining the network parameters of the LSTM neural network, and obtaining the trained LSTM prediction model.
3. The method of claim 2, wherein the loss function value of the LSTM neural network is calculated according to the following formula: ; ; ; in, This represents the loss function value during the s-th iteration of training. Indicates the training of the s-th iteration. The orientation factor of each sample Indicates the training of the s-th iteration. The nth sample pair Influence factors of each network parameter; It is a constant. , For the s-th iteration, train the th... The predicted value for each sample For the first The true value corresponding to each sample; This indicates that the s-th iteration of training the LSTM neural network represents the... Network parameters, It is the curvature parameter trained in the s-th iteration. sinh represents the hyperbolic sine function, and sign represents the sign function. This represents the inverse hyperbolic tangent function.
4. The method of claim 3, wherein, the network parameters in the LSTM network include weights and biases of input gates, forgetting gates, cell units and output gates respectively; using gradient descent method to update the network parameters includes: ; ; in, For learning rate, This indicates that the LSTM network is trained in the (s+1)th iteration. Network parameters.
5. The method of claim 1, wherein, the method further comprises a step of pre-training the BP prediction model, specifically: obtaining historical data of the prices and corresponding price influencing factors to construct a second training sample set; initializing network parameters in the BP neural network; inputting input data of each sample in the second training sample set into the BP neural network to obtain prediction values, comparing the prediction values with true values of the samples, and calculating loss function values; The gradient descent method is adopted to update the network parameters and perform iterative training. When the loss function is less than the second preset value, the iteration is stopped, the network parameters in the BP neural network are determined, and the trained BP prediction model is obtained.
6. The method of claim 5, wherein, The loss function value of the BP neural network is calculated according to the following formula each time the training is performed: , , , in, Represents the total loss function. , The first loss function is respectively Second loss function Adaptive weights, For the second training sample set The true value of each sample For the first One predicted value, This represents the number of samples in the second training sample set.
7. The method of claim 6, wherein, After each iteration, update as follows , : , , in, The first loss function at the t-th iteration Adaptive weights of the function The first loss function at the t-th iteration The function value, The second loss function at the t-th iteration The function value; The first loss function at the (t+1)th iteration Adaptive weight values, The second loss function at iteration t+1. The adaptive weight value.
8. The method of claim 1, wherein, The state space, the action space and the reward function of each power plant in the multi-agent reinforcement learning in the electricity-carbon-certificate market are independent.
9. The method of any one of claims 1-8, wherein, The price influencing factors of the thermal power generation price include crude oil futures price, natural gas futures price, national coal price, thermal power generation installed capacity and thermal power generation capacity; the price influencing factors of the carbon quota price include carbon trading volume, monthly carbon futures price and thermal power generation price; the price influencing factors of the CCER price include carbon quota price; the price influencing factors of the photovoltaic power generation price include photovoltaic power generation installed capacity, photovoltaic power generation capacity and sunshine hours; the price influencing factors of the wind power generation price include wind power generation installed capacity, wind power generation capacity and wind speed; and the price influencing factors of the green certificate price include renewable energy consumption weight.
10. A green certificate and CCER coupled power dispatch decision optimization device, characterized in that, comprise: a data acquisition module configured to acquire historical data of a plurality of price influencing factors of an electricity-carbon-certificate market, dispatching information data of a plurality of power plant operators, and a coupling relationship between green certificates and CCERs; a second prediction module configured to predict, according to the historical data of each of the price influencing factors of the electricity-carbon-certificate market, prediction data of each of the price influencing factors in a future T time period by using a pre-trained LSTM prediction model; a second prediction module configured to predict, according to the historical data and the prediction data of each of the price influencing factors, prediction data of various prices of the electricity-carbon-certificate market in the future T time period by using a pre-trained BP prediction model, the prices including thermal power generation price, green power generation price, CCER price, carbon emission quota price and green certificate price; a dispatching optimization module configured to construct a state space, an action space and a reward function of multi-agent reinforcement learning of power plant operators in the electricity-carbon-certificate market by using the prediction data of various prices of the electricity-carbon-certificate market, the dispatching information data of the power plant operators and the coupling relationship between green certificates and CCERs, and to obtain optimal power dispatching decisions of the power plant operators by using a deep Q network algorithm for iterative optimization. The dispatching information data of the power plant operators includes green power generation capacity, load, green certificate demand, CCER demand, carbon quota demand and risk preference coefficient; the green power includes photovoltaic power generation and wind power generation; and the coupling relationship between green certificates and CCERs is a proportional exchange between green certificates and CCERs. The state space of each power plant operator is represented as: ; wherein, is the state space, is the photovoltaic power generation price, is the wind power generation price, is the load amount, and are the wind power generation amount and the photovoltaic power generation amount, respectively, is the green certificate price, is the green certificate demand amount, is the CCER price, is the CCER demand amount, is the carbon quota demand amount, is the carbon quota price, is the risk preference coefficient; the subscript t represents a time period t. The action space is represented as: , wherein, is the action space, is the electricity trade volume, is the green certificate trade volume, is the CCER trade volume, is the green certificate to CCER ratio; The reward function of the power plant is the combination of the rewards of its dispatch sub-tasks in the electricity market, the carbon market, and the green certificate market, specifically , the sub-task reward represents the income of the electricity market, the sub-task reward represents the cost of purchasing carbon quotas in the carbon market or the income of selling carbon quotas, represents the cost of purchasing green certificates in the green certificate market or the income of selling green certificates, is the weight corresponding to the sub-task reward, is the state of the power plant, is the action of the power plant, .
Citation Information
Patent Citations
Electricity-carbon-green evidence multi-market equilibrium analysis method based on multi-agent reinforcement learning
CN117314040A
Carbon price prediction method and system based on multivariable screening strategy, and storage medium
CN117787487A
Optimal decision obtaining method for consumption subject in electricity-carbon-certificate multi-market
CN118195834A
Comprehensive energy system optimal scheduling method considering carbon-green certificate joint transaction mechanism
CN118735575A
Power dispatching optimization method and system based on deep reinforcement learning
CN119378733A