Unit power generation plan rapid solving method fusing constraint identification and deep reinforcement learning

By integrating constraint identification and deep reinforcement learning, effective security constraints are selected and transformed into Markov decision processes. The near-end policy optimization algorithm is used to solve the problem of optimizing and scheduling generator generation plans in large-scale power systems, achieving fast and accurate generation plan optimization and improving the computational efficiency and security of the power grid.

CN121504656APending Publication Date: 2026-02-10CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511465969.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly and accurately handle complex generator scheduling optimization in large-scale power systems, especially after the integration of large-scale renewable energy sources. The models are large in scale, computationally inefficient, and difficult to adapt to the uncertainties and complexities of the power grid. Traditional methods are insufficient to meet the requirements of industrial applications.

Method used

By employing a method that integrates constraint identification and deep reinforcement learning, effective safety constraints are screened out through stacked denoising autoencoders (SDAE), transformed into Markov decision processes, and solved using the proximal policy optimization algorithm (PPO). Combined with reward function design, this approach enables fast and accurate optimization of power generation plans.

Benefits of technology

It significantly improves the solution efficiency of unit power generation plans, reduces computation time, can quickly respond to grid uncertainties, ensures the safe and economical operation of the grid, and enhances the model's adaptability and solution rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504656A_ABST
    Figure CN121504656A_ABST
Patent Text Reader

Abstract

The invention discloses a unit power generation plan rapid solving method fusing constraint identification and deep reinforcement learning. The method comprises the following steps: constructing a unit power generation plan optimization model; carrying out constraint identification by utilizing a stacked automatic encoder; converting the unit power generation plan optimization model into a Markov decision process; and solving the unit power generation plan optimization model converted into the Markov decision process by adopting a deep reinforcement learning method. According to the method, unit power generation plan scheduling decision making is carried out based on an artificial intelligence method, uncertain changes of a novel power system are rapidly and accurately responded, efficient identification and adaptive solution of an acting constraint set in a power generation plan optimization model are achieved, the complexity of the model is reduced, the calculation efficiency is improved, and safe and economical operation of a power grid is guaranteed. According to the method, a large number of constraints can be quickly processed, scheduling strategy decisions can be adaptively learned, the solving rate can be more efficiently improved, and higher uncertainty and complexity of a novel power system can be dealt with.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of power grid optimal scheduling, and in particular to a unit generation plan fast solving method fusing constraint identification and deep reinforcement learning. BACKGROUND

[0002] With the construction and development of China's new power system, the scale of the power grid expands rapidly, and the number of units and lines in the power system increases dramatically. In addition to considering the operation constraints of wind, light, water, fire and storage resources, the optimal scheduling of generation plans also needs to consider the network constraints, section constraints and safety constraints of the power system. At the same time, with the large-scale access of renewable energy, there is a large fluctuation and randomness, which further increases the difficulty of solving the unit commitment problem in the generation plan. On the one hand, the increase in a large number of unit constraints leads to a large model size, increasing the difficulty of generation plan optimization calculation and reducing the calculation efficiency. At the same time, large-scale renewable energy brings a large risk of power flow uncertainty, and the demand for safety constraint processing increases. On the other hand, with the improvement of power optimization scheduling requirements, the day-ahead and intraday generation plan requires solving in a shorter time. The existing simple call of commercial mathematical optimization software cannot meet the requirements of industrial applications, and the traditional method is difficult to realize adaptive optimization, which cannot effectively support the needs of future actual projects. Therefore, it is necessary to deeply embed human intelligent thinking mode in the algorithm process design to avoid the blindness of the optimization process and intelligently improve the solving efficiency of large-scale unit generation plan.

[0003] To improve the solving efficiency of large-scale unit generation plan, it is necessary to effectively identify a large number of constraints under the actual operation scenario of the power grid, and then combine the deep reinforcement learning algorithm to adaptively solve the optimization model to obtain a more rapid, accurate and intelligent generation plan, thereby better guaranteeing the safe, high-quality and economic operation of the power grid.

[0004] For this research problem, through investigation, it is found that the existing technology mainly identifies the safety constraints in the optimization model.

[0005] Document [1]: Chen Zirui, Liu Mingbo, Zeng Guihua, et al. K nearest neighbor algorithm applied to accelerate security constrained unit commitment problem [J]. Southern Power Grid Technology, 2024, 18(11): 48-57+78. A prediction method based on improved K nearest neighbor algorithm is proposed, which introduces a neighbor consensus threshold p to improve the voting mechanism of neighbors in KNN algorithm. A smaller neighbor consensus threshold is selected when solving the active transmission power constraint, and a larger threshold is selected when solving the partial integer variable value. Exceeding the threshold sets a fixed integer variable and active transmission power constraint, thereby constructing a simplified security constrained unit commitment model. Then, an optimization solver is applied to directly solve the model, thereby shortening the solution time of the security constrained unit commitment problem.

[0006] Document [2]: Jiang Wei, Feng Bin, Guo Chuangxin. Based on the role of the safety constraint identification method of graph neural network [J]. Electrical automation, 2023, 45 (02): 106-108. A method for identifying the role of safety constraints based on graph neural network is proposed. The branch classification model based on graph neural network is used to classify the branch power flow constraints in the base state and fault state. The input data includes node features and branch features. The output is the probability of the role of the branch safety constraint. The role of the constraint is identified, and the model is solved in combination with the role of the constraint.

[0007] Document [3]: Wang Qin, Chen Siyuan, Xu Jian, et al. Robust decision-making method for power system forward scheduling based on redundant constraint fast identification [J / OL]. Power system automation, 1-14. A robust decision-making method for power system forward scheduling based on redundant constraint fast identification is proposed. Graph convolutional neural network and long short-term memory network are used to learn spatial and temporal features, respectively, to identify and reduce redundant constraints. Finally, the column and constraint generation algorithm is used to solve the model.

[0008] Document [4]: Zhu Zhengchun, Yang Zhifang, Yu Juan, et al. Data-driven fast calculation method for security constraint economic dispatch for small sample scenarios [J]. Proceedings of the Chinese Society of Electrical Engineering, 2022, 42 (12): 4430-4440. For small sample scenarios, a fast calculation method for security constraint economic dispatch based on Gaussian process is proposed, including sample pre-classification and Gaussian process. First, based on marginal unit feature clustering, sample pre-classification is performed. Second, a Gaussian classifier is used to establish a load-output mapping for solving the output prediction. Finally, the predicted unit output is checked for N-1, and the active constraint set is obtained.

[0009] In the technical solutions of the above documents, the traditional method is still used for solving process, and the collaborative control of each dispatching resource in the power system is not thoroughly explored. The self-adaptability of the solution is poor, and as the complexity and uncertainty of the system increase, it is difficult to establish an accurate model for fast solving. SUMMARY

[0010] To overcome the shortcomings of the prior art, the present application proposes a unit generation plan fast solving method combining constraint identification and deep reinforcement learning. Based on artificial intelligence method, the unit generation plan dispatching decision is made, the uncertainty change of new type power system is responded quickly and accurately, the active constraint set in the generation plan optimization model is efficiently identified and adaptively solved, the complexity of the model is reduced, the calculation efficiency is improved, and the safe and economic operation of the power grid is ensured. This method can quickly process a large number of constraints and adaptively learn the dispatching strategy decision, more efficiently improve the solving rate, and cope with higher uncertainty and complexity of new type power system.

[0011] The technical scheme adopted by the present application is: A method for quickly solving unit generation scheduling by combining constraint recognition and deep reinforcement learning, comprising the following steps: Step 1: constructing a unit generation scheduling optimization model; Step 2: using a stacked auto-encoder for constraint recognition; Step 3: converting the unit generation scheduling optimization model into a Markov decision process; Step 4: using a deep reinforcement learning method to solve the unit generation scheduling optimization model converted into a Markov decision process.

[0012] The overall implementation process of the above-mentioned method for quickly solving unit generation scheduling by combining constraint recognition and deep reinforcement learning is as shown in Figure 1 , and specifically as follows: In step 1, the constructed unit generation scheduling optimization model includes a unit combination part and an economic dispatch part, and the optimization objectives are as follows: a. Objective function of the unit combination part: (1) In formula (1): is the operating cost of thermal power units g, hydroelectric units h and energy storage s; and are the marginal generation cost quotes of thermal power units g and hydroelectric units h, respectively; and are the operating outputs of thermal power units g and hydroelectric units h at time t, respectively; and are the start-up and shutdown costs of thermal power units g, respectively; and are the start-up and shutdown variables of thermal power units g at time t, respectively; and are the charging and discharging unit costs of energy storage s, respectively; and are the charging and discharging powers of energy storage s, respectively.

[0013] b. Objective function of the economic dispatch part: (2) In formula (2): is the standby optimization cost of thermal power units g, hydroelectric units h and energy storage s, is the probability of error scenario m, and m is the number of scenarios; and are the positive and negative standby capacity costs of thermal power units g, respectively; and are the positive and negative standby capacity costs of hydroelectric units h, respectively; and These are the upper and lower reserve capacities reserved for thermal power units in scenario m, respectively. and These are the upper and lower backup capacities reserved for the hydropower unit in scenario m, respectively. and These are the unit costs for charging, discharging, and backup of energy storage, respectively. and These are the standby capacities for energy storage charging and discharging; and The costs of curtailment for wind and solar power, respectively; and The power rationing is for wind and solar power, respectively. As a penalty for loss of load, This represents the amount of load loss.

[0014] The objective function of the unit combination part and the constraint conditions of the objective function of the economic dispatch part are mainly referenced in [1]: Li Jianzhao, Xie Min, Li Shujia, et al. Joint optimization of unit combination and multi-scenario reserve decision considering CVaR [J]. Southern Energy Construction, 2021, 8(04): 50-65. In step 2, constraint identification is performed using a stacked automatic encoder (SDAE). The specific steps are as follows: The constraint identification refers to the process of quickly identifying over-limit safety constraints in the power flow of a line and obtaining the effective safety constraints in advance.

[0015] The constraint identification process employs a stacked denoising autoencoder (SDAE), including offline training and online testing phases. Figure 2 As shown: ①. In the offline training phase, the input vector is first determined to be the fluctuation of new energy and load demand, and the output vector is the unit output. The correlation mapping between the input and output is established by using a stacked noise reduction autoencoder. The stacked noise reduction autoencoder is trained offline using a large amount of collected input and output data to obtain the trained stacked noise reduction autoencoder (SDAE).

[0016] ②. During the online testing phase, the current fluctuations in new energy sources and load demand are input into the trained stacked noise reduction autoencoder (SDAE) to obtain the optimal output of the unit; then, N-1 verification analysis is performed on the generated optimal output to screen out the over-limit constraints, i.e., the effective safety constraints.

[0017] The stacked noise reduction autoencoder (SDAE) is as follows Figure 3As shown, the deep neural network model is composed of stacked noise-reducing autoencoders (DAEs). Through the encoding-decoding transformation learning of multiple accumulated autoencoders, it can extract more compact and meaningful high-order features from the input new energy and load demand fluctuations, and further establish an accurate mapping between them and the optimal output of the unit.

[0018] The calculation formula for identifying the effective constraints using a stacked noise-reducing autoencoder (SDAE) is as follows: (3); In formula (3): The input feature vector of the model is the sum of the renewable energy output P and the load demand. Fluctuations between | |.

[0019] The output vector of the model, i.e., the optimal unit output. ; f is the feedforward function of the SDAE neural network; R is the activation function; is a vector consisting of neural network parameters; W is the weight vector of the neurons; This is the bias vector of the neuron.

[0020] The generated optimal output is subjected to N-1 verification analysis to screen out the constraints that exceed the limits. The screening formula is as follows: (4); like If the value is within the range of this formula, then it is an effective safety constraint.

[0021] In equation (4): and Let be the minimum and maximum transmission capacities of branches i and j under the c-th anticipated fault, respectively; and These are the unit output vector and load demand vector for each node, respectively; This is the active power sensitivity coefficient matrix of each node in the entire network to branch i and j under the c-th fault scenario; Specifically, a relaxation factor is introduced. The value ranges from 0 to 1, which can expand the range of the initial effective constraint set and adjust the judgment accuracy of line safety constraints.

[0022] In step 3, the unit power generation plan optimization model is transformed into a Markov decision process, and the state space, action space, and reward function considering safety constraints are designed, specifically including: 1) Problem transformation: This invention considers using deep reinforcement learning methods for adaptive model solving. A typical reinforcement learning decision-making process can usually be represented by a quintuple. To describe this, the unit power generation plan optimization model is first transformed into a Markov decision process (MDP), which includes the design state space S, action space A, reward function R, and discount factor. .

[0023] 2) Design state space S: The state space S is the set of all possible states when the agent interacts with the environment, containing all the necessary information for the model input. The state space of this invention sequentially includes load forecast values. Forecasted wind and solar power output P, ​​and unit start-up / shutdown status. Energy storage state of charge and power flow of transmission lines As in equation (5): (5); 3) Design motion space A: In generator set problems, actions represent the start-up and shutdown decisions and power output of a generator set at future moments; action space A is the space of start-up and shutdown decisions that all generator sets can make, and action space A is subject to constraints such as continuous start-up and shutdown. The action space of this invention describes the specific actions of decision variables, including, in order, generator power output. Start and stop Top-mounted preparation and underspin preparation Energy storage charging and discharging power As in equation (6): (6); 4) Design the reward function R: In the process of solving the power generation plan, the reward will affect the evaluation result of the state action value function, and the state action value function will in turn affect the choice of action, thus determining the direction of optimization adjustment. Therefore, designing a reasonable reward mechanism is the key to guiding the power generation plan to gradually adjust to the target state. The scheduling objective of this invention is to minimize the operating cost while ensuring the safe operation of the system. However, the objective of reinforcement learning is to maximize the reward. Therefore, the negative value of the objective function is used as part of the reward function. The reward function R is designed according to the objective function, including the unit operating cost, reserve capacity cost, load shedding penalty cost, wind curtailment and solar curtailment penalty cost, and safety constraints, as shown in the following equation (7): (7); In equation (5): R1 represents the reward that takes into account the total cost, which includes operating costs. Backup costs and penalty costs R2 represents the bonus considering line safety constraints; For line transmission power; and These represent the minimum and maximum transmission power limits for the line, respectively; R3 represents the bonus considering power balance constraints. Indicates the unit's output; Indicates the unit's standby output; and This represents the scaling factor of the reward function. n and L represent the number of generator units and the number of lines, respectively.

[0024] 5) Discount Factor : Discount factor The value range of is [0,1]. If , A value close to 1 indicates a focus on future rewards, while a value close to 0 indicates a focus on immediate rewards. An appropriate value can balance short-term and long-term returns, accelerating learning convergence. This invention determined this value to be 0.78 through continuous adjustments.

[0025] In step 4, a deep reinforcement learning method is used to solve the power generation planning and scheduling problem considering safety constraints. This invention selects a near-end policy optimization algorithm to solve the unit power generation planning optimization model, which is transformed into a Markov decision process. Specific steps are as follows: Figure 4 As shown, it includes: 1) During the offline training phase, the agent is trained based on the Actor-Critic architecture. The Actor part outputs actions based on the input state, and the Critic part evaluates the actions output by the Actor part. Ultimately, an optimal Actor function is found, meaning the neural network representing the Actor has its hyperparameters adjusted based on a large amount of input and output data. The input to the neural network is the state S, and the output is the probability distribution of the actions. Thus, the optimal action A of the scheduling strategy is obtained, so that the adjustment strategy of the output can always obtain the maximum expected reward value, that is, the maximum R value in equation (7).

[0026] Specifically, during the offline training phase, the agent continuously interacts with the environment, stores sampled data in the experience pool, and updates the neural network parameters until training is complete and the neural network parameters are determined.

[0027] 2) During the online execution phase, distributed execution is performed based on the trained Actor-Critic architecture. Only the updated Actor network is needed to solve the optimization strategies of each agent, and the Critic network does not need to participate. During the execution, the agent obtains the current environmental state S, selects the scheduling action A according to the strategy obtained by the Actor network, as shown in Equation (6), and then moves to the next state after execution. At the same time, the current reward value R is calculated, as shown in Equation (7).

[0028] Continue acquiring environmental data until the entire decision-making process is complete.

[0029] This invention provides a method for rapidly solving unit power generation plans by integrating constraint identification and deep reinforcement learning, with the following technical advantages: 1) This invention integrates constraint identification and deep reinforcement learning methods by pre-selecting a set of effective constraints and substituting this set of safe constraints into the reward function for solving. This enables faster and more accurate solution of the agent's optimal policy, avoids blind optimization, and greatly saves computation time.

[0030] 2) This invention uses the SDAE model to screen effective safety constraints. During the offline training phase, model parameters are fine-tuned, and high-dimensional features of the input load demand data are continuously extracted through a continuous encoding process to fit and output the optimal unit output. In the online application phase, the optimal unit output can be directly obtained based on the SDAE model, and the effective safety constraints can then be analyzed. Compared to traditional methods, this approach more easily meets the needs of online applications.

[0031] 3) This invention innovatively uses Markov decision processes to describe the power generation planning optimization problem and employs the Proximal Policy Optimization (PPO) algorithm to solve the power generation planning optimization model. At the same time, it utilizes the algorithm mode of "centralized training and distributed execution" to train a Critic with better team collaboration characteristics to guide the policy update of each Actor. In subsequent execution, each Actor can adaptively execute its own actions based on its own local observation information, which better ensures the speed and accuracy of online real-time decision-making.

[0032] 4) This invention incorporates system network power flow constraints into the scheduling model, taking into account the overall security performance of the system. Compared with single-objective optimization scheduling schemes that only consider operational economic benefits and security performance, this invention further comprehensively considers the security performance and economic benefits of system operation, which has significant engineering and practical implications for the future development of new power systems.

[0033] 5) The present invention can be configured or programmed to execute all methods using computers and other equipment, and can be adapted to new power system dispatch and control centers. Attached Figure Description

[0034] The present invention will be further described below with reference to the accompanying drawings and examples; Figure 1 This is a general framework diagram of the method for rapidly solving unit power generation plans that integrates constraint identification and deep reinforcement learning as described in this invention.

[0035] Figure 2 This is a schematic diagram of the effective constraint identification method described in this invention.

[0036] Figure 3 This is a flowchart of the solution process for the SDAE model that addresses the active constraints as described in this invention.

[0037] Figure 4 This is an architecture diagram of the deep reinforcement learning solution method described in this invention.

[0038] Figure 5 This is a graph showing the predicted day-ahead load, wind, and solar power output.

[0039] Figure 6(a) shows the accuracy curve of the SDAE model training; Figure 6(b) shows the loss curve of the SDAE model training.

[0040] Figure 7 To evaluate the change curve of reward value during the training phase.

[0041] Figure 8(a) shows the start-up and shutdown status of the thermal power unit; Figure 8(b) shows the unit output plan of the thermal power unit; Figure 8(c) shows the total output of wind, solar, hydro, thermal and storage units. Detailed Implementation

[0042] A method integrating constraint identification and deep reinforcement learning is proposed for fast solution of unit power generation plan. First, a unit power generation plan optimization model is constructed. Second, effective safety constraints are screened based on the trained SDAE model. Then, the optimization scheduling physical model is transformed into a Markov decision process, and effective safety constraints are added to the reward function design. Finally, the improved deep reinforcement learning algorithm, namely the proximal policy optimization algorithm, is used to solve the problem.

[0043] Example: 1. Case Study Introduction: This case study uses a real power grid example from a region in Northwest China, integrating 152 generating units and 25 energy storage units. The generation plan is solved by combining the day-ahead wind and solar power output forecasts and actual daily output for a typical day in this region, as well as load demand forecasts (including power transmitted through transmission channels). For example... Figure 5 The data shown are the day-ahead load, wind, and solar power output forecasts.

[0044] 2. To verify the effectiveness of the proposed constraint identification method, 10,000 different operating scenarios were first generated. Based on new energy sources, load demand, and unit output data, the model was trained offline, and its classification performance was tested and evaluated. Figures 6(a) and 6(b) show the changes in accuracy and loss value during SDAE model training, respectively. As can be seen from Figures 6(a) and 6(b), the model exhibits low loss value, high accuracy, and fast convergence speed during both training and testing, enabling it to well meet the identification requirements. Then, methods considering effective constraint identification and those not considering effective constraint identification were used to optimize the model solution, and the iterative performance under different scenarios was analyzed, as shown in Table 1.

[0045] Table 1 compares iterative performance under different scenarios.

[0046] It can be observed that the method proposed in this invention can converge in a maximum of 3 iterations, with a convergence rate of over 63%, which can effectively reduce the number of iterations for model solving and significantly improve the model's solution efficiency.

[0047] 3. To verify the effectiveness of the PPO model, evaluate the changes in reward values ​​during the training phase, such as... Figure 7 As shown, the reward curve fluctuates significantly in the initial 2000 events, converging somewhat after approximately 6000 training events. The trained PPO model can be directly used to solve power generation plan optimization, offering a significant advantage in solution time compared to mathematical programming methods that do not consider constraint handling.

[0048] 4. Regarding the solution results, in the process of solving the power generation plan, wind power, photovoltaic and hydropower units are set to be always on, and only the start-up and shutdown changes of thermal power units are considered. The start-up and shutdown status and unit output plan of 152 thermal power units are obtained by using the trained PPO model, as well as the total output value of wind, solar, hydro, thermal and storage units, as shown in Figure 8(a), Figure 8(b) and Figure 8(c).

[0049] As shown in Figure 8(a), thermal power units bear the main load demand of the system. By starting and stopping the units and adjusting their output, they can better adapt to the fluctuations in wind and solar power output. At the same time, the added energy storage further improves the load peak-valley difference, reduces the number of unit start-ups and shutdowns, and reduces wind and solar curtailment losses. Especially during periods of high wind and solar power generation, thermal power units quickly reduce their output to absorb renewable energy as much as possible. Therefore, the optimized scheduling strategy obtained above can better meet the energy supply and absorption needs of the system operation and achieve optimized coordinated scheduling among multiple units.

[0050] As shown in Figures 8(b) and 8(c), during periods of low load and high wind and solar power generation, the unit output plan can not only further increase unit output but also proactively reduce it, and even shut down some less economically viable units, thereby further increasing wind and solar power consumption. Furthermore, by reducing the number of operating units, the overall unit output becomes more concentrated. Simultaneously, considering cost analysis based on risk measurement, the losses from power curtailment and load shedding are reduced to zero.

Claims

1. A method for rapidly solving unit power generation plans by integrating constraint identification and deep reinforcement learning, characterized in that... Includes the following steps: Step 1: Construct an optimization model for the unit's power generation plan; Step 2: Constraint identification using a stacked automatic encoder; Step 3: Transform the unit power generation plan optimization model into a Markov decision process; Step 4: Use deep reinforcement learning to solve the unit power generation plan optimization model, which is transformed into a Markov decision process.

2. The method for rapidly solving unit power generation plans by integrating constraint identification and deep reinforcement learning according to claim 1, characterized in that: In step 1, the constructed unit power generation plan optimization model includes a unit combination part, the objective function of which is as follows: (1) In formula (1): The operating costs of thermal power unit g, hydropower unit h, and energy storage s; and These are the marginal generation cost quotes for thermal power unit g and hydropower unit h, respectively. and The output power of thermal power unit g and hydropower unit h at time t are respectively. and These represent the start-up and shutdown costs of thermal power unit g, respectively. and Let g be the start-up and shutdown variables of thermal power unit g at time t; and These are the unit costs of charging and discharging energy storage s; and These are the charging and discharging power of the energy storage s, respectively.

3. The method for rapidly solving unit power generation plans by integrating constraint identification and deep reinforcement learning according to claim 2, characterized in that: In step 1, the constructed unit power generation plan optimization model includes an economic dispatch component, the objective function of which is: (2); In formula (2): Optimize the standby costs for thermal power unit g, hydropower unit h, and energy storage s. Let m be the probability of error scenario m, where m is the number of scenarios; and These represent the positive and negative standby capacity costs of thermal power unit g, respectively. and These represent the positive and negative standby capacity costs of the hydropower unit h, respectively. and These are the upper and lower reserve capacities reserved for thermal power units in scenario m, respectively. and These are the upper and lower backup capacities reserved for the hydropower unit in scenario m, respectively. and These are the unit costs for charging, discharging, and backup of energy storage, respectively. and These are the standby capacities for energy storage charging and discharging; and The costs of curtailment for wind and solar power, respectively; and The power rationing is for wind and solar power, respectively. As a penalty for loss of load, This represents the amount of load loss.

4. The method for rapidly solving unit power generation plans by integrating constraint identification and deep reinforcement learning according to claim 3, characterized in that: In step 2, constraint identification is performed using a stacked automatic encoder (SDAE). Specifically, constraint identification refers to obtaining the active safety constraints in advance by quickly identifying the over-limit safety constraints in the line power flow. The constraint identification process employs a stacked noise-reducing autoencoder (SDAE), which includes offline training and online testing phases.

5. The method for rapidly solving unit power generation plans by integrating constraint identification and deep reinforcement learning according to claim 4, characterized in that: In the offline training phase, the input vector is first determined to be the fluctuation of new energy and load demand, and the output vector is the unit output. The correlation mapping between the input and output is established by using a stacked noise reduction autoencoder. The stacked noise reduction autoencoder is then trained offline using a large amount of collected input and output data to obtain the trained stacked noise reduction autoencoder SDAE.

6. The method for rapidly solving unit power generation plans by integrating constraint identification and deep reinforcement learning according to claim 5, characterized in that: During the online testing phase, the current fluctuations in new energy sources and load demand are input into the trained stacked noise reduction autoencoder (SDAE) to obtain the optimal output of the unit. Then, N-1 verification analysis is performed on the generated optimal output to screen out the over-limit constraints, i.e., the effective safety constraints.

7. The method for rapidly solving unit power generation plans by integrating constraint identification and deep reinforcement learning according to claim 6, characterized in that: The calculation formula for identifying the effective constraints using a stacked noise-reducing autoencoder (SDAE) is as follows: (3); In formula (3): The input feature vector of the model is the sum of the renewable energy output P and the load demand. Fluctuations between | |; The output vector of the model, i.e., the optimal unit output. f is the feedforward function of the SDAE neural network; R is the activation function. is a vector consisting of neural network parameters; W is the weight vector of the neurons; This is the bias vector of the neuron.

8. The method for rapidly solving unit power generation plans by integrating constraint identification and deep reinforcement learning according to claim 6, characterized in that: The generated optimal output is subjected to N-1 verification analysis to screen out the constraints that exceed the limits. The screening formula is as follows: (4); like If the value is within the range of this formula, then it is an effective safety constraint; In equation (4): and Let be the minimum and maximum transmission capacities of branches i and j under the c-th anticipated fault, respectively; and These are the unit output vector and load demand vector for each node, respectively; This is the active power sensitivity coefficient matrix of each node in the entire network to branch i and j under the c-th fault scenario; Introducing relaxation factors The value ranges from 0 to 1, which can expand the range of the initial effective constraint set.

9. The method for rapidly solving unit power generation plans by integrating constraint identification and deep reinforcement learning according to claim 7, characterized in that: In step 3, the unit power generation plan optimization model is transformed into a Markov decision process, and the state space, action space, and reward function considering safety constraints are designed, specifically including: 1) Problem transformation: Consider using deep reinforcement learning methods for adaptive model solving. The reinforcement learning decision-making process uses a quintuple. To describe this, the unit power generation plan optimization model is first transformed into a Markov decision process (MDP), which includes the design state space S, action space A, reward function R, and discount factor. ; 2) Design state space S: The state space includes, in sequence, load forecast values. Forecasted wind and solar power output P, ​​and unit start-up / shutdown status Energy storage state of charge and power flow of transmission lines As in equation (5): (5); 3) Design motion space A: Action space A is used to describe the specific actions of decision variables, including, in order, unit power output. Start and stop Top-mounted preparation and underspin preparation Energy storage charging and discharging power As in equation (6): (6); 4) Design the reward function R: The objective function is used to design the reward function R, which includes the unit operating cost, reserve capacity cost, load shedding penalty cost, wind and solar curtailment penalty cost, and safety constraints, as shown in the following equation (7): (7); In equation (5): R1 represents the reward that takes into account the total cost, which includes operating costs. Backup costs and penalty costs R2 represents the bonus considering line safety constraints; For line transmission power; and These represent the minimum and maximum transmission power limits for the line, respectively; R3 represents the bonus considering power balance constraints. Indicates the unit's output; Indicates the unit's standby output; and This represents the scaling factor of the reward function; n and L represent the number of generator units and the number of lines, respectively. 5) Discount Factor : Discount factor The value range of is [0,1]. If , A value close to 1 indicates a focus on future rewards, while a value close to 0 indicates a focus on immediate rewards.

10. The method for rapidly solving unit power generation plans by integrating constraint identification and deep reinforcement learning according to claim 9, characterized in that: In step 4, the near-end strategy optimization algorithm is selected to solve the unit power generation plan optimization model transformed into a Markov decision process, including: 1) During the offline training phase, the agent is trained based on the Actor-Critic architecture. The Actor part outputs actions based on the input state, and the Critic part evaluates the actions output by the Actor part. Ultimately, an optimal Actor function is found, meaning the neural network representing the Actor has its hyperparameters adjusted based on a large amount of input and output data. The input to the neural network is the state S, and the output is the probability distribution of the actions. Thus, the optimal action A of the scheduling strategy is obtained, so that the adjustment strategy of the output can always obtain the maximum expected reward value, that is, the R value in equation (7) is maximized. During the offline training phase, the agent continuously interacts with the environment, stores the sampled data in the experience pool, and updates the neural network parameters until training is complete and the neural network parameters are determined. 2) During the online execution phase, distributed execution is performed based on the trained Actor-Critic architecture. Only the updated Actor network is needed to solve the optimization strategy for each agent, and the Critic network does not need to participate. During the execution, the agent obtains the current environmental state S, selects the scheduling action A according to the strategy obtained by the Actor network, and moves to the next state after execution. At the same time, the current reward value R is calculated. Continue acquiring environmental data until the entire decision-making process is complete.