Deep reinforcement learning power grid reconstruction method and device based on high-precision wind power prediction

By combining high-precision wind power forecasting and deep reinforcement learning methods with wind farm output and electrical attribute constraints, the grid reconfiguration is optimized, solving the problem of insufficient participation of new energy sources in traditional grid reconfiguration technology and improving the reliability and economy of power system restoration.

CN120933923APending Publication Date: 2025-11-11NORTH CHINA ELECTRIC POWER UNIV +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511067631.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional grid reconfiguration technology fails to effectively consider the participation of new energy sources in power system restoration and neglects the improvement of power transmission paths, resulting in poor reconfiguration effects.

Method used

A deep reinforcement learning method based on high-precision wind power forecasting is adopted. The shortest distance between unrecovered nodes and generator units is measured by the action selection function. Combined with the constraints of wind farm output, system frequency regulation capability and grid voltage support capability, the power transmission path is optimized.

Benefits of technology

Effectively introducing new energy sources to participate in power system restoration optimizes power transmission paths, enhances grid reconfiguration, and improves the reliability and economy of system restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120933923A_ABST
    Figure CN120933923A_ABST
Patent Text Reader

Abstract

The invention provides a deep reinforcement learning power grid reconstruction method and device based on high-precision wind power prediction. The method comprises the following steps: selecting an action for a grid reconstruction model by utilizing an action selection function; the action selection function is used for measuring the selected priority of an unrecovered node based on the shortest distance between the unrecovered node which is not selected in the target power system and the node of the unrecovered unit; if the action meets the action constraint condition, executing the action, rewarding the net rack reconstruction model based on the shortest distance between the recovered node and the node of the unrecovered unit, and continuing iteration; the action constraint condition comprises a wind power plant output constraint, a system frequency modulation capability constraint constructed based on a first electrical attribute parameter of a new energy unit, and a grid voltage support capability constraint constructed based on a second electrical attribute parameter of a new energy station grid-connected node. According to the invention, new energy can be introduced to participate in power system recovery and optimize the power transmission path, and the net rack reconstruction effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power systems, and more specifically, to a deep reinforcement learning-based grid reconfiguration method and apparatus for high-precision wind power forecasting. Background Technology

[0002] Power system restoration is a complex, multi-objective, multi-constraint, multi-timescale, and nonlinear process. Based on different restoration objectives and operational tasks, the system restoration process is divided into three stages: black start, grid reconfiguration, and load restoration. These stages are interconnected and progressive, ultimately leading to normal system operation. Among these, grid reconfiguration is the most complex and crucial stage, serving as a crucial link between the previous and subsequent stages. Grid reconfiguration refers to optimizing the grid's operational performance and improving power supply reliability, economy, and security by altering the grid's topology (such as adjusting line connections and switch states) during power system restoration.

[0003] Traditional grid reconfiguration technology only considers the recovery of traditional generating units, without addressing the participation of renewable energy sources in system recovery, and neglects improvements to power transmission paths. Therefore, traditional grid reconfiguration technology is not suitable for power systems that have introduced renewable energy sources, and its reconfiguration effect is poor. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a deep reinforcement learning-based grid reconfiguration method and device based on high-precision wind power forecasting, which can introduce new energy sources to participate in power system restoration and optimize power transmission paths, effectively solving the limitations of traditional grid reconfiguration technology and improving the grid reconfiguration effect.

[0005] In a first aspect, embodiments of this application provide a deep reinforcement learning-based grid reconfiguration method based on high-precision wind power prediction, the method comprising: An action selection function is used to select the action for the current iteration of the grid reconfiguration model. The action selection function is based on the shortest distance between the unselected unrestored nodes and the unrestored units in the target power system to measure the selection priority of the unrestored nodes. Determine whether the action satisfies the action constraints of the grid reconfiguration model; the action constraints include wind farm output constraints, system frequency regulation capability constraints constructed based on the first electrical attribute parameters of new energy units, and grid voltage support capability constraints constructed based on the second electrical attribute parameters of new energy power station grid connection nodes; If the conditions are not met, then jump to the step of using the action selection function to select the action for the latest space frame reconstruction model in this iteration, and continue execution; If the conditions are met, the action is executed; and the reward value of the action is determined based on the shortest distance between the node recovered in this iteration and the node of the unrecovered unit. After updating the parameters of the grid reconfiguration model based on the reward value, the next iteration is performed until all units in the target power system have been started.

[0006] In one possible implementation, determining the reward value of the action based on the shortest distance between the nodes recovered in the current iteration and the nodes of the unrecovered units includes: Determine whether the node recovered in this iteration of the action is a bus node directly connected to the generator; If the node restored in this iteration of the action is a bus node directly connected to the generator, then the reward value of the action is calculated based on the generator's power generation within a given time period. If the node restored in this iteration of the action is not a bus node directly connected to the generator, then the reward value of the action is determined based on the shortest distance between the node restored in this iteration and the node of the unrestored unit.

[0007] In one possible implementation, determining the reward value of the action based on the shortest distance between the nodes recovered in the current iteration and the nodes of the unrecovered units includes: Substitute the shortest distance between the node recovered in this iteration and the node of the unrecovered unit in the following formula to obtain the reward value of the action; ; ; in, The reward value for the action. The reward for restoring the unit node. The attenuation coefficient is... This refers to the shortest distance between the node recovered in this iteration of the action and the node of the unrecovered unit. The preset constrained radiation distance, The first behavior of the reconstructed space frame model is a constraint parameter. The recovery time required for the nodes recovered in this iteration. The reward value for the node to be recovered in this iteration is selected, where k is the iteration number corresponding to this iteration.

[0008] In one possible implementation, the expression for the wind farm output constraint is: ; in, This is a collection of all wind farms that are connected to the grid. This represents the minimum active power output of the wind farm w. This represents the maximum active power output of the wind farm w. For the active power output of the wind farm w.

[0009] In one possible implementation, the active power output of any wind farm is predicted through the following steps: Meteorological data of the wind farm are obtained according to a preset forecast cycle; Extract a predetermined number of meteorological feature components from the meteorological data; The meteorological feature components are input into the corresponding wind power prediction model to obtain the initial prediction value; the wind power prediction model is obtained by optimizing it using the frost and ice optimization algorithm. The sum of all initial predictions is used as the target prediction.

[0010] In one possible implementation, the action selection function is represented by the following formula: ; in, Priority for selecting unrecovered nodes. As the gravitational factor, For the farthest distance in the target power system environment, The second behavior of the reconstructed space frame model is a constraint parameter. The preset constrained radiation distance, This represents the shortest distance between nodes that have not been restored and nodes of units that have not been restored.

[0011] In one possible implementation, the expression for the system frequency modulation capability constraint is: ; in, Let be the active power of the new energy unit i at time t. The maximum allowable steady-state frequency deviation of the target power system. For the collection of all conventional units, This represents the start-up and shutdown status of a conventional unit g at time t. For a conventional unit g, the active power is... This is the frequency response value of the unit. A collection of new energy generating units. The time required for the reconfiguration of the space frame.

[0012] Secondly, embodiments of this application also provide a deep reinforcement learning-based grid reconfiguration device for high-precision wind power prediction, the device comprising: The selection module is used to select the action for the current iteration of the network reconfiguration model using the action selection function; the action selection function is based on the shortest distance between the unselected unrecovered nodes and the unrecovered units in the target power system to measure the selection priority of the unrecovered nodes; The judgment module is used to determine whether the action satisfies the action constraints of the grid reconfiguration model; the action constraints include wind farm output constraints, system frequency regulation capability constraints constructed based on the first electrical attribute parameters of the new energy units, and grid voltage support capability constraints constructed based on the second electrical attribute parameters of the grid-connected nodes of the new energy power station. The jump module is used to jump to the action selection function to select the action for the current iteration of the latest grid reconstruction model if the condition is not met, and continue execution. An execution determination module is used to execute the action if the condition is met; and to determine the reward value of the action based on the shortest distance between the node recovered in this iteration and the node of the unrecovered unit in the action. The iteration module is used to update the parameters of the grid reconfiguration model according to the reward value and then perform the next iteration until all units in the target power system have been started.

[0013] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the deep reinforcement learning grid reconfiguration method based on high-precision wind power prediction as described in any of the first aspects.

[0014] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the deep reinforcement learning-based grid reconfiguration method based on high-precision wind power prediction as described in any of the first aspects.

[0015] This application provides a deep reinforcement learning-based grid reconfiguration method and apparatus based on high-precision wind power forecasting. The method includes: selecting actions for the current iteration of the grid reconfiguration model using an action selection function; the action selection function measures the selection priority of unrecovered nodes based on the shortest distance between unrecovered nodes and unrecovered generating units in the target power system; if the action satisfies the action constraints of the grid reconfiguration model, the action is executed, and a reward value is determined based on the shortest distance between the recovered node and the unrecovered generating unit, thereby updating the grid reconfiguration model parameters and continuing the iteration until all generating units are started up; the action constraints include wind farm output constraints, system frequency regulation capability constraints constructed based on the first electrical attribute parameters of new energy generating units, and grid voltage support capability constraints constructed based on the second electrical attribute parameters of new energy power station grid-connected nodes. This application enables the introduction of new energy sources to participate in power system restoration and optimize power transmission paths, effectively solving the limitations of traditional grid reconfiguration technologies and improving the grid reconfiguration effect. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart of a deep reinforcement learning-based power grid reconfiguration method based on high-precision wind power prediction, provided in an embodiment of this application, is shown. Figure 2 A flowchart illustrating the prediction of active power output of a wind farm provided in an embodiment of this application is shown. Figure 3 This illustration shows a schematic diagram of a deep reinforcement learning-based power grid reconfiguration device based on high-precision wind power prediction, provided in an embodiment of this application. Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0019] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0020] To enable those skilled in the art to utilize the content of this application, and in conjunction with the specific application scenario of "power system field," the following implementation is provided. For those skilled in the art, the general principles defined herein can be applied to other embodiments and application scenarios without departing from the spirit and scope of this application. Although this application is primarily described within the "power system field," it should be understood that this is merely an exemplary embodiment.

[0021] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0022] The following is a detailed description of a deep reinforcement learning-based power grid reconfiguration method based on high-precision wind power prediction, provided by an embodiment of this application.

[0023] This application employs machine learning for network structure reconstruction, the specific process of which is as follows: Step 1: Set the objective function and constraints for the space frame reconstruction model.

[0024] In this embodiment, the objective function is constructed with the goal of maximizing the total power generation of the power system within a given time period. Constraints include constraints on the startup of generating units to be restored and constraints on system operation. Constraints on the startup of generating units to be restored include constraints on unit startup power and critical time constraints on thermal power units. System operation constraints include constraints on conventional unit output, wind farm output, node power balance, line power flow transmission, node voltage and phase angle, system frequency regulation capability, and grid voltage support capability.

[0025] A. Objective function: In this embodiment, to minimize losses caused by power outages, dispatchers must restore the power system as quickly as possible. Grid reconfiguration, as a crucial step in the restoration process, has a significant impact on power system recovery; the more generating units restored within a given timeframe, the faster the system recovery can be achieved. Therefore, this embodiment aims to maximize the total power generation of the power system within a given timeframe, with the specific objective function as follows: ; Where T is the grid reconfiguration time; n is the number of all generating units in the power system; This refers to the start-up and shutdown status of the i-th unit at time t (referring to the start-up state (represented by the value 1) and the stop state (represented by the value 0) of the unit during operation). Let be the active power generated by the i-th generating unit at time t. Let be the starting power of the i-th unit.

[0026] Furthermore, the force function is expressed by the following formula: ; in, This is the start-up time of the i-th generating unit; Let be the time from startup to the start of ramp-up for the i-th unit; Let be the time it takes for the i-th unit to start ramping up and reach its rated power. Let be the rated power of the i-th unit; Let be the ramp rate of the i-th unit.

[0027] B. Constraints: a. Start-up constraints for units to be restored (used to constrain the units to be restored connected to the nodes selected for restoration in the network reconfiguration model). (1) Unit start-up power constraints The starting power constraint of generating units has a significant impact on the optimal startup decision-making process. The power system provides the starting power for generating units, typically from units already connected to the grid. Therefore, the starting power constraint of the i-th generating unit can be expressed as: ; in, The power that a black-start generator in a power system can provide at time t. This refers to the number of generating units in the power system that are connected to the grid, excluding black-start units. Let be the output function used to describe the output power of the j-th generating unit at time t. Let j be the starting power of the j-th unit. For the first The starting power of the unit (i.e., the unit to be restored connected to the node in the action selected for restoration in the grid reconfiguration model).

[0028] (2) Critical time constraint of thermal power unit (used to constrain the thermal power units to be restored connected to the nodes selected for restoration in the grid reconstruction model) Considering that there are two methods for starting thermal power units: hot start and cold start, hot start requires the thermal power unit to have a maximum critical hot start time. Startup must be completed within the specified timeframe. If missed, it must be completed within the minimum critical startup time specified for cold start. Then, startup is achieved. In summary, the thermal power units to be restored are connected to the nodes that are being restored during the action at time t. i Start-up should meet the following critical time constraints for thermal power units: ; ; in, Less than .

[0029] Here, if the thermal power unit does not meet the minimum critical start-up time constraint, the thermal power unit will not be able to start. If the thermal power unit does not meet the maximum critical hot start-up time constraint, although the thermal power unit can start, it will prolong the required time and affect the recovery process.

[0030] b. System operational constraints (1) Output constraints of conventional units In this application's embodiments, conventional generating units refer to generator sets that use traditional energy sources (such as coal, natural gas, fuel oil, etc.) as fuel. The output constraints of conventional generating units are as follows: ; ; ; ; in, G represents the set of black-start generating units, and G represents the set of all conventional generating units that have been connected to the grid. The active power generated, This refers to the reactive power generated by a conventional generating unit (g). For conventional units that are not black-start units, the fixed output value of g is... This represents the minimum active power output of the black-start generator unit g. This represents the maximum active power output of the black-start generator unit g. This represents the minimum reactive power output of the black-start generator unit g. This represents the maximum active power output of the black starter unit g.

[0031] (2) Wind farm output constraints ; in, This is a collection of all wind farms that are connected to the grid. For wind farm The minimum value of the active power output. For wind farm The maximum value of active power. Contribute to the wind farm.

[0032] (3) Node power balance constraints ; ; in, A collection of grid-connected power outflow lines. For the collection of grid-connected power inflow lines, It is the set of all nodes connected to the grid in the power system. Let i be the set of units connected to node i. Let i be the set of wind farms connected to node i. Let be the active power of the line ij between the i-th node and the j-th node. Let g be the active power of unit g. For the active power output of wind farm w. Let be the reactive power of the line ij between the i-th node and the j-th node. The reactive power of unit g is... Let w be the reactive power output of the wind farm, and L be the set of all grid-connected circuits.

[0033] (4) Line power flow transmission constraints ; ; ; ; in, For sufficiently large positive numbers, For binary decision variables, This is used to indicate whether the line ij between the i-th node and the j-th node has been restored (the value corresponding to restoration is 1, and the value corresponding to non-restoration is 0). As the reference power, Let be the susceptance of line ij. Let i be the phase angle of the bus corresponding to node i. Let j be the phase angle of the bus corresponding to node j. Let be the voltage of the bus corresponding to node i. Let be the voltage of the bus corresponding to node j. Let be the conductance of line ij. This represents the minimum active power flowing through line ij. Let be the impedance of line ij. This represents the maximum active power flowing through line ij. This represents the minimum reactive power flowing through line ij. This represents the maximum reactive power flowing through line ij.

[0034] (5) Node voltage and phase angle constraints ; ; in, Let be the minimum voltage that the bus corresponding to node i can withstand. The maximum voltage that node i can withstand on the corresponding bus. Let be the minimum phase angle that can be withstood on the bus corresponding to node i. This represents the maximum phase angle that can be borne by the bus corresponding to node i.

[0035] (6) System frequency regulation capability constraints In this embodiment of the application, to ensure that the system frequency regulation capability is sufficient to withstand the power surge when new energy sources are connected, the system frequency regulation capability constraint that should be met under the premise that new energy generator units start up at different times is as follows: ; in, Let be the active power of the new energy unit i at time t. The maximum allowable steady-state frequency deviation of the target power system. This is a collection of all conventional units. This represents the start-up and shutdown status of a conventional unit g at time t. For a conventional unit g, the active power is... This is the frequency response value of the unit. A collection of new energy generating units. The time required for the reconfiguration of the space frame.

[0036] (7) Voltage support capacity constraints of the power grid In this embodiment, the voltage support strength of the restored grid can be measured using the Multiple Renewable Energy Station Short Circuit Ratio (MRSCR) index. This ensures that the grid connection of each renewable energy unit matches the grid strength of the power system, guaranteeing safe system restoration. The grid connection node of a renewable energy unit refers to the busbar in a renewable energy power generation system (such as a wind farm or photovoltaic power station) used to collect and connect the electrical energy generated by renewable energy power generation equipment (such as wind turbines and photovoltaic panels) to the power grid.

[0037] The current injected into the power system by the grid-connected bus of each new energy unit is , ,⋯, The voltage at the grid connection node of each new energy unit is... , ,⋯, The equivalent impedance matrix is: The equivalent impedance matrix includes the equivalent impedance between the grid connection nodes of every two new energy units. The relationship between current, voltage, and equivalent impedance is shown below: ; in, Injecting current from the power system into the grid connection node of the nth renewable energy unit. Let be the voltage at the grid connection node of the nth renewable energy unit. Let n be the equivalent impedance between the grid connection nodes of the nth renewable energy unit and the grid connection nodes of the nth renewable energy unit, where n is the number of renewable energy units.

[0038] The following formula represents the short-circuit ratio (MRSCR) of the i-th renewable energy unit across multiple power plants (an important indicator for evaluating the voltage support strength of the power system after multiple renewable energy power plants are connected to the power system). i : ; in, Let be the nominal voltage of the grid connection node of the i-th renewable energy unit. is the voltage of the grid connection node of the i-th new energy unit.

[0039] Multiply both the numerator and denominator of the above formula by We can get: ; Where, is the complex power conversion factor between the grid connection node of the i-th new energy unit and the j-th new energy unit, and * represents conjugate transpose; is the apparent power of the i-th new energy unit (which refers to the maximum apparent power that a new energy power generation system (such as a wind farm, a photovoltaic power station, etc.) can provide under specific conditions. Apparent power is an important parameter in the power system, which comprehensively considers the influence of active power and reactive power).

[0040] If the impedance ratio X / R of the grid connection node of the new energy unit > 10 (X is impedance and R is resistance) and it is assumed that the voltage phase angles between the new energy units are similar, the above formula can be simplified to obtain the calculation formula for the multi-station short-circuit ratio index MRSCR i of the i-th new energy unit: ; Where, S aci is the three-phase short-circuit capacity of the grid connection node of the i-th new energy unit, P REi is the active power of the grid connection node of the i-th new energy unit, is the voltage of the i-th new energy unit without direction, is the equivalent impedance between the grid connection node of the i-th new energy unit without direction and the grid connection node of the j-th new energy unit.

[0041] Therefore, the grid voltage support ability constraint can be expressed as: ; Where, is the multi-station short-circuit ratio index of the i-th new energy unit. Generally, it is considered that when MRSCRi ≥ 3, the receiving-end system is a strong system; when 2 < MRSCRi < 3, the receiving-end system is a weak system; when MRSCRi ≤ 2, the receiving-end system is an extremely weak system.

[0042] For example, the formula for calculating the multi-station short-circuit ratio (MRSCR) of the i-th renewable energy unit represents the ratio of the system short-circuit capacity to the apparent power of the renewable energy capacity, taking into account the coupling effect of multiple renewable energy stations. During system recovery, the node impedance matrix changes with the network topology. Simultaneously, the uncertainty of renewable energy output causes dynamic changes in the active power injected into the grid-connected nodes. The multiplication of these two factors creates a highly nonlinear relationship, leading to greater computational complexity and making it difficult to solve using traditional methods. However, deep reinforcement learning algorithms, through autonomous learning and interaction with the environment, can effectively solve the problem of strong nonlinear coupling, providing a solution for the efficient calculation of MRSCR and improving the robustness of the recovery strategy.

[0043] Step 2: Construct a grid reconstruction model based on Markov model.

[0044] In this embodiment of the application, based on the Markov assumption grid reconstruction model, the current state depends only on the previous state.

[0045] The Markov model consists of a state space, an action space, a reward function, and transition probabilities. Furthermore, considering that the APF-DQN (Artificial Potential Field Deep Q-Network) algorithm is a model-free deep reinforcement learning algorithm, it does not rely on prior knowledge of transition probabilities. The state space, action space, and reward function in the network reconstruction stage will be defined below.

[0046] A. State space: The state space contains all possible states encountered in the grid reconfiguration model. During the grid reconfiguration phase, the state includes the restored generating units in the power system, the output of the restored generating units (the electrical power generated by the generator units under specific operating conditions. It reflects the actual generating capacity of the units at a certain moment, such as active power and reactive power), and the restored power transmission paths.

[0047] To avoid resource waste due to excessively high spatial dimensionality, the state space consists of restored line nodes. When the k-th node recovers, the power system enters the k-th decision cycle, and the state... As shown below.

[0048] ; in, The state of node i (a node can only be connected to one generator set, and the state of node i represents whether the connected generator set or the node itself has recovered). A node is a connection point in a power system used to connect different components (such as generators, transformers, transmission lines, loads, etc.).

[0049] B. Movement Space: The action space encompasses all feasible actions. During the grid reconfiguration phase, an action represents the behavior of restoring a specific unit or node and its corresponding lines. Once an action is completed, the state space changes, entering a new decision cycle. A power supply path is formed from any restored node in the power system to the action node, restoring the nodes involved. The action in the k-th decision cycle... As shown in the formula.

[0050]

[0051] in, For multidimensional phasors, This represents the action taken on node i. Actions include restoring (represented by the value 1) and not changing (represented by the value 0).

[0052] Here, the execution of the action must satisfy three constraints: topology constraints, power constraints, and state constraints. Topology constraints require that there be a connecting transmission line between the restored node and the node to be restored (i.e., the node whose action is restoration). Power constraints require that when there is a unit to be restored at the node to be restored, the power system must meet the starting power constraint of the unit to be restored in order to restore the node. Conversely, it can only be started after the starting power constraint is met. State constraints stipulate that the node's state can only be from 0 to 1, representing that the action can only restore an unrestored node, and cannot shut down a restored node.

[0053] C. Reward Function: Rewards are timely feedback from the environment on actions taken. The reward function significantly impacts the algorithm's efficiency and result quality. The reward function is set with reference to the objective function of power system restoration. Taking into account the overall power system restoration situation, a piecewise reward function is adopted, with parameters used to adjust the reward, thereby increasing the adaptability of the reward function to the grid reconfiguration model. The reward function is shown below.

[0054] ; in, The reward value for the action. This represents the amount of electricity generated by the generator directly connected to the node restored during the action within a given time period. The first row of the space frame reconstruction model represents the constraint parameters. is the time required to restore the nodes recovered in the action, and k is the iteration number corresponding to this iteration. The second row of the space frame reconstruction model represents the constraint parameters. and The value is adjusted based on the actual training results.

[0055] Here, the reward function is based on whether the node recovered by the agent satisfies the constraints and whether... The generators are configured separately. If... If the constraints are not met, the agent will be penalized. Conversely, if the constraints are met, a judgment should be made. whether Generators require a reward for connection. Since the speed of generator start-up and ramp-up reflects the generator's output characteristics, the reward is set as the pre-processed power generation within a given time period; when the action... Node not When generators are activated, they will be penalized. This penalty helps guide the agent to start all units as quickly as possible, reducing the total recovery time during the grid reconfiguration phase.

[0056] Furthermore, reasonable feedback enables the agent to better explore the system environment. However, the reward function setting exhibits significant discreteness, which to some extent limits the agent's environmental exploration ability. Therefore, this embodiment improves the reward function using the Artificial Potential Field (APF) method to avoid excessive function discreteness. Simultaneously, considering that the agent's initial action selection during wind turbine participation in power system restoration is too random and consumes significant computational resources, the APF algorithm guides action selection, improving the agent's exploration efficiency and accelerating algorithm convergence. Compared to existing research that only considers turbine startup order optimization, the proposed method optimizes the restoration path based on existing research. In addition, the traditional DQN algorithm relies heavily on random exploration and environmental interaction, resulting in low sample efficiency and difficulty in quickly learning complex strategies. Therefore, this embodiment adopts the DM-DQN algorithm with a competitive network structure and introduces the APF algorithm on top of it. The algorithm decouples action selection and action evaluation, enabling a faster learning rate and allowing for more full utilization of early environmental exploration experience.

[0057] A. DM-DQN Algorithm: The competing network structure used in DM-DQN (Dueling Munchausen Deep Q-Network) updates the values ​​of all other actions simultaneously while updating the Q-value (a numerical measure of the expected reward (or utility) of taking a particular action in a given state). This more frequent updating allows for more accurate state value estimation when the action space has a high dimensionality, making the competing network structure superior in performance.

[0058] The DM-DQN network structure can be divided into two parts. One part is the value function that depends only on the state s. V (s, w, ) The other part is related to the state.s and actions a Relevant Advantage Function A(s, a, w, ) The network output can be represented as: ; in, These are common parameters of the value function V and the advantage function A. and These are the parameters of the value function V and the advantage function A, respectively. The value of the value function V can be seen as the state... s The average of the Q values. The A value is constrained to have a mean of 0, and the sum of the V value and the A value is the original Q value.

[0059] In M-DQN, when the Q-value of an action needs to be updated, the Q-network is updated directly, thereby improving the Q-value of that action. However, due to the constraint that the sum of the advantage functions must be zero, DM-DQN prioritizes updating the V-value. The V-value represents the average level of Q-values. Adjusting the V-value is equivalent to simultaneously updating the Q-values ​​of all actions in that state. This reduces the number of numerical updates, speeds up convergence, and allows for faster learning of the optimal decision.

[0060] When DM-DQN is applied to power system network reconfiguration, the value function handles the case where the agent fails to restore bus nodes not directly connected to generators, while the dominance function handles the case where the agent restores to nodes not directly connected to generators. To address the identifiability issue, the dominance function is centrally processed.

[0061] .

[0062] B. Artificial potential field method: The artificial potential field method is a path planning algorithm based on a virtual potential field. It treats the environment as a potential field space and uses nodes not directly connected to the generator and bus nodes directly connected to the generator as repulsive and attractive force sources, respectively, generating opposing and positive forces on the agent. Taking all factors into account, the agent will move according to the direction of the resultant force.

[0063] Assume the intelligent agent exists in a two-dimensional environment. For the position of the agent, The gravitational potential field function of the location of the bus node directly connected to the generator. It can be represented as: ; in, It is a proportionality coefficient of the gravitational potential field. This represents the distance between the agent and the bus nodes of the generators in the surrounding unit nodes, with the direction pointing towards the target point. Surrounding unit nodes are those that have an gravitational interaction with the agent.

[0064] ; in, This is the gravitational function.

[0065] Accordingly, based on the agent's connection to the node not directly connected to the generator. distance Determine the repulsive potential field function. As shown below: ; in, The distance threshold at which nodes excluding generators experience repulsive force effects.

[0066] Repulsive potential field function Differentiating yields the repulsive force function As shown below: ; Where k is the repulsive potential field coefficient.

[0067] Therefore, the resultant force F is as shown below.

[0068] .

[0069] (1) Optimization of reward function: To avoid low learning efficiency due to a discrete reward function, this embodiment improves the reward function using the aforementioned artificial potential field method. The unit within the system is considered a gravitational center, capable of releasing radiating rewards to the surrounding area. This reward is as follows: ; Furthermore, the improved reward function of the artificial potential field method is shown below: ; in, The reward value for the action. The reward for restoring the unit node. The attenuation coefficient is... This refers to the shortest distance between the node that was restored during the operation and the node that was not restored. The preset constrained radiation distance, The first row of the space frame reconstruction model represents the constraint parameters. This is the time required to restore the nodes recovered during the action. The reward value for the node restored during the selected action.

[0070] Traditional DQN algorithms often use absolutely random sampling during Q-table initialization, which makes the agent's action selection lack directionality in the early stages of training, resulting in slow convergence. In the early stages of training, if the environment is large, a large number of invalid iterations are likely to occur. As training progresses, the agent gradually understands the environment information, and the algorithm will gradually tend towards convergence, with the speed increasing accordingly.

[0071] To address the aforementioned issues, this invention employs the Artificial Potential Field (APF) algorithm to improve the action selection function. By treating the generator unit as a gravitational source, it corrects the action selection, making choices that are more conducive to achieving the objective. The optimized action selection function is shown below.

[0072] ; in, Priority for selecting unrecovered nodes. As the gravitational factor, For the farthest distance in the target power system environment, The second behavior of the reconstructed space frame model is a constraint parameter. The preset constrained radiation distance, This represents the shortest distance between unrecovered nodes and nodes connected to unrecovered units.

[0073] Step 3: After offline training of the space frame reconstruction model, apply the trained space frame reconstruction model online.

[0074] In this embodiment, the space frame reconstruction model is trained offline until the preset model performance is achieved. The process of training the space frame reconstruction model once can be referred to the model application process in S101-S105.

[0075] Furthermore, referring to Figure 1 The diagram shown is a flowchart of a deep reinforcement learning-based power grid reconfiguration method based on high-precision wind power prediction provided in an embodiment of this application. The exemplary steps of this embodiment are described below: S101. Use the action selection function to select the action for the current iteration of the space frame reconstruction model.

[0076] In this embodiment, the state of the grid reconfiguration model is first initialized. Then, during any iteration of the grid reconfiguration, an ε-greedy strategy is adopted, and an action selection function is used to select the action for the current iteration of the grid reconfiguration model. The action selection function measures the selection priority of unrecovered nodes based on the shortest distance between unselected unrecovered nodes in the target power system and nodes with unrecovered units.

[0077] Among them, the ε-greedy policy is a commonly used exploration-exploitation policy in reinforcement learning. It selects actions based on probability. Randomly select an action to explore, with probability. The strategy of selecting the best known action to exploit plays a crucial role in balancing exploration and exploitation.

[0078] S102. Determine whether the action satisfies the action constraints of the space frame reconstruction model.

[0079] In the embodiments of this application, the action constraints include wind farm output constraints, system frequency regulation capability constraints constructed based on the first electrical attribute parameters of the new energy units, and grid voltage support capability constraints constructed based on the second electrical attribute parameters of the grid-connected nodes of the new energy power stations.

[0080] The action constraints of the space frame reconstruction model have been described above and will not be repeated here.

[0081] S103. If not satisfied, jump to the action selection function to select the action for this iteration of the latest grid reconstruction model and continue execution.

[0082] In this embodiment of the application, if the conditions are not met, the second behavior constraint parameter of the grid reconstruction model is determined as the reward value of the action based on the reward function and stored in the experience pool. Then, the model is updated based on the reward value. Then, the process jumps to selecting the action for the current iteration of the latest grid reconstruction model using the action selection function, so as to reselect the action.

[0083] S104. If satisfied, execute the action; and determine the reward value of the action based on the shortest distance between the node recovered in this iteration and the node of the unrecovered unit.

[0084] In this embodiment of the application, the action is performed to restore the node restored in the current iteration of the action.

[0085] Specifically, the reward value for an action is determined based on the shortest distance between the nodes recovered in the current iteration and the nodes of the unrecovered units, including: Step 1: Determine whether the node recovered in this iteration is a bus node directly connected to the generator.

[0086] Step 2: If the node restored in this iteration of the action is a bus node directly connected to the generator, then calculate the reward value of the action based on the generator's power generation within a given time period.

[0087] In this embodiment of the application, based on the reward function, the reward value of the action is calculated using the following formula based on the generator's power generation within a given time period: .

[0088] Step 3: If the node restored in this iteration is not a bus node directly connected to the generator, then the reward value of the action is determined based on the shortest distance between the node restored in this iteration and the node of the unrestored unit.

[0089] In the embodiments of this application, based on the reward function, if the node restored in the current iteration of the action is not a bus node directly connected to the generator, then the shortest distance between the node restored in the current iteration of the action and the node of the unrestored unit is substituted into the following formula to obtain the reward value of the action. ; ; in, The reward value for the action. The reward for restoring the unit node. The attenuation coefficient is... This represents the shortest distance between the nodes recovered in this iteration of the action and the nodes of the unrecovered units. The preset constrained radiation distance, The first row of the space frame reconstruction model represents the constraint parameters. The recovery time required for the nodes recovered in this iteration. The reward value for the node to be recovered in this iteration is selected, where k is the iteration number corresponding to this iteration.

[0090] S105. After updating the parameters of the grid reconfiguration model based on the reward value, proceed to the next iteration until all units in the target power system have been started.

[0091] In this embodiment, the reward value is stored in the experience pool, the state is updated, and then the parameters of the grid reconstruction model are updated based on the reward value before the next iteration is performed.

[0092] Furthermore, the wind power prediction model is trained through the following steps to predict the active power output of the wind farm.

[0093] Step 1: Obtain raw data. The raw data consists of historical weather data and active power output data. First, fill in the missing values ​​and remove outliers. Perform preliminary cleaning to avoid affecting subsequent research.

[0094] Step 2: In order to further explore the information hidden in the historical active power output data of wind farms, the historical wind power sequence in the original data is decomposed into k modal components (IMFs) representing different meteorological data vector characteristics through Feature Mode Decomposition (FMD), and the corresponding original data matrix Mi (i = 1, 2, ⋯,k) is formed.

[0095] In the embodiments of this application, the algorithm of the wind power prediction model and the quality of the data used during training both affect the prediction accuracy of the active power output of the wind farm. Therefore, data decomposition technology is indispensable for the wind power prediction process. Commonly used methods such as empirical mode decomposition and variational mode decomposition analyze the potential patterns in the data, thereby simplifying the data structure and enhancing the reliability of the prediction results. Compared with commonly used decomposition methods, which are easily affected by filter shape, bandwidth, and filter center frequency, the decomposition mode used by Feature Mode Decomposition (FMD) is different.

[0096] Eigenmode decomposition (EMD) is a non-recursive signal processing algorithm that automatically selects different mode components by establishing an FIR filter bank and updating the filter coefficients. A modified maximum CK deconvolution method is proposed to address this issue. This method obtains the optimal FIR filter coefficients without requiring fault cycles as prior knowledge. The original signal is denoted as x(N), with a length of N. FMD theory will provide a solution for the constrained problem shown below.

[0097] ; st ; in, It is the nth value of the kth mode. M It is the number of shifts. Ts It is the input cycle. f represents the number of signal value decompositions. k It is the first k One filter, The first in the original signal Each signal value.

[0098] The above-mentioned constrained problem can be solved using an iterative eigenvalue decomposition algorithm. The decomposition process is as follows: .

[0099] in, , ; .

[0100] The correlation kurtosis CK of the decomposition mode can be expressed as: ; Where H represents the conjugate transpose operation. Used to control the weighted correlation matrix.

[0101] ; In summary, the calculations yield the following results: ; in, For the weighted correlation matrix, This is the correlation matrix. , These are equivalent parameters.

[0102] The FMD theory estimates the fault period by exploiting the characteristic that the autocorrelation spectrum of a signal produces local maxima at periodic locations. Let x(n) be the autocorrelation function, as shown below: ; in, This refers to the lag time.

[0103] To eliminate modal mixing, FMD prioritizes the mode with the highest correlation coefficient (CC). Simultaneously, it discards the mode with the smaller CK value among the two modes with the highest CC, thereby retaining the mode with more fault information. (Two modes) and CC is calculated by the following formula: ; in, and They are Modality and The average value of the modes.

[0104] Step 3: Initialize the RIME parameters of the wind power prediction model. With the minimum root mean square error as the optimization objective, the optimization is achieved through soft frost search strategy, hard frost puncture strategy and positive greedy mechanism. The particle position is continuously updated to optimize the parameters of the wind power prediction model using the Bidirectional Long Short-Term Memory Neural Network (BiLSTM) algorithm.

[0105] In this embodiment, the RIME and BiLSTM parameters are first initialized; the population is updated using a soft frost search strategy; information exchange is performed using a hard frost puncture strategy; the optimal parameters are found using a forward greedy mechanism; if the optimal parameters are obtained, the wind power prediction model is trained; if the optimal parameters are not obtained, the population is updated again using the soft frost search strategy.

[0106] Here, if only a simple model structure is used, the accuracy of the predicted active power output is often low, and the parameters have a significant impact on the model's performance but it is difficult to determine the precise values. However, optimizing the model's parameters through optimization algorithms can further improve the model's prediction accuracy. Therefore, this application uses the Rime Optimization Algorithm (RIME) to optimize the parameters of the bidirectional long short-term memory neural network, thereby obtaining the bidirectional long short-term memory neural network structure with the highest prediction accuracy, i.e., the wind power prediction model.

[0107] A. Frost Ice Optimization Algorithm The Frost Optimization Algorithm is a metaheuristic algorithm based on the growth mechanism of frost, possessing excellent global optimization and local adjustment capabilities. The core principles of the algorithm are a soft frost search strategy, a hard frost penetration mechanism, and a forward greedy selection mechanism.

[0108] Soft frost search strategy: In a light breeze, free-floating frost particles move according to specific patterns. Their movement efficiency is highly random and easily influenced by environmental factors. When they move near soft frost, they condense with internal particles, essentially covering the surface of an object, but their diffusion speed in the same direction is relatively slow. This characteristic allows the algorithm to quickly cover the entire search space and avoid getting trapped in local optima. The following formula can be used to calculate the position of frost particles.

[0109] ; in, R represents the position of the j-th particle in the i-th frost body agent after the update. best,j The j-th particle is the optimal frost body agent; parameter r1 is a random number in the range (-1, 1); r1 and The direction of particle motion is determined by: h is the adhesion degree; Ubij and Lbij are the upper and lower bounds of the escape space; E is the adhesion coefficient; r2 and E determine whether the particle position is updated.

[0110] ; ; ; Where t is the number of iterations; T is the maximum iteration limit; β Environmental factors; w The step size of environmental factors is determined; E is the adhesion coefficient, which affects the aggregation probability of individuals and increases with the number of iterations; r2 is a random number in the range of (0, 1), which, together with E, controls whether the particle position is updated. It will change as the number of iterations changes.

[0111] Hard frost puncture mechanism: Under strong winds, soft frost will gradually form hard frost structures in the same direction, and information from different frost bodies will continuously cross and exchange, thus avoiding the algorithm from getting trapped in local optima. The formula for particle replacement is as follows.

[0112] ; Where r3 is a random number in (-1, 1); Let be the probability that the i-th frost body is selected. It is a frost ice crystal.

[0113] Forward Greedy Selection Mechanism: This improves the selection mechanism of the metaheuristic algorithm by introducing a forward greedy selection mechanism. This mechanism aims to retain the original advantages while enhancing the ability to explore and mine the global solution space. It determines whether to update an individual and its corresponding solution by comparing the fitness values ​​before and after the update. While achieving the optimal result, it can further discover individuals with potentially superior solutions, achieving better global optimization.

[0114] B. Bidirectional Long Short-Term Memory Neural Network Bidirectional Long Short-Term Memory (BiLSTM)

[34] is a recurrent neural network improved on the basis of LSTM. Traditional LSTM is trained on data based only on the information of the forward time series, which makes the algorithm lack the ability to capture early learning features. However, BiLSTM can overcome the shortcomings of traditional LSTM and capture past and future state information. BiLSTM is composed of forward LSTM units and backward LSTM units. In each time step, the network can learn using the forward and backward data from the input time series. Based on the above structure, BiLSTM can better capture data features and enhance the learning ability of the algorithm, thereby improving the wind power output prediction. The algorithm controls the propagation error by adjusting the forward and backward propagation coefficients.

[0115] The parameter update mechanism within a one-way LSTM layer is the same as that of a traditional LSTM. Taking a single LSTM with hidden unit h as an example, backpropagation can be expressed as: (15) (16) (17) (18) (19) (20) in, Let C be the backpropagation error, N be the number of cell sets, and U and W be the weight arrays. Let E be the error derivative, y be the output error, g be the input gate, f be the forget gate, q be the output gate, and S be the output gate. h Here, h represents the state variable of cell h, and t represents the positive time step. This is the reverse time step.

[0116] Step 4: Input each Mi into the corresponding wind power prediction model for training, thereby enhancing the model's ability to capture fluctuations in the active power output of wind farms, and obtaining the corresponding predicted active power output value Yi (i = 1, 2, ...). (k), which are superimposed to obtain the final predicted value of the total active power output of the wind farm.

[0117] Furthermore, referring to Figure 2 The diagram shown is a flowchart of the active power output prediction process for a wind farm provided in this application embodiment. The steps are explained below: S201. Obtain meteorological data of the wind farm according to the preset forecast cycle.

[0118] In this application, considering factors such as recovery efficiency, recovery safety, and solution speed, a rolling optimization-based online decision-making method for system recovery is proposed. Wind power control is achieved through predictive control using a wind power forecasting model to ensure active power balance during power system recovery. Furthermore, with the goal of starting all generating units, a look-ahead rolling mechanism is adopted, using the period from the current time step to the end of recovery as the optimization time domain, and uniformly optimizing operations across all recovery periods within the recovery decision time domain.

[0119] During rolling optimization, the optimization time domain length, the recovery time required for components, and the reactive power generated by line charging all influence the formulation of power system grid reconfiguration decision schemes. Based on connectivity constraints, the recoverable components in the previous time step determine the recoverable components in the current time step. Only through long-term observation and accumulation can the impact of different decision schemes on the recovery process be determined. Therefore, an appropriate rolling optimization time domain should be determined when formulating decision schemes to ensure the foresight of wind power prediction models, avoiding both excessively large domains that affect wind power prediction accuracy and excessively small domains that consume resources and increase computational burden. Relevant regulations require that actual wind farm ultra-short-term forecasts be updated at least every 15 minutes, and that future wind farm output forecast data be reported to the power dispatch center to provide data support for rolling optimization and feedback correction.

[0120] S202. Extract a preset number of meteorological feature components from meteorological data.

[0121] S203. Input the meteorological feature components into the corresponding wind power prediction model to obtain the initial prediction value; the wind power prediction model is obtained by optimizing it using the frost and ice optimization algorithm.

[0122] S204. The sum of all initial predicted values ​​is determined as the target predicted value.

[0123] This application provides a deep reinforcement learning-based power grid reconfiguration method based on high-precision wind power forecasting. This method models the grid reconfiguration stage as an MDP (Markov Decision Process) model and uses a DM-DQN algorithm, improved from the Artificial Potential Field (APF) algorithm, to solve the model, thereby optimizing the system's generator start-up sequence and power transmission paths. Simultaneously, considering the impact of wind power output uncertainty, an online decision-making framework for the power system based on rolling optimization model predictive control is proposed, combining offline training and online decision-making to ensure scheme reliability and improve system recovery efficiency.

[0124] Based on the same inventive concept, this application also provides a deep reinforcement learning-based grid reconstruction device corresponding to the deep reinforcement learning-based grid reconstruction method based on high-precision wind power prediction. Since the principle of the device in this application is similar to the deep reinforcement learning-based grid reconstruction method based on high-precision wind power prediction described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0125] Reference Figure 3 The diagram shown is a schematic of a deep reinforcement learning-based grid reconfiguration device for high-precision wind power prediction provided in an embodiment of this application. The deep reinforcement learning-based grid reconfiguration device for high-precision wind power prediction includes: Selection module 301 is used to select the action for the current iteration of the grid reconfiguration model using the action selection function; the action selection function measures the selection priority of unrecovered nodes based on the shortest distance between the unselected unrecovered nodes in the target power system and the nodes with unrecovered units. The judgment module 302 is used to determine whether the action satisfies the action constraint conditions of the grid reconstruction model; the action constraint conditions include wind farm output constraints, system frequency regulation capability constraints constructed based on the first electrical attribute parameters of the new energy units, and grid voltage support capability constraints constructed based on the second electrical attribute parameters of the grid-connected nodes of the new energy stations; Jump module 303 is used to jump to the action selection function to select the action for the latest grid reconstruction model in this iteration if the condition is not met, and continue execution. The execution determination module 304 is used to execute the action if the condition is met; and to determine the reward value of the action based on the shortest distance between the node recovered in this iteration and the node of the unrecovered unit in the action. The iteration module 305 is used to update the parameters of the grid reconfiguration model according to the reward value and then perform the next iteration until all units in the target power system have been started.

[0126] like Figure 4 As shown in the embodiment of this application, an electronic device 400 includes a processor 401, a memory 402, and a bus. The memory 402 stores machine-readable instructions that can be executed by the processor 401. When the electronic device is running, the processor 401 communicates with the memory 402 through the bus. The processor 401 executes the machine-readable instructions to perform the steps of the deep reinforcement learning power grid reconfiguration method based on high-precision wind power prediction as described above.

[0127] Specifically, the memory 402 and processor 401 mentioned above can be general-purpose memory and processor, without any specific limitations. When the processor 401 runs the computer program stored in the memory 402, it can execute the above-mentioned deep reinforcement learning power grid reconfiguration method based on high-precision wind power prediction.

[0128] Corresponding to the above-mentioned deep reinforcement learning-based grid reconfiguration method based on high-precision wind power prediction, this application embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-mentioned deep reinforcement learning-based grid reconfiguration method based on high-precision wind power prediction.

[0129] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.

[0130] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0131] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0132] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the information processing methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0133] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A deep reinforcement learning-based power grid reconfiguration method based on high-precision wind power forecasting, characterized in that, The method includes: An action selection function is used to select the action for the current iteration of the grid reconfiguration model. The action selection function is based on the shortest distance between the unselected unrestored nodes and the unrestored units in the target power system to measure the selection priority of the unrestored nodes. Determine whether the action satisfies the action constraints of the grid reconfiguration model; the action constraints include wind farm output constraints, system frequency regulation capability constraints constructed based on the first electrical attribute parameters of new energy units, and grid voltage support capability constraints constructed based on the second electrical attribute parameters of new energy power station grid connection nodes; If the conditions are not met, then jump to the step of using the action selection function to select the action for the latest space frame reconstruction model in this iteration, and continue execution; If the conditions are met, the action is executed; and the reward value of the action is determined based on the shortest distance between the node recovered in this iteration and the node of the unrecovered unit. After updating the parameters of the grid reconfiguration model based on the reward value, the next iteration is performed until all units in the target power system have been started.

2. The deep reinforcement learning-based power grid reconfiguration method based on high-precision wind power prediction according to claim 1, characterized in that, The determination of the reward value for the action based on the shortest distance between the nodes recovered in this iteration and the nodes of the unrecovered units includes: Determine whether the node recovered in this iteration of the action is a bus node directly connected to the generator; If the node restored in this iteration of the action is a bus node directly connected to the generator, then the reward value of the action is calculated based on the generator's power generation within a given time period. If the node restored in this iteration of the action is not a bus node directly connected to the generator, then the reward value of the action is determined based on the shortest distance between the node restored in this iteration and the node of the unrestored unit.

3. The deep reinforcement learning-based power grid reconfiguration method based on high-precision wind power prediction according to claim 1 or 2, characterized in that, The determination of the reward value for the action based on the shortest distance between the nodes recovered in this iteration and the nodes of the unrecovered units includes: Substitute the shortest distance between the node recovered in this iteration and the node of the unrecovered unit in the following formula to obtain the reward value of the action; ; ; in, The reward value for the action. The reward for restoring the unit node. The attenuation coefficient is... This refers to the shortest distance between the node recovered in this iteration of the action and the node of the unrecovered unit. The preset constrained radiation distance, The first behavior of the reconstructed space frame model is a constraint parameter. The recovery time required for the nodes recovered in this iteration. The reward value for the node to be recovered in this iteration is selected, where k is the iteration number corresponding to this iteration.

4. The deep reinforcement learning-based power grid reconfiguration method based on high-precision wind power prediction according to claim 1, characterized in that, The expression for the wind farm output constraint is: ; in, This is a collection of all wind farms that are connected to the grid. This represents the minimum active power output of the wind farm w. This represents the maximum active power output of the wind farm w. For the active power output of the wind farm w.

5. The deep reinforcement learning-based power grid reconfiguration method based on high-precision wind power prediction according to claim 4, characterized in that, Predict the active power output of any wind farm using the following steps: Meteorological data of the wind farm are acquired according to a preset forecast cycle; Extract a predetermined number of meteorological feature components from the meteorological data; The meteorological feature components are input into the corresponding wind power prediction model to obtain the initial prediction value; the wind power prediction model is obtained by optimizing it using the frost and ice optimization algorithm. The sum of all initial predictions is used as the target prediction.

6. The deep reinforcement learning-based power grid reconfiguration method based on high-precision wind power prediction according to claim 1, characterized in that, The action selection function is expressed by the following formula: ; in, Priority for selecting unrecovered nodes. As the gravitational factor, For the farthest distance in the target power system environment, The second behavior of the reconstructed space frame model is a constraint parameter. The preset constrained radiation distance, This represents the shortest distance between nodes that have not been restored and nodes of units that have not been restored.

7. The deep reinforcement learning-based power grid reconfiguration method based on high-precision wind power prediction according to claim 1, characterized in that, The expression for the system frequency modulation capability constraint is: ; in, Let be the active power of the new energy unit i at time t. The maximum allowable steady-state frequency deviation of the target power system. For the collection of all conventional units, This represents the start-up and shutdown status of a conventional unit g at time t. For a conventional unit g, the active power is... This is the frequency response value of the unit. A collection of new energy generating units. The time required for the reconfiguration of the space frame.

8. A deep reinforcement learning-based grid reconfiguration device based on high-precision wind power forecasting, characterized in that, The device includes: The selection module is used to select the action for the current iteration of the network reconfiguration model using the action selection function; the action selection function is based on the shortest distance between the unselected unrecovered nodes and the unrecovered units in the target power system to measure the selection priority of the unrecovered nodes; The judgment module is used to determine whether the action satisfies the action constraints of the grid reconfiguration model; the action constraints include wind farm output constraints, system frequency regulation capability constraints constructed based on the first electrical attribute parameters of the new energy units, and grid voltage support capability constraints constructed based on the second electrical attribute parameters of the grid-connected nodes of the new energy power station. The jump module is used to jump to the action selection function to select the action for the current iteration of the latest grid reconstruction model if the condition is not met, and continue execution. An execution determination module is used to execute the action if the condition is met; and to determine the reward value of the action based on the shortest distance between the node recovered in this iteration and the node of the unrecovered unit in the action. The iteration module is used to update the parameters of the grid reconfiguration model according to the reward value and then perform the next iteration until all units in the target power system have been started.

9. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the deep reinforcement learning-based grid reconfiguration method based on high-precision wind power prediction as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the deep reinforcement learning-based grid reconfiguration method based on high-precision wind power prediction as described in any one of claims 1 to 7.