Adjusting method and device of power system and electronic equipment

By building the first and second sets of the power system, combining the deep reinforcement learning model and preset equations, the problem of low stability of the power system after the access of new energy equipment is solved, and efficient intelligent regulation and stable operation of the power system are achieved.

CN120341877APending Publication Date: 2025-07-18STATE GRID BEIJING ELECTRIC POWER CO +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510335719.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

With the increase in the proportion of new energy equipment, changes in line characteristics in the power system lead to low operating stability, and the existing model driving methods have insufficient regulation accuracy, making it difficult to ensure the stable operation of the power system.

Method used

By collecting power system operation data, building the first set and the second set, determining the first function and preset equation, and adjusting it in combination with the deep reinforcement learning model to ensure the stable relationship between the node voltage and branch power, and achieving intelligent regulation.

Benefits of technology

It improves the accuracy and stability of power system regulation, ensures that the bus voltage meets safety constraints, and maintains the stable operation of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120341877A_ABST
    Figure CN120341877A_ABST
Patent Text Reader

Abstract

The invention discloses an adjusting method and device of a power system and electronic equipment. The method comprises the steps that a first set and a second set are determined according to operation data of the power system, the data in the first set are used for representing the operation state of the power system, and the data in the second set are used for representing the adjustment range of the power system; a first function is determined according to the operation data and constraint information of the power system, the first function is used for measuring the performance of the power system, and the constraint information is used for carrying out safety constraint on bus voltage in the power system; according to the first set, the second set, the first function and a preset equation, the power system is adjusted, and the preset equation is used for describing the relation between the node voltage and the branch power in the stable operation process of the power system. The technical problems that after the power system is connected to the new energy equipment, the line characteristics change, and the operation stability of the power system is low are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of electric power, and in particular, to a method and device for regulating an electric power system and an electronic device. Background Art

[0002] In the technical field of electric power, the interconnected power grid, distributed power sources, new energy systems, and electrical equipment are usually managed as an integrated power system. With the development of new energy technologies, the proportion of distributed energy sources in the power system is increasing continuously, and the coupling degree among various energy sources in the power system is also increasing continuously.

[0003] The existing technical solutions can regulate the power system through the constructed traditional models. The traditional model-driven method is to construct a detailed physical model based on accurate power grid parameters. The formulation of the regulation plan in the power system is the core link for the stable operation of the power grid. The existing models solve the multi-period optimization problem to achieve the regulation decision-making for the power grid.

[0004] However, with the increase in the proportion of new energy equipment connected to the power system, the characteristic values of some buses and branches in the power system will change, resulting in the problem of an increase in the calculation scale of the physical model. In addition, with the popularization and application of large-scale new energy equipment, taking wind energy and solar energy as representative equipment of indirect new energy, they will be affected by meteorological factors such as temperature, wind direction, wind speed, and sunlight. Therefore, the high uncertainty brought by the power generation of new energy equipment will cause low accuracy of the regulation decision-making based on the traditional model-driven method, and further lead to the problem of low operation stability of the regulated power system.

[0005] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0006] The present application provides a method and device for regulating an electric power system and an electronic device, so as to at least solve the technical problems that the line characteristics change after new energy equipment is connected to the power system and the operation stability of the power system is low.

[0007] According to one aspect of the present application, a method for regulating a power system is provided, including: determining a first set and a second set according to the operation data of the power system, wherein the data in the first set is used to characterize the operation state of the power system, and the data in the second set is used to characterize the range of regulating the power system; determining a first function according to the operation data of the power system and constraint information, wherein the first function is used to measure the performance of the power system, and the constraint information is used to perform safety constraints on the bus voltages in the power system; regulating the power system according to the first set, the second set, the first function and a preset equation, wherein the preset equation is used to describe the relationship between the node voltages and the branch powers during the stable operation of the power system.

[0008] Optionally, determining the first set according to the operation data of the power system includes: determining the power information of the generator nodes, the power information of the load nodes, the line load rate and the bus voltage information in the power system according to the operation data; determining the first set according to the power information of the generator nodes, the power information of the load nodes, the line load rate and the bus voltage information.

[0009] Optionally, determining the second set according to the operation data of the power system includes: determining the power adjustment amount of the generator nodes, the voltage amplitude adjustment amount of the buses and the voltage phase angle adjustment amount of the buses according to the operation data and the first set; determining the second set according to the power adjustment amount of the generator nodes, the voltage amplitude adjustment amount of the buses and the voltage phase angle adjustment amount of the buses.

[0010] Optionally, determining the first function according to the operation data of the power system and the constraint information includes: determining the operation cost and the new energy consumption rate of the power system according to the operation data of the power system; determining the first function according to the operation cost, the new energy consumption rate and the constraint information.

[0011] Optionally, regulating the power system according to the first set, the second set, the first function and the preset equation includes: determining a second function according to the preset equation, wherein the second function is used to predict the operation state of the power system at a future moment; constructing a target model based on a preset framework, the first function and the second function, wherein the target model is a deep reinforcement learning model obtained by training an initial neural network based on L sets of historical operation data of the power system, and L is a positive integer; inputting the first set and the second set into the target model to obtain a target policy, wherein the target policy at least includes the adjustment amount for actually regulating the nodes in the power system; regulating the power system according to the target policy.

[0012] Optionally, constructing a target model based on a preset framework, a first function, and a second function, including: initializing a neural network based on the preset framework, the first function, the second function, and a first strategy to obtain a first model, where the first strategy is used to evaluate the security of the power system when any node in the power system fails, and the first model is used to interact with the operation data of the power system; constructing the target model based on the first model.

[0013] Optionally, constructing the target model based on the first model, including: updating the first model according to a second strategy to obtain a second model, where the second strategy is used to determine the failure probability of nodes in the power system according to the topological structure of the power system; constructing a third model based on a feedforward neural network and the second model, where the input data of the third model is the output data of the second model, and the output data of the third model is used to represent all adjustment schemes that the power system can select based on the current operating state; updating the third model based on the proximal policy optimization algorithm to obtain the target model.

[0014] According to another aspect of the present application, there is provided an adjustment device for a power system, including: a first determination unit configured to determine a first set and a second set according to the operation data of the power system, where the data in the first set is used to represent the operating state of the power system, and the data in the second set is used to represent the adjustment range of the power system; a second determination unit configured to determine a first function according to the operation data of the power system and constraint information, where the first function is used to measure the performance of the power system, and the constraint information is used to perform safety constraints on the bus voltage in the power system; an adjustment unit configured to adjust the power system according to the first set, the second set, the first function, and a preset equation, where the preset equation is used to describe the relationship between the node voltage and the branch power during the stable operation of the power system.

[0015] According to another aspect of the present application, there is provided a computer program product, where the computer program product includes a computer program, and when the computer program runs, it controls the computer program product to execute the above-mentioned adjustment method for the power system.

[0016] According to another aspect of the present application, there is provided an electronic device, including one or more processors and a memory, where the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the above-mentioned adjustment method for the power system.

[0017] In this application, first, a first set and a second set are determined based on the operation data of the power system. Among them, the data in the first set is used to characterize the operation state of the power system, and the data in the second set is used to characterize the range of adjustment to the power system. Then, this application determines a first function based on the operation data of the power system and constraint information. The first function is used to measure the performance of the power system, and the constraint information is used to perform safety constraints on the bus voltages in the power system. Finally, this application adjusts the power system based on the first set, the second set, the first function, and a preset equation. The preset equation is used to describe the relationship between the node voltages and branch powers during the stable operation of the power system.

[0018] As can be seen from the above, this application provides a method for adjusting a power system with embedded physical constraints to solve the problem that the line characteristics change after new energy devices are connected to the power system. This application first collects data that can characterize the operation state of the system (i.e., the data in the first set), so as to achieve the purpose of determining the adjustable range of the system based on the current operation state data (i.e., the data in the second set). Subsequently, this application determines the objective function for adjusting the power system (i.e., the first function) based on the operation state data of the system and physical constraint information. Finally, this application jointly adjusts the power system based on the first set, the second set, the first function, and a preset equation. While improving the accuracy of adjusting the power system, this application also ensures the stable operation of the power system through the relationship between the node voltages and branch powers described by the preset equation, ensuring that the bus voltages obtained after subsequent adjustment can meet the safety constraint conditions, thereby improving the stability of the adjusted power system.

[0019] It can be seen that in this application, the method of jointly adjusting the power system based on the first set, the second set, the first function, and a preset equation ensures the stable operation of the power system through the relationship between the node voltages and branch powers described by the preset equation, thereby achieving the technical effects of improving the accuracy of adjusting the power system and the stability of the adjusted power system, and further solving the technical problems of the change in line characteristics after new energy devices are connected to the power system and the low stability of the power system operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of this application and constitute a part of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0021] Figure 1 is a flowchart of an optional method for adjusting a power system according to an embodiment of this application;

[0022] Figure 2It is a structural diagram of an optional feedforward neural network according to an embodiment of the present application;

[0023] Figure 3 It is a flowchart of an optional proximal policy optimization algorithm according to an embodiment of the present application;

[0024] Figure 4 It is a flowchart of another optional method for regulating a power system according to an embodiment of the present application;

[0025] Figure 5 It is a structural diagram of an optional deep reinforcement learning model according to an embodiment of the present application;

[0026] Figure 6 It is a schematic diagram of an optional device for regulating a power system according to an embodiment of the present application;

[0027] Figure 7 It is a schematic diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0028] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0030] It should also be noted that the relevant information (including but not limited to the information for display and analysis) and data (including but not limited to the first data and the second data) involved in this application are all information and data authorized by the user or fully authorized by all parties. For example, an interface is set up between this system and relevant users or institutions. Before obtaining relevant information, a request for acquisition needs to be sent to the aforementioned users or institutions through the interface, and after receiving the consent information feedback from the aforementioned users or institutions, the relevant information can be obtained.

[0031] In addition, in the processes of collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant information and relevant data involved in this application, all comply with the relevant laws, regulations, and standards of the relevant regions, and necessary confidentiality measures are taken, without violating public order and good customs. In addition, this application provides corresponding operation entrances for users to choose to consent to authorization or refuse authorization. If the user chooses to refuse authorization, the corresponding expert decision-making process will be entered.

[0032] The explanations of professional terms related to the technical solution of this application are as follows:

[0033] RL (Reinforcement Learning) is one of the important technologies in the field of machine learning. Reinforcement learning technology has powerful potential for automatic grid operation. The core concept of reinforcement learning is that through interaction with the environment, the intelligent agent (i.e., the target model in this application) learns how to take actions to maximize a certain form of cumulative reward. This learning method simulates the natural process of animals and humans learning behaviors through trial and feedback, and is an important method for realizing autonomous decision-making and control in the field of artificial intelligence; Deep Reinforcement Learning combines the perception ability of deep learning and the decision-making ability of reinforcement learning, and can achieve optimal control from input to output through end-to-end training and learning. It can perceive high-dimensional data, perform real-time feedback and regulation, has good learning ability and generalization, and has a relatively fast inference speed. At the same time, it has strong versatility and can be applied to problems such as optimal control, resource allocation, power generation prediction, and regulation management in power systems.

[0034] Intelligent agent: The main body of reinforcement learning, responsible for learning and decision-making, making decision actions by observing the environmental state, and at the same time optimizing its own strategy according to environmental rewards to obtain the maximum cumulative reward.

[0035] Environment: Everything outside the intelligent agent, including all elements that can affect the behavior of the intelligent agent.

[0036] State: A quantitative representation of the environment where the intelligent agent is located, \(s_t\) t represents the state of the intelligent agent at time \(t\).

[0037] State space: The set \(S\) composed of all possible states.

[0038] Action: A quantitative representation of the decision made by the agent, a t represents the action executed by the agent at time t.

[0039] Action space: At a certain state s t The set A composed of all optional actions.

[0040] Reward: The feedback signal given by the environment after executing the agent's action, r t represents the reward feedback by the environment after executing the action a t under the state s t , which is used for the agent to measure its performance according to the task objective.

[0041] Policy: The agent maps the environmental state to an action, and this mapping relationship is called a policy. That is, the policy provides the action that the agent should take under a certain state. π represents the policy.

[0042] Transition function: The environment obtains the next state through the state transition probability. That is, when in the current state s t the agent takes the action a under the policy π t to obtain the next state s t+1 with a probability p(s t+1 |s t , a t ).

[0043] Cumulative return: The sum of a series of feedback rewards obtained by the agent from time t to the end decision-making process time t end -1.

[0044] According to an embodiment of the present application, an embodiment of a method for regulating a power system is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And, although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0045] The present application provides a regulation system for a power system (abbreviated as the regulation system) for executing the method for regulating the power system in the present application. Figure 1 is a flowchart of an optional method for regulating a power system according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:

[0046] Step S101, determining a first set and a second set according to the operation data of the power system.

[0047] In step S101, the data in the first set is used to characterize the operating state of the power system, and the data in the second set is used to characterize the range of adjustment to the power system.

[0048] Specifically, the operating data at least includes the active power values of the unit nodes and load nodes in the current period of the power system, the load rate of each line, the bus voltage amplitude, the bus phase angle, and the up / down adjustable space of the active power of the unit nodes. The operating data is the basis for the agent to learn and make decisions, and reflects the real-time operating state of the power system.

[0049] Optionally, the adjustment system is connected to Grid2Op (Grid to Operator, an open-source power grid operation simulation platform). The operating data of the power system is obtained through Grid2Op. Moreover, Grid2Op provides a simulation environment that can describe the Markov decision process, simulate some buses and branches whose characteristic values change, and perform training and evaluation on the known IEEE (Institute of Electrical and Electronics Engineers) synthetic power grids (such as IEEE14 and IEEE118) to test and optimize the performance of the agent (i.e., the target model constructed in this application) under different power grid conditions.

[0050] Optionally, the first set includes the power information of the unit nodes, the power information of the load nodes, the line load rate, and the voltage information of the buses, etc. The data in the first set is the real-time state information that the agent needs to consider when making decisions, that is, the first set is the state space defined in the Grid2Op simulation environment.

[0051] Optionally, the second set includes the power adjustment amount of the unit nodes, the voltage amplitude adjustment amount of the buses, and the voltage phase angle adjustment amount of the buses. The data in the second set defines the adjustment action space that the agent can take, that is, the second set is the action space defined in the Grid2Op simulation environment.

[0052] Step S102, determine the first function according to the operating data and constraint information of the power system.

[0053] In step S102, the first function is used to measure the performance of the power system, where the constraint information is used to perform safety constraints on the bus voltage in the power system.

[0054] Optionally, in a power system, the constraint information usually refers to the security constraints of the power system, such as the upper and lower limits of the bus voltage, the limit of the line load rate, etc. In this application, by imposing security constraints on the voltage amplitude, voltage phase angle value, and line load rate of the power system, the purpose of ensuring that the power system operates within a safe range is achieved, thereby improving the stability of the power system.

[0055] Optionally, the above first function quantifies key performance indicators of the power system, such as the operating cost of thermal power units in the power system, the accommodation rate of new energy, the stability of bus voltage and phase angle (i.e., constraint information), etc., into a numerical form of reward signals, and through these reward signals, reflects the direct impact of the actions taken by the agent on the system performance. Then, based on the reward signals feedback by the first function, the effect of the agent's actions is judged. When the actions taken can reduce the operating cost of thermal power units, increase the accommodation rate of new energy, and / or maintain the stability of bus voltage and phase angle, the agent will receive a higher reward feedback; otherwise, the agent will receive a lower reward or punishment.

[0056] As can be seen from the above, by guiding the agent to optimize its strategy through the first function to seek an action sequence that can maximize the cumulative reward, in this application, the first function serves as part of the reward mechanism, thereby achieving the purpose of guiding the agent to make action decisions that can improve the system performance, and further achieving the technical effects of reducing the operating cost of the power system, increasing the utilization rate of new energy, and ensuring the stable operation of the power grid.

[0057] Step S103, adjust the power system according to the first set, the second set, the first function, and the preset equation.

[0058] In step S103, the preset equation is used to describe the relationship between the node voltage and the branch power during the stable operation of the power system.

[0059] Optionally, the preset equation can be set as the power flow balance equation, which is based on circuit theory and takes into account the topological structure, electrical parameters, and operating conditions of the power system, and is the basis for power system analysis and control.

[0060] Optionally, in polar coordinate form, the power flow balance equation is shown as the following formula (1).

[0061]

[0062] In the above formula, P j and Q j respectively represent the active and reactive power injections at bus j, θ jk is the voltage phase angle difference between bus j and k, B jk and G jkrespectively represent the real part and the imaginary part of the (j,k)-th element in the bus admittance matrix.

[0063] In step S103, the agent adjusts the power system according to the first set, the second set, the first function, and the preset equation as follows:

[0064] First, the agent takes one or a series of actions based on the current system state (the first set) and the adjustable range (the second set); then, calculates the impact of these actions on the system operating state through the preset equation (power flow equation) to predict the future state; next, evaluates the power system performance of the future state according to the first function to provide a reward or punishment signal; finally, the agent adjusts its strategy according to the received reward signal, expecting to take better actions in future decisions, so as to achieve the goal of improving system performance and meeting safety constraints.

[0065] In an alternative embodiment, the traditional model-driven method constructs a detailed physical model based on accurate grid parameters, and realizes the regulation decision by solving the multi-period optimization problem. Among them, the physical knowledge contained in the historical data satisfies the safety constraints and can make the AC power flow converge, with strong interpretability. However, in a power system containing intermittent new energy, the characteristic values of some buses and branches will change, and the calculation scale based on the physical model is very large. The high uncertainty brought by renewable energy generation will cause problems such as difficult to accurately model and too long solution time.

[0066] In an alternative embodiment, the pure data-driven deep reinforcement learning method has a strong dependence on the grid operation data. The agent lacks the guidance of domain knowledge in regulation and needs a lot of time for trial and error, resulting in difficult convergence. At the same time, for scenarios that have not occurred in history, its decision is prone to large extrapolation errors, and even the situation of infeasible decisions may occur. For example, when the number of units is large, for the learning task with a high-dimensional action space, the large action exploration space will lead to a decrease in learning efficiency, resulting in poor regulation strategies.

[0067] In the embodiment of the present application, step S103 combines the operation data, constraint information, physical model, and deep reinforcement learning technology of the power system, so as to control the agent to be able to perform effective state prediction and policy optimization, and realize the efficient and intelligent regulation of the power system. Especially in the face of an uncertain environment where the characteristic values of some buses and branches change, the agent can quickly adapt and make the best adjustment decision, so as to ensure that the power grid is in a stable operation state. Through the iterative optimization of reinforcement learning, the agent can continuously learn and improve its strategy, improve the accuracy and precision of regulating the power system, and finally achieve the goal of maximizing the power system performance under the premise of meeting safety constraints.

[0068] As can be seen from the above, the present application provides a regulation method for a power system with embedded physical constraints, which is used to solve the problem that the line characteristics change after new energy equipment is connected to the power system. The present application first collects data that can characterize the operating state of the system (i.e., the data in the first set), so as to achieve the purpose of determining the adjustable range of the system based on the current operating state data (i.e., the data in the second set). Subsequently, the present application determines the objective function for regulating the power system (i.e., the first function) based on the operating state data of the system and the physical constraint information. Finally, the present application jointly regulates the power system based on the first set, the second set, the first function, and the preset equation. While improving the accuracy of regulating the power system, the present application also ensures the stable operation of the power system by the relationship between the node voltage and the branch power described by the preset equation, ensuring that the bus voltage obtained after subsequent regulation can meet the safety constraint conditions, thereby improving the stability of the regulated power system.

[0069] It can be seen that in the present application, the method of jointly regulating the power system based on the first set, the second set, the first function, and the preset equation ensures the stable operation of the power system through the relationship between the node voltage and the branch power described by the preset equation, thereby achieving the technical effects of improving the accuracy of regulating the power system and the stability of the regulated power system, and further solving the technical problems that the line characteristics change after the new energy equipment is connected to the power system and the stability of the power system operation is low.

[0070] In an optional embodiment, the regulation system first determines the power information of the unit nodes, the power information of the load nodes, the line load rate, and the voltage information of the busbars in the power system based on the operating data. Then, the regulation system determines the first set based on the power information of the unit nodes, the power information of the load nodes, the line load rate, and the voltage information of the busbars.

[0071] Optionally, the power information of the unit nodes includes at least the active power output P of the unit in the current period t 、the adjustable upper space of the active power output of the unit in the current period and the adjustable lower space of the active power output of the unit in the current period The power information of the load nodes includes at least the active power value D of the load in the current period t and the predicted active power value D of the load in the next period t+1 .

[0072] Optionally, the line load rate is represented by ρ t The line load rate refers to the ratio of the actual load of the transmission line to the maximum load capacity. It is directly related to the stability and safety of the power system. An excessively high line load rate may cause the line to be overloaded and affect the reliability of the power grid operation.

[0073] Optionally, the voltage information of the busbar includes at least the voltage amplitude V of the busbar in the current period t and the phase angle θ of the busbar in the current period t .

[0074] Combining the above, the obtained first set is expressed as Through the deep reinforcement learning technology, this application constructs a state space (the first set) based on the operation data of the power system (unit power, load power, line load rate, busbar voltage, etc.), achieving the purpose of providing necessary information for subsequent agent learning and optimizing control strategies.

[0075] In an optional embodiment, the regulation system first determines the power adjustment amount of the unit node, the voltage amplitude adjustment amount of the busbar, and the voltage phase angle adjustment amount of the busbar according to the operation data and the first set. Then, the regulation system determines the second set according to the power adjustment amount of the unit node, the voltage amplitude adjustment amount of the busbar, and the voltage phase angle adjustment amount of the busbar.

[0076] Optionally, the second set is expressed as a t =[△P 1,t ,…,△P g,t ,△V 1,t ,…△V b,t ,△θ 1,t ,…,△θ b,t . In the formula, △P i,t =P i,t+1 -P i,t represents the output adjustment value of unit i from time period t to time period t + 1; △V j,t =V j,t+1 -V j,t represents the voltage amplitude change value of busbar j from time period t to time period t + 1; △θ j,t =△θ j,t+1 -θ j,t represents the phase angle change value of busbar j from time period t to time period t + 1.

[0077] Optionally, the determination of the second set is a collective representation of the agent's control actions. The second set contains all possible control actions of the power system and is the concretization of the agent's action space. The functions of the second set include:

[0078] (1) Defining the action space: The second set clearly defines all actions that the agent can take during the regulation process, including but not limited to the adjustment amount of the power of each unit, the adjustment amount of the voltage amplitude and phase angle of each busbar. This provides clear action options for the agent and is the basis for the training of the deep reinforcement learning algorithm and the decision-making of the agent.

[0079] (2) Action Limitation: Through the second set, the limitations of the agent's actions can be set, such as the range of unit power adjustment, the range of bus voltage adjustment, etc. This helps the agent to abide by the actual operating rules of the power system when making decisions and avoid taking unrealistic actions or actions that may cause system instability.

[0080] (3) Policy Optimization: As part of the agent's action space, the second set participates in the process of policy optimization. The agent will explore different action sequences in the second set and learn how to take optimal regulation actions (actions in the second set) in the first set (state space) through interaction with the environment and feedback from the reward mechanism to achieve the goals of the power system regulation policy, such as reducing operating costs, increasing the new energy consumption rate, and ensuring the safe and stable operation of the power grid.

[0081] Combining the above content, it can be seen that through deep reinforcement learning technology, the agent can not only guide its regulation decisions based on the operating data and state information of the power system (the first set), but also quantify its regulation actions (the second set) and limit them within the actual feasible range, thus ensuring the effectiveness and safety of the regulation policy. This is a key link in realizing the intelligent optimal regulation of the power system and can help the agent make more reasonable and accurate decisions when dealing with complex and uncertain operating scenarios of the power system.

[0082] In an optional embodiment, the analysis system first determines the operating cost and new energy consumption rate of the power system based on the operating data of the power system. Then, the analysis system determines the first function based on the operating cost, new energy consumption rate, and constraint information.

[0083] Optionally, the training objective of the agent's continuous reinforcement learning is to find the optimal policy to maximize the expected value of the cumulative reward. In this application, the reward function (i.e., the first function) is defined as the weighted sum of the operating cost of thermal power units, the new energy consumption rate, and safety constraints.

[0084] Optionally, the functional expression of the first function is shown in the following formula (2).

[0085]

[0086] r = α1r1 + α2r2 + α3r3 (2)

[0087] In the above formula, and respectively represent the set of thermal power units and the set of wind power units, and c0, c1, and c2 respectively represent the constant term, linear coefficient, and quadratic coefficient related to the operating cost. It represents the total number of units. The lower the operating cost of thermal power units, the greater the reward. The higher the new energy consumption rate, the greater the reward. Safety constraints are imposed based on the stability of voltage amplitude and phase angle. β1 represents the voltage amplitude stability coefficient, β2 represents the phase angle stability coefficient, and V nom,j represents the nominal voltage of bus j. N represents the total number of buses. Buses m and n are located at both ends of the same line, represents the total number of lines. a1, a2, and a3 respectively represent the weights of rewards r1, r2, and r3.

[0088] In an optional embodiment, the regulation system first determines a second function according to a preset equation, where the second function is used to predict the operating state of the power system at a future moment. Then, the regulation system constructs a target model based on a preset framework, a first function, and the second function, where the target model is a deep reinforcement learning model obtained by training an initial neural network based on L sets of historical operating data of the power system, and L is a positive integer. Then, the regulation system inputs the first set and the second set into the target model to obtain a target policy, where the target policy at least includes the adjustment amount for actual adjustment of the nodes in the power system. Finally, the regulation system adjusts the power system according to the target policy.

[0089] Optionally, the second function is the transfer function corresponding to the agent (i.e., the target model) constructed according to the preset equation and the first function. An example of the application of this transfer function in the decision-making process of the agent is as follows: The transfer function can be used for the training environment. After the control agent executes an action, the unit adjustment value, the change value of the bus voltage amplitude, and the change value of the bus phase angle are input into the environment. The environment obtains the injection power of the nodes in the next time period based on the operating point of the current time period. Then the environment reads the load prediction value to obtain the injection power of the nodes in the next time period, and solves it through the power flow balance equation (i.e., the preset equation) in the power flow solver to obtain a new power grid state, that is, the voltages of all nodes and the branch powers.

[0090] Optionally, the preset framework is the Teacher-Tutor-Junior-Senior framework.

[0091] Optionally, applying the deep reinforcement learning model (i.e., the target model) embedded with physical knowledge to a power system where the characteristic values of some buses and branches change, so as to achieve the purpose of optimizing the control strategy of the power system. Then, select the optimal control strategy of the power system determined by the model, and then give an accurate control decision plan, thereby achieving the technical effect of maintaining the stable operation of the power grid.

[0092] As can be seen from the above, the combined use of the transfer function and the reward function in this application can support the iterative optimization of the agent's strategy. In each training iteration, the agent takes actions based on the current strategy, observes the state changes through the transfer function, and then adjusts its strategy parameters according to the feedback of the reward function. This process is repeated continuously until the agent learns how to adopt the optimal control strategy to maximize the comprehensive objective defined by the reward function under different operating conditions. By constructing the transfer function, this application improves the learning efficiency of the agent and the applicability of the agent's decision-making strategy.

[0093] In an alternative embodiment, the adjustment system first initializes a neural network based on a preset framework, a first function, a second function, and a first strategy to obtain a first model. The first strategy is used to evaluate the security of the power system when a fault occurs at any node in the power system, and the first model is used to interact with the operation data of the power system. Then, the adjustment system constructs a target model based on the first model.

[0094] Specifically, in the process of the adjustment system constructing the target model based on the first model: the adjustment system first updates the first model according to a second strategy to obtain a second model. The second strategy is used to determine the fault probability of nodes in the power system according to the topological structure of the power system. Then, the adjustment system constructs a third model based on a feedforward neural network and the second model. The input data of the third model is the output data of the second model, and the output data of the third model is used to represent all adjustment schemes that the power system can select based on the current operating state. Then, the adjustment system updates the third model based on the proximal policy optimization algorithm to obtain the target model.

[0095] Optionally, Figure 2 is a structural diagram of an alternative feedforward neural network according to an embodiment of the present application. As Figure 2 shown, the feedforward neural network structure includes: an input layer, a fully connected layer, a Dropout layer, and an output layer.

[0096] Optionally, the input layer is the first layer of the network and is responsible for receiving state information from the environment (Grid2Op environment in this application). The state information may include real-time parameters of the power grid, such as the active power output of units, the active power load value, the line load rate, the bus voltage amplitude, and the phase angle. The number of neurons in the input layer matches the dimension of the state space.

[0097] Optionally, each neuron in the fully connected layer is connected to every neuron in the previous layer. The fully connected layer passes information by weighted summing the input signals and adding a non-linear activation function (such as ReLU, sigmoid, or tanh). In Figure 2Among them, the network includes multiple fully connected layers, and each layer is connected by a Dropout layer. The Dropout layer is a technique for randomly deactivating neurons, which is used to reduce overfitting and improve the generalization ability of the model.

[0098] Optionally, the Dropout layer is located between the fully connected layers. During the training process, the Dropout layer randomly "drops out" a part of the neurons, that is, sets their outputs to 0, so as to simulate the sparse connection of the neural network, prevent neurons from relying too much on other neurons, and enhance the robustness and generalization performance of the network. The Dropout layer randomly drops out a part of the neurons during the training phase, while in the testing or inference phase, all neurons will be activated, but their weights will be adjusted to reflect the influence during training.

[0099] Optionally, the output layer is the last layer of the network and is responsible for generating the decisions of the agent. In the context of this application, the output layer includes all possible actions for regulating the power grid, such as the adjustment amount of unit output, the change amount of bus voltage amplitude, and the change amount of bus phase angle. The number of neurons in the output layer matches the dimension of the action space. It predicts the best action that the agent should take based on the input state and the processing results of the previous layers through an activation function (such as a linear activation function).

[0100] For example, in the case where the preset framework is the Teacher-Tutor-Junior-Senior framework, the first strategy is the N-1 fault strategy, and the second strategy is the topological regression strategy, the construction process of the target model is as follows:

[0101] (1) In the environment provided by Grid2Op, use the Teacher-Tutor-Junior-Senior framework proposed by Binbinchen for L2RPN (Learn to Run a Power Network, power system reinforcement learning). In the Teacher model part, the agent interacts with the environment, executes the selected actions, calculates the transfer function using the power flow balance equation, observes the state changes of the system, provides feedback rewards through the reward function to evaluate the effects of the actions, and selects actions that can keep the system stable. At the same time, use the N-1 fault strategy to ensure the stability of the power grid.

[0102] (2) In the Tutor model part, collect experience data by observing the environmental state and the selected actions for subsequent training of the Junior model, which helps to make effective decisions in the subsequent learning process, and set the maximum line load rate ρ max,t make a decision and select actions that can perform well in terms of power grid stability. If ρ max,t is lower than ρ tutorIf the threshold value is 0.9, the agent cannot interact with the grid. If ρ max,t is higher than ρ tutor with a threshold value of 0.9, the agent can interact with the grid. Meanwhile, a topological regression strategy is used to enhance the power grid stability and optimize the decision-making.

[0103] (3) In the Junior model part, by imitating the Tutor, it learns how to make decisions in power grid management, predicts the correct actions based on the input states, and thus provides the initial weights and experience for the subsequent Senior model. A feed-forward neural network structure is used, where the input layer corresponds to the observations of the Tutor, and the output layer corresponds to all possible topological actions.

[0104] (4) In the Senior model part, it has the same layer and neuron structure as the Junior, as well as the same types and quantities of inputs and outputs. It is trained by the PPO (Proximal Policy Optimization) algorithm. First, it is initialized with the weights and experience provided by the Junior model. When the maximum line load rate ρ max,t is lower than ρ senior with a threshold value of 0.9, the agent chooses not to execute an action. When the maximum line load rate ρ max,t is higher than ρ senior with a threshold value of 0.9, the agent chooses to execute an action, learns how to make effective decisions in different states, updates the agent's policy, and performs iterative optimization until the calculated value of the power flow balance equation converges to the target. Meanwhile, the N-1 fault strategy is used to enhance the power grid stability and optimize the decision-making.

[0105] Optionally, Figure 3 is a flowchart of an optional proximal policy optimization algorithm according to an embodiment of the present application. As Figure 3 shown, when the agent starts to interact with the environment, first, according to the current policy and the environmental state, the agent selects an action (such as adjusting the generator output or the bus voltage). Then, this action is sent to the environment, and the environment calculates the new state according to physical knowledge (such as the power flow equation) and gives a reward according to the effect of this action. Then, the agent records the state, action, reward, and new state in this interaction. These data are used to evaluate the performance of the current policy and serve as the input for subsequent training. The core of the PPO algorithm lies in that when updating the policy parameters, a method called "clip" is used to limit the amplitude of the policy update, avoiding too large changes in the policy, so as to maintain the "approximate" of the current policy and the updated policy, which helps to stabilize the learning process and avoid the training falling into the dilemma of local optimality.

[0106] Optionally, the Proximal Policy Optimization (PPO) algorithm is a policy gradient-based reinforcement learning algorithm that provides an efficient and easy-to-implement method for solving complex and high-dimensional control tasks. It aims to address issues such as unstable updates and low sample efficiency in traditional policy gradient algorithms and exhibits good performance when dealing with problems in continuous action spaces. Its core lies in optimizing a surrogate objective function through stochastic gradient ascent, using interaction sampling data with the environment to improve the policy. The PPO algorithm supports multiple mini-batch updates instead of only single updates for each sample.

[0107] Optionally, the full name of the N-1 contingency strategy is "N-1 Contingency Plan" or "N-1 Reliability Criterion". The N-1 contingency strategy is a core criterion for the safe operation of power systems. When the characteristic values of some buses and branches change, it can be used to evaluate the stability and reliability of the system and maintain the stable operation of the power system.

[0108] Optionally, the Topology Regression Strategy in the power system is a method that uses the network topology structure to optimize the system performance. When the characteristic values of some buses and branches change, by analyzing the topological characteristics of the power system, it can help improve the stability, reliability, and efficiency of the system.

[0109] In this application, the N-1 contingency strategy and the topology regression strategy are combined and applied to the deep reinforcement learning algorithm to optimize the control strategy of the power system. By analyzing the topological characteristics of the power grid, the agent can more accurately predict the impact of control actions on the system stability and select actions that can improve the system stability and optimize the decision-making. Especially when facing the challenge of changes in the characteristic values of some buses and branches, the topology regression strategy can help the agent make more effective and safer decisions in power grid management. This deep reinforcement learning method combined with physical knowledge can improve the learning efficiency and decision-making quality of the agent, ensuring that the power system can maintain stable operation under complex and changing operating conditions while optimizing costs and the accommodation of new energy.

[0110] Figure 4 is a flowchart of another optional power system regulation method according to an embodiment of the present application, as Figure 4 shown. The method includes three stages: selecting and defining the environment, generating a deep reinforcement learning model, and obtaining power system control decisions.

[0111] Optionally, in the stage of selecting and defining the environment, the present application first selects the Grid2Op environment to provide a platform for the development of the reinforcement learning agent. Then, in the Grid2Op environment, the state space, action space, reward function, and transition function are defined.

[0112] Optionally, in the stage of generating the deep reinforcement learning model, the present application controls the interaction between the agent and the environment. After that, the present application updates the regulation strategy of the power system according to the reward function and the transition function. After multiple iterative optimizations and trainings of the regulation strategy, when the output result of the power flow balance equation converges, the training is stopped to obtain the deep reinforcement learning model. Figure 5 It is a structural diagram of an optional deep reinforcement learning model according to an embodiment of the present application.

[0113] Optionally, in the stage of obtaining the power system regulation decision, the present application applies the deep reinforcement learning model embedded with physical knowledge to the power system in which the characteristic values of some buses and branches change, optimizes the regulation strategy of the power system, and then selects the optimal regulation strategy of the power system output by the model as the actual regulation strategy of the power system.

[0114] As can be seen from the above, the present application provides a regulation method for a power system embedded with physical constraints, which is used to solve the problem that the line characteristics change after new energy equipment is connected to the power system. The present application first collects data that can characterize the operating state of the system (i.e., the data in the first set), so as to achieve the purpose of determining the adjustable range of the system according to the current operating state data (i.e., the data in the second set). Subsequently, the present application determines the objective function for regulating the power system (i.e., the first function) according to the operating state data of the system and the physical constraint information. Finally, the present application jointly regulates the power system according to the first set, the second set, the first function, and the preset equation. While improving the accuracy of regulating the power system, the present application also ensures the stable operation of the power system by the relationship between the node voltage and the branch power described by the preset equation, and ensures that the bus voltage obtained after subsequent regulation can meet the safety constraint conditions, thereby improving the stability of the regulated power system.

[0115] Thus, in the present application, the method of jointly regulating the power system according to the first set, the second set, the first function, and the preset equation ensures the stable operation of the power system by the relationship between the node voltage and the branch power described by the preset equation, thereby achieving the technical effects of improving the accuracy of regulating the power system and the stability of the regulated power system, and further solving the technical problem that the line characteristics change after new energy equipment is connected to the power system and the stability of the power system operation is low.

[0116] According to another aspect of the embodiments of the present application, there is also provided a regulation device for a power system. Figure 6 It is a schematic diagram of an optional regulation device for a power system according to an embodiment of the present application, as Figure 6As shown in the figure, the regulating device of the power system includes: a first determination unit 601, a second determination unit 602, and a regulation unit 603.

[0117] Optionally, the first determination unit is configured to determine a first set and a second set according to the operation data of the power system, where the data in the first set is used to characterize the operation state of the power system, and the data in the second set is used to characterize the range of regulation of the power system; the second determination unit is configured to determine a first function according to the operation data of the power system and constraint information, where the first function is used to measure the performance of the power system, and the constraint information is used to perform safety constraints on the bus voltage in the power system; the regulation unit is configured to regulate the power system according to the first set, the second set, the first function, and a preset equation, where the preset equation is used to describe the relationship between the node voltage and the branch power during the stable operation of the power system.

[0118] In an alternative embodiment, the first determination unit includes: a first determination subunit and a second determination subunit.

[0119] Optionally, the first determination subunit is configured to determine the power information of the generator nodes, the power information of the load nodes, the line load rate, and the voltage information of the buses in the power system according to the operation data; the second determination subunit is configured to determine the first set according to the power information of the generator nodes, the power information of the load nodes, the line load rate, and the voltage information of the buses.

[0120] In an alternative embodiment, the first determination unit further includes: a third determination subunit and a fourth determination subunit.

[0121] Optionally, the third determination subunit is configured to determine the power adjustment amount of the generator nodes, the voltage amplitude adjustment amount of the buses, and the voltage phase angle adjustment amount of the buses according to the operation data and the first set; the fourth determination subunit is configured to determine the second set according to the power adjustment amount of the generator nodes, the voltage amplitude adjustment amount of the buses, and the voltage phase angle adjustment amount of the buses.

[0122] In an alternative embodiment, the second determination unit includes: a fifth determination subunit and a sixth determination subunit.

[0123] Optionally, the fifth determination subunit is configured to determine the operation cost and the new energy consumption rate of the power system according to the operation data of the power system; the sixth determination subunit is configured to determine the first function according to the operation cost, the new energy consumption rate, and the constraint information.

[0124] In an alternative embodiment, the regulation unit includes: a seventh determination subunit, a construction subunit, an input subunit, and a regulation subunit.

[0125] Optionally, a seventh determination subunit is configured to determine a second function according to a preset equation, where the second function is used to predict the operating state of the power system at a future moment; a construction subunit is configured to construct a target model based on a preset framework, a first function, and the second function, where the target model is a deep reinforcement learning model obtained by training an initial neural network based on L sets of historical operating data of the power system, and L is a positive integer; an input subunit is configured to input a first set and a second set into the target model to obtain a target policy, where the target policy at least includes an adjustment amount for actually adjusting nodes in the power system; an adjustment subunit is configured to adjust the power system according to the target policy.

[0126] In an alternative embodiment, the construction subunit includes: an initialization module and a construction module.

[0127] Optionally, the initialization module is configured to initialize a neural network based on a preset framework, a first function, a second function, and a first policy to obtain a first model, where the first policy is used to evaluate the security of the power system when any node in the power system fails, and the first model is used to interact with the operating data of the power system; the construction module is configured to construct a target model according to the first model.

[0128] In an alternative embodiment, the construction module further includes: a first update sub-module, a construction sub-module, and a second update sub-module.

[0129] Optionally, the first update sub-module is configured to update the first model according to a second policy to obtain a second model, where the second policy is used to determine the failure probability of nodes in the power system according to the topological structure of the power system; the construction sub-module is configured to construct a third model based on a feedforward neural network and the second model, where the input data of the third model is the output data of the second model, and the output data of the third model is used to represent all adjustment schemes that the power system can select based on the current operating state; the second update sub-module is configured to update the third model based on the proximal policy optimization algorithm to obtain the target model.

[0130] As can be seen from the above, the present application provides a regulation method for a power system with embedded physical constraints, which is used to solve the problem that the line characteristics change after new energy equipment is connected to the power system. The present application first collects data that can characterize the operating state of the system (i.e., the data in the first set), so as to achieve the purpose of determining the adjustable range of the system based on the current operating state data (i.e., the data in the second set). Subsequently, the present application determines the objective function for regulating the power system (i.e., the first function) based on the operating state data of the system and the physical constraint information. Finally, the present application jointly regulates the power system according to the first set, the second set, the first function, and the preset equation. While improving the accuracy of regulating the power system, the present application also ensures the stable operation of the power system by the relationship between the node voltage and the branch power described by the preset equation, and ensures that the bus voltage obtained after subsequent regulation can meet the safety constraint conditions, thereby improving the stability of the regulated power system.

[0131] Thus, in the present application, the method of jointly regulating the power system according to the first set, the second set, the first function, and the preset equation ensures the stable operation of the power system by the relationship between the node voltage and the branch power described by the preset equation, thereby achieving the technical effects of improving the accuracy of regulating the power system and the stability of the regulated power system, and further solving the technical problems that the line characteristics change after new energy equipment is connected to the power system and the stability of the power system operation is low.

[0132] According to another aspect of the embodiments of the present application, there is also provided a computer program product, which includes a stored computer program. When the computer program runs, it controls the computer program product to execute the regulation method of the power system in any one of the above.

[0133] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the regulation method of the power system in any one of the above by executing the executable instructions.

[0134] Optionally, Figure 7 is a schematic diagram of an optional electronic device according to an embodiment of the present application. As Figure 7 shown, the embodiments of the present application provide an electronic device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the regulation method of the power system in any one of the above.

[0135] The above-described embodiments or examples disclosed in this application are not exhaustive. They are only illustrations of some of the embodiments or examples and do not constitute specific limitations on the scope of protection disclosed in this application. Without conflict, each step in a certain embodiment or example in this application can be implemented as an independent embodiment, and the steps can be combined arbitrarily. For example, a solution obtained by removing some steps in a certain embodiment or example can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment or example can be arbitrarily exchanged. Additionally, the optional ways or optional examples in a certain embodiment or example can be combined arbitrarily; furthermore, the various embodiments or individual examples can be combined arbitrarily. For example, some or all of the steps of different embodiments or examples can be combined arbitrarily, and a certain embodiment or example can be combined arbitrarily with the optional ways or optional examples of other embodiments or examples.

[0136] In the above embodiments of this application, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0137] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of this application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0138] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0139] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1Steps of the functions specified in one or more boxes.

[0140] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory. The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0141] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0142] It should also be noted that the term "comprises", "comprising", or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, commodity, or device that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, commodity, or device that comprises the element.

[0143] Those skilled in the art will understand that the embodiments of the present application can be provided as a method, system, or computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0144] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A regulation method for a power system, characterized in that, Including: Determine a first set and a second set based on the operation data of the power system, where the data in the first set is used to characterize the operation state of the power system, and the data in the second set is used to characterize the adjustment range for the power system; Determine a first function based on the operation data and constraint information of the power system, where the first function is used to measure the performance of the power system, and the constraint information is used to perform safety constraints on the bus voltage in the power system; Adjust the power system according to the first set, the second set, the first function, and a preset equation, where the preset equation is used to describe the relationship between the node voltage and the branch power during the stable operation of the power system.

2. The adjustment method of the power system according to claim 1, characterized in that Determining a first set based on the operation data of the power system includes: Determine the power information of the unit nodes, the power information of the load nodes, the line load rate, and the voltage information of the buses in the power system according to the operation data; Determine the first set according to the power information of the unit nodes, the power information of the load nodes, the line load rate, and the voltage information of the buses.

3. The adjustment method of the power system according to claim 2, characterized in that, Determining a second set based on the operation data of the power system includes: Determine the power adjustment amount of the unit nodes, the voltage amplitude adjustment amount of the buses, and the voltage phase angle adjustment amount of the buses according to the operation data and the first set; Determine the second set according to the power adjustment amount of the unit nodes, the voltage amplitude adjustment amount of the buses, and the voltage phase angle adjustment amount of the buses.

4. The adjustment method of the power system according to claim 1, characterized in that Determining a first function based on the operation data and constraint information of the power system includes: Determine the operation cost and the new energy consumption rate of the power system according to the operation data of the power system; Determine the first function according to the operation cost, the new energy consumption rate, and the constraint information.

5. The adjustment method of the power system according to claim 1, characterized in that, Adjusting the power system according to the first set, the second set, the first function, and a preset equation includes: Determine a second function according to the preset equation, where the second function is used to predict the operation state of the power system at a future moment; Construct a target model based on a preset framework, the first function, and the second function, where the target model is a deep reinforcement learning model obtained by training an initial neural network based on L sets of historical operation data of the power system, and L is a positive integer; Input the first set and the second set into the target model to obtain a target policy, where the target policy at least includes the adjustment amount for actually adjusting the nodes in the power system; Adjust the power system according to the target policy.

6. The adjustment method of the power system according to claim 5, characterized in that, Constructing a target model based on a preset framework, the first function, and the second function includes: Initialize a neural network based on the preset framework, the first function, the second function, and a first policy to obtain a first model, where the first policy is used to evaluate the security of the power system when any node in the power system fails, and the first model is used to interact with the operation data of the power system; Construct the target model according to the first model.

7. The adjustment method of the power system according to claim 6, characterized in that, Constructing the target model according to the first model includes: Updating the first model according to a second strategy to obtain a second model, where the second strategy is used to determine the fault probability of nodes in the power system according to the topological structure of the power system; Constructing a third model according to a feedforward neural network and the second model, where the input data of the third model is the output data of the second model, and the output data of the third model is used to represent all adjustment schemes that can be selected by the power system based on the current operating state; Updating the third model based on the proximal policy optimization algorithm to obtain the target model.

8. An adjustment device for a power system, characterized in that, including: A first determination unit for determining a first set and a second set according to the operation data of the power system, where the data in the first set is used to represent the operation state of the power system, and the data in the second set is used to represent the range of adjustment to the power system; A second determination unit for determining a first function according to the operation data and constraint information of the power system, where the first function is used to measure the performance of the power system, and the constraint information is used to perform safety constraints on the bus voltage in the power system; An adjustment unit for adjusting the power system according to the first set, the second set, the first function, and a preset equation, where the preset equation is used to describe the relationship between the node voltage and the branch power during the stable operation of the power system.

9. A computer program product, characterized in that, The computer program product includes a computer program, where when the computer program runs, it controls the computer program product to execute the adjustment method of the power system according to any one of claims 1 to 7.

10. An electronic device, characterized in that, including one or more processors and a memory, the memory is used to store one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors implement the adjustment method of the power system according to any one of claims 1 to 7.

Citation Information

Cited By

  • Power grid fault handling method and system based on large model and multiple agents

    CN120746348A