System control method and device based on power system, storage medium and computer device
By acquiring target state information of the power system and using target update models and strategies to identify and respond to multi-source disturbances, the problem of high frequency sensitivity of the power system after renewable energy grid connection is solved, thereby improving the stability and security of the power system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTHEASTERN UNIV CHINA
- Filing Date
- 2026-01-19
- Publication Date
- 2026-04-10
AI Technical Summary
After renewable energy is connected to the grid, the power system becomes more sensitive to multi-source disturbances, which can lead to a rapid drop in frequency and potentially cause large-scale power outages. Existing control methods are difficult to effectively deal with multi-source coupled disturbances, resulting in frequency fluctuations and economic losses.
By acquiring target state information of the power system, and using target update models and strategies to identify and respond to engine mechanical disturbances, load disturbances, and grid power disturbances, the system can upgrade its proactive countermeasures and passive compensation capabilities, thereby improving its adaptability to complex operating conditions.
It effectively suppresses large fluctuations in system frequency, reduces the probability of low-frequency load shedding, reduces load loss, and improves the operational stability and reliability of the power system.
Smart Images

Figure CN121529994B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power control, and in particular to a system control method and device based on a power system, a storage medium and a computer device. BACKGROUND
[0002] With the continuous development of renewable energy, the electric energy generated by the power generation system based on renewable energy will be connected to the public power grid to realize the transmission of electric energy. However, after the renewable energy is connected to the grid, the sensitivity of the system frequency to the disturbance will be greatly improved.
[0003] Therefore, if the power system is subjected to multiple source disturbances such as engine mechanical disturbance and load disturbance, the system frequency will rapidly decrease, thereby causing a large-area power failure accident and seriously affecting social stability and energy security. SUMMARY
[0004] Therefore, the embodiments of the present application provide a system control method and device based on a power system, a storage medium and a computer device, which can improve the anti-disturbance ability of the power system and ensure the safe and reliable operation of the power system.
[0005] In a first aspect, the present application provides a system control method based on a power system, comprising:
[0006] obtaining target state information of the power system; wherein the target state information is system load or system frequency;
[0007] updating the target state information according to a target update model; wherein the target update model comprises a target update strategy, and the target update strategy is used to enable the target update model to identify and respond to multiple source disturbances, the multiple source disturbances including at least two of engine mechanical disturbance, load disturbance and power grid power disturbance;
[0008] controlling the operation of the power system according to the updated target state information.
[0009] In a second aspect, the present application provides a system control device based on a power system, comprising:
[0010] a state acquisition module configured to obtain target state information of the power system; wherein the target state information is system load or system frequency;
[0011] a state update module configured to update the target state information according to a target update model; wherein the target update model comprises a target update strategy, and the target update strategy is used to enable the target update model to identify and respond to multiple source disturbances, the multiple source disturbances including at least two of engine mechanical disturbance, load disturbance and power grid power disturbance;
[0012] A system control module is configured to control the power system according to the updated target state information.
[0013] In a third aspect, the present application provides a storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the system control method.
[0014] In a fourth aspect, the present application provides a computer device, comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, wherein the processor implements the system control method when executing the program.
[0015] According to the above technical solution, the system control method, device, storage medium and computer device provided by the embodiments of the present application can effectively identify and cope with multi-source coupled disturbances and improve the adaptability of the target update model to complex working conditions by updating the target state information of the power system through the target update model comprising a target update strategy. Then, the power system is regulated based on the updated target state information, which can compensate for power imbalance caused by multi-source disturbances in a timely manner, suppress large fluctuations in system frequency, and further improve the operation stability and safety and reliability of the power system.
[0016] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application, the above description can be implemented according to the content of the specification, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 A flowchart of a system control method based on a power system is shown;
[0019] Figure 2 A flowchart of a method for determining a target update strategy is shown;
[0020] Figure 3 A flowchart of an adjustment process of an estimation update strategy is shown;
[0021] Figure 4 A flowchart of a method for determining a target disturbance strategy is shown;
[0022] Figure 5 A flowchart of a method for determining a target value function is shown.
[0023] Figure 6 A structural diagram of a system control device based on a power system is shown.
[0024] Figure 7 A structural diagram of another system control device based on a power system is shown. DETAILED DESCRIPTION
[0025] To facilitate the description of the embodiments of the present application, some technical terms and technical means related to the embodiments of the present application, and the application scenarios of the embodiments of the present application will be introduced first.
[0026] After large-scale grid connection of wind power, photovoltaic and other renewable energy sources, in order to accommodate the power generation of renewable energy, the grid will gradually reduce the starting capacity and power generation output of traditional power sources such as thermal power and hydro power that rely on synchronous generators, that is, it will lead to a decrease in the proportion of synchronous generators in the power system. But with the decrease in the proportion of synchronous generators, the source devices that can provide inertia in the power system are reduced, which in turn leads to a significant reduction in the equivalent inertia of the power system.
[0027] The greater the system inertia, the stronger the system frequency's ability to respond to engine mechanical disturbances, load disturbances and other disturbances, and the slower the frequency change rate. Therefore, if the equivalent inertia of the power system is reduced, the buffering effect of the synchronous generator rotor kinetic energy will be lacking, which in turn greatly increases the sensitivity of the system frequency to disturbances, that is, the stability of the system frequency is poor when facing various disturbances.
[0028] Further, when facing engine mechanical disturbances, load disturbances, power grid power disturbances and other multi-source disturbances, the system frequency will rapidly drop, which may trigger Under-Frequency Load Shedding (UFLS) to cause a large amount of load loss, or even cause a large-area power outage accident, which seriously affects social stability and energy security. Among them, engine mechanical disturbances are caused by fluctuations in wind turbine output, load disturbances are caused by random changes in residential and industrial electricity consumption, and power grid power disturbances are caused by parameter drift of power transmission lines and transformers.
[0029] In some cases, a load shedding action is triggered by pre-setting static parameters. In this way, although the above-mentioned multi-source disturbances can be coped with, excessive reliance on pre-set static parameters may not be able to adapt to the time-varying characteristics of multi-source coupled disturbances, which is prone to cause problems such as frequency out-of-limit due to insufficient load shedding or economic loss caused by excessive load shedding, and it is difficult to achieve optimal control effect.
[0030] In some cases, a fast frequency response (FFR) technique is used, i.e., through energy storage equipment charging and discharging, power electronic equipment adjustment, or demand side response, to achieve fast balance for power shortage, and then to stabilize the system frequency. In this way, although fast response can be achieved when the system frequency fluctuates, passive power balance compensation is used, and adjustment lag may occur in a fast time-varying disturbance scenario, and the collaborative optimization capability of different control means (energy storage, demand side response, etc.) is insufficient.
[0031] In some other cases, a virtual inertia control (VIC) technique is used, i.e., through power electronic converter control, to let wind power, photovoltaic, energy storage, and other power electronic interface power sources simulate the inertia and damping characteristics of synchronous generators, to quickly adjust active power during frequency disturbance, to suppress the frequency change rate (df / dt), and to improve the system frequency stability. However, the VIC technique is based on the frequency change rate to adjust the electromagnetic power, which may cause secondary frequency fluctuations, and does not consider the dynamic interaction of control input constraints and disturbances, making it difficult to balance system frequency safety and control economy.
[0032] Therefore, in order to improve the anti-disturbance ability of the power system and ensure the safe and reliable operation of the power system, the embodiments of the present application provide a system control method based on a power system, which can be used in the system control scenario of the power system facing multi-source disturbances. In the method, target state information of the power system is obtained. The target state information is system load or system frequency. Then, the target state information is updated according to a target update model. The target update model includes a target update strategy, and the target update strategy is used to enable the target update model to identify and respond to multi-source disturbances, including at least two of engine mechanical disturbance, load disturbance, and power grid power disturbance. Then, the power system is controlled to operate according to the updated target state information.
[0033] In the embodiments of the present application, after obtaining the target state information of the power system, the target state information is updated through the target update model including the target update strategy, which can effectively identify and respond to multi-source coupled disturbances and improve the adaptability of the target update model to complex working conditions. Then, based on the updated target state information, the power system is regulated to operate, which can compensate for power imbalance caused by multi-source disturbances, suppress large fluctuations in system frequency, reduce the probability of UFLS triggering, reduce load loss, and thus improve the operation stability and safety and reliability of the power system.
[0034] In some examples, the computer device in the embodiments of the present application can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) \ virtual reality (VR) device, and the like, which has a state information updating function. The embodiments of the present application do not specially limit the specific form of the computer device.
[0035] In other examples, the subject performing the embodiments of the present application can also be a server, which can be a single server, a server cluster, a distributed server, a centralized server, a cloud server, or a computer, and the like, and the specific form is not limited.
[0036] Embodiments of the present application will be described in more detail below with reference to the accompanying drawings. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0037] The embodiments provide a system control method based on a power system, as shown in Figure 1 The method comprises:
[0038] S101, obtaining target state information of the power system.
[0039] The power system can be a low-inertia power system. The low-inertia power system is a power system in which the equivalent rotational inertia of the system is significantly reduced, the buffering capacity of power disturbance is greatly weakened, and the frequency dynamics and stability are facing severe challenges after a high proportion of renewable energy and power electronic equipment are connected to the grid.
[0040] In some embodiments of the present application, the target state information can be system load or system frequency. The system load (or system load) is the total active power or total energy consumption of the power user from the power grid at a certain time or time period, which is used to reflect the operation state of the power grid in the power system. The system frequency refers to the alternating frequency of the alternating current in the power grid, which is used to measure the operation stability of the power system.
[0041] In other embodiments of the present application, the target state information can also be the system speed. The system speed is the rotational speed of the rotor of the synchronous generator in the power system, which is used to determine the frequency of the alternating current generated by the synchronous generator.
[0042] S102, updating the target state information according to the target update model.
[0043] The target updating model is composed of an inherent dynamics term, a multi-source disturbance term, and an information updating term. The multi-source disturbance term includes a disturbance input matrix and a target disturbance strategy, and the target disturbance strategy is used to actively suppress the multi-source disturbance by the target updating model. The information updating term includes an updating input matrix and a target updating strategy. The target updating strategy is used to enable the target updating model to identify and respond to the multi-source disturbance. The multi-source disturbance includes at least two of the engine mechanical disturbance, the load disturbance, and the power grid power disturbance.
[0044] In the embodiments of the present application, by including the above multi-source disturbance into the game system and supporting it with the accurate target updating model, the upgrade from passive disturbance compensation to active confrontation is realized, and the response capability of the power system to the multi-source disturbance is greatly improved, thereby improving the updating accuracy of the target state information.
[0045] In an implementation manner, the target updating model can be characterized by expression one as shown below:
[0046] Expression one;
[0047] wherein, is the updated target state information, that is, the target state information at the moment k; is the state vector corresponding to the obtained target state information, that is, the target state information at the moment k; , , are all smooth and differentiable system model functions; is an inherent dynamics term (or referred to as an inherent dynamics function), which is used to describe the state change rule without updating and disturbance; is an updating input matrix, which is used to characterize the influence of the updating input on the target state information; is a target updating strategy (or referred to as a bounded updating input), which is used to enable the target updating model to identify and respond to the multi-source disturbance; is a disturbance input matrix, which is used to characterize the influence of the multi-source disturbance on the target state information; is a target disturbance strategy, which is used to characterize the multi-source disturbance.
[0048] It should be noted that, in , and in the case that the updating input matrix f(0)=0 of the above target updating model, the power system is in a balanced state, and the target updating model satisfies the Lipschitz continuity on a compact set containing the origin.
[0049] S103, controlling the power system to operate according to the updated target state information.
[0050] Specifically, corresponding to the target state information update completion, the updated target state information can be installed, and the power system operation is controlled. In this way, the power imbalance caused by multi-source disturbance can be compensated in time, the system frequency fluctuation is suppressed, and the operation stability and safety and reliability of the power system are improved.
[0051] The above introduces how to update the target state information to control the safe and stable operation of the power system. The determination process of the target update strategy involved in the state information update process will be described in detail below. For example Figure 2 As shown in the figure, the determination process of the target update strategy can specifically include S200-S206:
[0052] S200, the target control device adds multi-source disturbance to the power system.
[0053] The target control device can be an adaptive dynamic programming (ADP) control device.
[0054] In some cases, in order to improve the accuracy of the determination of the above target update strategy, a target scene can be constructed. Specifically, the center of inertia frequency of the power system and the optimization target of the target update model are set. The center of inertia (COI) frequency is used to represent the overall dynamic frequency of the power system. The optimization target includes minimizing the frequency deviation and load shedding cost.
[0055] In one implementation, the COI frequency can be obtained by expression two as follows:
[0056] Expression two;
[0057] Wherein, is the COI frequency at time k; n is the number of generators; is the inertia of the i-th generator; is the speed of the i-th generator at time k.
[0058] In another implementation, the optimization target of the target update model can be obtained by expression three as follows:
[0059] Expression three;
[0060] Wherein, is the optimization target, i.e. minimizing the frequency deviation and load shedding cost; is the minimum frequency of the power system; is the steady-state frequency threshold; is a weighting coefficient; m is the number of load shedding nodes; is a load importance factor of the load shedding node, which is used to quantify the importance of the load; is the load shedding amount of the jth node at the kth moment.
[0061] In some cases, in order to ensure that the above update strategy can meet the physical limit of power grid operation in the power system, thereby reducing the unsafe or infeasible load shedding action, the load shedding amount constraint may be set. is the maximum allowed load shedding amount of the jth node, which reduces the occurrence of local large-area power outage caused by excessive load shedding of a single node, and controls the upper limit of the cost.
[0062] Further, the power system can be added with a generator mechanical disturbance according to a target dynamic model of the generator in the power system.
[0063] In an implementation manner, the target dynamic model of the generator can be characterized by expression four as shown in the following:
[0064] ,
[0065] ,
[0066] ,
[0067] Expression four;
[0068] wherein, is a generator d-axis voltage (kV); is a voltage after d-axis transient reactance of the generator; is a d-axis transient reactance (Ω) of the generator; is a stator resistance (Ω) of the generator; is a d-axis component of the stator current of the generator (A); is a generator q-axis voltage (kV); is a voltage after q-axis transient reactance of the generator; is a q-axis transient reactance (Ω) of the generator; is a q-axis component of the stator current of the generator (A); is a power angle of the ith generator, which is used to represent an included angle between a stator magnetic field and a rotor magnetic field; is a sampling interval (s), i.e., an update period of the update strategy; is a rotating speed (rad / s) of the ith generator at the kth moment; is a synchronous rotating speed (rad / s), which is a constant (such as 314 rad / s); mechanical power of the i-th generator (MW); electromagnetic power of the i-th generator at time k; damping coefficient of the i-th generator (MW(rad / s)), which is used to represent the damping effect when the rotating speed changes; mechanical power disturbance of the i-th generator, such as the mechanical power change caused by the fluctuation of the wind turbine output.
[0069] Further, the power system can be added with a load disturbance according to a target load model.
[0070] In an implementation manner, the above target load model can be represented by expression five shown as follows:
[0071] ,
[0072] Expression five;
[0073] wherein, active power load injection of the j-th node (MW); original active power load of the j-th node (MW); active power load shedding amount of the j-th node (MW); active power random fluctuation amount of the j-th node (MW), which is used to represent the load random change disturbance; reactive power load injection of the j-th node (MVar); original reactive power load of the j-th node (MVar); reactive power load shedding amount of the j-th node (MVar); reactive power random fluctuation amount of the j-th node (MVar), which is used to represent the load random change disturbance.
[0074] Further, the power grid power disturbance scenario can be set according to the static power flow constraint and the target increment constraint. Specifically, according to the static power flow constraint, the active power injection and the reactive power injection of the power system at time k are determined. The active power injection is used to calculate the total active power increment. The reactive power injection is used to calculate the total reactive power increment. Then, according to the total active power increment and the total reactive power increment of the power system at time k and the target increment constraint, the power grid power disturbance is added to the power system.
[0075] In an implementation manner, the above static power flow constraint can be represented by expression six shown as follows:
[0076] ,
[0077] Expression six;
[0078] wherein, Pi is the active power injection (MW) of the i-th node; Vi is the voltage magnitude (kV) of the i-th node; Vj is the voltage magnitude (kV) of the j-th node; Gi,j is the mutual conductance (s) between node i and node j, which is determined by the power grid topology and parameters; θi,j(k) is the phase angle difference (rad) between node i and node j at time k; Bi,j is the susceptance (s) between node i and node j, which is determined by the power grid topology and parameters; Qi is the reactive power injection (MVar) of the i-th node.
[0079] In another implementation, the above target incremental constraint can be characterized by expression seven as shown below:
[0080] Expression seven;
[0081] wherein, ΔP(k) is the total active power increment (MW) of the power system at time k; J(k) is the Jacobian matrix of the power flow equation at time k, which is used to characterize the mapping relationship between voltage, phase angle change and power change; Δθ(k) is the relative increment vector of system node voltage and phase angle (rad) at time k; ΔPagg(k) is the total active power aggregation disturbance vector (MW) of the whole network at time k, which is used to integrate power grid power disturbances such as network parameter drift; ΔQ(k) is the total reactive power increment (MVar) of the power system at time k; ΔV(k) is the relative increment vector of system node voltage and magnitude at time k; ΔQagg(k) is the total reactive power aggregation disturbance vector (MVar) of the whole network at time k, which is used to integrate power grid power disturbances such as network parameter drift.
[0082] In the embodiments of the present application, the multi-source disturbance including at least two of the generator mechanical power disturbance, random load fluctuation and grid parameter drift is uniformly included in the zero-sum game framework and taken as the opponent, and the optimization target of the target control device is set as the collaborative optimization of frequency safety and load shedding cost. In this way, the integrated design of disturbance active confrontation and control optimization can be realized, the problem of incomplete multi-source disturbance coupling characterization and lack of exclusive system model support is solved, and necessary conditions are provided for subsequent accurate updating of system state information.
[0083] S201, the power system adds a multi-source disturbance, and sends system state information at a first time to a target control device.
[0084] S202, in a case where the target control device receives the system state information of the first time sent by the power system, the target control device takes the product between the transpose matrix of the first estimated weight matrix and the activation function vector of the first network as an estimation update strategy.
[0085] In an implementation manner, the estimation update strategy can be obtained by expression eight as shown below:
[0086] Expression eight;
[0087] wherein, is the estimation update strategy; is the transpose matrix of the first estimated weight matrix ; and is the activation function vector of the first network, and the first network can be an Actor neural network.
[0088] In other cases, after receiving the system state information of the first time, the target control device can take the product between the transpose matrix of the second estimated weight matrix and the activation function vector of the second network as an estimation disturbance strategy. Specifically, the estimation disturbance strategy can be obtained by expression nine as shown below:
[0089] Expression nine;
[0090] wherein, is the estimation disturbance strategy; is the transpose matrix of the second estimated weight matrix ; and is the activation function vector of the second network, and the second network can be a Disturbance neural network.
[0091] Then, the target control device can send the estimation disturbance strategy to the power system, and continue to run according to the estimation update strategy and the estimation disturbance strategy to generate the system state information of the second time.
[0092] In yet other cases, after receiving the system state information of the first time, the target control device can take the product between the transpose matrix of the third estimated weight matrix and the activation function vector of the third network as an approximate value function. Specifically, the approximate value function can be obtained by expression ten as shown below:
[0093] Expression ten;
[0094] wherein, is the approximate value function; is the transpose matrix of the third estimated weight matrix ; and The activation function vector of the third network is a vector of the third network, and the third network can be a Critic neural network.
[0095] S203, the target control device sends an estimation update strategy to the power system.
[0096] S204, the power system continues to run according to the estimation update strategy and generates system state information at the second time point after receiving the estimation update strategy of the target control device.
[0097] In some cases, the power system can also generate a running cost after running according to the estimation update strategy. The running cost refers to the sum of various costs and potential losses generated in the whole link of power generation, power transmission, power distribution and dispatching to maintain safe, stable and economic power supply, which can be divided into direct running cost and indirect cost (or called hidden cost). For example, the direct running cost can be fuel and operation and maintenance cost on the power generation side, loss and operation and maintenance cost on the power transmission and distribution side, etc. The indirect cost can be reliability loss cost, environmental and social cost, etc.
[0098] S205, the power system sends the system state information at the second time point to the target control device.
[0099] S206, the target control device adjusts the estimation update strategy according to the system state information at the first time point and the system state information at the second time point after receiving the system state information at the second time point sent by the power system, until the received system state information is the same as the system state information at the first time point, and obtains a target update strategy.
[0100] Specifically, after receiving the system state information at the second time point, the estimation update strategy can be adjusted according to the system state information at the second time point and the system state information at the first time point.
[0101] In an implementation mode, as shown in Figure 3 the adjustment process of the estimation update strategy can specifically include:
[0102] S2061, input the system state information at the first time point and the system state information at the second time point into the Hamilton function, and perform gradient solution on the update strategy included in the Hamilton function to obtain a theoretical update strategy.
[0103] Specifically, after receiving the system state information at the second time, the system state information at the second time and the system state information at the first time can be input into a Hamilton function. The Hamilton function is used to solve a theoretical policy pair (or a zero-sum differential game). The theoretical policy pair can include a theoretical update policy and a theoretical disturbance policy. That is, the Hamilton function can be used to obtain the theoretical update policy and the theoretical disturbance policy.
[0104] In an implementation manner, the Hamilton function can be represented by expression XI as shown below:
[0105] Expression XI;
[0106] wherein, is a Hamilton function, which is used to represent an instantaneous comprehensive effect of system state transition, update input and disturbance input at the kth time; is a value function corresponding to the system state information at the second time; is a value function corresponding to the system state information at the first time; is a state penalty term, which is used to represent a penalty of state deviating from an ideal value; is an input cost term, which is used to represent an economic cost of the update policy; is a target coupling term (or an update-disturbance coupling term), which is used to construct a penalty constraint for representing an association between the update policy and the multi-source disturbance; is a disturbance penalty term, which is used to suppress extreme disturbance of the power system.
[0107] Further, in response to the input of the system state information at the second time and the system state information at the first time being completed, the target control device can perform gradient solving on the update policy u included in the Hamilton function to obtain the theoretical update policy.
[0108] In an implementation manner, the theoretical update policy can be represented by expression XII as shown below:
[0109] Expression XII;
[0110] wherein, is the theoretical update policy; is an inverse matrix of the weight matrix R, which is used to adjust sensitivity of the update policy; is a transpose matrix of the update input matrix ; is a gradient of the theoretical value function at the second time, which is used to represent an influence of the state value at the second time on the current decision.
[0111] S2062, taking the difference between the theoretical update strategy and the estimated update strategy as a network error vector of the first network.
[0112] In an implementation manner, the network error vector of the first network can be obtained by expression thirteen as follows:
[0113] Expression thirteen;
[0114] wherein, is the network error vector of the first network, which is used to reflect the difference between the theoretical update strategy and the estimated update strategy; is the estimated update strategy; is the negative value of the theoretical update strategy.
[0115] S2063, inputting the network error vector of the first network into a loss function of the first network to obtain a first loss value.
[0116] In an implementation manner, the loss function of the first network can be characterized by expression fourteen as follows:
[0117] Expression fourteen;
[0118] wherein, is the loss value (or referred to as the first loss value) of the first network; is the transpose vector of the network error vector of the first network.
[0119] S2064, updating the first estimated weight matrix according to the first loss value until a first preset condition is met, obtaining a target first estimated weight matrix, and taking the product between the transpose matrix of the target first estimated weight matrix and an activation function vector of the first network as a target update strategy.
[0120] Specifically, after obtaining the first loss value, gradient descent can be performed on the first loss value to obtain a first weight update function, and the first estimated weight matrix is updated according to the first weight update function until a first preset condition is met, obtaining a target first estimated weight matrix. The first preset condition can include that the number of updates is greater than a first preset number, and / or the network error vector of the first network is less than a first preset error vector (such as 0).
[0121] In an implementation manner, the first weight update function can be characterized by expression fifteen as follows:
[0122] Expression fifteen;
[0123] wherein, a first weight update matrix for a next time instant; a first weight update matrix for a current time instant; a network learning rate of the first network, which is used to control a step size of the offloading instruction (or referred to as an update strategy) optimization; a network error vector of the first network a transpose vector of the network error vector.
[0124] After obtaining the target update strategy, the target update model can be constructed according to the target update strategy.
[0125] Further, the target control device can also adjust the estimated disturbance strategy to obtain a target disturbance strategy. Specifically, as shown in Figure 4 the determination process of the target disturbance strategy can include:
[0126] S401, inputting the system state information at the first time instant and the system state information at the second time instant into the Hamilton function, and performing gradient solving on the disturbance strategy included in the Hamilton function to obtain a theoretical disturbance strategy.
[0127] Specifically, after receiving the system state information at the second time instant, the system state information at the second time instant and the system state information at the first time instant can be input into the Hamilton function. Then, in response to the input of the system state information at the second time instant and the system state information at the first time instant being completed, the target control device can perform gradient solving on the disturbance strategy w included in the Hamilton function to obtain a theoretical disturbance strategy.
[0128] In an implementation manner, the theoretical disturbance strategy can be represented by expression sixteen as shown below:
[0129] Expression sixteen;
[0130] wherein, the theoretical disturbance strategy; an inverse matrix of a disturbance penalty matrix which is used to adjust the sensitivity of the disturbance strategy; a transpose matrix of a coupling weight matrix ; and a transpose matrix of a disturbance input matrix ; and a gradient of the theoretical value function at the second time instant, which is used to represent the influence of the state value at the second time instant on the current decision.
[0131] S402, taking the product between the transpose matrix of the second estimated weight matrix and the activation function vector of the second network as the estimated disturbance strategy.
[0132] Specifically, after receiving the system state information at the second time point, a product between a transpose matrix of the second estimated weight matrix and an activation function vector of the second network can be taken as the estimated perturbation strategy.
[0133] S403, taking a difference between the estimated perturbation strategy and the theoretical perturbation strategy as a network error vector of the second network.
[0134] In an implementation manner, the network error vector of the second network can be obtained by expression seventeen as follows:
[0135] Expression seventeen;
[0136] wherein, is the network error vector of the second network, which is used to reflect a difference between the theoretical perturbation strategy and the estimated perturbation strategy, that is, the smaller the network error vector of the second network, the closer the simulated perturbation to the real perturbation characteristic; is the estimated perturbation strategy; is the theoretical perturbation strategy.
[0137] S404, inputting the network error vector of the second network into a loss function of the second network to obtain a second loss value.
[0138] In an implementation manner, the loss function of the second network can be characterized by expression eighteen as follows:
[0139] Expression eighteen;
[0140] wherein, is the loss value (or referred to as the second loss value) of the second network; is the network error vector of the second network is a transpose vector of the network error vector of the second network.
[0141] S405, updating the second estimated weight matrix according to the second loss value until a second preset condition is met, obtaining a target second estimated weight matrix, and taking a product between a transpose matrix of the target second estimated weight matrix and an activation function vector of the second network as a target perturbation strategy.
[0142] Specifically, after obtaining the second loss value, gradient descent can be performed on it to obtain a second weight update function. Based on this function, the second estimated weight matrix is updated until a second preset condition is met, resulting in the target second estimated weight matrix. This second preset condition may include an update count greater than a second preset number of updates, and / or the network error vector of the second network being less than a second preset error vector (e.g., 0). In other words, this embodiment of the application achieves the process of approximating a theoretical disturbance strategy through online iteration of the second network, improving the robustness of state updates under multi-source disturbances, thereby enhancing the operational stability and reliability of the power system.
[0143] In one implementation, the second weight update function described above can be represented by the following expression nineteen:
[0144] Expression 19;
[0145] in, Update the second weight matrix for the next time step; Update the second weight matrix at the current time step; is the network learning rate of the second network, which is used to control the optimization step size for perturbation simulation accuracy; The network error vector of the second network The transpose of .
[0146] After obtaining the target perturbation strategy, the target update model can be constructed based on the target perturbation strategy and the target update strategy.
[0147] Furthermore, the target control device can also adjust the aforementioned approximate function to obtain the target value function. Specifically, for example... Figure 5 As shown, the process of determining the above objective function may specifically include:
[0148] S501, the reward term is obtained based on the product of the difference between the third estimated weight matrix and the activation function vector and the output value of the approximate Hamiltonian function.
[0149] The approximate Hamiltonian function is a Hamiltonian function that includes both the estimated perturbation policy and the estimated update policy. The activation function vector difference is the difference between the activation function vector at the second time step and the activation function vector at the first time step. It can be understood that the output value of the approximate Hamiltonian function is obtained by inputting the system state information at the first time step and the system state information at the second time step into the approximate Hamiltonian function.
[0150] In one implementation, the above approximate Hamiltonian function can be characterized by the following expression:
[0151] Expression twenty;
[0152] wherein, is an output value of the approximate Hamiltonian function; is a third estimated weight matrix; is an activation function vector difference, which is a difference between an activation function vector at a second time and an activation function vector at a first time; is a reward term.
[0153] S502, according to the transpose matrix of the third estimated weight matrix, the activation function vector difference, and the reward term, an auxiliary error vector is calculated.
[0154] wherein, the auxiliary error vector is used to represent a gap between an approximate value function and a theoretical value function. The approximate value function is a product between the transpose matrix of the third estimated weight matrix and an activation function vector of the third network. The theoretical value function is a function obtained after the theoretical update strategy and the theoretical perturbation strategy are brought back to the Hamiltonian function. In some cases, the theoretical value function can also be referred to as an HJI equation, which is a class of first-order nonlinear partial differential equations.
[0155] In an implementation manner, the theoretical value function can be represented by Expression twenty-one as shown below:
[0156] Expression twenty-one;
[0157] wherein, is a theoretical value function; is a state penalty term, which is used to represent a penalty of a state deviating from an ideal value; is a weight matrix; is a coupling weight matrix; is a perturbation penalty matrix.
[0158] Specifically, the auxiliary error vector can be obtained by Expression twenty-two as shown below:
[0159] Expression twenty-two;
[0160] wherein, is an auxiliary error vector; is a transpose matrix of the third estimated weight matrix ; and is an activation function vector difference set, which includes an activation function vector difference at a current time and an activation function vector difference at a j time before the current time; is a reward term set, which includes a reward term at the current time and a reward term at the j time before the current time.
[0161] S503, input the auxiliary error vector into a loss function of the third network to obtain a third loss value.
[0162] In an implementation manner, the loss function of the third network can be represented by expression twenty-three as shown in the following:
[0163] Expression twenty-three;
[0164] wherein, is the loss value of the third network (or referred to as the third loss value); is a transpose vector of the network error vector of the third network .
[0165] S504, update the third estimated weight matrix according to the third loss value until a third preset condition is met to obtain a target third estimated weight matrix, and take a product between a transpose matrix of the target third estimated weight matrix and an activation function vector of the third network as a target value function.
[0166] Specifically, after the third loss value is obtained, gradient descent can be performed on the third loss value to obtain a third weight update function.
[0167] In an implementation manner, the third weight update function can be represented by expression twenty-four as shown in the following:
[0168] Expression twenty-four;
[0169] wherein, is the third weight update matrix at a next moment; is the third weight update matrix at a current moment; is a network learning rate of the third network; is an activation function vector difference set, which includes an activation function vector difference at the current moment and at a j-th moment before the current moment; is a transpose vector of the auxiliary error vector of the third network .
[0170] Further, the third estimated weight matrix is updated according to the third weight update function until a third preset condition is met to obtain a target third estimated weight matrix. The third preset condition can include that the number of updates is greater than a third preset number, and / or the network error vector of the first network is less than a third preset error vector (such as 0). Then, a product between a transpose matrix of the target third estimated weight matrix and an activation function vector of the third network is taken as a target value function. The target value function is used to evaluate the performance of the target updated model.
[0171] In an implementation manner, the above-mentioned target value function can be characterized by expression twenty-five as shown below:
[0172] Expression twenty-five;
[0173] wherein, is a target value function, which is used to characterize the comprehensive performance of the long-term operation of the power system from the k time, and the k time refers to the time when the system state information is acquired in step S101; is a state penalty term, which is used to characterize the penalty of the state deviating from the ideal value, that is, the greater the state information deviation, the greater the output value of the state penalty term; is an input cost term, which is used to characterize the economic cost of the update strategy; is a target coupling term (or an update-disturbance coupling term), which is used to construct a penalty constraint, which is used to characterize the association between the update strategy and the multi-source disturbance, that is, to quantify the dynamic interaction relationship between the update strategy and the multi-source disturbance, so as to adapt to the control requirements of the low-inertia power system, so as to ensure the consistency of the constraints and the system dynamics; is a coupling weight matrix, which is used to adjust the influence strength of the interaction between the update strategy and the multi-source disturbance; is a disturbance penalty term, which is used to suppress the extreme disturbance of the power system; is a disturbance penalty matrix, which is used to suppress the influence of excessive disturbance on the power system.
[0174] It can be understood that, since the above-mentioned target value function includes the target coupling term, both the safety boundary of the load shedding amount and the dynamic association between the update strategy and the multi-source disturbance in the power system can be adapted, and the target value function can be deeply adapted to the above-mentioned target update model, effectively balancing the satisfaction degree of the input constraint and the system operation performance, and providing a necessary condition for the subsequent accurate update of the system state information.
[0175] In an implementation manner, the above-mentioned input cost term can be obtained by expression twenty-six as shown below:
[0176] Expression twenty-six;
[0177] wherein, is a penalty strength, which is used to control the cost weight matrix to adjust the control cost, that is, the greater the penalty strength, the higher the cost corresponding to the same load shedding amount; is a bounded strictly monotonic odd function (such as ), and the saturation constraint of the update strategy is realized through the inverse mapping thereof.
[0178] It should be noted that the embodiments of the present disclosure can include a plurality of steps, in order to facilitate description, these steps are numbered, but these numbers are not a limitation on the execution time slot, execution order between steps; these steps can be implemented in any order, the embodiments of the present disclosure do not limit this.
[0179] Further, the embodiment provides a system control device based on a power system, as shown in the figure, the device comprises a state acquisition module 601, a state update module 602 and a system control module 603. Figure 6
[0180] The state acquisition module 601 is configured to acquire target state information of the power system; wherein the target state information is system load or system frequency;
[0181] The state update module 602 is configured to update the target state information according to a target update model; wherein the target update model comprises a target update strategy, and the target update strategy is used to enable the target update model to identify and respond to multi-source disturbance, and the multi-source disturbance comprises at least two of engine mechanical disturbance, load disturbance and power grid power disturbance;
[0182] The system control module 603 is configured to control the power system to operate according to the updated target state information.
[0183] Further, in a possible implementation manner of the embodiment, as shown in the figure, the system control device further comprises a model training module 604. Figure 7
[0184] The model training module 604 is configured to add the multi-source disturbance to the power system, receive system state information of a first time sent by the power system, and take the product between the transpose matrix of the first estimation weight matrix and the activation function vector of the first network as an estimation update strategy;
[0185] Send the estimation update strategy to the power system, and receive system state information of a second time sent by the power system;
[0186] Adjust the estimation update strategy according to the system state information of the first time and the system state information of the second time until the received system state information is the same as the system state information of the first time, and obtain a target update strategy;
[0187] According to the target update strategy, a target update model is constructed.
[0188] Further, in a possible implementation manner of the embodiment, as shown in the figure, Figure 7 As shown, the model training module 604 is further configured to input the system state information at the first time and the system state information at the second time into the Hamilton function, and perform gradient solving on an update strategy included in the Hamilton function to obtain a theoretical update strategy.
[0189] A difference between the theoretical update strategy and the estimated update strategy is taken as a network error vector of the first network.
[0190] The network error vector of the first network is input into a loss function of the first network to obtain a first loss value.
[0191] The first estimated weight matrix is updated according to the first loss value until a first preset condition is met, a target first estimated weight matrix is obtained, and a product between a transposed matrix of the target first estimated weight matrix and an activation function vector of the first network is taken as a target update strategy.
[0192] Further, in a possible implementation manner of the embodiment, as shown in Figure 7 As shown, the model training module 604 is further configured to input the system state information at the first time and the system state information at the second time into the Hamilton function, and perform gradient solving on a disturbance strategy included in the Hamilton function to obtain a theoretical disturbance strategy.
[0193] A product between a transposed matrix of the second estimated weight matrix and an activation function vector of the second network is taken as an estimated disturbance strategy.
[0194] A difference between the estimated disturbance strategy and the theoretical disturbance strategy is taken as a network error vector of the second network.
[0195] The network error vector of the second network is input into a loss function of the second network to obtain a second loss value.
[0196] The second estimated weight matrix is updated according to the second loss value until a second preset condition is met, a target second estimated weight matrix is obtained, and a product between a transposed matrix of the target second estimated weight matrix and an activation function vector of the second network is taken as a target disturbance strategy.
[0197] According to the target update strategy and the target disturbance strategy, a target update model is constructed.
[0198] Further, in a possible implementation manner of the embodiment, as shown in Figure 7 As shown, the model training module 604 is further configured to input the system state information at the first time and the system state information at the second time into the Hamilton function, and perform gradient solving on an update strategy included in the Hamilton function to obtain a theoretical update strategy.
[0199] The reward term is obtained according to a product of a third estimated weight matrix and an activation function vector difference and an output value of the approximate Hamiltonian function; wherein the activation function vector difference is a difference between the activation function vector at the second moment and the activation function vector at the first moment;
[0200] The auxiliary error vector is calculated according to a transposed matrix of the third estimated weight matrix, the activation function vector difference and the reward term;
[0201] The third loss value is obtained by inputting the auxiliary error vector into a loss function of the third network;
[0202] The third estimated weight matrix is updated according to the third loss value until a third preset condition is met, a target third estimated weight matrix is obtained, and a product of a transposed matrix of the target third estimated weight matrix and an activation function vector of the third network is taken as a target value function; wherein the target value function is used to evaluate the performance of the target updated model.
[0203] Further, in a possible implementation manner of the embodiment, as shown in The model training module 604 is further configured to set an inertial center frequency of the power system and an optimization target of the target updated model; wherein the optimization target includes minimizing the frequency deviation and the load shedding cost;
[0204] According to the target dynamic model of the generator in the power system, an engine mechanical disturbance is added to the power system;
[0205] According to the target load model, a load disturbance is added to the power system;
[0206] According to the static flow constraint and the target incremental constraint, a power grid power disturbance is added to the power system.
[0207] It should be noted that the foregoing explanation and description of the method embodiments are also applicable to the device of the present embodiment, and the principles are the same, which will not be limited in the present embodiment.
[0208] The features of the embodiments corresponding to the system control device based on the power system can be referred to the related description of the embodiments corresponding to the system control method based on the power system, which will not be repeated here.
[0209] The embodiment of the present application further provides a computer device, which can be a personal computer, a server, a network device and the like. The computer device comprises a bus, a processor, a memory and a communication interface, and can further comprise an input / output interface and a display device. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal through network connection. The computer program is executed by the processor to implement the steps in the method embodiments.
[0210] Those skilled in the art can understand that the structure of the computer device described above is only part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. A specific computer device can comprise more or fewer components, or combine certain components, or have a different component arrangement.
[0211] In one embodiment, a computer readable storage medium is provided, which can be non-volatile or volatile, and stores a computer program. The computer program is executed by a processor to implement the steps in the method embodiments described above.
[0212] In one embodiment, a computer program product is provided, which comprises a computer program. The computer program is executed by a processor to implement the steps in the method embodiments described above.
[0213] It should be noted that the user personal information involved in the embodiments of the present application is authorized (known and agreed) by the relevant object or fully authorized by all parties. The execution subject can obtain it through various legal and compliant ways. The collection, storage, use, processing, transmission, provision and disclosure of the information, data and signals involved comply with the relevant legal regulations of the relevant countries and regions, and do not violate public order and good customs. It should be noted that if a software tool or component of a company other than the present company appears in the embodiments of the present application, it is only used for example introduction, and does not represent actual use.
[0214] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0215] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
[0216] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A system control method based on a power system, characterized by, The method comprises the following steps: adding a multi-source disturbance to a power system, receiving system state information at a first time sent by the power system, and taking the product between the transpose matrix of a first estimation weight matrix and the activation function vector of a first network as an estimation update strategy; wherein the multi-source disturbance comprises at least two of engine mechanical disturbance, load disturbance and power grid power disturbance; sending the estimation update strategy to the power system, and receiving system state information at a second time sent by the power system; adjusting the estimation update strategy according to the system state information at the first time and the system state information at the second time until the received system state information is the same as the system state information at the first time, to obtain a target update strategy; wherein the target update strategy is used to enable a target update model to identify and cope with the multi-source disturbance; constructing the target update model according to the target update strategy; obtaining target state information of the power system; wherein the target state information is system load or system frequency; updating the target state information according to the target update model; controlling the power system to operate according to the updated target state information; wherein the adjustment process of the estimation update strategy comprises the following steps: inputting the system state information at the first time and the system state information at the second time into a Hamilton function, and performing gradient solving on the update strategy included in the Hamilton function to obtain a theoretical update strategy; taking the difference between the theoretical update strategy and the estimation update strategy as a network error vector of the first network; inputting the network error vector of the first network into a loss function of the first network to obtain a first loss value; updating the first estimation weight matrix according to the first loss value until a first preset condition is met, to obtain a target first estimation weight matrix, and taking the product between the transpose matrix of the target first estimation weight matrix and the activation function vector of the first network as the target update strategy.
2. The method of claim 1, wherein, The target update model further comprises a target disturbance strategy, which is used to enable the target update model to actively suppress the multi-source disturbance. After receiving the system state information at the second time sent by the power system, the method further comprises the following steps: inputting the system state information at the first time and the system state information at the second time into a Hamilton function, and performing gradient solving on the disturbance strategy included in the Hamilton function to obtain a theoretical disturbance strategy; taking the product between the transpose matrix of a second estimation weight matrix and the activation function vector of a second network as an estimation disturbance strategy; taking the difference between the estimation disturbance strategy and the theoretical disturbance strategy as a network error vector of the second network; inputting the network error vector of the second network into a loss function of the second network to obtain a second loss value; Based on the second loss value, the second estimated weight matrix is updated until the second preset condition is met to obtain the target second estimated weight matrix, and the product between the transpose of the target second estimated weight matrix and the activation function vector of the second network is used as the target perturbation strategy. The step of constructing the target update model according to the target update strategy includes: The target update model is constructed based on the target update strategy and the target perturbation strategy.
3. The method of claim 2, wherein, The method further includes: The system state information at the first time step and the system state information at the second time step are input into the approximate Hamiltonian function to obtain the output value of the approximate Hamiltonian function; wherein, the approximate Hamiltonian function is a Hamiltonian function that includes the estimated perturbation strategy and the estimated update strategy; The reward term is obtained based on the product of the third estimated weight matrix and the difference between the activation function vectors and the output value of the approximate Hamiltonian function; wherein, the difference between the activation function vectors is the difference between the activation function vector at the second time step and the activation function vector at the first time step. The auxiliary error vector is calculated based on the transpose of the third estimated weight matrix, the difference in the activation function vector, and the reward term; The auxiliary error vector is input into the loss function of the third network to obtain the third loss value; The third estimated weight matrix is updated based on the third loss value until the third preset condition is met, resulting in a target third estimated weight matrix. The product between the transpose of the target third estimated weight matrix and the activation function vector of the third network is used as the target value function. The target value function is used to evaluate the performance of the target update model.
4. The method of claim 3, wherein, The objective value function includes a state penalty term, an input cost term, an objective coupling term, and a disturbance penalty term. The state penalty term is used to characterize the penalty for the state deviating from the ideal value. The input cost term is used to characterize the economic cost of the update strategy. The objective coupling term is used to construct penalty constraints. The penalty constraints are used to characterize the correlation between the update strategy and the multi-source disturbance. The disturbance penalty term is used to suppress extreme disturbances in the power system.
5. The method according to any one of claims 1-4, characterized in that, The addition of multi-source disturbances to the power system includes: The inertial center frequency of the power system and the optimization objective of the target update model are set; wherein, the optimization objective includes minimizing frequency deviation and load reduction cost; Based on the target dynamic model of the generator in the power system, the engine mechanical disturbance is added to the power system; The load disturbance is added to the power system according to the target load model; Based on static power flow constraints and target increment constraints, the power grid power disturbance is added to the power system.
6. A system control device based on a power system, characterized by comprising: include: The model training module is used to add multi-source disturbances to the power system, receive the system state information sent by the power system at the first moment, and use the product between the transpose of the first estimated weight matrix and the activation function vector of the first network as the estimation update strategy; wherein, the multi-source disturbances include at least two of the following: engine mechanical disturbances, load disturbances, and grid power disturbances; The model training module is further configured to send the estimated update strategy to the power system, and receive system state information of a second time sent by the power system; The model training module is further configured to adjust the estimated update strategy according to the system state information of the first time and the system state information of the second time until the received system state information is the same as the system state information of the first time, and obtain a target update strategy; wherein the target update strategy is used to enable a target update model to identify and cope with the multi-source disturbance; The model training module is further configured to construct the target update model according to the target update strategy; The state acquisition module is configured to acquire target state information of the power system; wherein the target state information is system load or system frequency; The state update module is configured to update the target state information according to the target update model; The system control module is configured to control the power system to operate according to the updated target state information; The adjustment process of the estimated update strategy specifically includes: inputting the system state information of the first time and the system state information of the second time into a Hamilton function, and performing gradient solving on an update strategy included in the Hamilton function to obtain a theoretical update strategy; taking a difference between the theoretical update strategy and the estimated update strategy as a network error vector of the first network; inputting the network error vector of the first network into a loss function of the first network to obtain a first loss value; updating the first estimated weight matrix according to the first loss value until a first preset condition is met, obtaining a target first estimated weight matrix, and taking a product between a transpose matrix of the target first estimated weight matrix and an activation function vector of the first network as the target update strategy.
7. A storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the method in any one of claims 1 to 5.
8. A computer device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, The processor executes the computer program to implement the method in any one of claims 1 to 5.
Citation Information
Patent Citations
Power grid frequency regulation and control method and device and storage medium
CN116667387A