A multi-source reactive power compensation control method, system, program, medium and device
By using the CNN-LSTM network in the power grid to predict the reactive power shortage interval and construct a power distribution model of a multi-source heterogeneous reactive compensation integrated system, the problem of coordination and control of multi-source heterogeneous reactive compensation equipment in the prior art is solved, and the stability of power grid operation and voltage stability are improved.
Patent Information
- Application Number
- CN202411744842.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-02
AI Technical Summary
The existing reactive power compensation technology is difficult to effectively coordinate multi-source heterogeneous reactive power compensation equipment in a complex and changeable power grid environment, resulting in control failure or instability, and the compensation ability of various types of equipment cannot be fully utilized. There are problems such as insufficient optimization, waste of resources or ineffective compensation.
By extracting the dominant influence characteristics of the reactive power of the power grid, the CNN-LSTM network is used to predict the reactive power shortage interval of the power grid, and a power distribution model of the multi-source heterogeneous reactive compensation integrated system is built, combining software and hardware collaboration solutions to achieve coordinated control and real-time optimization between devices.
It improves the stability of the power grid operation and voltage stability, enhances the flexibility and adaptability of the power grid, improves the utilization rate and operation efficiency of reactive compensation equipment, and reduces the voltage fluctuations and harmonic distortion of the power grid.
Smart Images

Figure CN119231559B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power grid systems, and in particular to a multi-source reactive power compensation control method, system, program, medium and equipment. Background Art
[0002] With the development of my country's power system, highly interconnected power grids and long-distance and large-scale power transmission have become the norm. The load center power grid determined by resource endowment is highly dependent on the power supply mode of large-scale and long-distance power transmission, forming a high-proportion load center power grid with a high degree of "power hollowing out". With the rapid growth of load levels brought about by the rapid development of the economy and society, the dynamic reactive power of the high-proportion "power hollowing out" power grid is seriously lacking, and the risk of transient voltage instability is becoming increasingly prominent, which has become the most critical constraint threatening the safe and stable operation of the power grid. With the accelerated construction of new power systems, the proportion of new energy such as wind power and photovoltaics in the power system has gradually increased, and the power structure has undergone profound changes, which will bring huge challenges to the risk analysis and prevention and control strategy research of voltage instability in the power system. Therefore, insufficient voltage support capacity has become an important factor restricting the power receiving capacity of large urban power grids and the large-scale new energy transmission capacity.
[0003] Reactive power compensation is essential to maintain the stability of grid voltage. However, with the continuous expansion of the scale of power systems and the increasing complexity of their structures, as well as the large-scale access of new energy sources and the increasing uncertainty of loads, traditional reactive power compensation equipment and control strategies face new challenges. The prediction of reactive power change trends in the power grid will provide more accurate data support for the dispatch prediction of the power system, thereby improving the economy of power system dispatch. It is also of great significance for optimizing the operation of the power system, improving the quality of power, and ensuring the stable operation of the power system. However, existing prediction models usually cannot fully consider the complexity and dynamic changes of the power grid operation status, resulting in a large room for improvement in prediction accuracy.
[0004] Existing reactive power compensation technologies mainly rely on static compensation equipment (such as static VAR compensators (SVCs), conventional capacitive reactance, etc.) and dynamic compensation equipment (such as synchronous phase condensers). Although these devices can provide basic reactive power support, their response speed and adjustment capabilities are limited when facing frequent fluctuations in the power grid, which can easily lead to problems such as system voltage fluctuations and harmonic distortion. In addition, a single type of reactive power compensation equipment cannot fully cope with reactive power requirements under different working conditions, especially in complex and changeable operating environments. The mutual coupling and control interaction problems between reactive power compensation devices are more prominent, which may cause control failure or instability. The use of grid-type static VAR generators SVG together with other reactive power compensation devices such as SVCs, conventional SVGs, synchronous phase condensers, conventional capacitive reactance, etc. will better meet the comprehensive needs in practical applications.
[0005] With the increasing demand for flexible transmission technology and new energy grid connection, current research has gradually shifted to the dynamic coordinated control of multi-source heterogeneous reactive compensation devices. Using the high-performance computing capabilities of computer equipment and the instruction storage function of readable storage media, domestic scholars have developed a variety of reactive compensation control algorithms, especially in improving voltage support capabilities under weak grid conditions. Significant progress has been made. However, most of the existing research is based on a single device or a specific scenario, and research on the collaboration and real-time optimization between multi-source heterogeneous devices is still insufficient.
[0006] In response to these problems, current research has gradually introduced artificial intelligence technology, especially neural network technology, to predict the reactive power shortage of the power grid through deep learning of the factors affecting reactive power. In addition, for multi-source heterogeneous reactive compensation equipment, the existing coordinated control strategies are mostly based on linear control models, which makes it difficult to fully utilize the compensation capabilities of various types of equipment, and there are problems such as insufficient optimization, waste of resources or ineffective compensation. Therefore, how to combine the characteristics of multi-source heterogeneous equipment to design a more efficient reactive compensation power allocation strategy has become an important research direction for improving the stability of power grid operation. How to achieve real-time optimization and collaboration of multi-source heterogeneous reactive compensation devices by combining computer programs, readable storage media, and computer devices to execute complex control algorithms can not only meet the actual needs of modern power grid operation, but also promote the large-scale application of reactive compensation devices in engineering practice, providing important support for the intelligent development of power systems. Summary of the invention
[0007] Based on the multi-source heterogeneous reactive compensation equipment in the prior art, it is difficult to fully utilize the compensation capabilities of various types of equipment, and there are problems such as insufficient optimization, waste of resources or ineffective compensation. The purpose of the present invention is to provide a multi-source reactive compensation control method, system, program, medium and equipment. First, based on the influencing factors of the reactive power change of the power grid, the prediction of the time-series reactive power shortage interval of the power grid is realized, and a reference is provided for the total output of the multi-source heterogeneous reactive compensation integrated system; then, considering the power grid voltage drop, network loss, voltage stability and the stability of the multi-source reactive compensation integrated system as the objective function, a reactive distribution model for various reactive devices is constructed.
[0008] In addition, this study focuses on the coordinated control of multi-source heterogeneous reactive compensation devices, combining the functional characteristics of modern computer equipment, memory and computer-readable storage media, and proposes a solution based on the combination of software and hardware. By using memory to store control algorithms and operating data, computer-readable storage media to store program instructions, and computer equipment to run the above instructions, complex reactive compensation control logic can be efficiently implemented. This study uses a programmed approach to support coordinated control between devices, improves the reliability of instruction storage and transmission through storage media, and combines the high-performance computing capabilities provided by computing devices to promote the realization of efficient operation of multi-source heterogeneous devices in power systems. This study demonstrates the important value of software and hardware collaboration in improving the intelligence level of complex power control tasks, and provides a technical path for further realizing large-scale intelligent reactive compensation.
[0009] The present invention is achieved through the following technical solutions:
[0010] In a first aspect, the present application provides a multi-source reactive power compensation control method, comprising the following steps:
[0011] The dominant influencing features of the reactive power of the power grid are extracted as the input of the CNN-LSTM network, the reactive power shortage interval of the power grid is predicted, and it is used as the power allocation target of the multi-source reactive power compensation equipment;
[0012] Taking the minimization of system network loss and voltage drop and the stability of multi-source heterogeneous reactive compensation integrated system as the optimization goals, a power distribution model of multi-source heterogeneous reactive compensation integrated system is constructed.
[0013] Furthermore, the historical load data, historical generator excitation system data, and node voltage data are used as the input of CNN, and the state space of the input data is:
[0014]
[0015] Where L(t) is the historical load data, F L is the characteristic number of load data; G(t) is the data of historical generator excitation system, F G is the characteristic number of the generator excitation system data; V(t) is the node voltage data, F V is the characteristic number of node voltage data; Q(t) is the historical reactive power data, F Q is the characteristic number of historical reactive power data; N is the number of samples, T is the time step, and R represents a real number set.
[0016] Furthermore, CNN is used to capture the multi-dimensional local features of the input data through convolution operations. By convolving and pooling the input data, the multi-dimensional state features of different nodes or devices in the power grid are extracted to reflect the reactive power changes of the power grid. The convolution and pooling process is formally expressed as:
[0017]
[0018] In the formula, Z 1 is the input data after convolution processing; P 1 is the output data of the CNN network, that is, the multi-dimensional state characteristics of different nodes or devices in the power grid; CONV1D is the one-dimensional convolution layer in the convolutional neural network, indicating the convolution operation; ReLU is the activation function of the one-dimensional convolution layer; MaxPooling1D is the pooling operation; is the normalized input data state space; W 1 is the convolution kernel with a shape of (K, F, C), where K is the convolution kernel size, F is the number of input features, and C is the number of output channels; b 1 is the bias; p is the pooling window size. Further, the output data of the CNN network is processed using the Dropout regularization technique to prevent the model from overfitting the training data, and then the flattening operation is used to process the above data to achieve the dimension alignment of the feature data of the CNN-LSTM network. The data processing process is:
[0019]
[0020] Where D 1 is the data processed by regularization technology; F LSTM is the data processed by the flattening operation; d 1 is the Dropout rate; Dropout is the regularization operation; Flatten is the flattening operation.
[0021] Furthermore, the LSTM layer performs time series modeling on the flattened feature tensor and generates the final reactive power shortage interval prediction result through the fully connected layer:
[0022] H t =LSTM(H t-1 , F LSTM ; W h , W x , b)
[0023] Y=W out ·Y dense +b out ;
[0024] In the formula, H t and H t-1 are the hidden states of the t-th and t-1-th time steps respectively; Y represents the output data; that is, the prediction result; W h and W x are the weight matrix of the hidden state and the weight matrix input to the hidden state; b is the bias vector; Ydense is the output of the fully connected layer, W out and b out are the weight matrix and bias vector of the output layer respectively.
[0025] Furthermore, the reactive power loss function is:
[0026]
[0027] In the formula, Q loss is the reactive power loss in the system, M is the total number of lines in the system, Q i is the reactive power of the ith node or line, R i is the resistance of the ith circuit, V i The voltage of the ith line.
[0028] Furthermore, the power quality of the power grid is characterized by voltage drop, and the calculation formula is:
[0029]
[0030] Where ΔV i is the voltage drop at the ith node; P i is the active power of the ith node; X i is the reactance of the ith node.
[0031] Furthermore, the calculation formula of the power grid short-circuit ratio is:
[0032]
[0033] In the formula, S sc is the short-circuit capacity of a node; S r is the rated capacity of the access device.
[0034] Furthermore, the objective function of the power allocation model is:
[0035]
[0036] In the formula, is the stability evaluation index; 0 is the expected stability margin of the system; γ is the actual stability margin of the system.
[0037] Furthermore, the calculation formula for the total reactive power output of the multi-source heterogeneous reactive power compensation integrated system is:
[0038] Q total =Q SVG +Q SVC +Q C +Q IND +Q STC ≥Q 缺额 ;
[0039]
[0040] In the formula, Q total is the total reactive power output of the multi-source heterogeneous reactive power compensation integrated system; Q SVG , Q SVC , Q C , Q IND , Q STC They are the reactive output of SVG, SVC, capacitor, reactance and synchronous condenser; Q 缺额 is the reactive power deficit of the grid; is the rated reactive capacity of SVG; is the rated reactive capacity of the SVC; is the rated reactive capacity of the capacitor; is the rated reactive capacity of the reactance; It is the rated reactive capacity of the synchronous condenser.
[0041] Furthermore, the voltage U of the grid-connected node of the multi-source heterogeneous reactive power compensation system i Need to meet: U min ≤U i ≤U max , U min and U max These are the minimum and maximum allowable voltage values, respectively.
[0042] Furthermore, a power allocation model of a multi-source heterogeneous reactive power compensation integrated system is constructed based on the TD3 deep reinforcement learning algorithm;
[0043] TD3's deep reinforcement learning algorithm process includes the input state space, the update process of the Critic network, the update process of the Actor network, and the update process of the target network.
[0044] Furthermore, the input state space is formulated as:
[0045] Ω t =[S t ,A t ,R t , S t+1 ];
[0046] S t = {Y t ,V t , Q i,t};
[0047] In the formula, Ω t is the input state space; S tA is the input state quantity at time t, including the reactive power shortage of the power grid, the voltage data of the power grid nodes and the reactive power distribution state of each reactive compensation device; t For the Actor network based on the state S t The resulting action, R t To perform action A t The resulting action reward; S t+1 is the input state at time t+1; Y t V is the reactive power shortage interval of the power grid; t is the corresponding node voltage data; Q i,t is the reactive power distribution state of each reactive power compensation device at time t.
[0048] Furthermore, the update process of the Critic network is as follows:
[0049] First, use the target Actor network to calculate the action in state s':
[0050] a′=μ′(s′+θ μ′ );
[0051] In the formula, a′ represents the action generated according to the state at the next moment; μ′ represents the target Actor network; θ μ′ Represents the parameters of the target Actor network;
[0052] Secondly, based on the idea of dual networks, calculate the target value:
[0053]
[0054] In the formula, y represents the target value; Q′ i Represents the target Critic network; represents the parameters of the target Critic network; γ represents the discount factor, which is used to represent the difference in benefits caused by the order of states;
[0055] Finally, the gradient descent algorithm is used to minimize the error between the evaluation value and the target value, thereby updating the parameters in the Critic1 and Critic2 networks:
[0056]
[0057] In the formula, s is the input state; a is the execution action; represents the gradient of the Critic network update; Q i represents the Critic network; Represents the parameters of the Critic1 network.
[0058] Furthermore, the update process of the Actor network is as follows:
[0059] First, use the Actor network to calculate the action in state s:
[0060] a=μ(s|θ μ );
[0061] Where μ represents the Actor network; θ μ Represents the parameters of the Actor network;
[0062] Secondly, use the Critic1 or Critic2 network to calculate the evaluation value of the state-action pair (s, a). Here, it is assumed that the Critic1 network is used:
[0063]
[0064] In the formula, q represents the execution target reward calculated by the Ctitic network for action a; Q 1 represents the Ctitic network; Represents the parameters of the Ctitic network;
[0065] Finally, the gradient ascent algorithm is used to maximize q to complete the update of the Actor network.
[0066] Furthermore, the update process of the target network is as follows:
[0067] Introduce a learning rate τ or momentum, take the weighted average of the old target network parameters and the new corresponding network parameters, and then assign them to the target network:
[0068]
[0069] In the formula, θ μ′ Represents the parameters of the target Actor network; Represents the parameters of the target Ctitic network.
[0070] In a second aspect, the present application provides a coordinated control system for a multi-source heterogeneous reactive power compensation device, including a first module and a second module.
[0071] The first module is used to extract the dominant influencing features of the reactive power of the power grid as the input of the CNN-LSTM network, predict the reactive power shortage interval of the power grid, and use it as the power allocation target of the multi-source reactive power compensation equipment;
[0072] The second module is used to construct a power distribution model of a multi-source heterogeneous reactive compensation integrated system with the minimization of system network loss and voltage drop and the stability of the multi-source heterogeneous reactive compensation integrated system as optimization goals.
[0073] In a third aspect, the present application provides a computer program, which can implement the steps of the method described above.
[0074] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the method described above.
[0075] In a fifth aspect, the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method described above.
[0076] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0077] (1) In this application, by studying the influencing factors of the reactive power of the power grid, the influencing characteristics are used as the neural network input to predict the reactive power shortage interval of the power grid, which can reflect the spatiotemporal distribution characteristics of the reactive power of the power grid, provide the power grid operator with the changing trend of reactive demand, and take corresponding reactive compensation measures in time. In addition, the reactive power shortage interval of the power grid can provide data support for the operation of reactive compensation equipment (such as static synchronous compensator SVG, static VAR compensator SVC, synchronous phase condenser, etc.), and by predicting the future reactive demand changes in advance, a multi-source heterogeneous reactive compensation integrated system is constructed, the utilization rate and operation efficiency of the equipment are improved, the flexibility and adaptability of the power grid are enhanced, and the voltage stability and safety of the power grid are improved.
[0078] (2) In this application, considering the reactive power loss, power quality and voltage stability of the power grid, a variety of reactive compensation equipment is used to compensate for the reactive power shortage of the power grid. By optimizing the reactive output of each reactive compensation device, the voltage level of the power grid node is effectively adjusted to prevent voltage fluctuations, overvoltage or undervoltage caused by load changes or power grid disturbances, thereby enhancing the voltage stability and operational safety of the power grid. In addition, since the operating characteristics and control strategies of different reactive devices are different, and the control failure or instability may occur due to coupling interaction during operation, this patent considers the operational stability of the multi-source heterogeneous reactive compensation integrated system when constructing the objective function. By reasonably arranging the reactive output of various reactive devices, each type of equipment can play its role to the maximum extent according to its characteristics and operating status, avoiding excessive loading and frequent start and stop of the equipment, and improving the overall utilization efficiency of the equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative work. In the drawings:
[0080] Figure 1 It is a flow chart of a coordinated control method of a multi-source heterogeneous reactive power compensation device of the present invention;
[0081] Figure 2 It is a schematic diagram of reactive power prediction of power grid based on CNN-LSTM hybrid model in the present invention;
[0082] Figure 3 This is the structural flow of the TD3 algorithm in the present invention, where Epoch represents the number of iterations of the agent, and E represents the number of iterations that meet the conditions;
[0083] Figure 4 It is a schematic diagram of the overall process of the present invention. DETAILED DESCRIPTION
[0084] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.
[0085] The following detailed description specifically discloses the implementation methods of the present application with appropriate reference to the accompanying drawings. However, there may be cases where unnecessary detailed descriptions are omitted. For example, there are cases where detailed descriptions of well-known matters and repeated descriptions of actually the same structures are omitted. This is to avoid the following description from becoming unnecessarily lengthy and to facilitate the understanding of those skilled in the art. In addition, the drawings and the following description are provided for those skilled in the art to fully understand the present application and are not intended to limit the subject matter described in the claims. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0086] Unless otherwise specified, all embodiments and optional embodiments of the present application can be combined with each other to form a new technical solution.
[0087] Unless otherwise specified, all technical features and optional technical features of this application can be combined with each other to form a new technical solution.
[0088] The terms "comprises," "comprising," or any other variation thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0089] Example 1
[0090] like Figures 1 to 4 As shown, this embodiment provides a coordinated control method for a multi-source heterogeneous reactive power compensation device, which is performed in the following two steps:
[0091] S1. Extract the dominant influencing features of the reactive power of the power grid as the input of the convolution-long short-term memory neural network (CNN-LSTM network), predict the reactive power shortage interval of the power grid, and use it as the power allocation target of the multi-source reactive power compensation equipment;
[0092] S2. Taking the minimization of system network loss and voltage drop and the stability of multi-source heterogeneous reactive compensation integrated system as optimization goals, a power distribution model of multi-source heterogeneous reactive compensation integrated system is constructed.
[0093] The above two steps are described in detail below.
[0094] 1. Research on the prediction method of reactive power shortage interval of power grid based on CNN-LSTM network
[0095] The reactive power of the power grid is usually related to a variety of influencing factors, among which the load type, node voltage, and generator excitation system are the dominant influencing factors. Inductive loads (such as motors and transformers) consume reactive power, while capacitive loads (such as capacitors) provide reactive power; the node voltage level of the power grid directly affects the demand and distribution of reactive power; and the excitation control system of the generator determines the reactive power provided by the generator to the power grid. Therefore, the above influencing factors are used as the input of CNN, and the state space of the input data is:
[0096]
[0097] Where L(t) is the historical load data, F L is the characteristic number of load data; G(t) is the data of historical generator excitation system, F G is the characteristic number of the generator excitation system data; V(t) is the node voltage data, F V is the characteristic number of node voltage data; Q(t) is the historical reactive power data, F Qis the characteristic number of historical reactive power data; N is the number of samples, T is the time step (sequence length), and R represents a real number set.
[0098] Since the above historical data all have spatiotemporal characteristics, CNN is used to capture the multi-dimensional local features of the input data through convolution operations. By convolving and pooling the input data, the multi-dimensional state features of different nodes or devices in the power grid are extracted to reflect the reactive power changes of the power grid. The convolution and pooling process is formally expressed as:
[0099]
[0100] In the formula, Z 1 is the input data after convolution processing; P 1 is the output data of the CNN network, that is, the multi-dimensional state characteristics of different nodes or devices in the power grid; CONV1D is the one-dimensional convolution layer in the convolutional neural network, indicating the convolution operation; ReLU is the activation function of the one-dimensional convolution layer; MaxPooling1D is the pooling operation; is the normalized input data state space; W 1 is the convolution kernel with a shape of (K, F, C), where K is the convolution kernel size, F is the number of input features, and C is the number of output channels; b 1 is the bias; p is the pooling window size. CNN extracts features and reduces the dimension of data through convolution and pooling operations, which can effectively reduce data redundancy, improve computational efficiency, and reduce the risk of overfitting.
[0101] The Dropout regularization technique is used in both CNN and LSTM to prevent the model from overfitting the training data and improve the model's prediction ability on new data. The feature data dimensions of the two networks are aligned through flattening operations. The data processing process is as follows:
[0102]
[0103] Where D 1 is the data processed by regularization technology; F LSTM is the data processed by the flattening operation; d 1 is the Dropout rate; Dropout is the regularization operation; Flatten is the flattening operation.
[0104] CNN can extract local features of each node and device, and LSTM models the changing trend of these features in time series. The LSTM layer performs time series modeling on the flattened feature tensor, and generates the final reactive power shortage interval prediction result through the fully connected layer:
[0105] H t =LSTM(Ht-1 , F LSTM ; W h ,W x , b)
[0106] Y=W out ·Y dense +b out ;
[0107] In the formula, H t and H t-1 are the hidden states of the t-th and t-1-th time steps respectively; Y represents the output data; that is, the prediction result; W h and W x are the weight matrix of the hidden state and the weight matrix input to the hidden state; b is the bias vector; Y dense is the output of the fully connected layer, W out and b out are the weight matrix and bias vector of the output layer respectively.
[0108] The CNN-LSTM hybrid network is used to predict the reactive power shortage interval of the power grid because it can combine the spatial feature extraction capability of CNN and the time series modeling capability of LSTM, establish a mapping relationship between the reactive power change of the power grid and the historical input data, adapt to the complex characteristics and dynamic changes of the power grid system, and then realize the prediction of the reactive power shortage interval of the power grid. The CNN-LSTM hybrid model can provide more accurate and stable prediction results, help the power system better cope with the changes in reactive power demand, provide a reference for the power allocation of reactive compensation equipment, and thus reduce the voltage fluctuation and harmonic distortion of the power grid, and maintain the safe and stable operation of the power grid. Figure 2 Schematic diagram of grid reactive power prediction based on CNN-LSTM hybrid model.
[0109] 2. Construction of power distribution model for multi-source heterogeneous reactive power compensation integrated system
[0110] The present invention considers the coordinated operation of multiple reactive power compensation devices to perform reactive power compensation for the power grid shortage interval, mainly including static synchronous compensator (Static Var Generator, SVG), static VAR compensator (Static Var Compensator, SVC), capacitor, reactance and synchronous phase condenser. Since the operating characteristics and control strategies of different devices are different, and the control failure or instability may occur due to coupling interaction during operation, this patent considers the stability of the multi-source heterogeneous reactive power compensation integrated system on the basis of considering the reactive power loss, power quality and voltage stability of the power grid, and constructs the objective function of the power distribution model. By compensating for the reactive power shortage interval of the power grid predicted in the previous article, the reactive power compensation capacity of various devices can be maximized.
[0111] (1) Objective function
[0112] The reactive power loss of the power grid can reflect the load level, network structure, equipment status and other operating conditions of the power grid. Higher network loss means more energy waste and increases the operating cost of the power system. Therefore, reducing network loss is very important to improve the economic efficiency of the power grid. Reactive power loss mainly occurs in transmission lines and transformers, and its calculation formula is related to the resistance and reactance of the line and the transmitted reactive power.
[0113]
[0114] In the formula, Q loss is the reactive power loss in the system, M is the total number of lines in the system, Q i is the reactive power of the ith node or line, R i is the resistance of the ith circuit, V i The voltage of the ith line.
[0115] The power quality of the power grid is usually characterized by voltage drop. The calculation formula is:
[0116]
[0117] Where ΔV i is the voltage drop at the ith node; P i is the active power of the ith node; X i is the reactance of the ith node.
[0118] The short-circuit ratio is an important indicator used to measure the stability of the power grid voltage. When the short-circuit ratio is high, it means that the short-circuit capacity of the power grid at this node is relatively large, and the impact of the access equipment on the power grid voltage under disturbance is small. The formula for calculating the short-circuit ratio of the power grid is as follows:
[0119]
[0120] Where SCR is the short circuit ratio; S sc is the short-circuit capacity of a certain node, usually expressed as the short-circuit apparent power of the grid at a specific node;
[0121] S r It is the rated capacity of the access equipment, usually expressed as the rated apparent power of the generator or renewable energy power source.
[0122] The stability of the multi-source heterogeneous reactive compensation integrated system can be characterized by the stability margin. The grid impedance is obtained by grid flow calculation, and the impedance of each reactive compensation device is obtained based on the sweep frequency method. Then the impedance of the multi-source heterogeneous reactive compensation integrated system is calculated. The intersection frequency of the amplitude-frequency characteristic curve of the grid impedance and the reactive compensation system impedance is the stability margin of the reactive compensation system. The difference between the actual stability margin and the expected stability margin is used as the stability evaluation index of the reactive compensation system:
[0123]
[0124] In the formula, is the stability evaluation index; 0 is the expected stability margin of the system; γ is the actual stability margin of the system.
[0125] In summary, the objective function of the power allocation model of the multi-source heterogeneous reactive power compensation integrated system can be expressed as:
[0126]
[0127] (2) Constraints
[0128] The multi-source heterogeneous reactive power compensation integrated system compensates for the reactive power shortage interval of the power grid and needs to meet the actual reactive power constraints. In order to ensure the normal operation of various reactive compensation equipment, the capacity constraints of each reactive compensation equipment and the voltage level constraints of the grid-connected nodes should also be considered. The specific formula is as follows:
[0129] 1) The total reactive power output of each reactive power compensation device must meet the reactive power deficit of the power grid to maintain the target voltage level and power factor.
[0130] Q total =Q SVG +Q SVC +Q C +Q IND +Q STC ≥Q 缺额 ;
[0131] In the formula, Q total It is the total reactive power output of the multi-source heterogeneous reactive power compensation integrated system; Q SVG Q SVC Q C , Q IND Q STC They are the reactive output of SVG, SVC, capacitor, reactance and synchronous condenser; Q 缺额 It is the reactive power shortage of the power grid.
[0132] 2) The reactive output of each device cannot exceed its rated capacity to avoid overload.
[0133]
[0134] In the formula, is the rated reactive capacity of SVG; is the rated reactive capacity of the SVC; is the rated reactive capacity of the capacitor; is the rated reactive capacity of the reactance; It is the rated reactive capacity of the synchronous condenser.
[0135] 3) Voltage fluctuations may occur at the nodes where the multi-source heterogeneous reactive power compensation system is connected to the grid. Ensure that the voltage level of the grid-connected nodes is within the specified range to avoid overvoltage or undervoltage.
[0136] V min ≤V i ≤V max ;
[0137] Where U i is the voltage of the grid-connected node of the multi-source heterogeneous reactive power compensation system, U min and U max These are the minimum and maximum allowable voltage values, respectively.
[0138] (3) Power allocation model of multi-source heterogeneous reactive power compensation integrated system based on TD3 algorithm
[0139] The power allocation model of the multi-source heterogeneous reactive compensation integrated system is constructed based on the TD3 deep reinforcement learning algorithm. The reactive power shortage of the system has typical spatiotemporal distribution characteristics, and the spatial characteristics are characterized by the node voltage changes. Therefore, the reactive power shortage of the power grid in each period predicted in Section 1 and the corresponding system node voltage are used as the state input of the agent. In addition, the TD3 algorithm needs to judge the action at the next moment according to the current state of the input variable, and obtain the reward for executing the action to guide the update of the action. Therefore, the initial power state of each reactive compensation device is also required as the input state quantity. The action output of the agent is the reactive output of each reactive device, and finally a power allocation scheme for the multi-source heterogeneous reactive compensation integrated system is formed on a short-term time scale. Figure 3 This is the TD3 algorithm structure flow. In the figure, Epoch represents the number of agent iterations, and E represents the number of iterations that meet the conditions. The algorithm flow is formally expressed as follows:
[0140] 1) Input state space
[0141] Ω t =[S t ,A t , R t , S t+1 ];
[0142] In the formula, Ω tis the input state space; S t is the input state quantity at time t, including the reactive power shortage of the power grid, the voltage data of the power grid nodes and the reactive power distribution state of each reactive compensation device; A t For the Actor network based on the state S t The resulting action, R t To perform action A t The resulting action reward; S t+1 is the input state quantity at time t+1. t It can be expressed as: S t = {Y t ,V t , Q i,t );
[0143] Where Y t V is the reactive power shortage interval of the power grid; t is the corresponding node voltage data; Q i,t is the reactive power distribution state of each reactive power compensation device at time t. At the beginning of training, the initial state space is obtained and stored in the Replay Buffer, from which data is sampled for subsequent training.
[0144] 2) Update process of the evaluation (Critic) network
[0145] The TD3 algorithm introduces a dual critic network to solve the problem of Q value overestimation. Both the Critic1 and Critic2 networks are updated by minimizing the error between the evaluation value and the target value, and all target networks are updated in a soft way. During the training phase, a batch of data is sampled from the Replay Buffer. Assume that a piece of data sampled is (s, a, r, s′, done), where s is the current state, a is the current action, r is the action reward, s′ is the next action, and done is a flag indicating whether the training is finished. The update process of the Critic network is as follows:
[0146] Use the target Actor network to calculate the action in state s':
[0147] First, use the target Actor network to calculate the action in state s':
[0148] a′=μ′(s′+θ μ′ );
[0149] In the formula, a′ represents the action generated according to the state at the next moment; μ′ represents the target Actor network; θ μ′ Represents the parameters of the target Actor network;
[0150] Secondly, based on the idea of dual networks, calculate the target value:
[0151]
[0152] In the formula, y represents the target value; Q′ i Represents the target Critic network; represents the parameters of the target Critic network; r represents the discount factor, which is used to represent the difference in benefits caused by the order of states;
[0153] Finally, the gradient descent algorithm is used to minimize the error between the evaluation value and the target value, thereby updating the parameters in the Critic1 and Critic2 networks:
[0154]
[0155] In the formula, s is the input state; a is the execution action; represents the gradient of the Critic network update; Q i represents the Critic network; Represents the parameters of the Critic1 network.
[0156] 3) Actor network update process
[0157] The Actor network is updated by maximizing the cumulative expected return (deterministic policy gradient), and a delayed update strategy is used to ensure the training stability of the algorithm (after the Ctitic1 and Critic2 networks are updated d steps, the Actor network update is started).
[0158] First, use the Actor network to calculate the action in state S:
[0159] a=μ(s|θ μ );
[0160] Where μ represents the Actor network; θ μ Represents the parameters of an Actor network.
[0161] Secondly, use the Critic1 or Critic2 network to calculate the evaluation value of the state-action pair (s, a). Here, it is assumed that the Critic1 network is used:
[0162]
[0163] In the formula, q represents the execution target reward calculated by the Ctitic network for action a; Q 1 represents the Ctitic network; Represents the parameters of the Ctitic network;
[0164] Finally, the gradient ascent algorithm is used to maximize q to complete the update of the Actor network.
[0165] 4) Target network update process
[0166] The target network is updated using a soft update method. A learning rate (or momentum) τ is introduced to perform a weighted average of the old target network parameters and the new corresponding network parameters, and then assigned to the target network:
[0167]
[0168] In the formula, θ μ′ Represents the parameters of the target Actor network; Represents the parameters of the target Ctitic network.
[0169] The power distribution model of the multi-source heterogeneous reactive compensation integrated system based on the TD3 algorithm can model the reactive power distribution problem of multiple reactive compensation devices as a sequential decision problem, that is, the reactive power distribution planning and solution process is divided into a set of sequential state transition processes. Through state transition, the reactive output of each reactive compensation device is continuously optimized, and finally a reactive power distribution plan that meets the target conditions is formed.
[0170] Example 2
[0171] This embodiment provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the coordinated control method of a multi-source heterogeneous reactive compensation device in the above-mentioned embodiment 1.
[0172] The computer device may be a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may interact with the user through a keyboard, a mouse, a remote control, a touch pad, or a voice control device.
[0173] The storage includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or D interface display memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory may be an internal storage unit of the computer device, for example, a hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), etc. equipped on the computer device. Of course, the memory may also include both an internal storage unit of the computer device and an external storage device. In this embodiment, the memory is often used to store an operating system and various application software installed on the computer device, such as a program code for running the coordinated control method of the multi-source heterogeneous reactive compensation device, etc. In addition, the memory may also be used to temporarily store various types of data that have been output or are to be output.
[0174] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data, such as running the program code of the coordinated control method of the multi-source heterogeneous reactive compensation device.
[0175] Example 3
[0176] This embodiment provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the coordinated control method of a multi-source heterogeneous reactive compensation device in the above-mentioned embodiment 1.
[0177] Among them, the computer-readable storage medium stores an interface display program, and the interface display program can be executed by at least one processor so that the at least one processor executes the steps of the coordinated control method of a multi-source heterogeneous reactive compensation device in the above-mentioned embodiment 1.
[0178] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0179] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0180] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0181] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0182] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server or network device, etc.) to execute the methods described in each embodiment of the present application.
[0183] Throughout the specification, references to "one embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment of the present invention. Therefore, the phrases "one embodiment," "an embodiment," "an example," or "an example" appearing in various places throughout the specification do not necessarily all refer to the same embodiment or example. In addition, particular features, structures, or characteristics may be combined in one or more embodiments or examples in any suitable combination and / or sub-combination. In addition, it will be appreciated by those of ordinary skill in the art that the figures provided herein are for illustrative purposes and that the figures are not necessarily drawn to scale. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0184] The specific embodiments described above further describe the purpose, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only the specific embodiments of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention. For those skilled in the art, it is obvious that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or basic features of the present application. Therefore, no matter from which point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present application is limited by the attached claims rather than the above description, and it is intended to include all changes within the meaning and scope of the equivalent elements of the claims in the present application.
Claims
1. A multi-source reactive power compensation control method, characterized in that: The following steps are involved: The dominant influencing features of the reactive power of the power grid are extracted as the input of the CNN-LSTM network, the reactive power shortage interval of the power grid is predicted, and it is used as the power allocation target of the multi-source reactive power compensation equipment; Taking the minimization of system network loss and voltage drop and the stability of multi-source heterogeneous reactive compensation integrated system as the optimization objectives, a power allocation model of multi-source heterogeneous reactive compensation integrated system is constructed; The reactive power loss function is: In the formula, Q loss is the reactive power loss in the system, M is the total number of lines in the system, Q i is the reactive power of the ith node or line, R i is the resistance of the ith circuit, V i The voltage of the ith line; The power quality of the power grid is characterized by voltage drop, and the calculation formula is: In the formula, ΔV i is the voltage drop at the ith node; P i is the active power of the ith node; X i is the reactance of the ith node; The calculation formula of the power grid short circuit ratio is: Where SCR is the short circuit ratio; S sc is the short-circuit capacity of a node; S r is the rated capacity of the access equipment; The objective function of the power allocation model is: In the formula, is the stability evaluation index; γ0 is the expected stability margin of the system; γ is the actual stability margin of the system; Construct a power distribution model for a multi-source heterogeneous reactive power compensation integrated system based on TD3's deep reinforcement learning algorithm; The deep reinforcement learning algorithm process of TD3 includes the input state space, the update process of the Critic network, the update process of the Actor network, and the update process of the target network; The input state space is formulated as: Oh t =[S t ,A t ,R t ,S t+1 ]; S t ={Y t ,V t ,Q i,t }; In the formula, Ω t is the input state space; S t A is the input state quantity at time t, including the reactive power shortage of the power grid, the voltage data of the power grid nodes and the reactive power distribution state of each reactive compensation device; t The action network is based on the state S t The action generated, R t To perform action A t The resulting action reward; S t+1 is the input state at time t+1; Y t V is the reactive power shortage interval of the power grid; t is the corresponding node voltage data; Q i,t is the reactive power distribution state of each reactive power compensation device at time t; The update process of the Critic network is as follows: First, use the target Actor network to calculate the action in state s': a′=μ′(s′+θ μ′ ); In the formula, a′ represents the action generated according to the state at the next moment; μ′ represents the target Actor network; θ μ′ Represents the parameters of the target Actor network; Secondly, based on the idea of dual networks, calculate the target value: In the formula, y represents the target value; Q′ i Represents the target Critic network; represents the parameters of the target Critic network; γ represents the discount factor, which is used to represent the difference in benefits caused by the order of states; Finally, the gradient descent algorithm is used to minimize the error between the evaluation value and the target value, thereby updating the parameters in the Critic1 and Critic2 networks: In the formula, s is the input state; a is the execution action; represents the gradient of the Critic network update; Q i represents the Critic network; Represents the parameters of the Critic1 network; The update process of the Actor network is as follows: First, use the Actor network to calculate the action in state s: a=μ(s|θ μ ); Where μ represents the Actor network; θ μ Represents the parameters of the Actor network; Secondly, use the Critic1 or Critic2 network to calculate the evaluation value of the state-action pair (s, a). Here, it is assumed that the Critic1 network is used: Where q represents the execution target reward calculated by the Critic network for action a; Q1 represents the Critic network; Represents the parameters of the Critic network; Finally, the gradient ascent algorithm is used to maximize q, thereby completing the update of the Actor network; The update process of the target network is as follows: Introduce a learning rate τ or momentum, take the weighted average of the old target network parameters and the new corresponding network parameters, and then assign them to the target network: In the formula, θ μ′ Represents the parameters of the target Actor network; Represents the parameters of the target Critic network.
2. A multi-source reactive power compensation control method according to claim 1, characterized in that: The historical load data, historical generator excitation system data, and node voltage data are used as the input of CNN, and the state space of the input data is: Where L(t) is the historical load data, F L is the characteristic number of load data; G(t) is the data of historical generator excitation system, F G is the characteristic number of the generator excitation system data; V(t) is the node voltage data, F V is the characteristic number of node voltage data; Q(t) is the historical reactive power data, F Q is the characteristic number of historical reactive power data; N is the number of samples, T is the time step, and R represents a real number set.
3. A multi-source reactive power compensation control method according to claim 1, characterized in that: CNN is used to capture the multi-dimensional local features of the input data through convolution operations. By convolving and pooling the input data, the multi-dimensional state features of different nodes or devices in the power grid are extracted to reflect the reactive power changes of the power grid. The convolution and pooling process is formally expressed as: In the formula, Z1 is the input data after convolution processing; P1 is the output data of the CNN network, that is, the multi-dimensional state characteristics of different nodes or devices in the power grid; CONV1D is the one-dimensional convolution layer in the convolutional neural network, indicating the convolution operation; ReLU is the activation function of the one-dimensional convolution layer; MaxPooling1D is the pooling operation; is the normalized input data state space; W1 is the convolution kernel, and its shape is (K, F, C), where K is the convolution kernel size, F is the number of input features, and C is the number of output channels; b1 is the bias; p is the pooling window size; N is the number of samples, T is the time step, and R represents a real number set.
4. A multi-source reactive power compensation control method according to claim 1, characterized in that: The output data of the CNN network is processed using the Dropout regularization technique to prevent the model from overfitting the training data. Then the above data is processed using the flattening operation to achieve the dimension alignment of the feature data of the CNN-LSTM network. The data processing process is as follows: In the formula, D1 is the data processed by regularization technology; F LSTM is the data processed by the flattening operation; d1 is the Dropout rate; Dropout is the regularization operation; Flatten is the flattening operation; P1 is the output data of the CNN network, which is the multi-dimensional state characteristics of different nodes or devices in the power grid; N is the number of samples, T is the time step, R represents the set of real numbers; K is the convolution kernel size, p is the pooling window size, and C is the number of output channels.
5. A multi-source reactive power compensation control method according to claim 4, characterized in that: The LSTM layer performs time series modeling on the flattened feature tensor and generates the final reactive power shortage interval prediction result through the fully connected layer: H t =LSTM(H t-1 ,F LSTM ;W h ,W x ,b) Y=W out ·Y dense +b out ; In the formula, H t and H t-1 are the hidden states of the t-th and t-1-th time steps respectively; Y represents the output data; that is, the prediction result; W h and W x are the weight matrix of the hidden state and the weight matrix input to the hidden state; b is the bias vector; Y dense is the output of the fully connected layer, W out and b out are the weight matrix and bias vector of the output layer respectively.
6. A multi-source reactive power compensation control method according to claim 1, characterized in that: The calculation formula for the total reactive power output of the multi-source heterogeneous reactive power compensation integrated system is: Q total =Q SVG +Q SVC +Q C +Q IND +Q STC ≥Q 缺额 In the formula, Q total It is the total reactive power output of the multi-source heterogeneous reactive power compensation integrated system; Q SVG , Q SVC , Q C , Q IND , Q STC They are the reactive output of SVG, SVC, capacitor, reactance and synchronous condenser; Q 缺额 is the reactive power deficit of the grid; is the rated reactive capacity of SVG; is the rated reactive capacity of the SVC; is the rated reactive capacity of the capacitor; is the rated reactive capacity of the reactance; It is the rated reactive capacity of the synchronous condenser.
7. A multi-source reactive power compensation control method according to claim 1, characterized in that: Voltage U of the grid-connected node of multi-source heterogeneous reactive power compensation system i Need to meet: U min ≤U i ≤U max , U min and U max These are the minimum and maximum allowable voltage values, respectively.
8. A multi-source reactive power compensation control system, characterized in that: It includes a first module and a second module. The first module is used to extract the dominant influencing features of the reactive power of the power grid as the input of the CNN-LSTM network, predict the reactive power shortage interval of the power grid, and use it as the power allocation target of the multi-source reactive power compensation equipment; The second module is used to construct a power allocation model of a multi-source heterogeneous reactive compensation integrated system with the minimization of system network loss and voltage drop and the stability of the multi-source heterogeneous reactive compensation integrated system as optimization goals; The reactive power loss function is: In the formula, Q loss is the reactive power loss in the system, M is the total number of lines in the system, Q i is the reactive power of the ith node or line, R i is the resistance of the ith circuit, V i The voltage of the ith line; The power quality of the power grid is characterized by voltage drop, and the calculation formula is: In the formula, ΔV i is the voltage drop at the ith node; P i is the active power of the ith node; X i is the reactance of the ith node; The calculation formula of the power grid short circuit ratio is: Where SCR is the short circuit ratio; S sc is the short-circuit capacity of a node; S r is the rated capacity of the access equipment; The objective function of the power allocation model is: In the formula, is the stability evaluation index; γ0 is the expected stability margin of the system; γ is the actual stability margin of the system; Construct a power distribution model for a multi-source heterogeneous reactive power compensation integrated system based on TD3's deep reinforcement learning algorithm; The deep reinforcement learning algorithm process of TD3 includes the input state space, the update process of the Critic network, the update process of the Actor network, and the update process of the target network; The input state space is formulated as: Oh t =[S t ,A t ,R t ,S t+1 ]; S t ={Y t ,V t ,Q i,t }; In the formula, Ω t is the input state space; S t A is the input state quantity at time t, including the reactive power shortage of the power grid, the voltage data of the power grid nodes and the reactive power distribution state of each reactive compensation device; t The action network is based on the state S t The action generated, R t To perform action A t The resulting action reward; S t+1 is the input state at time t+1; Y t V is the reactive power shortage interval of the power grid; t is the corresponding node voltage data; Q i,t is the reactive power distribution state of each reactive power compensation device at time t; The update process of the Critic network is as follows: First, use the target Actor network to calculate the action in state s': a′=μ′(s′+θ μ′ ); In the formula, a′ represents the action generated according to the state at the next moment; μ′ represents the target Actor network; θ μ′ Represents the parameters of the target Actor network; Secondly, based on the idea of dual networks, calculate the target value: In the formula, y represents the target value; Q′ i Represents the target Critic network; represents the parameters of the target Critic network; γ represents the discount factor, which is used to represent the difference in benefits caused by the order of states; Finally, the gradient descent algorithm is used to minimize the error between the evaluation value and the target value, thereby updating the parameters in the Critic1 and Critic2 networks: In the formula, s is the input state; a is the execution action; represents the gradient of the Critic network update; Q i represents the Critic network; Represents the parameters of the Critic1 network; The update process of the Actor network is as follows: First, use the Actor network to calculate the action in state S: a=μ(s|θ μ ); Where μ represents the Actor network; θ μ Represents the parameters of the Actor network; Secondly, use the Critic1 or Critic2 network to calculate the evaluation value of the state-action pair (s, a). Here, it is assumed that the Critic1 network is used: Where q represents the execution target reward calculated by the Critic network for action a; Q1 represents the Critic network; Represents the parameters of the Critic network; Finally, the gradient ascent algorithm is used to maximize q, thereby completing the update of the Actor network; The update process of the target network is as follows: Introduce a learning rate τ or momentum, take the weighted average of the old target network parameters and the new corresponding network parameters, and then assign them to the target network: In the formula, θ μ′ Represents the parameters of the target Actor network; Represents the parameters of the target Critic network.
9. A computer program product, characterized in that The computer program can implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 7.
11. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Intelligent reactive compensation method with minimum line loss as optimization target
CN115759321A
Multistage power grid real-time scheduling strategy adjustment method, system and device and storage medium
CN115940294A