A dual-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method, device, computer equipment and medium for power system
Through the double closed-loop decomposition prediction fractional-order reinforcement learning power generation control method, using technologies such as complete adaptive noise decomposition and Transformer model, the parameter sensitivity and training overfitting problems of the new power system frequency deviation control are solved, and more efficient power system frequency control is achieved.
Patent Information
- Application Number
- CN202510813853.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing new power system frequency deviation control method has the problems of high sensitivity to system preset parameters and general robustness. In addition, the long short-term memory network signal prediction has the problem of linear growth in data calculation and parameter quantity, which may lead to training overfitting problems.
A double closed-loop decomposition prediction fractional-order reinforcement learning power generation control method is adopted. Through complete adaptive noise complete ensemble empirical mode decomposition, Transformer model, k-means clustering, Q(λ) learning and fractional-order proportional integral differential method, the power system frequency deviation sequence is decomposed for prediction and control action optimization.
It improves control accuracy, reduces frequency fluctuations and power generation energy consumption of the power system, improves power quality, reduces parameter coupling, enhances robustness, and avoids training overfitting.
Smart Images

Figure CN120357556B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of new energy, intelligent power generation control of new power systems, artificial intelligence, deep learning and reinforcement learning, and relates to an artificial intelligence power generation control method, which is suitable for intelligent power generation control of new power systems. Background Art
[0002] The existing single closed-loop control framework for frequency deviation control of new power systems has the problems of high sensitivity to system preset parameters and general robustness.
[0003] The existing power system frequency deviation control method that uses long short-term memory networks for signal prediction has the problem of linear growth in data calculation and parameter quantity, and may cause training overfitting. Summary of the Invention
[0004] Based on this, it is necessary to provide a dual closed-loop decomposition prediction fractional-order reinforcement learning power generation control method, device, computer equipment, computer-readable storage medium and computer program product for the above technical problems.
[0005] In a first aspect, the present application provides a dual-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for a power system. The method comprises:
[0006] The frequency deviation sequence of the power system within a period of time is decomposed using the complete adaptive noise complete ensemble empirical mode decomposition method to obtain the intrinsic mode components and residual components. The intrinsic mode components are input into the Transformer model for prediction to obtain the future intrinsic mode component prediction values predicted by the Transformer model. The future intrinsic mode component prediction values predicted by the Transformer model are clustered using k-means clustering. The k-means clustering category is set to 2 categories. The sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a large sum value is recorded as a large fluctuation signal, and the sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a small sum value is recorded as a small fluctuation signal.
[0007] The Q(λ) learning method is used to reinforce the large fluctuation signal to obtain the output control action of the Q(λ) learning method. In the Q(λ) learning method, the state and action The corresponding Q value is updated as: , is the state at time t+1, that is, the large fluctuation signal at time t+1; is the state at time t; is the action at time t, that is, the output control action of the Q(λ) learning method at time t; is the discount factor; is the reward obtained at time t+1, ; For the next state All possible actions The maximum value of Q value; is the learning rate; Status and action The corresponding Q value; In state and actions The qualification trace factor under
[0008] The outer loop fractional-order proportional-integral-differential method is used to follow the obtained small fluctuation signal and obtain the output control action of the outer loop fractional-order proportional-integral-differential method;
[0009] The output control action of the Q(λ) learning method and the output control action of the outer loop fractional-order proportional-integral-differential method are added together to obtain the outer loop power generation command of the power system unit at the current moment;
[0010] The outer loop power generation instruction of the power system unit at the current moment is subtracted from the inner loop power generation instruction of the power system unit at the previous moment, and then the inner loop power generation instruction of the power system unit at the current moment is obtained through the inner loop fractional-order proportional integral differential method.
[0011] In a second aspect, the present application also provides a dual-closed-loop decomposition prediction fractional-order reinforcement learning power generation control device for a power system. The device comprises:
[0012] The frequency deviation decomposition and clustering module is used to decompose the frequency deviation sequence of the power system within a period of time using the complete adaptive noise complete set empirical mode decomposition method to obtain the intrinsic mode components and residual components, input the intrinsic mode components into the Transformer model for prediction, and obtain the future intrinsic mode component prediction values predicted by the Transformer model. The future intrinsic mode component prediction values predicted by the Transformer model are clustered using k-means clustering, and the k-means clustering category is set to 2 categories. The sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a large sum value is recorded as a large fluctuation signal, and the sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a small sum value is recorded as a small fluctuation signal.
[0013] The reinforcement learning module is used to perform reinforcement learning on the large fluctuation signal using the Q(λ) learning method to obtain the output control action of the Q(λ) learning method. In the Q(λ) learning method, the state and action The corresponding Q value is updated as: , is the state at time t+1, that is, the large fluctuation signal at time t+1; is the state at time t; is the action at time t, that is, the output control action of the Q(λ) learning method at time t; is the discount factor; is the reward obtained at time t+1, ; For the next state All possible actions The maximum value of Q value; is the learning rate; Status and action The corresponding Q value; In state and actions The qualification trace factor under
[0014] The analytical control method module is used to perform follow-up control on the obtained small fluctuation signal using the outer loop fractional-order proportional-integral-differential method to obtain the output control action of the outer loop fractional-order proportional-integral-differential method;
[0015] An outer loop summation module is used to add the output control action of the Q(λ) learning method and the output control action of the outer loop fractional-order proportional-integral-differential method to obtain the outer loop power generation command of the power system unit at the current moment;
[0016] The inner loop power generation instruction solving module is used to subtract the outer loop power generation instruction of the power system unit at the current moment from the inner loop power generation instruction of the power system unit at the previous moment, and then obtain the inner loop power generation instruction of the power system unit at the current moment through the inner loop fractional-order proportional integral differential method.
[0017] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are performed:
[0018] The frequency deviation sequence of the power system within a period of time is decomposed using the complete adaptive noise complete ensemble empirical mode decomposition method to obtain the intrinsic mode components and residual components. The intrinsic mode components are input into the Transformer model for prediction to obtain the future intrinsic mode component prediction values predicted by the Transformer model. The future intrinsic mode component prediction values predicted by the Transformer model are clustered using k-means clustering. The k-means clustering category is set to 2 categories. The sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a large sum value is recorded as a large fluctuation signal, and the sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a small sum value is recorded as a small fluctuation signal.
[0019] The Q(λ) learning method is used to reinforce the large fluctuation signal to obtain the output control action of the Q(λ) learning method. In the Q(λ) learning method, the state and action The corresponding Q value is updated as: , is the state at time t+1, that is, the large fluctuation signal at time t+1; is the state at time t; is the action at time t, that is, the output control action of the Q(λ) learning method at time t; is the discount factor; is the reward obtained at time t+1, ; For the next state All possible actions The maximum value of Q value; is the learning rate; Status and action The corresponding Q value; In state and actions The qualification trace factor under
[0020] The outer loop fractional-order proportional-integral-differential method is used to follow the obtained small fluctuation signal and obtain the output control action of the outer loop fractional-order proportional-integral-differential method;
[0021] The output control action of the Q(λ) learning method and the output control action of the outer loop fractional-order proportional-integral-differential method are added together to obtain the outer loop power generation command of the power system unit at the current moment;
[0022] The outer loop power generation instruction of the power system unit at the current moment is subtracted from the inner loop power generation instruction of the power system unit at the previous moment, and then the inner loop power generation instruction of the power system unit at the current moment is obtained through the inner loop fractional-order proportional integral differential method.
[0023] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps:
[0024] The frequency deviation sequence of the power system within a period of time is decomposed using the complete adaptive noise complete ensemble empirical mode decomposition method to obtain the intrinsic mode components and residual components. The intrinsic mode components are input into the Transformer model for prediction to obtain the future intrinsic mode component prediction values predicted by the Transformer model. The future intrinsic mode component prediction values predicted by the Transformer model are clustered using k-means clustering. The k-means clustering category is set to 2 categories. The sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a large sum value is recorded as a large fluctuation signal, and the sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a small sum value is recorded as a small fluctuation signal.
[0025] The Q(λ) learning method is used to reinforce the large fluctuation signal to obtain the output control action of the Q(λ) learning method. In the Q(λ) learning method, the state and action The corresponding Q value is updated as: , is the state at time t+1, that is, the large fluctuation signal at time t+1; is the state at time t; is the action at time t, that is, the output control action of the Q(λ) learning method at time t; is the discount factor; is the reward obtained at time t+1, ; For the next state All possible actions The maximum value of Q value; is the learning rate; Status and action The corresponding Q value; In state and actions The qualification trace factor under
[0026] The outer loop fractional-order proportional-integral-differential method is used to follow the obtained small fluctuation signal and obtain the output control action of the outer loop fractional-order proportional-integral-differential method;
[0027] The output control action of the Q(λ) learning method and the output control action of the outer loop fractional-order proportional-integral-differential method are added together to obtain the outer loop power generation command of the power system unit at the current moment;
[0028] The outer loop power generation instruction of the power system unit at the current moment is subtracted from the inner loop power generation instruction of the power system unit at the previous moment, and then the inner loop power generation instruction of the power system unit at the current moment is obtained through the inner loop fractional-order proportional integral differential method.
[0029] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the following steps:
[0030] The frequency deviation sequence of the power system within a period of time is decomposed using the complete adaptive noise complete ensemble empirical mode decomposition method to obtain the intrinsic mode components and residual components. The intrinsic mode components are input into the Transformer model for prediction to obtain the future intrinsic mode component prediction values predicted by the Transformer model. The future intrinsic mode component prediction values predicted by the Transformer model are clustered using k-means clustering. The k-means clustering category is set to 2 categories. The sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a large sum value is recorded as a large fluctuation signal, and the sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a small sum value is recorded as a small fluctuation signal.
[0031] The Q(λ) learning method is used to reinforce the large fluctuation signal to obtain the output control action of the Q(λ) learning method. In the Q(λ) learning method, the state and action The corresponding Q value is updated as: , is the state at time t+1, that is, the large fluctuation signal at time t+1; is the state at time t; is the action at time t, that is, the output control action of the Q(λ) learning method at time t; is the discount factor; is the reward obtained at time t+1, ; For the next state All possible actions The maximum value of Q value; is the learning rate; Status and action The corresponding Q value; In state and actions The qualification trace factor under
[0032] The outer loop fractional-order proportional-integral-differential method is used to follow the obtained small fluctuation signal and obtain the output control action of the outer loop fractional-order proportional-integral-differential method;
[0033] The output control action of the Q(λ) learning method and the output control action of the outer loop fractional-order proportional-integral-differential method are added together to obtain the outer loop power generation command of the power system unit at the current moment;
[0034] The outer loop power generation instruction of the power system unit at the current moment is subtracted from the inner loop power generation instruction of the power system unit at the previous moment, and then the inner loop power generation instruction of the power system unit at the current moment is obtained through the inner loop fractional-order proportional integral differential method.
[0035] The above-mentioned double closed-loop decomposition prediction fractional-order reinforcement learning power generation control method, device, computer equipment, storage medium and computer program product of the power system decompose the frequency deviation sequence of the power system within a period of time by using the complete adaptive noise complete set empirical mode decomposition method to obtain the intrinsic mode component and the residual component, input the intrinsic mode component into the Transformer model for prediction, and obtain the future intrinsic mode component prediction value predicted by the Transformer model, and use k-means clustering to cluster the obtained future intrinsic mode component prediction value predicted by the Transformer model, set the k-means clustering category to 2 categories, and the sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a large sum value is recorded as a large fluctuation signal, and the sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a small sum value is recorded as a small fluctuation signal; use the Q(λ) learning method to perform reinforcement learning on the obtained large fluctuation signal to obtain the output control action of the Q(λ) learning method, in the Q(λ) learning method, the state and action The corresponding Q value is updated as: , is the state at time t+1, that is, the large fluctuation signal at time t+1; is the state at time t; is the action at time t, that is, the output control action of the Q(λ) learning method at time t; is the discount factor; is the reward obtained at time t+1, ; For the next state All possible actions The maximum value of Q value; is the learning rate; Status and action The corresponding Q value; In state and actions The qualification trace factor under the condition of Q(λ) is obtained; the outer loop fractional-order proportional integral differential method is used to follow the small fluctuation signal obtained to obtain the output control action of the outer loop fractional-order proportional integral differential method; the output control action of the Q(λ) learning method and the output control action of the outer loop fractional-order proportional integral differential method are added to obtain the outer loop power generation instruction of the power system unit at the current moment; the outer loop power generation instruction of the power system unit at the current moment is subtracted from the inner loop power generation instruction of the power system unit at the previous moment, and then the inner loop power generation instruction of the power system unit at the current moment is obtained through the inner loop fractional-order proportional integral differential method; the control accuracy and power quality can be improved, the frequency fluctuation of the power system can be reduced, and the power generation energy consumption can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a power system power generation control framework diagram of the method of the present invention.
[0037] Figure 2 It is a workflow diagram of the Transformer model of the method of the present invention.
[0038] Figure 3 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0040] In one embodiment, Figure 1 As shown, a dual closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for a power system is provided. This embodiment uses the method applied to a terminal as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server.
[0041] A dual-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for power systems can combine a complete adaptive noise complete set empirical mode decomposition method, a Transformer model, k-means clustering, a fractional-order proportional integral differential method, and a Q(λ) learning method using a dual-closed-loop control architecture for controlling power system frequency deviation. In this embodiment, the method includes the following steps:
[0042] Step (1) uses the complete adaptive noise complete set empirical mode decomposition method to decompose the frequency deviation sequence of the power system within a period of time to obtain intrinsic mode components and residual components, inputs the intrinsic mode components into the Transformer model for prediction, obtains the future intrinsic mode component prediction values predicted by the Transformer model, uses k-means clustering to cluster the obtained future intrinsic mode component prediction values predicted by the Transformer model, sets the k-means clustering category to 2 categories, and the sum of the future intrinsic mode component prediction value signals predicted by the Transformer model in the cluster with a large sum value is recorded as a large fluctuation signal, and the sum of the future intrinsic mode component prediction value signals predicted by the Transformer model in the cluster with a small sum value is recorded as a small fluctuation signal.
[0043] In region A, obtain the frequency deviation sequence of the power system over a period of time , , is the reference time point, For arrive 128 consecutive time points;
[0044] The frequency deviation sequence is analyzed by using the complete adaptive noise complete set empirical mode decomposition method. Decompose, set the number of iterations to K, the original signal and the initial residual Both , construct a decomposition sequence :
[0045] (1)
[0046] Where, is the signal-to-noise ratio of the first decomposition; It is the first component after EMD decomposition; is the i-th white noise added;
[0047] Calculate the first residual component and the first eigenmode component :
[0048] (2)
[0049] (3)
[0050] Where, is the mean value function, To generate the local mean function of the signal;
[0051] Calculate the second residual component and the second eigenmode component :
[0052] (4)
[0053] (5)
[0054] Where, is the second decomposition signal-to-noise ratio; It is the second component after EMD decomposition;
[0055] Similarly, the kth residual component and the kth eigenmode component :
[0056] (6)
[0057] (7)
[0058] Where, is the k-1th residual component; is the k-th decomposition signal-to-noise ratio; is the kth component after EMD decomposition;
[0059] Thus, after K iterations, K intrinsic mode components and the Kth residual component are obtained. The proposed complete adaptive noise complete set empirical mode decomposition method is different from the fully integrated empirical mode decomposition and adaptive noise CEEMDAN algorithm; Formula (1), Formula (4) and Formula (6) are the main differences between the proposed complete adaptive noise complete set empirical mode decomposition method and the fully integrated empirical mode decomposition and adaptive noise CEEMDAN algorithm. The decomposition sequence constructed by the fully integrated empirical mode decomposition and adaptive noise CEEMDAN algorithm is for Secondly, formula (3), formula (5) and formula (7) are also different from the fully integrated empirical mode decomposition and adaptive noise CEEMDAN algorithm. The CEEMDAN algorithm is sometimes also translated as the complete adaptive noise ensemble empirical mode decomposition method.
[0060] Input K intrinsic mode components into the Transformer model for prediction, and obtain K future intrinsic mode component prediction values predicted by the Transformer model;
[0061] The Transformer model consists of an embedding layer, an encoder, and a decoder. The encoder consists of N l The input signal is converted into a fixed-length vector through the embedding layer and then added with position encoding. The input signal is then output to the decoder through the multi-head attention mechanism of the encoder. The decoder consists of N l The decoder consists of a decoding unit, each of which contains a multi-head attention mechanism, a feedforward layer, a residual connection, a layer normalization layer, and a mask attention mechanism. The output of the decoder passes through a fully connected layer and an exponential normalization layer to obtain the output signal of the Transformer model;
[0062] IMF input k The (T) sequence is converted into a fixed-length vector through the embedding layer and the position encoding process is added as follows:
[0063] (8)
[0064] (9)
[0065] Where pos is the current signal in IMF k (T) Position in the sequence; d represents the dimension value of the Transformer model set by PE; 2i represents the index of the even dimension, and 2i+1 represents the index of the odd dimension; is a sine function; is the cosine function; is the value of the encoding vector corresponding to the posth position in the position encoding in the even dimension; The value of the encoding vector corresponding to the posth position in the position encoding in odd dimensions;
[0066] The multi-head attention mechanism includes the query matrix Q, the key-value matrix H and the value matrix V;
[0067] The query matrix Q, key matrix H and value matrix V are transformed by matrix x and the set linear transformation matrix W. Q 、W H With W V Perform linear transformation to obtain;
[0068] The linear transformation of the query matrix Q is:
[0069] (10)
[0070] Where Q is the query matrix; W Q is the linear transformation matrix; is matrix multiplication;
[0071] The linear transformation of the key-value matrix H is:
[0072] (11)
[0073] Where H is the key value matrix; W H is the linear transformation matrix;
[0074] The linear transformation of the value matrix V is:
[0075] (12)
[0076] Where V is the value matrix; W V is the linear transformation matrix;
[0077] The output matrix of the self-attention mechanism for:
[0078] (13)
[0079] Where, is the transposed matrix of the key-value matrix; is the number of columns of the key-value matrix, that is, the column vector dimension; for The square root of yes The activation function of
[0080] The encoder uses a successive linear superposition mechanism to perform weighted fusion of latent state parameters and latent state parameters during the iterative calculation process, thereby achieving a high-dimensional feature representation of the matrix x; the output of the encoder is:
[0081] (14)
[0082] in, is the output function of the encoder, that is, the layer normalization layer output function; is the matrix output after the matrix x passes through the attention mechanism;
[0083] After the Transformer model prediction, a total of K future intrinsic mode component prediction values predicted by the Transformer model are obtained , the predicted values of the future intrinsic mode components predicted by the K Transformer models are As the input of k-means clustering, the future intrinsic mode component prediction values of the K Transformer models obtained by k-means clustering are Perform clustering, set the k-means clustering category to 2 categories, and record the sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a large sum value as a large fluctuation signal, and record the sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a small sum value as a small fluctuation signal;
[0084] from Randomly select 2 samples as the initial cluster centers, denoted as and , set the category of k-means clustering to 2, that is, the number of categories of k-means clustering obtained is 2, for For each sample in , calculate The Euclidean distance between each sample in and the two cluster centers is calculated, and the Each sample in is divided into the class corresponding to the cluster center with the smallest distance;
[0085] For 2 categories, recalculate the cluster centers of the 2 categories:
[0086] (15)
[0087] Where, is the current cluster center, is the new cluster center, and x is the signal in the current category;
[0088] Repeat The steps of classifying each sample and recalculating the cluster center in are repeated until the position of the cluster center no longer changes, that is, two classifications are obtained;
[0089] The predicted values of the future intrinsic mode components of the same category after clustering are summed up, and the predicted values of the future intrinsic mode components predicted by the Transformer model in the cluster with the largest sum value are The sum is recorded as a large fluctuation signal, and the future eigenmode component prediction value predicted by the Transformer model in the cluster with a small sum value The sum of is recorded as a small fluctuation signal;
[0090] Step (2), using the Q(λ) learning method to perform reinforcement learning on the obtained large fluctuation signal, and obtaining the output control action of the Q(λ) learning method;
[0091] In the Q(λ) learning method, in the state and actions The eligibility trace factor is:
[0092] (16)
[0093] Where, In state and actions The qualification trace factor under
[0094] Status and action The corresponding Q value is updated as:
[0095] (17)
[0096] Where, is the state at time t+1, that is, the large fluctuation signal at time t+1; is the state at time t; is the action at time t, that is, the output control action of the Q(λ) learning method at time t; is the discount factor; is the reward obtained at time t+1, ; For the next state All possible actions The maximum value of Q value; is the learning rate; Status and action The corresponding Q value;
[0097] After the large fluctuation signal is reinforced by the Q(λ) learning method, the output control action of the Q(λ) learning method is output ;
[0098] Step (3), using the outer loop fractional order proportional integral differential method to perform follow control on the obtained small fluctuation signal, and obtain the output control action of the outer loop fractional order proportional integral differential method;
[0099] In the outer loop fractional-order proportional-integral-differential method, the transfer function of the outer loop fractional-order proportional-integral-differential method is for:
[0100] (18)
[0101] Where, is the outer ring proportional adjustment coefficient; is the outer loop integral adjustment coefficient; is the outer loop differential adjustment coefficient; is the outer loop fractional integration order; is the fractional differential order of the outer loop;
[0102] Output control action of the outer loop fractional-order proportional-integral-differential method for:
[0103] (19)
[0104] Where, is the output control action of the outer loop fractional-order proportional-integral-differential method; is the inverse Laplace transform function; is the Laplace transform of the small fluctuation signal;
[0105] The input signal of the outer loop is the frequency deviation of the power system. The transmission path of the input signal of the outer loop passes through the complete adaptive noise complete set empirical mode decomposition method, Transformer model, k-means clustering and fractional proportional integral differential method or Q(λ) learning method;
[0106] In step (4), the output control action of the Q(λ) learning method and the output control action of the outer loop fractional-order proportional integral differential method are added together to obtain the outer loop power generation instruction of the power system unit at the current moment.
[0107] Step (5) is to subtract the inner loop power generation instruction of the power system unit at the previous moment from the outer loop power generation instruction of the power system unit at the current moment, and then obtain the inner loop power generation instruction of the power system unit at the current moment through the inner loop fractional-order proportional integral differential method.
[0108] The output control action of the Q(λ) learning method and the output control action of the fractional-order proportional-integral-differential method After addition, it serves as the output signal of the outer loop and also as the input signal of the inner loop;
[0109] In the inner loop fractional-order proportional-integral-differential method, the transfer function of the inner loop fractional-order proportional-integral-differential method is for:
[0110] (20)
[0111] Where, is the inner ring proportional adjustment coefficient; is the inner loop integral adjustment coefficient; is the inner loop differential adjustment coefficient; is the fractional integration order of the inner loop; is the fractional differential order of the inner loop;
[0112] The inner loop power generation instructions of the power system units at the current moment are:
[0113] (twenty one)
[0114] Where, It is the inner loop power generation instruction of the power system unit at the previous moment; It is the inner loop power generation instruction of the power system unit at the current moment.
[0115] The present invention has the following advantages and effects compared to the prior art:
[0116] (1) The existing single closed-loop control framework for frequency deviation control of new power systems has the problems of high sensitivity to system preset parameters and general robustness. The dual closed-loop architecture proposed in the present invention can reduce parameter coupling, adapt to parameter changes, and has high robustness.
[0117] (2) The existing power system frequency deviation control method that uses long short-term memory networks for signal prediction has the problem of linear growth in the amount of data calculation and the number of parameters, which may lead to overfitting in training. The Transformer model in the present invention can reduce excessive reliance on local features of the data and effectively solve and prevent the problem of overfitting in training. Figure 1 This is a diagram of the power system power generation control framework of the method of the present invention. First, frequency deviations are collected from the power system and decomposed into multiple signals using the complete adaptive noise complete set empirical mode decomposition method. Then, the decomposed multiple signals are input into the Transformer model for prediction, and the predicted values of the multiple future intrinsic mode components are classified using k-means clustering into large fluctuation signals and small fluctuation signals. The large fluctuation signals are learned and followed using Q(λ) learning; the small fluctuation signals are learned and followed using fractional-order proportional integral differential. The output values of Q(λ) learning and fractional-order proportional integral differential are summed to serve as the outer-loop power generation command of the power system unit at the current moment. Finally, the result of subtracting the inner-loop power generation command of the power system unit at the previous moment from the outer-loop power generation command of the power system unit at the current moment is input into the fractional-order proportional integral differential for learning and following. The inner-loop power generation command of the power system unit is output in real time every 4 seconds to the power system unit for real-time control.
[0118] Figure 2 This is the workflow diagram of the Transformer model of the method of the present invention. The Transformer model consists of an embedding layer, an encoder and a decoder. The encoder consists of N l The input signal is converted into a fixed-length vector through the embedding layer and then added with position encoding. The input signal is then output to the decoder through the multi-head attention mechanism of the encoder. The decoder consists of N l The decoder consists of decoding units, each of which contains a multi-head attention mechanism, a feedforward layer, a residual connection, a layer normalization layer, and a mask attention mechanism. The output of the decoder passes through a fully connected layer and an exponential normalization layer to obtain the output signal of the Transformer model.
[0119] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 3 As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless communication, and the wireless communication can be achieved via Wi-Fi, mobile cellular networks, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a dual-closed-loop decomposition predictive fractional-order reinforcement learning power generation control method for power systems. The display unit of the computer device is used to produce visual images and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0120] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0121] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0122] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0123] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0124] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0125] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0126] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0127] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A dual-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for power systems, characterized by: The method comprises: The frequency deviation sequence of the power system within a period of time is decomposed using the complete adaptive noise complete ensemble empirical mode decomposition method to obtain the intrinsic mode components and residual components. The intrinsic mode components are input into the Transformer model for prediction to obtain the future intrinsic mode component prediction values predicted by the Transformer model. The future intrinsic mode component prediction values predicted by the Transformer model are clustered using k-means clustering. The k-means clustering category is set to 2 categories. The sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a large sum value is recorded as a large fluctuation signal, and the sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a small sum value is recorded as a small fluctuation signal. The Q(λ) learning method is used to reinforce the large fluctuation signal to obtain the output control action of the Q(λ) learning method. In the Q(λ) learning method, the state and action The corresponding Q value is updated as: , is the state at time t+1, that is, the large fluctuation signal at time t+1; is the state at time t; is the action at time t, that is, the output control action of the Q(λ) learning method at time t; is the discount factor; is the reward obtained at time t+1, ; For the next state All possible actions The maximum value of Q value; is the learning rate; Status and action The corresponding Q value; In state and actions The qualification trace factor under The outer loop fractional-order proportional-integral-differential method is used to follow the obtained small fluctuation signal and obtain the output control action of the outer loop fractional-order proportional-integral-differential method; The output control action of the Q(λ) learning method and the output control action of the outer loop fractional-order proportional-integral-differential method are added together to obtain the outer loop power generation command of the power system unit at the current moment; The outer loop power generation instruction of the power system unit at the current moment is subtracted from the inner loop power generation instruction of the power system unit at the previous moment, and then the inner loop power generation instruction of the power system unit at the current moment is obtained through the inner loop fractional-order proportional integral differential method.
2. The double closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for the power system according to claim 1 is characterized in that: The Transformer model consists of an embedding layer, an encoder, and a decoder. The encoder consists of N l The input signal is converted into a fixed-length vector through the embedding layer and then added with position encoding. The input signal is then output to the decoder through the multi-head attention mechanism of the encoder. The decoder consists of N l The decoder consists of decoding units, each of which contains a multi-head attention mechanism, a feedforward layer, a residual connection, a layer normalization layer, and a mask attention mechanism. The output of the decoder passes through a fully connected layer and an exponential normalization layer to obtain the output signal of the Transformer model. The multi-head attention mechanism includes a query matrix, a key-value matrix, and a value matrix.
3. The double closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for the power system according to claim 1 is characterized in that: In the outer loop fractional order proportional integral differential method, the transfer function of the outer loop fractional order proportional integral differential method is for: , where is the outer ring proportional adjustment coefficient; is the outer loop integral adjustment coefficient; is the outer loop differential adjustment coefficient; is the outer loop fractional integration order; is the fractional differential order of the outer loop.
4. The double closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for the power system according to claim 1 is characterized in that: In the inner loop fractional order proportional integral differential method, the transfer function of the inner loop fractional order proportional integral differential method is for: , where is the inner ring proportional adjustment coefficient; is the inner loop integral adjustment coefficient; is the inner loop differential adjustment coefficient; is the fractional integration order of the inner loop; is the fractional differential order of the inner loop.
5. The double closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for the power system according to claim 1 is characterized in that: The steps of using the complete adaptive noise complete set empirical mode decomposition method to decompose the frequency deviation sequence of the power system within a period of time to obtain the intrinsic mode components and residual components are as follows: In region A, obtain the frequency deviation sequence of the power system over a period of time , , is the reference time point, For arrive 128 consecutive time points; Set the number of iterations to K, the original signal and the initial residual Both , construct a decomposition sequence : , where is the signal-to-noise ratio of the first decomposition; It is the first component after EMD decomposition; is the i-th white noise added; Calculate the first residual component and the first eigenmode component : , , where is the mean value function, To generate the local mean function of the signal; Calculate the second residual component and the second eigenmode component : , , where is the second decomposition signal-to-noise ratio; It is the second component after EMD decomposition; Similarly, the kth residual component and the kth eigenmode component : , , where is the k-1th residual component; is the k-th decomposition signal-to-noise ratio; is the kth component after EMD decomposition; Thus, after K iterations, K intrinsic mode components and the Kth residual component are obtained.
6. A dual-closed-loop decomposition prediction fractional-order reinforcement learning power generation control device for power systems, characterized in that: The device comprises: The frequency deviation decomposition and clustering module is used to decompose the frequency deviation sequence of the power system within a period of time using the complete adaptive noise complete set empirical mode decomposition method to obtain the intrinsic mode components and residual components, input the intrinsic mode components into the Transformer model for prediction, and obtain the future intrinsic mode component prediction values predicted by the Transformer model. The future intrinsic mode component prediction values predicted by the Transformer model are clustered using k-means clustering, and the k-means clustering category is set to 2 categories. The sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a large sum value is recorded as a large fluctuation signal, and the sum of the future intrinsic mode component prediction values predicted by the Transformer model in the cluster with a small sum value is recorded as a small fluctuation signal. The reinforcement learning module is used to perform reinforcement learning on the large fluctuation signal using the Q(λ) learning method to obtain the output control action of the Q(λ) learning method. In the Q(λ) learning method, the state and action The corresponding Q value is updated as: , is the state at time t+1, that is, the large fluctuation signal at time t+1; is the state at time t; is the action at time t, that is, the output control action of the Q(λ) learning method at time t; is the discount factor; is the reward obtained at time t+1, ; For the next state All possible actions The maximum value of Q value; is the learning rate; Status and action The corresponding Q value; In state and actions The qualification trace factor under The analytical control method module is used to perform follow-up control on the obtained small fluctuation signal using the outer loop fractional-order proportional-integral-differential method to obtain the output control action of the outer loop fractional-order proportional-integral-differential method; An outer loop summation module is used to add the output control action of the Q(λ) learning method and the output control action of the outer loop fractional-order proportional-integral-differential method to obtain the outer loop power generation command of the power system unit at the current moment; The inner loop power generation instruction solving module is used to subtract the outer loop power generation instruction of the power system unit at the current moment from the inner loop power generation instruction of the power system unit at the previous moment, and then obtain the inner loop power generation instruction of the power system unit at the current moment through the inner loop fractional-order proportional integral differential method.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the double closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for the power system described in any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the double closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for the power system described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Power grid frequency intelligent control method based on empirical mode decomposition
CN112398142A
Intelligent power grid voltage control method based on prediction fractional order reinforcement learning
CN115659804A