Double-closed-loop decomposition prediction fractional order reinforcement learning power generation control method and device of power system, computer equipment and medium

Through the double closed-loop decomposition prediction fractional-order reinforced learning power generation control method, the complete adaptive noise decomposition and Transformer model classification signal is used, combined with Q(λ) learning and fractional-order integral differential method, the robustness and training overfitting problems of frequency deviation control in the new power system are solved, and high-precision power system frequency stability and energy consumption optimization are achieved.

CN120357556AActive Publication Date: 2025-07-22GUANGXI UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510813853.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-22
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The existing new power system frequency deviation control method has high sensitivity and average robustness in the system preset parameters, and the data calculation amount and parameter amount of long and short-term memory network signal prediction have linear growth, which may cause training overfitting.

Method used

The double closed-loop decomposition prediction fractional-order reinforcement learning power generation control method is adopted, and the frequency deviation sequence is decomposed using the complete adaptive noise complete ensemble empirical mode decomposition method, and signal classification is performed by combining the Transformer model and k-mean clustering. Reinforcement learning and follow-up control are performed through Q(λ) learning and fractional-order proportional integral differential method to output power generation instructions of the power system unit.

Benefits of technology

It improves control accuracy, reduces frequency fluctuations in the power system, improves power quality, and reduces power generation energy consumption, and enhances the robustness and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357556A_ABST
    Figure CN120357556A_ABST
Patent Text Reader

Abstract

The invention provides a double-closed-loop decomposition prediction fractional order reinforcement learning power generation control method and device of a power system, computer equipment and a medium. A control framework of the method comprises an outer ring and an inner ring, and the frequency deviation of the power system is used as the input of the outer ring; firstly, a complete adaptive noise complete set empirical mode decomposition method is adopted to perform mode decomposition on frequency deviation; secondly, a Transform model is introduced to carry out prediction on the decomposed signals; classifying the signals on the basis of k-means clustering; and finally, through the synergistic effect of fractional order proportional integral differential and Q (lambda) learning, outputting an outer ring power generation instruction of the power system unit and taking the outer ring power generation instruction as the input of an inner ring. And the inner ring adopts a fractional order proportional integral differential method to learn and follow and output an inner ring power generation instruction of the power system unit. The method can solve the problem of uncertainty of power generation control in a novel power system, reduces frequency deviation, and improves control precision and system robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of new energy, intelligent power generation control of new power systems, artificial intelligence, deep learning, and reinforcement learning, and relates to a power generation control method of artificial intelligence, which is applicable to the intelligent power generation control of new power systems. Background Art

[0002] In the existing single-closed-loop control framework for frequency deviation control of new power systems, there are problems of high sensitivity to system preset parameters and general robustness.

[0003] In the existing power system frequency deviation control method using long short-term memory networks for signal prediction, there are problems of linear growth in data calculation volume and the number of parameters, and possible overfitting during training. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a dual-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method, device, computer device, computer-readable storage medium, and computer program product for a power system.

[0005] In a first aspect, the present application provides a dual-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for a power system. The method includes:

[0006] Using the complete adaptive noise complete ensemble empirical mode decomposition method to decompose the frequency deviation sequence of the power system over a period of time to obtain the intrinsic mode components and the residual component, inputting the intrinsic mode components into the Transformer model for prediction to obtain the predicted values of the future intrinsic mode components predicted by the Transformer model, using k-means clustering to cluster the predicted values of the future intrinsic mode components predicted by the Transformer model, setting the number of categories of k-means clustering to 2, and denoting the sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the clustering with the larger sum value as the large fluctuation signal, and denoting the sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the clustering with the smaller sum value as the small fluctuation signal;

[0007] Using the Q(λ) learning method to perform reinforcement learning on the obtained large fluctuation signal to obtain the output control action of the Q(λ) learning method;

[0008] Using the outer-loop fractional-order proportional integral derivative method to perform follow-up control on the obtained small fluctuation signal to obtain the output control action of the outer-loop fractional-order proportional integral derivative method;

[0009] Adding the output control action of the Q(λ) learning method and the output control action of the outer-loop fractional-order proportional integral derivative method to obtain the outer-loop power generation command of the power system unit;

[0010] Subtract the inner-loop power generation command of the power system unit at the previous moment from the outer-loop power generation command of the power system unit, and then obtain the inner-loop power generation command of the power system unit through the inner-loop fractional-order proportional integral derivative method.

[0011] In a second aspect, the present application also provides a double-loop decomposition prediction fractional-order reinforcement learning power generation control device for a power system. The device includes:

[0012] A frequency deviation decomposition and clustering module, which is used to decompose the frequency deviation sequence of the power system over a period of time by using the complete ensemble empirical mode decomposition method with complete adaptive noise to obtain the intrinsic mode components and the residual components, input the intrinsic mode components into the Transformer model for prediction to obtain the predicted values of the future intrinsic mode components predicted by the Transformer model, perform k-means clustering on the predicted values of the future intrinsic mode components predicted by the obtained Transformer model, set the number of categories of k-means clustering to 2, and record the sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the cluster with the larger sum value as the large fluctuation signal, and record the sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the cluster with the smaller sum value as the small fluctuation signal;

[0013] A reinforcement learning module, which is used to perform reinforcement learning on the obtained large fluctuation signal by using the Q(λ) learning method to obtain the output control action of the Q(λ) learning method;

[0014] An analytical control method module, which is used to perform follow-up control on the obtained small fluctuation signal by using the outer-loop fractional-order proportional integral derivative method to obtain the output control action of the outer-loop fractional-order proportional integral derivative method;

[0015] An outer-loop summation module, which is used to add the output control action of the Q(λ) learning method and the output control action of the outer-loop fractional-order proportional integral derivative method to obtain the outer-loop power generation command of the power system unit;

[0016] An inner-loop power generation command solving module, which is used to subtract the inner-loop power generation command of the power system unit at the previous moment from the outer-loop power generation command of the power system unit, and then obtain the inner-loop power generation command of the power system unit through the inner-loop fractional-order proportional integral derivative method.

[0017] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0018] The complete adaptive noise complete ensemble empirical mode decomposition method is used to decompose the frequency deviation sequence of the power system over a period of time to obtain the intrinsic mode components and the residual components. The intrinsic mode components are input into the Transformer model for prediction to obtain the predicted values of the future intrinsic mode components predicted by the Transformer model. K-means clustering is used to cluster the predicted values of the future intrinsic mode components predicted by the Transformer model. The number of categories of k-means clustering is set to 2. The sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the clustering with the larger sum value is denoted as the large fluctuation signal, and the sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the clustering with the smaller sum value is denoted as the small fluctuation signal;

[0019] The Q(λ) learning method is used to perform reinforcement learning on the obtained large fluctuation signal to obtain the output control action of the Q(λ) learning method;

[0020] The outer-loop fractional-order proportional-integral-derivative method is used to perform following control on the obtained small fluctuation signal to obtain the output control action of the outer-loop fractional-order proportional-integral-derivative method;

[0021] The output control action of the Q(λ) learning method and the output control action of the outer-loop fractional-order proportional-integral-derivative method are added together to obtain the outer-loop power generation command of the power system unit;

[0022] The outer-loop power generation command of the power system unit is subtracted from the inner-loop power generation command of the power system unit at the previous moment, and then through the inner-loop fractional-order proportional-integral-derivative method, the inner-loop power generation command of the power system unit is obtained.

[0023] Fourthly, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the following steps are implemented:

[0024] The complete adaptive noise complete ensemble empirical mode decomposition method is used to decompose the frequency deviation sequence of the power system over a period of time to obtain the intrinsic mode components and the residual components. The intrinsic mode components are input into the Transformer model for prediction to obtain the predicted values of the future intrinsic mode components predicted by the Transformer model. K-means clustering is used to cluster the predicted values of the future intrinsic mode components predicted by the Transformer model. The number of categories of k-means clustering is set to 2. The sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the clustering with the larger sum value is denoted as the large fluctuation signal, and the sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the clustering with the smaller sum value is denoted as the small fluctuation signal;

[0025] The obtained large fluctuation signal is subjected to reinforcement learning using the Q(λ) learning method to obtain the output control action of the Q(λ) learning method;

[0026] The obtained small fluctuation signal is subjected to following control using the outer-loop fractional-order proportional-integral-derivative method to obtain the output control action of the outer-loop fractional-order proportional-integral-derivative method;

[0027] The output control action of the Q(λ) learning method and the output control action of the outer-loop fractional-order proportional-integral-derivative method are added to obtain the outer-loop power generation command of the power system unit;

[0028] The outer-loop power generation command of the power system unit is subtracted from the inner-loop power generation command of the power system unit at the previous moment, and then through the inner-loop fractional-order proportional-integral-derivative method, the inner-loop power generation command of the power system unit is obtained.

[0029] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0030] The complete adaptive noise complete ensemble empirical mode decomposition method is used to decompose the frequency deviation sequence of the power system over a period of time to obtain the intrinsic mode components and the residual component. The intrinsic mode components are input into the Transformer model for prediction to obtain the predicted values of the future intrinsic mode components predicted by the Transformer model. K-means clustering is used to cluster the predicted values of the future intrinsic mode components predicted by the obtained Transformer model. The number of categories of k-means clustering is set to 2. The sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the clustering with the larger sum is denoted as the large fluctuation signal, and the sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the clustering with the smaller sum is denoted as the small fluctuation signal;

[0031] The obtained large fluctuation signal is subjected to reinforcement learning using the Q(λ) learning method to obtain the output control action of the Q(λ) learning method;

[0032] The obtained small fluctuation signal is subjected to following control using the outer-loop fractional-order proportional-integral-derivative method to obtain the output control action of the outer-loop fractional-order proportional-integral-derivative method;

[0033] The output control action of the Q(λ) learning method and the output control action of the outer-loop fractional-order proportional-integral-derivative method are added to obtain the outer-loop power generation command of the power system unit;

[0034] The outer-loop power generation command of the power system unit is subtracted from the inner-loop power generation command of the power system unit at the previous moment, and then through the inner-loop fractional-order proportional-integral-derivative method, the inner-loop power generation command of the power system unit is obtained.

[0035] The double - closed - loop decomposition prediction fractional - order reinforcement learning power generation control method, device, computer device, storage medium and computer program product of the above - mentioned power system decompose the frequency deviation sequence of the power system over a period of time by using the complete adaptive noise complete ensemble empirical mode decomposition method to obtain the intrinsic mode components and the residual component. The intrinsic mode components are input into the Transformer model for prediction to obtain the predicted values of the future intrinsic mode components predicted by the Transformer model. K - means clustering is used to cluster the predicted values of the future intrinsic mode components predicted by the Transformer model. The number of categories of k - means clustering is set to 2. The sum of the predicted values of the future intrinsic mode components in the clustering with the larger sum value is recorded as the large - fluctuation signal, and the sum of the predicted values of the future intrinsic mode components in the clustering with the smaller sum value is recorded as the small - fluctuation signal. The Q(λ) learning method is used to perform reinforcement learning on the obtained large - fluctuation signal to obtain the output control action of the Q(λ) learning method. The outer - loop fractional - order proportional - integral - derivative method is used to perform follow - up control on the obtained small - fluctuation signal to obtain the output control action of the outer - loop fractional - order proportional - integral - derivative method. The output control action of the Q(λ) learning method and the output control action of the outer - loop fractional - order proportional - integral - derivative method are added together to obtain the outer - loop power generation command of the power system unit. The outer - loop power generation command of the power system unit is subtracted from the inner - loop power generation command of the power system unit at the previous moment, and then through the inner - loop fractional - order proportional - integral - derivative method, the inner - loop power generation command of the power system unit is obtained. It can improve the control accuracy, enhance the power quality, reduce the frequency fluctuation of the power system and reduce the power generation energy consumption. Brief Description of the Drawings

[0036] Figure 1 is the power system power generation control framework diagram of the method of the present invention.

[0037] Figure 2 is the working flow chart of the Transformer model of the method of the present invention.

[0038] Figure 3 is the internal structure diagram of a computer device in an embodiment. Detailed Embodiments

[0039] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further describes the present application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0040] In one embodiment, as Figure 1As shown, a double - closed - loop decomposition prediction fractional - order reinforcement learning power generation control method for a power system is provided. In this embodiment, taking the application of this method to a terminal as an example, it can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is realized through the interaction between the terminal and the server.

[0041] A double - closed - loop decomposition prediction fractional - order reinforcement learning power generation control method for a power system can combine the complete adaptive noise complete ensemble empirical mode decomposition method, Transformer model, k - means clustering, fractional - order proportional integral differential method, and Q(λ) learning method using a double - closed - loop control architecture for the control of power system frequency deviation. In this embodiment, the method includes the following steps:

[0042] Step (1): Use the complete adaptive noise complete ensemble empirical mode decomposition method to decompose the frequency deviation sequence of the power system over a period of time, obtaining the intrinsic mode components and the residual component. Input the intrinsic mode components into the Transformer model for prediction to obtain the predicted values of the future intrinsic mode components predicted by the Transformer model. Use k - means clustering to cluster the predicted values of the future intrinsic mode components predicted by the Transformer model. Set the number of categories of k - means clustering to 2. Denote the sum of the predicted values of the future intrinsic mode components in the clustering with the larger sum as the large - fluctuation signal, and denote the sum of the predicted values of the future intrinsic mode components in the clustering with the smaller sum as the small - fluctuation signal.

[0043] In region A, obtain the frequency deviation sequence of the power system over a period of time , , is the reference time point, is from to a continuous 128 time points;

[0044] Use the complete adaptive noise complete ensemble empirical mode decomposition method to decompose the frequency deviation sequence , set the number of iterations to K, the original signal and the initial residual are both , construct the decomposition sequence :

[0045] (1)

[0046] In the formula, is the signal - to - noise ratio of the first - stage decomposition signal; is the first component after EMD decomposition by the empirical mode decomposition method; is the i-th added white noise;

[0047] Calculate the first residual component and the first intrinsic mode component :

[0048] (2)

[0049] (3)

[0050] wherein, is the average value function, is the function for generating the local mean value of the signal;

[0051] Calculate the second residual component and the second intrinsic mode component :

[0052] (4)

[0053] (5)

[0054] wherein, is the signal-to-noise ratio of the second decomposition; is the second component after the empirical mode decomposition method EMD decomposition;

[0055] Similarly, the k-th residual component and the k-th intrinsic mode component :

[0056] (6)

[0057] (7)

[0058] wherein, is the (k - 1)-th residual component; is the signal-to-noise ratio of the k-th decomposition; is the k-th component after the empirical mode decomposition method EMD decomposition;

[0059] Thus, after K iterations, K intrinsic mode components and the K-th residual component are obtained. The proposed complete adaptive noise complete ensemble empirical mode decomposition method is different from the complete ensemble empirical mode decomposition with adaptive noise CEEMDAN algorithm; Formulas (1), (4) and (6) are the main differences between the proposed complete adaptive noise complete ensemble empirical mode decomposition method and the complete ensemble empirical mode decomposition with adaptive noise CEEMDAN algorithm. The decomposition sequence constructed by the complete ensemble empirical mode decomposition with adaptive noise CEEMDAN algorithm is Secondly, Formula (3), Formula (5) and Formula (7) are also different from the Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN) algorithm. The CEEMDAN algorithm is sometimes also translated as the Complete Adaptive Noise Ensemble Empirical Mode Decomposition method.

[0060] Input the K intrinsic mode components into the Transformer model for prediction to obtain the predicted values of the K future intrinsic mode components predicted by the Transformer model;

[0061] The Transformer model consists of an embedding layer, an encoder, and a decoder. The encoder consists of N l encoding units. Each encoding unit contains a multi-head attention mechanism, a feed-forward layer, a residual connection, and a layer normalization layer. After the input signal is converted into a fixed-length vector by the embedding layer and added with positional encoding, it is then output to the multi-head attention mechanism of the decoder through the encoder. The decoder consists of N l decoding units. Each decoding unit contains a multi-head attention mechanism, a feed-forward layer, a residual connection, a layer normalization layer, and a masked attention mechanism. The output of the decoder passes through a fully connected layer and an exponential normalization layer to obtain the output signal of the Transformer model;

[0062] The input IMF k (T) sequence is converted into a fixed-length vector by the embedding layer and added with positional encoding as follows:

[0063] (8)

[0064] (9)

[0065] In the formula, pos is the position of the current signal in the IMF k (T) sequence; d represents the dimension value of the PE setting of the Transformer model; 2i represents the index of the even dimension, and 2i + 1 represents the index of the odd dimension; is the sine function; is the cosine function; is the value of the encoding vector corresponding to the pos-th position in the even dimension in the positional encoding; is the value of the encoding vector corresponding to the pos-th position in the odd dimension in the positional encoding;

[0066] The multi-head attention mechanism includes a query matrix Q, a key-value matrix H, and a value matrix V;

[0067] The query matrix Q, the key-value matrix H, and the value matrix V are obtained through the matrix x and the set linear transformation matrices W Q 、W H and W VObtained by performing a linear transformation;

[0068] The linear transformation of the query matrix Q is:

[0069] (10)

[0070] Where Q is the query matrix; W Q is the linear transformation matrix; is matrix multiplication;

[0071] The linear transformation of the key-value matrix H is:

[0072] (11)

[0073] Where H is the key-value matrix; W H is the linear transformation matrix;

[0074] The linear transformation of the value matrix V is:

[0075] (12)

[0076] Where V is the value matrix; W V is the linear transformation matrix;

[0077] The output matrix of the self-attention mechanism is:

[0078] (13)

[0079] Where is the transpose matrix of the key-value matrix; is the number of columns of the key-value matrix, that is, the column vector dimension; is the arithmetic square root of; is the activation function of;

[0080] The encoder, through the successive linear superposition mechanism, performs weighted fusion of the hidden state parameters with the hidden state parameters during the iterative calculation process, so as to realize the high-dimensional feature representation of the matrix x; the output of the encoder is:

[0081] (14)

[0082] Among them, is the output function of the encoder, that is, the output function of the layer normalization layer; is the matrix output after the matrix x passes through the attention mechanism;

[0083] After being predicted by the Transformer model, a total of K future eigenmode component prediction values predicted by the Transformer model are obtained , the predicted future intrinsic mode component values predicted by the obtained K Transformer models are used as the input of k-means clustering. K-means clustering is used to cluster the predicted future intrinsic mode component values predicted by the obtained K Transformer models into clusters. Set the number of classes of k-means clustering to 2. The sum of the predicted future intrinsic mode component values predicted by the Transformer models in the cluster with the larger sum value is denoted as the large fluctuation signal, and the sum of the predicted future intrinsic mode component values predicted by the Transformer models in the cluster with the smaller sum value is denoted as the small fluctuation signal;

[0084] From , randomly select 2 samples as the initial cluster centers, denoted as and . Set the number of classes of k-means clustering to 2, that is, the number of classes of the finally obtained k-means clustering is 2. For each sample in , calculate the Euclidean distance from each sample in to the 2 cluster centers, and assign each sample in to the class corresponding to the cluster center with the smallest distance;

[0085] For the 2 classes, recalculate the cluster centers of the 2 classes:

[0086] (15)

[0087] In the formula, is the current cluster center, is the new cluster center, and x is the signal in the current class;

[0088] Repeat the steps of classifying each sample in and recalculating the cluster center until the positions of the cluster centers no longer change, that is, two classifications are obtained;

[0089] Perform a summation operation on the predicted future intrinsic mode component values of the same class after clustering. The sum of the predicted future intrinsic mode component values predicted by the Transformer models in the cluster with the larger sum value is denoted as the large fluctuation signal, and the sum of the predicted future intrinsic mode component values predicted by the Transformer models in the cluster with the smaller sum value is denoted as the small fluctuation signal;

[0090] Step (2), use the Q(λ) learning method to perform reinforcement learning on the obtained large fluctuation signal to obtain the output control action of the Q(λ) learning method;

[0091] In the Q(λ) learning method, in the state and the eligibility trace factor under the action is: The eligibility trace factor under the action is:

[0092] (16)

[0093] In the formula, is the eligibility trace factor under the state and the action ;

[0094] Then the Q-value corresponding to the state and the action is updated to:

[0095] (17)

[0096] In the formula, is the state at time t + 1, that is, the large fluctuation signal at time t + 1; is the state at time t; is the action at time t, that is, the output control action of the Q(λ) learning method at time t; is the discount factor; is the reward obtained at time t + 1, which is ; is the maximum value of the Q-values of all possible actions of the next state ; is the learning rate; is the Q-value corresponding to the state and the action ;

[0097] After the large fluctuation signal undergoes reinforcement learning through the Q(λ) learning method, the output control action of the Q(λ) learning method is output;

[0098] Step (3), the obtained small fluctuation signal is subjected to follow-up control by using an outer-loop fractional-order proportional-integral-derivative method to obtain the output control action of the outer-loop fractional-order proportional-integral-derivative method;

[0099] In the outer-loop fractional-order proportional-integral-derivative method, the transfer function of the outer-loop fractional-order proportional-integral-derivative method is:

[0100] (18)

[0101] In the formula, is the outer-loop proportional adjustment coefficient; is the outer-loop integral adjustment coefficient; is the outer-loop derivative adjustment coefficient; is the outer-loop fractional-order integral order; is the outer - loop fractional - order differential order;

[0102] The output control action of the outer - loop fractional - order proportional - integral - derivative method is:

[0103] (19)

[0104] where, is the output control action of the outer - loop fractional - order proportional - integral - derivative method; is the inverse Laplace transform function; is the Laplace transform of the small - fluctuation signal;

[0105] The input signal of the outer - loop is the frequency deviation of the power system. The transmission path of the input signal of the outer - loop passes through the complete adaptive noise complete ensemble empirical mode decomposition method, the Transformer model, the k - means clustering, and the fractional - order proportional - integral - derivative method or the Q(λ) learning method;

[0106] Step (4): Add the output control action of the Q(λ) learning method and the output control action of the outer - loop fractional - order proportional - integral - derivative method to obtain the outer - loop generation command of the power - system unit.

[0107] Step (5): Subtract the inner - loop generation command of the power - system unit at the previous moment from the outer - loop generation command of the power - system unit, and then pass it through the inner - loop fractional - order proportional - integral - derivative method to obtain the inner - loop generation command of the power - system unit.

[0108] The output control action of the Q(λ) learning method and the output control action of the fractional - order proportional - integral - derivative method are added as the output signal of the outer - loop and also as the input signal of the inner - loop;

[0109] In the inner - loop fractional - order proportional - integral - derivative method, the transfer function of the inner - loop fractional - order proportional - integral - derivative method is:

[0110] (20)

[0111] where, is the inner - loop proportional regulation coefficient; is the inner - loop integral regulation coefficient; is the inner - loop differential regulation coefficient; is the inner - loop fractional - order integral order; is the inner - loop fractional - order differential order;

[0112] The inner - loop generation command of the power - system unit is:

[0113] (21)

[0114] In the formula, is the inner-loop power generation command of the power system units at the previous moment; is the inner-loop power generation command of the power system units.

[0115] The present invention has the following advantages and effects compared with the prior art:

[0116] (1) In the existing single-loop control framework for frequency deviation control of a new power system, there are problems of high sensitivity to system preset parameters and general robustness. However, the double-loop architecture proposed by the present invention can reduce parameter coupling, adapt to parameter changes, and has high robustness.

[0117] (2) In the existing method for power system frequency deviation control that uses a long short-term memory network for signal prediction, there are problems that the data calculation amount and the number of parameters increase linearly, and overfitting in training may occur. However, the Transformer model in the present invention can reduce the over-reliance on local features of data and effectively solve and prevent the problem of overfitting in training. Figure 1 is the power system generation control framework diagram of the method of the present invention. First, the frequency deviation is collected from the power system, and the frequency deviation is decomposed into multiple signals by using the complete ensemble empirical mode decomposition method with complete adaptive noise. Then, the multiple decomposed signals are input into the Transformer model for prediction, and the predicted values of multiple future intrinsic mode components are classified using k-means clustering into large-fluctuation signals and small-fluctuation signals. The large-fluctuation signals are learned and followed by Q(λ) learning; the small-fluctuation signals are learned and followed by fractional-order proportional integral derivative, and the output values of Q(λ) learning and fractional-order proportional integral derivative are added as the outer-loop power generation command of the power system units. Finally, the result of subtracting the inner-loop power generation command of the power system units at the previous moment from the outer-loop power generation command of the power system units is input into the fractional-order proportional integral derivative for learning and following, and the inner-loop power generation command of the power system units is output to the power system units for real-time control every 4 seconds.

[0118] Figure 2 is the working flow chart of the Transformer model of the method of the present invention. The Transformer model consists of an embedding layer, an encoder, and a decoder. The encoder consists of N l encoding units. Each encoding unit contains a multi-head attention mechanism, a feed-forward layer, a residual connection, and a layer normalization layer. After the input signal is converted into a fixed-length vector by the embedding layer and added with a position encoding, it is output to the multi-head attention mechanism of the decoder through the encoder. The decoder consists of N lIt consists of several decoding units, and each decoding unit includes a multi-head attention mechanism, a feed-forward layer, a residual connection, a layer normalization layer, and a masked attention mechanism. The output of the decoder passes through a fully connected layer and an exponential normalization layer to obtain the output signal of the Transformer model.

[0119] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 3 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a double-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for a power system. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0120] Those skilled in the art can understand that Figure 3 the structure shown in

[0121] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0122] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the above-mentioned method embodiments.

[0123] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps in the above method embodiments.

[0124] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0125] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0126] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0127] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A double-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for a power system, characterized in that The method includes: Using the complete ensemble empirical mode decomposition with adaptive noise (CEEMDAN) method to decompose the frequency deviation sequence of the power system over a period of time, obtaining the intrinsic mode components and the residual component. Input the intrinsic mode components into the Transformer model for prediction to obtain the predicted values of the future intrinsic mode components predicted by the Transformer model. Use k-means clustering to cluster the predicted values of the future intrinsic mode components predicted by the Transformer model. Set the number of categories of k-means clustering to 2. Denote the sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the clustering with the larger sum as the large fluctuation signal, and denote the sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the clustering with the smaller sum as the small fluctuation signal; Using the Q(λ) learning method to perform reinforcement learning on the obtained large fluctuation signal to obtain the output control action of the Q(λ) learning method; Using the outer-loop fractional-order proportional-integral-derivative method to perform following control on the obtained small fluctuation signal to obtain the output control action of the outer-loop fractional-order proportional-integral-derivative method; Add the output control action of the Q(λ) learning method and the output control action of the outer-loop fractional-order proportional-integral-derivative method to obtain the outer-loop power generation command of the power system unit; Subtract the inner-loop power generation command of the power system unit at the previous moment from the outer-loop power generation command of the power system unit, and then pass it through the inner-loop fractional-order proportional-integral-derivative method to obtain the inner-loop power generation command of the power system unit.

2. The double-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for the power system according to claim 1, characterized in that, The Transformer model consists of an embedding layer, an encoder, and a decoder. The encoder consists of N l encoding units. Each encoding unit contains a multi-head attention mechanism, a feed-forward layer, a residual connection, and a layer normalization layer. After the input signal is converted into a fixed-length vector by the embedding layer and positional encoding is added, it is then passed through the encoder and output to the multi-head attention mechanism of the decoder. The decoder consists of N l decoding units. Each decoding unit contains a multi-head attention mechanism, a feed-forward layer, a residual connection, a layer normalization layer, and a masked attention mechanism. The output of the decoder passes through a fully connected layer and an exponential normalization layer to obtain the output signal of the Transformer model. The multi-head attention mechanism includes a query matrix, a key-value matrix, and a value matrix.

3. The double-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for the power system according to claim 1, wherein In the Q(λ) learning method, the state and the action The corresponding Q value is updated as follows: Wherein, is the state at time t+1, i.e., the large fluctuation signal at time t+1; is the state at time t; is the action at time t, i.e., the output control action of the Q(λ) learning method at time t; is the discount factor; is the reward obtained at time t+1, which is ; is the next state all possible actions of the maximum value of the Q-values; is the learning rate; is the state and the action the corresponding Q-value; is in the state and the action the eligibility trace factor under.

4. The double-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for the power system according to claim 1, wherein In the outer-loop fractional-order proportional-integral-derivative method, the transfer function of the outer-loop fractional-order proportional-integral-derivative method is as follows: wherein, is the outer loop proportional regulation coefficient; is the outer loop integral regulation coefficient; is the outer loop derivative regulation coefficient; is the outer loop fractional integral order; is the outer loop fractional derivative order.

5. The double-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for the power system according to claim 1, characterized in that, In the inner-loop fractional-order proportional-integral-derivative method, the transfer function of the inner-loop fractional-order proportional-integral-derivative method is as follows: Wherein, is the inner loop proportional regulation coefficient; is the inner loop integral regulation coefficient; is the inner loop differential regulation coefficient; is the inner loop fractional integral order; is the inner loop fractional differential order.

6. The double-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method for the power system according to claim 1, characterized in that The step of using the complete ensemble empirical mode decomposition with adaptive noise (CEEMDAN) method to decompose the frequency deviation sequence of the power system over a period of time, obtaining the intrinsic mode components and the residual component is: In region A, obtain the frequency deviation sequence of the power system over a period of time , , is the reference time point, is from to for 128 consecutive time points; Set the number of iterations to K, and the original signal and the initial residual are both , and construct the decomposition sequence : In the formula, is the signal-to-noise ratio of the first decomposition; is the first component after empirical mode decomposition (EMD); is the i-th added white noise; Calculate the first residual component and the first eigenmode component : In the formula, is the average value function, is the function for generating the local mean value of the signal; Calculate the second residual component and the second eigenmode component : In the formula, is the second decomposition signal-to-noise ratio; is the second component after decomposition by the empirical mode decomposition method EMD; Similarly, the k-th residual component and the k-th intrinsic mode function : wherein, is the (k - 1)th residual component; is the kth decomposition signal-to-noise ratio; is the kth component after empirical mode decomposition (EMD) by the empirical mode decomposition method; Thus, after K iterations, K intrinsic mode components and the Kth residual component are obtained.

7. A double-closed-loop decomposition prediction fractional-order reinforcement learning power generation control device for a power system, characterized in that The device includes: A frequency deviation decomposition and clustering module, which is used to use the complete ensemble empirical mode decomposition with adaptive noise (CEEMDAN) method to decompose the frequency deviation sequence of the power system over a period of time, obtaining the intrinsic mode components and the residual component. Input the intrinsic mode components into the Transformer model for prediction to obtain the predicted values of the future intrinsic mode components predicted by the Transformer model. Use k-means clustering to cluster the predicted values of the future intrinsic mode components predicted by the Transformer model. Set the number of categories of k-means clustering to 2. Denote the sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the clustering with the larger sum as the large fluctuation signal, and denote the sum of the predicted values of the future intrinsic mode components predicted by the Transformer model in the clustering with the smaller sum as the small fluctuation signal; A reinforcement learning module, which is used to use the Q(λ) learning method to perform reinforcement learning on the obtained large fluctuation signal to obtain the output control action of the Q(λ) learning method; An analytical control method module, which is used to use the outer-loop fractional-order proportional-integral-derivative method to perform following control on the obtained small fluctuation signal to obtain the output control action of the outer-loop fractional-order proportional-integral-derivative method; An outer-loop summation module, which is used to add the output control actions of the Q(λ) learning method and the output control actions of the outer-loop fractional-order proportional-integral-derivative method to obtain the outer-loop power generation command of the power system unit; An inner-loop power generation command solving module, which is used to subtract the inner-loop power generation command of the power system unit at the previous moment from the outer-loop power generation command of the power system unit, and then through the inner-loop fractional-order proportional-integral-derivative method, obtain the inner-loop power generation command of the power system unit.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the double-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method of the power system according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the double-closed-loop decomposition prediction fractional-order reinforcement learning power generation control method of the power system according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Output power prediction method, device and apparatus of photovoltaic power generation system and medium

    CN108985521A

  • Power grid frequency intelligent control method based on empirical mode decomposition

    CN112398142A

  • Ultra-short-term power prediction model establishment method considering time-space characteristics of multiple offshore wind power units

    CN115587525A

  • Intelligent power grid voltage control method based on prediction fractional order reinforcement learning

    CN115659804A

  • Intelligent control method for combining proportion and deep learning of integrated energy system

    CN116961144A