Transform and active-disturbance-rejection deep reinforcement learning-based photovoltaic maximum power point tracking method and device, computer equipment and medium
By combining the Transformer network and deep reinforcement learning with auto-disturbance rejection (ADR), the voltage deviation signal of the photovoltaic array is decomposed. A fractional-order ADR controller and the soft actor-critic method of Informer are used to solve the multi-peak problem of the existing photovoltaic maximum power point tracking method under complex conditions, and achieve fast and accurate photovoltaic maximum power point tracking.
Patent Information
- Application Number
- CN202510827083.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-23
AI Technical Summary
Existing photovoltaic maximum power point tracking methods cannot simultaneously take into account dynamic response speed and tracking accuracy, especially it is difficult to handle multi-peak problems under complex shadow conditions, and existing artificial intelligence-based methods have high computing resource and time requirements.
The Transformer network is combined with deep reinforcement learning for auto-disturbance rejection (ADR). By decomposing the voltage deviation signal of the photovoltaic array, a fractional-order ADR controller and the Informer soft actor-critic method are used to achieve fast and accurate tracking of the photovoltaic maximum power point.
It achieves fast and accurate tracking of the global maximum power point of the photovoltaic system under complex shadow conditions, reduces computing resource requirements, and improves tracking accuracy and dynamic response speed.
Smart Images

Figure CN120688576A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of new energy, photovoltaic power generation control, natural language processing and deep reinforcement learning, and relates to a control method based on Transformer and self-disturbance rejection deep reinforcement learning, which is suitable for the control of photovoltaic maximum power point tracking in power systems. Background Art
[0002] Existing photovoltaic maximum power point tracking methods struggle to balance dynamic response speed and tracking accuracy. Traditional photovoltaic maximum power point tracking control methods, such as the perturbation-observation method, are generally only suitable for photovoltaic systems under uniform lighting conditions and cannot handle the multiple peaks of photovoltaic systems under complex shadow conditions. Furthermore, these methods often suffer from oscillation after tracking the maximum power point.
[0003] Additionally, existing AI-based control methods, including fuzzy logic control, machine learning, and artificial neural networks, offer advantages in improving tracking accuracy and adapting to complex environments. However, machine learning and deep learning methods rely on high-quality and large-scale datasets for training. Control methods based on intelligent algorithms typically require significant computing resources and time to complete the optimization process. Summary of the Invention
[0004] Based on this, it is necessary to provide a photovoltaic maximum power point tracking method, device, computer equipment, computer-readable storage medium and computer program product based on Transformer and self-disturbance rejection deep reinforcement learning to address the above technical problems.
[0005] In the first aspect, the present application provides a photovoltaic maximum power point tracking method based on Transformer and self-disturbance rejection deep reinforcement learning. The method includes:
[0006] The irradiance matrix and temperature matrix of the photovoltaic array are used as the input of the Transformer network, and the maximum power point voltage of the photovoltaic array is used as the output of the Transformer network. The Transformer network is trained.
[0007] The fully adaptive noise ensemble empirical mode decomposition method is used to decompose the difference between the actual voltage of the photovoltaic array and the maximum power point voltage of the photovoltaic array predicted by the Transformer network based on irradiance and temperature conditions, and the intrinsic mode components and the residual components of the fully adaptive noise ensemble empirical mode decomposition are obtained. The obtained intrinsic mode components and the residual components of the fully adaptive noise ensemble empirical mode decomposition are partially summed as large fluctuation mode signals, and the remaining components are summed as small fluctuation mode signals.
[0008] The fractional-order active disturbance rejection controller is used to process the small fluctuation mode signal, and the output of the fractional-order active disturbance rejection controller for the small fluctuation mode signal is obtained;
[0009] The Informer-based soft actor-critic method is used to process large fluctuation modal signals, and the output of the Informer-based soft actor-critic method for large fluctuation modal signals is obtained;
[0010] The output of the fractional-order ADRC for small-fluctuation modal signals and the output of the Informer-based soft actor-critic method for large-fluctuation modal signals are added together as the output control instructions based on Transformer and ADRC deep reinforcement learning.
[0011] In a second aspect, the present application also provides a photovoltaic maximum power point tracking device based on Transformer and deep reinforcement learning for auto-disturbance rejection. The device comprises:
[0012] Acquire data and train the network module, which is used to form a matrix of the irradiance matrix and temperature matrix of the photovoltaic array as the input of the Transformer network, and the maximum power point voltage of the photovoltaic array as the output of the Transformer network, and train the Transformer network;
[0013] A voltage deviation decomposition and classification module is used to decompose the difference between the actual voltage of the photovoltaic array and the maximum power point voltage of the photovoltaic array predicted by the Transformer network based on irradiance and temperature conditions using a fully adaptive noise ensemble empirical mode decomposition method, obtain intrinsic mode components and residual components of the fully adaptive noise ensemble empirical mode decomposition, sum some components of the obtained intrinsic mode components and residual components of the fully adaptive noise ensemble empirical mode decomposition as large fluctuation mode signals, and sum the remaining components as small fluctuation mode signals;
[0014] A fractional-order active disturbance rejection control module is used to process the small fluctuation mode signal using a fractional-order active disturbance rejection controller to obtain the output of the fractional-order active disturbance rejection controller for the small fluctuation mode signal;
[0015] A deep reinforcement learning control module is used to process large-fluctuation modal signals using an informer-based soft actor-critic method, and obtain the output of the informer-based soft actor-critic method for large-fluctuation modal signals;
[0016] The summation output module is used to add the output of the fractional-order active disturbance rejection controller for small-fluctuation modal signals and the output of the informer-based soft actor-critic method for large-fluctuation modal signals as the output control instruction based on Transformer and active disturbance rejection deep reinforcement learning.
[0017] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are performed:
[0018] The irradiance matrix and temperature matrix of the photovoltaic array are used as the input of the Transformer network, and the maximum power point voltage of the photovoltaic array is used as the output of the Transformer network. The Transformer network is trained.
[0019] The fully adaptive noise ensemble empirical mode decomposition method is used to decompose the difference between the actual voltage of the photovoltaic array and the maximum power point voltage of the photovoltaic array predicted by the Transformer network based on irradiance and temperature conditions, and the intrinsic mode components and the residual components of the fully adaptive noise ensemble empirical mode decomposition are obtained. The obtained intrinsic mode components and the residual components of the fully adaptive noise ensemble empirical mode decomposition are partially summed as large fluctuation mode signals, and the remaining components are summed as small fluctuation mode signals.
[0020] The fractional-order active disturbance rejection controller is used to process the small fluctuation mode signal, and the output of the fractional-order active disturbance rejection controller for the small fluctuation mode signal is obtained;
[0021] The Informer-based soft actor-critic method is used to process large fluctuation modal signals, and the output of the Informer-based soft actor-critic method for large fluctuation modal signals is obtained;
[0022] The output of the fractional-order ADRC for small-fluctuation modal signals and the output of the Informer-based soft actor-critic method for large-fluctuation modal signals are added together as the output control instructions based on Transformer and ADRC deep reinforcement learning.
[0023] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps:
[0024] The irradiance matrix and temperature matrix of the photovoltaic array are used as the input of the Transformer network, and the maximum power point voltage of the photovoltaic array is used as the output of the Transformer network. The Transformer network is trained.
[0025] The fully adaptive noise ensemble empirical mode decomposition method is used to decompose the difference between the actual voltage of the photovoltaic array and the maximum power point voltage of the photovoltaic array predicted by the Transformer network based on irradiance and temperature conditions, and the intrinsic mode components and the residual components of the fully adaptive noise ensemble empirical mode decomposition are obtained. The obtained intrinsic mode components and the residual components of the fully adaptive noise ensemble empirical mode decomposition are partially summed as large fluctuation mode signals, and the remaining components are summed as small fluctuation mode signals.
[0026] The fractional-order active disturbance rejection controller is used to process the small fluctuation mode signal, and the output of the fractional-order active disturbance rejection controller for the small fluctuation mode signal is obtained;
[0027] The Informer-based soft actor-critic method is used to process large fluctuation modal signals, and the output of the Informer-based soft actor-critic method for large fluctuation modal signals is obtained;
[0028] The output of the fractional-order ADRC for small-fluctuation modal signals and the output of the Informer-based soft actor-critic method for large-fluctuation modal signals are added together as the output control instructions based on Transformer and ADRC deep reinforcement learning.
[0029] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the following steps:
[0030] The irradiance matrix and temperature matrix of the photovoltaic array are used as the input of the Transformer network, and the maximum power point voltage of the photovoltaic array is used as the output of the Transformer network. The Transformer network is trained.
[0031] The fully adaptive noise ensemble empirical mode decomposition method is used to decompose the difference between the actual voltage of the photovoltaic array and the maximum power point voltage of the photovoltaic array predicted by the Transformer network based on irradiance and temperature conditions, and the intrinsic mode components and the residual components of the fully adaptive noise ensemble empirical mode decomposition are obtained. The obtained intrinsic mode components and the residual components of the fully adaptive noise ensemble empirical mode decomposition are partially summed as large fluctuation mode signals, and the remaining components are summed as small fluctuation mode signals.
[0032] The fractional-order active disturbance rejection controller is used to process the small fluctuation mode signal, and the output of the fractional-order active disturbance rejection controller for the small fluctuation mode signal is obtained;
[0033] The Informer-based soft actor-critic method is used to process large fluctuation modal signals, and the output of the Informer-based soft actor-critic method for large fluctuation modal signals is obtained;
[0034] The output of the fractional-order ADRC for small-fluctuation modal signals and the output of the Informer-based soft actor-critic method for large-fluctuation modal signals are added together as the output control instructions based on Transformer and ADRC deep reinforcement learning.
[0035] The photovoltaic maximum power point tracking method, device, computer equipment, storage medium and computer program product based on Transformer and self-disturbance rejection deep reinforcement learning are as follows: the irradiance matrix and temperature matrix of the photovoltaic array are combined into a matrix as the input of the Transformer network, and the maximum power point voltage of the photovoltaic array is used as the output of the Transformer network to train the Transformer network; the fully adaptive noise set empirical mode decomposition method is used to decompose the difference between the actual voltage of the photovoltaic array and the maximum power point voltage of the photovoltaic array predicted by the Transformer network for irradiance and temperature conditions, and the intrinsic mode components and the residual components of the fully adaptive noise set empirical mode decomposition are obtained; the partial components of the obtained intrinsic mode components and the residual components of the fully adaptive noise set empirical mode decomposition are summed as large fluctuation mode signals, and the remaining components are summed as small fluctuation mode signals; the fractional order self-disturbance rejection controller is used to process the small fluctuation mode signal to obtain the output of the fractional order self-disturbance rejection controller for the small fluctuation mode signal; the soft actor critic method based on Informer is used to process the large fluctuation mode signal. Fluctuation modal signal processing is used to obtain the output of the Informer-based soft actor-critic method for large fluctuation modal signals; the output of the fractional-order active disturbance rejection controller for small fluctuation modal signals and the output of the Informer-based soft actor-critic method for large fluctuation modal signals are added together as the output control instructions based on Transformer and deep reinforcement learning for active disturbance rejection; the Transformer network and deep reinforcement learning for active disturbance rejection can be combined for photovoltaic maximum power point tracking control, which can achieve fast and stable tracking; first, the irradiance and temperature during actual operation of the system are measured, and the voltage reference value of the global maximum power point is predicted through the Transformer network. Then, the actual voltage value of the working photovoltaic array is measured and subtracted from the voltage reference value to obtain a voltage deviation sequence. The voltage deviation sequence is decomposed using the fully adaptive noise ensemble empirical mode method and classified according to large and small fluctuation signals. The large and small fluctuation signals are respectively given to the deep Q network method and the fractional-order active disturbance rejection controller for control. The output control instructions of the deep Q network method and the fractional-order active disturbance rejection controller are added to obtain the final control instruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 The diagram is a framework diagram of a photovoltaic maximum power point tracking control system in one embodiment.
[0037] Figure 2 The figure is a control flow chart of deep reinforcement learning based on Transformer and self-disturbance rejection in one embodiment.
[0038] Figure 3 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0040] In one embodiment, Figure 1 Figure 2 shows a framework diagram of a photovoltaic maximum power point tracking (MPPT) control system. A typical PV MPPT control system includes a photovoltaic array, a boost converter circuit, and a maximum power point tracking controller. The PV array consists of multiple PV modules connected in series and parallel, and the boost converter circuit consists of capacitors C1 and C2, an inductor L, a diode D, and a switch element Q.
[0041] In one embodiment, Figure 2 As shown, a control flow chart of deep reinforcement learning based on Transformer and self-disturbance rejection is provided. This embodiment uses the method applied to a terminal as an example. It is understandable that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0042] In step (1), the irradiance matrix and temperature matrix of the photovoltaic array are combined into a matrix as the input of the Transformer network, and the maximum power point voltage of the photovoltaic array is used as the output of the Transformer network to train the Transformer network.
[0043] A photovoltaic array matrix with M rows and N columns consisting of multiple individual photovoltaic modules connected in series and parallel for:
[0044] ,
[0045] in, It is the photovoltaic module in the first row and first column; It is the photovoltaic module in the first row and second column; It is the first row PV panels in series; It is the photovoltaic module in the second row and first column; It is the photovoltaic module in the second row and second column; It is the second row PV panels in series; It is PV panels in the first column of the row; It is PV panels in the second column of the row; It is Rank PV panels in series;
[0046] Get the irradiance matrix of the photovoltaic array with M rows and N columns , temperature matrix And the photovoltaic array with M rows and N columns in the irradiance matrix and temperature matrix Maximum power point voltage under ;
[0047] The irradiance matrix of the photovoltaic array with M rows and N columns is and temperature matrix Composition matrix As the input of the Transformer network, the maximum power point voltage of the photovoltaic array with M rows and N columns is As the output of the Transformer network, train the Transformer network.
[0048] The Transformer network includes an encoder consisting of an attention mechanism and a feedforward neural network. The output of the Transformer network's encoder is:
[0049] ,
[0050] in, is the output function of the encoder, i.e. the layer normalized output function;
[0051] The self-attention mechanism of the Transformer network encoder includes the query matrix , key matrix Sum Matrix :
[0052] ,
[0053] ,
[0054] ,
[0055] in, is the embedding representation of the input matrix of the Transformer network encoder, with dimension , is the sequence length, is the embedding dimension; 、 and They are respectively the trainable query weight matrix, key weight matrix and value weight matrix;
[0056] The attention score of the self-attention mechanism of the Transformer network encoder is:
[0057] .
[0058] in, is the attention score of the self-attention mechanism of the Transformer network encoder; is the transpose of the key matrix; is the number of columns of the key matrix, i.e. the vector dimension;
[0059] The attention weights of the self-attention mechanism of the Transformer network encoder are:
[0060] .
[0061] in, is the softmax activation function;
[0062] Apply the attention weights to the value matrix to get the output matrix of the self-attention mechanism of the Transformer network encoder for:
[0063] .
[0064] Step (2) uses the fully adaptive noise ensemble empirical mode decomposition method to decompose the difference between the actual voltage of the photovoltaic array and the maximum power point voltage of the photovoltaic array predicted by the Transformer network for irradiance and temperature conditions, and obtains the intrinsic mode component and the residual component of the fully adaptive noise ensemble empirical mode decomposition. The obtained intrinsic mode component and the residual component of the fully adaptive noise ensemble empirical mode decomposition are partially summed as the large fluctuation mode signal, and the remaining components are summed as the small fluctuation mode signal.
[0065] A total of t moments of voltage deviation time series The difference between the actual voltage of the photovoltaic array at time t and the maximum power point voltage of the photovoltaic array predicted by the Transformer network based on the irradiance and temperature conditions at time t:
[0066] ,
[0067] in, is the actual voltage of the photovoltaic array at time t; is the maximum power point voltage of the photovoltaic array predicted by the Transformer network based on the irradiance and temperature conditions at time t;
[0068] The voltage deviation time series with t moments is decomposed using the fully adaptive noise ensemble empirical mode decomposition method. ,Will Paired positive and negative Gaussian white noise is added to the voltage deviation time series at a total of t moments Go up and get A total of t time series of voltage deviations containing noise are formed, The voltage deviation time series of t moments with noise is:
[0069] ,
[0070] in, For the The first A total of t time series of voltage deviations containing noise; is the weight coefficient of the added white noise; For the Group-added paired Gaussian white noise;
[0071] The signal after adding white noise to each group Perform empirical mode decomposition and take The mean of the group decomposition results is obtained for:
[0072] ,
[0073] in, is the first eigenmode component obtained; For Carry out the The first Group modal components;
[0074] Thus, the first residual component for:
[0075] ,
[0076] In the first residual component Join again The paired positive and negative Gaussian white noises are iterated, and the added white noise is the auxiliary noise signal after empirical mode decomposition. Signal after group white noise for:
[0077] ,
[0078] in, is the empirical mode decomposition function;
[0079] Add the Signal after group white noise Perform empirical mode decomposition and take The mean of the group decomposition results is obtained for:
[0080] ,
[0081] in, is the second eigenmode component obtained; For Carry out the The first modal components;
[0082] Thus, the second residual component for:
[0083] ,
[0084] Repeat the iteration until the number of extreme points of the residual component curve is , that is, no more empirical mode decomposition can be performed; the maximum number of iterations is , then in progress After iterations, we get eigenmode components and The residual component after iterations is obtained eigenmode components and The residual component after iterations satisfies:
[0085] ,
[0086] in, For the eigenmode components; for The residual component after the iteration is the residual component of the fully adaptive noise set empirical mode decomposition;
[0087] The obtained intrinsic mode components and some components of the residual components of the fully adaptive noise set empirical mode decomposition are summed as large fluctuation mode signals, and the remaining components are summed as small fluctuation mode signals, that is:
[0088] Set the first threshold , the second threshold and the third threshold ;
[0089] Will The eigenmode components and one residual component are numbered from the 1st mode / residual component to the +1 modal / residual component;
[0090] The number of extreme points of the voltage deviation time series curve at t moments is ;
[0091] When the number of extreme points of the voltage deviation time series curve at a total of t moments is greater than the first threshold, that is, When the first mode / residual component is transferred to the ( +1)×0.8 mode / residual components are summed as large wave mode signals , the first ( +1)×0.8+1 modal / residual component to the +1 mode / residual components are summed as small fluctuation mode signals ;
[0092] When the number of extreme points of the voltage deviation time series curve at a total of t moments is less than or equal to the first threshold and greater than the second threshold, that is, When the first mode / residual component is transferred to the ( +1)×0.6 mode / residual components are summed as large wave mode signals , the first ( +1)×0.6+1 modal / residual component to the +1 mode / residual components are summed as small fluctuation mode signals ;
[0093] When the number of extreme points of the voltage deviation time series curve at a total of t moments is less than or equal to the second threshold and greater than the third threshold, that is, When the first mode / residual component is transferred to the ( +1)×0.4 mode / residual components are summed as large wave mode signals , the first ( +1)×0.4+1 modal / residual component to the +1 mode / residual components are summed as small fluctuation mode signals ;
[0094] When the number of extreme points of the voltage deviation time series curve at a total of t moments is less than or equal to the third threshold, that is, When the first mode / residual component is transferred to the ( +1)×0.2 mode / residual components are summed as large wave mode signals , the first ( +1)×0.2+1 modal / residual component to the +1 mode / residual components are summed as small fluctuation mode signals .
[0095] In step (3), a fractional-order active disturbance rejection controller is used to process the small fluctuation mode signal, and the output of the fractional-order active disturbance rejection controller for the small fluctuation mode signal is obtained.
[0096] The fractional-order ADRC uses fractional-order calculus to replace the integer-order integration and differentiation operations in the ADRC with fractional-order operations; the fractional-order order is The transfer function of the fractional-order active disturbance rejection controller is for:
[0097] ,
[0098] in, , , is the transfer function of the fractional-order ADRC; s is the complex variable in Laplace transform; is the proportional gain of the fractional-order active disturbance rejection controller; is the differential gain of the fractional-order active disturbance rejection controller; is the bandwidth of the set fractional-order ADRC;
[0099] After the inverse Laplace transform, the output of the fractional-order active disturbance rejection controller for the small fluctuation mode signal is obtained for:
[0100] ,
[0101] in, The fractional-order active disturbance rejection controller is used to control the small fluctuation modal signal. Output; is the inverse Laplace transform function; is a small wave mode signal The Laplace transform of .
[0102] In step (4), the large fluctuation modal signal is processed by the Informer-based soft actor critic method to obtain the output of the Informer-based soft actor critic method for the large fluctuation modal signal.
[0103] Using Informer-based soft actor-critic method to simulate large fluctuating modal signals Processing, get the output of the Informer-based soft actor critic method for large fluctuation modal signals ;
[0104] The Informer-based soft actor-critic consists of an actor network, a critic network, and a target-critic network;
[0105] The target critic network synchronizes the target critic network parameters through exponential sliding average, the actor network updates the actor network parameters through the gradient ascent algorithm, the critic network samples batch data from the experience replay buffer, and optimizes the critic network parameters through the mean square error loss function;
[0106] The target critic network parameters are synchronized by exponential sliding average as follows:
[0107] ,
[0108] in, represents the set of parameters of the critic network; represents the set of parameters of the target critic network; Represents the smoothing coefficient, the value range is , used to control the smoothness of parameter updates;
[0109] The actor network updates the actor network parameters through the gradient ascent algorithm, and the update formula is
[0110] ,
[0111] in, represents the set of parameters of the actor network; Represents actor-based network parameters Performance evaluation function of represents the gradient of the performance evaluation function with respect to the actor network parameters; Indicates that in the strategy The state distribution under Represents the state-action value function, evaluated at state Next action long-term value; Represents the actor network based on parameters In state The generated action; Represents the time step environmental conditions, For expectations;
[0112] Informer-based soft actor-critic approach in action The output of the Informer-based soft actor-critic method for large-volume modal signals; the state of the Informer-based soft actor-critic method is the observation input to the Informer-based soft actor-critic method, state Large wave mode signal Or a total of t moments of voltage deviation time series ;
[0113] The critic network samples batches of data from the experience replay buffer and optimizes the critic network parameters using the mean squared error loss function. The loss function is
[0114] ,
[0115] in, Represents the critic network parameters The loss function of Represents the experience replay buffer, which stores historical state transition samples; Indicates the current state; Indicates that the status The action to be performed next; Indicates execution of an action Immediate rewards after Indicates execution of an action The next state to transfer to; Represents the critic network based on parameters State-Action Pairs Estimated value of Represents the target critic network based on parameters For the next state and actions ; Represents the discount factor, the value range is , used to weigh the importance of immediate rewards and future rewards;
[0116] The actor network, critic network, and target critic network in the Informer-based soft actor-critic are all Informer networks. Informer contains an Informer encoder and an Informer decoder. Both the Informer encoder and the Informer decoder contain a multi-head probabilistic attention layer and a feedforward neural network layer.
[0117] In step (5), the output of the fractional-order ADRC for small-fluctuation modal signals and the output of the Informer-based soft actor-critic method for large-fluctuation modal signals are added together as the output control instructions based on Transformer and ADRC deep reinforcement learning.
[0118] The output of the fractional-order active disturbance rejection controller for small fluctuation mode signals and the output of the Informer-based soft actor-critic method for large-volume modal signals After addition, it is used as the output control instruction based on Transformer and self-disturbance rejection deep reinforcement learning , which is input to the switching element Q of the Boost converter in the photovoltaic array to control the photovoltaic array to operate at the maximum power point in real time.
[0119] The output control instructions based on Transformer and ADRC deep reinforcement learning are:
[0120] .
[0121] The present invention has the following advantages and effects compared to the prior art:
[0122] (1) Existing photovoltaic maximum power point tracking control methods based on disturbance self-optimization are only applicable to photovoltaic systems under uniform lighting conditions and cannot handle the multi-peak problem of photovoltaic systems under complex shadow conditions. However, the present invention combines the efficient feature extraction capability of the Transformer network with the dynamic optimization characteristics of self-disturbance rejection deep reinforcement learning to effectively cope with the complex PV characteristic curves of photovoltaic systems under local shadow conditions and quickly and accurately track the global maximum power point.
[0123] (2) The existing process of using neural networks to predict the maximum power point of photovoltaics is relatively complex and has limitations in regional positioning when facing multi-peak power curves. However, the present invention uses a Transformer network to predict the reference value of the maximum power point of photovoltaics. By constructing a multi-dimensional feature space mapping relationship through a self-attention mechanism, it can accurately predict the location of the reference value of the maximum power point of photovoltaics.
[0124] (3) The present invention uses a Transformer network to directly predict the reference voltage of the maximum power point, which is completely different from the existing photovoltaic maximum power point tracking method, and the input of the Transformer network is irradiance and temperature.
[0125] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 3As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be achieved via Wi-Fi, a mobile cellular network, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a photovoltaic maximum power point tracking method based on Transformer and deep reinforcement learning for auto-disturbance rejection. The display unit of the computer device is used to produce a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0126] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0127] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0128] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0129] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0130] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0131] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0132] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0133] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A photovoltaic maximum power point tracking method based on Transformer and auto-disturbance rejection deep reinforcement learning, characterized in that: The method comprises: The irradiance matrix and temperature matrix of the photovoltaic array are used as the input of the Transformer network, and the maximum power point voltage of the photovoltaic array is used as the output of the Transformer network. The Transformer network is trained. The fully adaptive noise ensemble empirical mode decomposition method is used to decompose the difference between the actual voltage of the photovoltaic array and the maximum power point voltage of the photovoltaic array predicted by the Transformer network based on irradiance and temperature conditions, and the intrinsic mode components and the residual components of the fully adaptive noise ensemble empirical mode decomposition are obtained. The obtained intrinsic mode components and the residual components of the fully adaptive noise ensemble empirical mode decomposition are partially summed as large fluctuation mode signals, and the remaining components are summed as small fluctuation mode signals. The fractional-order active disturbance rejection controller is used to process the small fluctuation mode signal, and the output of the fractional-order active disturbance rejection controller for the small fluctuation mode signal is obtained; The Informer-based soft actor-critic method is used to process large fluctuation modal signals, and the output of the Informer-based soft actor-critic method for large fluctuation modal signals is obtained; The output of the fractional-order ADRC for small-fluctuation modal signals and the output of the Informer-based soft actor-critic method for large-fluctuation modal signals are added together as the output control instructions based on Transformer and ADRC deep reinforcement learning.
2. The photovoltaic maximum power point tracking method based on Transformer and auto-disturbance rejection deep reinforcement learning according to claim 1 is characterized in that: The Transformer network includes an encoder consisting of an attention mechanism and a feedforward neural network. The output of the Transformer network encoder is: , in, is the output function of the encoder, i.e. the layer normalized output function; The output matrix of the self-attention mechanism of the Transformer network encoder is: ; in, is the query matrix, is the value matrix, is the bond matrix The transpose of is the number of columns of the key matrix, i.e. the vector dimension; is the softmax activation function.
3. The photovoltaic maximum power point tracking method based on Transformer and auto-disturbance rejection deep reinforcement learning according to claim 1 is characterized in that: The specific steps of summing the obtained intrinsic mode components and some components of the residual components of the fully adaptive noise set empirical mode decomposition as the large fluctuation mode signal and summing the remaining components as the small fluctuation mode signal are as follows: Will The eigenmode components and one residual component are numbered from the 1st mode / residual component to the +1 modal / residual component; A total of t moments of voltage deviation time series The difference between the actual voltage of the photovoltaic array at time t and the maximum power point voltage of the photovoltaic array predicted by the Transformer network based on the irradiance and temperature conditions at time t; The number of extreme points of the voltage deviation time series curve at t moments is ; When the number of extreme points of the voltage deviation time series curve at a total of t moments is greater than the first threshold When When the first mode / residual component is transferred to the ( +1)×0.8 mode / residual components are summed as large wave mode signals , the first ( +1)×0.8+1 modal / residual component to the +1 mode / residual components are summed as small fluctuation mode signals ; When the number of extreme points of the voltage deviation time series curve at a total of t moments is less than or equal to the first threshold and greater than the second threshold When When the first mode / residual component is transferred to the ( +1)×0.6 mode / residual components are summed as large wave mode signals , the first ( +1)×0.6+1 modal / residual component to the +1 mode / residual components are summed as small fluctuation mode signals ; When the number of extreme points of the voltage deviation time series curve at a total of t moments is less than or equal to the second threshold and greater than the third threshold When When the first mode / residual component is transferred to the ( +1)×0.4 mode / residual components are summed as large wave mode signals , the first ( +1)×0.4+1 modal / residual component to the +1 mode / residual components are summed as small fluctuation mode signals ; When the number of extreme points of the voltage deviation time series curve at a total of t moments is less than or equal to the third threshold When When the first mode / residual component is transferred to the ( +1)×0.2 mode / residual components are summed as large wave mode signals , the first ( +1)×0.2+1 modal / residual component to the +1 mode / residual components are summed as small fluctuation mode signals .
4. The photovoltaic maximum power point tracking method based on Transformer and auto-disturbance rejection deep reinforcement learning according to claim 1 is characterized in that: The transfer function of the fractional-order active disturbance rejection controller is for: , in, , ; s is the complex variable in Laplace transform; is the proportional gain of the fractional-order active disturbance rejection controller; is the differential gain of the fractional-order active disturbance rejection controller; is the bandwidth of the set fractional-order ADRC; is a small wave mode signal Laplace transform of is the fractional order of the fractional-order ADRC.
5. The photovoltaic maximum power point tracking method based on Transformer and auto-disturbance rejection deep reinforcement learning according to claim 1 is characterized in that: The Informer-based soft actor-critic consists of an actor network, a critic network, and a target-critic network; The target critic network synchronizes the target critic network parameters through exponential sliding average, the actor network updates the actor network parameters through the gradient ascent algorithm, the critic network samples batch data from the experience replay buffer, and optimizes the critic network parameters through the mean squared error loss function.
6. The photovoltaic maximum power point tracking method based on Transformer and auto-disturbance rejection deep reinforcement learning according to claim 1 is characterized in that: The Informer-based soft actor-critic approach in action The output of the Informer-based soft actor-critic method for large-volume modal signals; the state of the Informer-based soft actor-critic method Large wave mode signal Or a total of t moments of voltage deviation time series .
7. The photovoltaic maximum power point tracking method based on Transformer and auto-disturbance rejection deep reinforcement learning according to claim 1 is characterized in that: The actor network, critic network and target critic network in the Informer-based soft actor-critic are all Informer networks, which include an Informer encoder and an Informer decoder. Both the Informer encoder and the Informer decoder include a multi-head probabilistic attention layer and a feedforward neural network layer.
8. A photovoltaic maximum power point tracking device based on Transformer and auto-disturbance rejection deep reinforcement learning, characterized in that: The device comprises: Acquire data and train the network module, which is used to form a matrix of the irradiance matrix and temperature matrix of the photovoltaic array as the input of the Transformer network, and the maximum power point voltage of the photovoltaic array as the output of the Transformer network, and train the Transformer network; A voltage deviation decomposition and classification module is used to decompose the difference between the actual voltage of the photovoltaic array and the maximum power point voltage of the photovoltaic array predicted by the Transformer network based on irradiance and temperature conditions using a fully adaptive noise ensemble empirical mode decomposition method, obtain intrinsic mode components and residual components of the fully adaptive noise ensemble empirical mode decomposition, sum some components of the obtained intrinsic mode components and residual components of the fully adaptive noise ensemble empirical mode decomposition as large fluctuation mode signals, and sum the remaining components as small fluctuation mode signals; A fractional-order active disturbance rejection control module is used to process the small fluctuation mode signal using a fractional-order active disturbance rejection controller to obtain the output of the fractional-order active disturbance rejection controller for the small fluctuation mode signal; A deep reinforcement learning control module is used to process large-fluctuation modal signals using an informer-based soft actor-critic method, and obtain the output of the informer-based soft actor-critic method for large-fluctuation modal signals; The summation output module is used to add the output of the fractional-order active disturbance rejection controller for small-fluctuation modal signals and the output of the informer-based soft actor-critic method for large-fluctuation modal signals as the output control instruction based on Transformer and active disturbance rejection deep reinforcement learning.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the photovoltaic maximum power point tracking method based on Transformer and self-disturbance rejection deep reinforcement learning are implemented in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the photovoltaic maximum power point tracking method based on Transformer and self-disturbance rejection deep reinforcement learning are implemented as described in any one of claims 1 to 7.