Intelligent photovoltaic array reconstruction system based on Pareto optimization and method thereof
By optimizing the photovoltaic array reconstruction through the Pareto-optimized Bi-PPO network, the dynamic planning problem of the photovoltaic array under moving cloud occlusion is solved, the power output is maximized and the number of switching times is minimized, and the efficiency of the photovoltaic system and the equipment durability are improved.
Patent Information
- Application Number
- CN202510665961.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-12
AI Technical Summary
Existing photovoltaic array reconstruction methods are difficult to effectively optimize power output and reduce wear of switching devices when facing moving cloud obstruction, and heuristic algorithms are difficult to solve dynamic programming problems.
The Bi-PPO network based on Pareto optimization is used to optimize the Markov decision process. The parameters are updated alternately through the Pareto optimization algorithm. Combined with the real-time monitoring of the shading status of the photovoltaic array, the array configuration is dynamically adjusted to maximize power output and minimize the number of switching times.
It improves the power generation efficiency of the photovoltaic array, reduces the wear of the switch equipment, extends the equipment life, reduces the operation and maintenance costs, and realizes the intelligence and economy of the photovoltaic system.
Smart Images

Figure CN120638463A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of photovoltaic system optimization, and in particular to an intelligent photovoltaic array reconstruction system and method based on Pareto optimization. Background Art
[0002] Photovoltaic (PV) arrays often experience partial shading during daily operation, which is mainly caused by varying degrees of moving cloud cover. This situation can lead to the occurrence of hotspot effects and reduce the output power due to partial shading of the PV array. A popular method to mitigate the negative effects of partial sun shading conditions (PSC) is the reconfiguration of the PV array. The negative effects of PSC can be effectively mitigated by using PV array reconfiguration strategies. PV array reconfiguration optimizes system performance by dynamically adjusting the connections between PV cells. This approach allows the system to maintain high efficiency under different environmental conditions and reduce losses caused by shading.
[0003] The inventors have discovered that existing technologies have at least the following shortcomings: First, in real-world environments, certain shadows can change over time, primarily due to moving clouds. For example, when a moving cloud passes over a portion of a PV module or submodule, the total output power of the PV array can drop dramatically, resulting in rapid and significant changes. Currently, some research is utilizing heuristic algorithms to address the problem of reconfiguring PV arrays under the influence of moving clouds. However, the problem of moving clouds involves dynamic programming, requiring maximization of benefits over the entire power supply control (PSC) cycle, rather than the current optimal PSC solution. Limited by the characteristics of heuristic algorithms, they struggle to effectively solve such problems. In contrast, deep reinforcement learning (DRL) can solve dynamic programming problems in complex environments through interactive learning with the environment, optimizing long-term decision-making. Currently, reinforcement learning has been successfully applied to many dynamic programming problems, achieving full-cycle benefit maximization. Furthermore, PV array reconfiguration emphasizes maximizing power output while ignoring the control complexity and lifespan of switching devices. Focusing solely on the maximum output power of the PV array reconfiguration can lead to increased control complexity and shortened switching device lifespan. Summary of the Invention
[0004] The present invention aims to provide a Pareto-optimized intelligent photovoltaic array reconstruction method, comprising the following steps:
[0005] 1) Model the smart photovoltaic array reconfiguration problem as a Markov decision process;
[0006] 2) Using the Bi-PPO network to optimize the Markov decision process, the optimal intelligent photovoltaic array reconstruction scheme is obtained;
[0007] During the optimization process, the parameters Φ of the Bi-PPO network are updated alternately using the Pareto optimization algorithm p , θ v , ω i .
[0008] Furthermore, in step 1), the Markov decision process is defined as {S, A, R, T, γ}; S is the state; A is the action; R is the reward; T is the transition probability; γ is the discount factor.
[0009] Furthermore, the reward function ω i is the weight;
[0010] where the objective function r i is as follows:
[0011]
[0012] where ||F s | and |F s' | are the maximum irradiance differences of the pv groups before and after reconstruction; i = 1, 2; N step is the iteration step.
[0013] Furthermore, the Bi-PPO network includes an actor network and a critic network.
[0014] Furthermore, in step 2), the steps of optimizing the Markov decision process using the Bi-PPO network include:
[0015] 2.1) The actor network parameterizes the action A as Φp; the parameter Φp is used to generate the mean and standard deviation for the photovoltaic array;
[0016] 2.2) The actor network samples the action A through the stochastic policy π(a|s);
[0017] 2.3) The critic network calculates the reward of the action output by the actor network;
[0018] 2.4) Repeat steps 2.1) - step 2.3) to obtain the optimal action Φp with the maximum reward as the intelligent photovoltaic array reconstruction scheme.
[0019] Furthermore, the stochastic policy π(a|s) is updated through the policy gradient algorithm;
[0020] The objective function L PPO (Φ p ) for updating the stochastic policy π(a|s) is as follows:
[0021]
[0022] Where i∈{1,2}; ε is the parameter that limits the probability difference; 0<ε<1; ρ(Φ p ) represents the normal policy gradient; clip(ρ(Φ p ), 1-ε, 1+ε) is the shearing operation; E is the identifier for finding the overall expectation; This is the policy before the update.
[0023] Furthermore, after the participant network samples the action A through the random strategy π(a|s), the critic network calculates the reward of the participant network output action, and uses the Pareto optimization algorithm to optimize the parameters ω of the Bi-PPO network. i The steps for performing an alternating update include:
[0024] S1) Define the loss function H i (η), that is:
[0025]
[0026] Where η represents H i Parameters of (η); η = Φ p ,θ v ; represents the strategy loss and value loss; θ v Parameters representing the value network;
[0027] S2) Construct the function H(η), that is:
[0028]
[0029] S3) Construct the objective function for obtaining the optimal solution of ωi, namely:
[0030]
[0031] Where, Indicates H i The gradient of (η);
[0032] S4) Φ updated based on the policy network and value network p ,θ v , reconstruct the objective function and get:
[0033]
[0034] Where c i is a predefined boundary constraint, c i is a constant between 0 and 1; is the reconstruction weight;
[0035] S5) Let G be the gradient The superposition matrix of e is a vector whose elements are all 1, and c is c i The concatenated vectors of , we get:
[0036]
[0037] Where, is the weight vector;
[0038] S6) Construct the Lagrangian, that is:
[0039]
[0040] Where λ is the Lagrange multiplier;
[0041] S7) Based on the Lagrangian, solve the objective function (9) and get ω i Optimal solution;
[0042] Among them, the solution condition of objective function (9) is expressed as follows:
[0043]
[0044] A system based on the intelligent photovoltaic array reconstruction method, a Markov decision process modeling unit, an optimization unit, and an intelligent photovoltaic array reconstruction solution output unit;
[0045] The Markov decision process modeling unit models the smart photovoltaic array reconfiguration problem as a Markov decision process;
[0046] The optimization unit optimizes the Markov decision process using a Bi-PPO network to obtain an optimal intelligent photovoltaic array reconstruction solution;
[0047] The intelligent photovoltaic array reconstruction scheme output unit is used to output and display the optimal intelligent photovoltaic array reconstruction scheme.
[0048] The technical effect of the present invention is unquestionable. The present invention proposes a BiPPO (Bi-objective Pareto-based Photovoltaic reconfiguration Optimization) algorithm based on Pareto optimization. During the optimization process, the algorithm simultaneously considers the maximization of power output and the minimization of the number of switching times. Through the Pareto frontier search strategy, the weight parameter ωi is adaptively adjusted to avoid the inconsistency caused by manually setting the weight value, thereby improving the optimization accuracy and stability. In addition, the present invention adopts an intelligent optimization control strategy, combined with real-time monitoring of the shading status of the photovoltaic array, to dynamically adjust the array configuration to achieve optimal power output with the minimum number of switching times. This strategy not only improves the power generation efficiency of the photovoltaic array, but also effectively reduces the wear of the switching equipment, extends the service life of the photovoltaic system, and reduces the operation and maintenance costs. Compared with the traditional photovoltaic array reconstruction method, the present invention can more efficiently balance power output and equipment durability, making the photovoltaic system more intelligent and economical.
[0049] In summary, this paper proposes a photovoltaic array optimization method that balances maximum power output with minimization of switching times, overcoming the limitations of traditional methods. This technology can effectively improve the overall efficiency of photovoltaic systems, reduce maintenance costs, and extend equipment life, providing an innovative solution for the intelligent and efficient operation of photovoltaic power generation systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 Schematic diagram of the structure of the Bi-PPO method proposed in the present invention. DETAILED DESCRIPTION
[0051] The present invention will be further described below with reference to the following examples, but it should not be understood that the scope of the present invention is limited to the following examples. Without departing from the above technical ideas of the present invention, various substitutions and modifications can be made according to common technical knowledge and customary means in the art, and all should be included in the scope of protection of the present invention.
[0052] Example 1:
[0053] See also Figure 1 , a Pareto optimization-based intelligent photovoltaic array reconfiguration method, comprising the following steps:
[0054] 1) Model the smart photovoltaic array reconfiguration problem as a Markov decision process;
[0055] 2) Using the Bi-PPO network to optimize the Markov decision process, the optimal intelligent photovoltaic array reconstruction scheme is obtained;
[0056] During the optimization process, the Pareto optimization algorithm is used to alternately update the parameters Φ p , θ v , ω i of the Bi-PPO network.
[0057] In step 1), the Markov decision process is defined as {S, A, R, T, γ}; S is the state; A is the action; R is the reward; T is the transition probability; γ is the discount factor.
[0058] Reward function ω i is the weight;
[0059] Among them, the objective function r i is as follows:
[0060]
[0061] Among them, ||F s | and |F s' | are the maximum irradiance differences of the pv groups before and after reconstruction; i = 1, 2; N step is the iteration step.
[0062] The Bi-PPO network includes an actor network and a critic network.
[0063] In step 2), the steps of using the Bi-PPO network to optimize the Markov decision process include:
[0064] 2.1) The actor network parameterizes the action A as Φp; the parameter Φp is used to generate the mean and standard deviation for the photovoltaic array;
[0065] 2.2) The actor network samples the action A through the stochastic policy π(a|s);
[0066] 2.3) The critic network calculates the reward of the action output by the actor network;
[0067] 2.4) Repeat steps 2.1) - step 2.3) to obtain the optimal action Φp with the maximum reward as the intelligent photovoltaic array reconstruction scheme.
[0068] The stochastic policy π(a|s) is updated through the policy gradient algorithm;
[0069] The objective function L PPO (Φ p ) for updating the stochastic policy π(a|s) is as follows:
[0070]
[0071]
[0072] Where i∈{1,2}; ε is the parameter that limits the probability difference; 0<ε<1; ρ(Φ p ) represents the normal policy gradient; clip(ρ(Φ p ), 1-ε, 1+ε) is the shearing operation; E is the identifier for finding the overall expectation; This is the policy before the update.
[0073] After the participant network samples the action A through the random strategy π(a|s), the critic network calculates the reward of the participant network output action and uses the Pareto optimization algorithm to optimize the parameters ω of the Bi-PPO network. i The steps for performing an alternating update include:
[0074] S1) Define the loss function H i (η), that is:
[0075]
[0076] Where η represents H i Parameters of (η); η = Φ p ,θ v ; represents the policy loss (standard loss function of PPO) and value loss (mean square error); θ v Parameters representing the value network;
[0077] S2) Construct the function H(η), that is:
[0078]
[0079] S3) Construct the objective function for obtaining the optimal solution of ωi, namely:
[0080]
[0081] Where, Indicates H i The gradient of (η);
[0082] S4) Φ updated based on the policy network and value network p ,θ v , reconstruct the objective function and get:
[0083]
[0084] Where c i is a predefined boundary constraint, c i is a constant between 0 and 1; is the reconstruction weight;
[0085] S5) Let G be the gradient The superposition matrix of e is a vector whose elements are all 1, and c is c i The concatenated vectors of , we get:
[0086]
[0087] Where, is the weight vector;
[0088] S6) Construct the Lagrangian, that is:
[0089]
[0090] Where λ is the Lagrange multiplier;
[0091] S7) Based on the Lagrangian, solve the objective function (9) and get ω i Optimal solution;
[0092] Among them, the solution condition of objective function (9) is expressed as follows:
[0093]
[0094] Example 2:
[0095] A Pareto optimization-based intelligent photovoltaic array reconfiguration method includes the following steps:
[0096] 1) Model the smart photovoltaic array reconfiguration problem as a Markov decision process;
[0097] 2) Using the Bi-PPO network to optimize the Markov decision process, the optimal intelligent photovoltaic array reconstruction scheme is obtained;
[0098] During the optimization process, the Pareto optimization algorithm is used to alternately update the parameters αΦ, αθ, and ωi of the Bi-PPO network.
[0099] Example 3:
[0100] A smart photovoltaic array reconstruction method based on Pareto optimization, the technical content is the same as that of Example 2, further, in step 1), the Markov decision process is defined as {S, A, R, T, γ}; S is the state; A is the action; R is the reward; T is the transition probability; γ is the discount factor.
[0101] Example 4:
[0102] A Pareto-optimized intelligent photovoltaic array reconstruction method, the technical content of which is the same as any one of Embodiments 2-3, further, the reward function ω i is the weight;
[0103] r i are the objective functions with different objectives. In this paper, the goal of R1 is to reduce the irradiance difference between the two PV groups with the largest and smallest irradiance after reconstruction, and R2 is the number of changes in switching devices, which is expressed as n steps.
[0104] Objective function r i As shown below:
[0105]
[0106] Among them, |Fs| and |Fs'| are the maximum irradiance differences of the PV group before and after reconstruction.
[0107] Example 5:
[0108] A smart photovoltaic array reconstruction method based on Pareto optimization, the technical content of which is the same as any one of Examples 2-4, further, the Bi-PPO network includes a participant network and a critic network.
[0109] Example 6:
[0110] A smart photovoltaic array reconstruction method based on Pareto optimization, the technical content of which is the same as any one of Examples 2-5, further, in step 2), the step of optimizing the Markov decision process using a Bi-PPO network includes:
[0111] 2.1) The actor network parameterizes action A as Φp; the parameter Φp is used to generate the mean and standard deviation for the photovoltaic array;
[0112] 2.2) The participant network samples action A using a random strategy π(a|s);
[0113] 2.3) The critic network calculates the reward for the participant network’s output action;
[0114] 2.4) Repeat steps 2.1) to 2.3) to obtain the optimal action Φp with the maximum reward, which is used as the smart photovoltaic array reconstruction solution.
[0115] Example 7:
[0116] A Pareto-optimized intelligent photovoltaic array reconfiguration method, having the same technical content as any one of Examples 2-6, wherein the random strategy π(a|s) is updated using a policy gradient algorithm;
[0117] The objective function L updated by the random policy π(a|s) PPO (Φ p ) is as follows:
[0118]
[0119] Where i∈{1,2}; ε is a parameter that limits the probability difference; 0<ε<1; ρ(Φp) represents the normal policy gradient; clip(ρ(Φp), 1-ε, 1+ε) is the clipping operation; E is the identifier for finding the overall expectation;
[0120] Example 8:
[0121] A smart photovoltaic array reconfiguration method based on Pareto optimization, with the same technical content as any one of Examples 2-7, further comprising the step of alternately updating the parameters ωi of the Bi-PPO network using the Pareto optimization algorithm after the participant network samples the action A using the random strategy π(a|s) and before the critic network calculates the reward of the participant network output action:
[0122] S1) Define the loss function H i (η), that is:
[0123]
[0124] Where η represents H i Parameters of (η); η = Φ p ,θ v ; represents the strategy loss and value loss; θ v Parameters representing the value network;
[0125] S2) Construct the function H(η), that is:
[0126]
[0127] S3) Construct the objective function for obtaining the optimal solution of ωi, namely:
[0128]
[0129] Where, Indicates H i The gradient of (η);
[0130] S4) Based on the updated αΦ and αθ of the policy network and the value network, reconstruct the objective function and obtain:
[0131]
[0132] Where c i is a predefined boundary constraint, c i is a constant between 0 and 1; is the reconstruction weight;
[0133] S5) Let G be the gradient The superposition matrix of e is a vector whose elements are all 1, and c is ci The concatenated vectors of , we get:
[0134]
[0135] Where, is the weight vector;
[0136] S6) Construct the Lagrangian, that is:
[0137]
[0138] Where λ is the Lagrange multiplier;
[0139] S7) Based on the Lagrangian, solve the objective function (9) and get ω i Optimal solution;
[0140] Among them, the solution condition of objective function (9) is expressed as follows:
[0141]
[0142] Example 9:
[0143] A system based on the intelligent photovoltaic array reconstruction method according to any one of embodiments 1-8, comprising a Markov decision process modeling unit, an optimization unit, and an intelligent photovoltaic array reconstruction solution output unit;
[0144] The Markov decision process modeling unit models the smart photovoltaic array reconfiguration problem as a Markov decision process;
[0145] The optimization unit optimizes the Markov decision process using a Bi-PPO network to obtain an optimal intelligent photovoltaic array reconstruction solution;
[0146] The intelligent photovoltaic array reconstruction scheme output unit is used to output and display the optimal intelligent photovoltaic array reconstruction scheme.
[0147] Example 10:
[0148] A smart photovoltaic array reconfiguration method based on Pareto optimization is as follows:
[0149] The strategy determines the action a based on the observed state s. The main goal of the method is to obtain policyπ by maximizing the cumulative reward R. When the irradiance difference between the PV groups is kept to a minimum, the action a of the current trial round is the same as the previous round, or the reward R of the reconfiguration exceeds -1. It is considered to be the best reconfiguration scheme found under the current scheme preparation criterion. When we are in the mobile cloud state, we need to obtain the PSC situation at all times and reconfigure it to obtain the reward R. Subsequently, all instances of R are summed up, and our goal is to find the minimum value of R for the entire mobile cloud. As Figure 1 FIG. 1 is a schematic diagram of the structure of the Bi-PPO method proposed in the present invention, which includes the following steps:
[0150] Construct state s, action a, reward function R, the reward function R is:
[0151] Construct the PPO algorithm. PPO is an efficient policy gradient algorithm that is widely recognized for its effectiveness in handling control problems. Its actor-critic architecture can effectively handle complex, high-dimensional continuous state and action spaces. To model the action a, the actor network generates a series of Gaussian distributions, parameterized as Φp, for PV array reconstruction. These distributions generate corresponding means and standard deviations for the entire PV array, and use a random policy π(a|s) to sample the optimal action a. The random policy π(a|s) is updated by the policy gradient algorithm to maximize its clipping agent objective under the constraints of the policy update:
[0152]
[0153] Where i∈{1,2},ε(0<ε<1) is the core parameter used to limit the probability difference; ρ(Φp) represents the normal policy gradient, and clip(ρ(Φp), 1-ε, 1+ε) trims the policy gradient by limiting the probability ratio ρ(Φp) to the range [1-ε, 1+ε]. If ρ(Φp) deviates from the range [1-ε, 1+ε], the advantage function is clipped. To ensure the stability of the algorithm, this method solves the instability in the vanilla policy gradient algorithm. It does this by creating a probability ratio between the new and old policies and then limiting it to a stable range. In this way, the PPO policy update is kept within the trust region.
[0154] To achieve dual-objective optimization, we use the Pareto optimization algorithm, which utilizes scalarization techniques to optimize dual objectives. The dual-objective optimization algorithm consists of two steps: first, updating the model parameters αΦ and αθ, and then updating ωi. These two steps alternate and influence each other to achieve the best results. In this algorithm, Hi(η) is defined as consisting of the policy loss and value loss of the dual objective:
[0155]
[0156] Here, η represents the parameter of Hi(η). Here we use H(η) to combine Hi(η) and ωi:
[0157]
[0158] Considering the Karush-Kuhn-Tucker (KKT) conditions of the model parameters, the optimal solution of ωi is obtained by minimizing H(η):
[0159]
[0160] in represents the gradient of Hi(η). At this point, the parameter updates αΦ and αθ of the policy network and value network have been completed. Let ωi be ωi-ci, and the equation becomes:
[0161]
[0162] Where ci is a predefined boundary constraint, ci is a constant between 0 and 1, that is, ωi≥ci. Let G be the gradient The superposition matrix, e is a vector whose elements are all 1, and c is the concatenated vector of ci. Therefore, it can be written as:
[0163]
[0164] Applying the Lagrange multiplier, we obtain the Lagrangian.
[0165]
[0166] So the solution can be expressed as:
[0167]
[0168] At this time, the optimal weight ωi can be obtained.
Claims
1. A smart photovoltaic array reconstruction method based on Pareto optimization, characterized in that: The following steps are involved: 1) Model the smart photovoltaic array reconfiguration problem as a Markov decision process; 2) Using the Bi-PPO network to optimize the Markov decision process, the optimal intelligent photovoltaic array reconstruction scheme is obtained; During the optimization process, the Pareto optimization algorithm is used to optimize the parameters Φ of the Bi-PPO network. p ,θ v 、ω i Update alternately.
2. The Pareto optimization-based intelligent photovoltaic array reconstruction method according to claim 1, characterized in that: In step 1), the Markov decision process is defined as {S, A, R, T, γ}; S is the state; A is the action; R is the reward; T is the transition probability; γ is the discount factor.
3. The Pareto optimization-based intelligent photovoltaic array reconstruction method according to claim 2, characterized in that: Reward Function ω i is the weight; Among them, the objective function r i As shown below: where, ||F s | and |F s' | is the maximum irradiance difference of the pv group before and after reconstruction; i = 1, 2; N step is the iteration step size.
4. The Pareto optimization-based intelligent photovoltaic array reconstruction method according to claim 1, characterized in that: The Bi-PPO network includes a participant network and a critic network.
5. The Pareto optimization-based intelligent photovoltaic array reconstruction method according to claim 4, characterized in that: In step 2), the steps of optimizing the Markov decision process using the Bi-PPO network include: 2.1) The participant network parameterizes action A as Φ p ; parameter Φ p Used to generate mean and standard deviation for photovoltaic arrays; 2.2) The participant network samples action A using a random strategy π(a|s); 2.3) The critic network calculates the reward for the participant network’s output action; 2.4) Repeat steps 2.1)-2.3) to get the optimal action Φ with the maximum reward p , as a smart photovoltaic array reconstruction solution.
6. The Pareto optimization-based intelligent photovoltaic array reconstruction method according to claim 5, characterized in that: The random policy π(a|s) is updated using the policy gradient algorithm; The objective function L updated by the random policy π(a|s) PPO (Φ p ) is as follows: Where i∈{1,2}; ε is the parameter that limits the probability difference; 0<ε<1; ρ(Φ p ) represents the normal policy gradient; clip(ρ(Φ p ), 1-ε, 1+ε) is the shearing operation; E is the identifier for finding the overall expectation; This is the policy before the update.
7. The Pareto optimization-based intelligent photovoltaic array reconstruction method according to claim 5, characterized in that: After the participant network samples the action A through the random strategy π(a|s), the critic network calculates the reward of the participant network output action and uses the Pareto optimization algorithm to optimize the parameters ω of the Bi-PPO network. i The steps for performing an alternating update include: S1) Define the loss function H i (η), that is: Where η represents H i Parameters of (η); η = Φ p ,θ v ; represents the strategy loss and value loss; θ v Parameters representing the value network; S2) Construct the function H(η), that is: S3) Construct the objective function for obtaining the optimal solution of ωi, namely: Where, η H i (η) represents H i The gradient of (η); S4) Φ updated based on the policy network and value network p ,θ v , reconstruct the objective function and get: Where c i is a predefined boundary constraint, c i is a constant between 0 and 1; is the reconstruction weight; S5) Let G be the gradient ▽ η H i (η), e is a vector whose elements are all 1, c is the superposition matrix of c i The concatenated vectors of , we get: Where, is the weight vector; S6) Construct the Lagrangian, that is: Where λ is the Lagrange multiplier; S7) Based on the Lagrangian, solve the objective function (9) and get ω i Optimal solution; Among them, the solution condition of objective function (9) is expressed as follows:
8. A system based on the intelligent photovoltaic array reconstruction method according to any one of claims 1 to 7, characterized in that: Markov decision process modeling unit, optimization unit, and intelligent photovoltaic array reconstruction plan output unit; The Markov decision process modeling unit models the smart photovoltaic array reconfiguration problem as a Markov decision process; The optimization unit optimizes the Markov decision process using a Bi-PPO network to obtain an optimal intelligent photovoltaic array reconstruction solution; The intelligent photovoltaic array reconstruction scheme output unit is used to output and display the optimal intelligent photovoltaic array reconstruction scheme.