Hybrid RL-SSA-based photovoltaic grid-connected inverter controller parameter identification method

By using a hybrid RL-SSA algorithm combined with population diversity and iterative strategy optimization, the complexity and adaptability problems of photovoltaic inverter control parameter identification are solved, and efficient parameter identification and control optimization are achieved.

CN120638461APending Publication Date: 2025-09-12NINGXIA ELECTRIC POWER ENERGY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510640729.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing photovoltaic inverter control parameter identification methods have defects in complexity and adaptability, slow convergence process, and lack of clear regularity guidance, resulting in low parameter identification efficiency.

Method used

A parameter identification method for photovoltaic grid-connected inverter controller is proposed by hybrid reinforcement learning (RL) and sparrow search algorithm (SSA). Through continuous interaction between the intelligent agent and the environment and strategy iterative optimization, combined with population diversity, number of iterations and fitness distribution, the optimal state space of the RL-SSA algorithm is established, and a greedy strategy and reward mechanism are designed to optimize the control parameters.

Benefits of technology

The identification effect of photovoltaic grid-connected inverter controller parameters is significantly improved, local optimal solutions are avoided, stable search performance in high-dimensional space is maintained, and the adaptability and efficiency of the algorithm are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120638461A_ABST
    Figure CN120638461A_ABST
Patent Text Reader

Abstract

The invention discloses a hybrid RL-SSA-based photovoltaic grid-connected inverter controller parameter identification method, and the method comprises the following steps: S1, building a three-phase grid-connected power generation system simulation model, and recording the DC side voltage udc of an inverter, and the current inner loop control currents id and iq; s2, determining to-be-identified parameters kPu, kIu, kPi and kIi of the photovoltaic grid-connected inverter controller in combination with the control strategy; s3, using the to-be-identified parameters as a population by adopting an SSA algorithm, and performing division; s4, establishing an optimization state space of an RL-SSA algorithm by adopting a reinforcement learning thought according to the population characteristics; s5, designing an action space of the RL-SSA according to a greedy strategy; and S6, according to the current population diversity, the number of iterations and the fitness distribution condition, improving a reinforcement learning reward mechanism and performing optimization, and recording an optimization result of the RL-SSA as a final identification result of the parameter. According to the method, the unique advantages of the RL in the identification process are fully utilized, and the identification effect of the control parameters is remarkably improved through continuous interaction between the intelligent agent and the environment and strategy iterative optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of photovoltaic grid-connected inverter control technology, and in particular to a photovoltaic grid-connected inverter controller parameter identification method based on a hybrid RL-SSA. Background Art

[0002] When analyzing and studying the grid-connected characteristics of photovoltaic inverters, it is necessary to obtain their control parameters through identification methods to establish an accurate simulation model for the photovoltaic power generation system. Currently, traditional parameter identification methods have inherent flaws in terms of complexity and adaptability. Although the integration of intelligent algorithms has injected new vitality into the research field of electrical engineering and achieved significant application in the field of parameter identification, existing algorithm update strategies are still subject to significant randomness and limitations, lacking clear regularity guidance, resulting in a slow convergence process.

[0003] To address this situation, a hybrid reinforcement learning (RL)-sparrow search algorithm (SSA) method for photovoltaic grid-connected inverter controller parameter identification is proposed. RL is an advanced learning method based on trial and error, which iteratively optimizes decision-making strategies through continuous interaction between the agent and the environment. It is particularly suitable for highly autonomous decision-making and continuous learning tasks. Therefore, RL has demonstrated its unique potential and significant advantages in dealing with complex real-world problems. Among them, the Q-learning algorithm can formally describe the decision-making problem of the agent in the environment and maximize the expected value of the cumulative discounted reward through optimization, providing strong theoretical support for the accurate identification of photovoltaic grid-connected inverter controller parameters.

[0004] In summary, RL can deeply evaluate the state-action space of SSA individuals, provide a clear direction for algorithm updates, and effectively reduce random interference. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a parameter identification method for photovoltaic grid-connected inverter controller based on hybrid RL-SSA, which fully utilizes the unique advantages of RL in the identification process and significantly improves the identification effect of control parameters through continuous interaction between the intelligent agent and the environment and strategy iterative optimization.

[0006] The present invention provides a parameter identification method for a photovoltaic grid-connected inverter controller based on a hybrid RL-SSA, comprising the following steps:

[0007] Step S1. Establish a three-phase grid-connected power generation system simulation model, obtain low voltage ride-through conditions by setting a three-phase short circuit fault, and record the inverter DC side voltage u dc , the inner current loop controls the current i d and iq ;

[0008] Step S2: Determine the parameter k to be identified of the photovoltaic grid-connected inverter controller in combination with the control strategy Pu 、k Iu 、k Pi and k Ii ;

[0009] Step S3: Use the SSA algorithm to identify the parameter k Pu 、k Iu 、k Pi and k Ii The established initial population is divided into "discoverers", "joiners" and "alerts", and the waveform of reactive current during transient period is used as the objective function;

[0010] Step S4. Based on the population diversity, number of iterations and fitness distribution as features, the reinforcement learning idea is used to establish the optimization state space of the RL-SSA algorithm;

[0011] Step S5. Design the action space of RL-SSA according to the greedy strategy;

[0012] Step S6. Improve the RL reward mechanism according to the current population diversity, number of iterations and fitness distribution, and then use SSA to control the parameter k of the photovoltaic grid-connected inverter controller. Pu 、k Iu 、k Pi and k Ii Perform optimization and record the optimization result of RL-SSA as the final identification result of the parameter.

[0013] Preferably, in the three-phase grid-connected power generation system simulation model in step S1, the inverter circuit of the inverter link topology structure is composed of six IGBTs with anti-parallel diodes, and the DC input is converted into an AC output u by regularly repeating the switching elements. a 、u b 、u c ,i a 、i b and i c , and after filtering, the grid connection point voltage e is obtained a 、e b and e c After dq transformation, we get u d and u q 、i d and i q 、e d and e q .

[0014] Preferably, the control strategy of the photovoltaic grid-connected inverter controller in step S2 adopts a voltage-oriented dual closed-loop control, wherein the dual closed-loop control includes a voltage outer loop and a current inner loop, wherein:

[0015] Inverter DC side voltage u dc With the given reference voltage u dc * The comparison result is output by the outer loop PI controller as the current reference value i of the d-axis component of the current d * , as shown in formula (1). The reference value of the current q-axis component during normal operation of the grid-connected point is recorded as i q * =0,

[0016]

[0017] Where k Pu and k Iu are the proportional coefficient and integral coefficient of the voltage outer loop respectively;

[0018] Current inner loop control current i d 、i q Its reference value i d * 、i q * The control target u of the inverter DC side voltage is output through the inner loop PI controller d * 、u q * , whose expression is:

[0019]

[0020] Where: k Pi is the current inner loop PI controller proportional coefficient, k Ii is the integral coefficient of the current inner loop PI controller;

[0021] When a grid fault occurs, the voltage at the grid connection point drops. When the voltage returns to normal, the dynamic reactive current output by the inverter needs to continuously track the grid connection point voltage. The per-unit value of the grid connection point voltage is u * , record the effective value of the grid-connected current as i, corresponding to i q * and i d * The response expressions are:

[0022]

[0023] Where u l and u hThe low voltage ride-through judgment threshold and high voltage ride-through judgment threshold, k l and k h They are respectively low voltage ride-through i q Support factor and HVRT q Support coefficient.

[0024] Preferably, the SSA algorithm in step S3 includes the following specific steps:

[0025] Step S3.1 Population parameter initialization:

[0026] Randomly generate N sparrow positions {n1,n2,...,n N}, where the position of each sparrow is a multidimensional vector used to represent the control parameters to be identified;

[0027] Step S3.2: Population classification:

[0028] According to the individual at position {n1,n2,...,n N The fitness value of} divides the population into three categories: "discoverers", "joiners" and "guardians". Among them, the y sparrows with better fitness are recorded as "discoverers", responsible for conducting global search in the population and leading the population to move to a more optimal area; the remaining Ny sparrows are recorded as "joiners", which need to conduct local search with the discoverers; at the same time, some sparrows are randomly selected as "guardians", responsible for monitoring dangers and escaping from the local optimal solution in the optimization process;

[0029] Step S3.3 sets the sparrow position update method:

[0030] The "discoverer" role is responsible for finding the global optimal area. Its warning value is set to R2. When R2 is lower than the threshold ST, it indicates that there is no predator threat around the current foraging environment. At this time, the "discoverer" can perform extensive search behavior. On the contrary, if R2 reaches or exceeds ST, it means that some sparrows in the population have detected predators and sounded the alarm. At this time, all sparrows must immediately migrate to other safe areas to continue foraging. Based on this, the "discoverer" position update follows the following formula:

[0031]

[0032] Where n i t and n i t+1 represents the position of sparrow i at the tth iteration and the t+1th iteration, R2∈[0,1] is a random number that simulates the sparrow's alertness to the environment, α is the step coefficient, and θ is a random number that follows a normal distribution; L is a row vector whose elements are all 1 and corresponds to the dimension of the parameter to be identified;

[0033] The "joiner" will move closer to the location with the best fitness among the discoverers, but there is also a certain probability that it will forage alone. The result of its position update is n i t+1 It can be expressed as:

[0034]

[0035] A + =A T (AA T ) -1 (7)

[0036] Where n worst t Indicates the position with the worst fitness in the population after iteration t, n p t+1 is the position with the best fitness among the discoverers that have completed iteration t+1 times. A represents a matrix whose elements are randomly 1 or -1. When i>N / 2, it means that the joiner i with a lower fitness value has not obtained food and needs to fly to other places to find food.

[0037] The “alert” escapes the current area with a certain probability to avoid falling into the local optimum. i >F best This means that the sparrows are at the edge of their population and are extremely vulnerable to predators. best t This means that the sparrow at this position is the best position in the population and is also very safe; F i =F best This indicates that the sparrows in the middle of the population are aware of the danger and need to move closer to other sparrows to minimize their risk of being preyed upon. The update formula is:

[0038]

[0039] Where n best t is the position with the best fitness in the population after t iterations, β∈[-1,1] is the step length control parameter of the sparrow movement, and it is a random number that obeys the normal distribution, F i represents the fitness of sparrow i, F best and F worst is the global optimal fitness and the worst fitness;

[0040] Step S3.4 calculates the fitness values ​​of all sparrows in the new location:

[0041] If it is better than the original position, update its position, otherwise keep it unchanged. i As an example, the objective function of formula (9) is used in this evaluation process,

[0042]

[0043] Preferably, in step S3.2, individuals with better positions in the sparrow population are regarded as "discoverers" accounting for 20%-30% of the total; at the same time, 10%-20% of the individuals are randomly assigned as "alerts"; in addition, the safety threshold is set to ST∈[0.5,1] to simulate whether the individual discovers environmental danger.

[0044] Preferably, establishing the RL-SSA optimization algorithm state space in step S4 includes the following specific steps:

[0045] Step S4.1: The state and potential action corresponding to the individuals in the population are recorded as s and a respectively. Then the five elements required to apply RL to them are: state space, action space, state transition probability p[s t |(s,a)], reward function R[s t |(s,a)] and discount factor γ, where the state space and action space represent all environmental states and potential actions of the individual, respectively, and p[s t |(s,a)] and R[s t |(s,a)] respectively represent the individual in state s transferred to state s under the action of action a t The probability and reward obtained, γ is the discount factor that balances the importance of current rewards and future rewards;

[0046] For SSA, after t iterations, the individual has state s t , you need to choose an action a t As an update strategy, the new state s is fed back through the environment t+1 and instant reward R t , through continuous trial and learning, the accumulated rewards are maximized;

[0047] Step S4.2: The individual diversity D of the current iteration number t is t , the fitness distribution entropy H in the population t , the relationship between the number of iterations and the maximum number of iterations t / T max The features are used as indicators as the basis for dividing the states and establishing the Q learning table.

[0048]

[0049] Where D max is the length of the diagonal of the search space;

[0050] According to the distribution histogram of individual fitness at the tth iteration according to the objective function statistics, the entropy value is calculated using formula (11):

[0051]

[0052] In the formula, k is the number of intervals of individual fitness, p k is the probability that an individual falls into the kth interval;

[0053] Step S4.3 For individual diversity, when D t When ≥1.5, it is recorded as high diversity, indicating that individuals are in the exploration stage and the degree of dispersion is large. When 0.2≤D t When D < 1.5, it is recorded as medium diversity, and individuals are in the balance and development stage, with local aggregation. t When H<0.2, it is recorded as low diversity, indicating that the population diversity is low, in a highly convergent state, and there is a risk of being in a local optimum. For the fitness distribution entropy, when H t When H is ≥1.5, it is recorded as high entropy, indicating that the individual fitness values ​​are dispersed and the distribution range is wide. t <1.5 is recorded as low entropy, and individuals need to enhance development; for iterative progress, 0<(t / T max )<0.3、0.3≤(t / T max )<0.7 and 0.7≤(t / T max )<1.0 is recorded as the early iteration, mid-iteration and late iteration, and the update strategy also needs to be adjusted to ensure the optimization effect;

[0054] Under extreme conditions, when D t Code value = 0, H t Code value = 0 and t / T max When the encoding value = 0, it corresponds to the typical population state at the beginning of the iteration, with high diversity and large degree of disorder in fitness distribution, and the search intensity needs to be increased. t Code value = 2, H t Code value = 1 and t / T max When the encoding value = 2, it corresponds to the typical population state in the late iteration, and the search range is reduced to perform optimization in a small range;

[0055] Table 1 Population characteristic coding values

[0056]

[0057] Table 2 Population characteristic coding values

[0058]

[0059] According to the index values ​​in Tables 1 and 2, the individual states are coded and quantified. There are six individual state parameters s, which are recorded as s1, s2, s3, s4, s5 and s6 in Table 3;

[0060] Table 3 Code values ​​corresponding to status.

[0061]

[0062] Preferably, the action space design of RL-SSA in step S5 uses equations (12)-(13) to improve the RL-SSA algorithm:

[0063] If the mutation probability e>0.5, then:

[0064] n i t+1 =n i t+1 +p m1 (n i1 -n i2 )+p m2 (n i3 -n i4 )(12)

[0065] If the mutation probability e<0.5, that is, 1-e>0.5:

[0066] n i temp =n best +p m1 (n i1 -n i2 )+p m2 (n i3 -n i4 )(13)

[0067]

[0068] Where n i1 、n i2 、n i3 and n i4 are random individuals i1, i2, i3 and i4, and satisfy i1≠i2≠i3≠i4≠i, p m1 and p m2 is a random factor in the range [-1,1], used to adjust the sparrow position, n i temp is the mutation result of e<0.5.

[0069] Preferably, the reward mechanism corresponding to different states in step S6 is set as follows:

[0070] Individuals use Q-learning tables to calculate the reward values ​​for choosing different actions under different conditions, and dynamically update the Q-learning table through iteration, thereby gradually optimizing the action-value function so that individuals can make the best decision when facing a specific state. The Q-learning table can be expressed as:

[0071]

[0072] Where a1, a2, a3, a4, a5 and a6 are the action spaces corresponding to the current state;

[0073] in:

[0074] Rewards for population diversity characteristics r Dt Denoted as:

[0075]

[0076] The reward r obtained by the number of population iterations t Denoted as:

[0077]

[0078] Where, F best (t) and F best (t-1) is the optimal fitness of the population at iteration t and t-1 respectively;

[0079] The reward r obtained by the population fitness entropy characteristic Ht Denoted as:

[0080]

[0081] Where sgn represents the sign function, F best0 is the historical optimal fitness;

[0082] Combining equations (16)-(18), the direct reward r obtained after executing action a is expressed as:

[0083]

[0084] The Q-value iterative update formula based on the Bellman equation is as follows:

[0085] Q(s,a)=Q(s,a)+ω[r+γQ max -Q(s,a)] (20)

[0086] Where Q(s,a) is the action value function for executing action a in the current state s, ω is the weight that controls the new information to cover the old value (0<ω≤1), γ is the discount factor that balances the importance of current and future rewards (0≤γ<1), and Q max is the maximum Q value of all possible actions in the next state s.

[0087] Beneficial effects of the present invention:

[0088] (1) SSA excels in global search capabilities, avoiding the problem of local optimal solutions and quickly locating near the global optimal solution, thus providing a powerful solution for high-dimensional optimization problems. In high-dimensional space, the algorithm can maintain stable search performance and will not significantly reduce search efficiency due to the increase in dimension.

[0089] (2) Through the quantitative coding of diversity, entropy and number of iterations, the dynamic changes of the population during the iteration process can be accurately characterized, providing an objective and measurable standard for algorithm performance evaluation, effectively avoiding the bias of subjective judgment, and providing clear guidance for the optimization of the algorithm's optimization strategy.

[0090] (3) The greedy strategy gives individuals the necessary flexibility in unknown or dynamic environments. It explores unknown actions to discover potential high-value strategies while taking into account the continuity of proven effective actions to maintain efficiency. This strategy combines practicality and flexibility, which not only guarantees the performance bottom line but also promotes individuals' effective learning and decision-making in complex and uncertain environments. Its dual attributes ensure that individuals are not limited to the use of existing knowledge, but also enhance the adaptability of strategies through continuous exploration. This balancing mechanism not only helps individuals find better solutions in complex environments, but also improves the efficiency of the learning process, ensuring that individuals can efficiently accumulate and utilize information at different learning stages. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] Figure 1 It is the topology structure of the inverter link;

[0092] Figure 2 This is a schematic diagram of the control strategy switching of the grid-connected inverter;

[0093] Figure 3 This is the SSA parameter identification flow chart;

[0094] Figure 4 is a schematic diagram of MDP;

[0095] Figure 5 This is the flow chart of RL-SSA parameter identification;

[0096] Figure 6 i is the disturbance condition d ;

[0097] Figure 7 i is the fault condition d ;

[0098] Figure 8 i is the fault condition q . DETAILED DESCRIPTION

[0099] In order to make the technical solution of the present invention easier to understand, the technical solution of the present invention is now clearly and completely described in the form of specific embodiments in conjunction with the accompanying drawings.

[0100] Example 1:

[0101] The photovoltaic grid-connected inverter controller parameter identification method based on hybrid RL-SSA of this embodiment includes the following steps:

[0102] Step S1. Establish a three-phase grid-connected power generation system simulation model, obtain low voltage ride-through conditions by setting a three-phase short circuit fault, and record the inverter DC side voltage u dc , the inner current loop controls the current i d and i q , a simulation model of a three-phase grid-connected power generation system. The inverter circuit of its inverter link topology consists of six IGBTs with anti-parallel diodes. Through regular repeated switching of switching elements such as transistors, the DC input is converted into an AC output u a 、u b 、u c ,i a 、i b and i c , and after filtering, the grid connection point voltage e is obtained a 、e b and e c After dq transformation, we get u d and u q 、i d and i q 、e d and e q ;

[0103] Step S2: Determine the parameter k to be identified of the photovoltaic grid-connected inverter controller in combination with the control strategy Pu 、k Iu 、k Pi and k Ii The control strategy of the photovoltaic grid-connected inverter controller adopts a voltage-oriented dual closed-loop control, which includes a voltage outer loop and a current inner loop, wherein:

[0104] Inverter DC side voltage u dc With the given reference voltage u dc * The comparison result is output by the outer loop PI controller as the current reference value i of the d-axis component of the current d * , as shown in formula (1). The reference value of the current q-axis component during normal operation of the grid-connected point is recorded as i q * =0,

[0105]

[0106] Where k Pu and k Iu are the proportional coefficient and integral coefficient of the voltage outer loop respectively;

[0107] Current inner loop control current i d 、i q Its reference value i d * 、i q * The control target u of the inverter DC side voltage is output through the inner loop PI controller d * 、u q * , whose expression is:

[0108]

[0109] Where: k Pi is the current inner loop PI controller proportional coefficient, k Ii is the integral coefficient of the current inner loop PI controller;

[0110] When a grid fault occurs, the voltage at the grid connection point drops. When the voltage returns to normal, the dynamic reactive current output by the inverter needs to continuously track the grid connection point voltage. The per-unit value of the grid connection point voltage is u * , record the effective value of the grid-connected current as i, corresponding to i q * and i d * The response expressions are:

[0111]

[0112] Where u l and u h The low voltage ride-through judgment threshold and high voltage ride-through judgment threshold, k l and k h They are respectively low voltage ride-through i q Support factor and HVRT q Support coefficient. In addition, the current control strategy considering the limiting effect is as follows Figure 2 shown.

[0113] Step S3: Use the SSA algorithm to identify the parameter k Pu 、k Iu 、k Pi and k IiThe initial population is divided into "discoverers," "joiners," and "guardians," with the reactive current waveform during transient periods serving as the objective function. SSA is a swarm intelligence optimization algorithm based on the foraging and vigilance behaviors of sparrow colonies. Its core concept is to simulate the collaborative behavior of the three roles of "discoverers," "joiners," and "guardians" in a sparrow colony, and to find the global optimal solution by dynamically adjusting the search strategy. The SSA algorithm includes the following specific steps:

[0114] Step S3.1 Population parameter initialization:

[0115] Randomly generate N sparrow positions {n1,n2,...,n N}, where the position of each sparrow is a multidimensional vector used to represent the control parameters to be identified;

[0116] Step S3.2: Population classification:

[0117] According to the individual at position {n1,n2,...,n N The fitness value of} divides the population into three categories: "discoverers", "joiners" and "guardians", among which: the y sparrows with better fitness are recorded as "discoverers", responsible for global search in the population and leading the population to move to a better area; the remaining Ny sparrows are recorded as "joiners", which need to conduct local search with the discoverers; at the same time, some sparrows are randomly selected as "guardians", responsible for monitoring dangers and escaping from the local optimal solution in the optimization process, such as Figure 3 As shown;

[0118] Step S3.3 sets the sparrow position update method:

[0119] The "discoverer" role is responsible for finding the global optimal area. Its warning value is set to R2. When R2 is lower than the threshold ST, it indicates that there is no predator threat around the current foraging environment. At this time, the "discoverer" can perform extensive search behavior. On the contrary, if R2 reaches or exceeds ST, it means that some sparrows in the population have detected predators and sounded the alarm. At this time, all sparrows must immediately migrate to other safe areas to continue foraging. Based on this, the "discoverer" position update follows the following formula:

[0120]

[0121] Where n i t and n i t+1 represents the position of sparrow i at the tth iteration and the t+1th iteration, R2∈[0,1] is a random number that simulates the sparrow's alertness to the environment, α is the step coefficient, and θ is a random number that follows a normal distribution; L is a row vector whose elements are all 1 and corresponds to the dimension of the parameter to be identified;

[0122] The "joiner" will move closer to the location with the best fitness among the discoverers, but there is also a certain probability that it will forage alone. The result of its position update is n i t+1 It can be expressed as:

[0123]

[0124] A + =A T (AA T ) -1 (7)

[0125] Where n worst t Indicates the position with the worst fitness in the population after iteration t, n p t+1 is the position with the best fitness among the discoverers that have completed iteration t+1 times. A represents a matrix whose elements are randomly 1 or -1. When i>N / 2, it means that the joiner i with a lower fitness value has not obtained food and needs to fly to other places to find food.

[0126] The “alert” escapes the current area with a certain probability to avoid falling into the local optimum. i >F best This means that the sparrows are at the edge of their population and are extremely vulnerable to predators. best t This means that the sparrow at this position is the best position in the population and is also very safe; F i =F best This indicates that the sparrows in the middle of the population are aware of the danger and need to move closer to other sparrows to minimize their risk of being preyed upon. The update formula is:

[0127]

[0128] Where n best t is the position with the best fitness in the population after t iterations, β∈[-1,1] is the step length control parameter of the sparrow movement, and it is a random number that obeys the normal distribution, F i represents the fitness of sparrow i, F best and F worst is the global optimal fitness and the worst fitness;

[0129] Step S3.4 calculates the fitness values ​​of all sparrows in the new location:

[0130] If it is better than the original position, update its position, otherwise keep it unchanged. i As an example, the objective function of formula (9) is used in this evaluation process,

[0131]

[0132] In step S3.2, individuals with better positions in the sparrow population are regarded as "discoverers", accounting for 20%-30% of the total; at the same time, 10%-20% of the individuals are randomly assigned as "alerts"; in addition, the safety threshold is set to ST∈[0.5,1] to simulate whether the individual discovers environmental danger.

[0133] Step S4. Based on population diversity, number of iterations, and fitness distribution as features, the RL-SSA algorithm's optimization state space is established using reinforcement learning. In reinforcement learning theory, Q-learning is a model-free reinforcement learning algorithm based on the MDP framework. The core of the RL-SSA algorithm is to use Q-learning to improve individual iteration strategies, thereby achieving an adaptive optimization process.

[0134] Establishing the state space of the RL-SSA optimization algorithm includes the following specific steps:

[0135] Step S4.1: The state and potential action corresponding to the individuals in the population are recorded as s and a respectively. Then the five elements required to apply RL to them are: state space, action space, state transition probability p[s t |(s,a)], reward function R[s t |(s,a)] and discount factor γ, where the state space and action space represent all environmental states and potential actions of the individual, respectively, and p[s t |(s,a)] and R[s t |(s,a)] respectively represent the individual in state s transferred to state s under the action of action a t The probability and reward obtained, γ is the discount factor that balances the importance of current rewards and future rewards;

[0136] For SSA, after t iterations, the individual has state s t , you need to choose an action a t As an update strategy, the new state s is fed back through the environment t+1 and instant reward R t ,like Figure 4 As shown, through continuous trial and learning, the accumulated rewards are maximized;

[0137] Step S4.2: The individual diversity D of the current iteration number t is t , the fitness distribution entropy H in the population t , the relationship between the number of iterations and the maximum number of iterations t / T max The features are used as indicators as the basis for dividing the states and establishing the Q learning table.

[0138]

[0139] Where D max is the length of the diagonal of the search space;

[0140] According to the distribution histogram of individual fitness at the tth iteration according to the objective function statistics, the entropy value is calculated using formula (11):

[0141]

[0142] In the formula, k is the number of intervals of individual fitness, p k is the probability that an individual falls into the kth interval;

[0143] Step S4.3 For individual diversity, when D t When ≥1.5, it is recorded as high diversity, indicating that individuals are in the exploration stage and the degree of dispersion is large. When 0.2≤D t When D < 1.5, it is recorded as medium diversity, and individuals are in the balance and development stage, with local aggregation. t When H<0.2, it is recorded as low diversity, indicating that the population diversity is low, in a highly convergent state, and there is a risk of being in a local optimum. For the fitness distribution entropy, when H t When H is ≥1.5, it is recorded as high entropy, indicating that the individual fitness values ​​are dispersed and the distribution range is wide. t <1.5 is recorded as low entropy, and individuals need to enhance development; for iterative progress, 0<(t / T max )<0.3、0.3≤(t / T max )<0.7 and 0.7≤(t / T max )<1.0 is recorded as the early iteration, mid-iteration and late iteration, and the update strategy also needs to be adjusted to ensure the optimization effect;

[0144] Under extreme conditions, when D t Code value = 0, H t Code value = 0 and t / T max When the encoding value = 0, it corresponds to the typical population state at the beginning of the iteration, with high diversity and large degree of disorder in fitness distribution, and the search intensity needs to be increased. t Code value = 2, H t Code value = 1 and t / T max When the encoding value = 2, it corresponds to the typical population state in the late iteration, and the search range is reduced to perform optimization in a small range;

[0145] Table 1 Population characteristic coding values

[0146]

[0147] Table 2 Population characteristic coding values

[0148]

[0149] According to the index values ​​in Tables 1 and 2, the individual states are coded and quantified. There are six individual state parameters s, which are recorded as s1, s2, s3, s4, s5 and s6 in Table 3;

[0150] Table 3 Code values ​​corresponding to status.

[0151]

[0152] Step S5. Design the action space of RL-SSA according to the greedy strategy and improve the RL-SSA algorithm using equations (12)-(13):

[0153] If the mutation probability e>0.5, then:

[0154] n i t+1 =n i t+1 +p m1 (n i1 -n i2 )+p m2 (n i3 -n i4 )(12)

[0155] If the mutation probability e<0.5, that is, 1-e>0.5:

[0156] n i temp =n best +p m1 (n i1 -n i2 )+p m2 (n i3 -n i4 )(13)

[0157]

[0158] Where n i1 、n i2 、n i3 and n i4 are random individuals i1, i2, i3 and i4, and satisfy i1≠i2≠i3≠i4≠i, p m1 and p m2 is a random factor in the range [-1,1], used to adjust the sparrow position, n i temp is the mutation result of e<0.5.

[0159] Adaptively selecting the optimal search strategy based on the optimization state of the parameters to be identified can effectively improve optimization efficiency. At the individual decision-making stage, the greedy strategy is a widely used approach, focusing on striking a balance between acquiring new information and leveraging existing knowledge. Within this strategy framework, individuals do not always choose the currently optimal option. Instead, they explore based on a preset probability e, randomly selecting actions. This approach allows them to explore actions that have not yet been tried or fully understood, hoping to discover more optimal strategies. Furthermore, the optimal action, based on existing experience, is selected with a probability of 1-e, ensuring that individuals fully utilize their accumulated knowledge to achieve stable and reliable results. This approach further reduces the risk of the results falling into local optima.

[0160] Step S6. Improve the RL reward mechanism according to the current population diversity, number of iterations and fitness distribution, and then use SSA to control the parameter k of the photovoltaic grid-connected inverter controller. Pu 、k Iu 、k Pi and k Ii Perform optimization and record the optimization result of RL-SSA as the final identification result of the parameter.

[0161] Individuals use Q-learning tables to calculate the reward values ​​for choosing different actions under different conditions, and dynamically update the Q-learning table through iteration, thereby gradually optimizing the action-value function so that individuals can make the best decision when facing a specific state. The Q-learning table can be expressed as:

[0162]

[0163] Where a1, a2, a3, a4, a5 and a6 are the action spaces corresponding to the current state;

[0164] During individual learning, high diversity in the exploration phase leads to high rewards, while in the mid-term, equilibrium phase, diversity provides moderate rewards. However, if diversity is too low in the late phase, it will be penalized to avoid premature convergence.

[0165] in:

[0166] Rewards for population diversity characteristics r Dt Denoted as:

[0167]

[0168] The reward r obtained by the number of population iterations t Denoted as:

[0169]

[0170] Where, F best (t) and F best(t-1) is the optimal fitness of the population at iteration t and t-1 respectively;

[0171] The reward r obtained by the population fitness entropy characteristic Ht Denoted as:

[0172]

[0173] Where sgn represents the sign function, F best0 is the historical optimal fitness;

[0174] Combining equations (16)-(18), the direct reward r obtained after executing action a is expressed as:

[0175]

[0176] The Q-value iterative update formula based on the Bellman equation is as follows:

[0177] Q(s,a)=Q(s,a)+ω[r+γQ max -Q(s,a)] (20)

[0178] Where Q(s,a) is the action value function for executing action a in the current state s, ω is the weight that controls the new information to cover the old value (0<ω≤1), γ is the discount factor that balances the importance of current and future rewards (0≤γ<1), and Q max is the maximum Q value of all possible actions in the next state s.

[0179] In order to further verify the effectiveness and accuracy of the proposed method for control parameter identification, the accuracy of control parameter identification is tested using the disturbance conditions and fault conditions shown in Table 4. The identification results are shown in Table 4. Figure 6-Figure 8 shown.

[0180] Table 4 Simulation model test conditions

[0181]

[0182] It should be noted that the embodiments described herein are only some embodiments of the present invention, not all implementations of the present invention. The embodiments are merely illustrative and serve only to provide a more intuitive and clear way to understand the contents of the present invention, rather than to limit the technical solutions described in the present invention. Without departing from the concept of the present invention, all other implementations that can be thought of by ordinary technicians in this field without creative work, as well as other simple replacements and various variations of the technical solutions of the present invention, are within the scope of protection of the present invention.

Claims

1. A photovoltaic grid-connected inverter controller parameter identification method based on hybrid RL-SSA, characterized in that: The following steps are involved: Step S1. Establish a three-phase grid-connected power generation system simulation model, obtain low voltage ride-through conditions by setting a three-phase short circuit fault, and record the inverter DC side voltage u dc , the inner current loop controls the current i d and i q ; Step S2: Determine the parameter k to be identified of the photovoltaic grid-connected inverter controller in combination with the control strategy Pu 、k Iu 、k Pi and k Ii ; Step S3: Use the SSA algorithm to use the parameter k to be identified Pu 、k Iu 、k Pi and k Ii The initial population is divided into "discoverers", "joiners" and "alerts", and the waveform of reactive current during transient period is used as the objective function; Step S4. Based on the population diversity, number of iterations and fitness distribution as features, the reinforcement learning idea is used to establish the optimization state space of the RL-SSA algorithm; Step S5. Design the action space of RL-SSA according to the greedy strategy; Step S6. Improve the RL reward mechanism according to the current population diversity, number of iterations and fitness distribution, and then use SSA to control the parameter k of the photovoltaic grid-connected inverter controller. Pu 、k Iu 、k Pi and k Ii Perform optimization and record the optimization result of RL-SSA as the final identification result of the parameter.

2. The photovoltaic grid-connected inverter controller parameter identification method based on hybrid RL-SSA according to claim 1, characterized in that: The three-phase grid-connected power generation system simulation model in step S1 has an inverter circuit with an inverter link topology consisting of six IGBTs with anti-parallel diodes. The DC input is converted into an AC output u by regularly switching the switching elements. a 、u b 、u c ,i a 、i b and i c , and after filtering, the grid connection point voltage e is obtained a 、e b and e c , after dq transformation, we get u d and u q 、i d and i q 、e d and e q .

3. The photovoltaic grid-connected inverter controller parameter identification method based on hybrid RL-SSA according to claim 2, characterized in that: The control strategy of the photovoltaic grid-connected inverter controller in step S2 adopts a voltage-oriented dual closed-loop control, which includes a voltage outer loop and a current inner loop, wherein: Inverter DC side voltage u dc With the given reference voltage u dc * The comparison result is output by the outer loop PI controller as the current reference value i of the d-axis component of the current d * , as shown in formula (1), and the reference value of the current q-axis component during normal operation of the grid point is recorded as i q * =0, Where k Pu and k Iu are the proportional coefficient and integral coefficient of the voltage outer loop respectively; Current inner loop control current i d 、i q Its reference value i d * 、i q * The control target u of the inverter DC side voltage is output through the inner loop PI controller d * 、u q * , whose expression is: Where: k Pi is the proportional coefficient of the current inner loop PI controller, k Ii is the integral coefficient of the current inner loop PI controller; When a grid fault occurs, the voltage at the grid connection point drops. When the voltage returns to normal, the dynamic reactive current output by the inverter needs to continuously track the grid connection point voltage. The per-unit value of the grid connection point voltage is u * , record the effective value of the grid-connected current as i, corresponding to i q * and i d * The response expressions are: Where u l and u h The low voltage ride-through judgment threshold and high voltage ride-through judgment threshold, k l and k h They are respectively low voltage ride-through i q Support factor and HVRT q Support coefficient.

4. The photovoltaic grid-connected inverter controller parameter identification method based on hybrid RL-SSA according to claim 1, characterized in that: The SSA algorithm in step S3 includes the following specific steps: Step S3.1 Population parameter initialization: Randomly generate N sparrow positions {n1,n2,...,n N }, where the position of each sparrow is a multidimensional vector used to represent the control parameters to be identified; Step S3.2: Population classification: According to the individual at position {n1,n2,...,n N The fitness values ​​of the y sparrows are divided into three categories: "discoverers", "joiners" and "guardians". Among them, the y sparrows with better fitness are recorded as "discoverers", responsible for conducting global search in the population and leading the population to move to a more optimal area; the remaining Ny sparrows are recorded as "joiners", which need to conduct local search with the discoverers; at the same time, some sparrows are randomly selected as "guardians", responsible for monitoring dangers and escaping from the local optimal solution during the optimization process; Step S3.3 sets the sparrow position update method: The "discoverer" role is responsible for finding the global optimal area. Its warning value is set to R2. When R2 is lower than the threshold ST, it indicates that there is no predator threat in the current foraging environment, and the "discoverer" can now conduct extensive search behavior. Conversely, if R2 reaches or exceeds ST, it means that some sparrows in the population have detected a predator and sounded the alarm. At this time, all sparrows must immediately migrate to other safe areas to continue foraging. Based on this, the "discoverer" position update follows the following formula: Where n i t and n i t+1 represents the position of sparrow i at the tth iteration and the t+1th iteration, R2∈[0,1] is a random number that simulates the sparrow's alertness to the environment, α is the step coefficient, and θ is a random number that follows a normal distribution; L is a row vector whose elements are all 1 and corresponds to the dimension of the parameter to be identified; The "joiner" will move closer to the location with the best fitness among the discoverers, but there is also a certain probability that it will forage alone. The result of its position update is n i t+1 It can be expressed as: A + =A T (CHALLENGE ACCEPTED T ) -1 (7) Where n worst t Indicates the position with the worst fitness in the population after iteration t, n p t+1 is the position with the best fitness among the discoverers that have completed iteration t+1 times. A represents a matrix whose elements are randomly 1 or -1. When i>N / 2, it means that the joiner i with a lower fitness value has not obtained food and needs to fly to other places to find food. The "alert" escapes the current area with a certain probability to avoid falling into the local optimum. i >F best This means that the sparrows are at the edge of their population and are extremely vulnerable to predators. best t This means that the sparrow at this position is the best position in the population and is also very safe; F i =F best This indicates that the sparrows in the middle of the population are aware of the danger and need to move closer to other sparrows to minimize their risk of being preyed upon. The update formula is: Where n best t is the position with the best fitness in the population after t iterations, β∈[-1,1] is the step length control parameter of the sparrow movement, and it is a random number that obeys the normal distribution, F i represents the fitness of sparrow i, F best and F worst is the global optimal fitness and the worst fitness; Step S3.4 calculates the fitness values ​​of all sparrows in the new location: If it is better than the original position, update its position, otherwise it remains unchanged. i , the objective function of formula (9) is used in this evaluation process, 5. The photovoltaic grid-connected inverter controller parameter identification method based on hybrid RL-SSA according to claim 4, characterized in that: In step S3.2, individuals with better positions in the sparrow population are regarded as "discoverers", accounting for 20%-30% of the total; at the same time, 10%-20% of the individuals are randomly assigned as "alerts"; in addition, the safety threshold is set to ST∈[0.5,1] to simulate whether the individual discovers environmental danger.

6. The photovoltaic grid-connected inverter controller parameter identification method based on hybrid RL-SSA according to claim 1, characterized in that: The establishment of the RL-SSA optimization algorithm state space in step S4 includes the following specific steps: Step S4.1: The state and potential action corresponding to the individuals in the population are recorded as s and a respectively. Then the five elements required to apply RL to them are: state space, action space, state transition probability p[s t |(s,a)], reward function R[s t |(s,a)] and discount factor γ, where the state space and action space represent all environmental states and potential actions of the individual, respectively, and p[s t |(s,a)] and R[s t |(s,a)] respectively represent the individual in state s transferred to state s under the action of action a t The probability and reward obtained, γ is the discount factor that balances the importance of current rewards and future rewards; For SSA, after t iterations, the individual has state s t , you need to choose an action a t As an update strategy, the new state s is fed back through the environment t+1 and instant reward R t , through continuous trial and learning, the accumulated rewards are maximized; Step S4.2: The individual diversity D of the current iteration number t is t , the fitness distribution entropy H in the population t , the relationship between the number of iterations and the maximum number of iterations t / T max The features are used as indicators as the basis for dividing the states and establishing the Q learning table. Where D max is the length of the diagonal of the search space; According to the distribution histogram of individual fitness at the tth iteration according to the objective function statistics, the entropy value is calculated using formula (11): In the formula, k is the number of intervals of individual fitness, p k is the probability that an individual falls into the kth interval; Step S4.3 For individual diversity, when D t When ≥1.5, it is recorded as high diversity, indicating that individuals are in the exploration stage and the degree of dispersion is large. When 0.2≤D t When D < 1.5, it is recorded as medium diversity, and individuals are in the balance and development stage, with local aggregation. t When H<0.2, it is recorded as low diversity, indicating that the population diversity is low, in a highly convergent state, and there is a risk of being in a local optimum. For the fitness distribution entropy, when H t When H is ≥1.5, it is recorded as high entropy, indicating that the individual fitness values ​​are dispersed and the distribution range is wide. t <1.5 is recorded as low entropy, and individuals need to enhance development; for iterative progress, 0<(t / T max )<0.3、0.3≤(t / T max )<0.7 and 0.7≤(t / T max )<1.0 is recorded as the early iteration, mid-iteration and late iteration, and the update strategy also needs to be adjusted to ensure the optimization effect; Under extreme conditions, when D t Code value = 0, H t Code value = 0 and t / T max When the encoding value = 0, it corresponds to the typical population state at the beginning of the iteration, with high diversity and large degree of disorder in fitness distribution, and the search intensity needs to be increased. t Code value = 2, H t Code value = 1 and t / T max When the encoding value = 2, it corresponds to the typical population state in the late iteration, and the search range is reduced to perform optimization in a small range; Table 1 Population characteristic coding values Table 2 Population characteristic coding values According to the index values ​​in Tables 1 and 2, the individual states are coded and quantified. There are six individual state parameters s, which are recorded as s1, s2, s3, s4, s5 and s6 in Table 3; Table 3 Code values ​​corresponding to status.

7. The photovoltaic grid-connected inverter controller parameter identification method based on hybrid RL-SSA according to claim 1, characterized in that: The action space design of RL-SSA in step S5 uses equations (12)-(14) to improve the RL-SSA algorithm: If the mutation probability e>0.5, then: n i t+1 =n i t+1 +p m1 (n i1 -n i2 )+p m2 (n i3 -n i4 )(12) If the mutation probability e<0.5, that is, 1-e>0.5: n i temp =n best +p m1 (n i1 -n i2 )+p m2 (n i3 -n i4 )(13) Where n i1 、n i2 、n i3 and n i4 are random individuals i1, i2, i3 and i4, and satisfy i1≠i2≠i3≠i4≠i, p m1 and p m2 is a random factor in the range [-1,1], used to adjust the sparrow position, n i temp is the mutation result of e<0.

5.

8. The photovoltaic grid-connected inverter controller parameter identification method based on hybrid RL-SSA according to claim 1, characterized in that: The reward mechanism corresponding to different states in step S6 is set as follows: Individuals use Q-learning tables to calculate the reward values ​​for choosing different actions under different conditions, and dynamically update the Q-learning table through iteration, thereby gradually optimizing the action-value function so that individuals can make the best decision when facing a specific state. The Q-learning table can be expressed as: Where a1, a2, a3, a4, a5 and a6 are the action spaces corresponding to the current state; in: Rewards for population diversity characteristics r Dt Denoted as: The reward r obtained by the number of population iterations t Denoted as: Where, F best (t) and F best (t-1) is the optimal fitness of the population at iteration t and t-1 respectively; The reward r obtained by the population fitness entropy characteristic Ht Denoted as: Where sgn represents the sign function, F best0 is the historical optimal fitness; Combining equations (16)-(18), the direct reward r obtained after executing action a is expressed as: The Q-value iterative update formula based on the Bellman equation is as follows: Q(s,a)=Q(s,a)+ω[r+γQ max -Q(s,a)] (20) Where Q(s,a) is the action value function for executing action a in the current state s, ω is the weight that controls the new information to cover the old value (0<ω≤1), γ is the discount factor that balances the importance of current and future rewards (0≤γ<1), and Q max is the maximum Q value of all possible actions in the next state s.