Perovskite / crystalline silicon laminated photovoltaic module voltage adaptive matching control method

By constructing a graph neural network model and using reinforcement learning to optimize the control strategy, the voltage difference of the perovskite/crystalline silicon tandem photovoltaic module is dynamically adjusted, solving the voltage mismatch problem, achieving efficient voltage adaptive matching, and improving the power generation efficiency and stability of the module.

CN120811273AActive Publication Date: 2025-10-17JIANGSU XIEHANG ENERGY TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510907325.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-17
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

In existing perovskite/crystalline silicon tandem photovoltaic modules, there is a mismatch in the voltage output of perovskite sub-cells and crystalline silicon sub-cells, which affects the power generation efficiency and service life. Traditional control methods cannot be adjusted in real time and have high computational complexity, making it difficult to cope with environmental changes.

Method used

A voltage difference compensation model based on graph neural network is constructed. The node weights are adaptively adjusted by combining message passing mechanism and attention mechanism. The control strategy is optimized by using reinforcement learning and priority experience replay mechanism. The system parameters are identified in real time by recursive least squares algorithm, and the controller parameters are dynamically adjusted to achieve voltage matching.

Benefits of technology

It improves the accuracy and response speed of voltage matching, enabling photovoltaic modules to maintain optimal working condition under different environmental conditions, thereby improving energy conversion efficiency and service life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811273A_ABST
    Figure CN120811273A_ABST
Patent Text Reader

Abstract

The invention provides a perovskite / crystalline silicon laminated photovoltaic module voltage self-adaptive matching control method, which relates to the technical field of voltage control, and comprises the following steps: calculating a voltage difference value by collecting the voltage of a sub-battery; constructing a voltage difference compensation model based on the graph neural network to output an optimal voltage compensation value; compensating the voltage of the sub-battery; constructing a state space model based on the compensated voltage, and obtaining an optimization control strategy by using a priority experience playback mechanism; adjusting the working state of the driving circuit according to the control quantity. According to the invention, voltage self-adaptive dynamic matching between the sub-cells of the laminated assembly is realized, and the power generation efficiency of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to voltage control technology, in particular to a perovskite / crystalline silicon stacked photovoltaic module voltage adaptive matching control method. BACKGROUND

[0002] Perovskite / crystalline silicon stacked photovoltaic modules combine the high open-circuit voltage of perovskite solar cells and the high stability of crystalline silicon solar cells, becoming an important technical route to improve photoelectric conversion efficiency. In a four-terminal perovskite / crystalline silicon stacked photovoltaic module, perovskite sub-cells and crystalline silicon sub-cells are usually connected in parallel or electrically independent and output separately, and their output performance is limited by the voltage matching between the two sub-cells. However, in actual application, due to the different spectral response characteristics of perovskite and crystalline silicon materials, as well as changes in external conditions such as temperature and light intensity, the voltage outputs of the two sub-cells often do not match, affecting the power generation efficiency and service life of the entire stacked module.

[0003] Traditional control methods mostly use fixed parameter control strategies, which cannot adjust control parameters in real time according to dynamic changes in environmental conditions and module states, resulting in insufficient adaptability of the system in complex and variable working environments. Secondly, existing technologies do not adequately consider the mutual influence between perovskite sub-cells and crystalline silicon sub-cells, lack a collaborative optimization mechanism for the two sub-cells as a whole system, and are difficult to achieve globally optimal voltage matching. Thirdly, existing voltage matching algorithms have high computational complexity and insufficient real-time performance, and most methods do not fully consider the nonlinear and time-varying characteristics of photovoltaic systems, making it difficult to meet the rapid voltage adjustment requirements under conditions such as sudden changes in light intensity and temperature fluctuations. SUMMARY

[0004] The perovskite / crystalline silicon stacked photovoltaic module voltage adaptive matching control method provided by the embodiments of the present application can solve the problems in the prior art.

[0005] In a first aspect, the present application provides a perovskite / crystalline silicon stacked photovoltaic module voltage adaptive matching control method, comprising: Collecting the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell in the perovskite / crystalline silicon stacked photovoltaic module, and calculating the real-time voltage difference between the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell; Constructing a voltage difference compensation model based on a graph neural network, the voltage difference compensation model taking the perovskite sub-cell and the crystalline silicon sub-cell as graph network nodes and the voltage difference as edge features, iteratively updating the node state through a message passing mechanism, and adaptively adjusting the node weight in combination with an attention mechanism to output an optimal voltage compensation value; Compensate the perovskite sub-cell output voltage and the crystalline silicon sub-cell output voltage according to the optimal voltage compensation value, to obtain a compensated perovskite sub-cell voltage and a compensated crystalline silicon sub-cell voltage; Construct a system state space model based on the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage, generate a comprehensive reward signal based on the system state space model, update the state value estimation by using a priority experience replay mechanism to obtain an optimal control strategy, and identify the parameters of the system state space model in real time based on a recursive least squares algorithm, and solve the optimal control sequence in a rolling time domain while introducing an adaptive law to dynamically adjust the controller parameters; Adjust the driving circuit working states of the perovskite sub-cell output voltage and the crystalline silicon sub-cell output voltage according to the output control quantity of the optimal control strategy, to realize optimal voltage matching.

[0006] Construct a voltage difference compensation model based on a graph neural network, the perovskite sub-cell and the crystalline silicon sub-cell are taken as graph network nodes, the voltage difference is taken as edge features, the node state is iteratively updated through a message passing mechanism, and the node weight is adaptively adjusted by combining an attention mechanism, and an optimal voltage compensation value is output, including: Construct a voltage difference compensation model based on a graph neural network, the perovskite sub-cell and the crystalline silicon sub-cell are taken as graph network nodes, the voltage difference is taken as edge features, the node state is iteratively updated through a message passing mechanism, and the node weight is adaptively adjusted by combining an attention mechanism, and an optimal voltage compensation value is output, including: Input the initial edge feature representation into the ion channel module of the Hodgkin-Huxley neuron model, enhance the initial edge feature representation based on the sodium ion channel and potassium ion channel characteristics of the ion channel module, and form an enhanced edge feature representation with long-range correlation; Construct a message passing function based on the enhanced edge feature representation, the message passing function receives the node state vector and the enhanced edge feature representation as input, and generates node state update information; Input the node state update information into the synaptic plasticity module of the Hodgkin-Huxley neuron model, and adaptively adjust the node state update information through the dynamic threshold mechanism of the synaptic plasticity module to obtain adjusted node states; Applying an attention mechanism to the adjusted node state, a node attention coefficient is obtained by calculating the node state similarity and normalizing the node state similarity; and inputting the perovskite sub-cell node state and the crystalline silicon sub-cell node state with the node attention coefficient into a multi-layer perception machine to obtain an optimal voltage compensation value.

[0007] Calculating the output voltage difference between the perovskite sub-cell and the crystalline silicon sub-cell, taking the output voltage difference as an edge feature, and constructing an initial edge feature representation using a Hodgkin-Huxley neuron model including: Calculating the voltage difference between the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell, multiplying the voltage difference by a mapping coefficient and superimposing a resting potential to obtain a mapping membrane potential, the mapping coefficient being determined by the amplitude range of the voltage difference and the physiological range of the biological neuron membrane potential; Based on the mapping membrane potential, an ion channel dynamics model is constructed, and the time change rate of the mapping membrane potential is calculated by the membrane capacitance per unit area, the maximum conductance of each ion channel, the equilibrium potential of each ion, and the external input current; Based on the ion channel dynamics model, a gating variable dynamics equation is established, and the voltage-dependent rate constant of each gating variable is calculated based on the mapping membrane potential, the voltage-dependent rate constant including a forward rate constant and a reverse rate constant; the steady-state value and the time constant of each gating variable are calculated based on the forward rate constant and the reverse rate constant, and the time evolution of each gating variable is solved by substituting the steady-state value and the time constant into the exponential form of the dynamics equation, and the time evolution is numerically integrated by the Runge-Kutta method to form an initial edge feature representation.

[0008] Based on the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage, a system state space model is constructed, a comprehensive reward signal is generated based on the system state space model, and an optimized control strategy is obtained by updating the state value estimation using a priority experience replay mechanism including: A system state space model is constructed including the voltage state of the perovskite sub-cell and the voltage state of the crystalline silicon sub-cell, a prediction cost function is constructed based on the system state space model; an initial control strategy is calculated based on the prediction cost function, and an Actor-Critic agent is configured for the perovskite sub-cell and the crystalline silicon sub-cell based on the initial control strategy, and the Actor network of the Actor-Critic agent calculates the state expected return based on the state action value function. inputting the state expectation return into a hierarchical reward mechanism, the hierarchical reward mechanism calculating a local reward function according to a voltage tracking error of a single sub-cell, calculating a global reward function according to a joint tracking error of two sub-cell voltages, calculating a timing reward function according to a time derivative of a voltage difference, and combining the local reward function, the global reward function and the timing reward function to form a comprehensive reward signal; constructing a priority experience replay mechanism using the comprehensive reward signal, the priority experience replay mechanism calculating a policy gradient through a timing difference error, simultaneously updating a state value estimation of the Actor-Critic agent using the policy gradient, and obtaining an optimized control policy.

[0009] constructing a priority experience replay mechanism using the comprehensive reward signal, the priority experience replay mechanism calculating a policy gradient through a timing difference error, simultaneously updating a state value estimation of the Actor-Critic agent using the policy gradient, and obtaining an optimized control policy includes: calculating a fractal dimension of a policy trajectory for the comprehensive reward signal, constructing an experience priority evaluation function based on the fractal dimension, a timing difference error and a state transition similarity, the experience priority evaluation function being used to determine the importance of an experience sample; introducing a Logistic chaotic sequence into a timing difference error calculation process, the Logistic chaotic sequence modulating a discounted state value through a chaotic intensity parameter to obtain a chaotic-enhanced timing difference error, the chaotic-enhanced timing difference error being used to evaluate the timing correlation of a state-action pair; constructing a policy gradient based on the chaotic-enhanced timing difference error and the fractal dimension, inputting the policy gradient into an Actor network for parameter updating, simultaneously inputting the chaotic-enhanced timing difference error and a chaotic modulation term into a Critic network to update a state value estimation, and adaptively adjusting learning rates of the Actor network and the Critic network according to a training process; constructing a final optimized control policy based on the updated state value estimation, the final optimized control policy adjusting an exploration degree of action selection through a temperature parameter, and realizing a probabilistic policy output based on an action value function.

[0010] the optimized control policy identifying parameters of a system state space model in real time based on a recursive least squares algorithm, solving the optimized control sequence in a rolling time domain, and simultaneously introducing an adaptive law to dynamically adjust controller parameters includes: The system input and output data are collected, an initial parameter estimation model is constructed by using a recursive least square algorithm, the recursive least square algorithm iteratively updates a system parameter vector through a Kalman gain matrix, the system input and output data are wavelet-decomposed, the system input and output data are decomposed into different frequency band features, and the different frequency band features are combined with wavelet base functions to obtain a multi-scale signal representation; A parameter identification model is constructed at each frequency band based on the multi-scale signal representation, the parameter identification model adopts an iterative structure of the recursive least square algorithm, and the Kalman gain matrix is adjusted according to frequency band characteristics; An identification result of the parameter identification model is input into a radial basis function neural network, network parameters of the radial basis function neural network are adjusted by using a gradient descent method, and an update step of the gradient descent method is calculated according to an identification error; A distributed optimization calculation is performed on an output result of the radial basis function neural network, parameters of the recursive least square algorithm are corrected based on a result of the distributed optimization calculation, a Lyapunov function is constructed based on the result of the distributed optimization calculation, an update equation of controller parameters is designed according to the Lyapunov function, and the update equation is solved through the Kalman gain matrix; An optimal control sequence is generated based on the controller parameters, the optimal control sequence is optimized by using the gradient descent method, and the optimal control sequence is used for online control of the system.

[0011] A distributed optimization calculation is performed on an output result of the radial basis function neural network, parameters of the recursive least square algorithm are corrected based on a result of the distributed optimization calculation, a Lyapunov function is constructed based on the result of the distributed optimization calculation, an update equation of controller parameters is designed according to the Lyapunov function, and the update equation is solved through the Kalman gain matrix; System input and output data are input into the radial basis function neural network for processing, the radial basis function neural network performs nonlinear mapping on the system input and output data through network weights, center vectors and expansion constants to obtain a neural network output result; A distributed optimization network is constructed based on the neural network output result, the distributed optimization network adopts a graph topology structure to connect multiple computing nodes, each computing node independently calculates a local gradient based on neighborhood information, and transmits local gradient information between the computing nodes through a consistency protocol, and the neural network output result is iteratively optimized according to the local gradient and the local gradient information; The parameter of the recursive least square algorithm is corrected according to the neural network output result, and the parameter correction process comprises: constructing a parameter error covariance matrix based on the neural network output result, calculating a Kalman gain by using the parameter error covariance matrix, updating a parameter estimation value according to the Kalman gain, and reconstructing the parameter error covariance matrix based on the parameter estimation value; A Lyapunov function is constructed by using the reconstructed parameter error covariance matrix, and the construction process of the Lyapunov function comprises: mapping the neural network output result and the reconstructed parameter error covariance matrix to an error space to obtain a state error vector, constructing a quadratic function of the state error vector to obtain a system error term, constructing a parameter error term for the parameter estimation error of the recursive least square algorithm, and combining the system error term and the parameter error term to form a complete Lyapunov function.

[0012] In a second aspect, the embodiment of the present application provides an electronic device, comprising: a processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described above.

[0013] In a third aspect, the embodiment of the present application provides a computer readable storage medium having computer program instructions stored thereon, wherein the computer program instructions are executed by a processor to implement the method described above.

[0014] The beneficial effects of the present application are as follows: By constructing a voltage difference compensation model based on a graph neural network, the voltage difference between perovskite sub-cells and crystalline silicon sub-cells can be accurately captured, and adaptive dynamic compensation of voltage can be realized by using a message passing mechanism and an attention mechanism, thereby effectively solving the problem of insufficient compensation accuracy of traditional methods under complex lighting conditions.

[0015] Based on the reinforcement learning and the priority experience replay mechanism, the system can continuously optimize the control strategy and adaptively learn the optimal control parameters, thereby significantly improving the accuracy and response speed of voltage matching, and enabling the photovoltaic module to maintain the best working state under different environmental conditions.

[0016] The recursive least square algorithm is used to identify system parameters in real time and solve the optimal control sequence in a rolling time domain, and the controller parameters are dynamically adjusted by using an adaptive law, so that the entire control system has strong robustness and adaptive ability, and the energy conversion efficiency and service life of the perovskite / crystalline silicon stacked photovoltaic module are effectively improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1A flowchart of a voltage self-adaptive matching control method for a perovskite / crystalline silicon tandem photovoltaic module according to an embodiment of the present application; Figure 2 A comparison diagram of enhanced effects of a Hodgkin-Huxley neuron model according to an embodiment of the present application; Figure 3 A performance comparison histogram of an optimization control strategy for a perovskite / crystalline silicon twin cell according to an embodiment of the present application; Figure 4 A flowchart of an optimization control strategy based on a recursive least square algorithm according to an embodiment of the present application; Figure 5 A distributed optimization efficiency comparison histogram of a radial basis function neural network according to an embodiment of the present application; Figure 6 (a) is a structural design diagram of a mechanical stacking photovoltaic module by a four-terminal perovskite sub-module and a crystalline silicon sub-module through a guide rail and a clamp according to an embodiment of the present application, N1 (black minus sign) and P1 (red plus sign) represent a negative electrode and a positive electrode of the perovskite sub-module respectively, and N2 (black minus sign) and P2 (red plus sign) represent a negative electrode and a positive electrode of the crystalline silicon sub-module respectively; Figure 6 (b) is a structural design diagram of a four-terminal perovskite sub-cell string and a crystalline silicon sub-cell string packaged into a tandem photovoltaic module by a film according to an embodiment of the present application, N1 (black minus sign) and P1 (red plus sign) represent a negative electrode and a positive electrode of the perovskite sub-cell string respectively, and N2 (black minus sign) and P2 (red plus sign) represent a negative electrode and a positive electrode of the crystalline silicon sub-cell string respectively. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0019] The technical solutions of the present application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described in some embodiments.

[0020] Figure 1 A flowchart of a voltage self-adaptive matching control method for a perovskite / crystalline silicon tandem photovoltaic module according to an embodiment of the present application, as shown in Figure 1 the method comprises: Collecting output voltages of perovskite sub-cells and output voltages of crystalline silicon sub-cells in a perovskite / crystalline silicon stacked photovoltaic module, and calculating real-time voltage difference values of the output voltages of the perovskite sub-cells and the output voltages of the crystalline silicon sub-cells; A voltage difference compensation model based on a graph neural network is constructed, the perovskite sub-cells and the crystalline silicon sub-cells are taken as graph network nodes, the voltage difference values are taken as edge features, the node states are iteratively updated through a message passing mechanism, and the node weights are adaptively adjusted in combination with an attention mechanism to output optimal voltage compensation values; The output voltages of the perovskite sub-cells and the output voltages of the crystalline silicon sub-cells are compensated according to the optimal voltage compensation values to obtain compensated perovskite sub-cell voltages and compensated crystalline silicon sub-cell voltages; A system state space model is constructed based on the compensated perovskite sub-cell voltages and the compensated crystalline silicon sub-cell voltages, an integrated reward signal is generated based on the system state space model, an optimized control strategy is obtained by updating state value estimation using a priority experience replay mechanism, the optimized control strategy identifies parameters of the system state space model in real time based on a recursive least squares algorithm, and solves the optimized control sequence in a rolling time domain, while introducing an adaptive law to dynamically adjust controller parameters; According to the output control quantity of the optimized control strategy, the working states of the drive circuits of the output voltages of the perovskite sub-cells and the output voltages of the crystalline silicon sub-cells are adjusted to realize optimal voltage matching.

[0021] In an optional implementation, constructing a voltage difference compensation model based on a graph neural network includes: The perovskite sub-cells and the crystalline silicon sub-cells are constructed as graph network nodes to generate node state vectors; the output voltage difference values of the perovskite sub-cells and the crystalline silicon sub-cells are calculated, the output voltage difference values are taken as edge features, an initial edge feature representation is constructed using a Hodgkin-Huxley neuron model, and the initial edge feature representation is dynamically represented by membrane capacitance, ion channel conductance and reversal potential; The initial edge feature representation is input into an ion channel module of the Hodgkin-Huxley neuron model, the initial edge feature representation is enhanced based on the characteristics of sodium ion channels and potassium ion channels of the ion channel module to form an enhanced edge feature representation with long-range correlation; construct a message passing function based on the enhanced edge feature representation, the message passing function receiving the node state vector and the enhanced edge feature representation as input, generating node state update information; input the node state update information into a synaptic plasticity module of the Hodgkin-Huxley neuron model, and adaptively adjust the node state update information through a dynamic threshold mechanism of the synaptic plasticity module to obtain an adjusted node state; apply an attention mechanism to the adjusted node state, calculate the node state similarity and perform normalization processing to obtain an inter-node attention coefficient; input the perovskite sub-cell node state and the crystalline silicon sub-cell node state with the node attention coefficient into a multi-layer perceptron to obtain an optimal voltage compensation value.

[0022] The voltage difference compensation model based on the graph neural network first constructs perovskite sub-cells and crystalline silicon sub-cells as graph network nodes. For each perovskite sub-cell node, a 32-dimensional vector is used to represent its features, including short-circuit current density (20 mA / cm 2 ), open-circuit voltage (1.1 V), fill factor (0.75), spectral response range (300-800 nm), temperature coefficient (-0.2% / ℃), and other parameters. Similarly, for each crystalline silicon sub-cell node, a 32-dimensional vector is used to represent its features, including short-circuit current density (42 mA / cm 2 ), open-circuit voltage (0.7 V), fill factor (0.82), spectral response range (400-1100 nm), temperature coefficient (-0.3% / ℃), and other parameters. These parameters are obtained through a real-time acquisition system with a sampling frequency of 1 Hz, and are normalized to scale all feature values to the [-1, 1] interval.

[0023] The output voltage difference between the perovskite sub-cell and the crystalline silicon sub-cell is calculated as an edge feature. For example, under standard test conditions (1000 W / m² illumination, 25℃ temperature), the output voltage of the perovskite sub-cell is 1.05 V, the output voltage of the crystalline silicon sub-cell is 0.68 V, and the voltage difference is 0.37 V. This voltage difference dynamically changes with environmental conditions and is characterized by the Hodgkin-Huxley neuron model.

[0024] The Hodgkin-Huxley neuron model processes the voltage difference signal by simulating the electrophysiological characteristics of biological neurons. The model includes a membrane capacitance parameter set to 1.0 μF / cm 2 , a maximum sodium channel conductance of 120 mS / cm 2 , a maximum potassium channel conductance of 36 mS / cm 2 , and a leakage current conductance of 0.3 mS / cm 2The sodium ion reversal potential is set to 50 mV, the potassium ion reversal potential is set to -77 mV, and the leakage current reversal potential is -54.4 mV. These parameters are optimized through a grid search method to minimize the root mean square error of the model on the validation set.

[0025] The initial edge feature represents the ion channel module of the Hodgkin-Huxley neuron model, which simulates the switching dynamics of sodium and potassium channels. When the input voltage difference exceeds the threshold value (set to 0.2V), the sodium channel is rapidly activated and reaches a peak (about 1.5ms), followed by the activation of the potassium channel (about 5ms). This timing characteristic enables the model to capture the rapid change characteristics of the voltage difference. In practical applications, when the light intensity increases from 600W / m 2 to 900W / m 2 , the voltage difference changes from 0.25V to 0.33V, and the feature representation after ion channel processing can reflect the time dynamic characteristics of this change, forming a 16-dimensional enhanced edge feature representation.

[0026] Based on the enhanced edge feature representation, a message passing function is constructed, which receives the node state vector and the enhanced edge feature representation as input. The specific implementation adopts a three-layer fully connected network, with an input layer dimension of (32+16)=48, a hidden layer dimension of 64, and an output layer dimension of 32. The activation function uses ReLU. This message passing process is iteratively executed 3 times, and the node state information is updated each iteration. In actual tests, the message passing process can effectively fuse the node features under different light conditions (200-1200W / m 2 ) and temperature conditions (10-45℃), enhancing the model's adaptability to environmental changes.

[0027] The node state update information is adaptively adjusted through the synaptic plasticity module of the Hodgkin-Huxley neuron model. This module implements a dynamic threshold mechanism, with an initial threshold set to 0.5 and dynamically adjusted according to the historical input signal intensity. When continuously receiving strong signals (voltage difference > 0.3V), the threshold is raised to 0.65; when receiving weak signals (voltage difference < 0.1V), the threshold is lowered to 0.35. This mechanism simulates the adaptability of biological neurons, making the model more stable at different operating points. For example, under weak light conditions (200W / m 2 ), the model can respond sensitively to small voltage differences (0.08V); while under strong light conditions (1200W / m 2 ), it can avoid overreaction to large voltage differences (0.45V).

[0028] The attention mechanism is applied to the adjusted node state to calculate the state similarity between the perovskite sub-cell node and the crystalline silicon sub-cell node. The similarity calculation adopts dot product operation and is normalized by a Softmax function to obtain an attention coefficient. In a typical scenario, when the perovskite sub-cell temperature rises to 40℃ and the crystalline silicon sub-cell remains at 25℃, the attention mechanism automatically gives the perovskite sub-cell node a weight of 0.65 and the crystalline silicon sub-cell node a weight of 0.35, reflecting their different contributions to the final compensation value.

[0029] Finally, the perovskite sub-cell node state and the crystalline silicon sub-cell node state with the attention coefficient are connected and input into a multi-layer perception. The multi-layer perception includes three layers, with an input layer of 64 dimensions, an intermediate layer of 32 dimensions, and an output layer of 1 dimension, and the activation function adopts a combination of Tanh and ReLU. The model output represents the optimal voltage compensation value, ranging from [-0.5V, 0.5V]. In actual tests, the model can increase the output power of the stacked cell system by 8.2% to 13.5%, which is 5.3 percentage points higher than the traditional fixed compensation method, and when the environmental conditions change rapidly (such as cloud cover causing the light to drop from 900W / m 2 to 400W / m 2 ), the response time is less than 150ms, meeting the real-time compensation requirements.

[0030] Figure 2 Comparison diagram of Hodgkin-Huxley neuron model enhancement effect for embodiments of the present application: The figure shows the performance comparison of three different methods in the feature distance and time response two dimensions. The left figure shows the decay trend of the feature distance with the increase of the dimension, and the technical solution (GNN+HH model) maintains at 0.63 when the feature distance is 4, which is obviously better than the 0.32 of the traditional GCN method and the 0.25 of the no-membrane sub-channel enhancement method, indicating that the scheme has better feature preservation ability. The right figure shows the time response characteristics of different ions, in which the peak response of sodium ion (Na+) reaches 0.75 at 7.5ms, while the response peak of the traditional GCN method is only 0.65. The technical solution shows faster response speed and higher response intensity throughout the time sequence, especially in the 15-20ms time window, its response curve smoothness and stability are better than the comparison methods. These results fully demonstrate the significant advantages of the technical solution of fusing GNN and HH model in feature extraction ability and dynamic response performance, providing a more accurate and efficient solution for ion channel dynamics modeling.

[0031] In an alternative embodiment, the output voltage difference between the perovskite sub-cell and the crystalline silicon sub-cell is calculated, and the output voltage difference is taken as an edge feature. An initial edge feature representation is constructed using a Hodgkin-Huxley neuron model, including: The voltage difference between the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell is calculated. The voltage difference is multiplied by a mapping coefficient and superimposed with a resting potential to obtain a mapping membrane potential. The mapping coefficient is determined by the amplitude range of the voltage difference and the physiological range of the membrane potential of biological neurons. Based on the mapping membrane potential, an ion channel dynamics model is constructed. The time rate of change of the mapping membrane potential is calculated by the membrane capacitance per unit area, the maximum conductance of each ion channel, the equilibrium potential of each ion, and the external input current. Based on the ion channel dynamics model, a gating variable dynamics equation is established. The voltage-dependent rate constant of each gating variable is calculated based on the mapping membrane potential. The voltage-dependent rate constant includes the forward rate constant and the reverse rate constant. The steady-state value and the time constant of each gating variable are calculated based on the forward rate constant and the reverse rate constant. The steady-state value and the time constant are substituted into the exponential form of the dynamics equation to solve the time evolution of each gating variable. The time evolution is numerically integrated by the Runge-Kutta method to form an initial edge feature representation.

[0032] The output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell are obtained. For example, under certain light conditions, the output voltage of the perovskite sub-cell is 0.9 volts, and the output voltage of the crystalline silicon sub-cell is 0.6 volts. The system calculates the voltage difference between the two, which is 0.9 volts minus 0.6 volts, resulting in a voltage difference of 0.3 volts.

[0033] To map the voltage difference into the membrane potential range of biological neurons, the system needs to determine a mapping coefficient. The membrane potential of biological neurons usually varies between -70 millivolts and +30 millivolts, with a total range of 100 millivolts. In this embodiment, the amplitude range of the voltage difference is -0.5 volts to +0.5 volts. Therefore, the mapping coefficient can be set to 100 millivolts divided by 1 volt, i.e. 100 millivolts / volt. The system multiplies the voltage difference of 0.3 volts by the mapping coefficient of 100 millivolts / volt to obtain 30 millivolts, and then superimposes the neuron resting potential of -70 millivolts to finally obtain the mapping membrane potential of -40 millivolts.

[0034] Based on the mapping membrane potential, the system constructs an ion channel kinetics model. In this model, the time rate of change of the membrane potential is determined by multiple factors, including the membrane capacitance per unit area, the maximum conductance of each ion channel, the equilibrium potential of each ion, and the external input current. In this embodiment, the membrane capacitance per unit area is set to 1 microfarad per square centimeter, the maximum conductance of the sodium ion channel is 120 millisiemens per square centimeter, the maximum conductance of the potassium ion channel is 36 millisiemens per square centimeter, and the maximum conductance of the leakage current channel is 0.3 millisiemens per square centimeter. The equilibrium potentials of the ions are as follows: the equilibrium potential of the sodium ion is +50 millivolts, the equilibrium potential of the potassium ion is -77 millivolts, and the equilibrium potential of the leakage current is -54.387 millivolts. The external input current can be set according to the actual application scenario, and in this example it is set to 0 microampere per square centimeter.

[0035] The system establishes a gating variable kinetics equation based on the ion channel kinetics model. The sodium ion channel is controlled by two gating variables m and h, and the potassium ion channel is controlled by one gating variable n. For a mapping membrane potential of -40 millivolts, the system calculates the voltage-dependent rate constants of each gating variable. For example, the forward rate constant a m of the m gating variable is 0.1 millivolt (-1) ×(25-(-40)) / [exp((25-(-40)) / 10)-1]=0.1 millivolt (-1) ×65 / [exp(6.5)-1]≈0.1×65 / 665≈0.00977 milliseconds (-1) ; the reverse rate constant b m is 4×exp(-(-40) / 18)=4×exp(40 / 18)≈4×9.2≈36.8 milliseconds (-1) . Similarly, the forward and reverse rate constants of the h and n gating variables are calculated.

[0036] Based on the forward and reverse rate constants, the system calculates the steady-state value and time constant of each gating variable. The steady-state value m ∞ of the m gating variable is a m / (a m +b m )=0.00977 / (0.00977+36.8)≈0.000265; the time constant t m is 1 / (a m +b m )=1 / (0.00977+36.8)≈0.027 milliseconds. The steady-state value and time constant of the h and n gating variables are calculated in the same way. In actual implementation, the steady-state value h ∞ of the h gating variable is approximately 0.596, and the time constant t h is approximately 8.04 milliseconds; the steady-state value n ∞ of the n gating variable is approximately 0.318, and the time constant tn about 1.11 milliseconds.

[0037] The system substitutes the steady-state values and time constants into the exponential form of the kinetic equations to solve the time evolution of each gating variable. The kinetic equations are in exponential form, describing how the gating variable approaches its steady-state value over time. The system solves these equations by numerical integration using the Runge-Kutta method. In implementation, the system chooses the fourth-order Runge-Kutta method with a time step of 0.01 milliseconds and simulates for a total duration of 50 milliseconds.

[0038] Taking the m gating variable as an example, suppose the initial value m (0) = 0.05, the system calculates the value at t = 0.01 milliseconds by the Runge-Kutta method. First, calculate four intermediate values: k1 = f(m (0) ) = (m ∞ -m (0) ) / τ m = (0.000265-0.05) / 0.027≈-1.844; k2 = f(m (0) +k1×0.01 / 2); k3 = f(m (0) +k2×0.01 / 2); k4 = f(m (0) +k3×0.01). Then calculate m (0.01) = m (0) +(k1+2k2+2k3+k4)×0.01 / 6. In this way, the system calculates the complete time series of the m, h, and n gating variables within 50 milliseconds.

[0039] Through the above steps, the system constructs an initial edge feature representation based on the output voltage difference between the perovskite sub-cell and the crystalline silicon sub-cell using the Hodgkin-Huxley neuron model. This feature representation is essentially the complete sequence of the evolution of the gating variables m, h, and n over time, capturing the neuron dynamics characteristics induced by the input voltage difference. This biologically inspired feature representation method enhances the system's sensitivity to voltage differences, providing rich dynamic features for subsequent anomaly detection and pattern recognition.

[0040] In an alternative embodiment, a system state space model is constructed based on the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage, a comprehensive reward signal is generated based on the system state space model, and an optimized control strategy is obtained by updating the state value estimation using the priority experience replay mechanism, which includes: A system state space model of voltage states of the perovskite sub-cell and voltage states of the crystalline silicon sub-cell is constructed, a prediction cost function is constructed according to the system state space model; an initial control strategy is calculated based on the prediction cost function, and an Actor-Critic agent is configured for the perovskite sub-cell and the crystalline silicon sub-cell respectively according to the initial control strategy, and a state expected return of the Actor network of the Actor-Critic agent is calculated based on a state action value function. The state expected return is input into a hierarchical reward mechanism, the hierarchical reward mechanism calculates a local reward function according to a voltage tracking error of a single sub-cell, calculates a global reward function according to a joint tracking error of the voltages of the two sub-cells, calculates a timing reward function according to a time derivative of the voltage difference, and combines the local reward function, the global reward function and the timing reward function to form a comprehensive reward signal. A priority experience replay mechanism is constructed using the comprehensive reward signal, the priority experience replay mechanism calculates a policy gradient through a timing difference error, and simultaneously updates a state value estimation of the Actor-Critic agent using the policy gradient to obtain an optimized control strategy.

[0041] The specific implementation of constructing a system state space model based on the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage, then generating a comprehensive reward signal and updating a state value estimation using a priority experience replay mechanism to obtain an optimized control strategy is as follows.

[0042] In the hybrid tandem solar cell system, first, the original voltage values of the perovskite sub-cell and the crystalline silicon sub-cell are obtained, denoted as V perovskite original and V silicon original respectively. The original voltages of the two sub-cells are subjected to temperature compensation and irradiance compensation. Temperature compensation is performed by correcting the voltage with a temperature coefficient k temp When a difference is detected between the ambient temperature T and the standard test condition temperature T STC (25℃), the compensated voltage V temp comp =V original ×(1+k temp ×(T-T STC )), where the temperature coefficient k temp perovskite of the perovskite sub-cell is -0.0035 / ℃, and the temperature coefficient k temp silicon of the crystalline silicon sub-cell is -0.0025 / ℃. Irradiance compensation is performed by adjusting the voltage value based on an irradiance coefficient k irr When the actual irradiance G is different from the standard test condition irradiance G STC (1000W / m 2 ), the compensated voltage V irr comp =V temp compx (1 + k irr x log(G / G STC ), where the irradiance coefficient k irr perovskite of the perovskite sub-cell is 0.015, and the irradiance coefficient k irr silicon of the crystalline silicon sub-cell is 0.012.

[0043] After the compensation processing is completed, the compensated perovskite sub-cell voltage V perovskite comp and the compensated crystalline silicon sub-cell voltage V silicon comp are obtained. Based on the two compensated voltage values, a system state space model is constructed. The model takes the compensated sub-cell voltages as state variables, forming a state vector s t = [V perovskite comp , V silicon comp ]. At the same time, the adjustment amount of the maximum power point tracking is defined as an action vector a t = [ΔD perovskite , ΔD silicon ], where ΔD represents the adjustment amount of the corresponding sub-cell duty cycle. The system state transition relationship can be described as s (t+1) = f(s t , a t ), where f represents the system dynamic characteristics. In actual application, when the environmental temperature is 30°C and the irradiance is 800 W / m², the original voltage of the perovskite sub-cell at a certain moment is 0.85V, and after compensation, it is 0.827V; the original voltage of the crystalline silicon sub-cell is 0.62V, and after compensation, it is 0.607V.

[0044] Based on the constructed system state space model, a prediction cost function J is constructed next. The cost function takes into account the voltage tracking error, power extraction efficiency and control action amplitude, and is expressed as a function of state s and action a. The prediction cost function calculation considers the cumulative cost in the future N time steps, and the time step is selected as 100ms. By minimizing the cost function, the initial control strategy π initial is calculated. In the example, when the compensated perovskite sub-cell voltage is 0.827V, its maximum power point voltage target value is 0.85V, and the compensated crystalline silicon sub-cell voltage is 0.607V, its maximum power point voltage target value is 0.62V, the action given by the initial control strategy is ΔD perovskite = 0.015, ΔD silicon = 0.008.

[0045] According to the initial control strategy, the Actor-Critic agent is configured for the perovskite and crystalline silicon sub-cells respectively. Each agent contains two parts of Actor network and Critic network. The Actor network is responsible for outputting actions according to the current state, and its network structure contains two hidden layers, each layer contains 64 neurons, using ReLU activation function. The input layer receives the voltage state of the corresponding sub-cell, and the output layer uses the tanh activation function to generate the duty cycle adjustment amount in the range of-0.05 to 0.05. The Critic network is responsible for evaluating the value of the state or state-action pair, and its structure also contains two hidden layers, each layer contains 64 neurons, using ReLU activation function. The Actor network calculates the state expected return V(s) based on the state action value function Q(s, a), which is in the form of weighted average of Q values for all actions according to the policy probability.

[0046] The state expected return is input into the hierarchical reward mechanism, which consists of three reward functions. The local reward function R local According to the voltage tracking error calculation of a single sub-cell, when the voltage tracking error is less than the threshold value 0.02V, a positive reward is given, otherwise a negative reward is given, the value range is [-1, 1]. The global reward function R global According to the joint tracking error calculation of the voltages of the two sub-cells, the overall system performance is measured, the value range is [-2, 2]. The time reward function R temporal According to the time derivative calculation of the voltage difference, it reflects the stability of the voltage adjustment, when the voltage change rate is less than the threshold value 0.01V / s, a positive reward is given, otherwise a negative reward is given, the value range is [-0.5, 0.5]. The comprehensive reward signal R local is calculated by combining the three parts, which is R global = w1 x R temporal + w2 x R temporal + w3 x R temporal , where the weight parameters w1=0.3, w2=0.5, w3=0.2.

[0047] The priority experience replay mechanism is constructed using the comprehensive reward signal. This mechanism maintains an experience pool with a capacity of 10000, which stores transition samples in the form of (s t , a t , r t , s (t+1) ). Each sample is assigned a priority according to its time difference error TD error = r t + γ x V(s (t+1) )- (s_t) , γ is the discount factor, which is 0.95. The priority p i = |TD error | + ε, ε is a small constant 0.01, which ensures that all samples have a non-zero probability of being selected. The sampling probability is proportional to the priority, and the samples are selected according to pi α / ∑p j α The calculation is that a is a priority factor, and the value is 0.6. The samples with a batch size of 128 are sampled according to the priority from the experience pool for training. The Actor network parameters are updated by calculating the policy gradient, and the Critic network parameters are updated by minimizing the TD error. The learning rate is set to 0.001, and the target network is updated once every 100 time steps.

[0048] After 50000 iterations of training, the system can quickly track the maximum power point under dynamic environmental conditions. Experimental results show that when using this method, the average voltage tracking error of the perovskite sub-cell is 0.012V, and the average voltage tracking error of the crystalline silicon sub-cell is 0.009V, which is reduced by 37% and 42% respectively compared with the traditional method, and the overall system efficiency is improved by 4.3%.

[0049] Figure 3 The bar chart for comparing the performance of the perovskite / crystalline silicon dual sub-cell optimization control strategy of the embodiment of the application is as follows: The figure shows the comparison results of the technical scheme (comprehensive reward + priority replay) and the traditional DQN control method and the standard Actor-Critic method in four key performance indicators. In terms of voltage tracking accuracy, the technical scheme reaches 90.5%, which is significantly higher than the 73.2% of the traditional DQN and the 62.3% of the standard Actor-Critic; in terms of convergence speed, the technical scheme performs best, reaching 95.4%, while the traditional DQN is 67.5% and the standard Actor-Critic is only 55.2%; in terms of robust performance, the technical scheme reaches 87.3%, the traditional DQN is 75.1%, and the standard Actor-Critic is 68.4%; in terms of energy efficiency improvement, the technical scheme also performs well, reaching 93.6%, the traditional DQN is 77.2%, and the standard Actor-Critic is 65.3%. Through data comparison, it can be seen that the technical scheme of fusing the comprehensive reward signal and the priority experience replay has achieved significant advantages in all performance indicators, especially in the improvement of convergence speed and energy efficiency, fully proving the advancement and practical value of the technical scheme. These performance improvements are mainly due to the comprehensive evaluation of the system state by the comprehensive reward mechanism and the efficient use of key experiences by the priority replay.

[0050] In an optional embodiment, a priority experience replay mechanism is constructed using the comprehensive reward signal, the priority experience replay mechanism calculates the policy gradient through the temporal difference error, and the state value estimation of the Actor-Critic intelligent agent is updated using the policy gradient to obtain an optimized control strategy, including: The fractal dimension of the strategy trajectory calculated based on the comprehensive reward signal is calculated, an experience priority evaluation function is constructed based on the fractal dimension, the time series difference error and the state transition similarity, and the experience priority evaluation function is used to determine the importance of the experience sample; The Logistic chaotic sequence is introduced into the time series difference error calculation process, the Logistic chaotic sequence is modulated by a chaotic intensity parameter to obtain a discounted state value, and a chaotic enhanced time series difference error is obtained, which is used to evaluate the time correlation of the state action pair; Based on the chaotic enhanced time series difference error and the fractal dimension, a strategy gradient is constructed, the strategy gradient is input into the Actor network for parameter update, and the chaotic enhanced time series difference error and the chaotic modulation term are input into the Critic network to update the state value estimation, and the learning rate of the Actor network and the Critic network is adaptively adjusted according to the training process; Based on the updated state value estimation, a final optimized control strategy is constructed, the final optimized control strategy adjusts the exploration degree of action selection through a temperature parameter, and realizes a probabilistic policy output based on the action value function.

[0051] The priority experience replay mechanism calculates the strategy gradient through the time series difference error, and simultaneously updates the state value estimation of the Actor-Critic intelligent agent using the strategy gradient, thereby obtaining the optimized control strategy.

[0052] When calculating the fractal dimension of the strategy trajectory based on the comprehensive reward signal, the system first collects the strategy trajectory data in the interaction process of the intelligent agent and the environment, including the state sequence, the action sequence and the reward sequence. The box counting method is applied to the collected reward sequence for fractal dimension calculation, the reward sequence is divided into n equal size subintervals, and the number of data points in each subinterval is counted. By calculating the relationship between ln(N (ε) ) and ln(1 / ε), where N (ε) represents the number of non-empty subintervals, and ε represents the subinterval size, the fractal dimension D value can be obtained. In actual application, when the reward sequence shows high complexity, the fractal dimension D value is close to 1.8; when the reward sequence is relatively regular, the D value is close to 1.2.

[0053] An experience priority evaluation function is constructed based on the calculated fractal dimension, time difference error and state transition similarity. The function considers three factors: fractal dimension reflects the complexity of state space, time difference error reflects the size of prediction error, and state transition similarity reflects the similarity between current state transition and historical experience. The specific implementation of the experience priority evaluation function is a weighted combination of the three, with the fractal dimension weight being 0.3, the time difference error weight being 0.5, and the state transition similarity weight being 0.2. In practical applications, when the evaluation function value of a certain experience sample exceeds the threshold value 0.75, the sample is given a high priority; the evaluation function value is between 0.4 and 0.75, which is a medium priority; and lower than 0.4 is a low priority.

[0054] When the Logistic chaotic sequence is introduced into the time difference error calculation process, the system first generates a Logistic chaotic sequence. The sequence is generated by iterative calculation, with the initial value set to 0.4 and the control parameter value set to 3.9 to ensure that the system is in a chaotic state. The generated chaotic sequence is used to modulate the discount state value. The specific modulation method is to multiply the chaotic sequence value by the discount coefficient, and the chaotic intensity parameter is set to 0.15. The modulated discount coefficient fluctuates between 0.8 and 0.98, introducing nonlinear dynamic characteristics. In practical applications, this chaotic modulation significantly improves the adaptability of the algorithm in non-stationary environments, increasing the convergence speed by 25%.

[0055] When calculating the chaotic enhanced time difference error, the system combines the current reward, the chaotic modulated discount state value and the current state value estimate. In an actual application case, the immediate reward in a certain state is 8.5, the current state value estimate is 15.2, the next state value estimate is 20.3, and the chaotic modulated discount coefficient is 0.87. The calculated chaotic enhanced time difference error is 10.16. Compared with the traditional time difference error of 8.76, the chaotic enhanced version shows greater volatility and exploration.

[0056] When constructing the policy gradient based on the chaotic enhanced time difference error and the fractal dimension, the two are weighted and combined, with the fractal dimension weight being 0.25 and the time difference error weight being 0.75. The constructed policy gradient is used for Actor network parameter update, with the initial learning rate set to 0.003. At the same time, the chaotic enhanced time difference error and the chaotic modulation term are input into the Critic network to update the state value estimate, with the initial learning rate of the Critic network being 0.005. The learning rates of the two networks are adaptively adjusted according to the training process. When the cumulative reward change rate of the last 5 rounds is less than 3%, the learning rate is reduced to 0.85 times of the original; when the change rate is greater than 10%, the learning rate is increased to 1.1 times of the original, but the maximum is not more than 2 times of the initial learning rate, and the minimum is not less than 0.1 times of the initial learning rate.

[0057] When constructing the final optimization control strategy based on the updated state value estimation, the system adjusts the exploration degree of action selection through the temperature parameter. The initial value of the temperature parameter is set to 1.0, and gradually decreases to 0.2 as the training progresses, with a reduction rate of 0.05 per 1000 training steps. When the temperature value is low (close to 0.2), the system tends to select the action with the highest estimated value; when the temperature value is high (close to 1.0), the system tends to explore more. In practical applications, when the temperature parameter is 0.8, the probability ratio of selecting the optimal action and the suboptimal action is about 3:1; when the temperature parameter decreases to 0.3, the ratio increases to about 9:1.

[0058] The entire optimization control strategy realizes probabilistic policy output through the action value function, and the action selection probability is proportional to the action value. In the early stage of training, the system has strong exploration, and the temperature parameter is high, resulting in a relatively flat action selection distribution; as the training progresses, the temperature parameter decreases, and the action selection gradually concentrates on high-value actions. Finally, when the cumulative reward of the last 50 rounds is stable at more than 95% of the target value, it is considered that the policy optimization is complete.

[0059] In an optional implementation, the optimization control strategy identifies the parameters of the system state space model in real time based on a recursive least squares algorithm, and solves the optimization control sequence in a rolling time domain while introducing an adaptive law to dynamically adjust the controller parameters, including: Collecting system input and output data, an initial parameter estimation model is constructed using a recursive least squares algorithm, which iteratively updates the system parameter vector through a Kalman gain matrix; the system input and output data are decomposed into different frequency band features through wavelet decomposition, and the different frequency band features are combined with wavelet basis functions to obtain multi-scale signal representations; Based on the multi-scale signal representations, parameter identification models are constructed in each frequency band, which use the iterative structure of the recursive least squares algorithm and adjust the Kalman gain matrix according to the frequency band characteristics; The identification results of the parameter identification model are input into a radial basis function neural network, the network parameters of which are adjusted by a gradient descent method, and the update step of the gradient descent method is calculated based on the identification error; The output results of the radial basis function neural network are subjected to distributed optimization calculation, and the parameters of the recursive least squares algorithm are corrected based on the results of the distributed optimization calculation; a Lyapunov function is constructed based on the results of the distributed optimization calculation, and an update equation for the controller parameters is designed according to the Lyapunov function, which is solved through the Kalman gain matrix; generate an optimal control sequence based on the controller parameters, the optimal control sequence being optimized by the gradient descent method and used for online control of the system.

[0060] As shown in Figure 4 The method further comprises: Collecting system input and output data, including sensor measurements and control input signals. Take a temperature control system as an example, collect the measurement values of the temperature sensor as the system output, and the heater power set value as the system input. The sampling period is set to 0.1 seconds, and 300 data points are continuously collected to form an initial data set.

[0061] An initial parameter estimation model is constructed using the recursive least squares algorithm. This algorithm iteratively updates the system parameter vector through the Kalman gain matrix. In actual operation, the system order is selected to be 2, the forgetting factor is set to 0.98, the initial covariance matrix diagonal element is set to 100, and the initial parameter vector is set to a zero vector. Take the temperature control system as an example, apply the recursive least squares algorithm to the collected data to obtain the initial estimated value of the system parameters [0.82, -0.15, 0.67, 0.08]. These parameters describe the dynamic characteristics of the system.

[0062] The system input and output data is processed by wavelet decomposition. Using db4 wavelet basis function, the system input and output data is decomposed to 3 levels to obtain different frequency band features. The decomposed signals include low frequency approximation components and high frequency detail components. After wavelet decomposition of the input and output data of the temperature control system, the low frequency approximation component A3 and the high frequency detail components D1, D2, D3 are obtained. These different frequency band features are combined with the wavelet basis function to form a multi-scale signal representation of the system.

[0063] Based on the multi-scale signal representation, parameter identification models are constructed in each frequency band. The recursive least squares algorithm is applied to the data of each frequency band. The low frequency component A3 uses a forgetting factor of 0.99, the medium frequency components D3 and D2 use a forgetting factor of 0.97, and the high frequency component D1 uses a forgetting factor of 0.95. The Kalman gain matrix of each frequency band is adjusted according to the frequency band characteristics, with the high frequency component increasing the gain coefficient and the low frequency component decreasing the gain coefficient. The temperature control system identifies the parameters in the low frequency band A3 as [0.85, -0.12, 0.70, 0.06], in the medium frequency band D3 as [0.25, 0.35, 0.48, 0.12], in the medium frequency band D2 as [0.18, 0.22, 0.38, 0.15], and in the high frequency band D1 as [0.08, 0.05, 0.12, 0.04].

[0064] The identification results of the parameter identification model are input into a radial basis function neural network. The network contains 10 hidden layer neurons, the radial basis function is selected as a Gaussian function, and the width parameter is initially set to 0.8. The network parameters are adjusted by the gradient descent method, and the initial learning rate is set to 0.05. In practical applications, the update step of the gradient descent method is dynamically calculated according to the identification error. When the identification error is greater than 0.1, the update step is increased to 0.08, and when the identification error is less than 0.01, the update step is reduced to 0.02. Through this method, the parameters of the temperature control system are further optimized, and the network output parameters are [0.84, -0.13, 0.69, 0.07], with an error of less than 3% compared with the actual system parameters.

[0065] The output results of the radial basis function neural network are calculated by distributed optimization. The optimization task is distributed to 3 parallel computing units, and each computing unit processes part of the parameters. The first computing unit processes the autoregressive parameters, the second computing unit processes the control parameters, and the third computing unit processes the cross-term parameters. Each computing unit uses the alternating direction multiplier method for local optimization, and the results are summarized after 20 iterations. The optimized parameters of the temperature control system are [0.83, -0.14, 0.68, 0.075], and the optimization accuracy is further improved.

[0066] Based on the distributed optimization calculation results, a Lyapunov function is constructed. A quadratic function is selected as the Lyapunov function, and the product of the system state vector and the positive definite matrix is used. The initial value of the diagonal elements of the positive definite matrix is set to [2.0, 1.5]. According to the Lyapunov function, the controller parameter update equation is designed, which is solved by the Kalman gain matrix. During the update process, the adjustment amplitude of the controller parameters is limited to within 10% of the original value, ensuring the stability of the system. The controller parameters of the temperature control system are updated from the initial value [0.5, 0.3] to [0.54, 0.28].

[0067] Based on the updated controller parameters, an optimal control sequence is generated. The prediction time domain length is set to 10 steps, and the control time domain length is set to 3 steps. The optimal control sequence is optimized by the gradient descent method, with an initial step size of 0.03, which gradually decreases to 0.01 as the number of iterations increases. The target temperature of the temperature control system is set to 70°C, and the current temperature is 65°C. The optimized control sequence is [85%, 75%, 60%], indicating the set values of the heater power in the three control periods. This control sequence is applied to the online control of the system, making the system output smoothly transition to the target value, with an overshoot of less than 2% and a regulation time reduced by 30% compared with traditional PID control.

[0068] Through the above method, the accurate identification of system parameters and the generation of optimal control sequence are realized, and the system control performance is significantly improved. The entire control process runs on a standard industrial computer, and the single control cycle calculation time is less than 50 milliseconds, meeting the real-time control requirements.

[0069] In an alternative embodiment, a distributed optimization calculation is performed on the output result of the radial basis function neural network, and the parameters of the recursive least squares algorithm are corrected based on the result of the distributed optimization calculation; constructing a Lyapunov function based on the result of the distributed optimization calculation comprises: The system input and output data are sent to the radial basis function neural network for processing, and the radial basis function neural network performs nonlinear mapping on the system input and output data through network weights, center vectors and expansion constants to obtain a neural network output result; A distributed optimization network is constructed based on the neural network output result, the distributed optimization network adopts a graph topology structure to connect multiple computing nodes, each computing node independently calculates a local gradient based on neighborhood information, and transmits local gradient information between the computing nodes through a consistency protocol, and iteratively optimizes the neural network output result according to the local gradient and the local gradient information; The parameters of the recursive least squares algorithm are corrected according to the neural network output result, and the parameter correction process comprises: constructing a parameter error covariance matrix based on the neural network output result, calculating a Kalman gain using the parameter error covariance matrix, updating a parameter estimate value according to the Kalman gain, and reconstructing a parameter error covariance matrix based on the parameter estimate value; A Lyapunov function is constructed using the reconstructed parameter error covariance matrix, and the construction process of the Lyapunov function comprises: mapping the neural network output result and the reconstructed parameter error covariance matrix to an error space to obtain a state error vector, constructing a quadratic function of the state error vector to obtain a system error term, constructing a parameter error term for the parameter estimation error of the recursive least squares algorithm, and combining the system error term and the parameter error term to form a complete Lyapunov function.

[0070] System input and output data are collected, including temperature, pressure, speed and other physical quantity data collected by sensors, and corresponding system response data. The collected data are preprocessed and sent to the radial basis function neural network for processing.

[0071] The radial basis function neural network includes an input layer, a hidden layer, and an output layer. The input layer receives pre-processed system input data, the hidden layer includes a plurality of radial basis function nodes, each node having a specific center vector and a spread constant. In this embodiment, Gaussian radial basis functions are used as activation functions, the network includes 50 hidden layer nodes, the initial value of the center vector is determined by the K-means clustering method, and the spread constant is set to 1.5 times the average distance between adjacent center vectors. For input data, calculate the Euclidean distance between each center vector and the node output value converted by the Gaussian radial basis function. The output layer weights the sum of the outputs of the hidden layer nodes through a weight vector to obtain the final output result of the neural network.

[0072] After obtaining the neural network output result, a distributed optimization network is constructed to optimize the result. The distributed optimization network uses a graph topology to connect multiple computing nodes, and each node represents an independent computing unit. In this embodiment, a ring topology is used to connect 10 computing nodes, and each node only communicates with two adjacent nodes. Each computing node receives part of the output result of the neural network and independently calculates the local gradient. The local gradient calculation is based on the node local data and the objective function, which is defined as the mean square error between the neural network output and the actual system output.

[0073] After calculating the local gradient, each computing node exchanges local gradient information through a consistency protocol. The consistency protocol uses a weighted average method, and each node receives the gradient information of the adjacent nodes according to the preset weight and updates its own gradient value. In specific implementation, the weight matrix is designed to ensure convergence, and the sum of the rows is 1. In each iteration step, the node combines its own gradient with the received adjacent node gradient according to the weight to generate a new gradient estimate. Through multiple iterations, the gradient estimates of each node gradually converge, thereby realizing distributed optimization.

[0074] The iterative optimization process uses an adaptive step size strategy, with an initial step size of 0.01, which is dynamically adjusted according to the gradient change rate. When the gradient change of three consecutive iterations is less than the preset threshold of 0.0001, it is determined that the optimization converges. Experiments show that in a typical scenario, the gradient difference between nodes is less than 0.0005 after 10-15 iterations, and the optimization process successfully converges.

[0075] After optimization, the neural network output result is used to correct the parameters of the recursive least squares algorithm. Parameter correction first constructs a parameter error covariance matrix P, with an initial value of 100 times the unit matrix, representing the uncertainty of the initial estimate. According to the output result of the neural network optimization, the Kalman gain K is calculated. The Kalman gain calculation involves the product of the covariance matrix P and the system regression vector, and the numerical stability is enhanced through normalization processing.

[0076] The parameter estimation value is updated by using Kalman gain, and the parameter update amount is the product of Kalman gain and prediction error. The prediction error is defined as the difference between the actual system output and the predicted output of the current parameter estimation. After the parameter is updated, the parameter error covariance matrix is reconstructed based on the new parameter estimation value. The reconstruction process uses the matrix update formula to ensure the positive definiteness of the covariance matrix, and the diagonal elements of the covariance matrix represent the uncertainty of the parameter estimation.

[0077] After the parameter correction is completed, the reconstructed parameter error covariance matrix is used to construct a Lyapunov function. The construction process first maps the neural network output result and the reconstructed parameter error covariance matrix to the error space to obtain a state error vector e. The state error vector represents the deviation between the current state of the system and the expected state. Based on the state error vector, a quadratic function V1 = e T ×P×e is constructed as the system error term of the Lyapunov function.

[0078] The parameter estimation error of the recursive least squares algorithm The parameter error term V2 = (p-p T ×Γ (-1) × is defined as the difference between the true value and the estimated value of the parameter, and the parameter error term V2 = (p-p

[0079] In practical applications, the constructed Lyapunov function is used to analyze the stability of the system. When the time derivative of the Lyapunov function satisfies the negative condition, the system is asymptotically stable. By adjusting the forgetting factor (taking values from 0.95 to 0.99) and the learning rate parameter of the recursive least squares algorithm, the negative definiteness of the Lyapunov function is ensured, thereby ensuring the stable convergence of the system. Experimental results show that in typical industrial control applications, the proposed method reduces the parameter estimation error by about 25% compared to the traditional single recursive least squares algorithm, and the system response time is shortened by about 30%, effectively improving the stability and control accuracy of the system.

[0080] Figure 5 The radial basis function neural network distributed optimization efficiency comparison bar chart for the embodiments of the present application: The figure shows the performance comparison of three different methods in the three dimensions of calculation efficiency, convergence performance and robustness. The traditional RBF algorithm reaches 72.5% in terms of calculation efficiency, the distributed optimization algorithm improves to 89.3%, and the Lyapunov correction algorithm further improves to 91.2%, indicating that the optimized algorithm has a significant improvement in calculation speed. In terms of convergence performance, the traditional RBF algorithm is 68.4%, the distributed optimization algorithm improves to 85.7%, and the Lyapunov correction algorithm reaches 93.5%, indicating that the improved algorithm has better convergence characteristics. In terms of robustness, the traditional RBF algorithm is 75.8%, the distributed optimization algorithm improves to 82.6%, and the Lyapunov correction algorithm reaches 94.7%, proving that the optimized algorithm has stronger anti-interference ability and system stability. Overall, by introducing distributed optimization and Lyapunov stability analysis, the algorithm has significantly improved in the three key performance indicators, especially the Lyapunov correction algorithm performs best, fully verifying the effectiveness of the optimization scheme.

[0081] A perovskite / crystalline silicon stacked photovoltaic module is a mechanical stack stacked photovoltaic module structure connected by guide rails and clamps, and a stacked photovoltaic module packaged as a whole by adhesive film. The biggest advantage of this mechanical stack perovskite / crystalline silicon stacked photovoltaic module design connected by guide rails and clamps is that when any sub-module fails, the failed sub-module can be easily replaced by a new good sub-module. The advantage of packaging perovskite sub-cell string and crystalline silicon sub-cell string into a whole stacked photovoltaic module by high-temperature-resistant and insulation-resistant high-transmittance plate and adhesive film is that it has an organic whole, less shading and less material, so the conversion efficiency is high and the cost is lower. Even when the perovskite sub-cell string fails, the stacked module of this structure can still work independently by using the backside crystalline silicon cell to continue generating electricity. Figure 6 A four-terminal perovskite / crystalline silicon mechanical stack photovoltaic module structure is designed, including an upper perovskite sub-module and a lower crystalline silicon sub-module. As shown in Figure 6 (a), this mechanical stack perovskite / crystalline silicon stacked photovoltaic module is independently packaged as a sub-module by perovskite sub-cell string and crystalline silicon sub-cell string, and then connected by guide rails and clamps to form a whole module, which is a mechanical connection mode. Figure 6 ​(b) The figure shows another design of perovskite / crystalline silicon tandem photovoltaic module, which is to separate perovskite sub-cell string and crystalline silicon sub-cell string with high-temperature-resistant and insulating high-transmittance plate material, and then encapsulate them into a whole module with adhesive film. The upper perovskite module adopts wide-bandgap material, with output voltage range of 300-600V; the lower crystalline silicon module adopts narrow-bandgap material, with output voltage range of 200-400V. In this structural design, the perovskite sub-cell and the crystalline silicon sub-cell independently lead out positive and negative poles, forming a four-terminal independent output structure, and realizing electrical isolation of the two sub-modules.

[0082] The intelligent current-voltage dynamic matching system mainly consists of the following functional modules: The dual-channel independent MPPT control unit adopts time-division multiplexing MPPT algorithm, switches tracking channels between perovskite and crystalline silicon sub-modules every 15ms, and realizes real-time capture of the maximum power point of the two sub-modules. The control unit integrates a hybrid logic controller based on DSP, which dynamically adjusts the MPPT tracking step according to real-time collected irradiance data. In actual tests, the tracking accuracy of this control unit under weak light conditions (irradiance <200W / m 2 ) reaches 99.2%, which is 5.7 percentage points higher than that of the traditional MPPT algorithm.

[0083] The adaptive voltage matching network sets up bidirectional DC / DC converters between the four-terminal independent output ports, realizing dynamic step-up / down voltage matching. Through PWM duty cycle adjustment (which can be step-adjusted within the range of 0.1%-99.9%), the voltage matching network dynamically aligns the perovskite module output voltage (300-600V) and the crystalline silicon module output voltage (200-400V) to the input voltage range required by the inverter (500-800V). Tests show that the conversion efficiency of this voltage matching network is ≥98.5%, and the voltage adjustment speed is ≤10ms, which can effectively cope with scenarios of rapid changes in light intensity.

[0084] The intelligent power distribution logic works based on a pre-established irradiance-current mapping model. This module preferentially outputs the power of the high-irradiance end, and the energy generated by the low-irradiance end is temporarily stored in supercapacitors before being output uniformly. Actual application data shows that this intelligent power distribution strategy improves the daily average efficiency of the system by 12-18%.

[0085] The spectral response range of the irradiance sensor is 300-1200nm, with a measurement accuracy of ±2%, which can monitor changes in environmental parameters in real time and feed back to the control system. The supercapacitor energy storage unit adopts 100F / 100V specifications, with a charge / discharge cycle life of more than 10 6 times. In the case of sudden changes in irradiance (such as from 1000W / m 2 to 500W / m 2), the system power fluctuation can be controlled within the range of ±3%. The communication module adopts a Zigbee+PLC hybrid communication mode, and the communication delay is less than 5ms, which can seamlessly adapt to the existing photovoltaic power station cascade control system.

[0086] The dynamic matching algorithm realizes that the double-end power dynamic balance time of the system is less than 50ms, and the matching error is controlled within 1.5%. When the irradiance mutates, the super capacitor buffer mechanism is used to smooth the power output, and under the condition of rainy weather, the system efficiency decay rate is reduced from 35% of the traditional design to 18%, which significantly improves the power generation efficiency under bad weather conditions.

[0087] The empirical data of the Three Gorges 50MW demonstration base in Qinghai Province show that the daily average power generation of the photovoltaic system using the technology is 23.7% higher than that of the traditional four-end component. At the same time, due to the use of the adaptive communication protocol, the system can be compatible with the existing photovoltaic power station communication architecture, and the transformation cost is reduced by more than 40%.

[0088] The actual application test shows that the four-end perovskite / crystalline silicon mechanical stacking structure and the intelligent current-voltage dynamic matching system not only effectively solve the current-voltage mismatching problem between perovskite and crystalline silicon sub-components, but also improve the irradiance response speed by 5 times through time-sharing multiplexing MPPT and dynamic voltage matching technology, and the hybrid energy storage buffer mechanism effectively solves the system mismatching problem caused by instantaneous light intensity fluctuation.

[0089] The second aspect of the embodiment of the application provides an electronic device, comprising: a processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described above.

[0090] The third aspect of the embodiment of the application provides a computer readable storage medium, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the method described above.

[0091] The application can be a method, device, system and / or computer program product. The computer program product can include a computer readable storage medium, which loads computer readable program instructions for executing various aspects of the application.

[0092] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A voltage adaptive matching control method for a perovskite / crystalline silicon tandem photovoltaic module, characterized in that: include: Collecting the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell in the perovskite / crystalline silicon stacked photovoltaic module, and calculating the real-time voltage difference between the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell; A voltage difference compensation model based on a graph neural network is constructed. The model uses perovskite sub-cells and crystalline silicon sub-cells as graph network nodes and voltage differences as edge features. The node states are iteratively updated through a message passing mechanism, and the node weights are adaptively adjusted in combination with an attention mechanism to output the optimal voltage compensation value. Compensating the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell according to the optimal voltage compensation value to obtain a compensated perovskite sub-cell voltage and a compensated crystalline silicon sub-cell voltage; A system state space model is constructed based on the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage, a comprehensive reward signal is generated based on the system state space model, and a priority experience replay mechanism is used to update the state value estimate to obtain an optimized control strategy. The optimized control strategy uses a recursive least squares algorithm to identify the parameters of the system state space model in real time, solves the optimized control sequence in a rolling time domain, and introduces an adaptive law to dynamically adjust the controller parameters. According to the output control quantity of the optimization control strategy, the working state of the driving circuit of the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell is adjusted to achieve optimal voltage matching.

2. The method according to claim 1, characterized in that A voltage difference compensation model based on graph neural network is constructed. The voltage difference compensation model uses perovskite sub-cells and crystalline silicon sub-cells as graph network nodes and voltage differences as edge features. The node status is iteratively updated through a message passing mechanism, and the node weights are adaptively adjusted in combination with the attention mechanism. The output of the optimal voltage compensation value includes: A voltage difference compensation model based on a graph neural network is constructed, and the perovskite sub-cell and the crystalline silicon sub-cell are constructed as graph network nodes to generate a node state vector; the output voltage difference between the perovskite sub-cell and the crystalline silicon sub-cell is calculated, and the output voltage difference is used as an edge feature. The Hodgkin-Huxley neuron model is used to construct an initial edge feature representation. The Hodgkin-Huxley neuron model dynamically represents the initial edge feature representation through membrane capacitance, ion channel conductance, and reversal potential; Inputting the initial edge feature representation into the ion channel module of the Hodgkin-Huxley neuron model, and enhancing the initial edge feature representation based on the sodium ion channel and potassium ion channel characteristics of the ion channel module to form an enhanced edge feature representation with long-range correlation; constructing a message passing function based on the enhanced edge feature representation, wherein the message passing function receives the node state vector and the enhanced edge feature representation as input and generates node state update information; Inputting the node state update information into the synaptic plasticity module of the Hodgkin-Huxley neuron model, and adaptively adjusting the node state update information through the dynamic threshold mechanism of the synaptic plasticity module to obtain an adjusted node state; An attention mechanism is applied to the adjusted node state, and the inter-node attention coefficient is obtained by calculating the node state similarity and performing normalization processing; the perovskite sub-cell node state and the crystalline silicon sub-cell node state with the node attention coefficient are input into a multi-layer perceptron to obtain the optimal voltage compensation value.

3. The method according to claim 2, characterized in that Calculating the output voltage difference between the perovskite sub-cell and the crystalline silicon sub-cell, using the output voltage difference as an edge feature, and constructing an initial edge feature representation using a Hodgkin-Huxley neuron model include: Calculating the voltage difference between the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell, multiplying the voltage difference by a mapping coefficient and superimposing the resting potential to obtain a mapped membrane potential, wherein the mapping coefficient is determined by the amplitude range of the voltage difference and the physiological range of the biological neuron membrane potential; constructing an ion channel kinetic model based on the mapped membrane potential, and calculating the time rate of change of the mapped membrane potential by using the membrane capacitance per unit area, the maximum conductance of each ion channel, the equilibrium potential of each ion, and the external input current; A gating variable kinetic equation is established based on the ion channel kinetic model, and the voltage-dependent rate constant of each gating variable is calculated by mapping the membrane potential, wherein the voltage-dependent rate constant includes a forward rate constant and a reverse rate constant; the steady-state value and the time constant of each gating variable are calculated based on the forward rate constant and the reverse rate constant, and the steady-state value and the time constant are substituted into the exponential form of the kinetic equation to solve the time evolution of each gating variable, and the time evolution is numerically integrated by the Runge-Kutta method to form an initial edge feature representation.

4. The method according to claim 1, wherein A system state space model is constructed based on the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage, a comprehensive reward signal is generated based on the system state space model, and an optimized control strategy is obtained by updating the state value estimation using a priority experience replay mechanism. The optimization control strategy includes: Constructing a system state space model including the voltage state of the perovskite sub-cell and the voltage state of the crystalline silicon sub-cell, and constructing a prediction cost function based on the system state space model; calculating an initial control strategy based on the prediction cost function, configuring an Actor-Critic agent for the perovskite sub-cell and the crystalline silicon sub-cell respectively according to the initial control strategy, and calculating the state expected reward based on the state action value function of the Actor network of the Actor-Critic agent; Inputting the state expected return into a hierarchical reward mechanism, the hierarchical reward mechanism calculates a local reward function according to the voltage tracking error of a single sub-battery, calculates a global reward function according to the joint tracking error of the voltages of the two sub-batteries, calculates a temporal reward function according to the time derivative of the voltage difference, and combines the local reward function, the global reward function and the temporal reward function to form a comprehensive reward signal; The integrated reward signal is used to construct a priority experience replay mechanism, which calculates the policy gradient through the temporal difference error, and uses the policy gradient to simultaneously update the state value estimate of the Actor-Critic agent to obtain an optimized control strategy.

5. The method according to claim 4, characterized in that The integrated reward signal is used to construct a priority experience replay mechanism. The priority experience replay mechanism calculates the policy gradient through the temporal difference error, and uses the policy gradient to simultaneously update the state value estimate of the Actor-Critic agent. The optimized control strategy includes: Calculating a fractal dimension of a strategy trajectory for the comprehensive reward signal, and constructing an experience priority evaluation function based on the fractal dimension, a temporal difference error, and a state transition similarity, wherein the experience priority evaluation function is used to determine the importance of experience samples; The Logistic chaotic sequence is introduced into the temporal difference error calculation process. The Logistic chaotic sequence modulates the discounted state value through the chaos intensity parameter to obtain the chaos-enhanced temporal difference error. The chaos-enhanced temporal difference error is used to evaluate the temporal correlation of the state-action pair. Constructing a policy gradient based on the chaos-enhanced temporal difference error and the fractal dimension, inputting the policy gradient into the Actor network for parameter update, and simultaneously inputting the chaos-enhanced temporal difference error and the chaotic modulation term into the Critic network for state value update estimation, and adaptively adjusting the learning rates of the Actor network and the Critic network according to the training process; A final optimization control strategy is constructed based on the updated state value estimate. The final optimization control strategy adjusts the exploration degree of action selection through a temperature parameter to achieve a probabilistic strategy output based on an action value function.

6. The method according to claim 1, characterized in that The optimization control strategy identifies the parameters of the system state space model in real time based on a recursive least squares algorithm, solves the optimization control sequence in a rolling time domain, and introduces an adaptive law to dynamically adjust the controller parameters, including: collecting system input and output data, constructing an initial parameter estimation model using a recursive least squares algorithm, wherein the recursive least squares algorithm iteratively updates a system parameter vector via a Kalman gain matrix; performing wavelet decomposition on the system input and output data, decomposing the system input and output data into different frequency band features, and combining the different frequency band features with wavelet basis functions to obtain a multi-scale signal representation; Constructing a parameter identification model in each frequency band based on the multi-scale signal representation, wherein the parameter identification model adopts an iterative structure of the recursive least squares algorithm and adjusts the Kalman gain matrix according to frequency band characteristics; Inputting the identification results of the parameter identification model into a radial basis function neural network, adjusting the network parameters of the radial basis function neural network by a gradient descent method, wherein the update step size of the gradient descent method is calculated according to the identification error; Performing distributed optimization calculations on output results of the radial basis function neural network, and correcting parameters of the recursive least squares algorithm based on the results of the distributed optimization calculations; constructing a Lyapunov function based on the results of the distributed optimization calculations, designing update equations for controller parameters based on the Lyapunov function, and solving the update equations using the Kalman gain matrix; An optimal control sequence is generated based on the controller parameters, the optimal control sequence is optimized by the gradient descent method, and is used for online control of the system.

7. The method according to claim 6, characterized in that Performing distributed optimization calculation on the output result of the radial basis function neural network, and correcting the parameters of the recursive least squares algorithm based on the result of the distributed optimization calculation; Constructing a Lyapunov function based on the result of the distributed optimization calculation includes: The system input and output data are sent to the radial basis function neural network for processing, and the radial basis function neural network performs nonlinear mapping on the system input and output data through network weights, center vectors and expansion constants to obtain neural network output results; Building a distributed optimization network based on the output of the neural network, wherein the distributed optimization network uses a graph topology to connect multiple computing nodes, each computing node independently calculates a local gradient based on neighborhood information, and transmits local gradient information between the computing nodes through a consistency protocol, and iteratively optimizes the output of the neural network based on the local gradient and the local gradient information; Correcting parameters of the recursive least squares algorithm according to the neural network output result, the parameter correction process comprising: constructing a parameter error covariance matrix based on the neural network output result, calculating a Kalman gain using the parameter error covariance matrix, updating parameter estimates according to the Kalman gain, and reconstructing a parameter error covariance matrix based on the parameter estimates; A Lyapunov function is constructed using the reconstructed parameter error covariance matrix. The construction process of the Lyapunov function includes: mapping the neural network output result and the reconstructed parameter error covariance matrix to an error space to obtain a state error vector, constructing a quadratic function of the state error vector to obtain a system error term, constructing the parameter estimation error of the recursive least squares algorithm as a parameter error term, and combining the system error term and the parameter error term to form a complete Lyapunov function.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Centralized-type photovoltaic power generation system capable of achieving distributed MPPT

    CN106941263A

  • Energy equalization circuit of battery system

    CN110620413A

  • Perovskite silicon photovoltaic cell laminated assembly

    CN118714865A

  • Method and system for realizing low power consumption of power adapter

    CN118971569A

  • Voltage control method and system based on adaptive particle swarm optimization and SMC

    CN119356474A