Perovskite / crystalline silicon stacked photovoltaic module voltage self-adaptive matching control method

By constructing a voltage difference compensation model using graph neural networks and reinforcement learning, the voltage difference of perovskite/crystalline silicon tandem photovoltaic modules is adjusted in real time, solving the voltage mismatch problem and improving the power generation efficiency and lifespan of the modules.

CN120811273BActive Publication Date: 2026-05-29JIANGSU XIEHANG ENERGY TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGSU XIEHANG ENERGY TECH CO LTD
Filing Date
2025-07-02
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing perovskite/crystalline silicon tandem photovoltaic modules, the voltage output mismatch between perovskite sub-cells and crystalline silicon sub-cells affects power generation efficiency and lifespan. Traditional control methods cannot adjust in real time and have high computational complexity, making it difficult to cope with environmental changes.

Method used

A voltage difference compensation model based on graph neural network is constructed. By combining attention mechanism and message passing mechanism, the voltage difference between perovskite and crystalline silicon sub-cells is adjusted in real time through optimized control strategy. Reinforcement learning and recursive least squares algorithm are used to adaptively adjust controller parameters to achieve optimal voltage matching.

Benefits of technology

It improves the accuracy and response speed of voltage matching, thereby enhancing the working efficiency and lifespan of photovoltaic modules under different environmental conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811273B_ABST
    Figure CN120811273B_ABST
Patent Text Reader

Abstract

The application provides a perovskite / crystalline silicon laminated photovoltaic module voltage self-adaptive matching control method, relates to the technical field of voltage control, and comprises the following steps: collecting sub-cell voltages, calculating voltage difference values; constructing a voltage difference value compensation model based on a graph neural network to output optimal voltage compensation values; compensating the sub-cell voltages; constructing a state space model based on the compensated voltages, and obtaining an optimized control strategy by using a priority experience replay mechanism; and adjusting the working state of a driving circuit according to the control quantity. The application realizes voltage self-adaptive dynamic matching between sub-cells of a laminated module, and improves the power generation efficiency of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to voltage control technology, and more particularly to a voltage adaptive matching control method for perovskite / crystalline silicon tandem photovoltaic modules. Background Technology

[0002] Perovskite / crystalline silicon tandem photovoltaic (PV) modules have become an important technological approach for improving photoelectric conversion efficiency due to their combination of the high open-circuit voltage of perovskite solar cells and the high stability of crystalline silicon solar cells. In four-terminal perovskite / crystalline silicon tandem PV modules, the perovskite and crystalline silicon sub-cells are typically connected in parallel or electrically independent, with their output performance limited by the voltage matching between the two types of sub-cells. However, in practical applications, due to the different spectral response characteristics of perovskite and crystalline silicon materials, as well as changes in external conditions such as ambient temperature and light intensity, voltage mismatch often occurs between the two types of sub-cells, affecting the overall power generation efficiency and lifespan of the tandem module.

[0003] Traditional control methods often employ fixed-parameter control strategies, which cannot adjust control parameters in real time according to dynamic changes in environmental conditions and component states, resulting in insufficient adaptability of the system in complex and ever-changing working environments. Secondly, existing technologies do not adequately consider the mutual influence between perovskite sub-cells and crystalline silicon sub-cells, and lack a collaborative optimization mechanism for the two sub-cells as a whole system, making it difficult to achieve globally optimal voltage matching. Thirdly, existing voltage matching algorithms have high computational complexity and insufficient real-time performance, and most methods fail to fully consider the nonlinear and time-varying characteristics of photovoltaic systems, making it difficult to cope with the rapid voltage adjustment requirements under conditions such as sudden changes in light intensity and temperature fluctuations. Summary of the Invention

[0004] This invention provides a voltage adaptive matching control method for perovskite / crystalline silicon tandem photovoltaic modules, which can solve the problems in the prior art.

[0005] A first aspect of the present invention provides a voltage adaptive matching control method for perovskite / crystalline silicon tandem photovoltaic modules, comprising:

[0006] Collect the output voltages of the perovskite sub-cells and the crystalline silicon sub-cells in the perovskite / crystalline silicon tandem photovoltaic module, and calculate the real-time voltage difference between the output voltages of the perovskite sub-cells and the crystalline silicon sub-cells.

[0007] A voltage difference compensation model based on a graph neural network is constructed. The voltage difference compensation model uses perovskite sub-cells and crystalline silicon sub-cells as graph network nodes, voltage difference as edge features, iteratively updates the node state through a message passing mechanism, and adaptively adjusts the node weights by combining an attention mechanism to output the optimal voltage compensation value.

[0008] The output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell are compensated according to the optimal voltage compensation value to obtain the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage.

[0009] A system state-space model is constructed based on the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage. A comprehensive reward signal is generated based on the system state-space model. The state value estimate is updated using a priority experience replay mechanism to obtain an optimized control strategy. The optimized control strategy identifies the parameters of the system state-space model in real time based on a recursive least squares algorithm and solves the optimized control sequence in the rolling time domain. At the same time, an adaptive law is introduced to dynamically adjust the controller parameters.

[0010] Based on the output control quantity of the optimized control strategy, the operating state of the drive circuits for the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell is adjusted to achieve optimal voltage matching.

[0011] A voltage difference compensation model based on a graph neural network is constructed. This model uses perovskite and crystalline silicon sub-cells as nodes in the graph network, and the voltage difference as edge features. The node states are iteratively updated through a message passing mechanism, and the node weights are adaptively adjusted using an attention mechanism. The optimal voltage compensation value is output as follows:

[0012] A voltage difference compensation model based on a graph neural network is constructed, in which the perovskite sub-cell and the crystalline silicon sub-cell are constructed as graph network nodes to generate node state vectors; the output voltage difference between the perovskite sub-cell and the crystalline silicon sub-cell is calculated, and the output voltage difference is used as an edge feature. An initial edge feature representation is constructed using a Hodgkin-Huxley neuron model. The Hodgkin-Huxley neuron model dynamically represents the initial edge feature representation through membrane capacitance, ion channel conductance, and reversal potential.

[0013] The initial edge feature representation is input into the ion channel module of the Hodgkin-Huxley neuron model. Based on the sodium and potassium ion channel characteristics of the ion channel module, the initial edge feature representation is enhanced to form an enhanced edge feature representation with long-range correlation.

[0014] A message passing function is constructed based on the enhanced edge feature representation. The message passing function receives the node state vector and the enhanced edge feature representation as input and generates node state update information.

[0015] The node state update information is input into the synaptic plasticity module of the Hodgkin-Huxley neuron model. The node state update information is adaptively adjusted through the dynamic threshold mechanism of the synaptic plasticity module to obtain the adjusted node state.

[0016] An attention mechanism is applied to the adjusted node states. By calculating the node state similarity and performing normalization, the attention coefficients between nodes are obtained. The perovskite sub-cell node states and crystalline silicon sub-cell node states with the node attention coefficients are input into a multilayer perceptron to obtain the optimal voltage compensation value.

[0017] The output voltage difference between the perovskite sub-cell and the crystalline silicon sub-cell is calculated. This output voltage difference is used as a side feature. An initial side feature representation is constructed using the Hodgkin-Huxley neuron model, including:

[0018] The voltage difference between the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell is calculated. The voltage difference is multiplied by a mapping coefficient and superimposed with the resting potential to obtain the mapped membrane potential. The mapping coefficient is determined by the amplitude range of the voltage difference and the physiological range of the membrane potential of biological neurons.

[0019] An ion channel dynamics model is constructed based on the mapped membrane potential, and the time rate of change of the mapped membrane potential is calculated using the membrane capacitance per unit area, the maximum conductance of each ion channel, the equilibrium potential of each ion, and the external input current.

[0020] Based on the ion channel kinetic model, a kinetic equation for the gated variables is established. The voltage-dependent rate constants of each gated variable are calculated using the mapped membrane potential. The voltage-dependent rate constants include forward rate constants and reverse rate constants. The steady-state values ​​and time constants of each gated variable are calculated based on the forward rate constants and the reverse rate constants. The steady-state values ​​and time constants are substituted into the exponential form kinetic equation to solve for the time evolution of each gated variable. The time evolution is numerically integrated using the Runge-Kutta method to form an initial edge feature characterization.

[0021] A system state-space model is constructed based on the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage. A comprehensive reward signal is generated based on the system state-space model. The optimized control strategy is obtained by updating the state value estimate using a priority experience replay mechanism.

[0022] A system state-space model is constructed, which includes the voltage states of the perovskite sub-cell and the crystalline silicon sub-cell. A prediction cost function is constructed based on the system state-space model. An initial control strategy is calculated based on the prediction cost function. An Actor-Critic agent is configured for the perovskite sub-cell and the crystalline silicon sub-cell according to the initial control strategy. The Actor network of the Actor-Critic agent calculates the expected state reward based on the state action value function.

[0023] The expected state reward is input into a hierarchical reward mechanism. The hierarchical reward mechanism calculates a local reward function based on the voltage tracking error of a single sub-cell, a global reward function based on the joint tracking error of the voltages of two sub-cells, and a time-series reward function based on the time derivative of the voltage difference. The local reward function, the global reward function, and the time-series reward function are then combined to form a comprehensive reward signal.

[0024] A priority experience replay mechanism is constructed using the comprehensive reward signal. The priority experience replay mechanism calculates the policy gradient through time-series differential error and uses the policy gradient to simultaneously update the state value estimate of the Actor-Critic agent to obtain an optimized control policy.

[0025] A priority experience replay mechanism is constructed using the comprehensive reward signal. This mechanism calculates the policy gradient through temporal differential error and simultaneously updates the state value estimate of the Actor-Critic agent using the policy gradient to obtain the optimized control policy, including:

[0026] The fractal dimension of the strategy trajectory is calculated for the comprehensive reward signal. An empirical priority evaluation function is constructed based on the fractal dimension, temporal differential error, and state transition similarity. The empirical priority evaluation function is used to determine the importance of empirical samples.

[0027] The Logistic chaotic sequence is introduced into the time-series difference error calculation process. The Logistic chaotic sequence is modulated and discounted with a chaos intensity parameter to obtain a chaotic enhanced time-series difference error. The chaotic enhanced time-series difference error is used to evaluate the temporal correlation of state-action pairs.

[0028] The policy gradient is constructed based on the chaotic enhanced temporal difference error and the fractal dimension. The policy gradient is input into the Actor network for parameter update. At the same time, the chaotic enhanced temporal difference error and the chaotic modulation term are input into the Critic network to update the state value estimate. The learning rates of the Actor network and the Critic network are adaptively adjusted according to the training process.

[0029] The final optimized control strategy is constructed based on the updated state value estimate. The final optimized control strategy adjusts the exploration degree of action selection through temperature parameters to achieve a probabilistic strategy output based on the action value function.

[0030] The optimized control strategy identifies the parameters of the system state-space model in real time based on the recursive least squares algorithm, solves the optimized control sequence in the rolling time domain, and introduces an adaptive law to dynamically adjust the controller parameters, including:

[0031] The system input and output data are collected, and an initial parameter estimation model is constructed using a recursive least squares algorithm. The recursive least squares algorithm iteratively updates the system parameter vector through the Kalman gain matrix. Wavelet decomposition is performed on the system input and output data to decompose the system input and output data into different frequency band features. The different frequency band features are combined with wavelet basis functions to obtain a multi-scale signal representation.

[0032] Based on the multi-scale signal representation, a parameter identification model is constructed in each frequency band. The parameter identification model adopts the iterative structure of the recursive least squares algorithm and adjusts the Kalman gain matrix according to the frequency band characteristics.

[0033] The identification results of the parameter identification model are input into the radial basis function neural network. The network parameters of the radial basis function neural network are adjusted by the gradient descent method. The update step size of the gradient descent method is calculated based on the identification error.

[0034] The output of the radial basis function neural network is subjected to distributed optimization calculation, and the parameters of the recursive least squares algorithm are corrected based on the results of the distributed optimization calculation; a Lyapunov function is constructed based on the results of the distributed optimization calculation, and the update equation of the controller parameters is designed according to the Lyapunov function, and the update equation is solved by the Kalman gain matrix;

[0035] An optimal control sequence is generated based on the controller parameters. This optimal control sequence is optimized using the gradient descent method and then used for the online control of the system.

[0036] Distributed optimization computation is performed on the output of the radial basis function neural network, and the parameters of the recursive least squares algorithm are corrected based on the results of the distributed optimization computation; the Lyapunov function is constructed based on the results of the distributed optimization computation, including:

[0037] The system input and output data are fed into the radial basis function neural network for processing. The radial basis function neural network performs a nonlinear mapping on the system input and output data through network weights, center vectors, and expansion constants to obtain the neural network output results.

[0038] A distributed optimization network is constructed based on the output of the neural network. The distributed optimization network uses a graph topology to connect multiple computing nodes. Each computing node independently calculates the local gradient based on neighborhood information and transmits the local gradient information between computing nodes through a consensus protocol. The output of the neural network is iteratively optimized based on the local gradient and the local gradient information.

[0039] The parameters of the recursive least squares algorithm are corrected based on the output of the neural network. The parameter correction process includes: constructing a parameter error covariance matrix based on the output of the neural network; calculating the Kalman gain using the parameter error covariance matrix; updating the parameter estimate based on the Kalman gain; and reconstructing the parameter error covariance matrix based on the parameter estimate.

[0040] The Lyapunov function is constructed using the reconstructed parameter error covariance matrix. The construction process of the Lyapunov function includes: mapping the output of the neural network and the reconstructed parameter error covariance matrix to the error space to obtain the state error vector; constructing a quadratic form function of the state error vector to obtain the system error term; constructing the parameter estimation error of the recursive least squares algorithm as the parameter error term; and combining the system error term and the parameter error term to form the complete Lyapunov function.

[0041] A second aspect of the present invention provides an electronic device, comprising:

[0042] processor;

[0043] Memory used to store processor-executable instructions;

[0044] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0045] A third aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0046] The beneficial effects of this application are as follows:

[0047] By constructing a voltage difference compensation model based on graph neural networks, the voltage difference between perovskite sub-cells and crystalline silicon sub-cells can be accurately captured. Adaptive dynamic voltage compensation is achieved by using message passing and attention mechanisms, which effectively solves the problem of insufficient compensation accuracy of traditional methods under complex lighting conditions.

[0048] Based on reinforcement learning and priority experience playback mechanism, the system can continuously optimize the control strategy and adaptively learn the best control parameters, which significantly improves the accuracy and response speed of voltage matching, enabling photovoltaic modules to maintain the best working state under different environmental conditions.

[0049] The recursive least squares algorithm is used to identify system parameters in real time and solve the optimized control sequence in the rolling time domain. Combined with the adaptive law, the controller parameters are dynamically adjusted, which makes the whole control system have strong robustness and adaptability, effectively improving the energy conversion efficiency and service life of perovskite / crystalline silicon tandem photovoltaic modules. Attached Figure Description

[0050] Figure 1 This is a schematic flowchart of the voltage adaptive matching control method for perovskite / crystalline silicon tandem photovoltaic modules according to an embodiment of the present invention;

[0051] Figure 2 This is a schematic diagram comparing the enhancement effects of the Hodgkin-Huxley neuron model in an embodiment of the present invention;

[0052] Figure 3 This is a bar chart comparing the performance of optimized control strategies for perovskite / crystalline silicon twin cells in embodiments of the present invention.

[0053] Figure 4 This is a flowchart of the optimization control strategy based on the recursive least squares algorithm in an embodiment of the present invention;

[0054] Figure 5 This is a bar chart comparing the distributed optimization efficiency of radial basis function neural networks in embodiments of the present invention.

[0055] Figure 6 (a) is a structural design diagram of a photovoltaic module mechanically stacked by a guide rail and a clamp according to an embodiment of the present invention, where N1 (black negative sign) and P1 (red positive sign) represent the negative and positive electrodes of the perovskite sub-module, respectively, and N2 (black negative sign) and P2 (red positive sign) represent the negative and positive electrodes of the crystalline silicon sub-module, respectively.

[0056] Figure 6 (b) is a structural design diagram of a tandem photovoltaic module formed by encapsulating four-terminal perovskite sub-cell strings and crystalline silicon sub-cell strings with an adhesive film according to an embodiment of the present invention. N1 (black negative sign) and P1 (red positive sign) represent the negative and positive electrodes of the perovskite sub-cell strings, respectively, and N2 (black negative sign) and P2 (red positive sign) represent the negative and positive electrodes of the crystalline silicon sub-cell strings, respectively. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0059] Figure 1 This is a flowchart illustrating the voltage adaptive matching control method for perovskite / crystalline silicon tandem photovoltaic modules according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0060] Collect the output voltages of the perovskite sub-cells and the crystalline silicon sub-cells in the perovskite / crystalline silicon tandem photovoltaic module, and calculate the real-time voltage difference between the output voltages of the perovskite sub-cells and the crystalline silicon sub-cells.

[0061] A voltage difference compensation model based on a graph neural network is constructed. The voltage difference compensation model uses perovskite sub-cells and crystalline silicon sub-cells as graph network nodes, voltage difference as edge features, iteratively updates the node state through a message passing mechanism, and adaptively adjusts the node weights by combining an attention mechanism to output the optimal voltage compensation value.

[0062] The output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell are compensated according to the optimal voltage compensation value to obtain the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage.

[0063] A system state-space model is constructed based on the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage. A comprehensive reward signal is generated based on the system state-space model. The state value estimate is updated using a priority experience replay mechanism to obtain an optimized control strategy. The optimized control strategy identifies the parameters of the system state-space model in real time based on a recursive least squares algorithm and solves the optimized control sequence in the rolling time domain. At the same time, an adaptive law is introduced to dynamically adjust the controller parameters.

[0064] Based on the output control quantity of the optimized control strategy, the operating state of the drive circuits for the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell is adjusted to achieve optimal voltage matching.

[0065] In one optional implementation, a voltage difference compensation model based on a graph neural network is constructed. This model uses perovskite and crystalline silicon sub-cells as graph network nodes, voltage difference as edge features, and iteratively updates node states through a message passing mechanism. It also adaptively adjusts node weights using an attention mechanism, outputting the optimal voltage compensation value, including:

[0066] A voltage difference compensation model based on a graph neural network is constructed, in which the perovskite sub-cell and the crystalline silicon sub-cell are constructed as graph network nodes to generate node state vectors; the output voltage difference between the perovskite sub-cell and the crystalline silicon sub-cell is calculated, and the output voltage difference is used as an edge feature. An initial edge feature representation is constructed using a Hodgkin-Huxley neuron model. The Hodgkin-Huxley neuron model dynamically represents the initial edge feature representation through membrane capacitance, ion channel conductance, and reversal potential.

[0067] The initial edge feature representation is input into the ion channel module of the Hodgkin-Huxley neuron model. Based on the sodium and potassium ion channel characteristics of the ion channel module, the initial edge feature representation is enhanced to form an enhanced edge feature representation with long-range correlation.

[0068] A message passing function is constructed based on the enhanced edge feature representation. The message passing function receives the node state vector and the enhanced edge feature representation as input and generates node state update information.

[0069] The node state update information is input into the synaptic plasticity module of the Hodgkin-Huxley neuron model. The node state update information is adaptively adjusted through the dynamic threshold mechanism of the synaptic plasticity module to obtain the adjusted node state.

[0070] An attention mechanism is applied to the adjusted node states. By calculating the node state similarity and performing normalization, the attention coefficients between nodes are obtained. The perovskite sub-cell node states and crystalline silicon sub-cell node states with the node attention coefficients are input into a multilayer perceptron to obtain the optimal voltage compensation value.

[0071] The voltage difference compensation model based on a graph neural network first constructs perovskite and crystalline silicon sub-cells as graph network nodes. For each perovskite sub-cell node, its features are represented by a 32-dimensional vector, including the short-circuit current density (20 mA / cm²). 2Parameters such as open-circuit voltage (1.1V), fill factor (0.75), spectral response range (300-800nm), and temperature coefficient (-0.2% / ℃) are also included. Similarly, for each crystalline silicon sub-cell node, its characteristics are represented by a 32-dimensional vector, including short-circuit current density (42mA / cm²). 2 The parameters include open-circuit voltage (0.7V), fill factor (0.82), spectral response range (400-1100nm), and temperature coefficient (-0.3% / ℃). These parameters are acquired through a real-time acquisition system with a sampling frequency of 1Hz and are normalized to scale all feature values ​​to the [-1,1] interval.

[0072] The output voltage difference between the perovskite and crystalline silicon sub-cells was calculated as a side feature. For example, under standard test conditions (1000 W / m² illumination, 25°C), the perovskite sub-cell output voltage was 1.05 V, and the crystalline silicon sub-cell output voltage was 0.68 V, with a voltage difference of 0.37 V. This voltage difference dynamically changes with environmental conditions and is characterized using the Hodgkin-Huxley neuron model.

[0073] The Hodgkin-Huxley neuron model processes voltage difference signals by simulating the electrophysiological characteristics of biological neurons. This model includes a membrane capacitance parameter set to 1.0 μF / cm. 2 The maximum conductivity of the sodium ion channel is 120 mS / cm 2 The maximum conductivity of the potassium ion channel is 36 mS / cm 2 The leakage current conductivity is 0.3 mS / cm. 2 The sodium ion reversal potential was set to 50 mV, the potassium ion reversal potential to -77 mV, and the leakage current reversal potential to -54.4 mV. These parameters were optimized using a grid search method to minimize the root mean square error of the model on the validation set.

[0074] The initial edge feature representation is processed by the ion channel module of the Hodgkin-Huxley neuron model, which simulates the switching dynamics of sodium and potassium ion channels. When the input voltage difference exceeds a threshold (set to 0.2V), the sodium ion channel is rapidly activated and reaches its peak (approximately 1.5ms), followed by the activation of the potassium ion channel (approximately 5ms). This temporal characteristic allows the model to capture rapid changes in voltage difference. In practical applications, when the light intensity increases from 600W / m²... 2 The power surged to 900W / m 2 When the voltage difference changes from 0.25V to 0.33V, the feature characterization after ion channel processing can reflect the time dynamics of this change, forming a 16-dimensional enhanced edge feature characterization.

[0075] A message passing function is constructed based on enhanced edge feature representations. This function receives node state vectors and enhanced edge feature representations as inputs. Specifically, a three-layer fully connected network is used: the input layer dimension is (32+16)=48, the hidden layer dimension is 64, and the output layer dimension is 32. The activation function is ReLU. This message passing process is executed iteratively three times, updating the node state information in each iteration. In actual testing, the message passing process effectively integrates different lighting conditions (200-1200W / m²). 2 The node characteristics under various temperature conditions (10-45℃) enhance the model's adaptability to environmental changes.

[0076] Node state updates are adaptively adjusted using the synaptic plasticity module of the Hodgkin-Huxley neuron model. This module implements a dynamic thresholding mechanism, with an initial threshold set to 0.5, which is dynamically adjusted based on the strength of historical input signals. When continuously receiving strong signals (voltage difference > 0.3V), the threshold increases to 0.65; when receiving weak signals (voltage difference < 0.1V), the threshold decreases to 0.35. This mechanism simulates the adaptability of biological neurons, making the model more stable at different operating points. For example, under low light conditions (200W / m²), the threshold remains stable. 2 Under normal conditions, the model is sensitive to small voltage differences (0.08V); however, under strong light conditions (1200W / m²), it can respond sensitively to small voltage differences (0.08V); 2 Under these conditions, it can avoid overreacting to large voltage differences (0.45V).

[0077] An attention mechanism is applied to the adjusted node states to calculate the state similarity between perovskite and crystalline silicon sub-cell nodes. The similarity calculation employs a dot product operation followed by Softmax normalization to obtain the attention coefficients. In a typical scenario, when the temperature of the perovskite sub-cell rises to 40°C while the crystalline silicon sub-cell remains at 25°C, the attention mechanism automatically assigns a weight of 0.65 to the perovskite sub-cell node and a weight of 0.35 to the crystalline silicon sub-cell node, reflecting their different contributions to the final compensation value.

[0078] Finally, the attention coefficient-based perovskite sub-cell node states and crystalline silicon sub-cell node states are concatenated and input into a multilayer perceptron. This multilayer perceptron consists of three layers: a 64-dimensional input layer, a 32-dimensional intermediate layer, and a 1-dimensional output layer. The activation function is a combination of Tanh and ReLU. The model output represents the optimal voltage compensation value, ranging from -0.5V to 0.5V. In actual testing, this model can increase the output power of the tandem cell system by 8.2% to 13.5%, a 5.3 percentage point improvement compared to traditional fixed compensation methods, and it also performs well under rapidly changing environmental conditions (such as cloud cover causing illumination to drop from 900W / m²). 2 The power dropped sharply to 400 W / m 2The response time is less than 150ms, which meets the requirements for real-time compensation.

[0079] Figure 2 This is a schematic diagram comparing the enhancement effects of the Hodgkin-Huxley neuron model in an embodiment of the present invention:

[0080] This figure compares the performance of three different methods in terms of feature distance and time response. The left figure shows the decay trend of feature distance with increasing dimension. Our proposed solution (GNN+HH model) maintains a feature distance of 0.63 at a feature distance of 4, which is significantly better than the 0.32 of the traditional GCN method and the 0.25 of the membrane-free subchannel enhancement method, indicating that our solution has better feature preservation ability. The right figure shows the time response characteristics of different ions. The peak response of sodium ions (Na+) reaches 0.75 at 7.5 ms, while the peak response of the traditional GCN method is only 0.65. Our proposed solution exhibits faster response speed and higher response intensity throughout the time series, especially in the 15-20 ms time window, where the smoothness and stability of its response curve are better than the comparative methods. These results fully demonstrate the significant advantages of the GNN and HH model fusion solution in feature extraction capability and kinetic response performance, providing a more accurate and efficient solution for ion channel kinetic modeling.

[0081] In one optional implementation, the output voltage difference between the perovskite sub-cell and the crystalline silicon sub-cell is calculated, and the output voltage difference is used as a side feature. The initial side feature representation is constructed using a Hodgkin-Huxley neuron model, including:

[0082] The voltage difference between the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell is calculated. The voltage difference is multiplied by a mapping coefficient and superimposed with the resting potential to obtain the mapped membrane potential. The mapping coefficient is determined by the amplitude range of the voltage difference and the physiological range of the membrane potential of biological neurons.

[0083] An ion channel dynamics model is constructed based on the mapped membrane potential, and the time rate of change of the mapped membrane potential is calculated using the membrane capacitance per unit area, the maximum conductance of each ion channel, the equilibrium potential of each ion, and the external input current.

[0084] Based on the ion channel kinetic model, a kinetic equation for the gated variables is established. The voltage-dependent rate constants of each gated variable are calculated using the mapped membrane potential. The voltage-dependent rate constants include forward rate constants and reverse rate constants. The steady-state values ​​and time constants of each gated variable are calculated based on the forward rate constants and the reverse rate constants. The steady-state values ​​and time constants are substituted into the exponential form kinetic equation to solve for the time evolution of each gated variable. The time evolution is numerically integrated using the Runge-Kutta method to form an initial edge feature characterization.

[0085] The system obtains the output voltages of the perovskite sub-cell and the crystalline silicon sub-cell. For example, under specific lighting conditions, the output voltage of the perovskite sub-cell is 0.9 volts, while the output voltage of the crystalline silicon sub-cell is 0.6 volts. The system calculates the voltage difference between the two, which is 0.9 volts minus 0.6 volts, resulting in a voltage difference of 0.3 volts.

[0086] To map the voltage difference to the membrane potential range of a biological neuron, the system needs to determine a mapping coefficient. The membrane potential of a biological neuron typically varies between -70 mV and +30 mV, with a total range of 100 mV. In this embodiment, the amplitude range of the voltage difference is -0.5 volts to +0.5 volts. Therefore, the mapping coefficient can be set to 100 mV divided by 1 volt, i.e., 100 mV / volt. The system multiplies the voltage difference of 0.3 volts by the mapping coefficient 100 mV / volt to obtain 30 mV, then adds the neuron's resting potential of -70 mV, finally obtaining a mapped membrane potential of -40 mV.

[0087] Based on the mapped membrane potential, an ion channel dynamics model is constructed. In this model, the time-varying rate of membrane potential is determined by several factors, including the membrane capacitance per unit area, the maximum conductance of each ion channel, the equilibrium potential of each ion, and the external input current. In this embodiment, the membrane capacitance per unit area is set to 1 μF / cm², the maximum conductance of the sodium ion channel is 120 mSiemens / cm², the maximum conductance of the potassium ion channel is 36 mSiemens / cm², and the maximum conductance of the leakage current channel is 0.3 mSiemens / cm². The equilibrium potentials of each ion are: sodium ion equilibrium potential is +50 mV, potassium ion equilibrium potential is -77 mV, and leakage current equilibrium potential is -54.387 mV. The external input current can be set according to the actual application scenario; in this example, it is set to 0 μA / cm².

[0088] The system establishes kinetic equations for gated variables based on an ion channel kinetic model. Sodium ion channels are controlled by two gated variables, m and h, while potassium ion channels are controlled by a gated variable n. For a mapped membrane potential of -40 mV, the system calculates the voltage-dependent rate constants of each gated variable. For example, the positive rate constant α of the m-gated variable... m 0.1 millivolts(-1) ×(25-(-40)) / [exp((25-(-40)) / 10)-1]=0.1 mV (-1) ×65 / [exp(6.5)-1]≈0.1×65 / 665≈0.00977 milliseconds (-1) Reverse rate constant β m The time is 4×exp(-(-40) / 18)=4×exp(40 / 18)≈4×9.2≈36.8 milliseconds. (-1) Similarly, calculate the forward and reverse rate constants for the gated variables h and n.

[0089] Based on the forward and reverse rate constants, the system calculates the steady-state values ​​and time constants of each gated variable. The steady-state value of the m-gated variable is given by [the formula is missing here]. ∞ =α m / (α m +β m =0.00977 / (0.00977+36.8)≈0.000265; Time constant τ m =1 / (α m +β m = 1 / (0.00977+36.8) ≈ 0.027 milliseconds. The steady-state values ​​and time constants of the h-gated variables and n-gated variables are calculated using the same method. In practical implementation, the steady-state value h of the h-gated variable... ∞ Approximately 0.596, time constant τ h Approximately 8.04 milliseconds; steady-state value of n-gated variable n ∞ Approximately 0.318, time constant τ n It takes approximately 1.11 milliseconds.

[0090] The system substitutes the steady-state values ​​and time constants into the exponential form of the dynamic equations to solve for the time evolution of each gated variable. The dynamic equations, expressed in exponential form, describe how the gated variables approach their steady-state values ​​over time. The system solves these equations using numerical integration via the Runge-Kutta method. In the specific implementation, the system selects the fourth-order Runge-Kutta method, sets the time step to 0.01 milliseconds, and the total simulation duration to 50 milliseconds.

[0091] Taking an m-gated variable as an example, assuming the initial value is m (0) =0.05, calculate the value at t=0.01 milliseconds using the Runge-Kutta method. First, calculate four intermediate values: k1=f(m (0) )=(m ∞ -m (0) ) / τ m =(0.000265-0.05) / 0.027≈-1.844;k2=f(m (0) +k1×0.01 / 2); k3=f(m (0)+k2×0.01 / 2); k4=f(m (0) +k3×0.01). Then calculate m. (0.01) =m (0) +(k1+2k2+2k3+k4)×0.01 / 6. And so on, the system calculates the complete time series of the three gated variables m, h, and n within 50 milliseconds.

[0092] Through the above steps, the system constructs an initial side feature representation based on the output voltage difference between the perovskite and crystalline silicon sub-cells using a Hodgkin-Huxley neuron model. This feature representation is essentially a complete sequence of the evolution of gating variables m, h, and n over time, capturing the neuronal dynamics induced by the input voltage difference. This bio-inspired feature representation method enhances the system's sensitivity to voltage differences, providing rich dynamic features for subsequent anomaly detection and pattern recognition.

[0093] In one optional implementation, a system state-space model is constructed based on the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage. A comprehensive reward signal is generated based on the system state-space model. The optimized control strategy is obtained by updating the state value estimate using a priority experience replay mechanism.

[0094] A system state-space model is constructed, which includes the voltage states of the perovskite sub-cell and the crystalline silicon sub-cell. A prediction cost function is constructed based on the system state-space model. An initial control strategy is calculated based on the prediction cost function. An Actor-Critic agent is configured for the perovskite sub-cell and the crystalline silicon sub-cell according to the initial control strategy. The Actor network of the Actor-Critic agent calculates the expected state reward based on the state action value function.

[0095] The expected state reward is input into a hierarchical reward mechanism. The hierarchical reward mechanism calculates a local reward function based on the voltage tracking error of a single sub-cell, a global reward function based on the joint tracking error of the voltages of two sub-cells, and a time-series reward function based on the time derivative of the voltage difference. The local reward function, the global reward function, and the time-series reward function are then combined to form a comprehensive reward signal.

[0096] A priority experience replay mechanism is constructed using the comprehensive reward signal. The priority experience replay mechanism calculates the policy gradient through time-series differential error and uses the policy gradient to simultaneously update the state value estimate of the Actor-Critic agent to obtain an optimized control policy.

[0097] The following is a detailed implementation of the optimized control strategy, which is constructed based on the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage, and then generates a comprehensive reward signal and updates the state value estimate using a priority experience replay mechanism.

[0098] In a hybrid tandem solar cell system, the initial voltage values ​​of the perovskite sub-cell and the crystalline silicon sub-cell are first obtained, denoted as V. perovskite original and V silicon original Temperature and irradiance compensation were applied to the original voltages of these two sub-cells. Temperature compensation was achieved through a temperature coefficient k. temp The voltage is calibrated when the ambient temperature T is detected to be different from the standard test condition temperature T. STC When there is a difference (at 25℃), the compensated voltage V temp comp =V original ×(1+k temp ×(TT STC The temperature coefficient k of the perovskite sub-cell is... temp perovskite The temperature coefficient k of the crystalline silicon sub-cell is -0.0035℃. temp silicon The value is -0.0025 / ℃. Irradiance compensation is based on the irradiance coefficient k. irr Adjust the voltage value so that the actual irradiance G is similar to the standard test condition irradiance G. STC (1000W / m 2 At the same time, the compensated voltage V irr comp =V temp comp ×(1+k irr ×log(G / G STC The irradiance coefficient k of the perovskite sub-cell is... irr perovskite The irradiance coefficient k of the crystalline silicon sub-cell is 0.015. irr silicon It is 0.012.

[0099] After the compensation process is completed, the compensated voltage V of the perovskite sub-cell is obtained. perovskite comp The voltage V of crystalline silicon sub-cell silicon comp A system state-space model is constructed based on these two compensated voltage values. This model uses the compensated sub-cell voltage as a state variable, forming the state vector s. t =[V perovskite comp V silicon comp Simultaneously, the adjustment amount for maximum power point tracking is defined as the motion vector a. t =[ΔD perovskite ,ΔD silicon ], where ΔD represents the adjustment amount of the duty cycle of the corresponding sub-cell. The system state transition relationship can be described as s (t+1) =f(s t ,a t), where f represents the dynamic characteristics of the system. In practical applications, when the ambient temperature is 30℃ and the irradiance is 800W / m², the original voltage of the perovskite sub-cell is 0.85V at a certain moment, and after compensation it is 0.827V; the original voltage of the crystalline silicon sub-cell is 0.62V, and after compensation it is 0.607V.

[0100] Based on the constructed system state-space model, the prediction cost function J is then constructed. This cost function comprehensively considers voltage tracking error, power extraction efficiency, and control action amplitude, and is expressed as a function of state s and action a. The prediction cost function calculation considers the cumulative cost over the next N time steps, with a time step size of 100 ms chosen. By minimizing this cost function, the initial control strategy π is calculated. initial In the example, when the compensated perovskite sub-cell voltage is 0.827V and its target maximum power point voltage is 0.85V, and the compensated crystalline silicon sub-cell voltage is 0.607V and its target maximum power point voltage is 0.62V, the initial control strategy gives the action ΔD. perovskite =0.015, ΔD silicon =0.008.

[0101] According to the initial control strategy, Actor-Critic agents are configured for both the perovskite and crystalline silicon sub-cells. Each agent consists of two parts: an Actor network and a Critic network. The Actor network is responsible for outputting actions based on the current state. Its network structure contains two hidden layers, each with 64 neurons, using the ReLU activation function. The input layer receives the voltage state of the corresponding sub-cell, and the output layer uses the tanh activation function to generate a duty cycle adjustment in the range of -0.05 to 0.05. The Critic network is responsible for evaluating the value of a state or state-action pair. Its structure also contains two hidden layers, each with 64 neurons, using the ReLU activation function. The Actor network calculates the expected reward V(s) of the state based on the state-action value function Q(s,a), which is a weighted average of the Q values ​​of all actions according to the policy probability.

[0102] The expected return of the state is input into a hierarchical reward mechanism, which consists of three parts: a local reward function R, and a local reward function R. local Based on the voltage tracking error of a single sub-cell, a positive reward is given when the voltage tracking error is less than a threshold of 0.02V, and a negative reward is given otherwise, with a value range of [-1, 1]. The global reward function R... global The overall system performance is measured by calculating the joint tracking error of the two sub-cell voltages, with a numerical range of [-2, 2]. The time-series reward function R... temporalThe voltage difference is calculated based on its time derivative, reflecting the smoothness of voltage adjustment. A positive reward is given when the voltage change rate is less than the threshold of 0.01V / s; otherwise, a negative reward is given, with a value range of [-0.5, 0.5]. The combined reward signal R, integrating these three parts, is calculated as R = w1 × R. local +w2×R global +w3×R temporal The weight parameters are w1=0.3, w2=0.5, and w3=0.2.

[0103] A priority experience replay mechanism is constructed using a comprehensive reward signal. This mechanism maintains an experience pool with a capacity of 10,000, stored in the form of (s t ,a t ,r t ,s (t+1) The transferred samples are based on their time-series difference error TD. error =r t +γ×V(s (t+1) )- (s_t) Assign priority, where γ is a discount factor with a value of 0.95. Priority p i =|TD error |+ε, where ε is a small constant of 0.01, ensures that all samples have a non-zero probability of being selected. The sampling probability is proportional to the priority, according to p. i α / ∑p j α Calculate α, a priority factor with a value of 0.6. Train the network by sampling batches of 128 samples from the experience pool according to priority. Update the Actor network parameters by calculating the policy gradient and the Critic network parameters by minimizing the TD error. Set the learning rate to 0.001 and update the target network every 100 time steps.

[0104] After 50,000 iterations of training, the system is able to quickly track the maximum power point under dynamic environmental conditions. Experimental results show that, using this method, the average voltage tracking error of the perovskite sub-cell is 0.012V, and the average voltage tracking error of the crystalline silicon sub-cell is 0.009V, which are reduced by 37% and 42% respectively compared with the traditional method, and the overall system efficiency is improved by 4.3%.

[0105] Figure 3 A bar chart comparing the performance of optimized control strategies for perovskite / crystalline silicon twin cells in embodiments of the present invention:

[0106] This figure illustrates the comparison of our proposed solution (comprehensive reward + priority playback) with the traditional DQN control method and the standard Actor-Critic method across four key performance indicators. In terms of voltage tracking accuracy, our proposed solution achieves 90.5%, significantly higher than the traditional DQN's 73.2% and the standard Actor-Critic's 62.3%. Regarding convergence speed, our proposed solution performs best at 95.4%, compared to the traditional DQN's 67.5% and the standard Actor-Critic's only 55.2%. In terms of robustness, our proposed solution achieves 87.3%, compared to the traditional DQN's 75.1% and the standard Actor-Critic's 68.4%. Finally, in terms of energy efficiency improvement, our proposed solution also demonstrates excellent performance, achieving 93.6%, compared to the traditional DQN's 77.2% and the standard Actor-Critic's 65.3%. Data comparison shows that the proposed solution, which integrates comprehensive reward signals and priority experience replay, achieves significant advantages across all performance metrics, particularly in convergence speed and energy efficiency, fully demonstrating its advanced nature and practical value. These performance improvements are primarily attributed to the comprehensive evaluation of system status provided by the comprehensive reward mechanism and the efficient utilization of key experiences through priority replay.

[0107] In one optional implementation, a priority experience replay mechanism is constructed using the comprehensive reward signal. This mechanism calculates the policy gradient using time-series differential error and simultaneously updates the state value estimate of the Actor-Critic agent using the policy gradient to obtain an optimized control policy, including:

[0108] The fractal dimension of the strategy trajectory is calculated for the comprehensive reward signal. An empirical priority evaluation function is constructed based on the fractal dimension, temporal differential error, and state transition similarity. The empirical priority evaluation function is used to determine the importance of empirical samples.

[0109] The Logistic chaotic sequence is introduced into the time-series difference error calculation process. The Logistic chaotic sequence is modulated and discounted with a chaos intensity parameter to obtain a chaotic enhanced time-series difference error. The chaotic enhanced time-series difference error is used to evaluate the temporal correlation of state-action pairs.

[0110] The policy gradient is constructed based on the chaotic enhanced temporal difference error and the fractal dimension. The policy gradient is input into the Actor network for parameter update. At the same time, the chaotic enhanced temporal difference error and the chaotic modulation term are input into the Critic network to update the state value estimate. The learning rates of the Actor network and the Critic network are adaptively adjusted according to the training process.

[0111] The final optimized control strategy is constructed based on the updated state value estimate. The final optimized control strategy adjusts the exploration degree of action selection through temperature parameters to achieve a probabilistic strategy output based on the action value function.

[0112] The priority experience replay mechanism calculates the policy gradient through temporal differential error and uses the policy gradient to simultaneously update the state value estimate of the Actor-Critic agent, thereby obtaining an optimized control policy.

[0113] When calculating the fractal dimension of the policy trajectory based on the comprehensive reward signal, the system first collects policy trajectory data during the agent's interaction with the environment, including state sequences, action sequences, and reward sequences. The collected reward sequences are then calculated using box counting, dividing the sequence into n equal-sized sub-intervals and counting the number of data points within each sub-interval. The fractal dimension is calculated by using ln(N...). (ε) The relationship between N and ln(1 / ε), where N (ε) Let ε represent the number of non-empty subintervals and ε represent the size of the subintervals, from which the fractal dimension D can be obtained. In practical applications, when the reward sequence exhibits high complexity, the fractal dimension D is close to 1.8; when the reward sequence is relatively regular, the D value is close to 1.2.

[0114] An empirical priority evaluation function is constructed based on the calculated fractal dimension, temporal difference error, and state transition similarity. This function comprehensively considers three factors: the fractal dimension reflects the complexity of the state space, the temporal difference error reflects the magnitude of the prediction error, and the state transition similarity reflects the degree of similarity between the current state transition and historical experience. The specific implementation of the empirical priority evaluation function is a weighted combination of the three factors, where the fractal dimension has a weight of 0.3, the temporal difference error has a weight of 0.5, and the state transition similarity has a weight of 0.2. In practical applications, when the evaluation function value of an empirical sample exceeds the threshold of 0.75, the sample is assigned high priority; when the evaluation function value is between 0.4 and 0.75, it is assigned medium priority; and when it is below 0.4, it is assigned low priority.

[0115] When introducing a Logistic chaotic sequence into the time-series difference error calculation process, the system first generates a Logistic chaotic sequence. This sequence is generated through iterative calculation, with an initial value set to 0.4 and a control parameter set to 3.9 to ensure the system is in a chaotic state. The generated chaotic sequence is used to modulate the discounted state value. Specifically, the modulation method involves multiplying the chaotic sequence value by a discount coefficient, with the chaos intensity parameter set to 0.15. The modulated discount coefficient fluctuates between 0.8 and 0.98, introducing nonlinear dynamic characteristics. In practical applications, this chaotic modulation significantly improves the algorithm's adaptability to non-stationary environments, increasing the convergence speed by 25%.

[0116] When calculating the time-series difference error of the chaos-enhanced version, the system combines the current reward, the discounted state value after chaos modulation, and the estimated current state value. In a practical application case, the immediate reward for a certain state is 8.5, the estimated current state value is 15.2, the estimated next state value is 20.3, and the discounted value after chaos modulation is 0.87. Therefore, the calculated time-series difference error of the chaos-enhanced version is 10.16. Compared to the traditional time-series difference error of 8.76, the chaos-enhanced version exhibits greater volatility and exploratory nature.

[0117] When constructing the policy gradient based on the chaotic-enhanced temporal difference error and fractal dimension, the two are weighted and combined, with the fractal dimension weighted at 0.25 and the temporal difference error weighted at 0.75. The constructed policy gradient is used for updating the parameters of the Actor network, with an initial learning rate of 0.003. Simultaneously, the chaotic-enhanced temporal difference error and the chaotic modulation term are jointly input into the Critic network to update the state value estimation, with an initial learning rate of 0.005 for the Critic network. The learning rates of both networks are adaptively adjusted according to the training process. When the cumulative reward change rate over five consecutive rounds is less than 3%, the learning rate is reduced to 0.85 times the original; when the change rate is greater than 10%, the learning rate is increased to 1.1 times the original, but the maximum cannot exceed twice the initial learning rate, and the minimum cannot be less than 0.1 times the initial learning rate.

[0118] When constructing the final optimized control strategy based on the updated state value estimates, the system adjusts the degree of exploration in action selection using a temperature parameter. The initial temperature parameter is set to 1.0, and gradually decreased to 0.2 during training, at a rate of 0.05 decrease per 1000 training steps. At low temperature values ​​(close to 0.2), the system tends to select the action with the highest estimated value; at high temperature values ​​(close to 1.0), the system is more inclined to make exploratory selections. In practical applications, when the temperature parameter is 0.8, the probability ratio of selecting the optimal action to the second-best action is approximately 3:1; when the temperature parameter decreases to 0.3, this ratio increases to approximately 9:1.

[0119] The entire optimization control strategy outputs a probabilistic strategy through an action-value function, where the action selection probability is proportional to the action value. In the early stages of training, the system is highly exploratory, and the temperature parameter is high, resulting in a relatively flat distribution of action selections. As training progresses and the temperature parameter decreases, action selection gradually concentrates on high-value actions. Finally, when the cumulative reward stabilizes at over 95% of the target value for 50 consecutive rounds, the strategy optimization is considered complete.

[0120] In one optional implementation, the optimized control strategy identifies the parameters of the system state-space model in real time based on a recursive least squares algorithm, solves the optimized control sequence in the rolling time domain, and introduces an adaptive law to dynamically adjust the controller parameters, including:

[0121] The system input and output data are collected, and an initial parameter estimation model is constructed using a recursive least squares algorithm. The recursive least squares algorithm iteratively updates the system parameter vector through the Kalman gain matrix. Wavelet decomposition is performed on the system input and output data to decompose the system input and output data into different frequency band features. The different frequency band features are combined with wavelet basis functions to obtain a multi-scale signal representation.

[0122] Based on the multi-scale signal representation, a parameter identification model is constructed in each frequency band. The parameter identification model adopts the iterative structure of the recursive least squares algorithm and adjusts the Kalman gain matrix according to the frequency band characteristics.

[0123] The identification results of the parameter identification model are input into the radial basis function neural network. The network parameters of the radial basis function neural network are adjusted by the gradient descent method. The update step size of the gradient descent method is calculated based on the identification error.

[0124] The output of the radial basis function neural network is subjected to distributed optimization calculation, and the parameters of the recursive least squares algorithm are corrected based on the results of the distributed optimization calculation; a Lyapunov function is constructed based on the results of the distributed optimization calculation, and the update equation of the controller parameters is designed according to the Lyapunov function, and the update equation is solved by the Kalman gain matrix;

[0125] An optimal control sequence is generated based on the controller parameters. This optimal control sequence is optimized using the gradient descent method and then used for the online control of the system.

[0126] like Figure 4 As shown, the method further includes:

[0127] The system collects input and output data, including sensor measurements and control input signals. Taking a temperature control system as an example, the system output is the temperature sensor measurement, and the system input is the heater power setpoint. The sampling period is set to 0.1 seconds, and 300 data points are continuously collected to form an initial dataset.

[0128] An initial parameter estimation model is constructed using a recursive least squares algorithm. This algorithm iteratively updates the system parameter vector using the Kalman gain matrix. In practice, a system order of 2 is chosen, the forgetting factor is set to 0.98, the diagonal elements of the initial covariance matrix are set to 100, and the initial parameter vector is set to zero. Taking a temperature control system as an example, the recursive least squares algorithm is applied to the collected data to obtain initial estimates of the system parameters [0.82, -0.15, 0.67, 0.08], which describe the dynamic characteristics of the system.

[0129] Wavelet decomposition is performed on the system's input and output data. Using the db4 wavelet basis function, the system's input and output data are decomposed into three levels to obtain features in different frequency bands. The decomposed signal includes low-frequency approximation components and high-frequency detail components. After wavelet decomposition, the input and output data of the temperature control system yields the low-frequency approximation component A3 and the high-frequency detail components D1, D2, and D3. These different frequency band features, combined with the wavelet basis function, form a multi-scale signal representation of the system.

[0130] Parameter identification models are constructed for each frequency band based on multi-scale signal representation. A recursive least squares algorithm is applied to the data in each frequency band. A forgetting factor of 0.99 is used for the low-frequency component A3, 0.97 for the mid-frequency components D3 and D2, and 0.95 for the high-frequency component D1. The Kalman gain matrix for each frequency band is adjusted according to the band characteristics, increasing the gain coefficient for high-frequency components and decreasing the gain coefficient for low-frequency components. The temperature control system identifies the parameters as follows: low-frequency A3 parameters [0.85, -0.12, 0.70, 0.06], mid-frequency D3 parameters [0.25, 0.35, 0.48, 0.12], mid-frequency D2 parameters [0.18, 0.22, 0.38, 0.15], and high-frequency D1 parameters [0.08, 0.05, 0.12, 0.04].

[0131] The identification results from the parameter identification model are input into a radial basis function neural network. This network contains 10 hidden neurons, with a Gaussian radial basis function and an initial width parameter set to 0.8. The network parameters are adjusted using gradient descent, with an initial learning rate of 0.05. In practical applications, the update step size of the gradient descent method is dynamically calculated based on the identification error. When the identification error is greater than 0.1, the update step size increases to 0.08; when the identification error is less than 0.01, the update step size decreases to 0.02. Through this method, the temperature control system parameters are further optimized, with the network output parameters [0.84, -0.13, 0.69, 0.07], reducing the error with the actual system parameters to within 3%.

[0132] Distributed optimization computation was performed on the output of the radial basis function neural network. The optimization task was assigned to three parallel computing units, each handling a subset of parameters. The first unit handled autoregressive parameters, the second handled control parameters, and the third handled cross-term parameters. Each unit used the alternating direction multiplier method for local optimization, and the results were summarized after 20 iterations. The parameters of the temperature control system after distributed optimization were [0.83, -0.14, 0.68, 0.075], further improving the optimization accuracy.

[0133] The Lyapunov function is constructed based on the results of distributed optimization computation. A quadratic form function is chosen as the Lyapunov function, which is the product of the system state vector and the positive definite matrix. The initial values ​​of the diagonal elements of the positive definite matrix are set to [2.0, 1.5]. The controller parameter update equation is designed based on the Lyapunov function, and this update equation is solved using the Kalman gain matrix. During the update process, the adjustment range of the controller parameters is limited to within 10% of the original value to ensure system stability. The controller parameters of the temperature control system are updated from the initial value [0.5, 0.3] to [0.54, 0.28].

[0134] The optimal control sequence is generated based on the updated controller parameters. The prediction time domain length is set to 10 steps, and the control time domain length is set to 3 steps. The optimal control sequence is optimized using gradient descent, with an initial step size of 0.03, which is gradually decreased to 0.01 as the number of iterations increases. The target temperature of the temperature control system is set to 70℃, and the current temperature is 65℃. The optimized control sequence is [85%, 75%, 60%], representing the setpoint of the heater power over three control cycles. This control sequence is applied to the online control of the system, enabling the system output to smoothly transition to the target value, with overshoot controlled within 2%, and the settling time shortened by 30% compared to traditional PID control.

[0135] Using the above methods, accurate identification of system parameters and generation of optimal control sequences were achieved, significantly improving system control performance. The entire control process runs on a standard industrial control computer, with a single control cycle calculation time of less than 50 milliseconds, meeting real-time control requirements.

[0136] In one optional implementation, the output of the radial basis function neural network is subjected to distributed optimization computation, and the parameters of the recursive least squares algorithm are corrected based on the results of the distributed optimization computation; constructing the Lyapunov function based on the results of the distributed optimization computation includes:

[0137] The system input and output data are fed into the radial basis function neural network for processing. The radial basis function neural network performs a nonlinear mapping on the system input and output data through network weights, center vectors, and expansion constants to obtain the neural network output results.

[0138] A distributed optimization network is constructed based on the output of the neural network. The distributed optimization network uses a graph topology to connect multiple computing nodes. Each computing node independently calculates the local gradient based on neighborhood information and transmits the local gradient information between computing nodes through a consensus protocol. The output of the neural network is iteratively optimized based on the local gradient and the local gradient information.

[0139] The parameters of the recursive least squares algorithm are corrected based on the output of the neural network. The parameter correction process includes: constructing a parameter error covariance matrix based on the output of the neural network; calculating the Kalman gain using the parameter error covariance matrix; updating the parameter estimate based on the Kalman gain; and reconstructing the parameter error covariance matrix based on the parameter estimate.

[0140] The Lyapunov function is constructed using the reconstructed parameter error covariance matrix. The construction process of the Lyapunov function includes: mapping the output of the neural network and the reconstructed parameter error covariance matrix to the error space to obtain the state error vector; constructing a quadratic form function of the state error vector to obtain the system error term; constructing the parameter estimation error of the recursive least squares algorithm as the parameter error term; and combining the system error term and the parameter error term to form the complete Lyapunov function.

[0141] The system acquires input and output data, including physical quantities such as temperature, pressure, and velocity collected by sensors, as well as corresponding system response data. After preprocessing, the acquired data is fed into a radial basis function neural network for further processing.

[0142] A radial basis function (RBF) neural network comprises an input layer, hidden layers, and an output layer. The input layer receives preprocessed system input data. The hidden layers contain multiple RBF nodes, each with a specific center vector and a spread constant. In this embodiment, a Gaussian RBF is used as the activation function. The network contains 50 hidden layer nodes. The initial values ​​of the center vectors are determined using K-means clustering, and the spread constant is set to 1.5 times the average distance between adjacent center vectors. For the input data, the Euclidean distance to each center vector is calculated and converted into node output values ​​using the Gaussian RBF. The output layer performs a weighted sum of the outputs of the hidden layer nodes using a weight vector to obtain the final output of the neural network.

[0143] After obtaining the neural network output, a distributed optimization network is constructed to optimize the result. The distributed optimization network uses a graph topology to connect multiple computing nodes, each node representing an independent computing unit. In this embodiment, a ring topology is used to connect 10 computing nodes, with each node communicating only with its two adjacent nodes. Each computing node receives a portion of the neural network's output and independently calculates its local gradient. The local gradient calculation is based on the node's local data and an objective function, defined as the mean squared error between the neural network output and the actual system output.

[0144] After calculating the local gradient, each computation node exchanges local gradient information through a consensus protocol. The consensus protocol uses a weighted averaging method, where each node receives gradient information from neighboring nodes according to preset weights and updates its own gradient value. In practice, the weight matrix is ​​designed to ensure convergence, and the sum of its rows is 1. In each iteration, a node combines its own gradient with the received gradients from neighboring nodes according to their weights to generate a new gradient estimate. Through multiple iterations, the gradient estimates of each node gradually converge, thus achieving distributed optimization.

[0145] The iterative optimization process employs an adaptive step-size strategy, with an initial step size set to 0.01, dynamically adjusted based on the gradient change rate. Optimization is considered complete when the gradient change over three consecutive iterations is less than a preset threshold of 0.0001. Experiments show that, in typical scenarios, after 10-15 iterations, the gradient difference between nodes is less than 0.0005, indicating successful convergence of the optimization process.

[0146] After optimization, the parameters of the recursive least squares algorithm are corrected using the neural network output. Parameter correction first constructs the parameter error covariance matrix P, with the initial value set to the identity matrix multiplied by a large coefficient of 100, representing the uncertainty of the initial estimate. Based on the optimized neural network output, the Kalman gain K is calculated. The Kalman gain calculation involves multiplying the covariance matrix P by the system regression vector, and normalization is used to enhance numerical stability.

[0147] The parameter estimates are updated using the Kalman gain, with the update amount being the product of the Kalman gain and the prediction error. The prediction error is defined as the difference between the actual system output and the predicted output of the current parameter estimate. After the parameter update, the parameter error covariance matrix is ​​reconstructed based on the new parameter estimates. The reconstruction process uses a matrix update formula to ensure the positive definiteness of the covariance matrix, where the diagonal elements represent the uncertainty of the parameter estimates.

[0148] After parameter correction, a Lyapunov function is constructed using the reconstructed parameter error covariance matrix. The construction process first maps the neural network output and the reconstructed parameter error covariance matrix to the error space, obtaining the state error vector e. The state error vector represents the deviation between the system's current state and the desired state. Based on the state error vector, a quadratic function V1 = e is constructed. T ×P×e serves as the systematic error term for the Lyapunov function.

[0149] Parameter estimation error of recursive least squares algorithm Defined as the difference between the true and estimated values ​​of the parameters, the parameter error term V2 is constructed. T ×Γ (-1) × Γ is a positive definite diagonal matrix, and the diagonal elements are the learning rates of each parameter, which are set to values ​​between 0.1 and 0.5 in this embodiment. The system error term V1 and the parameter error term V2 are combined to form the complete Lyapunov function V = V1 + V2.

[0150] In practical applications, the constructed Lyapunov function is used to analyze system stability. The system is asymptotically stable when the time derivative of the Lyapunov function satisfies the negative definite condition. By adjusting the forgetting factor (ranging from 0.95 to 0.99) and learning rate parameters of the recursive least squares algorithm, the negative definiteness of the Lyapunov function is ensured, thereby guaranteeing stable convergence of the system. Experimental results show that in typical industrial control applications, this method reduces parameter estimation error by approximately 25% and shortens system response time by approximately 30% compared to the traditional single recursive least squares algorithm, effectively improving system stability and control accuracy.

[0151] Figure 5 A bar chart comparing the distributed optimization efficiency of radial basis function neural networks in embodiments of the present invention:

[0152] This figure compares the performance of three different methods across three dimensions: computational efficiency, convergence performance, and robustness. The traditional RBF algorithm achieves 72.5% computational efficiency, the distributed optimization algorithm improves to 89.3%, and the Lyapunov correction algorithm further improves to 91.2%, indicating a significant improvement in computational speed after optimization. In terms of convergence performance, the traditional RBF algorithm achieves 68.4%, the distributed optimization algorithm improves to 85.7%, and the Lyapunov correction algorithm reaches 93.5%, demonstrating better convergence characteristics in the improved algorithms. Regarding robustness, the traditional RBF algorithm achieves 75.8%, the distributed optimization algorithm improves to 82.6%, and the Lyapunov correction algorithm reaches 94.7%, proving that the optimized algorithms have stronger anti-interference capabilities and system stability. Overall, by introducing distributed optimization and Lyapunov stability analysis, the algorithms achieve significant improvements in all three key performance indicators, with the Lyapunov correction algorithm showing the best performance, fully validating the effectiveness of the optimization scheme.

[0153] This invention discloses a mechanically stacked tandem photovoltaic module structure design using rails and clamps, and a tandem photovoltaic module encapsulated as a single unit using an encapsulating film. The biggest advantage of this mechanically stacked perovskite / crystalline silicon tandem photovoltaic module design using rails and clamps is that when any sub-module fails, it can be easily replaced with a new, functional sub-module. The advantage of encapsulating the perovskite and crystalline silicon sub-cell strings into a single tandem photovoltaic module using high-temperature resistant, insulating, and high-transmittance materials and an encapsulating film is that it creates an organic whole, resulting in less shading and less material usage, thus achieving higher conversion efficiency and lower cost. Even when a perovskite sub-cell string fails, the tandem module can be flipped over, and the crystalline silicon cells on the back side can still operate independently and continue generating electricity. Figure 6 As shown,

[0154] A four-terminal perovskite / crystalline silicon mechanically stacked photovoltaic module structure was designed, comprising an upper perovskite sub-module and a lower crystalline silicon sub-module. For example... Figure 6 As shown in (a), this mechanically stacked perovskite / crystalline silicon tandem photovoltaic module consists of perovskite sub-modules and crystalline silicon sub-modules independently packaged as sub-modules, which are then connected into a single module via rails and clamping blocks. This is a mechanical connection mode. Figure 6 (b) shows another design for a perovskite / crystalline silicon tandem photovoltaic module. This module separates the perovskite sub-cell strings and the crystalline silicon sub-cell strings using a high-temperature resistant, insulating, and high-transmittance sheet material, then encapsulates them into a single module using an encapsulating film. The upper perovskite module uses a wide bandgap material with an output voltage range of 300-600V; the lower crystalline silicon module uses a narrow bandgap material with an output voltage range of 200-400V. In this design, the perovskite and crystalline silicon sub-cells each have independently led-out positive and negative electrodes, forming a four-terminal independent output structure, achieving electrical isolation between the two sub-modules.

[0155] The intelligent current-voltage dynamic matching system mainly consists of the following functional modules:

[0156] The dual-channel independent MPPT control unit employs a time-division multiplexing MPPT algorithm, switching the tracking channel between the perovskite and crystalline silicon sub-modules every 15ms to achieve real-time capture of the maximum power point of both sub-modules. This control unit integrates a DSP-based hybrid logic controller, dynamically adjusting the MPPT tracking step size based on real-time irradiance data. In actual testing, this control unit performed well under low-light conditions (irradiance <200W / m²). 2 The tracking accuracy reached 99.2%, which is 5.7 percentage points higher than the traditional MPPT algorithm.

[0157] An adaptive voltage matching network employs a bidirectional DC / DC converter across four independent output ports to achieve dynamic boost / buck voltage matching. This network dynamically aligns the output voltage of the perovskite module (300-600V) with that of the crystalline silicon module (200-400V) to the inverter's required input voltage range (500-800V) through PWM duty cycle adjustment (in steps from 0.1% to 99.9%). Tests show that this voltage matching network achieves a conversion efficiency ≥98.5% and a voltage regulation speed ≤10ms, effectively handling scenarios with rapidly changing light intensity.

[0158] The intelligent power allocation logic operates based on a pre-established irradiance-current mapping model. This module prioritizes outputting power from the high irradiance end, while energy generated from the low irradiance end is temporarily stored in a supercapacitor before being output uniformly. Real-world application data shows that this intelligent power allocation strategy improves the system's daily efficiency by 12-18%.

[0159] The irradiance sensor has a spectral response range of 300-1200nm and a measurement accuracy of ±2%, enabling real-time monitoring of environmental parameter changes and feedback to the control system. The supercapacitor energy storage unit adopts a 100F / 100V specification and has a charge-discharge cycle life exceeding 10. 6 Next, under conditions of abrupt change in irradiance (e.g., from 1000 W / m²), 2 Mutation to 500W / m 2 It can control system power fluctuations within ±3%. The communication module adopts a Zigbee+PLC hybrid communication method, with a communication latency of less than 5ms, and can be seamlessly adapted to existing photovoltaic power plant cascade control systems.

[0160] The dynamic matching algorithm achieves dynamic power balancing time of less than 50ms at both ends of the system, with matching error controlled within 1.5%. When irradiance changes abruptly, the algorithm smooths power output through a supercapacitor buffering mechanism. Under cloudy and rainy weather conditions, the system efficiency degradation rate is reduced from 35% in the traditional design to 18%, significantly improving power generation efficiency under adverse weather conditions.

[0161] Empirical data from the 50MW Three Gorges Demonstration Base in Qinghai Province show that the photovoltaic system using this invention generates 23.7% more power per day than traditional four-terminal modules. Furthermore, due to the adoption of an adaptive communication protocol, the system is compatible with existing photovoltaic power plant communication architectures, reducing retrofit costs by over 40%.

[0162] Practical application tests show that the combination of this four-terminal perovskite / crystalline silicon mechanical stacking structure with an intelligent current-voltage dynamic matching system not only effectively solves the current-voltage mismatch problem between perovskite and crystalline silicon sub-modules, but also improves the irradiance response speed by 5 times through time-division multiplexing MPPT and dynamic voltage matching technology. The hybrid energy storage buffer mechanism effectively solves the system mismatch problem caused by instantaneous light intensity fluctuations.

[0163] A second aspect of the present invention provides an electronic device, comprising:

[0164] processor;

[0165] Memory used to store processor-executable instructions;

[0166] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0167] A third aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0168] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A voltage adaptive dynamic matching control method for perovskite / crystalline silicon tandem photovoltaic modules, characterized in that, include: Collect the output voltages of the perovskite sub-cells and the crystalline silicon sub-cells in the perovskite / crystalline silicon tandem photovoltaic module, and calculate the real-time voltage difference between the output voltages of the perovskite sub-cells and the crystalline silicon sub-cells. A voltage difference compensation model based on a graph neural network is constructed. This model uses perovskite and crystalline silicon sub-cells as nodes in the graph network, and the voltage difference as edge features. The node states are iteratively updated through a message passing mechanism, and the node weights are adaptively adjusted using an attention mechanism to output the optimal voltage compensation value. This includes: A voltage difference compensation model based on a graph neural network is constructed, in which the perovskite sub-cell and the crystalline silicon sub-cell are constructed as graph network nodes to generate node state vectors; the output voltage difference between the perovskite sub-cell and the crystalline silicon sub-cell is calculated, and the output voltage difference is used as an edge feature. An initial edge feature representation is constructed using a Hodgkin-Huxley neuron model. The Hodgkin-Huxley neuron model dynamically represents the initial edge feature representation through membrane capacitance, ion channel conductance, and reversal potential. The initial edge feature representation is input into the ion channel module of the Hodgkin-Huxley neuron model. Based on the sodium and potassium ion channel characteristics of the ion channel module, the initial edge feature representation is enhanced to form an enhanced edge feature representation with long-range correlation. A message passing function is constructed based on the enhanced edge feature representation. The message passing function receives the node state vector and the enhanced edge feature representation as input and generates node state update information. The node state update information is input into the synaptic plasticity module of the Hodgkin-Huxley neuron model. The node state update information is adaptively adjusted through the dynamic threshold mechanism of the synaptic plasticity module to obtain the adjusted node state. An attention mechanism is applied to the adjusted node states. The node state similarity is calculated and normalized to obtain the node attention coefficient. The perovskite sub-cell node states and crystalline silicon sub-cell node states with the node attention coefficient are input into a multilayer sensor to obtain the optimal voltage compensation value. The output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell are compensated according to the optimal voltage compensation value to obtain the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage. A system state-space model is constructed based on the compensated perovskite sub-cell voltage and the compensated crystalline silicon sub-cell voltage. A comprehensive reward signal is generated based on the system state-space model. The optimized control strategy is obtained by updating the state value estimate using a priority experience replay mechanism, including: A system state-space model is constructed, which includes the voltage states of the perovskite sub-cell and the crystalline silicon sub-cell. A prediction cost function is constructed based on the system state-space model. An initial control strategy is calculated based on the prediction cost function. An Actor-Critic agent is configured for the perovskite sub-cell and the crystalline silicon sub-cell according to the initial control strategy. The Actor network of the Actor-Critic agent calculates the expected state reward based on the state action value function. The expected state reward is input into a hierarchical reward mechanism. The hierarchical reward mechanism calculates a local reward function based on the voltage tracking error of a single sub-cell, a global reward function based on the joint tracking error of the voltages of two sub-cells, and a time-series reward function based on the time derivative of the voltage difference. The local reward function, the global reward function, and the time-series reward function are then combined to form a comprehensive reward signal. A priority experience replay mechanism is constructed using the comprehensive reward signal. This mechanism calculates the policy gradient through time-series differential error and simultaneously updates the state value estimate of the Actor-Critic agent using the policy gradient to obtain an optimized control policy. This optimized control policy identifies the parameters of the system's state-space model in real time based on a recursive least squares algorithm and solves for the optimized control sequence in the rolling time domain. Simultaneously, an adaptive law is introduced to dynamically adjust the controller parameters, including: The system input and output data are collected, and an initial parameter estimation model is constructed using a recursive least squares algorithm. The recursive least squares algorithm iteratively updates the system parameter vector through the Kalman gain matrix. Wavelet decomposition is performed on the system input and output data to decompose the system input and output data into different frequency band features. The different frequency band features are combined with wavelet basis functions to obtain a multi-scale signal representation. Based on the multi-scale signal representation, a parameter identification model is constructed in each frequency band. The parameter identification model adopts the iterative structure of the recursive least squares algorithm and adjusts the Kalman gain matrix according to the frequency band characteristics. The identification results of the parameter identification model are input into the radial basis function neural network. The network parameters of the radial basis function neural network are adjusted by the gradient descent method. The update step size of the gradient descent method is calculated based on the identification error. The output of the radial basis function neural network is subjected to distributed optimization calculation, and the parameters of the recursive least squares algorithm are corrected based on the results of the distributed optimization calculation; a Lyapunov function is constructed based on the results of the distributed optimization calculation, and the update equation of the controller parameters is designed according to the Lyapunov function, and the update equation is solved by the Kalman gain matrix; An optimal control sequence is generated based on the controller parameters. The optimal control sequence is optimized using the gradient descent method and used for online control of the system. Based on the output control quantity of the optimized control strategy, the operating state of the drive circuits for the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell is adjusted to achieve optimal voltage matching.

2. The method according to claim 1, characterized in that, The output voltage difference between the perovskite sub-cell and the crystalline silicon sub-cell is calculated. This output voltage difference is used as a side feature. An initial side feature representation is constructed using the Hodgkin-Huxley neuron model, including: The voltage difference between the output voltage of the perovskite sub-cell and the output voltage of the crystalline silicon sub-cell is calculated. The voltage difference is multiplied by a mapping coefficient and superimposed with the resting potential to obtain the mapped membrane potential. The mapping coefficient is determined by the amplitude range of the voltage difference and the physiological range of the membrane potential of biological neurons. An ion channel dynamics model is constructed based on the mapped membrane potential, and the time rate of change of the mapped membrane potential is calculated using the membrane capacitance per unit area, the maximum conductance of each ion channel, the equilibrium potential of each ion, and the external input current. Based on the ion channel kinetic model, a kinetic equation for the gated variables is established. The voltage-dependent rate constants of each gated variable are calculated using the mapped membrane potential. The voltage-dependent rate constants include forward rate constants and reverse rate constants. The steady-state values ​​and time constants of each gated variable are calculated based on the forward rate constants and the reverse rate constants. The steady-state values ​​and time constants are substituted into the exponential form kinetic equation to solve for the time evolution of each gated variable. The time evolution is numerically integrated using the Runge-Kutta method to form an initial edge feature characterization.

3. The method according to claim 1, characterized in that, A priority experience replay mechanism is constructed using the comprehensive reward signal. This mechanism calculates the policy gradient through temporal differential error and simultaneously updates the state value estimate of the Actor-Critic agent using the policy gradient to obtain the optimized control policy, including: The fractal dimension of the strategy trajectory is calculated for the comprehensive reward signal. An empirical priority evaluation function is constructed based on the fractal dimension, temporal differential error, and state transition similarity. The empirical priority evaluation function is used to determine the importance of empirical samples. The Logistic chaotic sequence is introduced into the time-series difference error calculation process. The Logistic chaotic sequence is modulated and discounted with a chaos intensity parameter to obtain a chaotic enhanced time-series difference error. The chaotic enhanced time-series difference error is used to evaluate the temporal correlation of state-action pairs. The policy gradient is constructed based on the chaotic enhanced temporal difference error and the fractal dimension. The policy gradient is input into the Actor network for parameter update. At the same time, the chaotic enhanced temporal difference error and the chaotic modulation term are input into the Critic network to update the state value estimate. The learning rates of the Actor network and the Critic network are adaptively adjusted according to the training process. The final optimized control strategy is constructed based on the updated state value estimate. The final optimized control strategy adjusts the exploration degree of action selection through temperature parameters to achieve a probabilistic strategy output based on the action value function.

4. The method according to claim 1, characterized in that, The output of the radial basis function neural network is subjected to distributed optimization calculation, and the parameters of the recursive least squares algorithm are corrected based on the results of the distributed optimization calculation. Constructing the Lyapunov function based on the results of the distributed optimization computation includes: The system input and output data are fed into the radial basis function neural network for processing. The radial basis function neural network performs a nonlinear mapping on the system input and output data through network weights, center vectors, and expansion constants to obtain the neural network output results. A distributed optimization network is constructed based on the output of the neural network. The distributed optimization network uses a graph topology to connect multiple computing nodes. Each computing node independently calculates the local gradient based on neighborhood information and transmits the local gradient information between computing nodes through a consensus protocol. The output of the neural network is iteratively optimized based on the local gradient and the local gradient information. The parameters of the recursive least squares algorithm are corrected based on the output of the neural network. The parameter correction process includes: constructing a parameter error covariance matrix based on the output of the neural network; calculating the Kalman gain using the parameter error covariance matrix; updating the parameter estimate based on the Kalman gain; and reconstructing the parameter error covariance matrix based on the parameter estimate. The Lyapunov function is constructed using the reconstructed parameter error covariance matrix. The construction process of the Lyapunov function includes: mapping the output of the neural network and the reconstructed parameter error covariance matrix to the error space to obtain the state error vector; constructing a quadratic form function of the state error vector to obtain the system error term; constructing the parameter estimation error of the recursive least squares algorithm as the parameter error term; and combining the system error term and the parameter error term to form the complete Lyapunov function.

5. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 4.

6. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 4.