Solar charging adaptive control method based on improved MPPT method
By combining generative adversarial networks and deep reinforcement learning policy networks, a virtual dataset is generated and trained in stages, solving the problems of low tracking accuracy and insufficient adaptability of existing MPPT methods in dynamic environments, and realizing efficient solar charging control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- POMPEY SMART TECH (JIANGSU) CO LTD
- Filing Date
- 2025-06-06
- Publication Date
- 2026-05-05
AI Technical Summary
Existing MPPT methods have low tracking accuracy in dynamic environments and lack adaptive control strategies, making them unable to adapt to changes in environmental complexity and battery aging conditions.
Generative adversarial networks are used to simulate dynamic environmental scenarios to generate virtual datasets. The action space is trained through deep reinforcement learning strategy networks, and the MPPT duty cycle and charging current are adjusted in stages. The control strategy is optimized by combining real-time environmental data.
It improves tracking accuracy and control strategy adaptability in dynamic environments, enhances dynamic response speed and charging efficiency, and solves the problem of strategy failure caused by environmental dynamism and hardware aging in existing methods.
Smart Images

Figure CN120872093B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of solar charging control technology, and in particular to an adaptive control method for solar charging based on an improved MPPT method. Background Technology
[0002] Existing MPPT methods (such as perturbation and observation (P&O) and incremental conductance method (IncCond) periodically sample the voltage and current of the photovoltaic (PV) panel to calculate the power gradient and determine the direction of duty cycle adjustment, thereby approximating the maximum power point (MPP). P&O periodically perturbs the operating voltage of the PV array (e.g., increasing or decreasing the duty cycle) and compares the output power change before and after the perturbation to determine if the perturbation direction is correct. Incremental conductance method (IncCond) calculates the difference between the instantaneous conductance (dI / dV) of the PV array and the load conductance to determine if the operating point is at the MPP. P&O and IncCond algorithms have clear logic, low hardware implementation costs, and are suitable for scenarios with slow changes in illumination (such as fixed PV power plants), thus they are widely used in industrial fields.
[0003] However, existing MPPT methods are usually based on fixed step size or preset adjustment period, which have shortcomings in terms of data-driven and dynamic strategy generation: First, fixed step size cannot adapt to changes in environmental complexity, resulting in low tracking accuracy in high dynamic scenarios; Second, control parameters (such as duty cycle adjustment range and charging step size) are static and cannot be dynamically adjusted according to battery aging state and environmental complexity, resulting in a lack of adaptability in the control strategy. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a solar charging adaptive control method based on an improved MPPT method, which solves the problems of low dynamic environment tracking accuracy and lack of adaptability in the control strategy.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, the present invention provides a solar charging adaptive control method based on an improved MPPT method, comprising,
[0008] Generative adversarial networks are used to simulate dynamic environment scenarios and generate virtual datasets.
[0009] The virtual dataset is divided into three course phases, and the environment complexity index and action space of each course phase are defined. The deep reinforcement learning policy network is trained in the order of the course phases. The action space includes the MPPT duty cycle adjustment range and the charging current adjustment step size.
[0010] Collect real-world environmental data, calculate environmental complexity indicators, match corresponding course stages through a deep reinforcement learning strategy network, and obtain MPPT duty cycle adjustment amount and charging current target value based on the action space associated with the course stage.
[0011] Based on the MPPT duty cycle adjustment and the target charging current, the switching frequency of the DC-DC converter and the charging threshold of the battery management unit in the solar charging hardware are controlled, and the operating data of the solar charging hardware is collected.
[0012] Based on the operational data of the solar charging hardware, the parameters of the generative adversarial network are updated, and the division of course phases and action space parameters are optimized.
[0013] As a preferred embodiment of the solar charging adaptive control method based on the improved MPPT method described in this invention, the generation of the virtual dataset refers to the generation of a virtual dataset based on a parameter space defined by a dynamic environment scenario and combined with a generator of a generative adversarial network.
[0014] As a preferred embodiment of the solar charging adaptive control method based on the improved MPPT method described in this invention, the step of dividing the virtual dataset into three course stages refers to calculating the comprehensive score of the virtual dataset based on the virtual dataset.
[0015] The overall score of the virtual dataset is divided into low, medium and high complexity ranges, and the virtual dataset is divided into low, medium and high course stages.
[0016] As a preferred embodiment of the solar charging adaptive control method based on the improved MPPT method described in this invention, wherein: the definition of the environmental complexity index and action space for each course stage refers to setting the environmental complexity index for low, medium and high course stages based on the virtual dataset and low, medium and high complexity ranges;
[0017] The action space is set for each course stage based on the environmental complexity index.
[0018] As a preferred embodiment of the solar charging adaptive control method based on the improved MPPT method described in this invention, the deep reinforcement learning policy network trained in the order of course stages refers to using virtual datasets and action spaces corresponding to low, medium and high course stages, and training the deep reinforcement learning policy network in three stages in the order of low, medium and high.
[0019] As a preferred embodiment of the solar charging adaptive control method based on the improved MPPT method described in this invention, the step of matching the corresponding course stage through a deep reinforcement learning strategy network refers to comparing the calculated environmental complexity index with the environmental complexity indices of low, medium, and high course stages to match the corresponding course stage.
[0020] As a preferred embodiment of the solar charging adaptive control method based on the improved MPPT method described in this invention, the specific steps for updating the parameters of the generative adversarial network are as follows:
[0021] The virtual dataset and the operating data of the solar charging hardware are input into the discriminator of the generative adversarial network, and the parameters of the discriminator of the generative adversarial network are updated by calculating the cross loss during the training process.
[0022] A virtual dataset is generated by the generator of a generative adversarial network (GAN). The adversarial loss between the generator and the discriminator of the GAN is calculated using cross-entropy loss. The parameters of the generator network are updated using the backpropagation algorithm.
[0023] As a preferred embodiment of the solar charging adaptive control method based on the improved MPPT method described in this invention, the specific steps for optimizing the course phase division and action space parameters are as follows:
[0024] A new virtual dataset is generated using the updated generative adversarial network generator, and a comprehensive score for the new virtual dataset is calculated.
[0025] Based on the comprehensive scores of the new virtual dataset, the courses are divided into low, medium, and high complexity ranges, and the low, medium, and high course stages are redefined.
[0026] The actual distribution of MPPT duty cycle adjustment values of the deep reinforcement learning strategy network output in each course stage was statistically analyzed, and the MPPT duty cycle adjustment range for the corresponding course stage was adjusted.
[0027] The execution error of the actual adjustment step size of the battery management unit in constant current charging mode is statistically analyzed, the mean of the execution error is calculated, and the charging current adjustment step size is adjusted.
[0028] In a second aspect, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the solar charging adaptive control method based on the improved MPPT method as described in the first aspect of the present invention.
[0029] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the solar charging adaptive control method based on the improved MPPT method as described in the first aspect of the present invention.
[0030] The beneficial effects of this invention are as follows: By generating virtual data through generative adversarial networks, the limitations of simulation models on the linear assumption of a single environmental variable are overcome, enabling the policy network to quickly optimize the duty cycle in complex environments during the pre-training stage, avoiding oscillations or misjudgments caused by traditional fixed step sizes, and improving the tracking accuracy of dynamic environments; by progressively training the deep reinforcement learning policy network in stages, real-time adaptation of the policy to hardware aging and environmental complexity is achieved, solving the problem of the lack of adaptability of existing control policies. Attached Figure Description
[0031] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This is a flowchart of an adaptive control method for solar charging based on an improved MPPT method.
[0033] Figure 2 This diagram illustrates the operation and update process of an adversarial network.
[0034] Figure 3 This is a diagram illustrating the course phases and training methods.
[0035] Figure 4 This is a schematic diagram of the architecture of a deep reinforcement learning policy network. Detailed Implementation
[0036] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0037] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0038] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0039] Reference Figures 1-4This paper presents an adaptive control method for solar charging based on an improved MPPT method, comprising the following steps:
[0040] S1: Use generative adversarial networks to simulate dynamic environment scenarios and generate virtual datasets.
[0041] The specific steps are as follows:
[0042] A generator and discriminator from a generative adversarial network (GAN) are loaded. An initial virtual dataset is generated using the generator from the GAN and historical environmental data. The discriminator iteratively optimizes the coupling correlation of the static parameters of the initial virtual dataset by combining the dynamic characteristics of the real environmental data, thus obtaining a dynamic environmental scenario. The generator learns parameters based on the static parameters in the historical environmental data. The discriminator establishes verification criteria based on the dynamic characteristics of the real environmental data. The real environmental feedback data refers to the real-time output power of the photovoltaic array, the battery charging and discharging efficiency, and the measured data of the temperature sensor collected during the operation of the solar charging hardware. The static parameters refer to the light intensity, temperature gradient, shading mode, and battery degradation law. The dynamic characteristics refer to sudden changes in light intensity, temperature fluctuations, shading distribution, and battery aging characteristics.
[0043] The parameter space required to generate a virtual dataset is defined based on the characteristics of sudden changes in light intensity, temperature gradient changes, local shading modes of photovoltaic arrays, and battery degradation curves in dynamic environmental scenarios.
[0044] The generator receives a Gaussian noise vector and combines it with the parameter space to generate a virtual dataset containing time-series data of light intensity, a temperature distribution matrix, shading location coordinates, and battery capacity decay curves.
[0045] The discriminator extracts the frequency of light intensity mutations, the spectrum of temperature fluctuations, the proportion of shaded area, and the trend of battery internal resistance from the virtual dataset. Based on the statistical distribution of light intensity mutations in real environmental data (such as the 95th percentile of historical data mutation amplitude), a mutation threshold is set. The number and magnitude of light intensity changes exceeding the mutation threshold per unit time in real environmental data and virtual dataset are counted. The spectral energy distribution of temperature time series data in real environmental data and virtual dataset are calculated by fast Fourier transform. The diffusion rate and spatial distribution pattern of shaded areas in real environmental data and virtual dataset are extracted by time series slope analysis combined with morphological image processing. The nonlinear coefficient of the battery capacity decay curve is fitted by nonlinear least squares method.
[0046] The Jensen-Shannon divergence metric was used to quantify the differences in the number of times light intensity changes exceeded the abrupt threshold and the distribution of temperature spectrum energy distribution between the real environmental data and the virtual dataset. The dynamic time bending distance of the shading diffusion rate and battery capacity decay curves in the real environmental data and the virtual dataset was calculated. The two-dimensional histogram matching method was used to evaluate the similarity of the spatial coverage area of the shading area coordinate matrix in the real environmental data and the virtual dataset.
[0047] By calculating the Pearson correlation coefficients of light abrupt changes, temperature fluctuations, shading diffusion, and battery aging characteristics, weighting coefficients are assigned to these characteristics. A consistency score is obtained by weighted summation of these characteristics.
[0048] A consistency score threshold is set based on the statistical distribution of the consistency score (e.g., the 95th percentile). If the consistency score is less than the consistency score threshold, the noise weight of the static parameters in the generator is adjusted, and the virtual dataset is iteratively generated until the consistency score is not less than the consistency score threshold.
[0049] It should also be noted that by simulating dynamic environment scenarios and generating virtual datasets through generative adversarial networks, the low accuracy of dynamic environment tracking caused by the static nature of environmental data in existing MPPT methods is solved. The adversarial training mechanism between the generator and discriminator can generate virtual scenes that highly match real-world environmental data (such as sudden changes in illumination, shading diffusion, battery aging, etc.), covering the coupling effects of multidimensional nonlinear environmental variables. This provides high-fidelity training samples for deep reinforcement learning policy networks and improves their generalization ability in dynamic environments.
[0050] S2: Divide the virtual dataset into three course phases, define the environment complexity index and action space for each course phase, and train the deep reinforcement learning policy network in the order of the course phases. The action space includes the MPPT duty cycle adjustment range and the charging current adjustment step size.
[0051] The specific steps are as follows:
[0052] The nonlinear coefficients of light abrupt change frequency, temperature spectrum energy distribution, shading diffusion rate, and battery capacity decay curve were extracted from the virtual dataset as dynamic characteristic parameters. The Z-score method was used to standardize the dynamic characteristic parameters. The dynamic characteristic parameters were weighted and summed based on the weight coefficients of light abrupt change, temperature fluctuation, shading diffusion, and battery aging to obtain a comprehensive score. Based on the quantile statistical distribution of the comprehensive score of the virtual dataset, it was divided into three complexity intervals: low, medium, and high, according to equal quantile intervals (such as 33% and 66% quantiles). The low, medium, and high quantile intervals correspond to the low, medium, and high complexity intervals, respectively. The virtual dataset was divided into low, medium, and high course stages according to the complexity intervals corresponding to the comprehensive score.
[0053] Set environmental complexity indicators (light change frequency, shading diffusion rate, and temperature fluctuation spectrum characteristics) for each course stage: For the lower course stage, set the light change frequency of the virtual dataset corresponding to the statistical upper limit of the lower quantile interval (e.g., the maximum value of the light change frequency in the virtual dataset with the lowest comprehensive score of 33%) as the light change frequency threshold of the lower course stage, set the average value of the shading diffusion rate in the lower quantile interval as the shading diffusion rate threshold of the lower course stage, and set the minimum value of the low-frequency energy proportion of temperature fluctuation in the lower quantile interval as the low-frequency energy proportion threshold of the lower course stage.
[0054] For the intermediate curriculum stage, the range from the minimum to the maximum value of the light change frequency in the intermediate quantile interval is set as the light change frequency screening range for the intermediate curriculum stage. The average value plus or minus the standard deviation of the shading diffusion rate in the intermediate quantile interval is used as the shading diffusion rate screening range for the intermediate curriculum stage. The minimum and maximum values of the low-frequency energy proportion of temperature fluctuation in the intermediate quantile interval are statistically analyzed to form the low-frequency energy proportion fluctuation range for the intermediate curriculum stage.
[0055] For the higher education stage, the light mutation frequency of the virtual dataset corresponding to the statistical lower limit of the high quantile interval (e.g., the minimum light mutation frequency in the virtual dataset with the highest comprehensive score of 33%) is set as the light mutation frequency threshold for the higher education stage. The peak value of the diffusion rate in the high quantile interval is used as the shading diffusion rate threshold for the higher education stage. The maximum value of the low-frequency energy proportion of temperature fluctuations in the high quantile interval is set as the low-frequency energy proportion threshold for the higher education stage.
[0056] Set the action space (MPPT duty cycle adjustment range and charging current adjustment step size) for each course stage: For the low course stage, select virtual datasets whose light change frequency does not exceed the light change frequency threshold, use the duty cycle range corresponding to the basic tracking capability of the DC-DC converter as the MPPT duty cycle adjustment range for the low course stage, and use the minimum controllable step size of the battery management unit as the charging current adjustment step size for the low course stage.
[0057] For the intermediate stage, virtual datasets with light change frequency within the light change frequency selection range are selected. The maximum safe duty cycle adjustment range of the DC-DC converter in the intermediate stage is tested and used as the MPPT duty cycle adjustment range in the intermediate stage. The average value of the nonlinear coefficient of the battery capacity decay curve in the middle quantile interval is used as the charging current adjustment step size in the intermediate stage.
[0058] For the advanced course stage, virtual datasets with light change frequency exceeding the light change frequency threshold are selected. The maximum allowable duty cycle range of the DC-DC converter is used as the MPPT duty cycle adjustment range for the advanced course stage, and the maximum safe adjustment step size of the battery management unit is used as the charging current adjustment step size for the advanced course stage.
[0059] Starting from the lower course stage, a deep reinforcement learning policy network is trained using the corresponding virtual dataset and action space: the virtual dataset of the lower course stage is input into the deep reinforcement learning policy network, and the action space is limited to the MPPT duty cycle adjustment range and charging current adjustment step size of the lower course stage; the average ratio of the actual output power of the photovoltaic array to the theoretical maximum power within the training cycle is calculated as the average MPPT efficiency; the proportion of battery charging over-limit times to the total number of charging times is counted as the battery overcharge rate; the training of the lower course stage is terminated when the average MPPT efficiency and battery overcharge rate meet the performance standards of the photovoltaic array and the battery safety protocol for 10 consecutive training cycles.
[0060] After training in the low-level course phase ends, the virtual dataset for the mid-level course phase is loaded, and the action space is expanded to include the MPPT duty cycle adjustment range and charging current adjustment step size for the mid-level course phase. The feature extraction layer of the deep reinforcement learning policy network is frozen, and only the fully connected layer parameters corresponding to the action space are trained. The shading diffusion rate tracking success rate is calculated within the training cycle. When the standard deviations of MPPT efficiency, battery overcharge rate, and shading diffusion rate tracking success rate are all less than the standard deviation threshold for 5 consecutive training cycles, the training in the mid-level course phase is terminated. The standard deviation threshold is set based on the convergence speed of the deep reinforcement learning policy network.
[0061] After the training in the intermediate course stage ends, the virtual dataset of the advanced course stage is loaded, and the action space is expanded to the MPPT duty cycle adjustment range and charging current adjustment step size of the advanced course stage. Extreme perturbation samples (such as peak frequency of sudden light change and extreme value of shading diffusion rate) are randomly injected into the virtual dataset of the advanced course stage. When the MPPT efficiency, battery overcharge rate and shading diffusion rate tracking success rate of the training cycle reach the safety threshold, and the standard deviation of the MPPT efficiency, battery overcharge rate and shading diffusion rate tracking success rate are in a convergent state, the training is completed. The safety threshold is set based on the safety limits of the DC-DC converter and battery management unit and the battery management protocol.
[0062] The architecture of the deep reinforcement learning policy network includes an input layer, a feature extraction layer, a policy decision layer, and a value function layer. The input layer receives and normalizes the virtual dataset. The feature extraction layer, consisting of a 1D convolutional layer, an LSTM layer, a 2D convolutional layer, and two fully connected layers, extracts features related to sudden changes in illumination, temperature distribution, and battery degradation. The policy decision layer fuses these features and generates action instructions for adjusting the MPPT duty cycle and charging current step size. The value function layer evaluates the value of the current state, providing a benchmark for calculating the advantage function for policy gradient optimization. The training algorithm for the deep reinforcement learning policy network employs proximal policy optimization (PPO).
[0063] It should also be noted that by employing a phased learning strategy (low → medium → high complexity), the problem of insufficient dynamic response speed caused by static control parameters in existing MPPT methods is resolved. Phased training allows the deep reinforcement learning policy network to gradually adapt to environmental scenarios from simple to complex, dynamically adjusting the action space (such as the duty cycle adjustment range and charging step size), avoiding gradient oscillation problems when directly training on high-complexity data, and ensuring that the control actions output by the policy network under different environmental complexities are always within a reasonable range, thus improving the policy convergence efficiency and stability in dynamic environments.
[0064] S3: Collect environmental data from the real environment, calculate the environmental complexity index, match the corresponding course stage through a deep reinforcement learning strategy network, and obtain the MPPT duty cycle adjustment amount and charging current target value based on the action space associated with the course stage.
[0065] The specific steps are as follows:
[0066] The photovoltaic array sensor group collects time-series data of light intensity, coordinate sequence of shaded areas, and temperature distribution matrix. After aligning by timestamp, the light intensity time-series data is standardized using the Z-score method. The mean distance of centroid movement between adjacent time points in the coordinate sequence of shaded areas is calculated to obtain the shading diffusion rate. The shading diffusion rate is linearly scaled to a unit interval, and the temperature distribution matrix is normalized using minimum-maximum normalization.
[0067] The cumulative number of times the light intensity change exceeds the mutation threshold within the statistical time window is used to obtain the mutation count. The mutation threshold is defined based on the maximum power point tracking sensitivity of the photovoltaic array and the standard deviation of the natural fluctuation of light intensity. The mutation count is converted into a frequency value per unit time to obtain the light mutation frequency. The geometric centroid coordinates (mean of two-dimensional plane coordinates) of the shading area coordinate sequence at each time point are calculated. The Euclidean distance between the geometric centroid coordinates of adjacent time points is calculated. The mean of the movement distance of the geometric centroid coordinates between all adjacent time points within the statistical time window is used to obtain the shading diffusion rate. The spatiotemporal data of the temperature distribution matrix is converted into time-series signals of each grid point. A Fast Fourier Transform (FFT) is performed on the time-series signal of each grid point to extract the energy value of the sampling frequency band. The sampling frequency band is determined according to the sampling rate of the temperature sensor and the temperature control requirements. The percentage of the total energy of all grid points in the sampling frequency band to the total energy of the frequency band is calculated to obtain the low-frequency energy ratio of temperature fluctuation.
[0068] The corresponding course stage is matched through a deep reinforcement learning strategy network: if the light change frequency does not exceed the light change frequency threshold of the low course stage, the shading diffusion rate does not exceed the shading diffusion rate threshold of the low course stage, and the low frequency energy proportion of temperature fluctuation is not less than the low frequency energy proportion threshold of the low course stage, then the current real environment is determined to be a low course stage.
[0069] If the light change frequency is within the light change frequency screening range of the intermediate curriculum stage, the shading diffusion rate is within the shading diffusion rate screening range of the intermediate curriculum stage, and the low-frequency energy ratio of temperature fluctuation is within the low-frequency energy ratio fluctuation range of the intermediate curriculum stage, then the current real environment is determined to be at the intermediate curriculum stage.
[0070] If the light change frequency is not less than the light change frequency threshold of the higher education stage, the shading diffusion rate is not less than the shading diffusion rate threshold of the higher education stage, and the proportion of low-frequency energy in temperature fluctuations does not exceed the low-frequency energy proportion threshold of the higher education stage, then the current real environment is determined to be the higher education stage.
[0071] By utilizing the motion space corresponding to each course stage, the MPPT duty cycle adjustment range and charging current adjustment step size are obtained.
[0072] Based on environmental data, the deep reinforcement learning strategy network outputs continuous action values within the duty cycle adjustment range of the corresponding course stage to determine the MPPT duty cycle adjustment amount.
[0073] The deep reinforcement learning strategy network generates discrete step size selection instructions based on the nonlinear coefficient of battery capacity decay and the current state of charge, under the constraint of the charging current adjustment step size in the corresponding course stage, and accumulates and calculates the target value of charging current.
[0074] It should also be noted that by calculating environmental complexity indicators (such as the frequency of sudden changes in illumination and the rate of shading diffusion) from real-world environmental data in real time and matching them with the corresponding course stages, the problem of the lack of adaptive adjustment capability in existing MPPT methods is solved. After dynamically matching the course stages, the deep reinforcement learning policy network generates precise MPPT duty cycle adjustments and charging current target values based on the associated action space, ensuring that control parameters (such as duty cycle and charging step size) are adapted to the dynamic characteristics of the environment in real time. This allows the system to quickly approach the maximum power point in scenarios such as sudden changes in illumination or battery aging, thereby improving dynamic response speed and charging efficiency.
[0075] S4: Control the switching frequency of the DC-DC converter and the charging threshold of the battery management unit in the solar charging hardware according to the MPPT duty cycle adjustment and the target value of the charging current, and collect the operating data of the solar charging hardware.
[0076] The specific steps are as follows:
[0077] The MPPT duty cycle adjustment is converted into the actual duty cycle of the PWM waveform signal of the DC-DC converter using a pulse width modulation (PWM) algorithm. Based on the actual duty cycle of the PWM waveform signal, the on and off times of the switching devices of the DC-DC converter are adjusted. The output voltage of the DC-DC converter is adjusted based on the volt-second balance principle to stabilize the output voltage of the photovoltaic array at the target operating point. The target value of the charging current and the charging current adjustment step size are sent to the battery management unit to trigger the constant current charging mode. The actual value of the charging current is adjusted in a discrete step manner according to the charging current adjustment step size to ensure that the rate of change of current meets the safety limits of the battery management unit.
[0078] Collect operational data from solar charging hardware: Real-time acquisition of output voltage, current, and power data from the DC-DC converter via voltage and current sensors; acquisition of actual battery state of charge, terminal voltage, temperature, and charging current values via the battery management unit.
[0079] A sliding window filtering algorithm is used to remove transient outliers caused by sensor noise from the operating data of the solar charging hardware. The removed transient outliers are replaced by the mean of the operating data of the solar charging hardware within the sliding window. The Z-score method is used to standardize the operating data of the solar charging hardware.
[0080] It should also be noted that by precisely controlling the switching frequency of the DC-DC converter and the charging threshold of the battery management unit through the PWM algorithm, and combining it with the sliding window filtering algorithm to eliminate sensor noise, the problem of reduced control accuracy caused by hardware response delay or data noise in existing MPPT methods is solved. Real-time acquisition and standardized processing of hardware operating data (such as voltage, current, and power) provides high-quality feedback for subsequent generative adversarial network parameter updates and curriculum strategy optimization, ensuring stability and reliability in dynamic environments.
[0081] S5: Based on the operational data of the solar charging hardware, update the parameters of the generative adversarial network and optimize the division of course phases and action space parameters.
[0082] The specific steps are as follows:
[0083] The operational data of the solar charging hardware is used as real samples and labeled as the "real" category (label 1). The virtual dataset is used as generated samples and labeled as the "fake" category (label 0). The real samples and generated samples are mixed in a 1:1 ratio to form the discriminator training batch.
[0084] The discriminator training batch is input into the discriminator, which extracts the feature vectors of the discriminator training batch through a multi-layer fully connected network. The feature vectors of the discriminator training batch are then subjected to nonlinear transformation and dimensionality reduction through multiple fully connected layers. After each fully connected layer, the ReLU activation function is used to enhance the nonlinear expressive power. In the last fully connected layer, the dimensionality-reduced feature vectors of the discriminator training batch are multiplied by the weight matrix, and a bias term is added to generate a scalar value.
[0085] The sigmoid function is used to map scalar values to the (0,1) interval, which is used as the probability that a sample belongs to the "true" category. A classification threshold is set based on the probability that a sample belongs to the "true" category. If the probability that a sample belongs to the "true" category is not less than the classification threshold, it is determined to be a true sample. If the probability that a sample belongs to the "true" category is less than the classification threshold, it is determined to be a generated sample.
[0086] The binary cross-entropy loss function is used to calculate the cross-loss between the probability of a sample belonging to the "true" category and label 1, and the cross-loss between the probability of a sample belonging to the "false" category and label 0. The average of the cross-loss between the probability of a sample belonging to the "true" category and label 1 and the cross-loss between the probability of a sample belonging to the "false" category and label 0 is calculated as the difference between the probability of a sample belonging to the "false" category being 1 and the probability of a sample belonging to the "true" category.
[0087] Based on the average loss, the backpropagation algorithm is used to calculate the gradient of the loss function with respect to the parameters of the discriminator. The gradient descent optimizer is then used to update the parameters of the discriminator according to the gradient direction of the learning rate and the gradient of the loss function with respect to the parameters of the discriminator, thereby improving the discriminator's ability to distinguish between real and generated data.
[0088] The discriminator parameters are fixed, and a virtual dataset is generated by the generator of a generative adversarial network. The virtual data is then input into the fixed discriminator to obtain the discriminator's judgment result. The adversarial loss between the discriminator's judgment result and the corresponding label is calculated using cross-entropy loss. The gradient of the generator's parameters is calculated using the backpropagation algorithm, and the generator's parameters are updated using the gradient descent optimizer based on the gradient direction of the generator's parameters and the learning rate.
[0089] A new virtual dataset is generated using the updated generative adversarial network generator. The comprehensive score of the new virtual dataset is calculated, the quantile distribution of the comprehensive score of the new virtual dataset is statistically analyzed, the low, medium and high complexity intervals are updated at equal quantile intervals, and the virtual dataset is re-divided into low, medium and high course stages based on the low, medium and high complexity intervals.
[0090] The actual value distribution of MPPT duty cycle adjustment of the deep reinforcement learning strategy network output in each course stage is statistically analyzed. If the actual value of MPPT duty cycle adjustment exceeds the MPPT duty cycle adjustment range of the corresponding course stage, the MPPT duty cycle adjustment range of the corresponding course stage is expanded or contracted to the range of the actual value distribution of MPPT duty cycle adjustment.
[0091] The execution error of the actual adjustment step size of the battery management unit in constant current charging mode is statistically analyzed, and the mean value of the execution error is calculated. If the mean value of the execution error exceeds the allowable value of the battery management protocol, the mean value of the execution error is added to the charging current adjustment step size.
[0092] It should also be noted that by continuously optimizing the generative adversarial network and curriculum strategy through a closed-loop feedback mechanism (virtual data → real hardware operation → parameter update), the problem of strategy failure caused by environmental dynamism and hardware aging in existing MPPT methods is solved. By updating GAN parameters and re-dividing the curriculum stages based on operational data, the virtual dataset and action space parameters can adapt to new environments (such as changes in photovoltaic module characteristics or deep battery degradation), thereby achieving adaptive iterative capability and maintaining a balance between control accuracy and battery life in long-term use.
[0093] An adaptive control method for solar charging based on an improved MPPT approach is based on a fitted MPPT architecture. The core principle of fitted MPPT includes: adjusting the duty cycle and topology parameters of the DC-DC converter in real time to solve the power loss caused by the mismatch between the output voltage of the photovoltaic panel and the battery voltage, thus ensuring efficient energy transmission.
[0094] Based on the IV curve of the photovoltaic array and the charging characteristics of the battery, the optimal voltage matching range is dynamically calculated so that the operating point is always in the high power-low loss region.
[0095] When there are drastic fluctuations in illumination, fuzzy control and neural network algorithms are used to smooth the trajectory of power point adjustment to avoid oscillations.
[0096] When the illumination is stable, the P&O / IncCond algorithm is used; when there is a sudden change, the transient optimization mode is switched to balance response speed and accuracy.
[0097] Dynamically manages photovoltaic charging and load power supply, supports MPPT+ boost / buck hybrid mode, and is compatible with off-grid hybrid systems;
[0098] Automatically identifies battery type (lead-acid / lithium battery, etc.) and adjusts the constant current-constant voltage-float charge curve to avoid overcharging / undercharging.
[0099] This embodiment also provides a computer device applicable to the solar charging adaptive control method based on the improved MPPT method, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the solar charging adaptive control method based on the improved MPPT method proposed in the above embodiment.
[0100] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0101] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the solar charging adaptive control method based on the improved MPPT method proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0102] In summary, this invention overcomes the limitation of simulation models' linear assumptions about single environmental variables by generating virtual data through generative adversarial networks. This allows the policy network to quickly optimize the duty cycle in complex environments during the pre-training stage, avoiding oscillations or misjudgments caused by traditional fixed step sizes and improving the tracking accuracy of dynamic environments. Furthermore, by progressively training the deep reinforcement learning policy network in stages, real-time adaptation of the policy to hardware aging and environmental complexity is achieved, solving the problem of the lack of adaptability in existing control policies.
[0103] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A solar charging adaptive control method based on an improved MPPT method, characterized in that: include, Generative adversarial networks are used to simulate dynamic environment scenarios and generate virtual datasets. The virtual dataset is divided into three course phases, and the environment complexity index and action space of each course phase are defined. The deep reinforcement learning policy network is trained in the order of the course phases. The action space includes the MPPT duty cycle adjustment range and the charging current adjustment step size. Collect real-world environmental data, calculate environmental complexity indicators, match corresponding course stages through a deep reinforcement learning strategy network, and obtain MPPT duty cycle adjustment amount and charging current target value based on the action space associated with the course stage. Based on the MPPT duty cycle adjustment and the target charging current, the switching frequency of the DC-DC converter and the charging threshold of the battery management unit in the solar charging hardware are controlled, and the operating data of the solar charging hardware is collected. Based on the operational data of the solar charging hardware, the parameters of the generative adversarial network are updated, and the division of course phases and action space parameters are optimized.
2. The solar charging adaptive control method based on the improved MPPT method as described in claim 1, characterized in that: The generation of virtual datasets refers to the generation of virtual datasets based on a parameter space defined by a dynamic environment scenario and combined with a generator from a generative adversarial network.
3. The solar charging adaptive control method based on the improved MPPT method as described in claim 1, characterized in that: The division of the virtual dataset into three course stages refers to calculating the comprehensive score of the virtual dataset based on the virtual dataset. The overall score of the virtual dataset is divided into low, medium and high complexity ranges, and the virtual dataset is divided into low, medium and high course stages.
4. The solar charging adaptive control method based on the improved MPPT method as described in claim 3, characterized in that: The definition of the environmental complexity index and action space for each course stage refers to setting environmental complexity indices for low, medium, and high course stages based on a virtual dataset and low, medium, and high complexity ranges. The action space is set for each course stage based on the environmental complexity index.
5. The solar charging adaptive control method based on the improved MPPT method as described in claim 4, characterized in that: The training of the deep reinforcement learning policy network according to the order of course stages refers to using virtual datasets and action spaces corresponding to low, medium and high course stages, and training the deep reinforcement learning policy network in three stages in the order of low, medium and high.
6. The solar charging adaptive control method based on the improved MPPT method as described in claim 1, characterized in that: The process of matching the corresponding course stage through a deep reinforcement learning strategy network refers to comparing the calculated environment complexity index with the environment complexity indexes of low, medium, and high course stages to match the corresponding course stage.
7. The solar charging adaptive control method based on the improved MPPT method as described in claim 1, characterized in that: The specific steps for updating the parameters of the generative adversarial network are as follows: The virtual dataset and the operating data of the solar charging hardware are input into the discriminator of the generative adversarial network, and the parameters of the discriminator of the generative adversarial network are updated by calculating the cross loss during the training process. A virtual dataset is generated by the generator of a generative adversarial network (GAN). The adversarial loss between the generator and the discriminator of the GAN is calculated using cross-entropy loss. The parameters of the generator network are updated using the backpropagation algorithm.
8. The solar charging adaptive control method based on the improved MPPT method as described in claim 7, characterized in that: The specific steps for optimizing the course phase division and action space parameters are as follows: A new virtual dataset is generated using the updated generative adversarial network generator, and a comprehensive score for the new virtual dataset is calculated. Based on the comprehensive scores of the new virtual dataset, the courses are divided into low, medium, and high complexity ranges, and the low, medium, and high course stages are redefined. The actual distribution of MPPT duty cycle adjustment values of the deep reinforcement learning strategy network output in each course stage was statistically analyzed, and the MPPT duty cycle adjustment range for the corresponding course stage was adjusted. The execution error of the actual adjustment step size of the battery management unit in constant current charging mode is statistically analyzed, the mean of the execution error is calculated, and the charging current adjustment step size is adjusted.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the solar charging adaptive control method based on the improved MPPT method as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the solar charging adaptive control method based on the improved MPPT method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Flexible control method and device for photovoltaic power generation energy storage system based on reinforcement learning
CN116345506A
Battery charging protection method and system based on intelligent control and storage medium
CN120016657A