Parameter self-calibration method and system of photovoltaic power generation power prediction model
Patent Information
- Application Number
- CN202611143968.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-30
- Publication Date
- 2026-10-09
AI Technical Summary
[0007]本发明旨在提供一种能够在仅2~3个月历史数据条件下,自动、精准、鲁棒地校准光伏物理模型参数的方法及系统,解决现有技术中少数据场景下模型无法建立、数据清洗不彻底、高功率段预测偏差大、优化参数易偏离物理边界等问题
1.少数据场景下的快速建模能力:本发明仅需2~3个月历史数据(约800~1300条白天样本),即可完成物理模型参数的自校准,特别适用于新建电站和分布式光伏项目。而传统数据驱动方法通常需要1年以上数据。
Smart Images

Figure CN122886408A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of photovoltaic power generation prediction technology, specifically relating to a parameter self-calibration method and system for a photovoltaic power generation prediction model, which is particularly suitable for rapid modeling and high-precision prediction in scenarios such as newly built power plants or where historical data is scarce. Background Technology
[0002] Accurate prediction of photovoltaic power generation is crucial for grid dispatch, renewable energy integration, and power plant operation and maintenance. Existing photovoltaic power prediction methods are mainly divided into two categories: pure data-driven models and traditional physical models, each with obvious technical shortcomings.
[0003] Pure data-driven models (such as LSTM, XGBoost, and LightGBM) rely on large amounts of high-quality labeled data, typically requiring more than a year of historical operational data to train an effective predictive model. These models have poor generalization ability, their prediction accuracy drops sharply after sudden weather changes or aging power plant equipment, and they lack physical interpretability. For newly built power plants, due to insufficient historical data accumulation, such methods cannot be effectively applied. Experiments show that with only 2-3 months of training data, the WMAPE of the LightGBM hybrid model on the validation set is 20.40%, which is lower than the 19.39% of the pure physics model.
[0004] Traditional physical models are based on component-based generator mechanisms ( Parameters (photovoltaic conversion efficiency α, temperature coefficient γ, etc.) are usually directly taken from the factory-specified values. However, in actual operation, factors such as component degradation, dust obstruction, and differences in installation angle cause the actual parameters to deviate significantly from the nominal values. Manually adjusting parameters is time-consuming and labor-intensive, and it is difficult to obtain the globally optimal solution.
[0005] The existing technology has the following technical defects: (1) The nominal value is used directly. The nominal value at the factory is significantly different from the actual operating conditions. On the actual power plant data of 1.3MW, the physical model using the nominal parameters can have an error of 8% to 15% on a stable sunny day, and the error is even greater on a cloudy or rainy day; (2) Simple statistical calibration only adjusts the overall loss coefficient and ignores the differences under different irradiance intervals, resulting in a large prediction deviation during high irradiance periods; (3) Equal weight error optimization treats the error of all power intervals equally, resulting in the optimization result may fit well in the low power interval, but the deviation is huge in the high power interval; (4) The data cleaning is insufficient. Conventional cleaning only performs simple threshold truncation and does not make sufficient use of the physical mechanism. In the case of few data, the incorrect removal of effective samples will seriously affect the model training effect; (5) The existing scheme ignores the special characteristics of the few data scenarios. Most existing technologies require at least one year of data to obtain reliable parameter calibration results. New power plants and distributed photovoltaic projects only have 2 to 3 months of data in the early stage of operation, and the existing methods cannot effectively cope with this.
[0006] Existing technologies mostly use MAE or RMSE as evaluation indicators, but these indicators are greatly affected by the dimensions and are difficult to compare horizontally between power plants of different capacities. Summary of the Invention
[0007] The present invention aims to provide a method and system for automatically, accurately and robustly calibrating photovoltaic physical model parameters under conditions of only 2 to 3 months of historical data, solving problems in the prior art such as the inability to build models in scenarios with limited data, incomplete data cleaning, large prediction deviations in high-power segments, and easy deviation of optimized parameters from physical boundaries.
[0008] The technical solution adopted by this invention to solve its technical problem is as follows: On the one hand, a parameter self-calibration method for a photovoltaic power generation prediction model is provided, including the following steps: Step S1: Obtain historical operating data of the photovoltaic power station and corresponding meteorological data, and perform multi-level physical constraint cleaning on the data to obtain cleaned data; Step S2: Construct a refined physical model for photovoltaic power generation prediction, wherein the physical model includes the photoelectric conversion efficiency parameter α and the temperature coefficient parameter γ to be calibrated; Step S3: Use the differential evolution algorithm to globally optimize the photoelectric conversion efficiency parameter α and the temperature coefficient parameter γ in the physical model, with the goal of minimizing the weighted average absolute percentage error, to obtain the optimal parameter combination. ; Step S4: For the optimal parameter combination Perform parameter boundary diagnosis to determine whether the optimal parameter combination touches a preset physical boundary, and output data quality feedback information based on the diagnosis results.
[0009] Furthermore, the parameter self-calibration method also includes step S5: multi-dimensional verification and visualization diagnosis; the multi-dimensional verification includes calculating the overall WMAPE of the verification set, daily detailed statistics of WMAPE, and weather stratification assessment based on the irradiance fluctuation coefficient; in the weather stratification assessment, the fluctuation coefficient is defined as the ratio of the standard deviation to the mean of irradiance during high irradiance periods (irradiance > 200 W / m²): , in, For time periods with irradiance > 200 W / m², μ and σ are the mean and standard deviation of irradiance for that time period, respectively. When the fluctuation coefficient < 0.35, it is classified as stable sunny; when 0.35 ≤ fluctuation coefficient ≤ 0.50, it is classified as cloudy and fluctuating; when the fluctuation coefficient > 0.50 or there are fewer than 10 effective high irradiance data, it is classified as cloudy / rainy / strong fluctuating. The average WMAPE under each weather type is calculated to clarify the applicable scenario boundaries of the physical model.
[0010] On the other hand, a parameter self-calibration system for a photovoltaic power generation prediction model is provided, including: The data cleaning module is used to acquire historical operating data of photovoltaic power plants and corresponding meteorological data, and to perform multi-level physical constraint cleaning on the data to obtain cleaned data. The physical model building module is used to build a refined physical model for photovoltaic power generation prediction. The physical model includes the photoelectric conversion efficiency parameter α and the temperature coefficient parameter γ to be calibrated. The differential evolution optimization module is used to globally optimize the photoelectric conversion efficiency parameter α and the temperature coefficient parameter γ in the physical model using the differential evolution algorithm, with the goal of minimizing the weighted average absolute percentage error, to obtain the optimal parameter combination. ; The boundary diagnosis module is used to analyze the optimal parameter combination. Perform parameter boundary diagnosis to determine whether the optimal parameter combination touches a preset physical boundary, and output data quality feedback information based on the diagnosis results.
[0011] Furthermore, the parameter self-calibration system also includes a verification and diagnostic module for multi-dimensional verification and visual diagnostics. The multi-dimensional verification includes calculating the overall WMAPE of the verification set, daily detailed WMAPE statistics, and weather stratification assessment based on the irradiance fluctuation coefficient. In the weather stratification assessment, the fluctuation coefficient is defined as the ratio of the standard deviation to the mean of irradiance during high-irradiance periods (irradiance > 200 W / m²). , in, For time periods with irradiance > 200 W / m², μ and σ are the mean and standard deviation of irradiance for that time period, respectively. When the fluctuation coefficient < 0.35, it is classified as stable sunny; when 0.35 ≤ fluctuation coefficient ≤ 0.50, it is classified as cloudy and fluctuating; when the fluctuation coefficient > 0.50 or there are fewer than 10 effective high irradiance data, it is classified as cloudy / rainy / strong fluctuating. The average WMAPE under each weather type is calculated to clarify the applicable scenario boundaries of the physical model.
[0012] One of the above technical solutions has the following advantages or beneficial effects: 1. Rapid modeling capability in scenarios with limited data: This invention requires only 2-3 months of historical data (approximately 800-1300 daytime samples) to complete the self-calibration of physical model parameters, making it particularly suitable for newly built power plants and distributed photovoltaic projects. Traditional data-driven methods typically require more than one year of data.
[0013] 2. Physical rationality and interpretability of parameters: The parameters optimized by differential evolution (α=0.9474, γ= The temperature coefficient of 0.0044 fully conforms to the physical characteristics of crystalline silicon photovoltaics, and falls within the actual range. 0.003~ The value of 0.005 ensures the physical interpretability and generalization ability of the model.
[0014] 3. Scientific validity of WMAPE evaluation: WMAPE naturally assigns greater weight to high-power periods, which aligns perfectly with the grid dispatching requirements for accuracy during high-power periods. Under stable, clear weather conditions, the validation set WMAPE can reach 5.47%–7.82%.
[0015] 4. Weather stratification diagnostic capability: The system automatically evaluates the data according to the irradiance fluctuation coefficient, clearly presenting the performance boundaries of the physical model under different weather conditions, and providing quantitative basis for users to understand the applicable scenarios of the model.
[0016] 5. Parameter boundary self-diagnosis: When the optimization parameters touch the physical boundary, the system will automatically issue an alarm and prompt the user to check the data quality or expand the training set to avoid the model producing unreliable results when the data is abnormal.
[0017] 6. Model self-calibration and continuous evolution: Supports periodic (e.g., monthly) automatic parameter calibration, updates parameters using new running data, and automatically tracks component degradation and seasonal changes without manual intervention.
[0018] 7. Synergistic Effect of Multiple Techniques: This invention organically integrates physical constraint data cleaning, refined physical modeling, differential evolution global optimization, WMAPE asymmetric evaluation, and parameter boundary diagnosis, with each technique working synergistically. Physical constraint data cleaning provides high-quality training samples for differential evolution optimization, avoiding noise interference with parameter identification in scenarios with limited data. The refined physical model ensures parameter identifiability in its simplest form, providing a structured optimization space for the differential evolution algorithm. The WMAPE evaluation index guides the differential evolution algorithm to prioritize the optimization of prediction accuracy in the high-power range, precisely aligning with the actual needs of power grid dispatch. Parameter boundary diagnosis utilizes prior physical knowledge to verify the rationality of the optimization results, forming a closed-loop feedback. This complete technical chain of "cleaning—modeling—optimization—evaluation—diagnosis" enables the acquisition of physically reasonable parameter calibration results with prediction accuracy superior to data-driven models, even under extreme conditions with only two months of training data. To verify the synergistic effect, the inventors conducted a comparative experiment: when using only differential evolution optimization (without physical constraint cleaning), the optimization parameters are prone to getting trapped in local optima, with a WMAPE of 22.3% on the validation set; when using only the physical model (without differential evolution), the WMAPE is 24.8%; while the complete solution of this invention achieves 19.39% on the same dataset, which is significantly better than the simple superposition of individual techniques. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating a parameter self-calibration method for a photovoltaic power generation prediction model according to an exemplary embodiment; Figure 2 This is a flowchart illustrating a parameter self-calibration method for another photovoltaic power generation prediction model according to an exemplary embodiment; Figure 3 This is a daily WMAPE distribution histogram of a validation set according to an exemplary embodiment; Figure 4 This is a comparison chart of actual and predicted power curves for a stable sunny day, according to an exemplary embodiment. Figure 5 This is a schematic diagram of the parameter self-calibration system structure of a photovoltaic power generation prediction model according to an exemplary embodiment; Figure 6 This is a schematic diagram of the parameter self-calibration system structure of another photovoltaic power generation prediction model according to an exemplary embodiment; Figure 7 This is a flowchart illustrating a specific implementation of a parameter self-calibration method for a photovoltaic power generation prediction model, according to an exemplary embodiment. Detailed Implementation
[0020] To more clearly illustrate the technical features of the present invention, the invention will be described in detail below through specific embodiments and in conjunction with the accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and arrangements of specific examples are described below. Of course, these are merely examples and are not intended to limit the invention.
[0021] The core idea of this invention is: physical constraint data cleaning → refined physical model → differential evolution global optimization → WMAPE asymmetric evaluation → multi-dimensional verification visualization.
[0022] Example 1 like Figure 1 As shown in the figure, the parameter self-calibration method for a photovoltaic power generation prediction model provided by the present invention includes the following steps: Step S1: Obtain historical operating data of the photovoltaic power station and corresponding meteorological data, and perform multi-level physical constraint cleaning on the data to obtain cleaned data.
[0023] The historical operating data includes the DC / AC power and timestamps of the photovoltaic power station, and the meteorological data includes the total irradiance of the tilted surface, ambient temperature, wind speed, humidity, and cloud cover. After the data is acquired, the photovoltaic power data is first resampled from the original sampling interval to a time resolution consistent with the meteorological data. The power values at non-hourly times are interpolated to the hourly times using a linear interpolation method. Then, the power data is concatenated and merged with the meteorological data within the timestamp to form a unified time series dataset.
[0024] The physical constraint multi-level cleaning includes at least basic hard constraint cleaning: setting all negative power values to zero and marking power values exceeding 1.5 times the AC installed capacity as abnormal and removing them; and dynamic anomaly detection based on theoretical maximum power, which calculates the theoretical maximum power using the component nominal parameters and executes zonal cleaning rules according to the irradiance range.
[0025] In the dynamic anomaly detection based on theoretical maximum power, the theoretical maximum power The calculation formula is: , Where G is the total irradiance of the inclined surface. The nominal conversion efficiency of the component. The nominal temperature coefficient, The ambient temperature is used to calculate the theoretical power limit to detect anomalies, in conjunction with the battery temperature in step S2. different; The zonal cleaning rule is as follows: when the irradiance G < 10W / m², the power is forcibly set to zero; when G ≥ 10W / m² and the actual power > 10W / m², the power is forcibly set to zero. When the value is ×1.5, it is marked as abnormal and removed; when G≥10W / m² and actual power< When the value is ×0.01, it is marked as suspicious data.
[0026] The physical constraint multi-level cleaning also includes layered filling of missing values: for isolated missing points, the mean of a rolling window with a window size of 3 is used for filling; for short-term continuous missing points with a duration of ≤3 hours, linear interpolation is used to interpolate along the time axis; for long-term continuous missing points with a duration of >3 hours or missing points during the nighttime period, they are directly filled with 0.
[0027] In step S1, if the photovoltaic power station has other power sources such as energy storage and diesel engines sharing the same bus with the photovoltaic power station, then the other power source components are stripped: after aligning the other power source data by time, the power of the other power sources is subtracted from the total active power to obtain the pure photovoltaic power; if the total active power is zero, then the pure photovoltaic power is directly set to zero.
[0028] In the specific implementation process, step S1 is implemented as follows.
[0029] 1.1 Data Acquisition: Obtain 2-3 months of historical operating data and corresponding meteorological data from the photovoltaic power station: Photovoltaic operation data: The raw data comes from the power plant monitoring system, with a sampling interval of 5 minutes. Recorded fields include timestamp, total active power, etc. Meteorological data: obtained through the meteorological service interface, with the original time resolution being hourly values. Fields include temperature, humidity, irradiance, wind speed, cloud cover, etc.
[0030] 1.2 Data Alignment and Resampling Strategies: To achieve time alignment between power and meteorological characteristics, resampling operations were performed according to Table 1.
[0031] Table 1 Resampling Operation and Instructions
[0032] 1.3 Basic hard constraint cleaning (physical constraints): Perform basic hard constraint cleaning as shown in Table 2.
[0033] Table 2 Cleaning Rules and Implementation Methods
[0034] The installed capacity is set at 1100kW (AC side), and the threshold is set at 1.5 times as a protection band to prevent false data from inverter overload.
[0035] 1.4 Dynamic anomaly detection based on theoretical maximum power: (1) Theoretical maximum power calculation model :
[0036] Where G is the total irradiance of the tilted surface (global_tilted_irradiance_instant), in W / ; This refers to the component's nominal conversion efficiency. , where is the temperature power coefficient ( / °C); The ambient temperature.
[0037] (2) Irradiance-based regional cleaning rules: The irradiance-based regional cleaning rules adopted are shown in Table 3.
[0038] Table 3 Area Cleaning Rules
[0039] 1.5 Hierarchical imputation strategy for missing values: A layered filling scheme was adopted, and the missing patterns were differentiated and handled according to the table 4.
[0040] Table 4 Layered Filling Strategy
[0041] 1.6 Stripping of other power components (if present): If other power sources (such as energy storage or diesel engines) share the same bus with the photovoltaic system on site, power decoupling is required. 1. Read data from "Other Power Sources" (sampling interval 5 minutes), and take the average or hourly value on the hour; 2. Aligned with the total photovoltaic power in time; 3. Calculate pure photovoltaic power: .
[0042] Step S2: Construct a refined physical model for predicting photovoltaic power generation, wherein the physical model includes the photoelectric conversion efficiency parameter α and the temperature coefficient parameter γ to be calibrated.
[0043] This invention employs the simplest form of physical model, abandoning complex but difficult-to-calibrate multi-parameter models, to ensure that parameters are identifiable under limited data conditions.
[0044] 2.1 Model Formula: , in: The total irradiance of the inclined surface is (W / m²). To comprehensively measure photoelectric conversion efficiency (including module efficiency, inverter efficiency, line loss, etc.); Temperature coefficient (% / °C); The operating temperature of the photovoltaic panel is (°C).
[0045] 2.2 Photovoltaic panel temperature calculation (NOCT model): , NOCT (Nominal Operating Cell Temperature) is a factory parameter for the module, typically set at 45°C for crystalline silicon modules. The ambient temperature.
[0046] This model is better suited for scenarios with limited data than the complex Faiman thermal model because it only relies on two variables: irradiance and temperature, without the need for additional parameters such as wind speed, thus avoiding the risk of overfitting caused by too many parameters.
[0047] 2.3 Model Physical Constraints: Forced application during optimization: ∈[0.5,1.0]: Efficiency cannot be lower than 50% (severe aging / failure) nor higher than 100% (violation of the second law of thermodynamics); ∈[-0.02,0.0]: The temperature coefficient must be negative (semiconductor characteristics) and the absolute value cannot be too large.
[0048] Parameters to be calibrated in the physical model α The preset physical boundary for γ is: ∈[0.5,1.0], The parameters are set to [-0.02, 0.0] to ensure that the parameters always conform to the physical characteristics of the photovoltaic module during the optimization process, thus avoiding the generation of physically feasible solutions.
[0049] Step S3: Use the differential evolution algorithm to globally optimize the photoelectric conversion efficiency parameter α and the temperature coefficient parameter γ in the physical model, with the goal of minimizing the weighted average absolute percentage error, to obtain the optimal parameter combination. .
[0050] The objective function for the weighted average absolute percentage error (WMAPE) is defined as follows: , Where WMAPE is the weighted average absolute percentage error, P actual For actual power, P pred The power predicted by the model is summed within the range of valid daytime samples in the training set with irradiance greater than 10 W / m² and actual power greater than 0.
[0051] The differential evolution algorithm employs a best / 1 mutation strategy and a binomial crossover strategy, with a population size of 20, a maximum number of iterations of 1500, a convergence tolerance of 0.005, and a fixed random seed to ensure reproducibility of results.
[0052] In the best / 1 mutation strategy, for each target vector X in the g-th generation population i (g) Generate mutation vector V i (g) The formula is: , in, The optimal individual with the minimum objective function value in the g-th generation population, where F is the scaling factor and in Adaptive random selection within the interval and For each unique individual, and not equal to i, use a random index. , Population size.
[0053] In the binomial crossover strategy, for each dimension parameter j, the trial vector is determined by the crossover probability Cr. The corresponding dimension comes from the mutation vector. Or the target vector And at least one dimension is guaranteed to come from the mutation vector, specifically: , in, For each dimension, j is a random number generated independently. rand For a randomly selected dimension index, Cr adaptively takes a random value in the range [0.5, 0.9].
[0054] The differential evolution algorithm employs a one-to-one greedy selection strategy: comparing trial vectors. and target vector ) Fitness (WMAPE value): , The experimental vector is only allowed to enter the next generation of the population if its fitness is not greater than that of the target vector; otherwise, the target vector is retained to ensure that the optimal fitness of the population does not decrease monotonically.
[0055] The iteration termination condition of the differential evolution algorithm is: reaching the maximum number of iterations of 1500 generations, or the absolute change of the WMAPE value of the best individual in the population within consecutive generations is less than the convergence tolerance of 0.005%. If either condition is met, the optimization will terminate and the current global optimal parameter combination will be output.
[0056] In the specific implementation process, step S3 is implemented as follows.
[0057] 3.1 Definition of the optimization problem: Finding the optimal parameters This minimizes WMAPE on the training set: .
[0058] 3.2 Differential Evolution Algorithm Configuration: This invention employs a differential evolution algorithm to globally optimize the physical model parameters. Its advantages include not requiring gradient information of the objective function, strong global search capability, and natural handling of parameter boundary constraints. Specific configurations are shown in Table 5.
[0059] Table 5. Mathematical definitions and implementation strategies for each key step.
[0060] (1) Definition of objective function (fitness function): This invention uses the weighted average absolute percentage error (WMAPE) on the training set as the optimization objective function to measure the fitting accuracy of the physical model. The decision variable is the component conversion efficiency. and temperature coefficient .
[0061] Given a set of candidate parameters The power predicted by the physical model is: , Where G(t) is the total irradiance of the inclined surface. The temperature of the photovoltaic panel is estimated using the NOCT model.
[0062] The objective function value (i.e., the fitness that needs to be minimized) is defined as: , Where N is the total number of valid daytime samples in the training set with irradiance greater than 10 W / m² and actual power greater than 0. The smaller the objective function value, the stronger the physical model's ability to interpret historical data.
[0063] (2) Population initialization method: When the differential evolution algorithm is started, an initial population containing NP (population size, taken as 20 here) individuals (candidate solutions) needs to be generated. .
[0064] This invention employs a uniform random initialization strategy to ensure that the initial solution is uniformly distributed throughout the feasible region, thereby enhancing the diversity of the initial population. For the j-th dimension parameter (j=1,2), its initialization formula is: , Where rand(0,1) is a random number that follows a uniform distribution in the range [0,1]; lb=[0.5, [0.02] and ub=[1.0,0.0] are the lower and upper bounds of the two decision variables, respectively, strictly satisfying the physical constraints of the photovoltaic module. ∈[0.5,1.0], γ∈[ 0.02,0].
[0065] (3) Mutation strategy (DE / best / 1): This invention employs a best / 1 mutation strategy, utilizing the best individual in the current population to guide the search direction and accelerate the convergence speed.
[0066] For each target vector in the g-th generation Generate the corresponding mutation vector : , in: The objective function value in the g-th generation population The smallest optimal individual; and It is a randomly selected, distinct index from the population that is not equal to i; F is the scaling factor, used to control the amplitude of the differential perturbation. This algorithm adopts an adaptive jitter mechanism in its implementation. (Random values are selected within the range) to enhance search flexibility.
[0067] This strategy combines the information of the optimal individual with the direction of random difference, thus ensuring both local mining capability (tending towards the optimal) and global exploration capability (random perturbation).
[0068] (4) Crossover strategy (binomial crossover): After mutation is complete, the target vector will be... Its corresponding mutation vector Perform crossover to generate test vectors .
[0069] This invention employs a binomial crossover strategy, or "bin". For each dimension parameter j∈{1,2}, the source of the experimental vector value is determined by the crossover probability Cr (in this invention, an adaptive value within the range of [0.5,0.9] is used): , in For each dimension, a random number is generated independently. The index is a randomly selected dimension (ensuring that at least one dimension comes from the mutation vector to prevent the population from stagnation).
[0070] (5) Selection strategy (greedy selection): It employs a one-to-one greedy selection mechanism, comparing trial vectors... and target vector The fitness value determines which individual enters the next generation of the population: , That is, only individuals with better or equal fitness (WMAPE) will be retained to the next generation, ensuring that the overall quality of the population does not decrease monotonically with the number of iterations.
[0071] (6) Iteration termination condition: During the algorithm iteration, the optimization will terminate and the current global optimal solution will be output when any of the following conditions are met: The number of iterations reaches the preset maximum value (1500 generations). If the absolute change in the objective function value of the optimal individual in the population is less than the convergence tolerance (0.005%) over consecutive generations, it is considered to have reached a stable convergence state.
[0072] 3.3 Robustness of optimization in scenarios with limited data: Even with only 2-3 months of data (approximately 800 valid daytime samples), the differential evolution algorithm still converges stably. Experiments show that: After optimization =0.9474, =-0.0044; It falls entirely within the true temperature coefficient range of crystalline silicon photovoltaics (-0.003 to -0.005). This shows that even with limited data, the algorithm can still identify physically reasonable parameters.
[0073] Step S4: For the optimal parameter combination Perform parameter boundary diagnosis to determine whether the optimal parameter combination touches a preset physical boundary, and output data quality feedback information based on the diagnosis results.
[0074] The parameter boundary detection includes: calculating the distance between each parameter in the optimal parameter combination and the preset physical boundary. , in and These are the preset lower and upper bounds for this parameter, respectively. The boundary is determined when the photoelectric conversion efficiency α touches the upper limit of 1.0 or the temperature coefficient γ touches the lower limit. If α is 0.02, the data is considered to have a systematic bias; if α touches the lower limit of 0.5, the component is considered to be severely aged or the data is abnormal; if no parameter touches the boundary, the data quality is considered to be good and the optimization results are reliable.
[0075] When a parameter is detected to be touching the boundary, the system automatically outputs an alarm message, prompting the user to check for abnormal dates in the original data and suggesting that the time range of the training data be expanded or the quality of the irradiance data be reviewed again; when no parameter touches the boundary, a data quality qualified signal is output and the optimal parameters are saved to the configuration file for the prediction system to call.
[0076] In the specific implementation process, step S4 is implemented as follows.
[0077] 4.1 Boundary Detection: After optimization, calculate the distance between each parameter and the physical boundary: , like It was determined to be "boundary reached".
[0078] The diagnostic threshold is determined based on the following criteria: The threshold for determining the boundary condition is set as follows: This value takes into account the numerical precision of double-precision floating-point operations (approximately...). Differential Evolutionary Algorithm Convergence Tolerance ( The relative scale between the physical boundary of the parameter (α interval width 0.5, The interval width is 0.02. This threshold is more than two orders of magnitude higher than the numerical rounding error, ensuring that misjudgments are not caused by computational inaccuracies; at the same time, it is less than 0.5% of the effective interval width of the parameter, ensuring a sensitive response to actual physical anomalies. In addition, this threshold is consistent with the boundary out-of-bounds truncation tolerance commonly used in differential evolution algorithms, which can accurately identify the trend of the algorithm breaking through physical constraints due to systematic biases in the data.
[0079] 4.2 Diagnostic Logic: like Touch limit (1.0) or Touching the lower limit (-0.02): Highly suspicious, indicating that there may be systematic bias in the data, and it is recommended to check the quality of the irradiance data; like Touching the lower limit (0.5): The component may be severely aged or there may be serious data problems; If no parameters touch the boundary: the data quality is good and the optimization results are reliable.
[0080] 4.3 Feedback and Alarms: When a parameter is detected to be touching a boundary, the system automatically outputs an alarm message, prompting the user to check the original data for the abnormal date and suggesting that the time range of the training data be expanded.
[0081] Example 2 Unlike Example 1, this embodiment of the invention provides a parameter self-calibration method for a photovoltaic power generation prediction model, such as... Figure 2 As shown, it also includes step S5: multi-dimensional verification and visual diagnosis.
[0082] The multi-dimensional validation includes calculating the overall WMAPE of the validation set, daily detailed statistics of WMAPE, and weather stratification assessment based on the irradiance fluctuation coefficient; in the weather stratification assessment, the fluctuation coefficient is defined as the ratio of the standard deviation to the mean of irradiance during high irradiance periods (irradiance > 200 W / m²). , in, For time periods with irradiance > 200 W / m², μ and σ are the mean and standard deviation of irradiance for that time period, respectively. When the fluctuation coefficient < 0.35, it is classified as stable sunny; when 0.35 ≤ fluctuation coefficient ≤ 0.50, it is classified as cloudy and fluctuating; when the fluctuation coefficient > 0.50 or there are fewer than 10 effective high irradiance data, it is classified as cloudy / rainy / strong fluctuating. The average WMAPE under each weather type is calculated to clarify the applicable scenario boundaries of the physical model.
[0083] In the specific implementation process, step S5 is implemented as follows.
[0084] 5.1 Overall WMAPE Assessment: The overall WMAPE of the validation set is calculated and used as the core indicator of model accuracy.
[0085] 5.2 Daily WMAPE Detailed Statistics: The system automatically calculates the WMAPE for each day of the validation set and outputs daily details to facilitate the identification of abnormal dates.
[0086] 5.3 Weather Stratification Assessment: To quantify the impact of weather conditions on the prediction accuracy of physical models, this invention proposes a weather stratification assessment method based on the irradiance fluctuation coefficient during high irradiance periods. This method divides the weather into three levels by calculating the irradiance fluctuation coefficient during daily high irradiance periods (tilted surface irradiance > 200 W / m²) in the validation set, and statistically analyzes the prediction error distribution of the physical model at each level, thereby clearly defining the applicable scenarios for the model.
[0087] (1) Definition of volatility coefficient: , in For periods with irradiance >200W / m², μ and σ are the mean and standard deviation of irradiance for that period, respectively. If there are fewer than 10 valid high-irradiance data points on a given day, the data is directly categorized as "cloudy / rainy / strong fluctuation".
[0088] (2) Weather stratification: The weather stratification criteria are shown in Table 6.
[0089] Table 6 Weather Stratification Standards
[0090] (3) Calculation example (using real data from June 3, 2026 as an example): The irradiance data for the high irradiance period of the day (09:00~17:00) are: 499.7, 716.4, 749.3, 821.1, 833.1, 693.0, 556.4, 522.3, 349.6 (W / m²).
[0091] Mean μ = 637.9 W / m²; Standard deviation σ = 157.3 W / m²; Volatility coefficient = 157.3 / 637.9 = 0.25 (<0.35 → stable sunny weather); WMAPE=25.32% on that day.
[0092] (4) Daily assessment results: Based on the validation set real data from June 1, 2026 to June 21, 2026, the daily classification and corresponding WMAPE are shown in Table 7, and the daily WMAPE distribution is as follows. Figure 3 As shown.
[0093] Table 7 WMAPE Daily Assessment Table:
[0094] (5) Stratified statistics: The stratified statistical results are shown in Table 8.
[0095] Table 8 Stratified Statistical Table
[0096] Stratified statistical conclusions: The physical model exhibits stable prediction accuracy under clear skies. During cloudy and fluctuating conditions, the randomness of cloud cover leads to slightly higher errors on some dates, but the overall accuracy remains manageable. Under overcast / rainy / low-light conditions, the model's accuracy is severely limited. The fluctuation coefficient is positively correlated with the prediction error and can be used as an indicator of the physical model's applicability. When the daily fluctuation coefficient exceeds 0.50, it is recommended to combine it with a data-driven model for error correction.
[0097] 5.4 Daily Power Curve Comparison Chart: Generate a separate image for the validation set each day, such as Figure 4 As shown, the curves showing the fit between actual power and predicted power throughout the day are presented, facilitating manual diagnosis.
[0098] 5.5 Validation instructions for scenarios with limited data: Experiments show that, with only two months of training data (772 daytime samples) from April to May, the physical model performs well on the validation set (21 days in June, 288 daytime samples): On stable sunny days (such as June 11, 13, and 14), WMAPE reached 5.47% to 7.82%; Partly cloudy with an average WMAPE of approximately 8% to 12%; The overall WMAPE of the validation set is 19.39%.
[0099] Stratified statistical results show that the pure physical model has extremely high accuracy under stable irradiance conditions, with the error mainly originating from unpredictable cloud fluctuations. This conclusion can be reached with only two months of training data.
[0100] Example 3 like Figure 5 As shown, to implement the parameter self-calibration method of the photovoltaic power generation prediction model described in Example 1, this embodiment of the invention provides a parameter self-calibration system for a photovoltaic power generation prediction model, comprising: The data cleaning module is used to acquire historical operating data of photovoltaic power plants and corresponding meteorological data, and to perform multi-level physical constraint cleaning on the data to obtain cleaned data. The physical model building module is used to build a refined physical model for photovoltaic power generation prediction. The physical model includes the photoelectric conversion efficiency parameter α and the temperature coefficient parameter γ to be calibrated. The differential evolution optimization module is used to globally optimize the photoelectric conversion efficiency parameter α and the temperature coefficient parameter γ in the physical model using the differential evolution algorithm, with the goal of minimizing the weighted average absolute percentage error, to obtain the optimal parameter combination. ; The boundary diagnosis module is used to analyze the optimal parameter combination. Perform parameter boundary diagnosis to determine whether the optimal parameter combination touches a preset physical boundary, and output data quality feedback information based on the diagnosis results.
[0101] Example 4 like Figure 6 As shown, the parameter self-calibration method for the photovoltaic power generation prediction model described in Example 2 is provided in this embodiment of the invention. The parameter self-calibration system for the photovoltaic power generation prediction model also includes a verification and diagnosis module for multi-dimensional verification and visualization diagnosis.
[0102] The multi-dimensional validation includes calculating the overall WMAPE of the validation set, daily detailed statistics of WMAPE, and weather stratification assessment based on the irradiance fluctuation coefficient; in the weather stratification assessment, the fluctuation coefficient is defined as the ratio of the standard deviation to the mean of irradiance during high irradiance periods (irradiance > 200 W / m²). , in, For time periods with irradiance > 200 W / m², μ and σ are the mean and standard deviation of irradiance for that time period, respectively. When the fluctuation coefficient < 0.35, it is classified as stable sunny; when 0.35 ≤ fluctuation coefficient ≤ 0.50, it is classified as cloudy and fluctuating; when the fluctuation coefficient > 0.50 or there are fewer than 10 effective high irradiance data, it is classified as cloudy / rainy / strong fluctuating. The average WMAPE under each weather type is calculated to clarify the applicable scenario boundaries of the physical model.
[0103] Example 5 This embodiment was implemented in a real-world 1.3MW rooftop photovoltaic power station: DC installed capacity approximately 1300kW and AC installed capacity 1100kW. Data spans from April 1, 2026 to June 21, 2026 (approximately 2.7 months), with a sampling interval of 1 hour. Meteorological data (including temperature, irradiance, wind speed, relative humidity, and cloud cover) was obtained from Open-Meteo, and photovoltaic power data was obtained from the power station's SCADA system.
[0104] like Figure 7 As shown, the specific implementation process of this embodiment of the invention is as follows.
[0105] Step 1: Data Loading and Multi-Level Cleaning Read the merged weather and power data (merged_weather_pv.csv), which contains approximately 2000 records, covering the period from April 1, 2026 to June 21, 2026.
[0106] S1.1: Negative power values are set to 0; there are 0 samples with power exceeding 1100kW. S1.2: Set the power to 0 for nighttime / dusk (irradiance <10W / m²); S1.3: Dynamic anomaly detection based on theoretical maximum power, no abnormal samples detected (thresholds are appropriately relaxed to adapt to scenarios with limited data). S1.4: After filling missing values in layers, there are approximately 2000 valid data entries.
[0107] Step 2: Training / Validation Set Partitioning: Training set: April 1, 2026 to May 31, 2026 (2 months), 786 valid daytime samples; Validation set: June 1, 2026 to June 21, 2026 (21 days), with 294 valid daytime samples.
[0108] Step 3: Perform differential evolution optimization: The code calls `scipy.optimize.differential_evolution` with the objective function WMAPE and parameter boundaries α∈[0.5,1.0], γ∈[-0.02,0.0]. It converges after 1500 iterations.
[0109] Step 4: Optimization Results: The optimal parameter output is shown in Table 9.
[0110] Table 9 Optimal Parameters
[0111] Neither parameter touched the boundary, indicating that the data quality is good.
[0112] The error indicators are shown in Table 10.
[0113] Table 10 Error Indicators
[0114] The results of the weather stratification assessment are shown in Table 11.
[0115] Table 11 Weather Stratification Assessment
[0116] The stratification results clearly show that the physical model has extremely high accuracy (<10%) under stable irradiance conditions, verifying the effectiveness of the present invention in scenarios with limited data.
[0117] Parameter boundary check: Neither of the two optimization parameters touched the boundary, and the system outputs "Data quality is good, optimization results are reliable".
[0118] Step 5: Comparison and Verification: To verify the superiority of this invention in scenarios with limited data, a comparison was made with a pure data-driven method (LightGBM hybrid model), and the comparison results are shown in Table 12.
[0119] Table 12 Comparison between the present invention and the LightGBM hybrid model:
[0120] Experimental results show that, with only two months of training data, the pure physics model of this invention has better accuracy than the data-driven hybrid model, and has significant advantages such as strong physical interpretability, clear parameter meaning, and no need for a large amount of labeled data.
[0121] Step 6: Deployment and Self-calibration Cycle: Save the optimal parameters (α=0.9474, γ=-0.0044) to the configuration file. The power plant power prediction system loads these parameters into the physical model each time it runs to obtain high-precision prediction results. It is recommended to automatically perform a parameter calibration process monthly to track component degradation and seasonal variations.
[0122] This invention uses the weighted average absolute percentage error (WMAPE) as the core evaluation index. WMAPE has the following technical advantages: (1) It is dimensionless and can be directly compared between power plants with different installed capacities; (2) The high power period has a significant contribution to WMAPE, which is highly consistent with the engineering needs of grid dispatch to focus on the high power period; (3) Compared with MAPE, WMAPE will not produce the numerical explosion problem caused by the denominator being too small when the power is close to zero.
[0123] Experiments show that the MAE of the same model on the validation set is 60.63 kW, but the MAE distribution is extremely uneven under different weather conditions (MAE is about 20 kW on sunny days and about 150 kW on cloudy and rainy days). MAE alone cannot comprehensively evaluate the model performance. WMAPE, on the other hand, can naturally weight the error according to the power generation, and more scientifically reflect the actual impact of the forecast on grid dispatch.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A parameter self-calibration method for a photovoltaic power generation prediction model, characterized in that, Includes the following steps: Step S1: Obtain historical operating data of the photovoltaic power station and corresponding meteorological data, and perform multi-level physical constraint cleaning on the data to obtain cleaned data; Step S2: Construct a refined physical model for predicting photovoltaic power generation, wherein the physical model includes the photoelectric conversion efficiency parameter α and the temperature coefficient parameter γ to be calibrated; Step S3: Use the differential evolution algorithm to globally optimize the photoelectric conversion efficiency parameter α and the temperature coefficient parameter γ in the physical model, with the goal of minimizing the weighted average absolute percentage error, to obtain the optimal parameter combination. ; Step S4: For the optimal parameter combination Perform parameter boundary diagnosis to determine whether the optimal parameter combination touches a preset physical boundary, and output data quality feedback information based on the diagnosis results.
2. The parameter self-calibration method for the photovoltaic power generation prediction model according to claim 1, characterized in that, The physical constraint multi-level cleaning includes at least basic hard constraint cleaning: setting all negative power values to zero, and marking power values exceeding 1.5 times the AC installed capacity as abnormal and removing them; It also includes dynamic anomaly detection based on theoretical maximum power, which calculates the theoretical maximum power using the component's nominal parameters and executes zoned cleaning rules according to the irradiance range.
3. The parameter self-calibration method for the photovoltaic power generation prediction model according to claim 1, characterized in that, The physical constraint multi-level cleaning also includes layered filling of missing values: for isolated missing points, the mean of a rolling window with a window size of 3 is used for filling; for short-term continuous missing points with a duration of ≤3 hours, linear interpolation is used to interpolate along the time axis; for long-term continuous missing points with a duration of >3 hours or missing points during the nighttime period, they are directly filled with 0.
4. The parameter self-calibration method for the photovoltaic power generation prediction model according to claim 1, characterized in that, The basic expression of the refined physical model is: , Where G is the total irradiance of the inclined surface. α γ represents the overall photoelectric conversion efficiency to be calibrated, and γ is the temperature coefficient to be calibrated. This refers to the operating temperature of the photovoltaic panel.
5. The parameter self-calibration method for the photovoltaic power generation prediction model according to claim 4, characterized in that, The operating temperature of the photovoltaic panel Calculations were performed using the nominal operating battery temperature model: , Wherein, NOCT is the nominal operating battery temperature. The ambient temperature.
6. The parameter self-calibration method for the photovoltaic power generation prediction model according to claim 1, characterized in that, The objective function for the weighted average absolute percentage error is defined as: , Where WMAPE is the weighted average absolute percentage error, P actual For actual power, P pred The power predicted by the model is summed within the range of valid daytime samples in the training set with irradiance greater than 10 W / m² and actual power greater than 0.
7. The parameter self-calibration method for the photovoltaic power generation prediction model according to claim 1, characterized in that, The differential evolution algorithm employs a best / 1 mutation strategy for each target vector X in the g-th generation population. i (g) Generate mutation vector V i (g) The formula is: , in, The optimal individual with the minimum objective function value in the g-th generation population, where F is the scaling factor and in Adaptive random selection within the interval and For each unique individual, and not equal to i, use a random index. , Population size.
8. The parameter self-calibration method for the photovoltaic power generation prediction model according to claim 1, characterized in that, The parameter boundary detection includes: calculating the distance between each parameter in the optimal parameter combination and the preset physical boundary. , in and These are the preset lower and upper bounds for this parameter, respectively. The boundary is determined when the photoelectric conversion efficiency α touches the upper limit of 1.0 or the temperature coefficient γ touches the lower limit. If α is 0.02, the data is considered to have a systematic bias; if α touches the lower limit of 0.5, the component is considered to be severely aged or the data is abnormal; if no parameter touches the boundary, the data quality is considered to be good and the optimization results are reliable.
9. The parameter self-calibration method for the photovoltaic power generation prediction model according to any one of claims 1 to 8, characterized in that, It also includes step S5: multi-dimensional verification and visualization diagnosis; the multi-dimensional verification includes calculating the overall WMAPE of the verification set, daily detailed statistics of WMAPE, and weather stratification assessment based on the irradiance fluctuation coefficient; in the weather stratification assessment, the fluctuation coefficient is defined as the ratio of the standard deviation of irradiance to the mean during high irradiance periods: , in, For time periods with irradiance > 200 W / m², μ and σ are the mean and standard deviation of irradiance for that time period, respectively. When the fluctuation coefficient < 0.35, it is classified as stable sunny; when 0.35 ≤ fluctuation coefficient ≤ 0.50, it is classified as cloudy and fluctuating; when the fluctuation coefficient > 0.50 or there are fewer than 10 effective high irradiance data, it is classified as cloudy / rainy / strong fluctuating. The average WMAPE under each weather type is calculated to clarify the applicable scenario boundaries of the physical model.
10. A parameter self-calibration system for a photovoltaic power generation prediction model, characterized in that, include: The data cleaning module is used to acquire historical operating data of photovoltaic power plants and corresponding meteorological data, and to perform multi-level physical constraint cleaning on the data to obtain cleaned data. The physical model building module is used to build a refined physical model for photovoltaic power generation prediction. The physical model includes the photoelectric conversion efficiency parameter α and the temperature coefficient parameter γ to be calibrated. The differential evolution optimization module is used to globally optimize the photoelectric conversion efficiency parameter α and the temperature coefficient parameter γ in the physical model using the differential evolution algorithm, with the goal of minimizing the weighted average absolute percentage error, to obtain the optimal parameter combination. ; The boundary diagnosis module is used to analyze the optimal parameter combination. Perform parameter boundary diagnosis to determine whether the optimal parameter combination touches a preset physical boundary, and output data quality feedback information based on the diagnosis results.