A photovoltaic power station power generation power prediction method and system

By constructing a physical sub-model and statistical sub-model database of photovoltaic power plants, and combining multi-dimensional measured data and historical error data, accurate prediction of photovoltaic power generation has been achieved. This solves the problems of poor scenario adaptability and insufficient prediction accuracy in existing technologies, and improves the reliability and adaptability of prediction.

CN122267708APending Publication Date: 2026-06-23TAUSHGAN DARYA HYDROPOWER BRANCH OF HUANENG XINJIANG ENERGY DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TAUSHGAN DARYA HYDROPOWER BRANCH OF HUANENG XINJIANG ENERGY DEVELOPMENT CO LTD
Filing Date
2026-01-28
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing methods for predicting the power generation of photovoltaic power plants suffer from poor adaptability to different scenarios, low reliability, and insufficient prediction accuracy. In particular, the prediction error increases significantly under sudden changes in lighting conditions or extreme weather.

Method used

We construct physical sub-models with different data processing functions and statistical sub-models for different scenarios, establish a model database, and make predictions through a physical and statistical fusion model. We also optimize the model by combining multidimensional measured data and historical error data.

Benefits of technology

It significantly improves the accuracy and reliability of photovoltaic power plant power generation prediction, enhances the adaptability to different lighting conditions and component types, and provides reliable support for the precise scheduling of the power system and the efficient operation and maintenance of photovoltaic power plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122267708A_ABST
    Figure CN122267708A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of photovoltaic power generation, and provides a photovoltaic power station power generation power prediction method and system. The method comprises the following steps: obtaining multi-dimensional measured data of a photovoltaic power station; constructing physical sub-models of different component data processing functions and statistical sub-models under different scenes respectively, and establishing a model database; determining an actual prediction scene, calling a target physical sub-model and a target statistical sub-model corresponding to the actual prediction scene from the model database, and constructing a physical and statistical fusion model according to the target physical sub-model and the target statistical sub-model; and obtaining a power generation power prediction result of the photovoltaic power station according to the multi-dimensional measured data and by using the physical and statistical fusion model. The scheme improves the accuracy of photovoltaic power station power generation power prediction, enhances the adaptability to different illumination conditions, component types and tracking modes, and provides reliable support for accurate dispatching of a power system, efficient operation and maintenance of a photovoltaic power station and large-scale grid-connected consumption of photovoltaic power generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of photovoltaic power generation technology, and in particular to a method and system for predicting the power generation of a photovoltaic power plant. Background Technology

[0002] As a core component of renewable energy, photovoltaic (PV) power generation has seen rapid and continuous growth in installed capacity, gradually increasing its share in the power system. However, the power output of PV power plants is susceptible to various complex factors, exhibiting significant intermittency, volatility, and randomness. These characteristics not only pose challenges to the stable operation and maintenance of PV power plants themselves but also impact the power balance, dispatch control, and safe and stable operation of the power grid. Therefore, accurately predicting the power output of PV power plants is crucial for the power system to formulate reasonable dispatch plans, optimize grid operation, and improve PV absorption capacity.

[0003] Currently, photovoltaic (PV) power generation prediction methods are mainly divided into two categories: physical model methods and statistical model methods. Regarding physical models, existing technologies mostly construct unified models based on the energy conversion principles of PV modules, resulting in poor model adaptability to various module types and installation methods. Regarding statistical models, existing technologies often employ single statistical models or simple scenario-based models, lacking refined segmentation of multi-dimensional scenarios. This leads to insufficient model adaptability to different scenarios, especially in special scenarios such as sudden changes in sunlight conditions or extreme weather. Due to the lack of targeted training data and model parameters, prediction errors increase significantly, making it difficult to meet the prediction needs of PV power plants with multiple module types.

[0004] It is evident that traditional photovoltaic power generation prediction methods suffer from technical problems such as poor scenario adaptability, low reliability, and insufficient prediction accuracy. Summary of the Invention

[0005] This invention provides a method and system for predicting the power generation of photovoltaic power plants, which solves the shortcomings of traditional photovoltaic power plant power generation prediction methods, such as poor scenario adaptability, low reliability and insufficient prediction accuracy.

[0006] On one hand, the present invention provides a method for predicting the power generation of a photovoltaic power plant, comprising: Obtain multi-dimensional measured data from photovoltaic power plants; Construct physical sub-models with different data processing functions for different components and statistical sub-models for different scenarios, and establish a model database. The actual prediction scenario is determined, and the target physical sub-model and target statistical sub-model corresponding to the actual prediction scenario are retrieved from the model database. Based on the target physical sub-model and the target statistical sub-model, a physical and statistical fusion model is constructed. Based on the multidimensional measured data and using the physical and statistical fusion model, the power generation prediction results of the photovoltaic power station are obtained.

[0007] According to the photovoltaic power generation prediction method provided by the present invention, multi-dimensional measured data of the photovoltaic power station are obtained, including: Acquire solar radiation data, photovoltaic module data, inverter data, environmental data, and power plant operation data to obtain multi-dimensional raw data; The multidimensional raw data is preprocessed at multiple levels to obtain multidimensional measured data.

[0008] According to the photovoltaic power generation prediction method provided by this invention, physical sub-models with different component data processing functions and statistical sub-models under different scenarios are constructed respectively, and a model database is established, including: Based on the core energy conversion principle of photovoltaic power plants, physical sub-models of data processing functions of different components are constructed. Based on initial basic training combined with scenario-specific training, statistical sub-models for different scenarios are constructed. Both the physical sub-model and the statistical sub-model are stored in the model database.

[0009] According to the photovoltaic power generation prediction method provided by this invention, based on the core energy conversion principle of a photovoltaic power station, physical sub-models with different component data processing functions are constructed, including: An irradiance correction sub-model is established based on the previously obtained irradiance forecast data and historical irradiance data; The effective irradiance is obtained by correcting the irradiance deviation of the photovoltaic module based on the irradiance correction sub-model, and the power calculation sub-model of various photovoltaic modules is established using the pre-established equivalent circuit model of the photovoltaic module. Based on the pre-obtained inverter efficiency curve, a sub-model of inverter conversion efficiency is established. The output power of all inverters in the photovoltaic power station is summarized to establish a total power aggregation sub-model of the photovoltaic power station; The irradiance correction sub-model, the power calculation sub-model, the conversion efficiency sub-model, and the total power aggregation sub-model are used as physical sub-models.

[0010] According to the photovoltaic power generation prediction method provided by this invention, based on initial basic training combined with scenario-specific training, statistical sub-models for different scenarios are constructed, including: Establish a basic model and a training sample set containing multi-dimensional input samples and power plant output power samples; The basic model is initially trained using the training sample set to obtain the trained basic model. Based on different types of lighting conditions, the trained base model is subjected to scene sub-training to obtain statistical sub-models for different scenes.

[0011] According to the photovoltaic power generation prediction method provided by the present invention, both the physical sub-model and the statistical sub-model are stored in a model database, including: The core scene features under different prediction scenarios are identified, and the core scene features are standardized and encoded to obtain scene feature labels for different prediction scenarios. Both the physical sub-model and the statistical sub-model are indexed according to their respective scene feature labels and stored in the model database.

[0012] According to the photovoltaic power generation prediction method provided by the present invention, the method retrieves the target physical sub-model and the target statistical sub-model corresponding to the actual prediction scenario from the model database, including: Determine the target scene feature label corresponding to the actual predicted scene; Using the target scene feature labels as query conditions, the scene feature labels of each sub-model in the model database are traversed to obtain the target physical sub-model and target statistical sub-model corresponding to the actual predicted scene.

[0013] According to the photovoltaic power generation prediction method provided by the present invention, a physical and statistical fusion model is constructed based on the target physical sub-model and the target statistical sub-model, including: Obtain the historical error data of the target physical sub-model and the target statistical sub-model respectively; Based on the historical error data, the model weights of the target physical sub-model and the target statistical sub-model are calculated respectively. According to the model weights, the target physical sub-model and the target statistical sub-model are fused to obtain a physical and statistical fusion model.

[0014] The photovoltaic power generation prediction method provided by the present invention further includes: The data deviation value between the predicted power generation result and the measured power generation result is determined according to a set period. If the data deviation value is higher than the set deviation threshold, then at least some of the model parameters of the physical and statistical fusion model are optimized.

[0015] On the other hand, the present invention also provides a photovoltaic power plant power generation prediction system, comprising: The acquisition module is used to acquire multi-dimensional measured data from photovoltaic power plants. Establish modules to build physical sub-models for different component data processing functions and statistical sub-models for different scenarios, and establish a model database; The module constructs a physical and statistical fusion model based on the actual prediction scenario, retrieves the target physical sub-model and target statistical sub-model corresponding to the actual prediction scenario from the model database, and constructs the physical and statistical fusion model based on the target physical sub-model and the target statistical sub-model. The prediction module is used to obtain the predicted power generation of the photovoltaic power station based on the multidimensional measured data and using the physical and statistical fusion model.

[0016] The photovoltaic power generation prediction method and system provided by this invention lays a data foundation for accurate prediction by acquiring multi-dimensional measured data. By constructing physical sub-models with different component data processing functions and statistical sub-models for different scenarios and establishing a model database, it achieves refined adaptation to multiple types of photovoltaic components and complex operating scenarios. Then, by accurately retrieving the target sub-model to construct a physical and statistical fusion model, it fully leverages the stability of the physical model based on the energy conversion principle and the scenario adaptability of the statistical model based on historical data mining. This effectively avoids the prediction limitations of a single model in specific scenarios. Finally, it outputs prediction results by combining multi-dimensional measured data. This not only significantly improves the accuracy of photovoltaic power generation prediction and enhances the adaptability to different lighting conditions, component types, and tracking methods, but also provides reliable support for precise power system scheduling, efficient operation and maintenance of photovoltaic power plants, and large-scale grid connection and consumption of photovoltaic power generation. It has outstanding practicality and application value. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the photovoltaic power generation prediction method provided in an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the specific composition of the multidimensional measured data in an embodiment of the present invention; Figure 3 This is a schematic diagram of the photovoltaic power generation prediction system provided in an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0020] The following is combined with Figures 1 to 3 This invention describes the detailed scheme of the photovoltaic power generation prediction method and system provided in the embodiments of the present invention.

[0021] like Figure 1 As shown, the photovoltaic power generation prediction method provided in this embodiment of the invention mainly includes the following steps: Step 110: Obtain multi-dimensional measured data of the photovoltaic power station.

[0022] In practical applications, sensors can be deployed and connected to meteorological platforms and monitoring equipment to collect multi-dimensional measured data such as key parameters of core components like solar radiation, photovoltaic modules, and inverters, environmental parameters, and power plant operating conditions, thereby covering the needs of different lighting conditions, module types, and extreme scenarios.

[0023] Step 120: Construct physical sub-models for different component data processing functions and statistical sub-models for different scenarios, and establish a model database.

[0024] In this embodiment, by constructing physical and statistical sub-models, effective data can be provided for the construction of subsequent fusion models, thereby improving the data reliability of the prediction process.

[0025] Step 130: Determine the actual prediction scenario, retrieve the target physical sub-model and target statistical sub-model corresponding to the actual prediction scenario from the model database, and construct the physical and statistical fusion model based on the target physical sub-model and target statistical sub-model.

[0026] It is understood that this embodiment uses a fusion of physical and statistical methods to generate the model. The resulting fusion model can effectively avoid the prediction limitations of a single model in a specific scenario, thereby achieving more accurate power generation prediction.

[0027] Step 140: Based on multidimensional measured data and using a physical and statistical fusion model, obtain the predicted power generation results of the photovoltaic power station.

[0028] The solution provided in this embodiment improves the scenario adaptability, prediction accuracy, and prediction reliability of the power generation prediction process, and can provide effective data support for the efficient operation and maintenance of photovoltaic power plants.

[0029] In one embodiment, such as Figure 2 As shown, multi-dimensional measured data of photovoltaic power plants are obtained, specifically including: First, we acquire solar radiation data, photovoltaic module data, inverter data, environmental data, and power plant operation data to obtain multi-dimensional raw data.

[0030] In practical applications, multiple sets of high-precision irradiance sensors can be deployed to cover different areas within the photovoltaic power station, recording solar radiation data including various core parameters such as total irradiance, diffuse irradiance, direct irradiance, solar altitude angle, and azimuth angle.

[0031] In the photovoltaic module data acquisition stage, basic parameters of various types of photovoltaic modules can be collected, such as rated power, open-circuit voltage, short-circuit current, maximum power point voltage, maximum power point current, temperature coefficient, module size, and encapsulation materials. Furthermore, through module-level monitoring equipment, real-time output voltage, current, power, module surface temperature, module cleanliness, and module aging degree of each photovoltaic module can be collected.

[0032] In the inverter data acquisition stage, basic inverter parameters such as rated output power, input voltage range, output voltage range, conversion efficiency curve, maximum conversion efficiency, and MPPT tracking accuracy can be collected, as well as real-time inverter operating data such as input power, output power, input voltage, output voltage, output current, power factor, inverter internal temperature, MPPT operating status, and fault alarm information.

[0033] In the environmental data collection phase, conventional environmental parameters such as ambient temperature, relative humidity, wind speed, wind direction, and atmospheric pressure can be collected through meteorological stations. Special environmental parameters such as precipitation intensity, snowfall, visibility, dust concentration, and cloud thickness can also be collected, thereby providing supplementary data for forecasting under extreme lighting conditions.

[0034] In the power plant operation data acquisition stage, basic operation data such as total power output, bus voltage, grid-connected current, power factor, energy storage equipment charging and discharging status, and reactive power compensation device operation status can be collected. Power plant operation and maintenance logs can also be recorded, including component cleaning time, maintenance records, fault handling status, component replacement records, etc.

[0035] Then, the multidimensional raw data is preprocessed at multiple levels to obtain multidimensional measured data.

[0036] In this embodiment, the multi-level preprocessing specifically involves core preprocessing operations such as data cleaning, standardization, and feature extraction. Specifically, in the data cleaning stage, the 3σ principle and box plot method can be used to identify abnormal data, such as zero values ​​and extreme values ​​caused by sensor malfunctions, and missing values ​​caused by data transmission interruptions. Abnormal values ​​are replaced or deleted using methods such as interpolation based on data from adjacent time points or replacement with the mean under similar operating conditions. For short-term missing data, linear interpolation can be used for completion. For long-term missing data, K-nearest neighbor imputation based on similar daily data can be used for completion, thereby ensuring data continuity.

[0037] In the standardization process, parameters with different dimensions can be standardized or normalized to eliminate the impact of dimensional differences on model training. In the feature extraction stage, time feature data such as year, month, day, hour, minute, season, solar term, solar altitude angle, and cosine and sine values ​​of azimuth angle can be extracted. Furthermore, derived feature data can be constructed by calculating the rate of change of irradiance, the difference between component temperature and ambient temperature, inverter load rate, and the ratio of irradiance to component power. In addition, based on key parameters such as solar altitude angle, cloud cover, and ambient temperature, Euclidean distance or cosine similarity algorithms can be used to screen historical dates similar to the predicted daily operating conditions, providing targeted training data for the model.

[0038] In one embodiment, physical sub-models with different component data processing functions and statistical sub-models for different scenarios are constructed respectively, and a model database is established, including: On the one hand, based on the core energy conversion principle of photovoltaic power plants, physical sub-models of data processing functions of different components are constructed.

[0039] In a specific implementation, based on the core energy conversion principle of a photovoltaic power plant, physical sub-models for the data processing functions of different components are constructed, specifically including: The first step is to establish an irradiance correction sub-model based on the previously obtained irradiance forecast data and historical irradiance data.

[0040] In practical applications, we can first integrate the pre-obtained irradiance forecast data and historical irradiance data to establish a basic comparison dataset. By calculating the deviation distribution between the forecast and measured values, we can construct a deviation correction function using linear regression or machine learning fitting methods to initially correct the systematic error of irradiance. At the same time, considering the topography of the power station, we can obtain data such as slope, aspect, and distribution of obstructions through on-site surveys, calculate the terrain correction coefficient, and spatially correct the irradiance distribution in different areas. Finally, combining the photovoltaic module tracking methods, we can establish adaptation algorithms. For fixed modules, we can calculate the solar incidence angle based on the installation tilt angle and azimuth angle to correct the irradiance. For single-axis and dual-axis tracking modules, we can combine the real-time angle of the tracking system or the preset running trajectory to accurately calculate the effective irradiance actually received by the module, forming a complete irradiance correction sub-model.

[0041] The second step involves correcting the irradiance deviation of the photovoltaic module using the irradiance correction sub-model to obtain the effective irradiance, and then using the pre-established equivalent circuit model of the photovoltaic module to establish power calculation sub-models for various photovoltaic modules.

[0042] In this step, firstly, core parameters such as rated power, open-circuit voltage, short-circuit current, temperature coefficient, and encapsulation materials of different types of photovoltaic modules, such as monocrystalline silicon, polycrystalline silicon, and thin-film, are pre-acquired to establish a module parameter library. Based on the physical mechanism of energy conversion of photovoltaic modules, a single-diode or dual-diode equivalent circuit model is selected as the basic framework, and the effective irradiance output from the first step of the irradiance correction sub-model and the module surface temperature collected by the infrared thermometer are used as core input parameters. By substituting the characteristic parameters of the corresponding type of module in the module parameter library, the open-circuit voltage, short-circuit current, and maximum power point of the module under the current operating conditions are calculated. At the same time, the attenuation coefficient fitted based on the module's operating years and historical attenuation data, as well as the cleanliness coefficient obtained through transmittance monitoring or image recognition, are introduced to correct the calculated power value. Finally, a dedicated power calculation sub-model is established for each type of photovoltaic module to ensure the accuracy of power calculation for different module types.

[0043] The third step is to establish a sub-model of the inverter's conversion efficiency based on the pre-obtained inverter efficiency curve.

[0044] In practical applications, basic parameters of the inverter, such as rated output power, input and output voltage range, conversion efficiency curves under different load rates, maximum conversion efficiency, and MPPT tracking accuracy, can be collected in advance. Using the inverter input power as the core input variable, and based on the inverter efficiency curve, a piecewise interpolation method is used to establish the mapping relationship between load rate and conversion efficiency, enabling preliminary calculations of conversion efficiency under different load conditions. Simultaneously, internal operating temperature data is collected through the inverter's built-in monitoring module, and a temperature correction coefficient model is established to analyze the impact of temperature on inverter conversion efficiency. This coefficient is then used to dynamically correct the preliminary calculated conversion efficiency. Furthermore, combined with the inverter's MPPT operating status monitoring data, when the MPPT is in an unstable operating state, a fluctuation correction factor is introduced to further optimize the conversion efficiency calculation results. Finally, a sub-model of inverter conversion efficiency that accurately reflects actual operating conditions is constructed.

[0045] The fourth step is to summarize the output power of all inverters in the photovoltaic power station and establish a total power aggregate sub-model of the photovoltaic power station.

[0046] In this embodiment, the real-time output power data of all inverters can be collected first through the power station monitoring system, i.e., the output results of the inverter conversion efficiency sub-model, and then summarized to form the total output power of the inverters. Subsequently, based on the electrical topology of the power station, the design parameters of the bus, transformer, and transmission line are collected, and a loss calculation model is established. The bus loss is calculated based on the bus current and resistance, the transformer loss is estimated using empirical formulas combined with the load rate and operating temperature, and the line loss is calculated based on the transmission current, resistance, and length. The total output power of the inverters is then deducted from the above-mentioned losses to obtain the preliminary total output power of the power station. If the power station is equipped with energy storage equipment or reactive power compensation devices, their operating strategies, such as energy storage charging and discharging plans and reactive power adjustment targets, need to be obtained in advance. Combined with their real-time operating status, the preliminary total power is then corrected a second time to finally form a total power aggregation sub-model that can reflect the actual grid connection capability of the power station.

[0047] The fifth step is to use the irradiance correction sub-model, power calculation sub-model, conversion efficiency sub-model, and total power aggregation sub-model as physical sub-models.

[0048] On the other hand, based on the initial basic training combined with scenario-specific training, statistical sub-models for different scenarios are constructed.

[0049] In a specific implementation, based on initial basic training combined with scenario-specific training, statistical sub-models for different scenarios are constructed, including: The first step is to establish a basic model and a training sample set containing multi-dimensional input samples and power plant output power samples.

[0050] In practical applications, firstly, based on the scenario adaptability requirements of photovoltaic power prediction, support vector machines, random forests, gradient boosting trees, or long short-term memory networks can be selected as the base model. Traditional machine learning models are suitable for scenarios with moderate data volume and well-defined features, while deep learning models are suitable for long-period data scenarios containing time-series features, ensuring that the base model possesses strong nonlinear fitting and generalization capabilities. Subsequently, a training sample set is constructed. The input sample dimensions cover preprocessed multi-dimensional measured data, and the output power sample selects the measured total active power data of the power plant, synchronized with the input sample time, ensuring the consistency of the sample time series. Simultaneously, the sample set undergoes quality screening, removing extreme outliers, and the training and validation sets are divided in a 7:3 ratio to provide high-quality data support for subsequent model training.

[0051] The second step is to perform initial basic training on the base model using the training sample set to obtain the trained base model.

[0052] In practical applications, the basic model can be initially trained using the 5-fold cross-validation method based on the constructed training sample set. Specifically, the training set can be randomly divided into 5 subsets, and 4 subsets can be selected as training data and 1 subset as validation data in turn, and the training and validation can be completed in 5 rounds.

[0053] After training, the generalization ability of the model can be verified using an independent validation set. If the validation error meets the preset threshold, it is determined as the base model after training; if it does not meet the threshold, the input features or hyperparameter range is readjusted and training is performed again until the target is met.

[0054] The third step involves training the trained base model for scene sub-models based on different types of lighting conditions, resulting in statistical sub-models for different scenes.

[0055] In this step, firstly, based on solar radiation intensity, cloud cover distribution, and weather type, the predicted scenarios can be divided into typical lighting scenarios, sunny days, cloudy days, overcast days, and extreme weather scenarios. Then, corresponding subsets are selected from the original training sample set according to scenario labels to form a dedicated training sample set for each scenario, ensuring sufficient sample size for each scenario. For each scenario's sample set, using the trained base model as the initial framework, some core parameters are frozen, and only the upper-layer adaptation parameters are fine-tuned for scenario-specific training. This preserves the general fitting ability of the base model while enhancing the model's adaptability to the data patterns of specific scenarios.

[0056] After training, the model performance is evaluated using validation subsets for each scenario. If the error of a model in a certain scenario is reduced by more than 30% compared to the base model, it is determined to be the statistical sub-model for that scenario. Finally, a statistical sub-model library covering all illumination scenarios and adapting to different irradiance conditions is formed.

[0057] Finally, both the physical sub-model and the statistical sub-model are stored in the model database.

[0058] In one specific implementation, both the physical sub-model and the statistical sub-model are stored in the model database, specifically including: The first step is to identify the core scene features under different prediction scenarios and to standardize and encode these core scene features to obtain scene feature labels for different prediction scenarios.

[0059] In this embodiment, core scene feature dimensions covering the adaptation requirements of both physical and statistical sub-models can be identified, including illumination conditions, photovoltaic module type, module tracking method, prediction time scale, and environmental level. This ensures that the feature dimensions comprehensively cover the core elements of scene adaptation for both types of sub-models. Subsequently, a unified standardized coding rule is established. For example, a combination of 4 digits and 1 letter can be used. The first digit represents the illumination conditions (e.g., 1 for sunny, 2 for cloudy, 3 for overcast, 4 for extreme weather); the second digit represents the module type (e.g., 1 for monocrystalline silicon, 2 for polycrystalline silicon, 3 for thin film); the third digit represents the tracking method (e.g., 1 for fixed, 2 for single-axis, 3 for dual-axis); the fourth digit represents the time scale (e.g., 1 for ultra-short-term, 2 for short-term, 3 for medium- to long-term); and the letter represents the environmental level (e.g., A for normal environment, B for extreme environment). This coding rule transforms each prediction scene into a unique scene feature label, achieving standardization and identifiability of scene features.

[0060] The second step is to index both the physical sub-model and the statistical sub-model according to their respective scene feature labels and store them in the model database.

[0061] In practical applications, the scene affiliation of all constructed physical and statistical sub-models can be determined first, clarifying the core scene features corresponding to each sub-model, and then matching the generated scene feature labels. Subsequently, a two-layer index system is established in the model database for the two types of sub-models. The first layer is the scene feature label index, which associates each sub-model with its corresponding scene feature label, forming a label-sub-model ID mapping table. The second layer is the sub-model attribute index, which records attribute information such as the ID, model type, construction time, historical error data, and applicable parameter range of each sub-model.

[0062] Finally, the complete files of the physical sub-model and statistical sub-model are stored in HDFS according to the directory structure of model type, scene feature label, and sub-model ID. At the same time, the index information is written to the MySQL database to achieve accurate association between the sub-model and the scene label, ensuring that the target sub-model can be quickly retrieved through the scene feature label during subsequent prediction.

[0063] In one embodiment, retrieving the target physical sub-model and target statistical sub-model corresponding to the actual prediction scenario from the model database specifically includes: First, determine the target scene feature labels corresponding to the actual prediction scenario.

[0064] In this embodiment, core data of the actual prediction scenario can be collected first through the real-time monitoring system of the photovoltaic power station. The currently operating photovoltaic module type and tracking method are retrieved from the power station equipment parameter database. The prediction timescale is determined based on actual scheduling or maintenance needs, and the environmental level is determined by combining the real-time ambient temperature range and the presence of extreme weather. Subsequently, according to the standardized coding rules established above, the actual scenario features are transformed into unique target scenario feature labels, completing the accurate determination of the target scenario feature labels and providing a unified query basis for subsequent sub-model query matching.

[0065] Then, using the target scene feature labels as query conditions, the scene feature labels of each sub-model in the model database are traversed to match the target physical sub-model and target statistical sub-model corresponding to the actual prediction scene.

[0066] In practical applications, the target scene feature labels are used as the core query conditions to initiate the retrieval process of the model database. First, the two-layer index system in the MySQL database is accessed. The first-layer label-to-submodel ID mapping table quickly filters out all submodel IDs associated with that label. Simultaneously, the second-layer submodel attribute index distinguishes between the IDs corresponding to physical submodels and statistical submodels. If a perfectly matching submodel ID is found, the corresponding target physical submodel and target statistical submodel are directly locked. If no perfectly matching result is found, fuzzy matching is performed based on priorities such as lighting conditions, component type, tracking method, time scale, and environment level. The submodel with the highest similarity and lowest historical error is selected as the candidate target submodel, and the matching similarity is recorded.

[0067] After determining the target sub-model ID, the storage path corresponding to the model type, scene feature label, and sub-model ID in HDFS is located through the index information. The complete sub-model file is loaded, and after integrity and validity verification, the qualified target physical sub-model and target statistical sub-model are loaded into the memory of the prediction computing node, thereby completing the accurate retrieval of the sub-model.

[0068] In one embodiment, a physical-statistical fusion model is constructed based on the target physical sub-model and the target statistical sub-model, specifically including: The first step is to obtain the historical error data for both the target physical sub-model and the target statistical sub-model.

[0069] In this embodiment, after the target physical sub-model and the target statistical sub-model are retrieved, the historical error data of the two types of sub-models can be extracted through the sub-model attribute index of the model database. Specifically, this includes the successive prediction error records under the same scenario in the past 30 days. The core indicators are the mean absolute percentage error and the root mean square error. At the same time, the working condition data corresponding to the error is associated.

[0070] Afterwards, the extracted error data can be filtered to remove extreme error values ​​caused by equipment failure or abnormal data transmission, and retain valid error samples. Finally, the effective historical error datasets of the target physical sub-model and the target statistical sub-model are summarized to provide a quantitative basis for weight calculation.

[0071] The second step is to calculate the model weights of the target physical sub-model and the target statistical sub-model based on historical error data.

[0072] In this step, historical error data is used as the core. The inverse error weighting method can be used to calculate the weights of the two types of sub-models, ensuring that the model with the smaller error receives a higher weight. First, the average MAPE is used as the core calculation indicator, taking into account both the intuitiveness and stability of the model's prediction accuracy, to preliminarily calculate the model weights of the target physical sub-model and the target statistical sub-model.

[0073] Subsequently, weighting constraints are applied. In extreme weather scenarios, the model weight of the target physical sub-model is forced to be above 0.6 to prevent the statistical sub-model from making inaccurate predictions due to insufficient extreme data. In normal scenarios, the model weights of both the target physical sub-model and the target statistical sub-model are ensured to be between 0.2 and 0.8 to prevent a single model from dominating the fusion result. Finally, the rationality of the weights is verified by RMSE. If the RMSE difference between the two sub-models exceeds 50%, the initial weights are fine-tuned to determine stable and reliable model weights.

[0074] The third step is to fuse the target physical sub-model and the target statistical sub-model according to the model weights to obtain the physical and statistical fusion model.

[0075] In this step, input data from the target physical sub-model and the target statistical sub-model can be received separately to ensure that the input data accurately matches the interface requirements of the two types of sub-models. Then, the two types of sub-models are loaded into the computing layer of the fusion architecture and run synchronously to obtain independent prediction results. The target physical sub-model outputs the physical predicted power, and the target statistical sub-model outputs the statistical predicted power. The fusion layer embeds the weight parameters determined in the second step, and the final fused predicted power is obtained by weighted summation. At the same time, based on the historical error distribution of the two types of sub-models, the confidence interval of the prediction result is calculated using the variance synthesis method. Finally, the fusion model structure is encapsulated, and the input and output interfaces, weight parameters, sub-model version information, and applicable scenarios are clarified to form a complete physical and statistical fusion model for subsequent photovoltaic power plant power prediction.

[0076] In one embodiment, the above-mentioned photovoltaic power plant power generation prediction method may further include: First, the data deviation between the predicted power generation and the actual power generation is determined according to the set cycle.

[0077] In practical applications, the deviation calculation process can be triggered according to a preset cycle. First, the measured power data synchronized with the power generation prediction result is extracted from the real-time monitoring system of the photovoltaic power station. This measured data needs to be preprocessed. Then, the absolute deviation between the prediction result and the measured result at the same time is calculated. Finally, the deviation data is summarized according to the set cycle. If it is a multi-time period prediction, the average relative deviation value within the cycle is calculated as the final data deviation value of the cycle, which provides a basis for model optimization.

[0078] Then, if the data deviation value is higher than the set deviation threshold, at least some of the model parameters of the physical and statistical fusion model are optimized.

[0079] In this embodiment, a deviation threshold can be preset, such as a relative deviation threshold of 5%, which can be dynamically adjusted according to the prediction time scale. The ultra-short-term prediction threshold is set to 3%-5%, the short-term prediction threshold is set to 5%-8%, and the medium- and long-term prediction threshold is set to 8%-10%. The data deviation value is compared with the set deviation threshold. If the data deviation value is higher than the set deviation threshold, the model parameter optimization process can be started. The weight parameters of the fusion model are optimized first. Based on the measured data of the current period and the prediction data of the two types of sub-models, the real-time error of the target physical sub-model and the target statistical sub-model is recalculated. The model weights are updated according to the error reciprocal weighting method, while retaining the weight constraint rules.

[0080] If the data deviation value still does not drop below the set deviation threshold after weight optimization, the core parameters of the sub-model are further optimized. For the target physical sub-model, key parameters such as irradiance correction coefficient and component temperature coefficient are fine-tuned; for the target statistical sub-model, hyperparameters are fine-tuned, and an incremental training set is formed using the measured data of the current period and historical similar scene data to conduct a small number of iterative trainings on the two types of sub-models. After optimization, the updated weight parameters and sub-model parameters are re-embedded into the fusion model, and parameter optimization logs are recorded to ensure that the model performance remains stable.

[0081] In summary, the photovoltaic power generation prediction method provided in this embodiment of the invention comprehensively collects multi-dimensional measured data, constructs refined physical and statistical sub-models and establishes a standardized indexed model database by combining different component characteristics and scenario requirements, and achieves accurate retrieval and dynamic weighted fusion of target sub-models. At the same time, through periodic deviation monitoring and parameter optimization mechanisms, the model performance is continuously iterated. This fully leverages the stability of the physical model and the scenario adaptability of the statistical model, effectively adapting to multiple types of photovoltaic modules, different light and environmental conditions, significantly improving the accuracy and reliability of power generation prediction. It also provides high-quality data support for precise power system scheduling and efficient operation and maintenance of photovoltaic power plants, facilitating the large-scale grid connection and consumption of photovoltaic power generation, and possesses strong practicality and application value.

[0082] Based on the same general inventive concept, this invention also protects a photovoltaic power generation prediction system. The photovoltaic power generation prediction system provided by this invention will be described below. The photovoltaic power generation prediction system described below can be referred to in correspondence with the photovoltaic power generation prediction method described above.

[0083] like Figure 3 As shown, the photovoltaic power generation prediction system provided in this embodiment of the invention specifically includes: The acquisition module 210 is used to acquire multi-dimensional measured data of photovoltaic power plants.

[0084] Module 220 is established to construct physical sub-models with different component data processing functions and statistical sub-models for different scenarios, and to establish a model database.

[0085] Module 230 is constructed to determine the actual prediction scenario, retrieve the target physical sub-model and target statistical sub-model corresponding to the actual prediction scenario from the model database, and construct a physical and statistical fusion model based on the target physical sub-model and target statistical sub-model.

[0086] The prediction module 240 is used to obtain the power generation prediction results of the photovoltaic power station based on multidimensional measured data and using a physical and statistical fusion model.

[0087] Regarding the system in the above embodiments, the specific ways in which each module performs operations have been described in detail in the embodiments of the relevant methods, and will not be elaborated further here.

[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting the power generation of a photovoltaic power plant, characterized in that, include: Obtain multi-dimensional measured data from photovoltaic power plants; Construct physical sub-models with different data processing functions for different components and statistical sub-models for different scenarios, and establish a model database. The actual prediction scenario is determined, and the target physical sub-model and target statistical sub-model corresponding to the actual prediction scenario are retrieved from the model database. Based on the target physical sub-model and the target statistical sub-model, a physical and statistical fusion model is constructed. Based on the multidimensional measured data and using the physical and statistical fusion model, the power generation prediction results of the photovoltaic power station are obtained.

2. The photovoltaic power generation prediction method according to claim 1, characterized in that, Obtain multi-dimensional measured data from photovoltaic power plants, including: Acquire solar radiation data, photovoltaic module data, inverter data, environmental data, and power plant operation data to obtain multi-dimensional raw data; The multidimensional raw data is preprocessed at multiple levels to obtain multidimensional measured data.

3. The photovoltaic power generation prediction method according to claim 1, characterized in that, Physical sub-models with different data processing functions and statistical sub-models for different scenarios are constructed separately, and a model database is established, including: Based on the core energy conversion principle of photovoltaic power plants, physical sub-models of data processing functions of different components are constructed. Based on initial basic training combined with scenario-specific training, statistical sub-models for different scenarios are constructed. Both the physical sub-model and the statistical sub-model are stored in the model database.

4. The photovoltaic power generation prediction method according to claim 3, characterized in that, Based on the core energy conversion principle of photovoltaic power plants, physical sub-models for data processing functions of different components are constructed, including: An irradiance correction sub-model is established based on the previously obtained irradiance forecast data and historical irradiance data; The effective irradiance is obtained by correcting the irradiance deviation of the photovoltaic module based on the irradiance correction sub-model, and the power calculation sub-model of various photovoltaic modules is established using the pre-established equivalent circuit model of the photovoltaic module. Based on the pre-obtained inverter efficiency curve, a sub-model of inverter conversion efficiency is established. The output power of all inverters in the photovoltaic power station is summarized to establish a total power aggregation sub-model of the photovoltaic power station; The irradiance correction sub-model, the power calculation sub-model, the conversion efficiency sub-model, and the total power aggregation sub-model are used as physical sub-models.

5. The photovoltaic power generation prediction method according to claim 3, characterized in that, Based on initial basic training combined with scenario-specific training, statistical sub-models for different scenarios are constructed, including: Establish a basic model and a training sample set containing multi-dimensional input samples and power plant output power samples; The basic model is initially trained using the training sample set to obtain the trained basic model. Based on different types of lighting conditions, the trained base model is subjected to scene sub-training to obtain statistical sub-models for different scenes.

6. The photovoltaic power generation prediction method according to claim 3, characterized in that, Both the physical sub-model and the statistical sub-model are stored in the model database, including: The core scene features under different prediction scenarios are identified, and the core scene features are standardized and encoded to obtain scene feature labels for different prediction scenarios. Both the physical sub-model and the statistical sub-model are indexed according to their respective scene feature labels and stored in the model database.

7. The photovoltaic power generation prediction method according to claim 6, characterized in that, Retrieve the target physical sub-model and target statistical sub-model corresponding to the actual prediction scenario from the model database, including: Determine the target scene feature label corresponding to the actual predicted scene; Using the target scene feature labels as query conditions, the scene feature labels of each sub-model in the model database are traversed to obtain the target physical sub-model and target statistical sub-model corresponding to the actual predicted scene.

8. The photovoltaic power generation prediction method according to claim 1, characterized in that, Based on the target physical sub-model and the target statistical sub-model, a physical and statistical fusion model is constructed, including: Obtain the historical error data of the target physical sub-model and the target statistical sub-model respectively; Based on the historical error data, the model weights of the target physical sub-model and the target statistical sub-model are calculated respectively. According to the model weights, the target physical sub-model and the target statistical sub-model are fused to obtain a physical and statistical fusion model.

9. The photovoltaic power generation prediction method according to claim 1, characterized in that, The method further includes: The data deviation value between the predicted power generation result and the measured power generation result is determined according to a set period. If the data deviation value is higher than the set deviation threshold, then at least some of the model parameters of the physical and statistical fusion model are optimized.

10. A photovoltaic power plant power generation prediction system, characterized in that, include: The acquisition module is used to acquire multi-dimensional measured data from photovoltaic power plants. Establish modules to build physical sub-models for different component data processing functions and statistical sub-models for different scenarios, and establish a model database; The module constructs a physical and statistical fusion model based on the actual prediction scenario, retrieves the target physical sub-model and target statistical sub-model corresponding to the actual prediction scenario from the model database, and constructs the physical and statistical fusion model based on the target physical sub-model and the target statistical sub-model. The prediction module is used to obtain the predicted power generation of the photovoltaic power station based on the multidimensional measured data and using the physical and statistical fusion model.