Wide-area photovoltaic power prediction method and system based on space-time residual feature fusion

By constructing a photovoltaic power prediction method that integrates spatiotemporal residual features, and utilizing multi-source data and multi-level models, the problem of unexplored spatiotemporal patterns in wide-area photovoltaic power prediction is solved, achieving high-precision regional photovoltaic power prediction, especially with adaptive prediction capabilities under complex weather conditions.

CN122026337AActive Publication Date: 2026-05-12SHANDONG JIANZHU UNIV
View PDF 12 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG JIANZHU UNIV
Filing Date
2026-04-14
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing photovoltaic power prediction methods fail to fully exploit the regional spatiotemporal patterns contained in the residuals over a wide area, lack modeling for geographically dispersed smoothing effects, and the existing general prediction framework is not closely integrated with the physical processes of macro-meteorological driving regional total power output, resulting in insufficient prediction accuracy.

Method used

A wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion is constructed. By acquiring multi-source heterogeneous data, regional macro-meteorological features are extracted, multiple first-level prediction models are used for parallel prediction and then aggregated to construct spatiotemporal residual features. Finally, the prediction value of the total photovoltaic power in the region is obtained through a second-level correction model.

Benefits of technology

It achieves in-depth modeling of the overall regional power output pattern, improving prediction accuracy. In particular, it has self-learning and adaptive capabilities under sudden weather conditions, continuously and dynamically improving prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122026337A_ABST
    Figure CN122026337A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of photovoltaic prediction, and provides a wide-area photovoltaic power prediction method and system based on space-time residual feature fusion, and the method comprises the steps: carrying out the parallel prediction of a plurality of trained first-stage prediction models based on regional macroscopic meteorological features, and then carrying out the aggregation, and obtaining a first-stage aggregation prediction value and a prediction residual sequence; constructing a corresponding space-time residual feature according to the prediction residual sequence, and fusing the space-time residual feature with the time-aligned regional macroscopic meteorological feature to obtain an enhanced feature set; based on the enhanced feature set, using the trained secondary correction model to perform prediction to obtain a final prediction power correction amount; and taking the sum of the final prediction power correction and the primary aggregation prediction value as a regional photovoltaic total power prediction value. According to the method, systematic prediction deviation caused by regional geographic dispersibility is obviously learned and corrected, so that dependence on bottom-layer mass data is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of photovoltaic prediction technology, specifically relating to a wide-area photovoltaic power prediction method and system based on spatiotemporal residual feature fusion. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] As an important component of clean energy, photovoltaic power generation has seen its installed capacity grow rapidly and continuously. However, photovoltaic power output is significantly affected by meteorological factors, exhibiting intermittency, volatility, and randomness, posing challenges to power system operation and market transactions. Therefore, high-precision power forecasting is of great significance for grid absorption, dispatch optimization, and operational safety.

[0004] Currently, photovoltaic power prediction mainly targets individual power plants, utilizing local historical data and weather forecasts to construct prediction models through physical models, statistical methods, or machine learning algorithms. With the large-scale deployment of photovoltaics across wide areas, the demand for regional prediction is becoming increasingly prominent. Existing regional prediction technologies mainly develop in two directions: first, optimizing single-point prediction models, such as using strategies like ensemble learning, mode decomposition, and residual correction to improve accuracy, where residuals are often used for real-time correction or training data selection; second, introducing dynamic weights and adaptive training mechanisms, adjusting model structure or hyperparameters through error evaluation to enhance model adaptability.

[0005] However, the above-mentioned schemes still have shortcomings in wide-area forecasting applications. Existing methods do not fully explore the regional spatiotemporal patterns contained in the residuals, the residual processing methods are relatively simple, and there is a lack of modeling for the geographical dispersion smoothing effect. Simply summing up the forecast results of each station ignores the spatial asynchronicity of power output in the region and its impact on the total power forecast. In addition, the existing general forecasting framework is not closely integrated with the physical process of macro-meteorological driving regional total power output, and there is a lack of matching hierarchical design. Summary of the Invention

[0006] To address the aforementioned issues, this invention proposes a wide-area photovoltaic power prediction method and system based on spatiotemporal residual feature fusion. This invention constructs spatiotemporal residual features that can reflect the spatial heterogeneity of regional power output and designs a two-level cascaded prediction framework. The second-level model specifically learns how to use these residual features to refine the first-level prediction, thereby achieving deep modeling of the overall regional power output pattern.

[0007] According to some embodiments, the first aspect of the present invention provides a wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion, employing the following technical solution: A wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion includes: Multi-source heterogeneous data from multiple sub-regions within the target area are acquired and preprocessed to obtain a regional historical database. Based on the regional historical database, regional macro-meteorological characteristics are extracted. Based on regional macro-meteorological characteristics, multiple trained first-level prediction models are used to make predictions in parallel and then aggregated to obtain first-level aggregated prediction values ​​and prediction residual sequences. Based on the predicted residual sequence, corresponding spatiotemporal residual features are constructed. The spatiotemporal residual features are then fused with time-aligned regional macro-meteorological features to obtain an enhanced feature set. Based on the enhanced feature set, a trained two-level correction model is used for prediction to obtain the final predicted power correction amount; The sum of the final predicted power correction and the first-level aggregated predicted value is used as the predicted total photovoltaic power for the region.

[0008] Furthermore, multi-source heterogeneous data from multiple sub-regions within the target area are acquired and preprocessed to obtain a regional historical database. Based on this regional historical database, regional macro-meteorological characteristics are extracted, including: Acquire measured meteorological data, remote sensing inversion data, weather forecast data, and historical data of total photovoltaic power generation in multiple sub-regions within the target area, as multi-source heterogeneous data for multiple sub-regions; Preprocessing multi-source heterogeneous data from multiple sub-regions yields a regional-level historical database. Based on the regional historical database, the average surface irradiance, spatial variation coefficient of irradiance, average temperature and humidity, cloud cover evolution index, and historical power statistics of the region are extracted to form a regional feature vector sequence, which serves as the regional macro-meteorological characteristics.

[0009] Furthermore, the average surface irradiance of the region is the arithmetic mean of the surface irradiance of all effective grid points or stations within the target region at a certain moment. The spatial variation coefficient of irradiance is the ratio of the standard deviation of irradiance in the region at that time to the average surface irradiance in the region. The average temperature and humidity of the region refer to the average ambient temperature and average relative humidity of the region.

[0010] Furthermore, the method of using multiple trained first-level prediction models in parallel to make predictions based on regional macro-meteorological characteristics and then aggregating them to obtain first-level aggregated prediction values ​​and prediction residual sequences corresponding to different first-level prediction models includes: Based on regional macro-meteorological characteristics, multiple trained first-level prediction models are used in parallel to make predictions, and the initial predicted values ​​of regional total photovoltaic power corresponding to different first-level prediction models are obtained. The initial predicted values ​​of total photovoltaic power in the region corresponding to different first-level prediction models are aggregated to obtain the first-level aggregated prediction value. The difference between the initial predicted value of the total photovoltaic power in the region corresponding to different first-level prediction models and the corresponding actual measured value of the total photovoltaic power in the region is used as the prediction residual of different first-level prediction models. Based on the prediction residuals of all first-level prediction models, a prediction residual sequence is constructed.

[0011] Furthermore, the spatiotemporal residual features include model difference residual features, regional output dispersion proxy features, and residual temporal statistical features.

[0012] Furthermore, the model variance residual feature is the standard deviation of the prediction residuals of all first-level prediction models at each time point; The regional output dispersion proxy feature is the product of the spatial variation coefficient of irradiance at each time point and the absolute value of the prediction residual of the optimal first-level prediction model at the current time point. The residual time-series statistical features are calculated using a sliding time window to measure the statistical value of the prediction residuals of each first-level prediction model within that window.

[0013] According to some embodiments, a second aspect of the present invention provides a wide-area photovoltaic power prediction system based on spatiotemporal residual feature fusion, employing the following technical solution: A wide-area photovoltaic power prediction system based on spatiotemporal residual feature fusion includes: The data processing module is configured to acquire multi-source heterogeneous data from multiple sub-regions within the target area and preprocess it to obtain a regional historical database, and extract regional macro-meteorological characteristics based on the regional historical database. The first-level prediction model is configured to make predictions in parallel using multiple trained first-level prediction models based on regional macro-meteorological characteristics, and then aggregate them to obtain the first-level aggregated prediction value and the prediction residual sequence. The spatiotemporal residual feature engineering module is configured to construct corresponding spatiotemporal residual features based on the predicted residual sequence, and then fuse the spatiotemporal residual features with time-aligned regional macro-meteorological features to obtain an enhanced feature set. The secondary correction module is configured to use the trained secondary correction model to make predictions based on the enhanced feature set, and obtain the final predicted power correction amount. The prediction fusion module is configured to use the sum of the final predicted power correction and the first-level aggregated prediction as the predicted total photovoltaic power for the region.

[0014] According to some embodiments, a third aspect of the present invention provides a computer-readable storage medium.

[0015] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion as described in the first scheme above.

[0016] According to some embodiments, a fourth aspect of the present invention provides a computer device.

[0017] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion as described in the first embodiment above.

[0018] According to some embodiments, a fifth aspect of the present invention provides a computer program product or computer program.

[0019] A computer program product or computer program includes computer instructions stored in a computer-readable storage medium, wherein a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in the wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion as described in the first embodiment above.

[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention constructs a dynamic feedback closed loop of prediction-residual-repreneurial; it introduces a dynamic error correction mechanism to achieve adaptive improvement in prediction accuracy; by using the residuals from the first round of prediction as key new features in the second round of training, the model can actively identify and correct its systematic prediction bias under specific weather patterns. This residual feedback mechanism endows the model with powerful self-learning and adaptive capabilities, especially in effectively dealing with sudden weather changes, thereby continuously and dynamically improving prediction accuracy, which is not available in existing technologies. Attached Figure Description

[0021] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0022] Figure 1 This is a diagram of the overall processing architecture of the wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion in this embodiment of the invention. Detailed Implementation

[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0024] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0025] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0026] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0027] Example 1 This embodiment provides a wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion. This embodiment uses the application of this method to a server as an example for illustration. It is understood that this method can also be applied to terminals, and can also be applied to systems including terminals, servers, and other components, and can be implemented through interaction between the terminal and the server. The server can be an independent physical server, a server cluster composed of multiple physical servers, or a distributed system. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network servers, cloud communication, middleware services, domain name services, CDN security services, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein. In this embodiment, the method includes the following steps: Multi-source heterogeneous data from multiple sub-regions within the target area are acquired and preprocessed to obtain a regional historical database. Based on the regional historical database, regional macro-meteorological characteristics are extracted. Based on regional macro-meteorological characteristics, multiple trained first-level prediction models are used to make predictions in parallel and then aggregated to obtain first-level aggregated prediction values ​​and prediction residual sequences. Based on the predicted residual sequence, corresponding spatiotemporal residual features are constructed. The spatiotemporal residual features are then fused with time-aligned regional macro-meteorological features to obtain an enhanced feature set. Based on the enhanced feature set, a trained two-level correction model is used for prediction to obtain the final predicted power correction amount; The sum of the final predicted power correction and the first-level aggregated predicted value is used as the predicted total photovoltaic power for the region.

[0028] like Figure 1 As shown, the overall processing procedure of the method described in this embodiment is as follows: Step S1: Obtain multi-source heterogeneous data from multiple sub-regions within the target area and preprocess them to obtain a regional historical database. Extract regional macro-meteorological characteristics based on the regional historical database. Step S1.1: Obtain multi-source heterogeneous data from multiple sub-regions within the target area and perform preprocessing; Step S1.1.1: Obtain measured meteorological data, remote sensing inversion data, weather forecast data, and historical data of total photovoltaic power generation in the target area for multiple sub-regions, as multi-source heterogeneous data for multiple sub-regions; The raw data can be accessed from at least one of the following data sources, and it is not mandatory to obtain individual data for every photovoltaic power station in the region.

[0029] (1) Ground meteorological station network data. Access the measured data of several (e.g., dozens to hundreds) meteorological monitoring stations that are relatively evenly distributed in the target area. The elements include, but are not limited to: total surface irradiance (GHI), ambient temperature, relative humidity, and wind speed.

[0030] (2) Satellite remote sensing inversion data. Access satellite remote sensing products covering the target area to obtain high spatiotemporal resolution cloud images, surface solar irradiance distribution and other data.

[0031] (3) Numerical weather forecast grid data. Access the numerical weather forecast (NWP) products released by the meteorological department to obtain forecast data of various meteorological elements on a regular geographic grid for future periods.

[0032] (4) Total photovoltaic power generation data in the region. Obtain historical data on the total grid-connected power of photovoltaic power plants in the entire target area, synchronized with the above meteorological data, from the regional power grid dispatch center or energy management platform.

[0033] Step S1.1.2: Preprocess the multi-source heterogeneous data from multiple sub-regions to obtain a regional historical database; Preprocessing refers to sequentially cleaning and imputing, standardizing and normalizing, and spatiotemporally gridding and aggregating multi-source heterogeneous data from multiple sub-regions to obtain a regional historical database. In other words, it involves preprocessing the first four types of raw data. Preprocessing results in higher data quality while maintaining the same data type. It is understood that the preprocessing process utilizes existing methods, including but not limited to the following steps: Cleaning and imputation: Detecting and processing missing values ​​and obvious outliers in multi-source heterogeneous data from multiple sub-regions. For spatial data, spatial imputation is performed using methods such as Kriging interpolation or inverse distance weighted interpolation; for time series data, imputation is performed using time series interpolation or methods based on the correlation of neighboring sites.

[0034] Standardization and normalization: Standardize meteorological data of different dimensions in multi-source heterogeneous data of multiple sub-regions (such as Z-score standardization), and normalize photovoltaic power to the installed capacity of the power station or the historical maximum value to eliminate the influence of dimensions.

[0035] Spatiotemporal gridding and aggregation: Data from different sources and spatial locations in multi-source heterogeneous data across multiple sub-regions are uniformly mapped onto a pre-defined regular regional grid using interpolation or statistical methods, or regional statistics are directly calculated. This ensures that all feature data and the total regional power data are strictly aligned in timestamps, forming a unified time-series sample.

[0036] Step S1.1.3: Based on the regional historical database, extract the regional average surface irradiance, irradiance spatial variation coefficient, regional average temperature and humidity, regional cloud cover evolution index, and regional historical power statistical characteristics to form a regional feature vector sequence as regional macro-meteorological characteristics. (1) Regional average surface irradiance. Calculate the arithmetic mean of the total surface irradiance of all effective grid points or stations within the target area at a given time, characterizing the overall solar radiation energy level received by the region. Assume that data was collected... Each region, then time Average surface irradiance of each region The calculation formula is:

[0037] in, It is the first Total surface irradiance of the region.

[0038] (2) Spatial variation coefficient of irradiance. Calculate the ratio of the standard deviation of irradiance in the region to the average surface irradiance of the region at that time. This feature is one of the key innovations, as it directly quantifies the spatial non-uniformity of illumination conditions in the region and is one of the fundamental physical factors that avoids errors from simple summation. Spatial variation coefficient of irradiance at time The calculation formula is:

[0039] in, express time The standard deviation of regional irradiance.

[0040] (3) Regional average temperature and humidity. Calculate the regional average ambient temperature using the same irradiance. and average relative humidity That is, the arithmetic mean of the ambient temperature and humidity of all effective grid points or stations in the region at a certain moment under the same irradiance.

[0041] (4) Regional cloud cover evolution index. Based on satellite cloud images or irradiance data, the proportion of a region covered by clouds is calculated, and a comprehensive index is constructed by combining its changing trends over time, such as movement speed and increase / decrease trends.

[0042] (5) Regional historical power statistics. Extract trend terms and periodic terms from historical power data, including daily periodicity and seasonal periodicity, as well as statistical features such as mean and variance within the sliding window.

[0043] The final output is a sequence of regional feature vectors, representing the regional macro-meteorological features. Each time point's feature vector contains multiple extracted regional macro-meteorological features and corresponds to the actual total regional photovoltaic power generation at that time point. For example: [time : , , , Historical power statistics characteristics, ..., actual value of total regional power ].

[0044] This embodiment does not require real-time individual data from every photovoltaic power station in the region. Instead, it builds a robust prediction model based on more readily available regional meteorological data and historical total power output data. This greatly reduces the requirements for data infrastructure and the dependence on the full amount of underlying data. It enables high-precision wide-area prediction even when a large amount of distributed power station data is unavailable or incomplete. This solves the data bottleneck problem faced by existing methods, has stronger engineering practicality and promotion value, and improves the feasibility and robustness of the method.

[0045] Step S2: Based on the regional macro-meteorological characteristics, multiple trained first-level prediction models are used to make predictions in parallel and then aggregated to obtain the first-level aggregated prediction value and the prediction residual sequence. The training process for the Level 1 prediction model is as follows: Step S2.1: Based on the multi-source nature of regional macro-meteorological characteristics data, obtain multiple basic prediction models and construct a basic prediction model pool; The basic prediction model pool includes at least machine learning algorithm models such as support vector machine, random forest, gradient boosting decision tree, and multilayer perceptron. This step employs a heterogeneous model parallel ensemble strategy, simultaneously deploying at least four core machine learning algorithm models, including Support Vector Machine (SVM), Random Forest (RF), Gradient Boosting Decision Tree (XGBoost), and Multilayer Perceptron (MLP), forming a complementary pool of basic prediction models. The roles of each model are as follows: Support Vector Machines (SVMs) focus on finding the optimal classification / regression hyperplane in a high-dimensional feature space. They have a certain ability to capture complex nonlinear relationships between features and are particularly robust on small to medium-sized datasets.

[0046] Random forests, by constructing a large number of decision trees and integrating their results, can effectively learn complex nonlinear relationships and, thanks to their out-of-bag estimation and random feature selection mechanisms, naturally possess strong resistance to overfitting.

[0047] XGBoost uses a forward-step additive model to iteratively correct the residuals of previous models, demonstrating excellent performance in handling structured data and feature interactions, and achieving high prediction accuracy.

[0048] As a classic feedforward neural network, the multilayer perceptron can approximate any complex continuous function through multilayer nonlinear transformations and is adept at automatically learning deep abstract feature representations from data.

[0049] Step S2.2: Using the regional macro-meteorological characteristics as the training dataset, train each basic prediction model in the basic prediction model pool independently in parallel to obtain different trained first-level prediction models and the corresponding initial prediction values ​​of the regional total photovoltaic power. Using regional macro-meteorological characteristics as the training dataset, the training dataset is divided into a training set, a validation set, and a test set in chronological order, with a ratio of 7:1:2. Each base model is independently tuned for hyperparameters based on its algorithm characteristics and trained on the training set.

[0050] After training, each basic prediction model performs forward inference on the same training and test sets to generate the corresponding initial predicted values ​​for the total photovoltaic power in the region. Assume there are a total of... training samples, For each test sample, the following prediction result set is generated:

[0051]

[0052]

[0053]

[0054] in, These represent the predicted total photovoltaic power obtained during the training and testing phases of the support vector product model, respectively. These represent the predicted total photovoltaic power obtained during the training and testing phases of the random forest model, respectively. These represent the predicted total photovoltaic power obtained during the training and testing phases of the gradient boosting decision tree model, respectively. These represent the predicted total photovoltaic power obtained during the training and testing phases of the multilayer perceptron model, respectively.

[0055] Step S2.3: Aggregate the initial predicted values ​​of the total photovoltaic power in the region corresponding to different first-level prediction models to obtain the first-level aggregated prediction value, and obtain the prediction residual sequence based on the difference between the initial predicted values ​​of the total photovoltaic power in the region corresponding to different first-level prediction models and the corresponding actual measured values ​​of the total photovoltaic power in the region. To provide a stable baseline forecast, initial forecasts of total photovoltaic power for each region were calculated. The weighted average or the initial predicted total photovoltaic power of the single region with the best performance is used as the first-level aggregated prediction value. .

[0056]

[0057] in, The model with the smallest MAE (mean absolute error) value is selected from SVM, RF, XGB, and MLP.

[0058] Simultaneously, the prediction residuals of each trained first-level prediction model are output. The prediction residuals are calculated by comparing the initial predicted total photovoltaic power of the region with the actual measured total photovoltaic power of the region, as determined by the corresponding first-level prediction model. Subtracting the two gives the following result:

[0059]

[0060]

[0061]

[0062] in, This represents the measured total photovoltaic power in the actual area. The initial predicted value of the total photovoltaic power in the region is obtained from the support vector product model. For the support vector product model residuals; This represents the initial predicted value of the total photovoltaic power in the region obtained from the random forest model. For the residuals of the random forest model; To obtain the initial predicted value of the total photovoltaic power in the light-emitting region from the gradient boosting decision tree model, To improve the residuals of the decision tree model using gradient boosting; This is the initial predicted value of the total photovoltaic power in the region obtained from the multilayer perceptron model. The residuals are from the multilayer perceptron model.

[0063] Based on the prediction residuals of all first-level prediction models, a prediction residual sequence is constructed. These data not only quantify the prediction bias of different primary prediction models on historical data, but also imply the cognitive differences and uncertainties of primary prediction models when facing weather patterns in specific regions. They are key information sources that drive the refinement of secondary correction models.

[0064] Step S3: Construct corresponding spatiotemporal residual features based on the predicted residual sequence, and fuse the spatiotemporal residual features with time-aligned regional macro-meteorological features to obtain an enhanced feature set; Step S3.1: Align the regional macro-meteorological characteristics with the predicted residual sequence according to time to obtain the time-aligned regional macro-meteorological characteristics.

[0065] Step S3.2: Calculate the model difference residual characteristics, regional output dispersion surrogate characteristics, and residual time series statistical characteristics of the predicted residual sequence to form the corresponding spatiotemporal residual characteristics; By performing in-depth processing on different prediction residuals, the following three types of spatiotemporal residual features with clear physical or statistical significance are constructed.

[0066] (1) Model consensus characteristics - Model difference residual characteristics Model difference residual characteristics: Calculate the standard deviation or range of the prediction residuals of different first-level prediction models at the same time to characterize the uncertainty of the prediction of the current weather pattern among the first-level prediction models.

[0067] This feature is used to quantify the inconsistency between predictions from different first-level forecasting models under specific weather conditions. This inconsistency often indicates a higher forecasting difficulty at that moment or the existence of complex patterns that are not fully understood. At each time point... Calculate the standard deviation of the prediction residuals for all first-level prediction models, using this as the model variance residual characteristic. The calculation formula is as follows:

[0068] in, This indicates calculating the standard deviation of the data within the parentheses. express The standard deviation of the predicted residual sequence at each time step. The larger the value, the greater the discrepancy between models, which may correspond to complex meteorological scenarios such as rapid changes in irradiance and moving edges of cloud clusters, and is a weak link in the first-level prediction.

[0069] (2) Regional output dispersion proxy characteristics Several spatially uniform meteorological stations within the region were selected, and their irradiance spatiotemporal variation coefficients were calculated. The correlation analysis between the irradiance spatiotemporal variation coefficients and the prediction residuals of the first-level prediction model was conducted to construct a synthetic feature that can indirectly reflect the simple summation error caused by uneven spatial output.

[0070] This feature is key to solving the simple summation bias problem; it aims to establish the correlation between regional spatial non-uniformity of illumination and first-order prediction residuals, enabling the model to explicitly learn geographic dispersion effects. (Time points...) spatial variation coefficient of irradiance , and at that moment The absolute values ​​of the prediction residuals from the best-performing Level 1 prediction models are combined to construct an interactive feature, which serves as a surrogate feature for regional output dispersion. The calculation formula is as follows:

[0071] in, To find the first-level prediction model with the minimum mean absolute error on the validation set The prediction residual at any given time. This feature directly amplifies the signal on samples with uneven spatial illumination and biased predictions by the first-level prediction model, guiding the second-level correction model to focus on and correct these typical errors caused by spatial smoothing effects.

[0072] (3) Statistical characteristics of residual patterns - statistical characteristics of residual time series Residual time-series statistical features extract sliding window statistics, such as mean and variance, from each predicted residual to describe the short-term fluctuation patterns of the predicted bias sequence. These features capture the short-term evolution patterns of the residuals over time, helping the secondary correction model identify persistent bias trends or abrupt changes. For each primary prediction model, the predicted residuals are used to calculate the statistics within a sliding time window, generating residual time-series statistical features. For example: The mean of the predicted residuals within the window, reflecting the direction of the recent average deviation.

[0073] The standard deviation of the predicted residuals within the window reflects the degree of fluctuation in recent bias.

[0074] The slope of the linear fit to the predicted residuals within the window reflects whether the bias is widening or converging.

[0075] Step S3.3: Fuse the spatiotemporal residual features with the time-aligned regional macro-meteorological features to obtain an enhanced feature set; This enhanced feature set not only includes the original meteorological factors driving photovoltaic output, but also adds in-depth information about the spatiotemporal conditions under which historical predictions are prone to errors and what the error patterns are.

[0076] This step aims to extract and construct high-order features from the prediction residuals of multiple primary prediction models to characterize the spatiotemporal patterns of regional photovoltaic power output response errors. These features will serve as navigation signals for the secondary correction model, enabling it to specifically learn and correct systematic prediction biases caused by complex factors such as regional geographical dispersion and abrupt weather changes.

[0077] Step S4: Based on the enhanced feature set, use the trained secondary correction model to make predictions and obtain the final predicted power correction amount; The training process for the second-level correction model is as follows: Construct a two-stage power correction model, including but not limited to LightGBM, LSTM, Transformer, and other models capable of handling complex feature interactions. The training objective of this model is not to directly predict the total power, but rather to predict the optimal power correction amount. During training, power correction amount The error can be determined in reverse by using grid search or optimization algorithms to minimize the final prediction error, or it can be simplified to using the negative value of the residual of the best-performing base model as an approximation.

[0078] An enhanced feature set containing rich spatiotemporal error pattern information is used to train a specialized machine learning model as a secondary correction model. This model does not directly predict the total photovoltaic power, but instead learns to predict an optimal power correction amount. This allows for precise and adaptive compensation for systematic biases in the primary prediction model.

[0079] Step S4.1: Use the enhanced feature set as training data.

[0080] The training data integrates original regional macro-meteorological characteristics with innovative spatiotemporal residual features. The training objective (label) of the second-level correction model is not the actual measured value of the total regional photovoltaic power, but rather the power correction amount. . Defined as the value that needs to be added to the first-level aggregated prediction value to achieve the optimal final prediction. That is, ideally, it satisfies:

[0081] in, It is a first-level aggregated prediction value. It is the actual measured value of the total photovoltaic power in the region.

[0082] Step S4.2: Determine the optimal power correction amount based on the predicted residual sequence or the first-level aggregated prediction value; To train the second-level correction model, a learning model is generated for each sample in the training set. This invention provides the following two feasible generation strategies: Strategy A (based on optimal basis model residuals): Based on the prediction residual sequence, the negative value of the prediction residual corresponding to the first-level prediction model with the smallest mean absolute error or root mean square error is used as the optimal power correction. ; The first-level prediction model with the smallest mean absolute error (MAE) or root mean square error (RMSE) on the first-level prediction model validation set is selected, and the negative value of its prediction residual on the training set is used as the optimal power correction. .Right now: This method is intuitive and effective, implicitly assuming that the residual pattern of the optimal basis model is the main component of the system bias.

[0083] Strategy B (Optimization Search): Iteratively search for the power correction amount until the overall error between the actual measured value of the total photovoltaic power in the region and the sum of the first-level aggregated prediction value and the power correction amount is minimized, thus obtaining the optimal power correction amount. ; With the goal of maximizing the final prediction accuracy, As an optimizable variable, the first-level aggregated predictions are fixed on a retained validation set. Find a set of optimal power correction values ​​using methods such as grid search or gradient descent. Make and The overall error (e.g., MAE) is minimized by this method. Theoretically, it is better, but the computational cost is slightly higher.

[0084] Step S4.3: Using the optimal power correction amount as the training objective, the enhanced feature set as the training data, and minimizing the error between the predicted power correction amount and the optimal power correction amount as the loss function, train the second-level correction model to obtain the trained second-level correction model. The secondary correction model needs to possess strong feature interaction learning capabilities and nonlinear fitting capabilities. A classic deep neural network structure is chosen. Such models excel at uncovering complex patterns from high-dimensional, heterogeneous sets of enhanced features.

[0085] Model training: using the generated {enhanced feature set, Data pairs to minimize predicted power correction. With optimal power correction The error between the two values ​​is used as the loss function to train the secondary correction model. During training, an independent validation set and strategies such as early stopping are required to prevent overfitting.

[0086] The final output is a trained and applicable second-level correction model. This model has learned how to determine and predict the extent to which the first-level base forecast needs to be corrected based on current macro-meteorological conditions and the spatiotemporal error characteristics derived from historical forecast residuals.

[0087] Step S5: The sum of the final predicted power correction and the first-level aggregated predicted value is used as the predicted value of the total photovoltaic power in the region. The calculation formula is:

[0088] in, yes The predicted total photovoltaic power in the region at a given time. yes The first-level aggregated prediction value at time 1. yes The final predicted power correction at time 1.

[0089] Output in time series The forecasts are formatted and packaged to generate standardized results. These results can be pushed in real time to the power grid dispatching system, power trading platform, or renewable energy management system via application programming interfaces (APIs), message queues, or files, providing crucial decision-making support for power balance, dispatching plans, and market transactions.

[0090] This embodiment constructs a prediction model directly at the regional level. By establishing a weather feature database covering the entire region and using wide-area meteorological information as direct input, the model can autonomously learn the spatiotemporal mapping relationship between the macro-weather system and the total regional power output, fundamentally avoiding the simple aggregation bias caused by geographical dispersion. It comprehensively utilizes multiple machine learning models, including Support Vector Machine (SVM), Random Forest (RF), Multivariate Perceptron (MLP), and Logistic Regression (LR), employing a heterogeneous model fusion strategy to enhance the model's generalization ability. This ensemble learning approach can capture the complex nonlinear relationship between weather features and photovoltaic power output, and through the complementary advantages of different models, forms a more robust and generalized regional prediction model, effectively avoiding overfitting or performance bottlenecks that may exist with a single model.

[0091] Example 2 This embodiment provides a wide-area photovoltaic power prediction system based on spatiotemporal residual feature fusion, including: The data processing module is configured to acquire multi-source heterogeneous data from multiple sub-regions within the target area and preprocess it to obtain a regional historical database, and extract regional macro-meteorological characteristics based on the regional historical database. The first-level prediction model is configured to make predictions in parallel using multiple trained first-level prediction models based on regional macro-meteorological characteristics, and then aggregate them to obtain the first-level aggregated prediction value and the prediction residual sequence. The spatiotemporal residual feature engineering module is configured to construct corresponding spatiotemporal residual features based on the predicted residual sequence, and then fuse the spatiotemporal residual features with time-aligned regional macro-meteorological features to obtain an enhanced feature set. The secondary correction module is configured to use the trained secondary correction model to make predictions based on the enhanced feature set, and obtain the final predicted power correction amount. The prediction fusion module is configured to use the sum of the final predicted power correction and the first-level aggregated prediction as the predicted total photovoltaic power for the region.

[0092] The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0093] The descriptions of each embodiment in the above embodiments have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0094] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and the division of modules described above is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.

[0095] Example 3 This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion as described in Embodiment 1 above.

[0096] Example 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion as described in Embodiment 1 above.

[0097] Example 5 This embodiment provides a computer program product or computer program, including computer instructions stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in the wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion described in Embodiment 1 above.

[0098] Those skilled in the art will understand that embodiments of the present invention can provide methods, systems, or computer program products. Therefore, the present invention can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0099] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0100] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0101] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0102] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0103] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion, characterized in that, include: Multi-source heterogeneous data from multiple sub-regions within the target area are acquired and preprocessed to obtain a regional historical database. Based on the regional historical database, regional macro-meteorological characteristics are extracted. Based on regional macro-meteorological characteristics, multiple trained first-level prediction models are used to make predictions in parallel and then aggregated to obtain first-level aggregated prediction values ​​and prediction residual sequences. Based on the predicted residual sequence, corresponding spatiotemporal residual features are constructed. The spatiotemporal residual features are then fused with time-aligned regional macro-meteorological features to obtain an enhanced feature set. Based on the enhanced feature set, a trained two-level correction model is used for prediction to obtain the final predicted power correction amount; The sum of the final predicted power correction and the first-level aggregated predicted value is used as the predicted total photovoltaic power for the region.

2. The wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion as described in claim 1, characterized in that, Multi-source heterogeneous data from multiple sub-regions within the target area are acquired and preprocessed to obtain a regional historical database. Based on this regional historical database, regional macro-meteorological characteristics are extracted, including: Acquire measured meteorological data, remote sensing inversion data, weather forecast data, and historical data of total photovoltaic power generation in multiple sub-regions within the target area, as multi-source heterogeneous data for multiple sub-regions; Preprocessing multi-source heterogeneous data from multiple sub-regions yields a regional-level historical database. Based on the regional historical database, the regional average surface irradiance, irradiance spatial variation coefficient, regional average temperature and humidity, regional cloud cover evolution index, and regional historical power statistical characteristics are extracted to form a regional feature vector sequence, which serves as the regional macro-meteorological characteristics.

3. The wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion as described in claim 2, characterized in that, The average surface irradiance of the region is the arithmetic mean of the surface irradiance of all effective grid points or stations within the target region at a certain moment. The spatial variation coefficient of irradiance is the ratio of the standard deviation of irradiance in the region at that time to the average surface irradiance in the region. The average temperature and humidity of the region refer to the average ambient temperature and average relative humidity of the region.

4. The wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion as described in claim 1, characterized in that, The method, based on regional macro-meteorological characteristics, utilizes multiple trained first-level prediction models to perform predictions in parallel and then aggregates them to obtain first-level aggregated prediction values ​​and prediction residual sequences corresponding to different first-level prediction models, including: Based on regional macro-meteorological characteristics, multiple trained first-level prediction models are used in parallel to make predictions, and the initial predicted values ​​of regional total photovoltaic power corresponding to different first-level prediction models are obtained. The initial predicted values ​​of total photovoltaic power in the region corresponding to different first-level prediction models are aggregated to obtain the first-level aggregated prediction value. The difference between the initial predicted value of the total photovoltaic power in the region corresponding to different first-level prediction models and the corresponding actual measured value of the total photovoltaic power in the region is used as the prediction residual of different first-level prediction models. Based on the prediction residuals of all first-level prediction models, a prediction residual sequence is constructed.

5. The wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion as described in claim 1, characterized in that, The spatiotemporal residual features include model difference residual features, regional output dispersion proxy features, and residual time series statistical features.

6. The wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion as described in claim 5, characterized in that, The model variance residual feature is the standard deviation of the prediction residuals of all first-level prediction models at each time point; The regional output dispersion proxy feature is the product of the spatial variation coefficient of irradiance at each time point and the absolute value of the prediction residual of the optimal first-level prediction model at the current time point. The residual time-series statistical features are calculated using a sliding time window to measure the statistical value of the prediction residuals of each first-level prediction model within that window.

7. A wide-area photovoltaic power prediction system based on spatiotemporal residual feature fusion, characterized in that, include: The data processing module is configured to acquire multi-source heterogeneous data from multiple sub-regions within the target area and preprocess it to obtain a regional historical database, and extract regional macro-meteorological characteristics based on the regional historical database. The first-level prediction model is configured to use multiple trained first-level prediction models to make predictions in parallel based on regional macro-meteorological characteristics and then aggregate them to obtain the first-level aggregated prediction value and the prediction residual sequence. The spatiotemporal residual feature engineering module is configured to construct corresponding spatiotemporal residual features based on the predicted residual sequence, and then fuse the spatiotemporal residual features with time-aligned regional macro-meteorological features to obtain an enhanced feature set. The secondary correction module is configured to use the trained secondary correction model to make predictions based on the enhanced feature set, and obtain the final predicted power correction amount. The prediction fusion module is configured to use the sum of the final predicted power correction and the first-level aggregated prediction as the predicted total photovoltaic power for the region.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion as described in any one of claims 1-6.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion as described in any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps in the wide-area photovoltaic power prediction method based on spatiotemporal residual feature fusion as described in any one of claims 1-6.