Forecasting method of solar radiation based on satellite data and multi-time and space scale super-resolution model

Through the solar radiation forecasting method based on satellite data and multi-time and space scale super-resolution models, the problems of spatiotemporal dependence separation and insufficient spatial resolution in regional solar radiation forecasting are solved, and efficient and accurate 0-4 hour solar radiation forecasting is achieved, thereby improving the stability and prediction accuracy of photovoltaic power generation.

CN120635598BActive Publication Date: 2025-10-21WUHAN INST OF TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511119662.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-10-21
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve high-precision regional solar radiation forecasts for the next 0-4 hours. The fragmentation of temporal and spatial dependencies, poor adaptability due to small samples, and insufficient spatial resolution affect photovoltaic power generation efficiency and grid-connected stability.

Method used

A solar radiation forecasting method based on satellite data and a multi-spatiotemporal scale super-resolution model is adopted. By constructing a multi-spatiotemporal scale super-resolution model and combining the GPT-2 model with the ResNet residual super-resolution module, the fusion of spatiotemporal information and high-resolution reconstruction are achieved. Small sample data is used for efficient training to dynamically adapt to changes in meteorological conditions.

Benefits of technology

It achieves high-precision regional solar radiation forecasts for the next 0-4 hours, improves the reliability of photovoltaic power forecasts and image generation quality, reduces computing resource requirements, and improves the model's adaptability and forecast accuracy under complex meteorological conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635598B_ABST
    Figure CN120635598B_ABST
Patent Text Reader

Abstract

A solar radiation prediction method based on satellite data and multi-time and space scale super-resolution model, comprising: acquiring satellite remote sensing data, and cutting the satellite image into the size of the prediction area; constructing a daily variation feature set and a seasonal variation feature set, reorganizing the data into a four-dimensional tensor, and dividing it into a training set, a validation set and a test set; linearly mapping the daily and monthly period features through a learnable parameter matrix to generate a comprehensive time embedding; adopting a hierarchical parameter control strategy, freezing the model parameters, and unfreezing the parameters; generating a low-resolution preliminary prediction result through a regression convolution layer; using a residual super-resolution module for upsampling, combining a skip connection and a residual scaling to reconstruct a high-resolution image, outputting single-channel solar radiation intensity data and denormalizing it to physical dimensions; and configuring an optimizer. It can realize future 0-4 hour regional high-precision solar radiation prediction, and solve the problems of time and space dependence fragmentation, poor sample adaptability and insufficient spatial resolution in solar radiation prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a regional solar radiation forecasting method, in particular to a solar radiation forecasting method based on satellite data and a multi-time and space scale super-resolution model. Background Art

[0002] The share of photovoltaic power generation has increased significantly in recent years. The accuracy of short-term solar radiation forecasts (0-4 hours) directly impacts photovoltaic power generation efficiency and grid connection stability. However, due to complex weather conditions and cloud cover fluctuations, the accuracy of short-term solar radiation forecasts has been relatively low, limiting the reliability of ultra-short-term photovoltaic power forecasts. With the large-scale integration of distributed photovoltaic power into the grid and the improvement of time-of-use trading mechanisms in the electricity spot market, hourly solar radiation and power forecasts are gaining increasing attention.

[0003] Currently, solar radiation forecasts are mainly derived from mesoscale numerical weather predictions (NWPs). However, because solar radiation has typical daily periodic variations and is significantly affected by different weather conditions, numerical weather forecasts often have difficulty accurately capturing hourly solar radiation variation characteristics. To address this issue, some studies have used ground-based solar radiation observations and related meteorological data to conduct research on the correction of solar radiation output from numerical weather forecasts. For example, some have used the WRF-SOLAR model to forecast solar radiation at eight stations in a certain province, and then used phase differences to correct the forecast results. Other studies have used satellite data to correct solar radiation, such as using the WRF model's total solar radiation forecast product, geostationary satellite total cloud cover, and ground-based meteorological data to correct errors in solar radiation forecast results.

[0004] Mesoscale numerical weather prediction (NWP) also has limitations, including low spatial resolution and high forecasting costs for distributed photovoltaic projects. Satellite observation data can capture the continuously changing characteristics of cloud parameters and solar radiation. Previous studies have used various geostationary satellites to conduct solar radiation forecasts. However, this has primarily focused on single-point predictions. Image-based forecasts typically use statistical extrapolation to predict cloud parameters and then use radiative transfer models to invert solar radiation. These methods can predict cloud movement and are effective under clear sky conditions. However, the difficulty in capturing the development and dissipation of cloud parameters, coupled with the influence of terrain effects, can increase forecast errors under complex weather conditions.

[0005] In recent years, some studies have combined artificial intelligence technology to predict global solar radiation on the horizontal surface. Research on regional solar radiation prediction mainly includes three types of methods.

[0006] One approach is based on early-generation AI models, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs). These models use CNNs to extract spatial features and RNNs to model time series dependencies. Prior art has proposed a solar radiation prediction framework based on LSTM, but its spatial feature extraction relies on fixed-scale convolution kernels, making it difficult to adapt to multi-scale meteorological phenomena (such as sudden changes in localized cloud clusters). Furthermore, the serial computational nature of RNNs results in inefficient modeling of long-term dependencies and insufficient representation of dynamic spatial correlations (such as inter-regional radiation transfer). Furthermore, the Unified Network (UNet), with its encoder-decoder architecture and skip connections, has also been widely used in image segmentation and prediction. UNet variants have also been proposed for solar radiation prediction, enhancing spatial detail modeling capabilities through multi-scale feature fusion. However, UNet's core design targets static image processing, making it weak in modeling continuous temporal evolution. It also lacks explicit mechanisms for spatial and temporal dependency interactions, making it difficult to capture the cross-temporal propagation of radiation intensity. While these early-generation AI models consume relatively few resources during training, their predictions generally have room for improvement, and the generated images are often blurry.

[0007] The second approach is based on the Spatiotemporal Graph Neural Network (STGCN). STGCN uses a graph structure to model spatial relationships between regions and combines it with temporal convolution to capture temporal patterns. However, this approach relies on a predefined static adjacency matrix and cannot dynamically adapt to changing meteorological conditions (such as the association of non-adjacent regions caused by cloud movement). Existing techniques have attempted to incorporate attention mechanisms to optimize graph structures, but this significantly increases model complexity and makes it prone to overfitting in low-sample scenarios.

[0008] The third approach is based on a generative adversarial network model. Because the first two methods typically fail to guarantee both accuracy and quality in solar radiation image prediction, a modified adversarial neural network model (DGMR) has been applied to solar radiation prediction for the first time. While this method can improve image accuracy and quality, it is difficult to train the model and requires significant computing resources.

[0009] The limitations of existing technologies are as follows: 1. Spatiotemporal dependency separation: Most models process spatial and temporal features independently, failing to achieve true spatiotemporal joint modeling. 2. Poor adaptability to small sample sizes: Spatiotemporal graph neural network optimization strategies increase model complexity and are prone to overfitting, while traditional CNN / RNN models rely on large amounts of labeled data. 3. Limited spatial resolution: Methods based on fixed convolution kernels or static graph structures struggle to improve local region prediction accuracy, and image prediction results are too blurry to be practically applied. Furthermore, the application of super-resolution technology has yet to be fully explored.

[0010] Related patent document: CN114492204A discloses a solar radiation forecasting method, device, equipment and medium, which obtain the total solar radiation time series and meteorological data corresponding to the observation site from the preset historical moment to the current moment; decompose the total solar radiation time series into nonlinear mode and linear mode through the ensemble empirical mode decomposition method; input the linear mode into the linear mode prediction model to predict solar radiation, and output the linear mode value of the total solar radiation at the future moment; input the nonlinear mode into the nonlinear mode prediction model to predict solar radiation, and output the nonlinear mode value of the total solar radiation at the future moment; sum the linear mode value and the linear mode value of the total solar radiation at the future moment to obtain the total solar radiation at the future moment, which improves the technical problem that the existing technology uses artificial neural network for solar energy forecasting, the fitting effect of the linear change part of solar energy is not ideal, resulting in low forecast accuracy.

[0011] The above technologies have failed to achieve high-precision regional solar radiation forecasts for the next 0-4 hours, and have not solved the problems of spatiotemporal dependency fragmentation, poor adaptability to small samples, and insufficient spatial resolution in solar radiation forecasts. Summary of the Invention

[0012] The purpose of the present invention is to provide a solar radiation forecasting method based on satellite data and a multi-time and space scale super-resolution model, which can achieve high-precision regional solar radiation forecasts for the next 0-4 hours, so as to solve the problems of spatiotemporal dependency separation, poor adaptability due to small samples, and insufficient spatial resolution in solar radiation forecasting.

[0013] In order to solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0014] A solar radiation forecasting method based on satellite data and a multi-time and space-scale super-resolution large-scale model (or a regional short-term solar radiation forecasting method based on satellite data and a multi-time and space-scale super-resolution large-scale model) includes the following steps:

[0015] S1: Acquire satellite remote sensing data, parse the time tag and convert it to Beijing time, extract the year, month, day, hour, minute, and second information; filter the valid data from sunrise to sunset and remove the invalid data at night; randomly sample the valid data by month as training samples, and crop the satellite image to the prediction area size;

[0016] S2: Construct a feature set of daily and seasonal variations of solar radiation. Use time normalization to map timestamps to the zero-one interval and use one-hot encoding to extract month information. Reorganize the data into a four-dimensional tensor [number of samples, time steps, number of features, number of channels], divide it into training, validation, and test sets, and perform data dimensionality reduction and standardization.

[0017] S3: Linearly map the diurnal and monthly cycle features of solar radiation through a learnable parameter matrix to generate a comprehensive time embedding; align the spatial information of satellite images with the time embedding, and use convolutional layers to map the spatiotemporal information of solar radiation to a channel space compatible with the GPT-2 model; concatenate the temporal, spatial, and projected features, generate a fused spatiotemporal feature matrix through compression operations, and convert it into the standard input format for large language models;

[0018] S4: Using a layered parameter control strategy, freeze the parameters of the first four layers of the GPT-2 model (the total number of layers can be 12), and unfreeze the parameters of the last two layers to learn the spatiotemporal correlation of solar radiation, forming a forecast model (large model) that integrates the spatiotemporal characteristics of solar radiation; generate low-resolution preliminary forecast results through the regression convolution layer;

[0019] S5: Use the ResNet residual super-resolution module for upsampling, combine skip connections and residual scaling to reconstruct a high-resolution image, output 0-4 hours single-channel solar radiation intensity data and denormalize it to physical dimensions;

[0020] S6: Configure the model's optimizer parameters, using a cosine annealing strategy to dynamically adjust the learning rate; set weight decay and gradient clipping with mean absolute error (MAE) as the loss function, monitor validation set metrics, and save the optimal model parameters through a dynamic early stopping mechanism.

[0021] In the above technical solutions, a preferred technical solution may be that the satellite remote sensing data (satellite data) is the surface downlink shortwave radiation product of the Fengyun-4 meteorological satellite. The step S1 specifically includes:

[0022] S1.1: The acquired satellite remote sensing data include the iso-latitude and longitude gridded data of the downlink shortwave solar radiation retrieval product, and its spatial resolution conforms to the standardized output format of the global atmospheric reanalysis product;

[0023] S1.2: Read solar radiation data from satellite data. By parsing the time tag in the file name, extract the UTC time and convert it to Beijing time. At the same time, obtain the year, month, day, hour, minute, and second information. The satellite data time resolution is 15 minutes.

[0024] S1.3: Based on the latitude and longitude of the forecast area and the date and time information, calculate the daily sunrise and sunset times, filter the valid data between sunrise and sunset, and eliminate the invalid data at night;

[0025] S1.4: To increase sample diversity, further compress the dataset, increase training speed, and reduce computing resource pressure, monthly random sampling is used. The data is grouped by month, and four days of data are randomly selected each month as training samples. The selected satellite image paths and corresponding timestamp information are stored in a local .csv file.

[0026] S1.5: Clip the satellite solar radiation images used in model training to the size of the prediction area to form a local spatial matrix to meet the model input requirements;

[0027] S1.6: Flatten the cropped two-dimensional radiation image matrix into a one-dimensional vector representing the number of features of each sample after expansion in the spatial dimension.

[0028] In the above technical solution, a preferred technical solution may also be that step S2 specifically includes:

[0029] S2.1: Use a time series normalization method to construct a diurnal variation feature set;

[0030] S2.2: Construct a seasonal variation feature set and extract the month information (one-hot encoding) in the sample time to capture the seasonal variation pattern;

[0031] S2.3: Finally, the multidimensional features containing diurnal and seasonal variations are embedded into the feature dataset to form a three-dimensional tensor [number of samples, number of features, number of channels] consisting of time features and solar radiation;

[0032] S2.4: Generate ordered time series data with 15-minute intervals. Suppose the input time offset is n , the output time offset is m , that is, using the past n Solar radiation images, predicting the future m solar radiation images, and then judge the temporal continuity of the sample sequence: n + m Among the consecutive samples, only the valid sequences with a time interval of 15 minutes between adjacent samples are retained. Finally, the time step dimension is added and the data dimension is reorganized into a four-dimensional tensor [number of samples, time steps, number of features, number of channels];

[0033] S2.5: Dataset construction and serialized storage. The dataset is divided into training set, validation set, and test set with the following ratios:

[0034] The ratio of training set: validation set: test set is 7:1.5:1.5. The first 70% of the data in the dataset is the training set, the middle 15% of the data in the dataset is the validation set, and the last 15% of the data in the dataset is the test set. The data is processed into the NPZ format, a compressed data storage format dedicated to the NumPy library, and stored in a local directory.

[0035] S2.6: Data dimensionality reduction. Before model training, we first load the NPZ dataset file and use AdaptiveAvgPool2d to reduce the original data size from a 128-pixel × 224-pixel grid to a 32-pixel × 56-pixel grid, preserving the spatial distribution characteristics.

[0036] S2.7: Data standardization: Perform Z-score standardization on the data based on the mean and standard deviation of the training set to ensure uniform input dimensions. The calculation formula is:

[0037] (1), Where x is the radiation value of the training set data, represents the mean radiation value of the training set, std represents the standard deviation of the radiation value of the training set, and z represents the standardized data value.

[0038] In the above technical solution, a preferred technical solution may also be that in step S2.1, a time series normalization method is adopted by truncating the sample timestamp to the current day's reference point and calculating the proportional difference, mapping the time dimension to a [0,1) continuous space, and processing the time series data according to the following steps:

[0039] (a) Baseline time alignment: Align the timestamp index to midnight of the current day to generate a base time series.

[0040] (b) Time difference calculation: Calculate the difference between the sample timestamp and the current day's base timestamp to obtain the length of time that has passed since the sample file was created;

[0041] (c) Proportional mapping: Divide the time difference by the total time unit corresponding to 24 hours and output the normalized time proportion value in the interval [0,1);

[0042] Assume that the sample timestamp is t, then the normalized value Tn satisfies: (2), in, Indicates the zero o'clock of the day corresponding to the sample timestamp. 1 Day corresponds to 24 hours.

[0043] In the above technical solution, a preferred technical solution may also be that the step S3 specifically includes:

[0044] S3.1: The multi-scale spatiotemporal super-resolution model receives historical solar radiation observation data tensors with the dimensions [batch size, time step, number of samples, number of channels], where the number of channels includes three channels: radiation intensity, daily characteristics, and monthly characteristics. A time embedding module is used to linearly map the daily and monthly cycle characteristics of the input data using a learnable parameter matrix, and then the two are added together to obtain a comprehensive time embedding, achieving temporal feature fusion.

[0045] S3.2: Embed the spatial information of the satellite image clipped by latitude and longitude into a matrix, where the rows and columns of the matrix correspond to the pixel coordinates of the image, and the matrix elements represent the solar radiation value of the corresponding pixel;

[0046] S3.3: Use the broadcast mechanism to expand the dimension of the spatial embedding vector to match the dimension of the temporal embedding vector. Align the expanded spatial embedding vector and the temporal embedding vector on the corresponding time dimension to ensure that the spatial information and temporal information corresponding to each time step match.

[0047] S3.4: Map the aligned data through the convolutional layer and convert it into a channel space compatible with the GPT-2 model;

[0048] S3.5: In the channel dimension, the convolutional mapped features are concatenated with the temporal embedding and spatial embedding features to form a feature matrix that integrates spatiotemporal information, i.e., a spatiotemporal feature fusion matrix is ​​formed.

[0049] S3.6: Use a convolutional layer to compress the fused feature matrix to reduce its dimension to a size that matches the GPT-2 model input dimension.

[0050] S3.7: Convert the compressed feature matrix into the standard input format of the large language model so that it can be input into the GPT-2 model for further processing and analysis;

[0051] In the above technical solution, a preferred technical solution may also be that in step S3.1, the time feature calculation formula is:

[0052] (3);

[0053] (4);

[0054] (5);

[0055] in, and Represent the daily and monthly periodic input feature tensors respectively, and Represent the daily cycle and monthly cycle embedding weight matrices, and are high-dimensional semantic embeddings of the solar and lunar cycles, respectively, used to capture temporal dynamics, and T is the integrated temporal representation after fusion.

[0056] In the above technical solution, a preferred technical solution may also be that in step S3.5, the calculation formula for spatiotemporal feature fusion is:

[0057] (6);

[0058] (7);

[0059] (8);

[0060] in, is the input data, P is the input projection, and is the learnable spatial position embedding parameter, P is the original input spatial feature embedding, S is the output spatial embedding feature, is the input convolution parameter, is the fusion convolution parameter, is the fused spatiotemporal feature.

[0061] In the above technical solution, a preferred technical solution may also be that the step S4 specifically includes:

[0062] S4.1: Freeze the multi-head attention mechanism and feed-forward network (FFN) parameters of the first four layers of the GPT-2 model to preserve the general temporal modeling capabilities of the pre-trained large language model;

[0063] S4.2: Unfreeze the parameters of the multi-head attention mechanism in the last two layers so that it can learn the spatiotemporal correlation of solar radiation through the back-propagation algorithm;

[0064] S4.3: Input the spatiotemporal fusion feature sequence into the GPT-2 model, which automatically captures non-local spatiotemporal correlation features such as cloud movement trajectories and regional radiation transfer effects;

[0065] S4.4: Update the unfrozen parameters through the backpropagation algorithm to optimize the model's ability to capture specific spatiotemporal correlation patterns, and ultimately output a hidden state that can describe the model's internal state information;

[0066] S4.5: Construct a regression convolution layer and input the spatiotemporal fusion feature sequence into this layer to generate a low-resolution preliminary solar radiation forecast result. The forecast result contains the regional solar radiation intensity distribution information for the next 16 time steps (0-4 hours), providing basic data for subsequent super-resolution reconstruction.

[0067] In the above technical solution, a preferred technical solution may also be that step S5 specifically includes:

[0068] S5.1: Construct a ResNet residual super-resolution module, which contains multiple residual block groups for extracting local texture features and upsamples the spatial resolution of the preliminary prediction results to the original high-resolution image size through sub-pixel convolution operations;

[0069] S5.2: Introduce skip connections and residual scaling, set the scaling factor to 0.15, and retain the global radiation distribution characteristics of the original predicted image while improving the spatial resolution, thus achieving detail enhancement of the predicted image.

[0070] S5.3: Extract single-channel solar radiation intensity data from the super-resolution reconstruction results, denormalize the extracted radiation intensity data, map the standardized radiation intensity value output by the model back to the actual physical dimension, and generate the final high-precision solar radiation image prediction results that can be used in business applications to meet actual application needs.

[0071] In the above technical solution, a preferred technical solution may also be that step S6 further includes:

[0072] S6.1: Optimizer parameter configuration of the model, including:

[0073] S6.11: AdamW optimizer is selected as the optimization algorithm for model training. This optimizer combines the advantages of the Adam optimization algorithm and weight decay technology to effectively prevent model overfitting.

[0074] S6.12: Set the initial learning rate to , a cosine annealing strategy is used to dynamically adjust the learning rate. The cycle is set to 50 epochs. In each cycle, the learning rate gradually decreases according to the shape of the cosine function, and then returns to a higher value in the next cycle. This can encourage the model to continuously jump out of the local optimum during training and find a better solution;

[0075] S6.13: Set the weight decay parameter to By applying L2 norm penalty to model weights during the optimization process, we can prevent model parameters from being too complex, thereby reducing the risk of model overfitting and improving the generalization ability of the model.

[0076] S6.14: Set the gradient clipping threshold to 5. In each training step, the calculated gradient is constrained by its norm. If the gradient norm exceeds the threshold, it is scaled to within the threshold range. This can effectively prevent the gradient explosion problem and ensure the stability of the model training process.

[0077] S6.2: Model preservation and dynamic early stopping, specifically including:

[0078] S6.21: The mean absolute error is used as the loss function of the model to measure the average absolute difference between the model prediction value and the true value;

[0079] S6.22: During model training, evaluation metrics are calculated at the end of each epoch, including mean absolute error, root mean square error (RMSE), mean absolute percentage error (MAPE), and weighted mean absolute percentage error (WMAPE).

[0080] S6.23: Continuously monitor the mean absolute error (MAE) on the validation set and use it as the primary basis for judging model performance.

[0081] S6.24: Set a patience value of 100 epochs. When the mean absolute error metric on the validation set does not improve within 100 consecutive epochs, trigger the early stopping mechanism and terminate the training process. This can avoid overfitting of the model in the later stages of training and save training resources.

[0082] S6.25: Save model parameters only when validation loss decreases. Specifically, whenever the mean absolute error metric on the validation set improves compared to the last save, save the current model parameters. Conversely, if validation loss does not decrease, do not save. This ensures that the saved model parameters correspond to the model state that performed best on the validation set.

[0083] In response to the above-mentioned problems in the background technology, the present invention proposes an STLLM_ResNet model that can be applied to regional-level short-term solar radiation prediction. Its core solution includes:

[0084] 1. This paper proposes a multi-scale super-resolution model, introducing the spatiotemporal large language model (STLLM) to the field of solar radiation prediction. Based on the GPT-2 (Generative Pretrained Transformer) architecture developed by OpenAI, this model integrates the diurnal and seasonal variations of solar radiation, as well as the spatial characteristics of regional solar radiation variations caused by cloud movement, to construct a multi-scale super-resolution model. Using a spatiotemporal embedding strategy, the radiation sequence is reconstructed into spatiotemporal labels. By dynamically integrating spatial position encoding with global temporal representation, this approach enables accurate modeling of non-local correlations such as cloud movement and inter-regional radiation transfer. Furthermore, a partial frozen attention mechanism is introduced to retain attention heads related to radiation physics, reducing redundant computational overhead and improving model interpretability.

[0085] 2. To address the existing models' dependence on large-scale data, this invention utilizes sampling learning techniques to efficiently utilize lightweight, small-sample datasets, achieving high-precision solar radiation prediction and high-quality image generation even with limited computing resources. Furthermore, this model combines the few-shot learning advantages of the Large Language Model (LLM) with the spatial refinement capabilities of the ResNet super-resolution model to construct a cross-modal generalization framework. The LLM improves prediction robustness in few-shot scenarios through pre-training knowledge transfer, while the ResNet enhances the accuracy of local radiation intensity details through super-resolution reconstruction. This design effectively overcomes the traditional model's dependence on data size and significantly improves the model's adaptability and prediction accuracy in complex meteorological conditions.

[0086] This invention aims to improve the accuracy and image quality of regional solar radiation forecasts using small sample historical data and limited computing resources. It achieves efficient spatiotemporal joint modeling to accurately capture cross-regional and cross-period dependencies; enhances generalization capabilities for small sample scenarios to maintain forecast robustness with limited data; optimizes high-resolution forecasts to capture details of local radiation intensity; and implements dynamic correlation adaptive learning to overcome the traditional constraints of spatial correlation modeling. By improving the accuracy of 0-4h solar radiation forecasts, the reliability of ultra-short-term photovoltaic power generation forecasts can be enhanced to meet the needs of power systems and power generation companies.

[0087] Compared with the prior art, the present invention has the following advantages and positive effects:

[0088] 1) Lightweight training architecture: By partially freezing the underlying parameters of the GPT-2 model, the number of model parameters is reduced, and the training time on a single NVIDIA RTX 4090 card is shortened from 24 hours to 14 hours. Video memory usage and resource consumption are significantly reduced, and computational efficiency and generalization are optimized, greatly enhancing the applicability of this method in practical applications.

[0089] 2) Improved Small-Sample Learning: Through a collaborative optimization framework combining dynamic spatiotemporal embedding with a partially frozen pre-trained large language model, the model effectively captures spatiotemporal feature correlations with limited labeled samples, significantly reducing its reliance on large labeled datasets. Experiments demonstrate that this method maintains comparable predictive performance to traditional fully supervised models (such as UNet) even with a small number of training samples, validating its robustness and generalization capabilities in small-sample learning tasks.

[0090] 3) Significantly improved forecasting accuracy. Long-sequence forecasting performance has been optimized by integrating GPT-2 time series modeling with a residual super-resolution network. In tests conducted on typical months of January, April, July, and October, regional forecast MAE indicators were significantly reduced compared to existing advanced large-scale time series forecasting models. The average reduction compared to traditional UNet models was 23.6%, and the RMSE was reduced by 18.2%. The most significant improvement in forecast error was seen in July, a month with complex lighting conditions, with a MAE reduction of 27.3%.

[0091] In summary, the present invention provides a solar radiation forecasting method based on satellite data and a multi-spatiotemporal-scale super-resolution model, introduces natural language processing and sequence modeling technology into the field of solar radiation forecasting, and realizes high-precision regional solar radiation forecasting for the next 0-4 hours by integrating spatiotemporal features with deep learning. It solves the problems of spatiotemporal dependency separation, poor adaptability due to small samples, and insufficient spatial resolution in solar radiation forecasting. This method is a more efficient and accurate regional short-term solar radiation forecasting method. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] Figure 1 This is a reference diagram (block diagram) of the solar radiation forecasting method based on satellite data and a multi-space-time scale super-resolution model of the present invention.

[0093] Figure 2 The present invention is a flow chart of the solar radiation forecasting method based on satellite data and multi-time and space scale super-resolution large model.

[0094] FIG3 is a schematic diagram of the structure of a photovoltaic power generation prediction system implemented in a 50MW photovoltaic power station in a certain province according to the present invention. DETAILED DESCRIPTION

[0095] Example 1: Figure 1 、 Figure 2 As shown, the solar radiation forecasting method based on satellite data and multi-time and space scale super-resolution large model of the present invention includes the following steps:

[0096] S1: Acquire satellite remote sensing data, parse the time tags, convert them to Beijing time, and extract the year, month, day, hour, minute, and second information; filter valid data from sunrise to sunset and discard invalid data at night; randomly sample valid data by month as training samples, and crop the satellite image to the size of the prediction area. The satellite remote sensing data (satellite data) is the surface downlink shortwave radiation product of the Fengyun-4 meteorological satellite. Step S1 specifically includes:

[0097] S1.1: The acquired satellite remote sensing data include the equal latitude and longitude gridded data of the downlink shortwave solar radiation retrieval product, whose spatial resolution conforms to the standardized output format of the global atmospheric reanalysis product.

[0098] S1.2: Read solar radiation data from satellite data. By parsing the time tag in the file name, extract the UTC time and convert it to Beijing time. At the same time, obtain the year, month, day, hour, minute, and second information. The satellite data time resolution is 15 minutes.

[0099] S1.3: Based on the latitude and longitude of the prediction area and the date and time information, calculate the daily sunrise and sunset times, filter the valid data from sunrise to sunset, and eliminate the invalid data at night.

[0100] S1.4: To increase sample diversity, further compress the dataset, increase training speed, and reduce computing resource pressure, monthly random sampling is used. The data is grouped by month, and four days of data are randomly selected each month as training samples. The selected satellite image paths and corresponding timestamp information are stored in a local .csv file.

[0101] S1.5: Crop the satellite-observed solar radiation images used in model training into the size of the prediction area to form a local spatial matrix to meet the model input requirements.

[0102] S1.6: Flatten the cropped two-dimensional radiation image matrix into a one-dimensional vector representing the number of features of each sample after expansion in the spatial dimension.

[0103] S2: Construct a feature set for daily and seasonal variations in solar radiation. Use time normalization to map timestamps to the zero-one interval and use one-hot encoding to extract month information. Reorganize the data into a four-dimensional tensor [number of samples, time steps, number of features, number of channels], divide it into training, validation, and test sets, and perform data dimensionality reduction and standardization. Step S2 specifically includes:

[0104] S2.1: Use a time series normalization method to construct a daily variation feature set.

[0105] In step S2.1, a time series normalization method is used to map the time dimension to the [0,1) continuous space by truncating the sample timestamps to the current day's reference point and calculating the proportional difference. The time series data is processed as follows:

[0106] (a) Baseline time alignment: Align the timestamp index to midnight of the current day to generate a base time series.

[0107] (b) Time difference calculation: Calculate the difference between the sample timestamp and the current day's base timestamp to obtain the length of time that has passed since the sample file was created;

[0108] (c) Proportional mapping: Divide the time difference by the total time unit corresponding to 24 hours and output the normalized time proportion value in the interval [0,1);

[0109] Assume that the sample timestamp is t, then the normalized value Tn satisfies:

[0110] (2), in, Indicates the zero o'clock of the day corresponding to the sample timestamp. 1 Day corresponds to 24 hours.

[0111] S2.2: Construct a seasonal variation feature set and extract the month information (one-hot encoding) in the sample time to capture the seasonal variation pattern.

[0112] S2.3: Finally, the multidimensional features containing diurnal and seasonal variations are embedded into the feature dataset to form a three-dimensional tensor [number of samples, number of features, number of channels] consisting of time features and solar radiation.

[0113] S2.4: Generate ordered time series data with 15-minute intervals. Suppose the input time offset is n , the output time offset is m , that is, using the past n Solar radiation images, predicting the future m solar radiation images, and then judge the temporal continuity of the sample sequence: n + m Among the consecutive samples, only the valid sequences with a time interval of 15 minutes between adjacent samples are retained. Finally, the time step dimension is added and the data dimension is reorganized into a four-dimensional tensor [number of samples, time steps, number of features, number of channels].

[0114] S2.5: Dataset construction and serialized storage. The dataset is divided into training set, validation set, and test set with the following ratios:

[0115] The ratio of training set: validation set: test set is 7:1.5:1.5. The first 70% of the data in the dataset is the training set, the middle 15% of the data in the dataset is the validation set, and the last 15% of the data in the dataset is the test set. The data is processed into the NPZ format, a compressed data storage format dedicated to the NumPy library, and stored in a local directory.

[0116] S2.6: Data dimensionality reduction. Before model training, we first load the NPZ dataset file and use adaptive average pooling (AdaptiveAvgPool2d) to reduce the original data size from a 128-pixel × 224-pixel grid to a 32-pixel × 56-pixel grid, preserving the spatial distribution characteristics.

[0117] S2.7: Data standardization: Perform Z-score standardization on the data based on the mean and standard deviation of the training set to ensure uniform input dimensions. The calculation formula is: (1), Among them, x is the radiation value of the training set data, represents the mean radiation value of the training set, std represents the standard deviation of the radiation value of the training set, and z represents the standardized data value.

[0118] S3: Linearly map the diurnal and monthly cycle features of solar radiation using a learnable parameter matrix to generate a comprehensive time embedding; align the spatial information of the satellite image with the time embedding, and use a convolutional layer to map the spatiotemporal information of solar radiation to a channel space compatible with the GPT-2 model; concatenate the temporal, spatial, and projected features, generate a fused spatiotemporal feature matrix through compression operations, and convert it into the standard input format of the large language model. Step S3 specifically includes:

[0119] S3.1: The multi-scale spatiotemporal super-resolution model receives a tensor of historical solar radiation observation data with the dimensions [batch size, time step, number of samples, number of channels], where the number of channels includes three channels: radiation intensity, daily characteristics, and monthly characteristics. Using a time embedding module, the daily and monthly cycle characteristics of the input data are linearly mapped using a learnable parameter matrix. The two are then added together to obtain a comprehensive time embedding, achieving time feature fusion. In step S3.1, the time feature calculation formula is:

[0120] (3);

[0121] (4);

[0122] (5);

[0123] in, and Represent the daily and monthly periodic input feature tensors respectively, and Represent the daily cycle and monthly cycle embedding weight matrices, and are high-dimensional semantic embeddings of the solar and lunar cycles, respectively, used to capture temporal dynamics, and T is the integrated temporal representation after fusion.

[0124] S3.2: Embed the spatial information of the satellite image clipped by longitude and latitude into a matrix, where the rows and columns of the matrix correspond to the pixel coordinates of the image, and the matrix elements represent the solar radiation value of the corresponding pixel.

[0125] S3.3: Use the broadcast mechanism to expand the dimension of the spatial embedding vector to match the dimension of the temporal embedding vector, and align the expanded spatial embedding vector and the temporal embedding vector on the corresponding time dimension to ensure that the spatial information and temporal information corresponding to each time step match.

[0126] S3.4: Map the aligned data through the convolutional layer and convert it into a channel space compatible with the GPT-2 model.

[0127] S3.5: In the channel dimension, the convolutional mapped features are concatenated with the temporal embedding and spatial embedding features to form a feature matrix that integrates spatiotemporal information, i.e., a spatiotemporal feature fusion matrix. In step S3.5, the spatiotemporal feature fusion calculation formula is:

[0128] (6);

[0129] (7);

[0130] (8);

[0131] in, is the input data, P is the input projection, and is the learnable spatial position embedding parameter, P is the original input spatial feature embedding, S is the output spatial embedding feature, is the input convolution parameter, is the fusion convolution parameter, is the fused spatiotemporal feature.

[0132] S3.6: Use a convolutional layer to compress the fused feature matrix to reduce its dimension to a size that matches the GPT-2 model input dimension.

[0133] S3.7: Convert the compressed feature matrix into the standard input format of the large language model so that it can be input into the GPT-2 model for further processing and analysis.

[0134] S4: Using a layered parameter control strategy, freeze the parameters of the first four layers of the GPT-2 model (the total number of layers can be 12), and unfreeze the parameters of the last two layers to learn the spatiotemporal correlation of solar radiation, forming a forecast model that integrates the spatiotemporal characteristics of solar radiation; generate low-resolution preliminary forecast results through the regression convolution layer. Step S4 specifically includes:

[0135] S4.1: Freeze the multi-head attention mechanism and feed-forward network (FFN) parameters of the first four layers of the GPT-2 model to preserve the general temporal modeling capabilities of the pre-trained large language model.

[0136] S4.2: Unfreeze the parameters of the multi-head attention mechanism in the last two layers so that it can learn the spatiotemporal correlation of solar radiation through the back-propagation algorithm.

[0137] S4.3: The spatiotemporal fusion feature sequence is input into the GPT-2 model, and the model automatically captures non-local spatiotemporal correlation features such as cloud movement trajectories and regional radiation transfer effects.

[0138] S4.4: Update the unfrozen parameters through the backpropagation algorithm to optimize the model's ability to capture specific spatiotemporal correlation patterns, and finally output the hidden state that can describe the internal state information of the model.

[0139] S4.5: Construct a regression convolution layer and input the spatiotemporal fusion feature sequence into this layer to generate a low-resolution preliminary solar radiation forecast result. The forecast result contains the regional solar radiation intensity distribution information for the next 16 time steps (0-4 hours), providing basic data for subsequent super-resolution reconstruction.

[0140] S5: Use the ResNet residual super-resolution module for upsampling, combine skip connections and residual scaling to reconstruct a high-resolution image, output 0-4 hour single-channel solar radiation intensity data and denormalize it to physical dimensions. Step S5 specifically includes:

[0141] S5.1: Construct a ResNet residual super-resolution module, which contains multiple residual block groups for extracting local texture features and upsamples the spatial resolution of the preliminary prediction results to the original high-resolution image size through sub-pixel convolution operations.

[0142] S5.2: Introduce skip connections and residual scaling, set the scaling factor to 0.15, and retain the global radiation distribution characteristics of the original predicted image while improving the spatial resolution, thereby achieving detail enhancement of the predicted image.

[0143] S5.3: Extract single-channel solar radiation intensity data from the super-resolution reconstruction results, denormalize the extracted radiation intensity data, map the standardized radiation intensity value output by the model back to the actual physical dimension, and generate the final high-precision solar radiation image prediction results that can be used in business applications to meet actual application needs.

[0144] S6: Configure the model's optimizer parameters, dynamically adjust the learning rate using a cosine annealing strategy; set weight decay and gradient clipping to use mean absolute error as the loss function, monitor validation set metrics, and save the optimal model parameters using a dynamic early stopping mechanism. Step S6 further includes:

[0145] S6.1: Optimizer parameter configuration of the model, including:

[0146] S6.11: The AdamW optimizer is selected as the optimization algorithm for model training. This optimizer combines the advantages of the Adam optimization algorithm and weight decay technology, which can effectively prevent model overfitting.

[0147] S6.12: Set the initial learning rate to , the cosine annealing strategy is used to dynamically adjust the learning rate. The cycle is set to 50 epochs. In each cycle, the learning rate gradually decreases according to the shape of the cosine function, and then returns to a higher value in the next cycle. This can encourage the model to continuously jump out of the local optimum during the training process and find a better solution.

[0148] S6.13: Set the weight decay parameter to By applying L2 norm penalty to model weights during the optimization process, the model parameters are prevented from being too complex, thereby reducing the risk of model overfitting and improving the generalization ability of the model.

[0149] S6.14: Set the gradient clipping threshold to 5. In each training step, the calculated gradient is constrained by its norm. If the gradient norm exceeds the threshold, it is scaled to within the threshold range. This can effectively prevent the gradient explosion problem and ensure the stability of the model training process.

[0150] S6.2: Model preservation and dynamic early stopping, specifically including:

[0151] S6.21: The mean absolute error is used as the loss function of the model, which is used to measure the average absolute difference between the model prediction value and the true value.

[0152] S6.22: During model training, evaluation metrics are calculated at the end of each epoch, including mean absolute error, root mean square error, mean absolute percentage error, and weighted mean absolute percentage error.

[0153] S6.23: Continuously monitor the mean absolute error metric on the validation set and use it as the primary basis for judging model performance.

[0154] S6.24: Set a patience value of 100 epochs. When the mean absolute error metric on the validation set does not improve within 100 consecutive epochs, trigger the early stopping mechanism and terminate the training process. This can avoid overfitting of the model in the later stages of training and save training resources.

[0155] S6.25: Save model parameters only when validation loss decreases. Specifically, whenever the mean absolute error metric on the validation set improves compared to the last save, save the current model parameters. Conversely, if validation loss does not decrease, do not save. This ensures that the saved model parameters correspond to the model state that performed best on the validation set.

[0156] The following are application examples of the present invention: Figure 3Figure 1: Provincial dispatch / local dispatch main station side 1, dispatch data network 2, meteorological station observation data 3, duty platform query terminal 4 (duty platform query terminal is duty platform client), power prediction server 5, photovoltaic power station side 6, forward isolation device 7, reverse isolation device 8, satellite data receiving server 9, Internet 10, satellite data publishing platform 11. Figure 3 As shown, the technical solution of the present invention was implemented in a 50MW photovoltaic power station in a certain province. During the implementation, a three-tier server architecture based on a physically isolated one-way transmission channel was constructed: the external network server deployed a meteorological data acquisition module, established a one-way data link with the (intranet) power prediction server 5 through a power-specific reverse isolation device 8, and the duty platform query terminal 4 was interconnected with the power prediction server 5 through an intranet dedicated line, forming a network topology that complies with the safety protection regulations of the power monitoring system. The following core steps are performed during the system operation cycle: (1) Satellite data transmission and preprocessing: The external satellite data on the satellite data publishing platform 11 is transmitted to the (extranet) satellite data receiving server 9 of the photovoltaic station through the Internet 10. The satellite data is then sent to the (intranet) power prediction server 5 through the reverse isolation device 8. The power prediction server 5 also receives the meteorological station observation data 3 from the station end. (2) Algorithm operation: On the power prediction server 5, the model proposed by the present invention (i.e., the prediction model integrating the spatiotemporal characteristics of solar radiation) is used to dynamically predict the solar radiation value. Then, the forecast result at the location of the photovoltaic station is extracted. The predicted solar radiation energy is converted into power using the power prediction algorithm to achieve a 0-4h ultra-short-term photovoltaic power generation power forecast. (3) Forecast result reporting, display and feedback: The forecast result is uploaded to the provincial / local dispatching master station side 1 through the dispatching data network 2. At the same time, the forecast result can be viewed at the duty platform query terminal 4. The forecast result and actual data can also be transmitted to the satellite data receiving server 9 through the forward isolation device 7 for algorithm back-calculation. The overall network topology conforms to the safety protection regulations of the power monitoring system.

[0157] In summary, the above embodiments of the present invention provide a solar radiation forecasting method based on satellite data and a multi-spatiotemporal scale super-resolution model, which introduces natural language processing and sequence modeling technology into the field of solar radiation prediction. By integrating spatiotemporal features and deep learning, it achieves high-precision regional solar radiation prediction for the next 0-4 hours, and solves the problems of spatiotemporal dependency separation, poor adaptability of small samples, and insufficient spatial resolution in solar radiation prediction. This method is a more efficient and accurate regional short-term solar radiation prediction method.

Claims

1. A solar radiation forecasting method based on satellite data and multi-time and space scale super-resolution large-scale model, characterized by It includes the following steps: S1: Acquire satellite remote sensing data, parse the time tag and convert it to Beijing time, extract the year, month, day, hour, minute, and second information; filter the valid data from sunrise to sunset and remove the invalid data at night; randomly sample the valid data by month as training samples, and crop the satellite image to the prediction area size; S2: Construct a feature set of daily and seasonal variations of solar radiation. Use time normalization to map timestamps to the zero-one interval and use one-hot encoding to extract month information. Reorganize the data into a four-dimensional tensor [number of samples, time steps, number of features, number of channels], divide it into training, validation, and test sets, and perform data dimensionality reduction and standardization. S3: Linearly map the diurnal and monthly cycle features of solar radiation through a learnable parameter matrix to generate a comprehensive time embedding; align the spatial information of satellite images with the time embedding, and use convolutional layers to map the spatiotemporal information of solar radiation to a channel space compatible with the GPT-2 model; concatenate the temporal, spatial, and projected features, generate a fused spatiotemporal feature matrix through compression operations, and convert it into the standard input format for large language models; S4: Using a hierarchical parameter control strategy, the parameters of the first four layers of the GPT-2 model are frozen, and the parameters of the last two layers are unfrozen to learn the spatiotemporal correlation of solar radiation, forming a forecast model that integrates the spatiotemporal characteristics of solar radiation; low-resolution preliminary forecast results are generated through the regression convolution layer; S5: Use the ResNet residual super-resolution module for upsampling, combine skip connections and residual scaling to reconstruct a high-resolution image, output 0-4 hours single-channel solar radiation intensity data and denormalize it to physical dimensions; S6: Configure the model's optimizer parameters, using a cosine annealing strategy to dynamically adjust the learning rate; set weight decay and gradient clipping with mean absolute error as the loss function, monitor validation set metrics, and save the optimal model parameters through a dynamic early stopping mechanism.

2. The solar radiation forecasting method based on satellite data and multi-time and space scale super-resolution large model according to claim 1 is characterized in that Step S1 includes: S1.1: The acquired satellite remote sensing data include the iso-latitude and longitude gridded data of the downlink shortwave solar radiation retrieval product, and its spatial resolution conforms to the standardized output format of the global atmospheric reanalysis product; S1.2: Read solar radiation data from satellite data. By parsing the time tag in the file name, extract the UTC time and convert it to Beijing time. At the same time, obtain the year, month, day, hour, minute, and second information. The satellite data time resolution is 15 minutes. S1.3: Based on the latitude and longitude of the forecast area and the date and time information, calculate the daily sunrise and sunset times, filter the valid data between sunrise and sunset, and eliminate the invalid data at night; S1.4: Use monthly random sampling to group the data by month. Randomly select 4 days of data each month as training samples. Save the selected satellite image paths and corresponding timestamp information into a local .csv file. S1.5: Clip the satellite solar radiation images used in model training to the size of the prediction area to form a local spatial matrix; S1.6: Flatten the cropped two-dimensional radiation image matrix into a one-dimensional vector, representing the number of features of each sample after expansion in the spatial dimension.

3. The solar radiation forecasting method based on satellite data and multi-time and space scale super-resolution large model according to claim 1 is characterized in that Step S2 includes: S2.1: Use a time series normalization method to construct a diurnal variation feature set; S2.2: Construct a seasonal variation feature set and extract the monthly information in the sample time; S2.3: Finally, the multidimensional features containing diurnal and seasonal variations are embedded into the feature dataset to form a three-dimensional tensor [number of samples, number of features, number of channels] consisting of time features and solar radiation; S2.4: Generate ordered time series data with 15-minute intervals. Assume the input time offset is n and the output time offset is m. That is, use the past n solar radiation images to predict the future m solar radiation images. Then, determine the temporal continuity of the sample sequence: among the n + m consecutive samples, only retain the valid sequences with a time interval of 15 minutes between adjacent samples. Finally, add the time step dimension and restructure the data into a four-dimensional tensor [number of samples, time step, number of features, number of channels]. S2.5: Dataset construction and serialized storage. The dataset is divided into training set, validation set, and test set with the following ratios: The training set: validation set: test set ratio is 7:1.5:1.

5. The data is processed into the NPZ format, a compressed data storage format specifically designed for the NumPy library, and stored in a local directory. S2.6: Data dimensionality reduction. Before model training, we first load the NPZ dataset file and use adaptive average pooling to reduce the original data size from a 128-pixel × 224-pixel grid to a 32-pixel × 56-pixel grid, preserving the spatial distribution characteristics. S2.7: Data standardization: Perform Z-score standardization on the data based on the mean and standard deviation of the training set to ensure uniform input dimensions. The calculation formula is: (1), Among them, x is the radiation value of the training set data, represents the mean radiation value of the training set, std represents the standard deviation of the radiation value of the training set, and z represents the standardized data value.

4. The solar radiation forecasting method based on satellite data and multi-time and space scale super-resolution large model according to claim 3 is characterized in that In step S2.1, a time series normalization method is used to map the time dimension to the [0,1) continuous space by truncating the sample timestamps to the current day's reference point and calculating the proportional difference. The time series data is processed as follows: (a) Baseline time alignment: Align the timestamp index to midnight of the current day to generate a base time series. (b) Time difference calculation: Calculate the difference between the sample timestamp and the current day's base timestamp to obtain the length of time that has passed since the sample file was created; (c) Proportional mapping: Divide the time difference by the total time unit corresponding to 24 hours and output the normalized time proportion value in the interval [0,1); Assume that the sample timestamp is t, then the normalized value Tn satisfies: (2), in, Indicates the zero o'clock of the day corresponding to the sample timestamp. 1 Day corresponds to 24 hours.

5. The solar radiation forecasting method based on satellite data and multi-time and space scale super-resolution large model according to claim 1 is characterized in that Step S3 includes: S3.1: The multi-scale spatiotemporal super-resolution model receives historical solar radiation observation data tensors of the dimensions [batch size, time step, number of samples, number of channels], where the number of channels includes three channels: radiation intensity, daily characteristics, and monthly characteristics. A time embedding module is used to linearly map the daily and monthly cycle characteristics of the input data using a learnable parameter matrix, and then sum the two to obtain a comprehensive time embedding, achieving temporal feature fusion. S3.2: Embed the spatial information of the satellite image clipped by latitude and longitude into a matrix, where the rows and columns of the matrix correspond to the pixel coordinates of the image, and the matrix elements represent the solar radiation value of the corresponding pixel; S3.3: Use the broadcast mechanism to expand the dimension of the spatial embedding vector to match the dimension of the temporal embedding vector. Align the expanded spatial embedding vector and the temporal embedding vector on the corresponding time dimension to ensure that the spatial information and temporal information corresponding to each time step match. S3.4: Map the aligned data through the convolutional layer and convert it into a channel space compatible with the GPT-2 model; S3.5: In the channel dimension, the convolutional mapped features are concatenated with the temporal embedding and spatial embedding features to form a feature matrix that integrates spatiotemporal information, i.e., a spatiotemporal feature fusion matrix is ​​formed. S3.6: Use a convolutional layer to compress the fused feature matrix to reduce its dimension to a size that matches the GPT-2 model input dimension. S3.7: Convert the compressed feature matrix into the standard input format of the large language model and input it into the GPT-2 model for further processing and analysis.

6. The solar radiation forecasting method based on satellite data and multi-time and space scale super-resolution large model according to claim 5 is characterized in that In step S3.1, the time characteristic calculation formula is: (3); (4); (5); in, and Represent the daily and monthly periodic input feature tensors respectively, and denote the embedding weight matrices of daily cycle and monthly cycle respectively, and are high-dimensional semantic embeddings of the solar and lunar cycles, respectively, capturing temporal dynamics, and T is the integrated temporal representation after fusion.

7. The solar radiation forecasting method based on satellite data and multi-time and space scale super-resolution large model according to claim 5 is characterized in that In step S3.5, the calculation formula for spatiotemporal feature fusion is: (6); (7); (8); in, is the input data, P is the input projection, and is the learnable spatial position embedding parameter, P is the original input spatial feature embedding, S is the output spatial embedding feature, is the input convolution parameter, is the fusion convolution parameter, is the fused spatiotemporal feature.

8. The solar radiation forecasting method based on satellite data and multi-time and space scale super-resolution large model according to claim 1 is characterized in that Step S4 includes: S4.1: Freeze the multi-head attention mechanism and feed-forward network (FFN) parameters of the first four layers of the GPT-2 model to preserve the general temporal modeling capabilities of the pre-trained large language model; S4.2: Unfreeze the parameters of the multi-head attention mechanism in the last two layers so that it can learn the spatiotemporal correlation of solar radiation through the back-propagation algorithm; S4.3: Input the spatiotemporal fusion feature sequence into the GPT-2 model, which automatically captures non-local spatiotemporal correlation features such as cloud movement trajectories and regional radiation transfer effects; S4.4: Update the unfrozen parameters through the backpropagation algorithm to optimize the model's ability to capture specific spatiotemporal correlation patterns, and ultimately output a hidden state that can describe the model's internal state information; S4.5: Construct a regression convolution layer and input the spatiotemporal fusion feature sequence into this layer to generate a low-resolution preliminary solar radiation forecast result. The forecast result contains the regional solar radiation intensity distribution information for the next 16 time steps (0-4 hours), providing basic data for subsequent super-resolution reconstruction.

9. The solar radiation forecasting method based on satellite data and multi-time and space scale super-resolution large model according to claim 1 is characterized in that Step S5 includes: S5.1: Construct a ResNet residual super-resolution module, which contains multiple residual block groups for extracting local texture features and upsamples the spatial resolution of the preliminary prediction results to the original high-resolution image size through sub-pixel convolution operations; S5.2: Introduce skip connections and residual scaling, set the scaling factor to 0.15, and retain the global radiation distribution characteristics of the original predicted image while improving the spatial resolution, thus achieving detail enhancement of the predicted image. S5.3: Extract single-channel solar radiation intensity data from the super-resolution reconstruction results, denormalize the extracted radiation intensity data, map the standardized radiation intensity value output by the model back to the actual physical dimension, and generate the final high-precision solar radiation image prediction results that can be used in business.

10. The solar radiation forecasting method based on satellite data and multi-time and space scale super-resolution large model according to claim 1 is characterized in that Step S6 includes: S6.1: Optimizer parameter configuration of the model, including: S6.11: AdamW optimizer is selected as the optimization algorithm for model training. This optimizer combines the advantages of the Adam optimization algorithm and weight decay technology. S6.12: Set the initial learning rate to , the cosine annealing strategy is used to dynamically adjust the learning rate. The cycle is set to 50 epochs. In each cycle, the learning rate gradually decreases according to the shape of the cosine function, and then returns to a higher value in the next cycle; S6.13: Set the weight decay parameter to , by applying an L2 norm penalty to the model weights during the optimization process; S6.14: Set the gradient clipping threshold to 5. In each training step, the calculated gradient is constrained by its norm. If the gradient norm exceeds the threshold, it is scaled to within the threshold range. S6.2: Model preservation and dynamic early stopping, specifically including: S6.21: The mean absolute error is used as the loss function of the model to measure the average absolute difference between the model prediction value and the true value; S6.22: During model training, evaluation metrics are calculated at the end of each epoch, including mean absolute error, root mean square error, mean absolute percentage error, and weighted mean absolute percentage error. S6.23: Continuously monitor the mean absolute error (MAE) on the validation set and use it as the primary basis for judging model performance. S6.24: Set a patience value of 100 epochs. When the mean absolute error on the validation set does not improve within 100 consecutive epochs, trigger the early stopping mechanism and terminate the training process. S6.25: Save the model parameters only when the validation loss decreases. Otherwise, do not save if the validation loss does not decrease.

Citation Information

Patent Citations

  • Solar radiation forecasting method, device, equipment and medium

    CN114492204A

  • Intelligent forecasting method for ground surface incident solar radiation based on wind cloud 4A satellite

    CN117808155A

  • Pre-training large model traffic flow prediction method based on double-activation domain bridging and space-time self-attention

    CN120260282A