Photovoltaic power short-term prediction method and device based on cloud evolution
Through the cloud evolution-based photovoltaic power short-term prediction method, combined with multiple data processing and models, the problem of insufficient photovoltaic power generation power prediction accuracy in existing technologies is solved, efficient and accurate photovoltaic power generation power prediction is achieved, and energy management and grid stability are optimized.
Patent Information
- Application Number
- CN202510691524.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-19
AI Technical Summary
Existing photovoltaic power generation prediction methods lack accuracy when dealing with multivariable and nonlinear relationships, and are particularly difficult to meet actual needs under dynamic cloud changes and complex meteorological conditions.
A short-term photovoltaic power prediction method based on cloud evolution is adopted. By acquiring real-time cloud image data and meteorological data, local binary pattern LBP and three-dimensional fuzzy adaptive extended Kalman filter 3D-FAEKF are combined for data preprocessing. The improved RepViTS-YOLOX model is used to process cloud image data, and the meteorological data is processed by combining the SAMformer model. The solar radiation is calculated by the Hargreaves equation, and finally the time series lightweight adaptive network TSLANet is used to predict photovoltaic power.
It significantly improves the accuracy and efficiency of photovoltaic power forecasting, optimizes photovoltaic power generation planning and energy storage scheduling, reduces the rate of abandoned light, improves energy utilization, provides a reliable basis for grid scheduling, and enhances grid stability.
Smart Images

Figure CN120670934A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of photovoltaic power prediction, and specifically relates to a method and device for short-term photovoltaic power prediction based on cloud evolution. Background Art
[0002] As a clean, renewable energy source, photovoltaic power generation plays a vital role in the global energy transition. However, PV power generation is affected by a variety of environmental factors, including cloud cover, meteorological conditions (such as temperature, humidity, wind speed, and direction), and solar radiation intensity. The complexity and uncertainty of these factors make PV power generation prediction a challenging task. Accurately predicting PV power generation is crucial to improving the efficiency and stability of PV power generation systems.
[0003] Currently, photovoltaic power generation forecasts primarily rely on historical and real-time meteorological data. Traditional methods typically employ statistical models (such as ARIMA and linear regression) or machine learning models (such as support vector machines and random forests) for forecasting. However, these methods are limited in their ability to handle multivariate and nonlinear relationships, particularly under dynamic cloud conditions and complex meteorological conditions, resulting in forecast accuracy that often fails to meet practical requirements.
[0004] In recent years, deep learning technologies have made significant progress in time series forecasting and image processing, providing new solutions for photovoltaic power forecasting. For example, convolutional neural networks (CNNs) and long short-term memory networks (LSTMs) have been widely used to process meteorological data and historical power data. Furthermore, architectures based on the Visual Transformer (ViT) have demonstrated excellent performance in image recognition tasks, providing new insights into cloud data processing. However, existing methods still suffer from shortcomings such as data quality issues, insufficient cloud data processing, difficulties in multivariate data fusion, and the trade-off between model complexity and accuracy. Therefore, effective short-term photovoltaic power forecasting is an important research topic. Summary of the Invention
[0005] Purpose of the Invention: To address the above technical issues, the present invention provides a method and device for short-term photovoltaic power prediction based on cloud evolution, thereby improving the accuracy and efficiency of photovoltaic power probabilistic prediction.
[0006] Technical solution: The present invention provides a method for short-term photovoltaic power prediction based on cloud evolution, comprising the following steps:
[0007] (1) Obtain real-time cloud map data, relevant meteorological data, historical photovoltaic power and solar radiation variable data of the area to be predicted;
[0008] (2) The cloud image data is preprocessed with local binary pattern (LBP) to extract texture features, and the three-dimensional fuzzy adaptive extended Kalman filter (3D-FAEKF) is used to reduce the noise of other variable data collected to make the data smoother;
[0009] (3) The pre-processed cloud image data is processed and trained using the RepViTS-YOLOX model based on the YOLO-improved VIT architecture applied to the CNN structure to obtain cloud image prediction data including cloud coverage and cloud change rate. The photovoltaic status is predicted by predicting the cloud cover's blocking of solar radiation;
[0010] (4) The relevant meteorological data are trained using a shallow lightweight model SAMformer that combines sharpness perception minimization and channel attention to obtain temperature and humidity prediction values;
[0011] (5) Calculate solar radiation based on the collected solar radiation variable data using the empirical formula Hargreaves equation;
[0012] (6) The cloud map prediction data, temperature prediction value, humidity prediction value, solar radiation and historical photovoltaic power are input into the time series lightweight adaptive network TSLANet for complex variable training, and finally the short-term prediction of photovoltaic power is obtained.
[0013] Furthermore, through satellite remote sensing and ground observation, visible light all-sky imaging technology is used to obtain real-time cloud map data of the predicted area, meteorological sensors are used to collect relevant meteorological data including relative humidity, temperature, wind speed and wind direction, and photovoltaic power stations are used to obtain historical photovoltaic power and solar radiation variable data including solar altitude angle, solar position, surface reflectivity and surface absorptivity.
[0014] Furthermore, in step (2), the texture feature extraction process of the cloud image data through local binary pattern LBP preprocessing is as follows:
[0015] (21) Convert the cloud image of the area to be predicted into a grayscale image to adapt to the LBP algorithm;
[0016] (22) The target pixel is used to define the area size according to the specific situation. The size of the neighborhood can be determined according to the specific situation. Common neighborhoods are 3×3, 5×5, or 8×8. The larger the neighborhood, the more the extracted texture features can reflect the texture information in a larger range, but the computational complexity will also increase accordingly.
[0017] (23) The grayscale value of the center pixel is used as the threshold, and the grayscale value of the area pixel is compared with it to generate a binary code, which is then converted into a decimal to obtain the LBP value of the center pixel;
[0018] (24) The frequency of LBP coding is accumulated throughout the image to construct a texture distribution histogram. Based on the histogram data, texture features are calculated and extracted to characterize the texture complexity and distribution characteristics. Common features include histogram mean, variance, energy, contrast, etc. These features can reflect the distribution and complexity of local textures in the image and provide a basis for further image analysis and processing.
[0019] Furthermore, in step (2), the 3D fuzzy adaptive extended Kalman filter 3D-FAEKF is used to perform noise reduction processing on the collected other variable data as follows:
[0020] (25) System Modeling:
[0021] Construct the state vector:
[0022] Combine other variable data into a state vector x k :
[0023]
[0024] Where, P pv,k is the photovoltaic power at the kth moment, H rh,k is the absolute humidity, T temp,k is the temperature, V wind,k is the wind speed, D wind,k It is the wind direction;
[0025] Establish the dynamic equation:
[0026] x k+1 =f(x k ,u k )+w k
[0027] Where u k is the control input, w k is the process noise;
[0028] Establish the observation equation:
[0029] z k =h(x k )+v k
[0030] Where z k is the observed value, v k is the observation noise;
[0031] (26) Fuzzy adaptive control:
[0032] Establish a fuzzy rule base and define fuzzy sets and fuzzy rules based on expert knowledge or experience; for example, define fuzzy sets such as "high", "medium", "low", etc., and establish corresponding fuzzy rules;
[0033] Dynamically adjust the covariance matrix Q of process noise and observation noise using fuzzy rules k and R k ; Dynamically adjust the Kalman gain K according to the system state and observation error k , to improve the filter's response capability;
[0034]
[0035] Among them, α and β are adaptive coefficients, and γ is the gain adjustment coefficient.
[0036] (27) Iterative prediction update:
[0037] Predict the current state based on the state and control input at the previous moment:
[0038]
[0039] Where, is the predicted state, u k is the control input;
[0040] Compute the covariance matrix of the predicted states:
[0041]
[0042] Where, F k is the Jacobian matrix of the state transfer function, Q k is the covariance matrix of the process noise;
[0043] Compute the expected observations based on the predicted states:
[0044]
[0045] Kalman gain calculation:
[0046]
[0047] Where H k is the Jacobian matrix of the observation model.
[0048] Combine the observed and predicted values to update the state estimate and perform a covariance update:
[0049]
[0050] (28) The fuzzy adaptive EKF sets the initial state estimate x0 and covariance matrix P0, the initial noise covariance matrix Q0 and R0. The loop is iterated, in which fuzzy logic control is added: according to the current state and observation error, the noise covariance matrix Q is adjusted using a fuzzy inference system. k and R kand the Kalman gain K k .
[0051] Furthermore, the implementation process of step (3) is as follows:
[0052] (31) Innovative backbone network: Based on YOLOX, the backbone network CSPDarknet of YOLOX is replaced with the RepViTS feature extraction network, which is composed of RepViT and SCConv modules; to enhance the model's target focus and feature extraction capabilities, and improve the model's detection effect on cloud cover blocking solar radiation.
[0053] (32) Feature processing: YOLOX inputs cloud map data, and the image is sliced through the Focus structure. Then, the feature map is divided into two parts through CSPDarknet. The main part passes through the residual module, and the other part is directly spliced with the main part. Finally, the feature map is converted into a fixed-size feature vector through spatial pyramid pooling (SPP).
[0054] The feature pyramid FPN and the path enhancement network PAN are connected and interacted to realize the exchange of semantic information of low-resolution feature maps and positioning information of high-resolution feature maps.
[0055] (33) The decoupling head design is adopted to process the two subtasks of target detection, positioning and classification, separately.
[0056] (34) Feature response visualization:
[0057] Select target feature layer: Select one or more key feature layers F from the RepViTS feature extraction network or the CsC-FPN feature fusion network. k , used to generate the heat map, select the feature layer F close to the detection head k , extract its characteristic response;
[0058] Global average pooling of multi-channel responses:
[0059]
[0060] Where H is the heat map of a single channel, m×m is the feature map f ki The spatial dimensions;
[0061] The generated heat map H is normalized so that its value range is [0,1].
[0062] Furthermore, the implementation process of step (4) is as follows:
[0063] SAMformer is a model based on the Transformer architecture that is suitable for time series forecasting tasks, such as photovoltaic power forecasting. The following are the steps to train SAMformer on variables such as temperature, humidity, wind speed, and wind direction:
[0064] (41) Data normalization:
[0065]
[0066] Where μ is the mean; σ is the standard deviation;
[0067] Split the data into training, validation, and test sets.
[0068] (42) Model architecture:
[0069] Convert the time series data into a high-dimensional vector through the embedding layer:
[0070] E=Embedding(X norm )
[0071] Where, Embedding(·) is the embedding layer that maps the input data to a high-dimensional space; E is the high-dimensional vector representation after embedding;
[0072] Position encoding adds position encoding to time series data to preserve time order information:
[0073] P=PositionalEncoding(E)
[0074] Where PositionalEncoding(·) is the position encoding function, which adds time sequence information to the time series data; P is the vector after adding position encoding, which retains the order information of the time series;
[0075] Use a multi-layer Transformer encoder to capture long-term dependencies in time series:
[0076] H = TransformerEncoder(P)
[0077] Where TransformerEncoder(·) is the Transformer encoder, which consists of a multi-head attention mechanism and a feedforward neural network; H is the hidden state output by the encoder, which captures the long-term dependencies in the time series;
[0078] The output layer outputs the prediction results through the fully connected layer:
[0079]
[0080] Where Linear(·) is a fully connected layer (linear layer) that maps the hidden state to the output space; is the predicted output of the model;
[0081] (43) Circuit training:
[0082] Loss function: Mean Squared Error (MSE) calculates the mean squared error between the predicted value and the true value:
[0083]
[0084] Where Y i is the true value, i.e. the actually observed photovoltaic power value; is the photovoltaic power value predicted by the model; N is the number of samples; MSE is the mean square error, which measures the difference between the predicted value and the true value;
[0085] Use Adam optimizer to update model parameters:
[0086]
[0087] Where η is the learning rate; is the gradient of the loss function with respect to the model parameters;
[0088] Training loop: forward propagation to compute model outputs Back propagation calculates the gradient of the loss function MSE to the model parameters Update the model parameters using the Adam optimizer; repeat the above steps until the model converges or the maximum number of training rounds is reached;
[0089] Perform validation set evaluation, use the validation set to calculate the model's performance metrics and adjust hyperparameters. Evaluate the final performance of the model on the test set to ensure that the model has good generalization ability.
[0090] (44) Prediction data: Input temperature, humidity, and historical photovoltaic power data of the new time step; output the predicted value of multivariate time series data of the future time step.
[0091] Furthermore, the implementation process of step (5) is as follows:
[0092] (51) Collect celestial and surface parameters: solar altitude angle α, solar position azimuth Φ, surface reflectivity ρ, surface absorptivity α s , Sun-Earth distance correction factor d r , extra-atmospheric solar radiation G0;
[0093] (52) Calculate solar radiation:
[0094] First calculate the solar radiation intensity G0 outside the atmosphere:
[0095] G0=G sc ·d r ·cos(θ z )
[0096] Where G sc is the solar constant, with a value of 1367W / m 2 ;d r is the correction factor for the distance between the Sun and the Earth;
[0097]
[0098] Where DOY is the day of the year; θ z is the solar zenith angle:
[0099] θ z =90°-α
[0100] Where α is the solar altitude angle.
[0101] Then calculate the surface solar radiation:
[0102] Calculate the surface solar radiation G using the Hargreaves equation:
[0103] G=G0·(1-ρ)·α s
[0104] Where G0 is the solar radiation intensity of the atmosphere; ρ is the surface reflectivity; α s is the surface absorption rate.
[0105] (52) Verification and calibration: Output the calculated actual received solar radiation and compare it with the measured data, and dynamically optimize and adjust the parameters.
[0106] Furthermore, the implementation process of step (6) is as follows:
[0107] The adaptive network TSLANet integrates adaptive spectrum blocks (ASBs) and interactive convolution blocks (ICBs), replacing Transformer self-attention with the lightweight ASBs. ASBs fuse the global spectrum (long-range dependencies) with local circular convolutions (short-term features) based on Fourier transforms and dynamically suppress high-frequency noise using adaptive thresholds. The model upgrades the feedforward network to interactive convolution blocks, enhancing its ability to parse complex patterns through a multi-scale convolution kernel mutual control mechanism. Combined with dataset self-supervised pre-training (such as mask reconstruction), this approach reduces computational complexity while maintaining Transformer scalability, improving the efficiency of extracting multi-scale temporal features and enhancing noise robustness.
[0108] The integrated adaptive spectrum block (ASB) uses Fourier analysis to convert time series data into the frequency domain and adopts adaptive thresholding to attenuate high-frequency noise and highlight relevant spectral features. After processing, IFFT reconstructs the time domain features. The interactive convolution block (ICB) uses different kernel sizes to interactively improve the features and balance the local and global time series feature extraction for time series analysis.
[0109] The time series data is divided into overlapping blocks, linearly projected into the embedding space, and a learnable positional encoding is added to obtain an enhanced block. After enhancing the feature representation through ASB, an interactive convolutional block (ICB) is proposed. The design of ICB uses a two-layer convolutional structure, including parallel convolutions with different kernel sizes to capture local features and longer-range dependencies, balancing spectral globality and temporal locality. The outputs of all blocks are spliced and then mapped to the photovoltaic power prediction value for the next T steps through a linear layer.
[0110] The present invention also provides a photovoltaic power short-term prediction system based on cloud evolution, comprising:
[0111] Multi-source data acquisition module, used to collect satellite remote sensing cloud images, meteorological sensor networks, and photovoltaic power station historical data;
[0112] The three-dimensional fuzzy adaptive extended Kalman filter noise reduction unit 3D-FAEKF performs noise reduction on other variable data collected;
[0113] The cloud cover prediction model based on RepViTS-YOLOX is used to detect the impact of cloud cover dynamic changes on solar radiation.
[0114] The SAMformer multivariate prediction module combines sharpness perception minimization and channel attention to accurately model the dynamic coupling relationship between temperature, humidity, and wind speed.
[0115] The Hargreaves equation solar radiation calculation unit is used to accurately estimate the actual solar radiation intensity received by the surface, providing key energy input parameters for the photovoltaic power prediction model.
[0116] The TSLANet fusion prediction network integrating adaptive spectrum block (ASB) and interactive convolution block (ICB) is used for real-time prediction of photovoltaic power.
[0117] The present invention also provides a computer device comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the programs are executed by the processors, the steps of the cloud evolution-based photovoltaic power short-term prediction method as described above are implemented.
[0118] Beneficial effects: Compared with the existing technology, the technical solution of the present invention has the following significant advantages: the meteorological data is denoised by a three-dimensional fuzzy adaptive extended Kalman filter (3D-FAEKF), and the cloud texture features are extracted by combining the local binary pattern (LBP), which significantly improves the data quality and feature characterization capabilities. The improved RepViTS-YOLOX model is used to process cloud data to accurately predict cloud cover, and the SAMformer model is used to efficiently process multivariate meteorological data and capture complex nonlinear relationships. The solar radiation intensity is dynamically calculated based on the Hargreaves equation, which further improves the accuracy of the input data. The time series lightweight adaptive network (TSLANet) is used to fuse multi-source data to achieve high robustness prediction in complex environments. While ensuring lightweight, the model has improved prediction accuracy compared to traditional methods. This method can optimize photovoltaic power generation planning and energy storage scheduling, reduce the abandonment rate, improve energy utilization, and provide a reliable basis for power grid scheduling and enhance power grid stability. Its versatility and scalability support multi-scenario applications and provide an innovative technical solution for the intelligent management of clean energy. BRIEF DESCRIPTION OF THE DRAWINGS
[0119] Figure 1 This is a schematic flow chart of a method for short-term photovoltaic power prediction based on cloud evolution according to the present invention;
[0120] Figure 2 This is a structural diagram of the RepViTS-YOLOX model of the application of the VIT architecture improved based on YOLO on the CNN structure in an embodiment of the present invention;
[0121] Figure 3 This is a flowchart of the shallow lightweight model SAMformer that combines sharpness perception minimization and channel attention in an embodiment of the present invention;
[0122] Figure 4 This is a flow chart of the time series lightweight adaptive network model TSLANet in an embodiment of the present invention. DETAILED DESCRIPTION
[0123] The technical solution of the present invention is further described in detail below through specific embodiments.
[0124] like Figure 1 As shown, the present invention provides a short-term photovoltaic power prediction method based on cloud evolution, comprising the following steps:
[0125] Step (1): Collect real-time cloud image data through satellite remote sensing and ground observation, obtain historical photovoltaic power through meteorological stations, and collect other variable data such as relative humidity, temperature, wind speed and direction, and solar radiation through meteorological sensors;
[0126] Step (2): To ensure data quality, the cloud image data is subjected to texture feature extraction through local binary pattern (LBP), and the other variable data collected are subjected to noise reduction processing using three-dimensional fuzzy adaptive extended Kalman filter (3D-FAEKF) to make the data smoother;
[0127] Step (3): The pre-processed cloud image data is processed and trained using the VIT architecture based on the YOLO improved CNN structure (RepViTS-YOLOX) model, which can predict the photovoltaic status by predicting the cloud cover's blocking of solar radiation;
[0128] Step (4): For other variable data such as temperature, humidity, wind speed, and wind direction, a shallow lightweight Transformer (SAMformer) model combining sharpness perception minimization and channel attention is used to train the predicted values of other variables;
[0129] Step (5): Calculate solar radiation using the empirical formula Hargreaves equation based on the collected data such as solar altitude angle, solar position, surface reflectivity, and surface absorptivity. Combine the calculated solar radiation data with other variable data to form a multivariate data set;
[0130] Step (6): The cloud map prediction, temperature and humidity prediction values measured by the multivariate data set, solar radiation and historical photovoltaic power are input into the time series lightweight adaptive network (TSLANet) for complex variable training, and the photovoltaic power generation prediction is finally obtained to improve the prediction accuracy of the overall model.
[0131] Specifically, the implementation process of step (1) is as follows:
[0132] Visible light all-sky imaging technology is used to obtain cloud evolution image data, and temperature, humidity, wind speed, wind direction and other meteorological data are obtained through temperature and humidity sensors, wind speed sensors, and wind direction sensors. Historical photovoltaic power, solar altitude angle, solar position, surface reflectivity, and surface absorptivity data are obtained through photovoltaic power stations.
[0133] Furthermore, the 3D-FAEKF implementation process in step (2) is as follows:
[0134] Carry out system modeling, define the state vector, and combine variables such as photovoltaic power, relative humidity, temperature, wind speed and wind direction into a state vector x k :
[0135]
[0136] Among them, P pv,k is the photovoltaic power at the kth moment, H rh,k is the absolute humidity at time k, Ttemp,k is the temperature at time k, V wind,k is the wind speed at time k, D wind,k is the wind direction at time k.
[0137] Establish the nonlinear dynamic equation of the state vector changing with time:
[0138] x k+1 =f(x k ,u k )+w k
[0139] Among them, u k is the control input, w k is the process noise.
[0140] Establish a nonlinear relationship between the observation value and the state vector:
[0141] z k =h(x k )+v k
[0142] Among them, z k is the observed value, v k is the observation noise.
[0143] Establish a fuzzy rule base based on expert knowledge or experience to handle uncertainty. For example, define fuzzy sets such as "high", "medium", "low", etc., and establish corresponding fuzzy rules;
[0144] Use fuzzy rules to reason about input data and generate fuzzy outputs. These outputs will be used to adjust the parameters of the filter;
[0145] Dynamically adjust the covariance matrix Q of process noise and observation noise according to actual observation data k and R k ; Dynamically adjust the Kalman gain K according to the system state and observation error k , in order to improve the filter's response capability.
[0146]
[0147] Among them, α and β are adaptive coefficients, and γ is the gain adjustment coefficient.
[0148] Predict the current state based on the state and control input at the previous moment:
[0149]
[0150] in, is the predicted state.
[0151] Compute the covariance matrix of the predicted states:
[0152]
[0153] Among them, F k is the Jacobian matrix of the state transfer function.
[0154] Compute the expected observations based on the predicted states:
[0155]
[0156] Kalman gain calculation:
[0157]
[0158] Among them, H k is the Jacobian matrix of the observation model.
[0159] Combine the observed and predicted values to update the state estimate and perform a covariance update:
[0160]
[0161] The fuzzy adaptive EKF sets the initial state estimate x0 and covariance matrix P0, the initial noise covariance matrix Q0 and R0. The loop is iterated, in which fuzzy logic control is added: according to the current state and observation error, the noise covariance matrix Q is adjusted using a fuzzy inference system. k and R k and the Kalman gain K k .
[0162] Specifically, the local binary pattern implementation process in step (2) is as follows:
[0163] Determine the region of interest and grayscale image: Identify the cloud image region where you want to extract texture features. If you want to extract features from the entire image, use the entire image as the region of interest. Convert the image to grayscale. If the cloud image is in color, convert it to grayscale for subsequent processing. Because the LBP algorithm operates primarily on grayscale values, grayscale images simplify computation.
[0164] Define pixel neighborhoods: For each pixel, define a neighborhood. The size of the neighborhood can be determined based on the specific situation; common neighborhoods include 3×3, 5×5, or 8×8. A larger neighborhood allows the extracted texture features to better reflect texture information within a larger area, but this also increases the computational complexity.
[0165] Calculating Local Binary Patterns: Original LBP operator: When the window is 3×3, the grayscale value of the center pixel of the window is used as the threshold and compared with the grayscale values of the adjacent 8 pixels. If the surrounding pixel value is greater than the center pixel value, the corresponding position is marked as 1, otherwise it is marked as 0, resulting in an 8-bit binary number, which is then converted to decimal to obtain the LBP value of the center pixel.
[0166] Circular LBP operator: When extracting texture features at different scales, the circular LBP operator can be used. Within a circular neighborhood with a radius of R, there are P sampling points. The grayscale value of the center pixel is compared with the grayscale values of these sampling points. Similarly, a binary number is obtained, which is then converted to decimal as the LBP value. When a sampling point is located at a pixel boundary, bilinear interpolation is used to calculate the pixel value at that point.
[0167] Statistical Local Binary Patterns: For each pixel in the image, the histogram of its local binary pattern is calculated. The histogram shows the frequency distribution of different binary codes in the image and can be used as a representation of texture features.
[0168] Feature extraction: Based on histogram data, a series of texture features can be calculated, including the histogram mean, variance, energy, contrast, etc. These features can reflect the distribution and complexity of local textures in the image, providing a basis for further image analysis and processing.
[0169] like Figure 2 As shown, the implementation process of step (3) is as follows:
[0170] The application of the improved VIT architecture based on YOLO in the CNN structure is based on YOLOX, replacing YOLOX's backbone network CSPDarknet with the RepViTS feature extraction network. The network consists of RepViT and SCConv modules to enhance the model's target attention and feature extraction capabilities, and improve the model's detection effect on cloud cover blocking solar radiation.
[0171] YOLOX inputs cloud image data, slices the image through the Focus structure, and then passes it through CSPDarknet's unique CSP structure, which divides the feature map into two parts. The backbone part passes through the residual module, and the other part is directly spliced with the backbone part. This improves feature reuse efficiency and enhances gradient propagation. Finally, spatial pyramid pooling (SPP) is used to convert the feature map into a fixed-size feature vector.
[0172] The feature fusion network consists of a feature pyramid (FPN) and a path augmentation network (PAN). The FPN constructs a feature pyramid by combining feature layers layer by layer and then connects it laterally to the PAN, enabling the exchange of semantic information from low-resolution feature maps with positioning information from high-resolution feature maps. The detection head uses a decoupled head design to separate the localization and classification subtasks of object detection.
[0173] RepViT re-examines the modular, micro-, and macro-design of standard lightweight CNNs, integrating them with the effective ViTs architecture from eight perspectives to gradually improve the performance of CNNs in vision tasks. Compared to the traditional ViTs architecture, RepViT achieves superior performance and lower resource usage on mobile devices through efficient design and structural reparameterization. The introduction of RepViT narrows the gap between lightweight CNNs and lightweight ViTs and highlights the potential of ViTs for mobile applications.
[0174] Feature response visualization:
[0175] Select target feature layer: Select one or more key feature layers F from the RepViTS feature extraction network or the CsC-FPN feature fusion network. k , used to generate heat maps, feature layers close to the detection head are selected because these feature layers contain richer semantic information and target location information.
[0176] Extract feature response: For the selected feature layer F k , extract its feature response. The feature response is a multi-channel feature map F k ={f k1 ,f k2 ,...,f kn}, each channel f ki Indicates the strength of the model's response to a specific feature.
[0177] Global average pooling: In order to convert the multi-channel feature map into a single-channel heat map, the feature response can be globally averaged. The feature map of each channel is spatially averaged to obtain a channel response value r ki , and then add the response values of all channels to obtain a single-channel heat map H:
[0178]
[0179] in, m×m is the feature map f ki space dimensions.
[0180] Normalization: The generated heat map H is normalized so that its value range is [0,1].
[0181] like Figure 3 As shown, the implementation process of step (4) is as follows:
[0182] SAMformer is a model based on the Transformer architecture that is suitable for time series forecasting tasks, such as photovoltaic power forecasting. The following are the steps to train SAMformer on variables such as temperature, humidity, wind speed, and wind direction:
[0183] Data normalization: Normalize the data to make them in the same dimension
[0184]
[0185] Where μ is the mean and σ is the standard deviation.
[0186] Split the data into training, validation, and test sets.
[0187] Model Architecture:
[0188] The time series data is converted into high-dimensional vectors through the embedding layer.
[0189] E=Embedding(X norm )
[0190] Among them, Embedding(·) is the embedding layer, which maps the input data to a high-dimensional space; E is the high-dimensional vector representation after embedding.
[0191] Position encoding adds position encoding to time series data to preserve time order information.
[0192] P=PositionalEncoding(E)
[0193] Among them, PositionalEncoding(·) is the position encoding function, which adds time sequence information to time series data; P is the vector after adding position encoding, which retains the order information of the time series.
[0194] Use a multi-layer Transformer encoder to capture long-term dependencies in time series.
[0195] H = TransformerEncoder(P)
[0196] Among them, TransformerEncoder(·) is the Transformer encoder, which consists of a multi-head attention mechanism and a feedforward neural network; H is the hidden state output by the encoder, which captures the long-term dependencies in the time series.
[0197] The output layer outputs the prediction results through the fully connected layer:
[0198]
[0199] Among them, Linear(·) is a fully connected layer (linear layer) that maps the hidden state to the output space; is the predicted output of the model.
[0200] Loss function: Mean Squared Error (MSE) calculates the mean squared error between the predicted value and the true value.
[0201]
[0202] Among them, Y i is the true value, i.e. the actually observed photovoltaic power value; is the photovoltaic power value predicted by the model; N is the number of samples; MSE is the mean square error, which measures the difference between the predicted value and the true value.
[0203] The Adam optimizer is used to update the model parameters.
[0204]
[0205] Where η is the learning rate; is the gradient of the loss function with respect to the model parameters.
[0206] Training process: forward propagation calculates model output Back propagation calculates the gradient of the loss function MSE to the model parameters Update the model parameters using the Adam optimizer; repeat the above steps until the model converges or the maximum number of training rounds is reached.
[0207] Perform validation set evaluation, use the validation set to calculate the model's performance metrics and adjust hyperparameters. Evaluate the final performance of the model on the test set to ensure that the model has good generalization ability.
[0208] Finally, the trained model is used to predict new data. The input is the temperature, humidity, wind speed, wind direction and other data of the new time step; the output is the predicted value of the multivariate time series data in the future time step.
[0209] The SAMformer model can effectively train and predict multivariate time series data such as temperature, humidity, wind speed, and wind direction. It can capture long-term dependencies in time series, process time series data in parallel, improve training efficiency, flexibly handle multiple input variables, and be extended to other time series prediction tasks.
[0210] Specifically, the implementation process of step (5) is as follows:
[0211] Collect the solar altitude angle (α), solar position azimuth (Φ), surface reflectivity (ρ), surface absorptivity (α s ), Sun-Earth distance correction factor (d r ), extra-atmospheric solar radiation (G0).
[0212] Calculate solar radiation:
[0213] First calculate the solar radiation intensity G0 outside the atmosphere:
[0214] G0=G sc ·d r ·cos(θ z )
[0215] Among them, G sc is the solar constant, with a value of 1367W / m 2 ;d r is the correction factor for the distance between the Sun and the Earth.
[0216]
[0217] Where DOY is the day of the year; θ z is the solar zenith angle:
[0218] θ z =90°-α
[0219] Where α is the sun's altitude angle.
[0220] Then calculate the surface solar radiation:
[0221] Calculate the surface solar radiation G using the Hargreaves equation:
[0222] G=G0·(1-ρ)·α s
[0223] Among them, G0 is the solar radiation intensity of the atmosphere; ρ is the surface reflectivity; α s is the surface absorption rate.
[0224] Output the calculated actual received solar radiation and verify it, and adjust the parameters and formulas according to the actual situation.
[0225] like Figure 4 As shown, the implementation process of step (6) is as follows:
[0226] TSLANet is an efficient time series modeling architecture that replaces Transformer self-attention with a lightweight Adaptive Spectral Block (ASB). ASB fuses the global spectrum (long-range dependencies) with local circular convolution (short-term features) based on Fourier transforms and dynamically suppresses high-frequency noise using adaptive thresholds. The model upgrades the feedforward network to interactive convolution blocks, enhancing the ability to parse complex patterns through a multi-scale convolution kernel mutual control mechanism. Combined with dataset self-supervised pre-training (such as mask reconstruction), it reduces computational complexity while maintaining the scalability of the Transformer, improving the efficiency of extracting multi-scale time series features and noise robustness.
[0227] The model integrates two new components, the Adaptive Spectral Block (ASB) and the Interactive Convolution Block (ICB), which form a single layer that can be extended to multiple layers. The ASB uses Fourier analysis to transform the time series data into the frequency domain, where we employ adaptive thresholding to attenuate high-frequency noise and highlight relevant spectral features. After processing, the IFFT reconstructs the time domain features, now with reduced noise and enhanced representation. The ICB is a streamlined convolution block that interactively refines the features using different kernel sizes, improving adaptability to temporal dynamics in the time series. Together, these components form a cohesive structure that balances local and global time series feature extraction for time series analysis.
[0228] Split the sequence into M overlapping blocks of length p, with overlapping step s = [p / 2]
[0229]
[0230] Project each block into the embedding space through a linear mapping
[0231]
[0232] Adding learnable positional encoding Get Enhanced Fast:
[0233] SPE i =P i '+E i
[0234] Adaptive Spectral Block (ASB) employs Fourier domain processing. This block is designed to learn spatial information through global circular convolution operations. In addition, it provides adaptive local filters to isolate the noisy high-frequency components of any time series data.
[0235] Fast Fourier Transform. Given a discrete time series x[n], we obtain its frequency domain representation x[k] by performing FFT along the spatial dimension. Given S PE , expressed as:
[0236] F i =F(SPE i )∈C d×L′
[0237] Here, F[·] represents a one-dimensional FFT operation, and L' represents the length of the transformed frequency domain sequence, which may differ from L depending on the FFT implementation and the nature of the time series data. Each channel of the time series is transformed independently, resulting in a comprehensive frequency domain representation F that encapsulates the spectral characteristics of the original time series across all channels.
[0238] An adaptive local filter model dynamically adjusts the filtering level based on the characteristics of the dataset and removes these high-frequency noise components. This is crucial when dealing with non-stationary data, as the spectrum may vary over time. The proposed filter adaptively sets the appropriate frequency threshold for each specific time series data.
[0239] Generate a binary mask through a learnable threshold θ to filter out high-frequency noise:
[0240] M filter =I(P>θ)∈{0,1} d×L′
[0241] Where I(·) is an exponential function, and the frequency domain characteristics after filtering out high-frequency noise are:
[0242] F filtered =F i ⊙M filter
[0243] After adaptively filtering the frequency domain data, the model employs two sets of learnable filters; a global filter learned from the original frequency domain data F and a local filter learned from the adaptively filtered data. Let WG and WL be the learnable global and local filters, respectively. The application of these filters is expressed as:
[0244] F G =W G ⊙F
[0245] F L =W L ⊙F filtered
[0246] These filtered features are integrated to capture comprehensive spectral details, i.e., F integrated =F G +F L .
[0247] Inverse Fourier Transform,To convert the integrated frequency domain data back to the time domain, we apply the Inverse Fast Fourier Transform (IFFT).,The resulting time domain signal S':
[0248] S′=F -1 [F integrated ]∈R C×p′
[0249] Among them, F -1 (·) represents the inverse FFT operation. IFFT ensures that the enhanced features are consistent with the original data structure of the input time series.
[0250] After enhancing feature representation through ASB, the interactive convolutional block (ICB) is proposed, which utilizes a two-layer convolutional structure. The design of ICB includes parallel convolutions with different kernel sizes to capture local features and longer-range dependencies. Specifically, the first convolutional layer aims to capture fine-grained, local patterns in the data using a smaller kernel. In contrast, the second layer aims to identify broader, longer-range dependencies with a larger kernel. ICB is designed so that the output of each layer modulates the feature extraction of another layer. Element-wise multiplication encourages interaction between features extracted at different scales, making it possible to better model complex relationships.
[0251] Two 1D convolutions (Conv1 and Conv2) with different kernel sizes are used to extract features:
[0252] A1=φ(Conv1(S′ i ))☉Conv2(S′ i )
[0253] A2=φ(Conv2(S′ i ))☉Convl(S′ i )
[0254] Where φ(·) is the GELU activation function;
[0255] The two convolution results are added together and further features are extracted through the third convolution layer Conv3:
[0256] O ICB =Conv3(A1+A2)
[0257] The output OICB representation is the enhanced feature prepared for the last layer in the network, which is represented by a customizable linear layer according to the task.
[0258] The outputs of multiple blocks are concatenated and mapped to the predicted value through a linear layer, and the outputs of all blocks are concatenated:
[0259]
[0260] Then it is mapped to the photovoltaic power prediction value in the future T steps through the linear layer:
[0261]
[0262] Among them, W o ∈R T×(d·M·) is a learnable weight matrix; b0∈R T For paranoid items.
[0263] This paper introduces a three-dimensional fuzzy adaptive extended Kalman filter (3D-FAEKF) to reduce the noise of meteorological data, uses local binary patterns (LBP) to extract cloud texture features, and combines the improved RepViTS-YOLOX model and SAMformer model to process cloud data and multivariate data, respectively. In addition, solar radiation is calculated using the Hargreaves equation, and all data are input into the time series lightweight adaptive network (TSLANet) for training, ultimately achieving high-precision photovoltaic power generation prediction. This method has significant advantages in data quality, feature extraction, multivariate fusion, and model performance, providing a new technical path for photovoltaic power generation prediction.
[0264] Based on the same technical concept as the method embodiment, the present invention also provides a photovoltaic power short-term prediction system based on cloud evolution, comprising:
[0265] Multi-source data acquisition module, used to collect satellite remote sensing cloud images, meteorological sensor networks, and photovoltaic power station historical data;
[0266] The three-dimensional fuzzy adaptive extended Kalman filter noise reduction unit 3D-FAEKF performs noise reduction on other variable data collected;
[0267] The cloud cover prediction model based on RepViTS-YOLOX is used to detect the impact of cloud cover dynamic changes on solar radiation.
[0268] The SAMformer multivariate prediction module combines sharpness perception minimization and channel attention to accurately model the dynamic coupling relationship between temperature, humidity, and wind speed.
[0269] The Hargreaves equation solar radiation calculation unit is used to accurately estimate the actual solar radiation intensity received by the surface, providing key energy input parameters for the photovoltaic power prediction model.
[0270] The TSLANet fusion prediction network integrating adaptive spectrum block (ASB) and interactive convolution block (ICB) is used for real-time prediction of photovoltaic power.
[0271] The present invention also provides a computer device comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the programs are executed by the processors, the steps of the cloud evolution-based photovoltaic power short-term prediction method as described above are implemented.
Claims
1. A photovoltaic power short-term prediction method based on cloud evolution, characterized in that: The following steps are involved: (1) Obtain real-time cloud map data, relevant meteorological data, historical photovoltaic power and solar radiation variable data of the area to be predicted; (2) The cloud image data is preprocessed with local binary pattern (LBP) to extract texture features, and the three-dimensional fuzzy adaptive extended Kalman filter (3D-FAEKF) is used to reduce the noise of other variable data collected to make the data smoother; (3) The pre-processed cloud map data is processed and trained using the RepViTS-YOLOX model based on the VIT architecture improved by YOLO on the CNN structure to obtain cloud map prediction data including cloud coverage and cloud change rate; (4) The relevant meteorological data are trained using a shallow lightweight model SAMformer that combines sharpness perception minimization and channel attention to obtain temperature and humidity prediction values; (5) Calculate solar radiation based on the collected solar radiation variable data using the empirical formula Hargreaves equation; (6) The cloud map prediction data, temperature prediction value, humidity prediction value, solar radiation and historical photovoltaic power are input into the time series lightweight adaptive network TSLANet for complex variable training, and finally the short-term prediction of photovoltaic power is obtained.
2. The method according to claim 1, characterized in that The implementation process of step (1) is as follows: Visible light all-sky imaging technology is used through satellite remote sensing and ground observations to obtain real-time cloud map data of the predicted area. Meteorological sensors are used to collect relevant meteorological data including relative humidity, temperature, wind speed and wind direction. Photovoltaic power stations are used to obtain historical photovoltaic power and solar radiation variable data including solar altitude angle, solar position, surface reflectivity and surface absorptivity.
3. The method according to claim 1, characterized in that In step (2), the texture feature extraction process of the cloud image data through local binary pattern LBP preprocessing is as follows: (21) Convert the cloud image of the area to be predicted into a grayscale image to adapt to the LBP algorithm; (22) Define the area size based on the target pixel according to the specific situation; (23) The grayscale value of the center pixel is used as the threshold, and the grayscale value of the area pixel is compared with it to generate a binary code, which is then converted into a decimal to obtain the LBP value of the center pixel; (24) The frequency of LBP coding is accumulated throughout the entire image to construct a texture distribution histogram. Based on the histogram data, texture features are calculated and extracted to characterize the texture complexity and distribution characteristics.
4. The method according to claim 2, characterized in that In step (2), the 3D fuzzy adaptive extended Kalman filter 3D-FAEKF is used to perform noise reduction processing on the collected other variable data as follows: (25) System Modeling: Construct the state vector: Combine other variable data into a state vector x k : Where, P pv,k is the photovoltaic power at time k, H rh,k is the absolute humidity at time k, T temp,k is the temperature at time k, V wind,k is the wind speed at time k, D wind,k is the wind direction at time k; Establish the dynamic equation: x k+1 =f(x k ,u k )+w k Where u k is the control input, w k is the process noise; Establish the observation equation: z k =h(x k )+v k Where z k is the observed value, v k is the observation noise; (26) Fuzzy adaptive control: Establish a fuzzy rule base and define fuzzy sets and fuzzy rules based on expert knowledge or experience; Dynamically adjust the covariance matrix Q of process noise using fuzzy rules k and the covariance matrix R of the observation noise k ; Dynamically adjust the Kalman gain K according to the system state and observation error k , to improve the filter's response capability; (27) Iterative prediction update: Predict the current state based on the state and control input at the previous moment: Where, is the predicted state, u k is the control input; Compute the covariance matrix of the predicted states: Where, F k is the Jacobian matrix of the state transfer function, Q k is the covariance matrix of the process noise; Compute the expected observations based on the predicted states: Kalman gain calculation: Where H k is the Jacobian matrix of the observation model; Combine the observed and predicted values to update the state estimate and perform a covariance update: (28) According to the current state and observation error, the process noise covariance matrix Q is adjusted using a fuzzy inference system. k and the observation noise covariance matrix R k and the Kalman gain K k .
5. The method according to claim 2, characterized in that The implementation process of step (3) is as follows: (31) Innovative backbone network: Based on YOLOX, the backbone network CSPDarknet of YOLOX is replaced with the RepViTS feature extraction network, which consists of RepViT and SCConv modules; (32) Feature processing: YOLOX inputs cloud image data, and the image is sliced through the Focus structure. Then, the feature map is divided into two parts through CSPDarknet. The main part passes through the residual module, and the other part is directly spliced with the main part. Finally, the feature map is converted into a fixed-size feature vector through spatial pyramid pooling (SPP). The feature pyramid FPN and the path enhancement network PAN are connected and interacted to realize the exchange of semantic information of low-resolution feature maps and positioning information of high-resolution feature maps. (33) The decoupling head design is adopted to process the two subtasks of target detection, positioning and classification, separately; (34) Feature response visualization: Select target feature layer: Select one or more key feature layers F from the RepViTS feature extraction network or the CsC-FPN feature fusion network. k , used to generate the heat map, select the feature layer F close to the detection head k , extract its characteristic response; Global average pooling of multi-channel responses: Where H is the heat map of a single channel, m×m is the feature map f ki The spatial dimensions; The generated heat map H is normalized so that its value range is [0,1].
6. The method according to claim 1, characterized in that The implementation process of step (4) is as follows: (41) Data normalization: Where μ is the mean; σ is the standard deviation; Split the data into training, validation, and test sets; (42) Model architecture: Convert the time series data into a high-dimensional vector through the embedding layer: E=Embedding(X norm ) Where, Embedding(·) is the embedding layer that maps the input data to a high-dimensional space; E is the high-dimensional vector representation after embedding; Position encoding adds position encoding to time series data to preserve time order information: P=PositionalEncoding(E) Where PositionalEncoding(·) is the position encoding function, which adds time sequence information to the time series data; P is the vector after adding position encoding, which retains the order information of the time series; Use a multi-layer Transformer encoder to capture long-term dependencies in time series: H = TransformerEncoder(P) Where TransformerEncoder(·) is the Transformer encoder, which consists of a multi-head attention mechanism and a feedforward neural network; H is the hidden state output by the encoder, which captures the long-term dependencies in the time series; The output layer outputs the prediction results through the fully connected layer: Where Linear(·) is a fully connected layer (linear layer) that maps the hidden state to the output space. is the predicted output of the model; (43) Circuit training: Loss function: Mean Squared Error (MSE) calculates the mean squared error between the predicted value and the true value: Where Y i is the true value, i.e. the actually observed photovoltaic power value; is the photovoltaic power value predicted by the model; N is the number of samples; MSE is the mean square error, which measures the difference between the predicted value and the true value; Use Adam optimizer to update model parameters: Where η is the learning rate; is the gradient of the loss function with respect to the model parameters; Training loop: forward propagation to compute model outputs Back propagation calculates the gradient of the loss function MSE to the model parameters Update the model parameters using the Adam optimizer; repeat the above steps until the model converges or the maximum number of training rounds is reached; Perform validation set evaluation, use the validation set to calculate the model's performance indicators and adjust hyperparameters; evaluate the final performance of the model on the test set to ensure that the model has good generalization ability; (44) Prediction data: Input temperature, humidity, and historical photovoltaic power data of the new time step; output the predicted value of multivariate time series data of the future time step.
7. The method according to claim 1, characterized in that The implementation process of step (5) is as follows: (51) Collect celestial and surface parameters: solar altitude angle α, solar position azimuth Φ, surface reflectivity ρ, surface absorptivity α s , Sun-Earth distance correction factor d r , extra-atmospheric solar radiation G0; (52) Calculate solar radiation: First calculate the solar radiation intensity G0 outside the atmosphere: G0=G sc ·d r ·cos(θ z ) Where G sc is the solar constant, with a value of 1367W / m 2 ;d r is the correction factor for the distance between the Sun and the Earth; Where DOY is the day of the year; θ z is the solar zenith angle: i z =90°-a Where α is the solar altitude angle; Then calculate the surface solar radiation: Use the Hargreaves equation to calculate the surface solar radiation G: G=G0·(1-ρ)·α s Where G0 is the solar radiation intensity of the atmosphere; ρ is the surface reflectivity; α s is the surface absorption rate; (52) Verification and calibration: Output the calculated actual received solar radiation and compare it with the measured data, and dynamically optimize and adjust the parameters.
8. The method according to claim 1, characterized in that The implementation process of step (6) is as follows: The adaptive network TSLANet includes an integrated adaptive spectrum block (ASB) and an interactive convolution block (ICB). The integrated adaptive spectrum block (ASB) uses Fourier analysis to convert time series data into the frequency domain, and adopts an adaptive threshold to attenuate high-frequency noise and highlight related spectral features. After processing, IFFT reconstructs the time domain features. The interactive convolution block (ICB) uses different kernel sizes to interactively improve features, balancing local and global time series feature extraction for time series analysis; the time series data is divided into overlapping blocks, linearly projected into the embedding space, and a learnable position code is added to obtain an enhanced block. After enhancing the feature representation through the adaptive spectrum block (ASB), an interactive convolution block (ICB) is proposed. The design of the ICB with a two-layer convolution structure includes parallel convolutions with different kernel sizes to capture local features and longer-range dependencies, balancing spectral globality and temporal locality. The outputs of all blocks are spliced and then mapped to the photovoltaic power prediction value for the next T steps through a linear layer.
9. A photovoltaic power short-term prediction system based on cloud evolution, characterized in that: include: Multi-source data acquisition module, used to collect satellite remote sensing cloud images, meteorological sensor networks, and photovoltaic power station historical data; The three-dimensional fuzzy adaptive extended Kalman filter noise reduction unit 3D-FAEKF performs noise reduction on other variable data collected; The cloud cover prediction model based on RepViTS-YOLOX is used to detect the impact of cloud cover dynamic changes on solar radiation. The SAMformer multivariate prediction module combines sharpness perception minimization and channel attention to accurately model the dynamic coupling relationship between temperature, humidity, and wind speed. The Hargreaves equation solar radiation calculation unit is used to accurately estimate the actual solar radiation intensity received by the surface, providing key energy input parameters for photovoltaic power prediction models; The TSLANet fusion prediction network integrating adaptive spectrum block (ASB) and interactive convolution block (ICB) is used for real-time prediction of photovoltaic power.
10. A computer device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and when the programs are executed by the processors, the steps of the photovoltaic power short-term prediction method based on cloud evolution according to any one of claims 1 to 8 are implemented.