Intra-hour photovoltaic power prediction method

By extracting multi-dimensional features from all-sky images and using dual-channel LSTM and Attention mechanisms for multi-modal fusion, the problem of insufficient feature extraction and fusion difficulties in the prior art is solved, and efficient and accurate photovoltaic power prediction is achieved, which is suitable for edge devices.

CN120510397APending Publication Date: 2025-08-19CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510608147.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

In the prediction of photovoltaic power within hours, insufficient feature extraction and difficulty in fusion of multi-dimensional feature lead to low prediction accuracy, and deep learning methods have large calculation volume and high hardware requirements, making it difficult to effectively deploy on edge devices.

Method used

By extracting color features, time features, sun position and cloud motion information from all-sky images, a dual-channel LSTM neural network is built, and multi-modal fusion is used to optimize hyperparameters by using step-by-step training and intelligent optimization algorithm IVYA to achieve efficient extraction and fusion of features.

Benefits of technology

It realizes high-precision and low-computation photovoltaic power prediction, which is suitable for edge computing devices, and can accurately track power fluctuations in complex weather conditions, reduce computing costs, and improve prediction performance and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510397A_ABST
    Figure CN120510397A_ABST
Patent Text Reader

Abstract

The invention provides an intra-hour photovoltaic power prediction method. The intra-hour photovoltaic power prediction method comprises the steps of extracting color features from an all-sky image sequence, extracting time features based on historical power generation data, determining sun position information, calculating cloud motion information and cutting local cloud cluster features according to the sun position and cloud motion; constructing a dual-channel LSTM neural network, and encoding the features; performing interactive fusion on the encoded different modal features by using a multi-modal fusion device based on an Attention mechanism; a step-by-step training mechanism is adopted, a dual-channel LSTM model is pre-trained, then a cross-modal attention module is combined to perform re-training, and an intelligent optimization algorithm IVYA is used to optimize network model hyper-parameters; and obtaining features of a to-be-predicted all-sky image sequence according to the feature extraction method, and inputting the features into the trained network model to obtain an intra-hour photovoltaic power prediction result. According to the method, high-precision prediction of the photovoltaic power can be realized, and the method has good robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of photovoltaic power generation technology, and in particular to a method for predicting photovoltaic power within an hour based on all-sky image derivation and using deep learning and multimodal fusion technology. Background Art

[0002] With the continuous depletion of fossil energy and the increasing severity of global climate change, renewable energy has garnered widespread attention as an effective solution to these problems. In recent years, the proportion of renewable energy in the global energy system has continued to increase. Solar energy, with its clean, safe, and inexhaustible characteristics, has occupied a significant proportion of the renewable energy system. However, photovoltaic power generation systems are susceptible to weather conditions, especially cloud cover. Complex and variable cloud cover can cause significant fluctuations in power generation, severely impacting the stability of photovoltaic power generation systems and, to a certain extent, limiting their further development. Effectively predicting photovoltaic power generation can help energy decision-makers adjust strategies, prevent energy waste, and reduce the need for energy storage facilities. This has significant implications for improving the global climate and resolving the global energy crisis.

[0003] All-sky imagery effectively reflects cloud conditions and is crucial for accurate photovoltaic power prediction. In recent years, image processing technology has advanced, resulting in the development of numerous photovoltaic power prediction methods based on all-sky imagery. These methods can be broadly categorized into two types: empirical methods, which extract important descriptors from all-sky imagery, such as cloud motion information, cloud cluster matching, and cloud cover index, for prediction. Examples include the methods described in the paper "Systematic Review on Ground-Based Cloud Tracking Methods for Photovoltaics Nowcasting." Deep learning methods, on the other hand, use images directly as input and automatically extract key features from all-sky imagery using CNNs or VITs for prediction. Examples include "Improving Ultra-Short-Term Photovoltaic Power Forecasting Using a Novel Sky-Image-Based Framework Considering Spatial-Temporal Feature Interaction." While empirical methods typically require low computing power and are simple to implement, they often lack comprehensive features, resulting in poor adaptability and suboptimal prediction performance. Deep learning methods are currently the mainstream approach, capable of automatically learning important features from massive amounts of feature data, often offering greater adaptability and higher prediction accuracy. However, the training process of deep learning methods often requires significant computing power, especially when using images as direct input. This consumes significant computing power and places high demands on hardware.

[0004] Based on the above situation, some experts have proposed extracting secondary features from images for neural network training, such as in "Deep Learning to Forecast Solar Irradiance Using a Six-Month UTSA SkyImager Dataset," to avoid the extensive computational overhead associated with using images directly as input. To fully capture sky conditions, it is often necessary to extract a large number of features from full-sky images. These features contain diverse and complex information, making them difficult for deep learning networks to directly learn. Multimodal fusion technology can effectively address this issue. By effectively interacting information from different modalities, it generates features that can be captured by deep learning networks, helping the network effectively learn full-sky imagery and achieve better predictions. However, when training a network, different types of network modules often require different hyperparameter settings. For example, the LSTM network used for feature capture and the Multi-Head Attention network used for multimodal fusion often require different learning rates. Directly training these two network types may result in insufficient network training, resulting in low prediction accuracy and poor adaptability.

[0005] Currently, there are few achievements in extracting secondary scalar features from all-sky images to train deep learning networks for intra-hour photovoltaic power prediction. This may be due to the following reasons: the extraction of secondary scalar features is not perfect; the extracted multimodal features are difficult to effectively fuse, making it difficult for the network to capture the features; and the network becomes more complex after the introduction of the multimodal fusion part, requiring a more effective training strategy. Summary of the Invention

[0006] In response to the problems of insufficient feature extraction and difficulty in multi-dimensional feature fusion in existing intra-hour photovoltaic power prediction, which lead to low prediction accuracy, the present invention provides a method for predicting intra-hour photovoltaic power based on all-sky image derivation, using deep learning and multimodal fusion technology, to solve the problems of insufficient secondary feature representation and low prediction accuracy in intra-hour photovoltaic power prediction.

[0007] The technical solution of the present invention:

[0008] In a first aspect, the present invention provides a method for predicting photovoltaic power within an hour, comprising:

[0009] Feature extraction stage: extract color features from the full sky image sequence, extract time features based on historical power generation data, determine the sun's position information, calculate cloud motion information, and crop local cloud cluster features based on the sun's position and cloud motion;

[0010] Feature encoding stage: A dual-channel LSTM neural network is constructed to encode features reflecting past photovoltaic power changes and features reflecting future photovoltaic power changes respectively;

[0011] Feature fusion stage: A multimodal fuser based on the Attention mechanism is used to interactively fuse the encoded features of different modalities.

[0012] Model training phase: A step-by-step training mechanism is used to pre-train the dual-channel LSTM model, then re-train it in conjunction with the cross-modal attention module. The intelligent optimization algorithm IVYA is used to optimize the network model hyperparameters.

[0013] Prediction stage: The features of the full-sky image sequence to be predicted are obtained according to the above feature extraction method, and the features are input into the trained network model to obtain the photovoltaic power prediction results within the hour.

[0014] Furthermore, the color feature extraction is to extract three types of scalar features, namely mean, variance, and entropy, from the four matrices of red, green, blue, and red-to-blue ratio of the full sky image, respectively, and a total of 12 scalar features are obtained. The calculation formula is:

[0015]

[0016] Among them, ν i Represents the elements in the matrix, N b refers to 100 evenly divided brightness intervals, and p i represents the frequency of the vector appearing in the i-th interval, μ, σ, e are the mean, variance, and entropy respectively.

[0017] Furthermore, the time feature extraction uses the month, day, hour, and minute information when the full sky image is taken as the time feature, and the formula is:

[0018]

[0019] Where Month, Day, Hour, and Minute correspond to the month, day, hour, and minute when the full-sky image was taken, respectively, and the time features are spliced with the color features.

[0020] Furthermore, the step of determining the sun position information is: removing fisheye distortion and grayscale processing on the full sky image, sliding the slider on the image, calculating the average brightness of the pixels in each slider position, and determining the center position of the block with the highest brightness as the sun position pixel.

[0021] Furthermore, the cloud motion information is calculated by using an optical flow map to calculate the direction and speed of cloud motion in continuous full-sky images.

[0022] Furthermore, the extraction of local cloud features is to crop out cloud image blocks that may block the sun in the future based on the sun's position and cloud motion information, and extract 12 scalar features reflecting the cloud cover conditions from the image blocks using the same method as color feature extraction.

[0023] Furthermore, in the dual-channel LSTM neural network, one channel processes features reflecting past photovoltaic power changes, including historical power generation data, time features, and global color features, while the other channel processes local cloud cluster features and related features reflecting future photovoltaic power changes based on cloud motion prediction.

[0024] Furthermore, the multimodal fuser based on the Attention mechanism realizes interactive fusion of features by calculating the attention weights between different modal features, and generates a fused feature vector containing rich information.

[0025] Furthermore, in the step-by-step training mechanism, the two LSTM models are trained separately in the pre-training stage and their hyperparameters are optimized. In the re-training stage, the pre-trained LSTM weights are loaded into the entire network including the cross-modal attention module for joint training.

[0026] Furthermore, the intelligent optimization algorithm IVYA optimizes the hyperparameters of the dual-channel LSTM in the pre-training stage, and optimizes the hyperparameters of the entire network including the attention mechanism in the retraining stage. The hyperparameters include the learning rate, the number of hidden layer neurons, and the training batch size.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] The present invention provides a method for predicting photovoltaic power within an hour based on multidimensional feature extraction and multimodal fusion. Through innovative feature engineering, network architecture, and training strategies, it effectively solves the problems of insufficient feature representation, low fusion efficiency, and insufficient prediction accuracy in the prior art. The specific beneficial effects are as follows:

[0029] (1) Multi-dimensional feature extraction comprehensively characterizes the relationship between sky conditions and power generation

[0030] Combination of global and local features: By extracting color features (mean, variance, entropy of red, green, blue and red-to-blue ratio matrices, a total of 12 scalar features) from the full sky image, it effectively reflects the overall brightness, uniformity and complexity of the cloud distribution; at the same time, based on the sun's position and cloud motion information, local cloud image blocks are cropped and local scalar features of the same dimension are extracted to accurately capture key cloud information that may block the sun in the future.

[0031] Temporal features enhance temporal correlation: By integrating the time features of the shooting time, such as month, day, hour, and minute, and combining them with historical power generation data, it can effectively characterize the periodic law of photovoltaic power changes over time under clear weather conditions, making up for the insufficient prediction of single image features for stable lighting scenes.

[0032] Sun Positioning and Cloud Motion Analysis: For complex scenarios where the fisheye camera is not positioned vertically, the sun's position is determined through image grayscale conversion and sliding window search. Cloud motion information is calculated using an optical flow algorithm to achieve real-time tracking of dynamic clouds blocking the sun, providing key input for short-term power fluctuation prediction.

[0033] (2) Multimodal fusion architecture improves feature interaction efficiency and prediction accuracy

[0034] Dual-channel LSTM temporal encoding: Independent LSTM encoders are constructed to process historical data reflecting past power changes (including time and global color features) and cloud motion characteristics reflecting future changes (including local cloud features). Utilizing LSTM's ability to model long sequence dependencies, the encoder captures power fluctuation patterns and cloud motion trends at different time scales, avoiding information ambiguity caused by the mixing of single-channel features.

[0035] Deep fusion of cross-modal Attention mechanism: Through an attention-based multimodal fuser, the attention weights between different modal features are calculated, dynamic interaction between historical power time series features and future cloud motion features is achieved, and a fused feature vector containing spatiotemporal correlations is generated. This significantly enhances the model's ability to capture power mutations in complex weather conditions (such as sudden cloudy weather and short periods of rain).

[0036] A step-by-step training strategy optimizes compatibility: A "pre-training-retraining" step-by-step training mechanism is adopted. The parameters of the dual-channel LSTM are first independently optimized to stabilize the temporal encoding capability. Then, cross-modal Attention modules are combined for joint training to resolve compatibility issues between LSTM and Attention caused by differences in gradient propagation. This ensures coordinated optimization of various network modules and improves overall prediction robustness.

[0037] (3) Intelligent optimization and lightweight design balance performance and practicality

[0038] IVYA algorithm hyperparameter optimization: Using the new intelligent optimization algorithm IVYA, the optimal hyperparameters of the dual-channel LSTM (such as the number of hidden layer neurons and learning rate) are automatically searched during the pre-training phase. During the retraining phase, the complete network parameters, including the Attention module, are globally optimized. This eliminates the need for manual parameter adjustment and fully unleashes the model's predictive potential, maintaining high accuracy, especially in scenarios with uneven data distribution.

[0039] Lightweight feature engineering reduces computing power requirements: Different from deep learning methods that directly input raw images, this invention extracts secondary scalar features (non-pixel-level raw data) as network input, which greatly reduces the amount of calculation and has low requirements for hardware equipment. It can be efficiently deployed on edge computing devices or small and medium-sized servers, and is adapted to the real-time prediction needs of step-by-step photovoltaic power stations.

[0040] Adaptability to a wide range of scenarios: Verified by ablation experiments, the present invention demonstrates better prediction performance than the baseline method under clear, cloudy, and complex weather conditions. In particular, it can accurately track power fluctuation trends in rapidly changing cloud cover scenarios, providing a reliable basis for grid scheduling and energy storage configuration.

[0041] (4) Technical solution innovation and industrial application value

[0042] End-to-end prediction framework: This framework automates the entire process from image feature extraction to power prediction, eliminating the need for manual feature screening. This framework achieves deep coupling between features and models through a data-driven approach, reducing reliance on domain expert knowledge and improving the applicability of the method.

[0043] Cost-effectiveness: While ensuring high accuracy, lightweight design and intelligent optimization significantly reduce computing costs and training time. This meets the dual needs of photovoltaic power plants for prediction timeliness and economy, and has broad engineering application prospects.

[0044] In summary, the present invention breaks through the existing accuracy bottleneck of intra-hour photovoltaic power prediction through the deep characterization of multi-dimensional features, the efficient fusion of multimodal information and intelligent optimization strategies, and provides key technical support for improving the photovoltaic energy absorption capacity and ensuring the stable operation of the power grid. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the following detailed description, they are used to explain the present invention but should not be construed as limiting the present invention. In the accompanying drawings:

[0046] Figure 1 It is the overall implementation roadmap of the present invention;

[0047] Figure 2 The solar pixel positioning roadmap of the present invention;

[0048] Figure 3 It is a schematic diagram of local feature extraction of the present invention;

[0049] Figure 4 Schematic diagram of the dual-channel LSTM of the present invention;

[0050] Figure 5 Schematic diagram of the multimodal fusion device based on the Attention mechanism of the present invention;

[0051] Figure 6 It is a schematic diagram of the step-by-step training of the present invention;

[0052] Figure 7 The data distribution of the experimental content of the present invention;

[0053] Figure 8 This is a comparison chart of the prediction performance of the present invention under different feature inputs and with and without the feature fusion step;

[0054] Figure 9 This is the actual rainy day prediction effect diagram of the present invention;

[0055] Figure 10 This is the actual prediction effect diagram of a sunny day according to the present invention. DETAILED DESCRIPTION

[0056] The following describes the specific embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention.

[0057] See also Figure 1 The present invention provides a method for predicting photovoltaic power within hours using deep learning technology and multimodal fusion technology, comprising the following steps:

[0058] Step S1: extracting color features from the full sky image sequence to reflect the cloud conditions in the sky. Step S1 specifically includes:

[0059] The extracted color features include three types of scalar features: mean, variance, and entropy. A total of 12 scalar features are extracted from the four matrices of red, green, blue, and red-to-blue ratio of the image. The process is performed using the following formula:

[0060]

[0061] where ν i Represents the elements in the matrix, N b refers to 100 evenly divided brightness intervals, and p i represents the frequency of the vector appearing in the i-th interval, μ, σ, e are the mean, variance, and entropy respectively.

[0062] The present invention achieves multi-dimensional quantification of cloud optical properties by extracting 12-dimensional color features (mean, variance, and entropy of the red / green / blue / red-blue ratio matrix) from the full sky image:

[0063] Multi-dimensional representation: The mean reflects the overall brightness, the variance reflects the uniformity of pixel distribution, and the entropy quantifies the complexity of the cloud layer. Combined with the four-matrix spectral information, it comprehensively depicts the cloud state in sunny, cloudy, rainy and other weather conditions;

[0064] Lightweight and efficient: Converts images into 12-dimensional scalar features, compressing data by over 99%, significantly reducing computational load and meeting the low-latency requirements of real-time prediction.

[0065] Standardization and robustness: Features are calculated using a unified formula (divided into 100 uniform brightness intervals), eliminating the influence of subjective thresholds, improving the comparability of data across devices and time periods, and enhancing the model's adaptability to complex weather conditions.

[0066] Step S2: Using historical power generation data, a time feature extraction method is developed and combined with color features, as follows:

[0067]

[0068] Where Month, Day, Hour, and Minute correspond to the month, day, hour, and minute when the full-sky image was taken, respectively. Extracting this type of feature can effectively reflect the power generation under clear skies.

[0069] This method extracts the month, day, hour, and minute of the full-sky image capture time as time features and combines them with color features to effectively enhance the model's ability to capture the periodicity of photovoltaic power. The specific advantages are as follows:

[0070] Accurately characterize the periodicity of light: Directly linking the solar altitude angle with the duration of sunshine, such as hourly characteristics reflecting changes in light intensity during the day (peak at noon, decay in the morning and evening), and monthly characteristics reflecting seasonal differences in light intensity (long days in summer, short days in winter). This enables the model to accurately predict the regular fluctuations in power over time in clear and cloudless scenes, reducing the basic prediction error.

[0071] Complementary enhancement of multimodal features: "Time-space" information coupling is formed with color features (reflecting real-time cloud conditions). For example, when color features tend to be stable in clear weather, time features can dominate power prediction, preventing single image features from being interfered with by instantaneous noise; in cloudy weather, the two work together to correct predictions and improve adaptability to complex scenarios.

[0072] Lightweight time series modeling support: Simple numerical features (month / day / hour / minute) replace complex time series processing. Without the need for additional time series analysis modules, this model can provide clear time scale information for time series networks such as LSTM, reducing model learning costs and improving training efficiency and real-time prediction.

[0073] Step S3: In order to adapt to the situation where the fisheye camera is not placed vertically and the tilt angle is unknown, which leads to the failure of the fisheye camera back projection method to find the sun position, a sun position positioning method based only on the image is proposed. The process can be done by the attached Figure 2 express:

[0074] Step S3 specifically includes: removing fisheye distortion and grayscale processing on the full sky image, sliding a slider on the image, calculating the average brightness of the pixels within each slider position, and determining the center position of the block with the highest brightness as the sun position pixel.

[0075] The pure image domain sun position positioning method proposed in this paper effectively solves the positioning problem when the fisheye camera is not installed vertically and the angle is unknown:

[0076] Strong environmental adaptability: No camera installation parameters are required. The system locates the sun directly from the image through dedistortion, grayscale conversion, and a slider to search for the area with the highest brightness. It is suitable for complex outdoor installation scenarios and the renovation of old equipment.

[0077] Efficient and accurate: Single-frame processing takes a short time, focusing on brightness features to resist cloud interference, accurately capturing the core area of the sun, providing a reliable spatial benchmark for subsequent cloud motion analysis, and improving the accuracy of local feature extraction.

[0078] Low-cost and easy to deploy: No external information such as geographic coordinates is required. The pure image algorithm reduces system dependence, is compatible with low-cost monitoring equipment, and significantly enhances the engineering practicality of the prediction system.

[0079] Step S4: This step can be performed by Figure 3 The calculation of cloud motion information is to use the optical flow map to calculate the direction and speed of cloud motion in continuous full-sky images.

[0080] The extraction of local cloud features is to crop out cloud image blocks that may block the sun in the future based on the sun's position and cloud motion information, and then extract 12 scalar features reflecting the cloud cover from the image blocks using the same method as color feature extraction.

[0081] This paper uses a local cloud feature extraction method that combines optical flow maps with the sun's position to accurately predict dynamic cloud occlusion. The specific advantages are as follows:

[0082] Dynamic cloud tracking: Utilizes optical flow maps to calculate cloud movement direction and speed, capturing dynamic cloud changes in real time. This provides key timing information for short-term power fluctuation forecasts and effectively responds to complex weather conditions such as sudden cloudy weather.

[0083] Targeted feature extraction: Based on the sun's position and cloud motion information, local cloud image blocks are cropped to focus on key areas that may block the sun in the future. 12-dimensional scalar features consistent with global features are extracted to enhance sensitivity to local cloud cover changes and improve the accuracy of power mutation prediction.

[0084] Spatiotemporal information coupling: Combining the time series characteristics of cloud motion (speed and direction) with the spatial characteristics of local cloud clusters (brightness and complexity) provides high-value input for the subsequent dual-channel LSTM and Attention mechanisms, significantly enhancing the model's ability to learn the "cloud-sun-power" relationship and reducing prediction errors in complex weather conditions.

[0085] Step S5: Construct a dual-channel LSTM neural network to encode the features reflecting past photovoltaic power changes and the features reflecting future photovoltaic power changes. The framework is shown in the attached figure. Figure 4 In the dual-channel LSTM neural network, one channel processes features reflecting past PV power changes, including historical power generation data, time features, and global color features, while the other channel processes local cloud cluster features and features related to future PV power changes based on cloud motion prediction.

[0086] This paper uses a dual-channel LSTM neural network architecture to achieve accurate encoding of features in different time dimensions. The specific advantages are as follows:

[0087] Feature decoupling enhances time series modeling: Historical power generation data, time features, global color features, and future cloud movement and local features are processed in separate channels to avoid mixing of information at different time scales. This allows the LSTM to focus on learning long-term trends in past power changes (such as daily / seasonal periodicity) and short-term dynamics of future cloud obstruction (such as minute-level cloud movement), improving the targeted nature of time series feature extraction.

[0088] Multimodal parallel processing efficiency: Independent channels encode two types of features, "past-future", respectively, retaining the power fluctuation patterns of historical data and key information such as the spatial position and movement speed of future cloud clusters. This provides structured input for subsequent cross-modal Attention fusion, enhances the model's ability to learn the interaction between "historical patterns and real-time cloud conditions", and reduces prediction errors in complex weather conditions.

[0089] Lightweight network design: Optimize LSTM parameters for different feature types, ensuring feature encoding quality while reducing redundant computations, and meeting the dual requirements of model efficiency and accuracy for intra-hour predictions.

[0090] Step S6: Develop a multimodal fusion method based on the Attention mechanism. This method solves the compatibility issue between LSTM and Attention through a step-by-step training sequence.

[0091] Step S6 includes:

[0092] Step S61: Construct a multimodal fusion engine based on the Attention mechanism. The multimodal fusion engine based on the Attention mechanism calculates the attention weights between different modal features to achieve interactive fusion of features and generate a fusion feature vector containing rich information. The framework is shown in the attached figure. Figure 5 shown.

[0093] Step S62: In the step-by-step training mechanism, the two LSTM models are trained and their hyperparameters are optimized in the pre-training phase. In the re-training phase, the pre-trained LSTM weights are loaded into the entire network including the cross-modal attention module for joint training. By training different network parts in steps, the compatibility problem of the network is solved. This process can be done by the attached Figure 6 express.

[0094] The multimodal fusion method based on the Attention mechanism developed by this invention has significant advantages:

[0095] Deep fusion of multimodal features: The multimodal fuser based on the Attention mechanism can calculate the attention weights between different modal features, effectively realizing interactive feature fusion. The generated fused feature vector contains rich information, allowing the model to more accurately capture the correlation between different features and improve the accuracy of photovoltaic power prediction.

[0096] Solving the compatibility problem: Using a step-by-step training mechanism, we first pre-trained two LSTM models and optimized their hyperparameters. We then loaded the pre-trained weights into a network containing a cross-modal attention module for joint training. This successfully solved the compatibility issue between LSTM and Attention, ensuring that all parts of the network work together.

[0097] Improve training efficiency and effectiveness: The step-by-step training method avoids problems such as gradient instability that may occur during overall training, allowing the network to converge faster and improving training efficiency. At the same time, the optimized model can better learn the characteristic patterns in the data, enhancing the stability and reliability of the prediction.

[0098] Step S7: Use the intelligent optimization algorithm IVYA to optimize the hyperparameters of the network model during training to further improve the prediction accuracy.

[0099] The intelligent optimization algorithm IVYA optimizes the hyperparameters of the dual-channel LSTM in the pre-training phase and optimizes the hyperparameters of the entire network including the attention mechanism in the retraining phase. The hyperparameters include the learning rate, the number of hidden layer neurons, and the training batch size.

[0100] This paper uses the intelligent optimization algorithm IVYA to optimize network hyperparameters in stages, significantly improving the model's prediction potential and generalization ability:

[0101] Automated hyperparameter tuning: Instead of manual trial and error, the IVYA algorithm automatically searches for the optimal parameters of the dual-channel LSTM (such as the number of hidden layer neurons and learning rate) during the pre-training phase. During the retraining phase, it globally optimizes the entire network including the Attention module, unleashing the model's maximum predictive power and avoiding the limitations of empirical parameter tuning.

[0102] Phased precision optimization: Dynamically adjust hyperparameters (such as batch size adaptation to feature dimension changes) to meet the different needs of pre-training (LSTM independent optimization) and retraining (multimodal fusion network overall optimization), solving the problem of training inefficiency caused by parameter coupling in complex network structures, and improving prediction accuracy by 10%-15%.

[0103] Enhanced model robustness: By optimizing key parameters such as the learning rate, the model's sensitivity to uneven data distribution and noise interference is reduced. Especially in complex weather scenarios such as cloudy weather, the stability of prediction results is improved by more than 20%, providing reliable parameter optimization support for practical engineering applications.

[0104] Step S8: The feature engineering method from steps S1 to S4 is used to extract rich feature representations and input them into the trained neural network for prediction. The photovoltaic power prediction results obtained in the prediction stage are multiple power prediction values divided into certain time intervals within the next hour.

[0105] The prediction stage of this invention achieves fine-grained and accurate prediction of photovoltaic power within an hour through standardized feature extraction and efficient model reasoning. The specific advantages are as follows:

[0106] Fine-grained time coverage: Outputs multiple power forecast values divided into fixed time intervals (such as 5 minutes and 10 minutes) within the next hour, providing high-density data support for grid scheduling and energy storage system control, facilitating real-time adjustment of energy distribution strategies, and reducing the impact of power fluctuations on grid stability.

[0107] Efficient inference adapted to real-time scenarios: Based on lightweight scalar features extracted in the early stage, prediction does not require processing of original image data. The low input dimension and small computational effort meet the stringent real-time requirements for prediction within an hour and can be seamlessly integrated into the photovoltaic power station monitoring system.

[0108] High-precision prediction supports decision-making: Relying on early multi-dimensional feature fusion and model optimization, the prediction results can accurately capture short-term power fluctuation trends, especially for sudden power surges / drops caused by sudden cloud cover, and issue reliable early warnings, helping to reduce redundant configuration of energy storage equipment and lower operation and maintenance costs.

[0109] Standardized processes ensure generalization capabilities: Feature extraction and model input formats are strictly aligned during the training phase to avoid errors caused by data format differences during prediction. This supports stable prediction performance across regions and seasons, significantly improving the method's universal applicability in engineering applications.

[0110] Although the flowcharts of the above embodiments show the steps in sequence as indicated by the arrows, these steps are not necessarily executed strictly in the order shown by the arrows. Unless expressly provided herein, there is no fixed restriction on the order in which the steps are executed and they can be adjusted according to actual needs. In addition, at least some of the steps in the flowcharts may include multiple sub-steps or stages, which are not necessarily completed at the same time, but can be executed at different time points. The order in which they are executed is not necessarily linear, but can be executed alternately or in parallel with other steps or parts of the steps within them.

[0111] Simulation and Experiment

[0112] (1) Simulation conditions

[0113] The simulation of the present invention was performed on a computer equipped with an NVIDIA4060Ti graphics processing unit (GPU) with 8GB of memory and the pytorch2.1 framework environment. The images and data used in the experiment were derived from the full-sky images taken by the Stanford University rooftop photovoltaic equipment and the Stanford University fisheye camera. The data contains photovoltaic power data for the three years of 2017, 2018, and 2019. Figure 7 The PV power distribution of the dataset is shown.

[0114] (2) Experimental content

[0115] In order to verify the practicality and effectiveness of the photovoltaic power prediction method proposed in this invention, the data from 2017 and 2018 are used as the training set, and the data from the whole year of 2019 is used as the validation set.

[0116] The IVYA algorithm automatically adjusts parameters during both pre-training and re-training of the network model. During the pre-training phase, data from two modalities is input, concatenated using the Concat method, and the dual-channel LSTM is trained. During the re-training phase, the pre-trained LSTM model hyperparameters are loaded into the LSTM of the re-trained model, and the entire network model is then trained.

[0117] In addition, in order to verify the effectiveness of the method proposed in this invention, we designed a set of ablation experiments, and verified the effectiveness of feature engineering and cross-modal attention of the invention and the accurate prediction effect of the invention by continuously increasing the feature types used for prediction and multimodal fusion methods. Three indicators are used to evaluate the prediction effect, namely MAE, RMSE, R 2 The results are as attached Figure 8 shown.

[0118] MAE (Mean Absolute Error): It directly reflects the average magnitude of the prediction error. The smaller the value, the closer the predicted value is to the true value.

[0119] RMSE (Root Mean Square Error): Measures the degree of dispersion of the deviation between the predicted value and the true value. The smaller the value, the higher the prediction accuracy.

[0120] R 2 (Coefficient of Determination): Its value ranges from 0 to 1. The closer it is to 1, the better the model fits the data, the higher the proportion of true value changes that can be explained, and the better the prediction performance.

[0121] Analysis Attachment Figure 8 It can be found that with the addition of multidimensional features, the prediction performance of the prediction method proposed in this invention is gradually improved, and after the introduction of the multimodal fusion method, its MAE, RMSE, R 2 It is optimal within 30 minutes and has the best performance.

[0122] Attachment Figure 9 、 10 The actual prediction effect of the present invention under different weather conditions is demonstrated. Figure 9 、 10 It is easy to see that this method can effectively predict future photovoltaic power generation on sunny days. Even in cloudy days with complex and changeable weather conditions, the prediction method proposed in the present invention can still effectively track the changing trend of photovoltaic power. Compared with the commonly used benchmark method Pers, the method proposed in the present invention can more effectively and accurately predict photovoltaic power generation and has better prediction performance.

[0123] The present invention provides a method for predicting photovoltaic power within an hour based on multidimensional feature extraction and multimodal fusion. Through innovative feature engineering, network architecture, and training strategies, it effectively solves the problems of insufficient feature representation, low fusion efficiency, and insufficient prediction accuracy in the prior art. The specific beneficial effects are as follows:

[0124] (1) Multi-dimensional feature extraction comprehensively characterizes the relationship between sky conditions and power generation

[0125] Combination of global and local features: By extracting color features (mean, variance, entropy of red, green, blue and red-to-blue ratio matrices, a total of 12 scalar features) from the full sky image, it effectively reflects the overall brightness, uniformity and complexity of the cloud distribution; at the same time, based on the sun's position and cloud motion information, local cloud image blocks are cropped and local scalar features of the same dimension are extracted to accurately capture key cloud information that may block the sun in the future.

[0126] Temporal features enhance temporal correlation: By integrating the time features of the shooting time, such as month, day, hour, and minute, and combining them with historical power generation data, it can effectively characterize the periodic law of photovoltaic power changes over time under clear weather conditions, making up for the insufficient prediction of single image features for stable lighting scenes.

[0127] Sun Positioning and Cloud Motion Analysis: For complex scenarios where the fisheye camera is not positioned vertically, the sun's position is determined through image grayscale conversion and sliding window search. Cloud motion information is calculated using an optical flow algorithm to achieve real-time tracking of dynamic clouds blocking the sun, providing key input for short-term power fluctuation prediction.

[0128] (2) Multimodal fusion architecture improves feature interaction efficiency and prediction accuracy

[0129] Dual-channel LSTM temporal encoding: Independent LSTM encoders are constructed to process historical data reflecting past power changes (including time and global color features) and cloud motion characteristics reflecting future changes (including local cloud features). Utilizing LSTM's ability to model long sequence dependencies, the encoder captures power fluctuation patterns and cloud motion trends at different time scales, avoiding information ambiguity caused by the mixing of single-channel features.

[0130] Deep fusion of cross-modal Attention mechanism: Through an attention-based multimodal fuser, the attention weights between different modal features are calculated, dynamic interaction between historical power time series features and future cloud motion features is achieved, and a fused feature vector containing spatiotemporal correlations is generated. This significantly enhances the model's ability to capture power mutations in complex weather conditions (such as sudden cloudy weather and short periods of rain).

[0131] A step-by-step training strategy optimizes compatibility: A distributed "pre-training-retraining" training mechanism is used to independently optimize the parameters of the dual-channel LSTM to stabilize temporal encoding capabilities. This is then combined with cross-modal Attention module training to address compatibility issues between LSTM and Attention caused by differences in gradient propagation. This ensures coordinated optimization of all network modules and improves overall prediction robustness.

[0132] (3) Intelligent optimization and lightweight design balance performance and practicality

[0133] IVYA algorithm hyperparameter optimization: Using the new intelligent optimization algorithm IVYA, the optimal hyperparameters of the dual-channel LSTM (such as the number of hidden layer neurons and learning rate) are automatically searched during the pre-training phase. During the retraining phase, the complete network parameters, including the Attention module, are globally optimized. This eliminates the need for manual parameter adjustment and fully unleashes the model's predictive potential, maintaining high accuracy, especially in scenarios with uneven data distribution.

[0134] Lightweight feature engineering reduces computing power requirements: Different from deep learning methods that directly input raw images, this invention extracts secondary scalar features (non-pixel-level raw data) as network input, which greatly reduces the amount of calculation and has low requirements for hardware equipment. It can be efficiently deployed on edge computing devices or small and medium-sized servers, and is adapted to the real-time prediction needs of step-by-step photovoltaic power stations.

[0135] Wide range of scene adaptability: Through ablation experiments, the proposed method shows better prediction performance than the baseline method in sunny, cloudy and complex weather conditions (such as Figure 8 ), especially for fast-changing cloud cover scenes (such as the attached Figure 9 Rainy day case), it can accurately track power fluctuation trends and provide a reliable basis for grid dispatching and energy storage configuration.

[0136] (4) Technical solution innovation and industrial application value

[0137] End-to-end prediction framework: This framework automates the entire process from image feature extraction to power prediction, eliminating the need for manual feature screening. This framework achieves deep coupling between features and models through a data-driven approach, reducing reliance on domain expert knowledge and improving the applicability of the method.

[0138] Cost-effectiveness: While ensuring high accuracy, lightweight design and intelligent optimization significantly reduce computing costs and training time. This meets the dual needs of photovoltaic power plants for prediction timeliness and economy, and has broad engineering application prospects.

[0139] In summary, the present invention breaks through the existing accuracy bottleneck of intra-hour photovoltaic power prediction through the deep characterization of multi-dimensional features, the efficient fusion of multimodal information and intelligent optimization strategies, and provides key technical support for improving the photovoltaic energy absorption capacity and ensuring the stable operation of the power grid.

[0140] Those skilled in the art will understand that all or part of the process steps for implementing the above-mentioned method embodiments can be controlled by a computer program to execute the relevant hardware. The computer program can be stored in a non-volatile computer-readable storage medium and implement the process steps of the above-mentioned method embodiments when it is run. In the various embodiments provided in this application, the memory, database, or other storage medium involved may include at least one of non-volatile memory and volatile memory. Non-volatile memory includes, but is not limited to, read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), and graphene memory. Volatile memory may include random access memory (RAM) or external cache, wherein RAM may take the form of static random access memory (SRAM) or dynamic random access memory (DRAM). In addition, the databases involved in the embodiments of this application may include relational databases and non-relational databases, wherein non-relational databases may include blockchain-based distributed databases, etc. In terms of processors, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices, data processing units based on quantum computing, etc. can be used, and the specific implementation is not limited to the above types.

[0141] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for predicting photovoltaic power within an hour, characterized in that: include: Feature extraction stage: extract color features from the full sky image sequence, extract time features based on historical power generation data, determine the sun's position information, calculate cloud motion information, and crop local cloud cluster features based on the sun's position and cloud motion; Feature encoding stage: A dual-channel LSTM neural network is constructed to encode features reflecting past photovoltaic power changes and features reflecting future photovoltaic power changes respectively; Feature fusion stage: A multimodal fuser based on the Attention mechanism is used to interactively fuse the encoded features of different modalities. Model training phase: A step-by-step training mechanism is used to pre-train the dual-channel LSTM model, then re-train it in conjunction with the cross-modal attention module. The intelligent optimization algorithm IVYA is used to optimize the network model hyperparameters. Prediction stage: The features of the full-sky image sequence to be predicted are obtained according to the above feature extraction method, and the features are input into the trained network model to obtain the photovoltaic power prediction results within the hour.

2. The method for predicting photovoltaic power within an hour according to claim 1, characterized in that: The color feature extraction is to extract three types of scalar features, namely mean, variance and entropy, from the red, green, blue and red-to-blue ratio matrices of the full sky image, respectively, and obtain a total of 12 scalar features. The calculation formula is: Among them, ν i Represents the elements in the matrix, N b refers to 100 evenly divided brightness intervals, and p i represents the frequency of the vector appearing in the i-th interval, μ, σ, e are the mean, variance, and entropy respectively.

3. The method for predicting photovoltaic power within an hour according to claim 1, characterized in that: The time feature extraction uses the month, day, hour, and minute information when the full sky image is taken as the time feature, and the formula is: Where Month, Day, Hour, and Minute correspond to the month, day, hour, and minute when the full-sky image was taken, respectively, and the time features are spliced with the color features.

4. The method for predicting photovoltaic power within an hour according to claim 1, characterized in that: The steps for determining the sun position information are: removing fisheye distortion and grayscale processing on the full sky image, sliding a slider on the image, calculating the average brightness of the pixels within each slider position, and determining the center position of the block with the highest brightness as the sun position pixel.

5. The method for predicting photovoltaic power within an hour according to claim 1, characterized in that: The cloud motion information is calculated by using an optical flow map to calculate the direction and speed of cloud motion in continuous full-sky images.

6. The method for predicting photovoltaic power within an hour according to claim 1 or 5, characterized in that: The extraction of local cloud features is to crop cloud image blocks that may block the sun in the future based on the sun's position and cloud motion information, and extract 12 scalar features reflecting the cloud cover from the image blocks using the same method as color feature extraction.

7. The method for predicting photovoltaic power within an hour according to claim 1, characterized in that: In the dual-channel LSTM neural network, one channel processes features reflecting past photovoltaic power changes, including historical power generation data, time features, and global color features, while the other channel processes local cloud cluster features and related features reflecting future photovoltaic power changes based on cloud motion prediction.

8. The method for predicting photovoltaic power within an hour according to claim 1, characterized in that: The multimodal fuser based on the Attention mechanism realizes interactive fusion of features by calculating the attention weights between different modal features, and generates a fused feature vector containing rich information.

9. The method for predicting photovoltaic power within an hour according to claim 1 or 8, characterized in that: In the step-by-step training mechanism, the two LSTM models are trained separately in the pre-training phase and their hyperparameters are optimized. In the re-training phase, the pre-trained LSTM weights are loaded into the entire network including the cross-modal attention module for joint training.

10. The method for predicting photovoltaic power within an hour according to claim 1, characterized in that: The intelligent optimization algorithm IVYA optimizes the hyperparameters of the dual-channel LSTM in the pre-training stage and optimizes the hyperparameters of the entire network including the attention mechanism in the retraining stage. The hyperparameters include the learning rate, the number of hidden layer neurons, and the training batch size.

Citation Information

Cited By

  • Photovoltaic module fault early warning method and system based on time sequence and space double-current model

    CN120806661A

  • Cloud layer motion trail analysis method and system based on Transform neural network

    CN121053170A