Short-term irradiance prediction method and system based on multi-modal data later feature fusion
By combining convolutional neural networks and long and short-term memory networks, extracting and fusing multimodal data features, the shortcomings of the existing technology in short-term irradiance prediction accuracy and adaptability are solved, and a more efficient and reliable prediction effect is achieved.
Patent Information
- Application Number
- CN202510040844.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-06-06
AI Technical Summary
The existing solar irradiance prediction methods have insufficient accuracy and adaptability, especially in short-term prediction and complex weather conditions, it is difficult to effectively extract and fuse multi-source information, resulting in a decrease in prediction accuracy.
The late feature fusion method based on multimodal data is adopted, and the features in image and text data are extracted through the combination of convolutional neural network (CNN) and long and short-term memory network (LSTM), and the feature weighting and fusion are performed through the gating mechanism to finally output the short-term irradiance prediction value.
It improves the accuracy and robustness of short-term irradiance prediction, can quickly adapt to weather and environment changes, and provides more reliable scheduling and optimization solutions for photovoltaic power generation systems.
Smart Images

Figure CN120105359A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of solar irradiance prediction, and in particular to a short-term irradiance prediction method and system based on late feature fusion of multimodal data. Background Art
[0002] As a clean and renewable energy source, solar energy has the significant advantages of abundant resources and environmental friendliness. Especially under the current global trend of responding to climate change and reducing carbon emissions, the importance of solar power generation has become increasingly prominent and has become a key force in promoting energy transformation. However, the volatility and randomness of solar irradiance make photovoltaic power generation systems face significant uncertainty in operation, which brings huge challenges to the stable operation of the system, efficient scheduling, and coordinated operation with the power grid. Accurate solar irradiance prediction is one of the important technical supports for promoting the transformation of photovoltaic power generation from an auxiliary energy source to a dominant energy source.
[0003] Global horizontal irradiance (GHI) prediction is crucial for the efficient operation of photovoltaic power generation systems and grid dispatching. Common prediction methods include physical models, statistical methods, and machine learning methods. Physical models are based on the principles of atmospheric physics and use models such as numerical weather prediction (NWP) to make medium- and long-term predictions, but they require high computing resources and have low accuracy under complex terrain. Statistical methods, such as ANN, SVR, RF, and LSTM, can capture nonlinear relationships and have good results for short- and medium-term predictions, but there is still room for improvement in accuracy. However, the above methods are still lacking in prediction accuracy. In order to improve prediction accuracy, fusion methods combine physical models, statistical models, and machine learning models to give full play to their respective advantages. For example, after obtaining preliminary prediction results using the NWP model, they are further optimized using machine learning methods. In addition, multimodal data fusion uses data from different sources, such as meteorological observations, satellite remote sensing, and historical power generation data, to effectively capture the spatial and temporal characteristics of GHI and reduce the uncertainty caused by a single data source. This fusion greatly improves the prediction accuracy under complex weather conditions. Although there are currently a variety of irradiance prediction methods, they still face many problems in application. For example, methods based on physical models rely on a large number of meteorological parameters and complex calculations, and data acquisition and model maintenance costs are high. It is difficult to update in real time and adapt to changing environmental conditions. Although data-driven statistical methods can achieve good prediction results in some cases, it is difficult to effectively extract and integrate multi-source information under complex weather conditions, resulting in reduced prediction accuracy. These problems seriously restrict the overall performance of photovoltaic power generation systems, especially when responding to weather and environmental changes quickly in a short period of time. Summary of the invention
[0004] The purpose of the present invention is to provide a short-term irradiance prediction method and system based on late-stage feature fusion of multimodal data in response to the above-mentioned problems, aiming to improve the accuracy of irradiance prediction. This method comprehensively utilizes information from multiple data sources, such as meteorological observations, historical power generation data, satellite images, etc., and fully captures the spatial and temporal characteristics of irradiance through the fusion of multimodal data, thereby improving the robustness and accuracy of the prediction model. Especially in short-term irradiance prediction, this method can quickly adapt to drastic changes in weather and environment, and provide more reliable scheduling and optimization solutions for photovoltaic systems. In this way, photovoltaic power generation can be better integrated into the existing energy system and gradually replace traditional fossil energy.
[0005] The technical solution of the present invention is as follows:
[0006] A short-term irradiance prediction method based on late feature fusion of multimodal data includes the following steps:
[0007] Data acquisition and preprocessing: Acquire all-sky image data and synchronized text data, and perform data preprocessing on the image data and text data. The text data includes irradiance data and weather data.
[0008] Feature extraction: A convolutional neural network ATT-CNN with an embedded channel attention mechanism is used to extract image features from image data, including cloud coverage, sun position, and cloud shape; a long short-term memory network LSTM is used to extract time series features from text data;
[0009] Multimodal late feature fusion: image features and time series features are fused through interaction and splicing to form a joint feature vector; during the fusion process, the gating mechanism is used to weight the features of different modalities;
[0010] Prediction output: The joint feature vector is predicted through a multi-layer feedforward artificial neural network MLP and a linear regression layer to output the short-term irradiance prediction value.
[0011] Furthermore, the feature extraction includes image feature extraction and text feature extraction. The ATT-CNN network includes a CNN convolution component and a channel attention mechanism component. The CNN convolution component includes at least two convolution layers. The feature extraction of the input image is performed through convolution operation to capture image features at different levels layer by layer. The channel attention mechanism component can automatically focus on important areas in the image, improve regional feature representation, and obtain image features.
[0012] Furthermore, the image feature extraction specifically includes the following steps:
[0013] The input full sky image data is convolved by the CNN convolution component to extract features, and a feature map is created for each layer, and the feature map is integrated into the global feature vector. The CNN convolution component includes a convolution layer with an activation function ReLU. The activation function expression is:
[0014]
[0015] Where σ is the activation function, f is the number of filters, c is the input channel index, and A i,j,f is the output activation value of the f filters at the input image position (i, j), K is the size of the convolution kernel, I i+m,j+n,c is the pixel value at (i+m,j+n) of the input feature map. The input feature map size is H×W×C, where H is the height, W is the width, and C is the number of channels. After the convolution layer, W m,n,c,f is the weight of the convolution kernel at position (m, n) of the input channel c and the filter f in the convolution kernel, b f is the bias term;
[0016] The convolution operation slides the filter on the input image and calculates the dot product of the pixels in the local area and the filter weight to obtain the output feature map at each position; after each convolution layer, the spatial dimension of the feature map is reduced by the maximum pooling layer. The expression of the maximum pooling operation is:
[0017] P i,j,f =max(A 2i,2j,f ,A 2i,2j+1,f ,A 2i+1.f ,A 2i+1,2j+1,f ),
[0018] Among them, P i,j,f represents the pooled output of the f-th filter at position (i, j);
[0019] After completing the convolution and pooling operations, the network transitions to the fully connected layer MLP, which performs weighted summation on the input of each neuron, obtains the output after the activation function, and converts the pooled output into a one-dimensional vector for flattening. k is the kth eigenvalue after flattening the feature map output by the pooling, then:
[0020]
[0021] A n =σ(z n ),
[0022] Among them, z n is the weighted sum of the inputs to the nth neuron in the fully connected layer, W k,nis the weight connecting the kth neuron in the previous layer to the nth neuron in the current layer, F is the feature dimension input to the current layer, and b n is the bias term of the nth neuron, A n is the activation value of the nth neuron in the fully connected layer, and is the output of the fully connected layer;
[0023] After five sets of convolution and maximum pooling operations, a channel attention mechanism is applied to the feature map. The channel attention mechanism highlights important channels and suppresses unimportant channels by applying weighted adjustments to each channel. First, global average pooling is used to capture the global information of each channel, and then the pooled feature vector is passed through two consecutive 1D convolution layers. The ReLU activation function is used after each convolution layer. After the second convolution layer, the weight of each channel is compressed to between 0 and 1. The weight represents the relative importance of each channel. The closer the channel is to 1, the more important it is, and the closer the channel is to 0, the less important it is. The attention vector is multiplied element by element with the original input feature to dynamically adjust the importance of each channel by highlighting important features and suppressing less important features. The obtained weighted features are enhanced by the attention mechanism and then forwarded to the subsequent layers of the neural network for further processing. The size of the input tensor X is H×W×C. The core step expression of channel attention is:
[0024] Global Average Pooling: Among them, T c is the global average pooling value of channel c;
[0025] The feature vector after global average pooling is passed through multiple 1x1 convolutional layers and the ReLU activation function is applied to further refine the feature representation; two layers of 1D convolution and ReLU activation are:
[0026]
[0027] Among them, W 1 , W 2 is the weight of the convolutional layer, b 1 , b 2 is the bias term, T c ′ It is the channel attention vector after two layers of convolution;
[0028] Apply the Sigmoid activation function to transform T c ′ Convert to a value between 0 and 1:
[0029]
[0030] Among them, σ represents the Sigmoid function, is the final attention value of channel c;
[0031] Finally, the attention value of each channel Multiply the input feature map X element by element to dynamically adjust the importance of each channel:
[0032]
[0033] Among them, Y is the weighted feature map, which represents the output after the channel attention mechanism is applied. The final expression of the channel attention mechanism is:
[0034]
[0035] Where X is an input tensor of shape (H, W, C), where H is the height, W is the width, and C is the number of channels. is the attention tensor obtained after dense layers and reshaping, with shape (1,1,C), where ⊙ represents element-wise multiplication.
[0036] Furthermore, the text feature extraction comprises the following steps:
[0037] Input the preprocessed text data into the LSTM network to extract the trend and periodic change characteristics in the time series;
[0038] After being processed by the ATT-CNN model, the image features are flattened into a one-dimensional vector and input into the fully connected layer MLP, combined with the temporal features of the text data extracted by LSTM to generate the final joint feature vector.
[0039] Furthermore, the multimodal late feature fusion specifically includes the following steps:
[0040] The extracted image text features are input into the MLP layer for modal fusion processing. The MLP includes a neural network with multiple fully connected layers, which can perform nonlinear mapping and feature extraction on the input features.
[0041] Applying the activation function σ to the output of the MLP limits the weight of each feature position to [0,1], and controls it using the gating mechanism in the long short-term memory network LSTM to generate a gate coefficient matrix G for gating control, which is expressed as:
[0042] G = σ(MLP(F input )),
[0043] Among them, F input Represents input features, including image features or text features, G and F input have the same dimensions, each element G i The value of is between [0,1], indicating the degree to which the position feature is retained or suppressed in subsequent processing;
[0044] The generated gate coefficient G is compared with the input feature F input Multiply element by element to get the weighted feature representation F weighted :
[0045] F weighted =G⊙F input ,
[0046] Among them, ⊙ represents the element-wise multiplication operation, that is, F weighted[i] =G i *F input[i] , which is equivalent to a dynamic selection and weighting of the input features, and the gated weighted image features F img and text features F text Perform splicing and fusion to obtain the joint feature representation F fused , used to capture the correlation and complementarity between multiple modalities:
[0047] F fused =Concat(F img ,F text ),
[0048] For the joint feature F fused Perform nonlinear transformation to further extract deep features and obtain the final output feature F output :
[0049] F output =MLP(F fused ),
[0050] MLP consists of multiple fully connected layers and uses nonlinear activation functions to increase the representation capability of the model, so that it can capture complex relationships between features and high-level abstract information.
[0051] Furthermore, the prediction output also includes model training and verification, and the model training and verification includes the following steps:
[0052] The training process includes: data set division, dividing the acquired data set into training set, validation set and test set. The training set is used to train the model, the validation set is used to adjust the model hyperparameters, and the test set is used to evaluate the performance of the model.
[0053] The mean square error (MSE) is used as the loss function, which measures the difference between the predicted value and the true value:
[0054]
[0055] Among them, y i Represents the true irradiance value, represents the predicted value of the model, N represents the number of samples,
[0056] The trained model needs to be evaluated on the test set. The evaluation indicators include mean square error RMSE, mean absolute error MAE and determination coefficient R 2 .
[0057] Furthermore, the data preprocessing of the image data and the text data comprises the following steps:
[0058] Image data is downsampled and compressed to balance all input weights: the collected all-sky image data is downsampled to 128×128, and the time shift problem caused by continuous shooting is screened, the data points with significant offset are deleted, and all RGB channels and numerical data are normalized to the [0,1] interval;
[0059] The text data is resampled to balance the data distribution: the text data is collected, and after data alignment and quality control, the sample size of sunny and non-sunny days is balanced; the continuous points of sunny and non-sunny days are filtered, including: if the CSI values of the current five data points and the next ten data points are continuously in the same fixed interval, it is a "sample imbalance" state, then one data point is excluded in this interval to balance the data interference.
[0060] Furthermore, the prediction output inputs the processed image data and text data into a pre-built short-term solar irradiance prediction model to output a short-term irradiance prediction value, wherein the short-term solar irradiance prediction model includes a data processing layer, a text image training layer, and a later cross-modal feature fusion prediction output layer.
[0061] The present application also includes a short-term irradiance prediction system based on late feature fusion of multimodal data, including:
[0062] Data collection module: collects all-sky image data and synchronized text data, and performs data preprocessing on the image data and text data. The text data includes irradiance data and weather data. The image data balances the weights of all inputs through downsampling and compression, and the text data balances the data distribution through resampling.
[0063] Feature extraction module: A convolutional neural network ATT-CNN with embedded channel attention mechanism is used to extract image features from image data; long short-term memory network LSTM is used to extract time series features from text data;
[0064] Feature fusion module: fuses image features and time series features through interaction and splicing to form a joint feature vector;
[0065] Model training and validation module: before making predictions, training and validation are performed using a large amount of data;
[0066] Prediction output module: predicts the joint feature vector through a multi-layer feedforward artificial neural network MLP and a linear regression layer, and outputs the short-term irradiance prediction value.
[0067] Furthermore, the feature extraction module includes an image feature extraction module and a text feature extraction module, the ATT-CNN network includes a CNN convolution module and a channel attention mechanism module, the CNN convolution module includes at least two convolution layers, and performs feature extraction on the input image through convolution operations to capture image features at different levels layer by layer; the channel attention mechanism module can automatically focus on important areas in the image, improve regional feature representation, and obtain image features.
[0068] Compared with the prior art, the present invention has the following beneficial effects:
[0069] 1. A short-term irradiance prediction method and system based on multimodal data post-feature fusion. This method processes all-sky images and data information through a hybrid network model of an image feature extraction network (ATT-CNN) and a text feature extraction network (LSTM), and then performs post-feature fusion through a feature fusion module to finally output an accurate short-term irradiance prediction value. This method combines image and data information, which can significantly improve the accuracy of short-term irradiance prediction;
[0070] 2. A short-term irradiance prediction method and system based on late-stage feature fusion of multimodal data, using an attention mechanism embedded convolutional neural network (ATT-CNN) model, which optimizes the prediction performance of solar irradiance by combining local and global features. This model effectively overcomes the limitations of existing CNN-based methods in cloudy weather and large-scale cloud movement, thereby improving the reliability of solar irradiance prediction and providing more accurate prediction results for practical applications;
[0071] 3. A short-term irradiance prediction method based on late-stage feature fusion of multimodal data. The features of different modes are screened and weighted through a gating architecture. The gating architecture is used to control the interaction and information flow between features, fuse multimodal features, and optimize the combination of features. The features after gating are further integrated and dimensionally reduced. Finally, the short-term irradiance prediction value is calculated and output through the regression layer. This achieves the effect of high-precision irradiance prediction within a short time scale. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 This is a flow chart of a short-term irradiance prediction method based on late feature fusion of multimodal data.
[0073] Figure 2 Schematic diagram of the model structure of ATT-CNN, a short-term irradiance prediction method based on late feature fusion of multimodal data.
[0074] Figure 3 This is a schematic diagram of the structure of an LSTM network for a short-term irradiance prediction method based on late feature fusion of multimodal data.
[0075] Figure 4 This is the architecture diagram of the multimodal data post-feature fusion of this application.
[0076] Figure 5 This is the overall framework flow chart of this application. DETAILED DESCRIPTION
[0077] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0078] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.
[0079] See also Figure 1-5 , a short-term irradiance prediction method based on late feature fusion of multimodal data, e.g. Figure 1 As shown, the following steps are included:
[0080] Data acquisition and preprocessing: Acquire all-sky image data and synchronized text data, and perform data preprocessing on the image data and text data. The text data includes irradiance data and weather data.
[0081] Feature extraction: A convolutional neural network ATT-CNN with an embedded channel attention mechanism is used to extract image features from image data, including cloud coverage, sun position, cloud shape, etc.; a long short-term memory network LSTM is used to extract time series features from text data;
[0082] Multimodal late feature fusion: image features and time series features are fused through interaction and splicing to form a joint feature vector; during the fusion process, the gating mechanism is used to weight the features of different modalities;
[0083] Prediction output: The joint feature vector is predicted through a multi-layer feedforward artificial neural network MLP and a linear regression layer to output the short-term irradiance prediction value.
[0084] Data preprocessing for image data and text data includes the following steps:
[0085] Image data is downsampled and compressed to balance all input weights: the collected all-sky image data is downsampled to 128×128, and the time shift problem caused by continuous shooting is screened, the data points with significant offset are deleted, and all RGB channels and numerical data are normalized to the [0,1] interval;
[0086] The text data is resampled to balance the data distribution: the text data is collected, and after data alignment and quality control, the sample size of sunny and non-sunny days is balanced; the continuous points of sunny and non-sunny days are filtered, including: if the CSI values of the current five data points and the next ten data points are continuously in the same fixed interval, it is a "sample imbalance" state, then one data point is excluded in this interval to balance the data interference.
[0087] Feature extraction includes image feature extraction and text feature extraction. The ATT-CNN network includes a CNN convolution component and a channel attention mechanism component. The CNN convolution component includes at least two convolution layers. It extracts features of the input image through convolution operations and captures image features at different levels layer by layer. The channel attention mechanism component can automatically focus on important areas in the image, improve regional feature representation, and obtain image features.
[0088] Image feature extraction specifically includes the following steps:
[0089] like Figure 2 As shown, the CNN convolution component performs a convolution operation on the input full sky image data to extract features, creates a feature map for each layer, and integrates the feature map into the global feature vector. The CNN convolution component includes a convolution layer with an activation function ReLU, and the activation function expression is:
[0090]
[0091] Where σ is the activation function, f is the number of filters, c is the input channel index, and A i,j,f is the output activation value of the f filters at the input image position (i, j), K is the size of the convolution kernel, I i+m,j+n,c is the pixel value at (i+m,j+n) of the input feature map. The input feature map size is H×W×C, where H is the height, W is the width, and C is the number of channels. After the convolution layer, W m,n,c,f is the weight of the convolution kernel at position (m, n) of the input channel c and the filter f in the convolution kernel, b f is the bias term;
[0092] The convolution operation slides the filter on the input image and calculates the dot product of the pixels in the local area and the filter weight to obtain the output feature map at each position; after each convolution layer, the spatial dimension of the feature map is reduced by the maximum pooling layer. The expression of the maximum pooling operation is:
[0093] P i,j,f =max(A 2i,2j,f ,A 2i,2j+1,f ,A 2i+1.f ,A 2i+1,2j+1,f ),
[0094] Among them, P i,j,f represents the pooled output of the f-th filter at position (i, j);
[0095] After completing the convolution and pooling operations, the network transitions to the fully connected layer MLP, which performs weighted summation on the input of each neuron, obtains the output after the activation function, and converts the pooled output into a one-dimensional vector for flattening. k is the kth eigenvalue after flattening the feature map output by the pooling, then:
[0096]
[0097] A n =σ(z n ),
[0098] Among them, z n is the weighted sum of the inputs to the nth neuron in the fully connected layer, W k,n is the weight connecting the kth neuron in the previous layer to the nth neuron in the current layer, F is the feature dimension input to the current layer, and b n is the bias term of the nth neuron, A n is the activation value of the nth neuron in the fully connected layer, and is the output of the fully connected layer;
[0099] After five sets of convolution and maximum pooling operations, a channel attention mechanism is applied to the feature map. The channel attention mechanism highlights important channels and suppresses unimportant channels by applying weighted adjustments to each channel. First, global average pooling is used to capture the global information of each channel, and then the pooled feature vector is passed through two consecutive 1D convolution layers. The ReLU activation function is used after each convolution layer. After the second convolution layer, the weight of each channel is compressed to between 0 and 1. The weight represents the relative importance of each channel. The closer the channel is to 1, the more important it is, and the closer the channel is to 0, the less important it is. The attention vector is multiplied element by element with the original input feature to dynamically adjust the importance of each channel by highlighting important features and suppressing less important features. The obtained weighted features are enhanced by the attention mechanism and then forwarded to the subsequent layers of the neural network for further processing. The size of the input tensor X is H×W×C. The core step expression of channel attention is:
[0100] Global Average Pooling: Among them, T c is the global average pooling value of channel c;
[0101] The feature vector after global average pooling is passed through multiple 1x1 convolutional layers and the ReLU activation function is applied to further refine the feature representation; two layers of 1D convolution and ReLU activation are:
[0102]
[0103] Among them, W 1 , W 2 is the weight of the convolutional layer, b 1 , b 2 is the bias term, T c ′ It is the channel attention vector after two layers of convolution;
[0104] Apply the Sigmoid activation function to transform T c ' Convert to a value between 0 and 1:
[0105]
[0106] Among them, σ represents the Sigmoid function, is the final attention value of channel c;
[0107] Finally, the attention value of each channel Multiply the input feature map X element by element to dynamically adjust the importance of each channel:
[0108]
[0109] Among them, Y is the weighted feature map, which represents the output after the channel attention mechanism is applied. The final expression of the channel attention mechanism is:
[0110]
[0111] Where X is an input tensor of shape (H, W, C), where H is the height, W is the width, and C is the number of channels. is the attention tensor obtained after dense layer and reshaping, with shape (1,1,C), ⊙ represents element-wise multiplication; Multiply the input tensor X channel by channel (i.e., dot product operation ⊙) to obtain the weighted feature tensor Y.
[0112] Text feature extraction includes the following steps:
[0113] like Figure 3 As shown, the preprocessed text data is input into the LSTM network to extract the trend and periodic change characteristics in the time series. The working principle of LSTM is:
[0114] f(t)=σ(W f ·[h t-1 ,x t ]+b f ),
[0115] i(t)=σ(W i ·[h t-1 ,x t ]+b i ),
[0116]
[0117]
[0118] O t =σ(W O ·[h t-1 ,x t ]+b O ),
[0119] h(t)=O t *tanh(C t ),
[0120]
[0121] σ represents the sigmoid function, which controls the information transmission state. When σ is 0, no information can pass through; when σ is 1, everything can pass through. f , W i , W c , W Ois the input weight, corresponding to b f 、b i 、b c 、b O is the bias, where t and t-1 represent the current and previous time states, x and h represent input and output, and C represents the cell state;
[0122] After being processed by the ATT-CNN model, the image features are flattened into a one-dimensional vector and input into the fully connected layer MLP, combined with the temporal features of the text data extracted by LSTM to generate the final joint feature vector.
[0123] The multimodal late feature fusion specifically includes the following steps:
[0124] like Figure 4 As shown, the extracted image text features are input into the MLP layer modality fusion processing. The MLP includes a neural network with multiple fully connected layers, which can perform nonlinear mapping and feature extraction on the input features.
[0125] Applying the activation function σ to the output of the MLP limits the weight of each feature position to [0,1], and controls it using the gating mechanism in the long short-term memory network LSTM to generate a gate coefficient matrix G for gating control, which is expressed as:
[0126] G = σ(MLP(F input )),
[0127] Among them, F input Represents input features, including image features or text features, G and F input have the same dimensions, each element G i The value of is between [0,1], indicating the degree to which the position feature is retained or suppressed in subsequent processing;
[0128] The generated gate coefficient G is compared with the input feature F input Multiply element by element to get the weighted feature representation F weighted :
[0129] F weighted =G⊙F input ,
[0130] Among them, ⊙ represents the element-wise multiplication operation, that is, F weighted[i] =G i *F input[i] , which is equivalent to a dynamic selection and weighting of the input features, and the gated weighted image features F img and text features F text Perform splicing and fusion to obtain the joint feature representation F fused , used to capture the correlation and complementarity between multiple modalities:
[0131] F fused =Concat(F img ,F text ),
[0132] For the joint feature F fused Perform nonlinear transformation to further extract deep features and obtain the final output feature F output :
[0133] F output =MLP(F fused ),
[0134] MLP consists of multiple fully connected layers and uses nonlinear activation functions to increase the representation capability of the model, so that it can capture complex relationships between features and high-level abstract information.
[0135] Model training and validation are also required before prediction output. Model training and validation include the following steps:
[0136] The training process includes: data set division, dividing the acquired data set into training set, validation set and test set. The training set is used to train the model, the validation set is used to adjust the model hyperparameters, and the test set is used to evaluate the performance of the model.
[0137] The mean square error (MSE) is used as the loss function, which measures the difference between the predicted value and the true value:
[0138]
[0139] Among them, y i Represents the true irradiance value, represents the predicted value of the model, N represents the number of samples,
[0140] The trained model needs to be evaluated on the test set. The evaluation indicators include mean square error RMSE, mean absolute error MAE and determination coefficient R 2 , the expression is:
[0141]
[0142] in, is the average of all true values. For short-term irradiance prediction, high R 2 A low MSE indicates that the model has good predictive performance.
[0143] The prediction output inputs the processed image data and text data into a pre-built short-term solar irradiance prediction model to output the short-term irradiance prediction value, where the short-term solar irradiance prediction model includes a data processing layer, a text image training layer, and a later cross-modal feature fusion prediction output layer.
[0144] This application also includes a short-term irradiance prediction system based on late feature fusion of multimodal data, such as Figure 5 As shown, including:
[0145] Data collection module: collects all-sky image data and synchronized text data, and performs data preprocessing on the image data and text data. The text data includes irradiance data and weather data. The image data balances the weights of all inputs through downsampling and compression, and the text data balances the data distribution through resampling.
[0146] Feature extraction module: A convolutional neural network ATT-CNN with embedded channel attention mechanism is used to extract image features from image data; long short-term memory network LSTM is used to extract time series features from text data;
[0147] Feature fusion module: fuses image features and time series features through interaction and splicing to form a joint feature vector;
[0148] Model training and validation module: before making predictions, training and validation are performed using a large amount of data;
[0149] Prediction output module: predicts the joint feature vector through a multi-layer feedforward artificial neural network MLP and a linear regression layer, and outputs the short-term irradiance prediction value.
[0150] The feature extraction module includes an image feature extraction module and a text feature extraction module. The ATT-CNN network includes a CNN convolution module and a channel attention mechanism module. The CNN convolution module includes at least two convolution layers. It extracts features of the input image through convolution operations and captures image features at different levels layer by layer. The channel attention mechanism module can automatically focus on important areas in the image, improve regional feature representation, and obtain image features.
[0151] In another specific embodiment, a short-term irradiance prediction method based on late feature fusion of multimodal data, such as Figure 1 As shown, the following steps are included:
[0152] Acquire all-sky image data and synchronized text data, and perform data preprocessing on the image data and text data, the text data including irradiance data and weather data;
[0153] The dataset used is from the Folsom Public Database in California, which contains raw full-sky images (1536 pixels × 1536 pixels), solar irradiance data, and weather data. These data are first time-aligned using timestamps, and then the corresponding clear sky irradiance is obtained from the Maclear Clear Sky Model. Subsequently, quality control filters are applied to each data;
[0154] In order to improve the efficiency and accuracy of the prediction model, the image data was first downsampled to 128×128 pixels. Due to the cumulative error caused by continuous shooting, the image dataset showed occasional time shift possibilities. Data points showing significant offsets (more than 15 seconds from the timestamp) were deleted. Reducing the image pixel size can reduce redundancy and speed up model training. The original resolution of 1536×1536 images was reduced to 128×128, and the compression ratio was:
[0155]
[0156] In order to balance the weights of all inputs for compression processing and obtain a compressed full-sky image, the numerical data of all RGB channels and images are normalized to the interval [0,1];
[0157] Meteorological data include temperature (unit: ℃), humidity (unit: %), wind speed (unit: m / s), etc. In preprocessing, continuous data is resampled to once an hour to ensure synchronization with all-sky image data. Sunny and non-sunny data are adjusted by sampling balance. The text data comes from the Folsom dataset. After the data alignment and quality control stage, due to the Folsom dataset, the sample size of sunny days is much larger than that of non-sunny days. If the CSI value of the first 5 data points and the last 10 data points is greater than 0.9 and less than 1.05, it is defined as a "sunny" state, and a data point will be excluded to balance the interference of data during sunny days. Irradiance and other meteorological data are normalized to ensure that different input features have balanced weights. In addition, the preprocessing process also includes screening for significantly offset image data points and resampling text data to balance the sample ratio of sunny and non-sunny days to avoid data imbalance from having a negative impact on the prediction results;
[0158] The image data, measured irradiance, and meteorological data are time-aligned by timestamps to ensure that all data are observed at the same time point. During time alignment, the timestamp of the Folsom data point is used as an index, and the Maclear model generates clear sky irradiance values based on the corresponding timestamps to ensure that the clear sky model values are strictly synchronized with the real-time measurement data;
[0159] For feature extraction, the convolutional neural network embedded with the attention mechanism (ATT-CNN) is used to extract image features of image data and capture key features, including cloud coverage, sun position, cloud shape, cloud movement direction, sky brightness distribution, and cloud transparency. The long short-term memory network (LSTM) extracts features of data information from the input data to capture the temporal trend, periodic changes, abnormal fluctuations and other time-related features of the data, and further understand the dynamic changes of weather conditions and irradiance.
[0160] ATT-CNN is used to extract image features. It contains a CNN module with multiple convolutional layers and a channel attention mechanism to capture key image features such as cloud cover and sun position.
[0161] See also Figure 2 , the CNN component consists of a series of convolutional layers with an activation function (ReLU), where the activation function expression is as follows:
[0162]
[0163] Where σ is the activation function, f is the number of filters, and c is the input channel index A i,j,f Refers to the output activation value of the f filters at the input image position (i, j), K is the size of the convolution kernel, I i+m,j+n,c is the pixel value at (i+m,j+n) of the input feature map (size H×W×C). After the convolution layer, W m,n,c,f is the weight of the convolution kernel at position (m, n) of the input channel c and the filter f in the convolution kernel, b f is the bias term. The convolution operation slides the filter on the input image and calculates the dot product of the pixels in the local area and the filter weight to obtain the output feature map at each position. After each convolution layer, the maximum pooling layer is used to reduce the spatial dimension of the feature map, while retaining the most important information and improving the model's robustness to deformations such as image translation. The maximum pooling operation is expressed as follows:
[0164] P i,j,f =max(A 2i,2j,f ,A 2i,2j+1,f ,A 2i+1.f ,A 2i+1,2j+1,f ),
[0165] P i,j,f represents the pooled output of the f-th filter at position (i, j). This formula indicates that the maximum value is selected within a 2×2 window to ensure that each pooled position retains only the most important features in the area;
[0166] After completing the convolution and pooling operations, the transition to the fully connected layer (MLP) is performed. The fully connected layer performs a weighted summation on the input of each neuron and obtains the output after the activation function. First, the pooled output is converted into a one-dimensional vector flattening step. Let P k is the kth eigenvalue after flattening the feature map output by the pooling, then:
[0167]
[0168] A n =σ(z n ),
[0169] Among them, z n is the weighted sum of the inputs to the nth neuron in the fully connected layer, W k,n is the weight connecting the kth neuron in the previous layer to the nth neuron in the current layer, F is the feature dimension input to the current layer, and b n is the bias term of the nth neuron, A n is the activation value of the nth neuron in the fully connected layer, and is the output of the fully connected layer;
[0170] After five sets of convolution and maximum pooling operations, a channel attention mechanism is applied to the feature map (output from the CNN convolution component). The channel attention mechanism highlights important channels and suppresses unimportant channels by applying weighted adjustments to each channel. First, global average pooling is used to capture the global information of each channel, and then the pooled feature vector passes through two consecutive 1D convolution layers. Each convolution layer is followed by a ReLU activation function to enhance the expressiveness of the model by introducing nonlinearity. Here, 1x1 convolution (i.e., each convolution operation is performed only within each channel) is used to reduce computational complexity and allow the network to focus on the relationship between channels. After the second convolution layer, the weights of each channel are compressed to between 0 and 1. These weights represent the relative importance of each channel. Channels closer to 1 represent more important channels, and channels closer to 0 represent less important channels. The attention vector is multiplied element-wise with the original input features to dynamically adjust the importance of each channel by highlighting important features and suppressing less important features. The obtained weighted features are enhanced by the attention mechanism and then forwarded to the subsequent layers of the neural network for further processing, thereby improving the overall performance and feature representation of the model. Assuming that the size of the input tensor X is H×W×C, the core step expression of channel attention is as follows:
[0171] Global Average Pooling: Among them, T c is the global average pooling value of channel c;
[0172] The feature vector after global average pooling is passed through multiple 1x1 convolutional layers and the ReLU activation function is applied to further refine the feature representation. Assume there are two layers of 1D convolution and ReLU activation:
[0173]
[0174] Among them, W 1 , W 2 is the weight of the convolutional layer, b 1 , b 2 is the bias term, T c ' It is the channel attention vector after two layers of convolution;
[0175] In order to obtain the attention value of each channel, the Sigmoid activation function is applied to T c ′ Convert to a value between 0 and 1:
[0176]
[0177] Among them, σ represents the Sigmoid function, is the final attention value of channel c;
[0178] Finally, the attention value of each channel Multiply the input feature map X element by element to dynamically adjust the importance of each channel:
[0179]
[0180] Among them, Y is the weighted feature map, which represents the output after the channel attention mechanism is applied. In general, the final expression of the channel attention mechanism is:
[0181]
[0182] Where X is an input tensor of shape (H, W, C), where H is the height, W is the width, and C is the number of channels. is the attention tensor obtained after dense layer and reshaping, with shape (1,1,C), ⊙ represents element-wise multiplication; Multiply the input tensor X channel by channel (i.e., dot product operation ⊙) to obtain the weighted feature tensor Y.
[0183] The weighted feature Y is the enhanced output of the channel attention mechanism, indicating that the importance of each channel has been optimized after dynamic adjustment through the channel attention, and then further processed by the subsequent layers of the neural network, thereby improving the overall performance and feature representation of the model. The channel attention mechanism acts like an intelligent filter, which analyzes each channel of the input (such as different colors) to understand their importance, and then shines brighter light on the most important information in each channel. The intelligent channel attention mechanism emphasizes the red channel of a clear sky, the orange channel of a sunset, and also emphasizes highlight areas with texture changes, which may indicate clouds.
[0184] Input text data into LSTM, see Figure 3 ,The preprocessed text data is input into the LSTM network to extract the trend and periodic change characteristics in the time series. LSTM can extract time series features;
[0185] The working principle of LSTM is as follows:
[0186] f(t)=σ(W f·[h t-1 ,x t ]+b f ),
[0187] i(t)=σ(W i ·[h t-1 ,x t ]+b i ),
[0188]
[0189]
[0190] O t =σ(W O ·[h t-1 ,x t ]+b O ),
[0191] h(t)=O t *tanh(C t ),
[0192]
[0193] Among them, σ represents the sigmoid function, which controls the information transmission state. When σ is 0, no information can pass; when σ is 1, everything can pass. f , W i , W c , W O is the input weight. The corresponding b f 、b i 、b c 、b O is the bias. Where t and t-1 represent the current and previous time states. x and h represent input and output, and C represents the cell state;
[0194] Multimodal late feature fusion, through the gating architecture to filter and weight the features of different modes, including a 64-dimensional ANN layer and a 64-dimensional gating factor layer to control the interaction and information flow between features. The gating architecture fuses multimodal features and optimizes the combination of features. The features after gating are further input into a 32-dimensional ANN layer for further feature integration and dimensionality reduction. Finally, the short-term irradiance prediction value is calculated and output through the regression layer;
[0195] See also Figure 4 , image and text features are later fused through a multi-layer perceptron (MLP). First, the gate coefficient matrix G is generated using the gating mechanism:
[0196] G = σ(MLP(Finput )),
[0197] Among them, F input Represents input features (image features or text features), G and F input have the same dimensions, each element G i The value of is between [0,1], indicating the degree to which the position feature is retained or suppressed in subsequent processing;
[0198] The generated gate coefficient G is compared with the input feature F input Multiply element by element to get the weighted feature representation F weighted :
[0199] F weighted =G⊙F input ,
[0200] Among them, ⊙ represents the element-wise multiplication operation, that is, F weighted[i] =G i *F input[i] , which is equivalent to a dynamic selection and weighting of the input features, and the gated weighted image features F img and text features F text Perform splicing and fusion to obtain the joint feature representation F fused , used to capture the correlation and complementarity between multiple modalities:
[0201] F fused =Concat(F img ,F text ),
[0202] For the joint feature F fused Perform nonlinear transformation to further extract deep features and obtain the final output feature F output :
[0203] F output =MLP(F fused ),
[0204] MLP consists of multiple fully connected layers and uses nonlinear activation functions to increase the representation capability of the model, so that it can capture complex relationships between features and high-level abstract information;
[0205] The gate structure generates gate coefficients with the same dimensions as the input features for each node in the artificial neural network (ANN). These gate coefficients control the activation degree of the feature channel. After control, the gate coefficients are converted into an attention weight map, which is then multiplied element-wise with the original input features to dynamically adjust the weights of the features. Through this weighted operation, the areas that the model should focus on can be highlighted, and the features that contribute to the short-term irradiance prediction can be enhanced, while the irrelevant areas can be ignored. On this basis, the fused feature sequence is input into the regression layer to obtain the final solar irradiance prediction result;
[0206] The acquired dataset is divided into training set (70%), validation set (15%) and test set (15%). The mean square error (MSE) is used as the loss function. MSE measures the difference between the predicted value and the true value and is defined as follows:
[0207]
[0208] Among them, y i is the true value, is the predicted value, N is the number of samples;
[0209] The trained model needs to be evaluated on the test set to verify its generalization ability. Evaluation indicators include mean square error (RMSE), mean absolute error (MAE) and determination coefficient R 2 , the expression is as follows:
[0210]
[0211] in is the average of all true values. For short-term irradiance prediction, high R 2 A low MSE indicates that the model has good predictive performance;
[0212] The prediction performance of the method was evaluated and the accuracy of the output future short-term irradiance prediction values was calculated.
[0213] In order to verify the performance of the short-term irradiance prediction method, a comparative experiment was conducted based on different model structures and feature combinations in the embodiment. The specific experimental settings are as follows:
[0214] Benchmark model: The benchmark models include the traditional linear regression model (LR), support vector regression model (SVR) and convolutional neural network without attention mechanism (CNN-LSTM), and the improved model: the short-term irradiance prediction model based on multimodal late feature fusion;
[0215] Hyperparameter settings:
[0216] Learning rate α: 0.001;
[0217] Batch Size: 64
[0218] Epochs: 100
[0219] Optimization algorithm: Adam optimizer;
[0220] Early stopping strategy: If the loss of the validation set does not decrease within 10 consecutive training rounds, stop training;
[0221] Regularization technique: To prevent overfitting, Dropout is used in the fully connected layer, and the Dropout ratio is set to 0.5.
[0222] Use mean square error (RMSE), mean absolute error (MAE) and coefficient of determination R 2 It is used as an indicator to compare the contrast models and evaluate the prediction effect.
[0223] Prediction output: The joint feature vector is predicted through a multi-layer feedforward artificial neural network MLP and a linear regression layer to output the short-term irradiance prediction value.
[0224] The above-mentioned embodiments only express the specific implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the protection scope of the present application. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the technical solution concept of the present application, and these all belong to the protection scope of the present application.
Claims
1. A short-term irradiance prediction method based on late feature fusion of multimodal data, characterized in that: The following steps are involved: Data acquisition and preprocessing: Acquire all-sky image data and synchronized text data, and perform data preprocessing on the image data and text data. The text data includes irradiance data and weather data. Feature extraction: A convolutional neural network ATT-CNN with an embedded channel attention mechanism is used to extract image features from image data, including cloud coverage, sun position, and cloud shape; a long short-term memory network LSTM is used to extract time series features from text data; Multimodal late feature fusion: image features and time series features are fused through interaction and splicing to form a joint feature vector; during the fusion process, the gating mechanism is used to weight the features of different modalities; Prediction output: The joint feature vector is predicted through a multi-layer feedforward artificial neural network MLP and a linear regression layer to output the short-term irradiance prediction value.
2. According to claim 1, a short-term irradiance prediction method based on late feature fusion of multimodal data is characterized in that: The feature extraction includes image feature extraction and text feature extraction. The ATT-CNN network includes a CNN convolution component and a channel attention mechanism component. The CNN convolution component includes at least two convolution layers. The feature extraction of the input image is performed through the convolution operation, and the image features at different levels are captured layer by layer. The channel attention mechanism component can automatically focus on important areas in the image, improve the regional feature representation, and obtain image features.
3. The short-term irradiance prediction method based on late feature fusion of multimodal data according to claim 2 is characterized in that: The image feature extraction specifically comprises the following steps: The input full sky image data is convolved by the CNN convolution component to extract features, and a feature map is created for each layer, and the feature map is integrated into the global feature vector. The CNN convolution component includes a convolution layer with an activation function ReLU. The activation function expression is: Where σ is the activation function, f is the number of filters, c is the input channel index, and A i,j,f is the output activation value of the f filters at the input image position (i, j), K is the size of the convolution kernel, I i+m,j+n,c is the pixel value at (i+m,j+n) of the input feature map. The input feature map size is H×W×C, where H is the height, W is the width, and C is the number of channels. After the convolution layer, W m,n,c,f is the weight of the convolution kernel at position (m, n) of the input channel c and the filter f in the convolution kernel, b f is the bias term; The convolution operation slides the filter on the input image and calculates the dot product of the pixels in the local area and the filter weight to obtain the output feature map at each position; after each convolution layer, the spatial dimension of the feature map is reduced by the maximum pooling layer. The expression of the maximum pooling operation is: P i,j,f =max(A 2i,2j,f ,A 2i,2j+1,f ,A 2i+1.f ,A 2i+1,2j+1,f ), Among them, P i,j,f represents the pooled output of the f-th filter at position (i, j); After completing the convolution and pooling operations, the network transitions to the fully connected layer MLP, which performs weighted summation on the input of each neuron, obtains the output after the activation function, and converts the pooled output into a one-dimensional vector for flattening. k is the kth eigenvalue after flattening the feature map output by the pooling, then: A n =σ(z n ), Among them, z n is the weighted sum of the inputs to the nth neuron in the fully connected layer, W k,n is the weight connecting the kth neuron in the previous layer to the nth neuron in the current layer, F is the feature dimension input to the current layer, and b n is the bias term of the nth neuron, A n is the activation value of the nth neuron in the fully connected layer, and is the output of the fully connected layer; After five sets of convolution and maximum pooling operations, a channel attention mechanism is applied to the feature map. The channel attention mechanism highlights important channels and suppresses unimportant channels by applying weighted adjustments to each channel. First, global average pooling is used to capture the global information of each channel, and then the pooled feature vector passes through two consecutive 1D convolution layers, and the ReLU activation function is used after each convolution layer. After the second convolution layer, the weight of each channel is compressed to between 0 and 1. The weight represents the relative importance of each channel. The attention vector is element-wise multiplied with the original input feature to dynamically adjust the importance of each channel by highlighting important features and suppressing less important features. The obtained weighted features are enhanced by the attention mechanism and then forwarded to the subsequent layers of the neural network for further processing. The size of the input tensor X is H×W×C. The core step expression of channel attention is: Global Average Pooling: Among them, T c is the global average pooling value of channel c; The feature vector after global average pooling is passed through multiple 1x1 convolutional layers and the ReLU activation function is applied to further refine the feature representation; two layers of 1D convolution and ReLU activation are: Among them, W1, W2 are the weights of the convolutional layer, b1, b2 are the bias terms, and T c ' It is the channel attention vector after two layers of convolution; Apply the Sigmoid activation function to transform T c ' Convert to a value between 0 and 1: Among them, σ represents the Sigmoid function, is the final attention value of channel c; Finally, the attention value of each channel Multiply the input feature map X element by element to dynamically adjust the importance of each channel: Among them, Y is the weighted feature map, which represents the output after the channel attention mechanism is applied. The final expression of the channel attention mechanism is: Where X is an input tensor of shape (H, W, C), where H is the height, W is the width, and C is the number of channels. is the attention tensor obtained after dense layers and reshaping, with shape (1,1,C), where ⊙ represents element-wise multiplication.
4. The short-term irradiance prediction method based on late feature fusion of multimodal data according to claim 2 is characterized in that: The text feature extraction comprises the following steps: Input the preprocessed text data into the LSTM network to extract the trend and periodic change characteristics in the time series; After being processed by the ATT-CNN model, the image features are flattened into a one-dimensional vector and input into the fully connected layer MLP, combined with the temporal features of the text data extracted by LSTM to generate the final joint feature vector.
5. The short-term irradiance prediction method based on late feature fusion of multimodal data according to claim 1 is characterized in that: The multimodal late feature fusion specifically includes the following steps: The extracted image text features are input into the MLP layer for modal fusion processing. The MLP includes a neural network with multiple fully connected layers, which can perform nonlinear mapping and feature extraction on the input features. Applying the activation function σ to the output of the MLP limits the weight of each feature position to [0,1], and controls it using the gating mechanism in the long short-term memory network LSTM to generate a gate coefficient matrix G for gating control, which is expressed as: G=σ(MLP(F input )), Among them, F input Represents input features, including image features or text features, G and F input have the same dimensions, each element G i The value of is between [0,1], indicating the degree to which the position feature is retained or suppressed in subsequent processing; The generated gate coefficient G is compared with the input feature F input Multiply element by element to get the weighted feature representation F weighted : F weighted =G⊙F input , Among them, ⊙ represents the element-wise multiplication operation, that is, F weighted[i] =G i *F input[i] , which is equivalent to a dynamic selection and weighting of the input features, and the gated weighted image features F img and text features F text Perform splicing and fusion to obtain the joint feature representation F fused , used to capture the correlation and complementarity between multiple modalities: F fused =Concat(F img ,F text ), For the joint feature F fused Perform nonlinear transformation to further extract deep features and obtain the final output feature F output : F output =MLP(F fused ), MLP consists of multiple fully connected layers and uses nonlinear activation functions to increase the representation capability of the model, so that it can capture complex relationships between features and high-level abstract information.
6. The short-term irradiance prediction method based on late feature fusion of multimodal data according to claim 1 is characterized in that: The prediction output also includes model training and verification, which includes the following steps: The training process includes: data set division, dividing the acquired data set into training set, validation set and test set. The training set is used to train the model, the validation set is used to adjust the model hyperparameters, and the test set is used to evaluate the performance of the model. The mean square error (MSE) is used as the loss function, which measures the difference between the predicted value and the true value: Among them, y i Represents the true irradiance value, represents the predicted value of the model, N represents the number of samples, The trained model needs to be evaluated on the test set. The evaluation indicators include mean square error RMSE, mean absolute error MAE and determination coefficient R 2 .
7. The short-term irradiance prediction method based on late feature fusion of multimodal data according to claim 1 is characterized in that: The data preprocessing of the image data and the text data comprises the following steps: Image data is downsampled and compressed to balance all input weights: the collected all-sky image data is downsampled to 128×128, and the time shift problem caused by continuous shooting is screened, the data points with significant offset are deleted, and all RGB channels and numerical data are normalized to the [0,1] interval; The text data is resampled to balance the data distribution: the text data is collected, and after data alignment and quality control, the sample size of sunny days and non-sunny days is balanced; the continuous points of sunny and non-sunny days are filtered, including: if the CSI values of the current five data points and the next ten data points are continuously in the same fixed interval, it is a "sample imbalance" state, then one data point is excluded in this interval to balance the data interference.
8. The short-term irradiance prediction method based on late feature fusion of multimodal data according to claim 1 is characterized in that: The prediction output inputs the processed image data and text data into a pre-built short-term solar irradiance prediction model to output a short-term irradiance prediction value, wherein the short-term solar irradiance prediction model includes a data processing layer, a text image training layer, and a later cross-modal feature fusion prediction output layer.
9. A short-term irradiance prediction system based on late feature fusion of multimodal data, characterized in that: include: Data collection module: collects all-sky image data and synchronized text data, and performs data preprocessing on the image data and text data. The text data includes irradiance data and weather data. The image data balances the weights of all inputs through downsampling and compression, and the text data balances the data distribution through resampling. Feature extraction module: A convolutional neural network ATT-CNN with embedded channel attention mechanism is used to extract image features from image data; long short-term memory network LSTM is used to extract time series features from text data; Feature fusion module: fuses image features and time series features through interaction and splicing to form a joint feature vector; Model training and validation module: before making predictions, training and validation are performed using a large amount of data; Prediction output module: predicts the joint feature vector through a multi-layer feedforward artificial neural network MLP and a linear regression layer, and outputs the short-term irradiance prediction value.
10. The short-term irradiance prediction system based on late feature fusion of multimodal data according to claim 9, characterized in that: The feature extraction module includes an image feature extraction module and a text feature extraction module. The ATT-CNN network includes a CNN convolution module and a channel attention mechanism module. The CNN convolution module includes at least two convolution layers. The feature extraction of the input image is performed through the convolution operation, and the image features at different levels are captured layer by layer. The channel attention mechanism module can automatically focus on important areas in the image, improve the regional feature representation, and obtain image features.
Citation Information
Patent Citations
Ultra-short-term solar irradiance prediction method based on multi-modal feature fusion
CN118470486A
Cited By
Photovoltaic power generation anomaly detection method and device and storage medium
CN120354315A
Soil entropy condition prediction method and system combined with multiple models
CN120688033A