Construction and application of multimodal diffusion model for ultra-short-term photovoltaic power probability prediction
By constructing a multimodal diffusion model and using a coupled U-shaped network to process sky images and photovoltaic power sequences, the problems of rapid cloud changes and uncertainty quantification in photovoltaic power prediction were solved, achieving higher accuracy and flexibility in ultra-short-term photovoltaic power prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU YIQI FUTURE ENERGY TECHNOLOGY CO LTD
- Filing Date
- 2025-06-25
- Publication Date
- 2026-05-12
AI Technical Summary
Existing photovoltaic power prediction models struggle to accurately capture rapid cloud changes and lack uncertainty quantification, resulting in inaccurate ultra-short-term photovoltaic power predictions, particularly in terms of multimodal data fusion and end-to-end uncertainty quantification.
A multimodal diffusion model is constructed using a coupled U-shaped network. By denoising the sky image sequence and photovoltaic power sequence branches, the probability distributions of multiple possible future sky image sequences and photovoltaic power sequences are generated, enabling cross-modal information interaction and joint prediction.
It significantly improves the accuracy and uncertainty quantification level of photovoltaic power prediction, and can flexibly handle multimodal information within a unified framework, thereby improving the accuracy and flexibility of prediction.
Smart Images

Figure CN120822406B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of photovoltaic forecasting, and in particular to the construction and application of a multimodal diffusion model for ultra-short-term photovoltaic power probability forecasting. Background Technology
[0002] Renewable energy, particularly solar photovoltaic (PV), is increasingly becoming a crucial pillar of future power systems. However, the inherent instability of PV power generation, especially the dramatic power fluctuations within seconds to 30 minutes caused by cloud cover changes (i.e., the "ultra-short-term" forecast range), poses a severe challenge to the stable operation of the power grid. Accurate ultra-short-term PV power probability forecasts are crucial for optimizing dispatch, reducing management costs, and improving system reliability. Failure to obtain timely and accurate ultra-short-term PV power probability forecasts will not only increase the complexity of grid dispatch but may also create additional economic and risk burdens in various application scenarios, including residential, commercial, and utility applications.
[0003] In recent years, sky images captured by ground-up cameras have become a key data source for monitoring cloud dynamics and assisting in the ultra-short-term probability prediction of photovoltaic power due to their high spatiotemporal resolution. The rapid development of deep learning in video prediction and time series analysis has also prompted more and more studies to try to use structures such as convolutional neural networks, recurrent neural networks (RNNs), and Transformers to automatically extract spatiotemporal features from sky images and combine them with historical photovoltaic data or other meteorological information to achieve end-to-end regression prediction of future photovoltaic power probability.
[0004] However, such end-to-end models typically face two major challenges: first, they struggle to accurately capture rapidly changing cloud dynamics, often exhibiting time-lag prediction errors when cloud cover changes drastically; second, they lack effective uncertainty quantification (UQ) capabilities, especially crucial for risk management in scenarios with a high proportion of renewable energy grid integration. While probabilistic solar forecasting has been extensively studied, most predictions are based on time-series photovoltaic measurements, and explorations using sky images for probabilistic forecasting remain relatively limited. In other words, there is currently a lack of generative models on the market capable of jointly generating different possible future sky images and their corresponding photovoltaic power sequence probability distributions within a unified end-to-end framework.
[0005] Diffusion models, due to their powerful probabilistic generation capabilities and advantages in modeling high-dimensional and complex data distributions, have shown great potential in computer vision and photovoltaic (PV) forecasting in recent years. Through progressive denoising, diffusion models not only achieve stable training but also flexibly characterize multimodal distributions and extreme values. In solar PV forecasting, studies by SHADECast et al. have applied diffusion models to probabilistic spatiotemporal prediction of satellite imagery, utilizing cloud evolution information to guide ensemble prediction. However, such works typically rely on single-modal inputs (e.g., satellite imagery) and fail to effectively integrate multimodal information such as PV power sequences. With the success of multimodal diffusion models in text-image and audio-visual generation, joint probabilistic prediction of visual data and physical quantities (e.g., PV power) within a unified diffusion framework is expected to significantly improve the prediction accuracy and UQ (unified quality) of PV power under extreme weather conditions, overcoming the shortcomings of existing methods in deep multimodal fusion and end-to-end uncertainty quantification.
[0006] However, constructing such an end-to-end multimodal photovoltaic power probability prediction diffusion model faces key challenges: (1) Data heterogeneity: Sky images (high-dimensional spatiotemporal RGB sequences) and photovoltaic power (one-dimensional time series) differ significantly in modality and representation. How to process them in parallel and effectively fuse them in a unified model is still undetermined. (2) Cross-modal alignment and interaction: The photovoltaic power sequence and the sky image sequence are closely related in time. The model needs to be able to capture their cross-modal dependencies and achieve bidirectional information interaction in joint prediction.
[0007] In summary, there is currently no effective solution for jointly generating possible future sky image sequences and their corresponding photovoltaic power sequence probability distributions from "photovoltaic power sequences + sky image sequences," resulting in inaccurate ultra-short-term probability predictions of photovoltaic power. Summary of the Invention
[0008] This application provides a construction and application of a multimodal diffusion model for ultra-short-term photovoltaic power probability prediction. This end-to-end multimodal diffusion model for ultra-short-term photovoltaic power probability prediction can jointly generate multiple possible future sky image sequences and corresponding photovoltaic power sequence probability distributions based on historical sky image sequences and historical photovoltaic power sequences, thereby achieving ultra-short-term photovoltaic prediction.
[0009] In a first aspect, embodiments of this application provide a method, the method comprising: acquiring a training dataset, wherein the training dataset includes multiple sets of training data, each set of training data including conditional pairing data and corresponding predicted pairing data of the conditional pairing data, wherein the conditional pairing data includes a first sky image sequence and a corresponding first photovoltaic power sequence, and the predicted pairing data includes a second sky image sequence and a corresponding second photovoltaic power sequence;
[0010] The training dataset is input into a multi-model diffusion framework to train a multimodal diffusion model. The multi-model diffusion framework is built based on the diffusion model and uses a coupled U-shaped network, which is coupled with a sky image sequence branch and a photovoltaic power sequence branch, for inverse denoising. The sky image sequence branch and the photovoltaic power sequence branch of the coupled U-shaped network share multiple coding layers, transition layers, and multiple decoding layers. Each coding layer of the sky image sequence branch has skip connections with the corresponding decoding layer, and each coding layer of the photovoltaic power sequence branch has skip connections with the corresponding decoding layer. The sky image sequence branch includes a sky image sequence encoder set before the coding layer and a sky image sequence decoder set after the decoding layer. The photovoltaic power sequence branch includes an image sequence encoder set before the coding layer and a photovoltaic power sequence decoder set after the decoding layer.
[0011] Secondly, embodiments of this application provide a method for probabilistic prediction of ultra-short-term photovoltaic power, comprising the following steps:
[0012] Input historical sky image sequences and corresponding historical photovoltaic data into the multimodal diffusion model, and output future sky image sequences and corresponding future photovoltaic power sequences;
[0013] Alternatively, input historical photovoltaic power sequences into a multimodal diffusion model and output future photovoltaic power sequences;
[0014] Alternatively, input historical sky sequences into a multimodal diffusion model and output future sky image sequences and future photovoltaic power sequences;
[0015] The multimodal diffusion model mentioned above is constructed according to the construction method of multimodal diffusion model.
[0016] The main contributions and innovations of this invention are as follows:
[0017] 1. This solution proposes a unified end-to-end multimodal diffusion model for "photovoltaic power sequence + sky image sequence". Within the unified end-to-end multimodal diffusion framework, different possible future sky images and their corresponding photovoltaic power distribution probability distributions are jointly generated, which fully explores the correlation of multimodal information and avoids the limitations of traditional "two-stage" models in information interaction and uncertainty transmission.
[0018] 2. The multimodal diffusion model of this scheme uses two coupled denoising autoencoders in the reverse denoising process. Based on historical sky image sequences and historical photovoltaic power sequences, it jointly generates multiple possible future sky image sequences and corresponding photovoltaic power sequence probability distributions. It iteratively optimizes by implicitly utilizing historical information and the generation results of the previous time step in each denoising step.
[0019] 3. The multimodal diffusion model designed in this scheme adaptively supports three input-output modes: "image + photovoltaic", "image only", and "photovoltaic only" through a unified network structure and hybrid loss strategy, thereby enhancing the flexibility and scalability of the multimodal diffusion model in real-world scenarios.
[0020] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description
[0021] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0022] Figure 1 This is an architectural diagram of a coupled U-shaped network according to an embodiment of this application.
[0023] Figure 2 This is an architecture diagram of an efficient multimodal module according to one embodiment of this application.
[0024] Figure 3 This is a schematic diagram of photovoltaic power sequences and sky image sequences in a PV multimodal attention network based on random offset.
[0025] Figure 4 This is a schematic diagram of the predicted scenario based on the multimodal diffusion model of this scheme.
[0026] Figures 5-7 The results are obtained by inputting the multimodal diffusion model and different benchmark models into the ultra-short-term photovoltaic power probability prediction.
[0027] Figure 8 This is a comparison chart of VGG cosine similarity between different models.
[0028] Figure 9 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0029] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0030] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in this specification in other embodiments. In some other embodiments, the methods may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; and multiple steps described in this specification may be combined into a single step in other embodiments.
[0031] Example 1
[0032] This solution constructs an end-to-end multimodal diffusion model for ultra-short-term photovoltaic (PV) power probability prediction. This model innovatively employs two coupled denoising autoencoders to jointly generate multiple possible future sky image sequences and corresponding PV power sequence probability distributions, conditioned on historical sky image sequences and historical PV power sequences. By implicitly utilizing historical information and the previous generation result for iterative optimization in each denoising step, the multimodal diffusion model can learn the joint distribution of multiple modes. Furthermore, a cross-modal attention module is designed within the model to enhance semantic synchronization and information interaction between the sky image sequences and PV power sequences at each time step, providing a comprehensive evaluation of model performance and demonstrating its flexibility. This multimodal diffusion model for ultra-short-term PV power probability prediction can significantly improve the efficient and resilient management of solar power generation, supporting a smooth transition to a grid dominated by renewable energy.
[0033] The method for constructing the multimodal diffusion model for ultra-short-term photovoltaic power probability prediction includes the following steps:
[0034] Obtain a training dataset, which includes multiple sets of training data. Each set of training data includes conditional pairing data and corresponding predicted pairing data. The conditional pairing data includes a first sky image sequence and a corresponding first photovoltaic power sequence, and the predicted pairing data includes a second sky image sequence and a corresponding second photovoltaic power sequence.
[0035] The training dataset is input into a multi-model diffusion framework to train a multimodal diffusion model. The multi-model diffusion framework is built based on the diffusion model and uses a coupled U-shaped network, which is coupled with a sky image sequence branch and a photovoltaic power sequence branch, for inverse denoising. The sky image sequence branch and the photovoltaic power sequence branch of the coupled U-shaped network share multiple coding layers, transition layers, and multiple decoding layers. Each coding layer of the sky image sequence branch has skip connections with the corresponding decoding layer, and each coding layer of the photovoltaic power sequence branch has skip connections with the corresponding decoding layer. The sky image sequence branch includes a sky image sequence encoder set before the coding layer and a sky image sequence decoder set after the decoding layer. The photovoltaic power sequence branch includes an image sequence encoder set before the coding layer and a photovoltaic power sequence decoder set after the decoding layer.
[0036] This scheme employs a coupled U-shaped network to jointly denoise the photovoltaic power sequence and the sky image sequence. The photovoltaic power sequence is input into the photovoltaic power sequence branch, and the sky image sequence is input into the sky image sequence branch, thereby more accurately predicting the future photovoltaic power sequence and sky image sequence during the denoising process. Multiple possible sky image sequences and their corresponding photovoltaic power sequences are presented in probability distribution.
[0037] Specifically, for each set of training data input into the multi-model diffusion model, the second photovoltaic power sequence is denoised using the first photovoltaic power sequence as a condition to obtain a photovoltaic denoised sequence, and the second sky image sequence is denoised using the first sky image sequence as a condition to obtain a sky image denoised sequence. The photovoltaic denoised sequence, the first photovoltaic power sequence, the sky image denoised sequence, and the first sky image sequence are input together into a coupled U-shaped network for coupled denoising to obtain a sky image denoised sequence and a photovoltaic denoised sequence.
[0038] Specifically, each set of training data is represented as a tensor pair to be input into the multi-model diffusion framework. The tensor pair is represented as:
[0039]
[0040] Where P represents a photovoltaic power sequence composed of a first photovoltaic power sequence and a second photovoltaic power sequence. C represents the number of channels, and T represents the time dimension;
[0041] S represents a sky image sequence consisting of a first sky image sequence and a second sky image sequence. T represents the time dimension, C represents the number of channels, and H and W represent the height and width of the sky image, respectively.
[0042] Regarding the noise addition process of the multi-model diffusion model, the second photovoltaic power sequence and the second sky image sequence are subjected to independent forward noise addition processing, and the first photovoltaic power sequence and the first sky image sequence remain unchanged during the forward noise addition process.
[0043] Furthermore, this scheme represents the first photovoltaic power sequence in the first training data as follows: The first sky image sequence is represented as The second photovoltaic power sequence is represented as The second sky image sequence is represented as The superscript indicates the time interval, and the subscript indicates the number of diffusion steps.
[0044] The forward noise-adding diffusion process of photovoltaic power sequences and sky image sequences in the multi-model diffusion framework is as follows:
[0045] For the training dataset (P,S) consisting of a photovoltaic power sequence P and a sky image sequence S, since the photovoltaic power sequence and the sky image sequence are distributed differently, the forward noise diffusion process can be considered independent of each other. It is worth noting that the first photovoltaic power sequence... and the first sky image sequence It remains unchanged during forward noise diffusion and is only used for the second photovoltaic power sequence. Second sky image sequence Noise is injected during the forward noise diffusion process.
[0046] For a photovoltaic power sequence P, the forward noise-adding diffusion process of the second photovoltaic power sequence at the t-th diffusion step is defined as:
[0047]
[0048] in It is the second photovoltaic power sequence at the t-th diffusion step. It is the second photovoltaic power sequence at the (t-1)th diffusion step. Let I be the variance scheduling sequence for the t-th diffusion step, where I is the identity matrix and T is the total number of diffusion steps. It is a Gaussian distribution for the photovoltaic power sequence, and q() is the probability distribution of the forward process.
[0049] For a sky image sequence S, the forward diffusion process of the second sky image sequence at the t-th diffusion step is defined as follows:
[0050]
[0051] in It is the second photovoltaic power sequence at the t-th diffusion step. It is the second photovoltaic power sequence at the (t-1)th diffusion step. Let I be the variance scheduling sequence for the t-th diffusion step, where I is the identity matrix and T is the total number of diffusion steps. It is a Gaussian distribution for the sky image sequence, and q() is the probability distribution of the forward process.
[0052] The reverse denoising process for photovoltaic power sequences and sky image sequences in a multi-model diffusion framework is as follows:
[0053] Unlike the forward process of modeling the second photovoltaic power sequence and the second sky image sequence independently, the reverse denoising stage not only needs to consider the overall correlation between the two modes, but also needs to take into account the correlation between the internal condition part and the prediction part of each mode.
[0054] For a photovoltaic power sequence P, the reverse denoising process of the photovoltaic noise-added sequence at the t-th diffusion step is defined as:
[0055] ;
[0056] in and These represent the first photovoltaic power sequence and the first sky image sequence, respectively (which remain unchanged during the reverse process). and Let represent the denoising results of the photovoltaic denoised sequence and the sky image denoised sequence at step t, respectively. The denoising result of the photovoltaic noise sequence at step t-1 is derived from the result of simultaneously depending on , , and The diffusion is generated by a Gaussian distribution, where t represents the number of diffusion steps. This represents a multimodal diffusion model. The probability distribution of the reverse process, where N() represents a Gaussian distribution. This represents the mean function of a Gaussian distribution.
[0057] For a sky image sequence S, the reverse denoising process of the second sky image sequence at the t-th diffusion step is defined as follows:
[0058]
[0059] in and These represent the first photovoltaic power sequence and the first sky image sequence, respectively (which remain unchanged during the reverse process). and Let represent the denoised results of the photovoltaic noise-added sequence and the sky image noise-added sequence at step t, respectively. For the noisy sequence of the sky image after denoising at step t-1, it is simultaneously dependent on , , and The diffusion is generated by a Gaussian distribution, where t represents the number of diffusion steps. This represents a multimodal diffusion model. The probability distribution of the reverse denoising process, where N() represents a Gaussian distribution. This represents the mean function of a Gaussian distribution.
[0060] To optimize the entire network, this solution adopts... - The prediction strategy allows the network to directly predict the noise added to the "future prediction" part (i.e., the second photovoltaic power sequence and the second sky image sequence). Since the conditional part (the first photovoltaic power sequence and the first sky image sequence) does not require denoising, the loss only applies to the prediction part.
[0061] The loss for the second photovoltaic power sequence is defined as follows:
[0062] ;
[0063] The loss for the second sky image sequence is defined as:
[0064] ;
[0065] in It is a multimodal diffusion model corresponding to the second photovoltaic power sequence or the second sky image sequence. The loss, This is for noise in noisy photovoltaic power sequences. The expectation operation of the Gaussian distribution, It is for noise in noisy sky image sequences. The expectation operation of the Gaussian distribution, It is noise predicted by the multimodal diffusion model. It is a time-step related weighting function, which is an optional weighting function. and These represent the first photovoltaic power sequence and the first sky image sequence, respectively (which remain unchanged during the reverse process). and Let denoise the second photovoltaic power sequence and the second sky image sequence at step t, respectively, where t is the diffusion step number. , It is an L2 norm.
[0066] Through the aforementioned multimodal diffusion framework, the trained multimodal diffusion model can not only more fully learn the semantic association between photovoltaic power sequences and sky image sequences in the multimodal feature space, but also comprehensively utilize the historical condition information of each modality. and This allows for the separate inference of their evolution at future moments. During the prediction phase, the "future" portions of the two modalities can complement and learn from each other, effectively mitigating the uncertainty caused by insufficient information from a single modality and ultimately achieving more accurate prediction results.
[0067] The architecture diagram of the coupled U-shaped network in this scheme is as follows: Figure 1 As shown, each encoding layer, each decoding layer, and each transition layer is a high-efficiency multimodal module. The architecture of the high-efficiency multimodal module is as follows: Figure 2 As shown.
[0068] Specifically, the coupled U-shaped network of this scheme includes four coding layers corresponding to different scales, two transition layers, and four decoding layers corresponding to different scales. Each coding layer and each decoding layer contains two efficient multimodal modules. It should be noted that self-attention and cross-modal attention for sky image sequences are only introduced at the second, third, and fourth scales.
[0069] Each efficient multimodal module includes a parallel sky image sequence sub-network, a step embedding branch, and a photovoltaic power sequence sub-network. The sky image sequence of the sky image sequence branch is input into the sky image sequence sub-network of the efficient multimodal module, and the photovoltaic power sequence of the photovoltaic power sequence branch is input into the photovoltaic power sequence sub-network of the efficient multimodal module. The discrete step count of the denoising process is input into the step embedding branch for linear processing and scaling offset to obtain the embedding vector. The embedding vector is applied to the sky image sequence sub-network and the photovoltaic power sequence sub-network, respectively. The output sequences of the sky image sequence sub-network and the photovoltaic power sequence sub-network are processed by a PV multimodal attention network based on random offset for cross-modal attention processing to obtain the sky image sequence of the corresponding sky image sequence branch and the photovoltaic power sequence of the corresponding photovoltaic power sequence branch.
[0070] Specifically, the step embedding branch of this scheme includes a linear layer and a scaling offset layer connected in sequence. The discrete step number is input into the step embedding branch and passes through the linear layer and the scaling offset layer in sequence to obtain the embedding vector. That is, the t-th step of the reverse denoising process is transformed into a floating-point vector through the step embedding branch.
[0071] The photovoltaic power sequence subnetwork of this scheme includes a group normalization layer, a sigmoid linear unit, a one-dimensional convolution, a one-dimensional self-attention layer, and a normalization layer connected in sequence. The photovoltaic power sequence of the photovoltaic power sequence branch is input into the photovoltaic power sequence subnetwork and processed by the group normalization layer, sigmoid linear unit, one-dimensional convolution, one-dimensional self-attention layer, and normalization layer in sequence. It is then connected to the photovoltaic power sequence residual of the original input photovoltaic power sequence subnetwork and embedded in the vector input normalization layer.
[0072] The sky image sequence subnetwork of this scheme includes a group normalization layer, a sigmoid linear unit, a two-dimensional spatial convolution, a one-dimensional temporal convolution, an upsampling layer, a one-dimensional to two-dimensional self-attention layer, and a normalization layer connected in sequence. The sky image sequence of the sky image sequence branch is input into the sky image sequence subnetwork and is processed by the group normalization layer, sigmoid linear unit, two-dimensional spatial convolution, one-dimensional temporal convolution, upsampling layer, one-dimensional to two-dimensional self-attention layer, and normalization layer in sequence. Then, it is residually connected with the sky image sequence input into the sky image sequence subnetwork and the vector is embedded into the input normalization layer.
[0073] It should be noted that the photovoltaic power sequence extracts dynamic features after undergoing one-dimensional convolution and one-dimensional self-attention processing in the photovoltaic power sequence sub-network, while the sky image sequence sequentially undergoes two-dimensional spatial convolution, one-dimensional temporal convolution, and one-dimensional-two-dimensional self-attention to replenish spatiotemporal correlation.
[0074] Subsequently, the photovoltaic power sequence after residual connection processing of the photovoltaic power sequence branch and the sky image sequence after residual connection processing of the sky image sequence branch are jointly input into the PV multimodal attention network based on random offset for alignment and adaptation processing.
[0075] Design of the sky image sequence subnetwork:
[0076] This scheme's efficient multimodal module captures both temporal and spatial information while reducing computational overhead. Specifically, in the sky image sequence sub-network, spatial and temporal convolutions are separated. Two-dimensional spatial convolutions are used to encode the spatial resolution of the input features, followed by stacking one-dimensional temporal convolutions along the temporal dimension, avoiding the high computational burden of directly using three-dimensional convolutions. Furthermore, to capture the changing characteristics of clouds across different dimensions, the sky image sequence sub-network introduces one-dimensional to two-dimensional self-attention; the one-dimensional self-attention corresponds to temporal attention, and the two-dimensional self-attention corresponds to spatial attention, thus enabling more effective modeling of the dynamic evolution of clouds over time and space.
[0077] Design of photovoltaic power sequence subnetwork:
[0078] Since it has only 16 time steps and only one feature value per time step, this scheme adopts the strategy of "one-dimensional convolution + one-dimensional self-attention". As the number of network layers increases, the number of channels is continuously increased, and downsampling is not performed on the time axis. This can preserve the key information in the photovoltaic power sequence to the maximum extent in the encoder stage, and achieve alignment with the sky image sequence sub-network at the multimodal attention level in the later stage.
[0079] Regarding PV multimodal attention networks based on random offsets:
[0080] When performing feature alignment and joint learning between the photovoltaic time series sub-network and the sky image sequence sub-network, the most direct approach is to perform global cross-modal attention on the features of both. However, due to the large size of the original attention maps for both the photovoltaic power sequence and the sky image sequence modalities, the computational complexity is extremely high. Furthermore, sky image sequences and photovoltaic power sequences generally exhibit redundancy in both temporal and spatial dimensions, and not all parts need to be included in cross-modal attention calculations. To address this issue, this scheme combines features from both photovoltaic power sequences and sky image sequences using a random offset-based PV multimodal attention network.
[0081] like Figure 3 As shown, for the photovoltaic power sequence and sky image sequence input into the PV multimodal attention network based on random offset, where the photovoltaic power sequence is a one-dimensional sequence corresponding to T frames and C channels, and the sky image sequence is a block sequence corresponding to T frames, C channels, H height and W width, cross-modal attention processing is performed on the photovoltaic power sequence and sky image sequence.
[0082] Regarding the operation of cross-modal attention processing: Set a window smaller than the number of frames and randomly sample offsets within the window. For each photovoltaic segment of the photovoltaic power sequence, use the offset, window, and number of frames to calculate the sky image sequence segment corresponding to the current photovoltaic segment. Perform cross-modal attention calculation on the photovoltaic segment and the corresponding sky image sequence segment.
[0083] Specifically, for the l-th layer of the coupled U-shaped network, the sky image sequence features are represented as follows: The block sequence, the photovoltaic power sequence characteristics are represented as The sequence is a one-dimensional sequence, where T represents the number of frames, H represents the height of the sky image, W represents the width of the sky image, and C represents the channel.
[0084] Set a window size S much smaller than the number of frames T, randomly sample offsets R in the interval [0, TS], and calculate the i-th photovoltaic segment in the photovoltaic power sequence features according to the following formula. Corresponding sky image sequence segmentation The sky image sequence is segmented. For frames arrive Sequence segmentation, i.e. :
[0085] ;
[0086]
[0087] Here, mod is the modulo operation. The purpose of this approach is to ensure that each photovoltaic value focuses on different local sky image sequences, thereby reducing ineffective cross-modal attention calculations.
[0088] Furthermore, for the i-th photovoltaic segment Segmentation with the corresponding sky image sequence The cross-modal attention between them is calculated to obtain the photovoltaic power sequence and sky image sequence output by the PV multimodal attention network based on random offset:
[0089] ;
[0090] in It is the dimension of K, and softmax() is the activation function. It is the query feature of the i-th photovoltaic segment. It is the query feature of the j-th sky image sequence segment. It is the key feature of the j-th sky image sequence segment. It is the key feature of the i photovoltaic segments. It is the value feature of the j-th sky image sequence segment. These are the value characteristics of the i photovoltaic segments. It is a segmentation of a sky image sequence. This is the i-th photovoltaic segment. `linear()` performs linear processing, and `flatten` performs flattening processing. It is a sequence of sky images output by a PV multimodal attention network based on random offset. It is the photovoltaic power sequence output by a PV multimodal attention network based on random offset.
[0091] This solution reduces computational complexity from the original by employing random offsets within a local window. Reduced to This significantly reduces the computational overhead required for cross-modal operations. Furthermore, the random offset design retains more fine-grained attention information between adjacent time segments, enabling more thorough cross-modal feature learning.
[0092] In addition, since cross-modal attention enables bidirectional interaction between features of photovoltaic power sequences and sky image sequences, the top layer of the coupled U-shaped network can use a smaller window to capture more nuanced feature correspondences, while the bottom layer of the coupled U-shaped network uses a larger window to capture higher-level semantic alignment relationships.
[0093] The coupled U-shaped network in this scheme can simultaneously focus on the feature representations of two modalities, photovoltaic and sky image, during the denoising process. This coupled U-shaped network can not only optimize the deep semantic association between photovoltaic power sequence and sky image sequence in multimodal feature space, but also infer their respective prediction components based on the conditional parts of the two modalities, while promoting interactive learning among these prediction components, thereby achieving a significant improvement in prediction accuracy in the final prediction stage.
[0094] In addition, the multimodal diffusion model provided by this scheme is compatible with three different data input-output modes: I. Input: historical sky image sequence and historical photovoltaic power sequence; Output: future sky image sequence and future photovoltaic power sequence; II. Input: historical sky sequence; Output: future sky image sequence and future photovoltaic power sequence; III. Input: historical photovoltaic power sequence; Output: future photovoltaic power sequence.
[0095] Specifically, the multimodal diffusion model in this scheme uses a "spaced-out" approach to handle missing modalities during both training and inference: if only historical sky image sequences are input, the historical photovoltaic power sequences are set to zero tensors; if only historical photovoltaic power sequences are input, the historical sky image sequences are set to zero tensors. Through this spaced-out operation, the multimodal diffusion model can uniformly handle three different data input-output modes during training and inference without the need to divide into multiple independent models. This improves flexibility and versatility while reducing the complexity of model management and ensuring the consistency of prediction results.
[0096] Correspondingly, the loss function of the coupled U-shaped network in the training process of this multimodal diffusion model is as follows:
[0097] ;
[0098] in It is the loss function of the coupled U-shaped network. This is for noise in noisy photovoltaic power sequences. The expectation operation of the Gaussian distribution, It is noise predicted by the multimodal diffusion model. and These represent the first photovoltaic power sequence and the first sky image sequence, respectively (which remain unchanged during the reverse process). and Let denoise the second photovoltaic power sequence and the second sky image sequence at step t, respectively, where t is the diffusion step number. , It is an L2 norm.
[0099] , , These correspond to three different data input-output modes: I. Input: Historical sky image sequence and historical photovoltaic power sequence; Output: Future sky image sequence and future photovoltaic power sequence; III. Input: Historical photovoltaic power sequence; Output: Future photovoltaic power sequence; II. Input: Historical sky sequence; Output: Future sky image sequence and future photovoltaic power sequence; To maintain fairness in the training process for each mode, each scenario accounts for a certain percentage of the total loss. The weight.
[0100] Example 2
[0101] Based on the same concept, this solution provides a method for constructing a multimodal diffusion model for ultra-short-term photovoltaic power probability prediction, as shown in Embodiment 1. This model is used to construct the multimodal diffusion model for ultra-short-term photovoltaic power probability prediction. Figure 4 As shown.
[0102] The application method of the multimodal diffusion model for ultra-short-term photovoltaic power probability prediction, namely, an ultra-short-term photovoltaic power probability prediction method, includes the following steps:
[0103] Input historical sky image sequences and corresponding historical photovoltaic data into the multimodal diffusion model, and output future sky image sequences and future photovoltaic power sequences;
[0104] Alternatively, input historical photovoltaic power sequences into a multimodal diffusion model and output future photovoltaic power sequences;
[0105] Alternatively, input historical sky sequences into a multimodal diffusion model and output future sky image sequences and future photovoltaic power sequences.
[0106] To verify the performance of the multimodal diffusion model for ultra-short-term photovoltaic power probability prediction, this approach uses a test dataset, which includes historical sky image sequences and historical photovoltaic power sequences. These datasets are input into the multimodal diffusion model for ultra-short-term photovoltaic power probability prediction and different benchmark models to obtain the following results: Figures 5-7 As shown, the sky image sequence corresponds to different cloud layer changes within 31 minutes: Figure 5 This is a scene that transitions from partly cloudy to overcast. Figure 6 This is a scene where the sky changes from overcast to partially cloudy. Figure 7The scene transitions from partially cloudy to sunny. While other models can generally capture the overall cloud movement, significant differences remain in image sharpness and details such as cloud texture, lighting, and shadows. These high-frequency details are crucial for solar energy prediction. Specifically, the deterministic model ConvLSTM maintains relatively clear cloud outlines at times T+1 and T+3, but from T+5 onwards, high-frequency structures and details are gradually lost, and the clouds become smoother, resulting in significantly blurred predictions. While SkyGPT maintains some sharpness and texture at longer prediction steps and accurately simulates the movement of large-scale cloud clusters, its overall color saturation is low, and it fails to capture small-scale lighting dapples (such as highlights scattered from cloud edges). In contrast, the multimodal diffusion model provided in this solution maintains higher fidelity cloud texture and lighting details throughout the entire prediction period. Figure 5 and Figure 6 In this scenario, the model not only accurately predicts the process of cloud density enhancement or reduction, but also preserves the hierarchical structure of cloud edges and the shadow movement trajectory between adjacent frames.
[0107] Furthermore, to further quantify the accuracy of the model-generated images, this approach uses the commonly used metric VGG Cosine Similarity (VGG CS) for evaluation on a validation set containing 4467 samples and a test set containing 2582 samples. VGG CS calculates the cosine similarity between the predicted result and the real image within the pre-trained VGG feature space, providing a better measure of the consistency of images in high-level perceptual semantics. Figure 8 As shown, the overall curve of VGG CS decreases as the prediction time increases from 1 minute to 15 minutes, indicating that long-term prediction is indeed challenging. However, the multimodal diffusion model in our proposal maintains the highest VGG CS value throughout the entire prediction time interval: it has higher similarity than SkyGPT and ConvLSTM in both short-term (1–3 minutes) and long-term (>7 minutes) time intervals, and its similarity decreases the slowest. This fully demonstrates the superior performance of the multimodal diffusion model in preserving high-level semantic features and achieving long-term prediction.
[0108] Example 3
[0109] This embodiment also provides an electronic device, see reference. Figure 9 It includes a memory 404 and a processor 402. The memory 404 stores a computer program, and the processor 402 is configured to run the computer program to execute the steps in the above-described method for constructing a multimodal diffusion model for ultra-short-term photovoltaic power probability prediction or the method for applying a multimodal diffusion model for ultra-short-term photovoltaic power probability prediction.
[0110] Specifically, the processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0111] The memory 404 may include a large-capacity memory 404 for data or instructions. The memory 404 can be used to store or cache various data files that need to be processed and / or communicated, as well as possible computer program instructions executed by the processor 402.
[0112] The processor 402 reads and executes the computer program instructions stored in the memory 404 to implement any of the steps in the above embodiments of the method for constructing a multimodal diffusion model for ultra-short-term photovoltaic power probability prediction or the method for applying a multimodal diffusion model for ultra-short-term photovoltaic power probability prediction.
[0113] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408, wherein the transmission device 406 is connected to the processor 402, and the input / output device 408 is connected to the processor 402.
[0114] The transmission device 406 can be used to receive or send data via a network. Specific examples of the network described above may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 406 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0115] The input / output device 408 is used to input or output information. In this embodiment, the input information may be a historical sky image sequence or a historical photovoltaic power sequence, and the output information may be a future photovoltaic power sequence, etc.
[0116] Optionally, in this embodiment, the processor 402 can be configured to perform the following steps via a computer program:
[0117] Obtain a training dataset, which includes multiple sets of training data. Each set of training data includes conditional pairing data and corresponding predicted pairing data. The conditional pairing data includes a first sky image sequence and a corresponding first photovoltaic power sequence, and the predicted pairing data includes a first sky image sequence and a corresponding second photovoltaic power sequence.
[0118] The training dataset is input into a multi-model diffusion framework to train a multimodal diffusion model. The multi-model diffusion framework is built based on the diffusion model and uses a coupled U-shaped network, which is coupled with a sky image sequence branch and a photovoltaic power sequence branch, for inverse denoising. The sky image sequence branch and the photovoltaic power sequence branch of the coupled U-shaped network share multiple coding layers, transition layers, and multiple decoding layers. Each coding layer of the sky image sequence branch has skip connections with the corresponding decoding layer, and each coding layer of the photovoltaic power sequence branch has skip connections with the corresponding decoding layer. The sky image sequence branch includes a sky image sequence encoder set before the coding layer and a sky image sequence decoder set after the decoding layer. The photovoltaic power sequence branch includes an image sequence encoder set before the coding layer and a photovoltaic power sequence decoder set after the decoding layer.
[0119] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.
[0120] Generally, various embodiments can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects of the invention can be implemented in hardware, while others can be implemented by firmware or software executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, by way of non-limiting example, these blocks, apparatuses, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0121] Embodiments of the present invention can be implemented by computer software, which may be executable by a data processor of a mobile device, such as a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. A computer program product may include one or more computer-executable components configured to perform embodiments when the program is run. One or more computer-executable components may be at least one piece of software code or a portion thereof. Additionally, it should be noted that any block in the logical flow of the figures may represent a program step, or interconnected logical circuitry, blocks and functions, or a combination of program steps and logical circuitry, blocks and functions. The software may be stored on physical media such as memory chips or blocks of storage implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs, etc. The physical medium is a non-transient medium.
[0122] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0123] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for constructing a multimodal diffusion model for ultra-short-term photovoltaic power probability prediction, characterized in that, Includes the following steps: Obtain a training dataset, which includes multiple sets of training data. Each set of training data includes conditional pairing data and corresponding predicted pairing data. The conditional pairing data includes a first sky image sequence and a corresponding first photovoltaic power sequence, and the predicted pairing data includes a second sky image sequence and a corresponding second photovoltaic power sequence. The training dataset is input into the multi-model diffusion framework to train a multimodal diffusion model. The multi-model diffusion framework is built based on the diffusion model and uses a coupled U-shaped network with a sky image sequence branch and a photovoltaic power sequence branch for inverse denoising. The sky image sequence branch and the photovoltaic power sequence branch of the coupled U-shaped network share multiple coding layers, transition layers and multiple decoding layers. Each coding layer of the sky image sequence branch has a skip connection with the corresponding decoding layer, and each coding layer of the photovoltaic power sequence branch has a skip connection with the corresponding decoding layer. The sky image sequence branch includes a sky image sequence encoder set before the coding layer and a sky image sequence decoder set after the decoding layer. The photovoltaic power sequence branch includes an image sequence encoder set before the coding layer and a photovoltaic power sequence decoder set after the decoding layer. Each encoding layer, decoding layer, and transition layer is a high-efficiency multimodal module. Each high-efficiency multimodal module includes a parallel sky image sequence sub-network, a step embedding branch, and a photovoltaic power sequence sub-network. The sky image sequence of the sky image sequence branch is input into the sky image sequence sub-network of the high-efficiency multimodal module, and the photovoltaic power sequence of the photovoltaic power sequence branch is input into the photovoltaic power sequence sub-network of the high-efficiency multimodal module. The discrete step count of the denoising process is input into the step embedding branch for linear processing and scaling offset to obtain the embedding vector. The embedding vector is applied to the sky image sequence sub-network and the photovoltaic power sequence sub-network, respectively. The output sequences of the sky image sequence sub-network and the photovoltaic power sequence sub-network are processed by a PV multimodal attention network based on random offset for cross-modal attention processing to obtain the sky image sequence of the corresponding sky image sequence branch and the photovoltaic power sequence of the corresponding photovoltaic power sequence branch.
2. The method for constructing a multimodal diffusion model for ultra-short-term photovoltaic power probability prediction according to claim 1, characterized in that, The second photovoltaic power sequence of each set of training data input into the multi-model diffusion model is denoised using the first photovoltaic power sequence as a condition to obtain a photovoltaic denoised sequence. The second sky image sequence is denoised using the first sky image sequence as a condition to obtain a sky image denoised sequence. The photovoltaic denoised sequence, the first photovoltaic power sequence, the sky image denoised sequence and the first sky image sequence are input together into a coupled U-shaped network for coupled denoising to obtain a sky image denoised sequence and a photovoltaic denoised sequence.
3. The method for constructing a multimodal diffusion model for ultra-short-term photovoltaic power probability prediction according to claim 1, characterized in that, The reverse denoising process is defined as: ; in and These represent the first photovoltaic power sequence and the first sky image sequence, respectively. and Let represent the denoising results of the photovoltaic denoised sequence and the sky image denoised sequence at step t, respectively. This represents the denoising result of the photovoltaic noise-added sequence at step t-1. For the noisy sequence of the sky image after denoising at step t-1, it is simultaneously dependent on , , and The diffusion is generated by a Gaussian distribution, where t represents the number of diffusion steps. This represents a multimodal diffusion model. The probability distribution of the reverse denoising process, where N() represents a Gaussian distribution. This represents the mean function of a Gaussian distribution.
4. The method for constructing a multimodal diffusion model for ultra-short-term photovoltaic power probability prediction according to claim 1, characterized in that, The photovoltaic power sequence subnetwork consists of a group normalization layer, a sigmoid linear unit, a one-dimensional convolution, a one-dimensional self-attention layer, and a normalization layer connected in sequence. The photovoltaic power sequence of the photovoltaic power sequence branch is input into the photovoltaic power sequence subnetwork and processed by the group normalization layer, sigmoid linear unit, one-dimensional convolution, one-dimensional self-attention layer, and normalization layer in sequence. It is then connected to the photovoltaic power sequence residual of the original input photovoltaic power sequence subnetwork and embedded in the vector input normalization layer.
5. The method for constructing a multimodal diffusion model for ultra-short-term photovoltaic power probability prediction according to claim 1, characterized in that, The sky image sequence subnetwork consists of a group normalization layer, a sigmoid linear unit, a two-dimensional spatial convolution, a one-dimensional temporal convolution, an upsampling layer, a one-dimensional to two-dimensional self-attention layer, and a normalization layer, all connected in sequence. The sky image sequence from the sky image sequence branch is input into the sky image sequence subnetwork and processed sequentially through the group normalization layer, sigmoid linear unit, two-dimensional spatial convolution, one-dimensional temporal convolution, upsampling layer, one-dimensional to two-dimensional self-attention layer, and normalization layer. After processing, it is residually connected to the sky image sequence input into the sky image sequence subnetwork and the embedded vector is input to the normalization layer.
6. The method for constructing a multimodal diffusion model for ultra-short-term photovoltaic power probability prediction according to claim 1, characterized in that, For the photovoltaic power sequence and sky image sequence input into the PV multimodal attention network based on random offset, a window smaller than the number of frames is set and the offset is randomly sampled within the window. For each photovoltaic segment with photovoltaic power sequence features, the sky image sequence segment corresponding to the current photovoltaic segment is calculated using the offset, window and frame number. The photovoltaic segment and the corresponding sky image sequence are segmented and cross-modal attention is calculated.
7. The method for constructing a multimodal diffusion model for ultra-short-term photovoltaic power probability prediction according to claim 1, characterized in that, If only historical sky image sequences are input, the historical photovoltaic power sequence is set to a zero tensor; if only historical photovoltaic power sequences are input, the historical sky image sequences are set to zero tensors. The loss function of the model coupled with the U-shaped network is as follows: ; in It is the loss function of the coupled U-shaped network. This is for noise in noisy photovoltaic power sequences. The expectation operation of the Gaussian distribution, It is noise in the photovoltaic power sequence. It is noise added to the multimodal model. and These represent the first photovoltaic power sequence and the first sky image sequence, respectively. and Let denoise the second photovoltaic power sequence and the second sky image sequence at step t, respectively, where t is the diffusion step number. , It is an L2 norm.
8. A method for probabilistic prediction of ultra-short-term photovoltaic power, characterized in that, Includes the following steps: Input historical sky image sequences and corresponding historical photovoltaic data into the multimodal diffusion model, and output future sky image sequences and future photovoltaic power sequences; Alternatively, input historical photovoltaic power sequences into a multimodal diffusion model and output future photovoltaic power sequences; Alternatively, input historical sky sequences into a multimodal diffusion model and output future sky image sequences and future photovoltaic power sequences; The multimodal diffusion model is constructed using the method described in any one of claims 1 to 7.
9. A readable storage medium, characterized in that, The readable storage medium stores a computer program, the computer program including program code for controlling a process to execute the process, the process including the method for constructing a multimodal diffusion model according to any one of claims 1 to 7 or the method for ultra-short-term photovoltaic power probability prediction according to claim 8.