Multi-mode ultra-short-term photovoltaic power generation power prediction system and method
Through the multimodal photovoltaic power generation prediction system, combined with image sequence prediction and multimodal prediction module, historical features are extracted using gated cycle and diffusion models and feature fusion, the problem of insufficient accuracy of photovoltaic power generation prediction in the existing technology is solved, and higher prediction accuracy and real-time performance are achieved.
Patent Information
- Application Number
- CN202510990983.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-08-15
AI Technical Summary
The existing ultra-short-term photovoltaic power generation prediction methods have the problem of poor prediction accuracy, especially the methods based on numerical information are difficult to reflect cloud motion, while the methods based on sky images are poor in real-time and difficult to meet the requirements of high real-time.
The multimodal ultra-short-term photovoltaic power generation prediction system is adopted to obtain and preprocess historical sky images and power generation power data through the data acquisition module. The historical image features are extracted using the gating cycle algorithm and diffusion model in the image sequence prediction module, and the feature fusion is performed through the multimodal prediction module, and the self-attention and cross-attention algorithm are used for prediction.
The accuracy and real-time prediction of photovoltaic power generation power is improved, and the generated image clarity and prediction results are more accurate, overcoming the problem of insufficient prediction accuracy in the prior art.
Smart Images

Figure CN120494222A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a multimodal ultra-short-term photovoltaic power prediction system and method. Background Art
[0002] Solar power generation relies on sunlight, which is not constant. It fluctuates due to factors such as cloud cover, time of day, and season. These fluctuations can lead to rapid and unpredictable changes in power generation, which in turn affects the stability of the power grid. Fluctuations in photovoltaic power generation are a direct indicator of grid stability. Real-time photovoltaic power forecasting can provide early insights into grid stability and mitigate the physical and economic losses caused by solar fluctuations.
[0003] Existing technologies use ultra-short-term photovoltaic power forecasting methods to leverage existing historical power generation data and other data to predict photovoltaic power generation within the short term (minutes to hours). These methods can be broadly categorized into two main types: those based on numerical data and those based on sky images or satellite remote sensing images.
[0004] Numerical prediction methods use numerical information, such as historical PV power generation data and weather forecasts, to estimate ultra-short-term future PV power generation. However, since cloud movement in the sky is difficult to reflect in numerical data, these methods struggle to produce accurate predictions.
[0005] Prediction methods based on sky images or satellite remote sensing images estimate future photovoltaic power generation based on sky images or satellite remote sensing images. However, due to the slow acquisition speed of satellite remote sensing images, they have the disadvantage of poor real-time performance, and thus cannot meet the high real-time requirements of short-term predictions. Therefore, this prediction method usually finds it difficult to fully extract information from the sky image sequence, and has the technical defect of poor prediction accuracy. Summary of the Invention
[0006] The technical problem to be solved by the present invention is how to overcome the technical defect of poor prediction accuracy in the existing ultra-short-term photovoltaic power generation prediction method. In order to overcome this technical defect, the present invention provides a multimodal ultra-short-term photovoltaic power generation prediction system and method, specifically including a multimodal ultra-short-term photovoltaic power generation prediction system and a multimodal ultra-short-term photovoltaic power generation prediction method.
[0007] The present invention provides a multi-modal ultra-short-term photovoltaic power generation power prediction system, comprising: a data acquisition module configured to pre-process the acquired historical sky image data and historical power generation data to obtain respective pre-processing results; an image sequence prediction module, electrically connected to the data acquisition module, configured to extract historical image features from preprocessing results of the historical sky image data using a gated loop algorithm, and obtain future sky camera images from the historical image features using a conditional diffusion model algorithm; The multimodal prediction module is electrically connected to the image sequence prediction module and is configured to perform feature fusion on the features extracted from the preprocessing results of the historical sky image data and the future sky camera image and the preprocessing results of the historical power generation data using a self-attention encoding and decoding algorithm, and to obtain a predicted value of photovoltaic power generation in the short term in the future using a cross-attention algorithm.
[0008] The multimodal ultra-short-term photovoltaic power prediction system disclosed in this invention utilizes an image sequence prediction module and a gated recurrent algorithm to gradually extract historical image features. This method is more suitable for processing sequence data than other methods, thereby obtaining more accurate historical motion features. Furthermore, the image sequence prediction module uses a diffusion model-based approach (i.e., a diffusion model algorithm) to predict future sky images (i.e., future sky images). Compared to other sky image prediction models, the diffusion model offers the advantages of accurate predictions and clearer generated images. Furthermore, by providing a multimodal prediction module and employing a self-attention encoding and decoding algorithm for feature fusion, the self-attention approach can better integrate features from different modalities. Meanwhile, the cross-attention layer integrates historical features with predicted sky image features. Consequently, the coordinated implementation of the image sequence prediction module and the multimodal prediction module yields more accurate power generation prediction results, thereby overcoming the technical drawback of the prior art, which suffers from poor prediction accuracy.
[0009] In a possible implementation, the data acquisition module includes: A fisheye camera, set up to acquire images of the sky in real time; a first preprocessing device electrically connected to the fisheye camera and configured to retrieve all sky images acquired by the fisheye camera during a period from a first sampling period before a current moment to the current moment to obtain the historical sky image data, and perform normalization processing on the historical sky image data using an image normalization processing method to obtain a preprocessing result of the historical sky image data; A power meter is configured to obtain photovoltaic power generation in real time; The second preprocessing device is electrically connected to the power meter and is configured to retrieve all photovoltaic power generation powers obtained by the power meter during the period from the second sampling cycle before the current moment to the current moment to obtain the historical power generation power data, and normalize the historical power generation power data through a normalization calculation formula to obtain a preprocessing result of the historical power generation power data.
[0010] The data acquisition module with the above structure can quickly acquire power and image data, and perform normalization processing according to the properties of the power and image data, thereby improving the overall processing and prediction efficiency.
[0011] In a possible implementation, the image sequence prediction module includes: a convolution unit, electrically connected to both the first preprocessing device and the second preprocessing device, configured to extract a feature map time series from the preprocessing results of the historical sky image data by a convolution operation, and then obtain motion features of the historical sky image based on differences between feature maps at adjacent times in the feature map time series; a gated recurrent unit, electrically connected to the convolution unit, and configured to obtain the historical image features from the motion features of the historical sky image by performing hidden state updates using a gated recurrent algorithm; A conditional diffusion unit is electrically connected to the gated recurrent unit and is configured to obtain future motion features using the historical image features through a neural network algorithm, and then obtain the future sky camera image using the future motion features through a conditional diffusion model algorithm.
[0012] The above scheme improves the efficiency of processing image sequence data by setting a gated recurrent unit to gradually extract historical image features based on a gated recurrent algorithm. It further improves the clarity of the generated image by setting a conditional diffusion unit to adopt a diffusion model-based method (i.e., a diffusion model algorithm) to predict future images.
[0013] In a possible implementation, a method of obtaining the future sky camera image using the future motion feature includes the following steps: A1: Define the forward noise addition process; A2: Obtaining a corresponding reverse denoising process according to the forward denoising process; A3: Generate the future sky camera image by utilizing the future motion features through the reverse denoising process.
[0014] In a possible implementation, in step A2, the inferred expression of the reverse denoising process is: , , , Where, Represents the total number of noise addition steps in the forward noise addition process; is a conditional term representing the future shown in the future motion feature The motion characteristics of the moment; Represents the output of the previous denoising step using this denoising step and conditional items Get the output of this denoising step The probability distribution of Represents the input term of the inverse denoising process and conditional items Get the future Future Sky Camera images predicted at any moment The probability distribution of represents a Gaussian distribution, where is the mean, is the standard deviation, is the identity matrix; represents the random Gaussian noise term; Represents a tensor concatenation operation.
[0015] The conditional diffusion model algorithm, which has the above-mentioned method steps and inference algorithm, generates future sky camera images by defining a forward denoising process and then a reverse denoising process. Combined with the inference calculation formula, it further improves the accuracy of the predicted image and makes the generated image clearer.
[0016] In one possible implementation, the multimodal prediction module includes: a first feature extraction unit, electrically connected to the conditional diffusion unit, and configured to extract features from the preprocessing results of the historical sky image data using a ResNet image feature extraction model algorithm to obtain historical image features; a multi-layer perception unit, electrically connected to the conditional diffusion unit, and configured to extract features from the preprocessing results of the historical power generation data using a multi-layer perception model algorithm to obtain historical power generation features; a second feature extraction unit, electrically connected to the conditional diffusion unit, and configured to extract features from the future sky camera image using a ResNet image feature extraction model algorithm to obtain predicted image features; a first encoding unit, electrically connected to both the first feature extraction unit and the multi-layer perception unit, configured to perform self-attention encoding layer mapping on the historical image features and the historical power generation features to obtain a first encoding feature; a second encoding unit, electrically connected to the second feature extraction unit, and configured to perform self-attention encoding layer mapping on the predicted image feature to obtain a second encoding feature; A cross attention unit is electrically connected to the first encoding unit and the second encoding unit at the same time, and is configured to fuse the first encoding feature and the second encoding feature through a mapping algorithm of a cross attention layer to obtain the photovoltaic power prediction value.
[0017] The multimodal prediction module with the above structure can not only better fuse features from different modalities, but also integrate historical features with predicted sky image features to obtain more accurate power generation prediction results.
[0018] In a possible implementation, the first encoding unit obtains the first encoding feature using the following calculation formula: , , , , Where, represents the first coding feature; represents the query matrix; representing said historical image features; representing the historical power generation characteristics; represents the bond matrix; Representative value matrix.
[0019] Another technical solution of the present invention is to provide a multi-modal ultra-short-term photovoltaic power generation power prediction method, comprising the following steps: S1: Acquire a historical sky image dataset and a historical power generation data set, and preprocess the acquired historical sky image dataset and the historical power generation data set to obtain respective preprocessing results; S2: constructing an image sequence prediction model based on a diffusion model, using the preprocessing results of the historical sky image dataset, and optimizing the parameters of the image sequence prediction model through a mean loss function to obtain an image sequence prediction module; S3: Construct a Transformer-based multimodal prediction model. Utilize the preprocessing results of the historical sky image dataset, the preprocessing results of the historical power generation data set, and the predicted future sky camera images. Optimize the parameters of the multimodal prediction model using a mean square error loss function to obtain a multimodal prediction module. S4: preprocessing the acquired historical sky image data and historical power generation data through the data acquisition module to obtain respective preprocessing results; S5: Obtaining a future sky camera image through the image sequence prediction module; S6: Obtaining a predicted value of photovoltaic power generation in the short term in the future through the multimodal prediction module.
[0020] The multimodal ultra-short-term photovoltaic power prediction method disclosed in the present invention constructs an image sequence prediction model based on a diffusion model, and then uses this method to predict future sky images. Compared to other sky image prediction models, the diffusion model has the characteristics of accurate predicted images and clear generated images. A Transformer-based method is used to construct a multimodal prediction model to obtain a multimodal prediction module. The self-attention layer of this method can better fuse features from different modalities. At the same time, the cross-attention layer can integrate historical features with predicted sky image features, thereby better extracting information from historical sky images. Historical power generation data is then combined with sky image data to accurately predict future photovoltaic power generation, resulting in more accurate power generation prediction results.
[0021] In a possible implementation, step S2 includes the following steps: S21: sequentially connecting a convolutional neural network, a gated recurrent unit, a neural network, and an image generation module based on a conditional diffusion model to construct the image sequence prediction model; S22: Utilizing the preprocessing results of the historical sky image dataset, the image sequence prediction model is optimized through a mean loss function to obtain an image sequence prediction module.
[0022] The above scheme uses a network based on gated neural units to construct an image sequence prediction model, which is more suitable for processing sequence data than other methods, and can therefore obtain more accurate historical motion features.
[0023] In a possible implementation, step S3 includes the following steps: S31: Set up two self-attention encoding layers and two ResNet image feature extraction models, connect the first ResNet image feature extraction model, the first self-attention encoding layer, and the multi-layer perceptron model in series, and communicate with the other self-attention encoding layer and the other ResNet image feature extraction model. And let the cross attention layer communicate with the two self-attention encoding layers at the same time to build a Transformer-based multimodal prediction model. S32: Utilizing the preprocessing results of the historical sky image dataset, the preprocessing results of the historical power generation data set, and the predicted future sky camera images, the multimodal prediction model is parameter optimized through a mean square error loss function to obtain a multimodal prediction module.
[0024] This multimodal prediction model built using a Transformer-based method has a self-attention layer that can better fuse features from different modalities, making the final power generation prediction results more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 Schematic diagram of the structure of a multi-modal ultra-short-term photovoltaic power generation prediction system disclosed in an embodiment of the present invention; Figure 2 This is a flow chart of the method disclosed in the embodiment of the present invention. DETAILED DESCRIPTION
[0026] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of the present application and are not intended to limit the scope of protection of the embodiments of the present application. Those skilled in the art may adjust them as needed to suit specific application scenarios.
[0027] In the description of the embodiments of the present application, it should be noted that, unless otherwise expressly specified or limited, the terms "electrically connected" and "establishing an electrical connection relationship" should be understood broadly, that is, it should be understood that two or more devices have an electrical relationship, which can be achieved through a wire connection, a wireless connection, or a combination of the two; and can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood in specific circumstances.
[0028] In the description of the embodiments of the present application, it should be noted that, unless otherwise clearly specified and limited, the term "forming a communication link structure" refers to the multiple communication elements or modules involved forming a network structure or a network link structure through communication connections, and communication or communication connection refers to the transmission of information between the first feature and the second feature. This information transmission can be either unidirectional or bidirectional, and the way to achieve the communication connection can be electrical connection of wires, radio connection, electrical connection of electromagnetic media (such as optical fibers, semiconductors), communication achieved by channels, etc.
[0029] At the same time, the "short term in the future" refers to the period from a few minutes to a few hours in the future. The specific mathematical expression is: ,in, Represents the current moment, , Not less than 1 minute, Less than 10 hours.
[0030] In the embodiments of the present application, unless otherwise expressly specified or limited, a first feature being "above" or "below" a second feature may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, a first feature being "above," "above," and "above" a second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. A first feature being "below," "below," and "below" a second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.
[0031] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0032] See also Figures 1 and 2 The present application discloses a multi-modal ultra-short-term photovoltaic power generation prediction system. Figure 1 This is a structural diagram of the power generation prediction system, which includes a data acquisition module, an image sequence prediction module and a multimodal prediction module, wherein the image sequence prediction module is electrically connected to the data acquisition module, and the multimodal prediction module is electrically connected to the image sequence prediction module.
[0033] See also Figure 1 In the power generation prediction system, the data acquisition module is configured to pre-process the acquired historical sky image data and historical power generation data to obtain respective pre-processing results. Figure 1 As shown, the data acquisition module includes a fisheye camera, a first preprocessing device, a power meter and a second preprocessing device, wherein the first preprocessing device is electrically connected to the fisheye camera, and the second preprocessing device is electrically connected to the power meter.
[0034] In the data acquisition module, the fisheye camera is installed on a fixed platform near the photovoltaic power station and extends vertically upward from the ground. It is configured to obtain sky images in real time. The first preprocessing device is configured to retrieve all sky images obtained by the fisheye camera from the first sampling period before the current moment to the current moment to obtain historical sky image data, and standardize the historical sky image data through image standardization processing to obtain the preprocessing results of the historical sky image data. Let's assume that the current moment is , the first sampling period is , then the period from the first sampling cycle before the current moment to the current moment is The power meter is installed in the photovoltaic power station and is configured to obtain photovoltaic power generation in real time. The second pre-processing device is configured to retrieve all photovoltaic power generation obtained by the power meter from the second sampling period before the current moment to the current moment to obtain historical power generation data, and normalize the historical power generation data through the normalization calculation formula to obtain the pre-processing result of the historical power generation data. Let's assume that the second sampling period is , then the period from the second sampling cycle before the current moment to the current moment is .
[0035] See also Figure 1 In this power generation prediction system, the image sequence prediction module is configured to extract historical image features from the preprocessing results of historical sky image data using a gated loop algorithm, and to obtain future sky camera images from the historical image features using a conditional diffusion model algorithm. In this embodiment, the image sequence prediction module includes a convolution unit, a gated loop unit, and a conditional diffusion unit, wherein the convolution unit is electrically connected to both the first preprocessing device and the second preprocessing device, the gated loop unit is electrically connected to the convolution unit, and the conditional diffusion unit is electrically connected to the gated loop unit. The specific electrical connection method is as follows: two data transmission channels are provided between the convolution unit and the gated loop unit, one responsible for transmitting the motion features of the historical sky image, and the other responsible for transmitting the preprocessing results of the historical power generation data; two data transmission channels are provided between the gated loop unit and the conditional diffusion unit, one responsible for transmitting the historical image features, and the other responsible for transmitting the preprocessing results of the historical power generation data.
[0036] In the image sequence prediction module, the convolution unit is configured to extract a feature map time series from the preprocessed results of the historical sky image data through a convolution operation, and then obtain the motion features of the historical sky image based on the difference between the feature maps at adjacent times in the feature map time series. The gated recurrent unit is configured to obtain the historical image features from the motion features of the historical sky image by updating the hidden state through a gated recurrent algorithm. The conditional diffusion unit is configured to use the historical image features to obtain future motion features through a neural network algorithm, and then use the future motion features to obtain future sky camera images through a conditional diffusion model algorithm. The neural network algorithm can be implemented using a small neural network to efficiently achieve parameter optimization.
[0037] In this embodiment, the conditional diffusion unit performs the following steps to utilize future motion features using a conditional diffusion model algorithm to obtain a future sky camera image: A1: defining a forward denoising process; A2: deriving a corresponding reverse denoising process based on the forward denoising process; A3: generating a future sky camera image using the future motion features through the reverse denoising process. In step A2, the inferred expression for the reverse denoising process is: , , , Where, Represents the total number of noise addition steps in the forward noise addition process; is a conditional term representing the future shown in the future motion feature The motion characteristics of the moment; Represents the output of the previous denoising step using this denoising step and conditional items Get the output of this denoising step The probability distribution of Represents the input term using the inverse denoising process and conditional items Get the future Future Sky Camera images predicted at any moment The probability distribution of represents a Gaussian distribution, where is the mean, is the standard deviation, is the identity matrix; and These two parameters are obtained through the U-Net model; represents the random Gaussian noise term; Represents tensor concatenation operations, for example, Representing the future Future Sky Camera images predicted at any moment and the random Gaussian noise term Tensor concatenation operation.
[0038] See also Figure 1In this power generation prediction system, a multimodal prediction module is configured to use a self-attention encoding and decoding algorithm to perform feature fusion on features extracted from the preprocessing results of historical sky image data and future sky camera images, as well as the preprocessing results of historical power generation data, and to use a cross-attention algorithm to obtain a predicted value of photovoltaic power generation in the short term. Specifically, in this embodiment, the multimodal prediction module includes a first feature extraction unit, a multi-layer perception unit, a first encoding unit, a second feature extraction unit, a second encoding unit, and a cross-attention unit. The first feature extraction unit is electrically connected to the conditional diffusion unit, the multi-layer perception unit is electrically connected to the conditional diffusion unit, the second feature extraction unit is electrically connected to the conditional diffusion unit, the first encoding unit is electrically connected to both the first feature extraction unit and the multi-layer perception unit, the second encoding unit is electrically connected to the second feature extraction unit, and the cross-attention unit is electrically connected to both the first encoding unit and the second encoding unit. The first feature extraction unit and the conditional diffusion unit are electrically connected in such a way that two data transmission channels are provided between the first feature extraction unit and the conditional diffusion unit: one channel transmits the preprocessing results of the historical sky image data, and the other channel transmits the future sky camera image.
[0039] In the multimodal prediction module, the first feature extraction unit is configured to extract features from the preprocessed results of historical sky image data and future sky camera images using the ResNet image feature extraction model algorithm to obtain image features. The multi-layer perception unit is configured to extract features from the preprocessed results of historical power generation data using the multi-layer perceptron model algorithm to obtain historical power generation features. The second feature extraction unit is configured to extract features from future sky camera images using the ResNet image feature extraction model algorithm to obtain predicted image features. The first encoding unit is configured to perform self-attention encoding layer mapping on the image features and historical power generation features to obtain first encoded features. The second encoding unit is configured to perform self-attention encoding layer mapping on the predicted image features to obtain second encoded features. The cross-attention unit is configured to fuse the first encoded features and the second encoded features using the cross-attention layer mapping algorithm to obtain a predicted photovoltaic power generation value.
[0040] Specifically in this embodiment, the calculation formula for obtaining the first coding feature by the first coding unit is as follows: , , , , Where, represents the first coding feature; represents the query matrix; Represents historical image features; Represents the historical power generation characteristics; represents the bond matrix; Representative value matrix.
[0041] For the matrix , The function definition is: .
[0042] See also Figure 2 The following will further disclose the prediction method corresponding to the multi-modal ultra-short-term photovoltaic power generation prediction system described in this embodiment. Figure 2 The overall flow chart of the method is as follows: S1: Acquire a historical sky image dataset and a historical power generation data set, and preprocess the historical sky image dataset and the historical power generation data set to obtain respective preprocessing results.
[0043] S2: Construct an image sequence prediction model based on the diffusion model, use the preprocessing results of the historical sky image dataset, and optimize the parameters of the image sequence prediction model through the mean loss function to obtain the image sequence prediction module.
[0044] Specifically in this embodiment, step S2 includes the following steps: S21: using the convolutional neural network as the convolution unit to be trained, using the combined network of the neural network and the image generation module based on the conditional diffusion model as the conditional diffusion unit to be trained, and connecting the convolutional neural network, the gated recurrent unit, the neural network and the image generation module based on the conditional diffusion model in series to construct an image sequence prediction model; S22: using the preprocessing results of the historical sky image dataset, optimizing the parameters of the image sequence prediction model through the mean loss function to obtain the image sequence prediction module. The calculation formula of the mean loss function is as follows: , Where, Represents the mean loss function value; represents the mean; represents random noise sampled from a standard Gaussian distribution; represents the noise prediction result; Represents the number of diffusion steps.
[0045] S3: Construct a Transformer-based multimodal prediction model. Utilize the preprocessing results of the historical sky image dataset, the preprocessing results of the historical power generation data set, and the predicted future sky camera images. Optimize the parameters of the multimodal prediction model using the mean square error loss function to obtain a multimodal prediction module.
[0046] Specifically in this embodiment, step S3 includes the following steps: S31: setting two self-attention coding layers as the first coding unit and the second coding unit to be trained respectively, setting two ResNet image feature extraction models as the first feature extraction unit and the second feature extraction unit to be trained respectively, setting a multi-layer perceptron model as the multi-layer perceptron unit to be trained, setting a cross-attention layer as the cross-attention unit to be trained, and connecting the first ResNet image feature extraction model, the first self-attention coding layer (i.e., the first coding unit to be trained) and the multi-layer perceptron model in series in sequence, and communicating with the other self-attention coding layer and the other ResNet image feature extraction model, and communicating with the cross-attention layer at the same time to construct a Transformer-based multimodal prediction model; S32: using the preprocessing results of the historical sky image dataset, the preprocessing results of the historical power generation data set and the predicted future sky camera image, the parameters of the multimodal prediction model are optimized by the mean square error loss function to obtain a multimodal prediction module.
[0047] S4: Preprocessing the acquired historical sky image data and historical power generation data through the data acquisition module to obtain respective preprocessing results.
[0048] It's important to note that the historical sky image dataset differs from the historical sky image data. The historical sky image dataset is retrieved from a database and contains not only historical sky image data but also random noise sampled from a standard Gaussian distribution. This random noise serves as the ideal output for the image generation module based on the conditional diffusion model. Furthermore, the historical power generation dataset originates from the database, not the data acquisition module, and serves as training data for the Transformer-based multimodal prediction model.
[0049] S5: Obtain future sky camera images through the image sequence prediction module.
[0050] S6: Obtain the predicted value of photovoltaic power generation in the short term in the future through the multimodal prediction module.
[0051] The technical effects of the multimodal ultra-short-term photovoltaic power generation prediction system of this embodiment will be described in detail below. In order to verify the effectiveness of the technical solution of the power generation prediction system, this embodiment uses 7 months of photovoltaic power generation data and sky image data of a certain area as the original data set for experiment. The experiment is implemented based on Python 3.8 and Pytorch 1.12, and the parameter optimization process is performed on the Nvidia A100 GPU. The input of the model is the photovoltaic power generation data and sky image data within the past 2 hours, and the output is the photovoltaic power generation power in the next 2 hours. The resolution of the input and output is 15 minutes. The methods involved in the comparison include a method based on numerical time series prediction (hereinafter referred to as method 1); a Transformer-based multimodal prediction model after removing the diffusion model-based image prediction model (hereinafter referred to as method 2); and a method after replacing the Transformer-based multimodal prediction method in this technical solution with an image regression model based on a convolutional neural network (hereinafter referred to as method 3). The experiment uses RMSE and MAE as evaluation indicators, and their calculation methods are: , , In the formula is the number of predicted points, For the The predicted value of photovoltaic output power, Table 1 shows the comparison results of this technical solution and other methods.
[0052] Table 1
[0053] Experimental results show that the technical solution of this embodiment has the best prediction effect compared to other methods. Compared with Method 1, the technical solution of this embodiment introduces sky image data into the ultra-short-term photovoltaic power generation prediction task, resulting in better prediction results, with an improvement of 9.3% in RMSE and 9.4% in MAE. Compared with Method 2, the technical solution of this embodiment uses a Transformer-based multimodal prediction model that takes data from multiple modalities as input, thereby simultaneously utilizing information from multiple aspects, resulting in an improvement of 22.4% in RMSE and 28.71% in MAE. Compared with Method 3, the technical solution of this embodiment uses a diffusion model-based image prediction model to predict future sky images, thereby better capturing the information in image sequence data, resulting in an improvement of 4.2% in RMSE and 6.5% in MAE. Therefore, the technical solution proposed in this embodiment has the best performance in the ultra-short-term photovoltaic power generation prediction problem.
[0054] In summary, the multimodal ultra-short-term photovoltaic power generation prediction system disclosed in this embodiment, by setting up an image sequence prediction module, adopts a gated loop algorithm to gradually extract historical image features. Compared with other methods, it is more suitable for processing sequence data, and thus can obtain more accurate historical motion features. In addition, the image sequence prediction module also adopts a diffusion model-based method (i.e., a diffusion model algorithm) to predict future sky images. Compared with other sky image prediction models, the diffusion model has the characteristics of accurate predicted images and clear generated images. Moreover, by setting up a multimodal prediction module and performing feature fusion with self-attention encoding and decoding algorithms, the self-attention method can better fuse features from different modalities. At the same time, the cross-attention layer can integrate historical features with predicted sky image features, thereby obtaining more accurate power generation prediction results.
[0055] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms such as "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present application.
[0056] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "in the present embodiment", "specific example", or "some examples" means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.
[0057] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A multi-modal ultra-short-term photovoltaic power generation prediction system, characterized in that: include: a data acquisition module configured to pre-process the acquired historical sky image data and historical power generation data to obtain respective pre-processing results; an image sequence prediction module, electrically connected to the data acquisition module, configured to extract historical image features from preprocessing results of the historical sky image data using a gated loop algorithm, and obtain future sky camera images from the historical image features using a conditional diffusion model algorithm; The multimodal prediction module is electrically connected to the image sequence prediction module and is configured to perform feature fusion on the features extracted from the preprocessing results of the historical sky image data and the future sky camera image and the preprocessing results of the historical power generation data using a self-attention encoding and decoding algorithm, and to obtain a predicted value of photovoltaic power generation in the short term in the future using a cross-attention algorithm.
2. The multi-modal ultra-short-term photovoltaic power generation prediction system according to claim 1, characterized in that: The data acquisition module includes: A fisheye camera, set up to acquire images of the sky in real time; a first preprocessing device electrically connected to the fisheye camera and configured to retrieve all sky images acquired by the fisheye camera during a period from a first sampling period before a current moment to the current moment to obtain the historical sky image data, and perform normalization processing on the historical sky image data using an image normalization processing method to obtain a preprocessing result of the historical sky image data; A power meter is configured to obtain photovoltaic power generation in real time; The second preprocessing device is electrically connected to the power meter and is configured to retrieve all photovoltaic power generation powers obtained by the power meter during the period from the second sampling cycle before the current moment to the current moment to obtain the historical power generation power data, and normalize the historical power generation power data through a normalization calculation formula to obtain a preprocessing result of the historical power generation power data.
3. The multi-modal ultra-short-term photovoltaic power generation prediction system according to claim 2, characterized in that: The image sequence prediction module includes: a convolution unit, electrically connected to both the first preprocessing device and the second preprocessing device, configured to extract a feature map time series from the preprocessing results of the historical sky image data by a convolution operation, and then obtain motion features of the historical sky image based on differences between feature maps at adjacent times in the feature map time series; a gated recurrent unit, electrically connected to the convolution unit, and configured to obtain the historical image features from the motion features of the historical sky image by performing hidden state updates using a gated recurrent algorithm; A conditional diffusion unit is electrically connected to the gated recurrent unit and is configured to obtain future motion features using the historical image features through a neural network algorithm, and then obtain the future sky camera image using the future motion features through a conditional diffusion model algorithm.
4. The multi-modal ultra-short-term photovoltaic power generation prediction system according to claim 3, characterized in that: The method of obtaining the future sky camera image by using the future motion feature comprises the following steps: A1: Define the forward noise addition process; A2: Obtaining a corresponding reverse denoising process according to the forward denoising process; A3: Generate the future sky camera image by utilizing the future motion features through the reverse denoising process.
5. The multi-modal ultra-short-term photovoltaic power generation prediction system according to claim 4, characterized in that: In step A2, the inferred expression of the reverse denoising process is: , , , Where, Represents the total number of noise addition steps in the forward noise addition process; is a conditional term representing the future shown in the future motion feature The motion characteristics of the moment; Represents the output of the previous denoising step using this denoising step and conditional items Get the output of this denoising step The probability distribution of Represents the input term of the inverse denoising process and conditional items Get the future Future Sky Camera images predicted at any moment The probability distribution of represents a Gaussian distribution, where is the mean, is the standard deviation, is the identity matrix; represents the random Gaussian noise term; Represents a tensor concatenation operation.
6. The multi-modal ultra-short-term photovoltaic power generation prediction system according to any one of claims 3 to 5, characterized in that: The multimodal prediction module includes: a first feature extraction unit, electrically connected to the conditional diffusion unit, and configured to extract features from the preprocessing results of the historical sky image data using a ResNet image feature extraction model algorithm to obtain historical image features; a multi-layer perception unit, electrically connected to the conditional diffusion unit, and configured to extract features from the preprocessing results of the historical power generation data using a multi-layer perception model algorithm to obtain historical power generation features; a second feature extraction unit, electrically connected to the conditional diffusion unit, and configured to extract features from the future sky camera image using a ResNet image feature extraction model algorithm to obtain predicted image features; a first encoding unit, electrically connected to both the first feature extraction unit and the multi-layer perception unit, configured to perform self-attention encoding layer mapping on the historical image features and the historical power generation features to obtain a first encoding feature; a second encoding unit, electrically connected to the second feature extraction unit, and configured to perform self-attention encoding layer mapping on the predicted image feature to obtain a second encoding feature; A cross attention unit is electrically connected to the first encoding unit and the second encoding unit at the same time, and is configured to fuse the first encoding feature and the second encoding feature through a mapping algorithm of a cross attention layer to obtain the photovoltaic power prediction value.
7. The multi-modal ultra-short-term photovoltaic power generation prediction system according to claim 6, characterized in that: The calculation formula for obtaining the first coding feature by the first coding unit is as follows: , , , , Where, represents the first coding feature; represents the query matrix; representing said historical image features; representing the historical power generation characteristics; represents the bond matrix; Representative value matrix.
8. A multi-modal ultra-short-term photovoltaic power generation power prediction method, characterized in that: The multimodal ultra-short-term photovoltaic power generation prediction system according to any one of claims 1 to 7 comprises the following steps: S1: Acquire a historical sky image dataset and a historical power generation data set, and preprocess the acquired historical sky image dataset and the historical power generation data set to obtain respective preprocessing results; S2: constructing an image sequence prediction model based on a diffusion model, using the preprocessing results of the historical sky image dataset, and optimizing the parameters of the image sequence prediction model through a mean loss function to obtain an image sequence prediction module; S3: Construct a Transformer-based multimodal prediction model. Utilize the preprocessing results of the historical sky image dataset, the preprocessing results of the historical power generation data set, and the predicted future sky camera images. Optimize the parameters of the multimodal prediction model using a mean square error loss function to obtain a multimodal prediction module. S4: preprocessing the acquired historical sky image data and historical power generation data through the data acquisition module to obtain respective preprocessing results; S5: Obtaining a future sky camera image through the image sequence prediction module; S6: Obtaining a predicted value of photovoltaic power generation in the short term in the future through the multimodal prediction module.
9. The multi-modal ultra-short-term photovoltaic power generation prediction method according to claim 8, characterized in that: The step S2 comprises the following steps: S21: sequentially connecting a convolutional neural network, a gated recurrent unit, a neural network, and an image generation module based on a conditional diffusion model to construct the image sequence prediction model; S22: Utilizing the preprocessing results of the historical sky image dataset, the image sequence prediction model is optimized through a mean loss function to obtain an image sequence prediction module.
10. The multi-modal ultra-short-term photovoltaic power generation prediction method according to claim 9, characterized in that: The step S3 comprises the following steps: S31: setting two self-attention coding layers and two ResNet image feature extraction models, connecting the first ResNet image feature extraction model, the first self-attention coding layer, and the multi-layer perceptron model in series, communicating another self-attention coding layer with another ResNet image feature extraction model, and allowing the cross attention layer to communicate with the two self-attention coding layers at the same time to construct the Transformer-based multimodal prediction model; S32: Utilizing the preprocessing results of the historical sky image dataset, the preprocessing results of the historical power generation data set, and the predicted future sky camera images, the multimodal prediction model is parameter optimized through a mean square error loss function to obtain a multimodal prediction module.
Citation Information
Patent Citations
Solar photovoltaic power generation short-term prediction method based on sky image
CN117410976A
Photovoltaic power generation power prediction method and system based on domain knowledge embedding model
CN117477551A
Traffic video anomaly detection method and system based on dual-condition diffusion model
CN117789087A
Ultra-short-term photovoltaic power prediction method and device based on random skyline video prediction and medium
CN120181281A
Photovoltaic power ultra-short-term probability prediction method based on multi-source spatio-temporal information representation
CN120259682A
Cited By
Photovoltaic power uncertainty prediction method and system
CN121688870A
Photovoltaic power uncertainty prediction method and system
CN121688870B
Multi-mode ultra-short-term photovoltaic power generation power prediction method and system
CN121688877A
A multi-modal ultra-short-term photovoltaic power generation power prediction method and system
CN121688877B