Ultra-short-term photovoltaic power prediction method and system based on physical constraint and multi-source data space-time fusion

By combining ViT and TCN with cross-attention mechanism and optical flow technology, multi-source data features of cloud maps and historical photovoltaic power are extracted, and various physical constraint losses are constructed. This solves the problem that existing photovoltaic power prediction methods do not accurately describe cloud movement on a minute-level time scale, and achieves high-precision and robust ultra-short-term photovoltaic power prediction.

CN121923095APending Publication Date: 2026-04-24SHANGHAI UNIVERSITY OF ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI UNIVERSITY OF ELECTRIC POWER
Filing Date
2025-12-29
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing photovoltaic power prediction methods are unable to accurately describe cloud movement on a minute-level timescale, lack the ability to predict sudden changes, and lack physical constraints in multi-source data fusion, resulting in unstable prediction results and significant biases, making it difficult to meet the high accuracy and robustness requirements of high-proportion photovoltaic grid connection.

Method used

We employ a visual Transformer (ViT) and temporal convolutional network (TCN) to extract multi-source data features from cloud images and historical photovoltaic power. By combining cross-attention mechanism and optical flow technology, we extract physical quantities of cloud motion, construct various physical constraint losses, and adjust the weights through Bayesian optimization to achieve ultra-short-term photovoltaic power prediction.

Benefits of technology

It significantly improves the accuracy and stability of minute-level photovoltaic power prediction, enhances the physical consistency of prediction results, and is suitable for grid dispatch and new energy consumption scenarios, especially maintaining the reliability and robustness of predictions under conditions of rapid cloud movement or sudden changes in irradiance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121923095A_ABST
    Figure CN121923095A_ABST
Patent Text Reader

Abstract

The invention relates to an ultra-short-term photovoltaic power prediction method and system based on physical constraint and multi-source data space-time fusion, and the method comprises the steps: carrying out the feature extraction through ViT, and obtaining a cloud picture feature; performing feature extraction by using TCN to obtain time sequence features; fusing the cloud picture features and the time sequence features based on a cross attention mechanism to obtain predicted photovoltaic power, and calculating basic prediction loss; extracting an instantaneous motion vector field of the cloud picture data to calculate the acceleration, divergence and optical flow edge intensity of the cloud picture, and matching a grade coefficient corresponding to each physical quantity; on the premise that only the loss corresponding to the single physical quantity is activated, training is carried out, and the optimal weight coefficient of each physical loss is obtained; taking the optimal weight coefficient of each physical loss as an initial point of a search space, and obtaining an optimal combined weight coefficient through Bayesian search; and physical constraint loss is calculated, and training of the ultra-short-term photovoltaic power prediction model is realized in combination with the basic prediction loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of photovoltaic power generation technology, and in particular to an ultra-short-term photovoltaic power prediction method and system based on physical constraints and spatiotemporal fusion of multi-source data. Background Technology

[0002] With the continuous expansion of photovoltaic (PV) installations, solar power generation has become an important component of the power system. However, PV power is significantly affected by moving cloud formations, especially on a minute-by-minute timescale. Cloud shading of solar radiation can cause rapid, drastic, and unpredictable fluctuations in irradiance and output power. These sudden fluctuations increase the pressure on grid frequency regulation and peak shaving, affecting the safe and stable operation of the grid and placing higher demands on the high proportion of PV grid connection.

[0003] Existing photovoltaic power forecasting methods mainly rely on historical power series, on-site meteorological data, and numerical weather prediction (NWP). However, NWP has inherent limitations in terms of temporal resolution, spatial accuracy, and update frequency, making it difficult to capture minute-level cloud movement characteristics. Statistical models or deep learning models based on historical series also struggle to cope with the non-stationarity caused by rapid and random changes in clouds. Under weather conditions such as cloudy skies, intermittent clouds, and low-level radiative disturbances, irradiance changes are often unrelated to historical trends, making traditional ultra-short-term forecasting models unstable on minute-scale operations and significantly increasing prediction bias.

[0004] With the development of ground-based imaging systems such as all-sky cloud images, cloud-based photovoltaic (PV) forecasting has become an important technological direction for addressing minute-level fluctuations. Cloud images can characterize cloud location, thickness, boundary morphology, and direction of movement in real time, providing highly relevant prior information for forecasting. However, cloud images have strong spatiotemporal coupling and a highly dynamic structure. The movement patterns of cloud clusters vary significantly under different weather conditions, making data-driven deep learning models prone to physical inconsistencies, prediction drift, or insufficient generalization ability.

[0005] Traditional models fail to explicitly utilize these physical constraints, making it difficult to maintain the stability and reliability of predictions under complex cloud conditions.

[0006] In summary, existing methods for minute-level photovoltaic power prediction still suffer from problems such as imprecise description of cloud motion, lack of abrupt change prediction capabilities, and lack of physical constraints in multi-source data fusion. These limitations make it difficult to meet the high-precision and robust ultra-short-term prediction requirements under conditions of high-proportion photovoltaic grid connection. There is an urgent need for an ultra-short-term photovoltaic power prediction method that can combine multi-source data and fully consider the physical characteristics of clouds for constraints. Summary of the Invention

[0007] The purpose of this invention is to overcome the defects of the prior art by providing an ultra-short-term photovoltaic power prediction method and system based on physical constraints and spatiotemporal fusion of multi-source data, so as to solve or partially solve the problems of imprecise cloud movement description, lack of abrupt change prediction capability, and lack of physical constraints in multi-source data fusion.

[0008] The objective of this invention can be achieved through the following technical solutions: One aspect of the present invention provides an ultra-short-term photovoltaic power prediction method based on physical constraints and spatiotemporal fusion of multi-source data. Cloud map data and historical photovoltaic power data are used as inputs to an ultra-short-term photovoltaic power prediction model to obtain ultra-short-term photovoltaic power prediction results. The training process of the ultra-short-term photovoltaic power prediction model includes the following steps: Obtain cloud map data, use ViT to extract features, and obtain cloud map features; Historical photovoltaic power data is acquired, and time-series features are extracted using TCN. The predicted photovoltaic power is obtained by fusing the cloud map features and time series features based on the cross-attention mechanism, and the basic prediction loss is calculated. Extract the instantaneous motion vector field of the cloud map data to calculate multiple physical quantities of the cloud map, including acceleration, divergence and optical flow edge intensity, and match the level coefficient corresponding to each physical quantity; Under the premise of activating only the loss corresponding to a single physical quantity, training is performed separately to obtain the optimal weight coefficients for each physical loss; The optimal weight coefficients of each physical loss are used as the initial points of the search space, and the optimal combined weight coefficients are obtained through Bayesian search. The physical constraint loss is calculated based on the grade coefficient and the optimal combination weight coefficient, and the training of the ultra-short-term photovoltaic power prediction model is achieved by combining the basic prediction loss.

[0009] As a preferred technical solution, the process of obtaining cloud map features includes the following steps: The preprocessed cloud image data is divided into multiple image blocks; For each image patch, flattening and linear projection are performed to form an embedding sequence with position encoding; The embedding sequence of each image patch is input into ViT, and the corresponding deep spatial features are extracted through a multi-head attention mechanism; For the deep spatial features of each image patch, average pooling is used to obtain the global spatial feature vector that characterizes the entire cloud map spatial structure, i.e., the cloud map features.

[0010] As a preferred technical solution, the process of obtaining the time-series features includes the following steps: The historical photovoltaic power data is normalized and length-aligned to form a fixed-length time series. The time series is input into a causal convolutional layer, and causal padding is performed on the right side of the convolutional kernel so that the convolutional operation depends only on the current and past power values ​​at any time point. Dilated convolution is performed on causal convolution to form a dilated convolution module. The receptive field of the convolution is expanded by progressively increasing dilation coefficients to capture power change features at different time scales. The dilation coefficients grow exponentially. In each dilated convolutional module, a normalization layer, a non-linear activation function, and a Dropout regularization structure are set sequentially, and the input features and output features are added through residual connections; The time-step features extracted by multiple dilated convolutional modules are aggregated by global average pooling to obtain a global temporal feature vector that characterizes the historical power evolution pattern, i.e., the temporal features.

[0011] As a preferred technical solution, the process of obtaining the predicted photovoltaic power includes the following steps: Perform linear projection on the cloud map features and the time series features; The time-series features after linear projection are used as the Query, and the features after linear projection are used as the Key and Value of the cloud map features. Attention weights are calculated through cross-attention. The cloud map features after linear projection are weighted and summed based on the attention weights to obtain the fused features; The fused features are then subjected to residual connections and layer normalization. Based on the processed fusion characteristics, the predicted photovoltaic power is obtained using a fully connected prediction head.

[0012] As a preferred technical solution, the process of calculating acceleration, divergence, and optical flow edge intensity includes the following steps: Based on the cloud image data, the input cloud image is converted into a grayscale image and then subjected to Gaussian filtering for smoothing. Based on the preset displacement of adjacent cloud maps, the grayscale values ​​after displacement are expanded. Based on the grayscale value expansion, the instantaneous motion vector field of the cloud map is extracted; Based on the instantaneous motion vector field, the acceleration, divergence, and optical flow edge intensity of the cloud map are calculated respectively.

[0013] As a preferred technical solution, the process of obtaining the optimal weight coefficients for each physical loss includes the following steps: A new loss function is constructed by summing the basic prediction loss and the physical loss. The acceleration loss, divergence loss and boundary loss are trained independently. The optimal weight coefficients for each physical quantity are obtained by activating only a single physical loss and freezing the other physical losses.

[0014] As a preferred technical solution, the process of obtaining the optimal combination weight coefficients includes the following steps: We construct a continuous search space for physical loss weights by taking the optimal weights for acceleration loss, divergence loss, and boundary loss as the initial points of the Bayesian optimization search space. The predictive performance of the model on the validation set is used as the objective function of Bayesian optimization, and the next set of candidate weights is selected based on the preset acquisition function. For each set of candidate weights, the model is trained and validated. Feedback and updates are performed based on the validation error. After multiple rounds of iterative optimization, the physical loss weight vector that optimizes the prediction performance of the validation set is output as the optimal combination of weight coefficients for the final physical loss function.

[0015] As a preferred technical solution, the total loss function of the ultra-short-term photovoltaic power prediction model is: in, , , , , These are the total loss, basic prediction loss, acceleration loss, divergence loss, and boundary loss, respectively. , , These are the optimal weighting coefficients for acceleration loss, divergence loss, and boundary loss, respectively. , , These are the level coefficients for acceleration, divergence, and boundary loss, respectively.

[0016] As a preferred technical solution, the acceleration loss, divergence loss, and boundary loss are as follows: in, and These represent the instantaneous motion vector field at... and acceleration components in the direction, It is a very small positive value. This represents the mathematical expectation calculated over all pixels. ,for The velocity component in the direction, ,for The velocity component in the direction, This indicates variance calculation. This represents the set of velocities of the instantaneous motion vector field at the image boundary pixels. It is the instantaneous motion vector field.

[0017] Another aspect of the present invention provides an ultra-short-term photovoltaic power prediction system based on physical constraints and spatiotemporal fusion of multi-source data, characterized in that, for implementing the aforementioned ultra-short-term photovoltaic power prediction method, the system comprises: The cloud map feature extraction module is used to acquire cloud map data and extract features using ViT to obtain cloud map features; The historical power time-series feature extraction module is used to acquire historical photovoltaic power data, extract features using TCN, and obtain the time-series features. The multimodal cross-fusion module is used to fuse the cloud map features and temporal features based on the cross-attention mechanism to obtain the predicted photovoltaic power and calculate the basic prediction loss. The optical flow and physical quantity extraction module is used to extract the instantaneous motion vector field of the cloud image data to calculate multiple physical quantities of the cloud image, including acceleration, divergence, and optical flow edge intensity. The cloud condition level assessment module is used to match the level coefficient corresponding to each physical quantity; The physical consistency loss construction module is used to construct the loss function for each physical quantity; The physical loss weight optimization module is used to train the system separately, under the premise of activating only the loss corresponding to a single physical quantity, to obtain the optimal weight coefficients for each physical loss. The adaptive weight adjustment module is used to take the optimal weight coefficients of each physical loss as the initial points of the search space and obtain the optimal combined weight coefficients through Bayesian search. The joint training and prediction module is used to calculate the physical constraint loss based on the rank coefficients and the optimal combined weight coefficients, and to train the ultra-short-term photovoltaic power prediction model by combining the basic prediction loss.

[0018] Compared with the prior art, the present invention has at least one of the following beneficial effects: (1) This invention achieves synchronous modeling of spatial cloud structure and historical power variation patterns by constructing a cloud map feature extraction module based on visual ViT and a power temporal feature extraction module based on temporal convolutional networks. Compared with existing technologies that rely on a single data source or low-dimensional feature expression, this invention can extract more comprehensive spatiotemporal features from cloud morphology, boundary structure, shading gradient and multi-scale power changes, providing refined data support for minute-level photovoltaic power prediction.

[0019] (2) This invention introduces a cross-attention mechanism to achieve deep fusion of cloud image features and power features, and combines optical flow technology to extract key physical quantities such as cloud motion velocity, acceleration, divergence, and boundary changes, and constructs a variety of physical constraints such as acceleration smoothing loss, divergence consistency loss, and boundary structure constraints. Compared with the traditional method that relies solely on deep neural networks for automatic learning, this invention can significantly enhance the physical consistency of prediction results, effectively suppress instantaneous prediction drift, and improve prediction stability and reliability under conditions of rapid cloud movement or sudden changes in irradiance.

[0020] (3) This invention employs Bayesian optimization to automatically optimize physical loss weights and combines them with an adaptive weight adjustment mechanism based on cloud condition levels, enabling the model to achieve optimal physical constraint strength and fusion strategy under different weather scenarios such as clear skies, thin clouds, cloudy skies, and fast-moving clouds. Simultaneously, this invention achieves joint driving of multi-source data and physical models, resulting in significantly better prediction accuracy, anti-interference capabilities, and generalization ability than existing technologies. It is particularly suitable for grid dispatching and new energy consumption scenarios involving minute-level ultra-short-term photovoltaic power prediction. Attached Figure Description

[0021] Figure 1 This is a flowchart of the ultra-short-term photovoltaic power prediction method based on physical constraints and spatiotemporal fusion of multi-source data in the embodiment. Figure 2 This is a schematic diagram of the ViT model structure used to extract spatial features of future cloud maps in the embodiment. Figure 3 This is a schematic diagram of the TCN model structure used to extract historical power data in the embodiment; Figure 4 This is a schematic diagram of the feature fusion process in the embodiment; Figure 5 This is a schematic diagram of an ultra-short-term photovoltaic power prediction system based on physical constraints and spatiotemporal fusion of multi-source data in the embodiment. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0023] Example 1 To address the problems of the aforementioned existing technologies, this embodiment provides an ultra-short-term photovoltaic power prediction method based on physical constraints and spatiotemporal fusion of multi-source data. Utilizing multi-source information such as sky images, optical flow physical quantities, and historical power data, it extracts cloud structure, movement speed, and change trends through a spatiotemporal deep network. Combined with physical constraints such as velocity continuity and acceleration smoothness, it achieves short-term extrapolation of future sky conditions. Subsequently, a cross-modal fusion mechanism is employed to jointly model sky image features and power features, thereby completing minute-level photovoltaic power prediction. This method can accurately capture irradiance fluctuations caused by rapid cloud changes and maintains stable prediction performance under complex weather conditions such as cloudy skies and rapid cloud movement. Compared to traditional models that rely solely on historical sequences or single images, this embodiment has significant advantages in prediction accuracy, stability, and physical consistency, providing more reliable technical support for photovoltaic grid-connected scheduling and new energy consumption.

[0024] The main principle of this invention is as follows: In terms of multi-source information preprocessing, considering the strong non-stationarity and significant random perturbations of minute-level photovoltaic power sequences, this invention simultaneously utilizes sky cloud images, optical flow fields, and historical power data to construct multi-source inputs. Specifically, the sky image extracts spatial structure features using a visual Transformer (ViT), optical flow characterizes cloud velocity, acceleration, and motion trends, and the historical power sequence utilizes a temporal convolutional network (TCN) to extract multi-scale temporal dependencies. This multi-source fusion strategy can distinguish between cloud driving factors and power change ontological information at the model input stage, providing physically consistent and lower-noise spatiotemporal features for subsequent predictions.

[0025] Regarding the construction of physical constraints, cloud motion conforms to physical characteristics such as continuity, smoothness, and conservation, which are difficult for traditional deep learning models to explicitly capture. This invention proposes a physical quantity extraction and constraint mechanism based on optical flow, including physical features such as cloud velocity field, acceleration field, divergence, and boundary structure. Acceleration loss, divergence loss, and boundary consistency loss are constructed to form a physical consistency constraint module specifically for photovoltaic scenarios. By introducing these physical features into the model, the prediction results are guided to maintain a reasonable cloud evolution trend in the future, reducing prediction drift caused by non-compliance with physical features.

[0026] In terms of predictive model construction, this invention proposes a spatiotemporal deep fusion model based on cross-modal interactive attention. This model dynamically couples cloud image features and power features through a cross-attention mechanism, enabling the model to simultaneously focus on both spatial cloud variations and temporal power dependence. Building upon this, an adaptive learning strategy for physical constraint weights is introduced. Bayesian optimization is used to determine the optimal combination coefficients of multiple physical losses, and the physical loss weights are dynamically adjusted based on different cloud condition levels, constructing a closed-loop system of "physical drive – data fusion – adaptive optimization." This mechanism significantly improves the model's robustness and predictive reliability under scenarios such as strong weather disturbances and rapid cloud movement.

[0027] See Figure 1 This method includes the following steps: S1. Spatial features are extracted from the input cloud map using the ViT model.

[0028] In this embodiment, a Visual Transformer (ViT) is used to extract high-dimensional spatial features from the input all-sky cloud image. The shape, thickness, cloud top structure, boundary texture, and local occlusion patterns of the clouds in the original cloud image are encoded into a unified vectorized representation. See [link to documentation]. Figure 2 For the ViT model graph used to extract spatial features of future cloud maps, feature extraction specifically includes the following sub-steps: S11. Preprocess the input raw cloud image by uniformly cropping or scaling it to a fixed size, and dividing the cloud image into several non-overlapping small image patches according to the preset patch size. The original cloud image is a three-channel RGB image with a size of [missing information]. The number of patches obtained after partitioning is: In the formula: Image height, Image width, The number of image channels. For the set image block size, This represents the total number of image blocks after partitioning. S12. Flatten each image patch and project it onto a fixed-dimensional vector space using a linear mapping to form a patch embedding sequence. Introduce a learnable positional code for each patch to preserve the spatial structure information of the cloud image. The linear mapping is performed according to the following formula: In the formula: Indicates the first i The embedding vector of each patch has a dimension of . d . W This represents the weight matrix, with the shape as follows: ,inP The height / width of the patch. C This represents the number of channels. Indicates the first i Each patch is flattened into a one-dimensional vector. b This represents the bias vector, with dimension . d This way, each patch becomes a... d A dimensional vector, all patches are concatenated into one. The characteristic sequence; S13. The patch embedding sequence with added position encoding is input into a multi-layer Transformer encoder structure. The global spatial relationship between different patches is modeled through a multi-head self-attention mechanism to extract deep spatial features of the cloud map. The multi-head self-attention module is used to calculate the correlation between patches, and the encoder employs residual connections and layer normalization structures to enhance training stability. S14. Perform average pooling on all patch features output by the Transformer encoder along the sequence dimension to obtain a global spatial feature vector representing the entire cloud map spatial structure, which serves as the spatial feature input for the subsequent multimodal fusion module. Through average pooling, the feature dimensionality can be reduced while preserving global spatial information, improving the stability and effectiveness of the subsequent fusion process.

[0029] S2. Use the TCN model to extract time-series features from historical power data.

[0030] In this embodiment, a Temporal Convolutional Network (TCN) is used to extract multi-scale time-series patterns from historical photovoltaic power sequences. Dilated convolution and residual structures are used to enhance the joint perception of short-term disturbances and long-term trends. (See [link to TCN]). Figure 3 The feature extraction process for the TCN model used to extract historical power data includes the following sub-steps: S21. Normalize and length-align the input historical power sequence to construct a fixed-length time series input, so as to ensure that the power data of different time periods participate in time series feature modeling under a unified scale and structure. S22. Input the historical power sequence into the causal convolutional layer. By performing causal padding on the right side of the convolutional kernel, the convolutional operation depends only on the current and past power values ​​at any point in time, thereby avoiding future information leakage and satisfying the requirement of temporal causality. S23. Based on causal convolution, an expanded convolutional structure is introduced. The receptive field of the convolution is expanded by progressively increasing expansion coefficients to capture power change features at different time scales. The expansion coefficients grow exponentially, enabling the output features to cover a longer time range and effectively model the multi-scale dynamics of photovoltaic power without increasing the network depth. S24. In each dilated convolutional module, a normalization layer, a non-linear activation function, and a Dropout regularization structure are set in sequence, and the input features and output features are added by residual connection to enhance the stability of network training and alleviate the gradient vanishing problem. S25. The time-step features extracted by the multi-layer dilated convolutional network are aggregated by global average pooling to obtain a global temporal feature vector that can characterize the historical power evolution law, which serves as the temporal feature input for the subsequent multimodal fusion module.

[0031] S3. Introduce a cross-attention mechanism to fuse multimodal features for preliminary future power prediction.

[0032] In this embodiment, a cross-modal cross-attention mechanism is used to jointly model cloud image feature sequences and power feature sequences. This allows the model to simultaneously focus on the temporal dependence of cloud spatial changes and historical power, thereby forming a more physically meaningful spatiotemporal fusion representation. See also Figure 4 This is a schematic diagram of the feature fusion process, which specifically includes the following sub-steps: S31. The spatial features extracted from the cloud map and the temporal features extracted from the historical power are respectively mapped to a feature space of the same dimension through linear projection to ensure that the two modes have a consistent vector structure when feature fusion. S32. The projected temporal features are used as the Query, and the projected spatial features are used as the Key and Value. These are input into the cross-attention module. The correlation between the two modal features is calculated through the attention score. The attention score calculation process is as follows: In the formula: The query matrix is ​​composed of time-series features. To query the number of steps, , These are the key matrix and value matrix, respectively, composed of spatial features. , These are the key steps and value steps, respectively. Let be the dimension of the matrix vector. Scaling factor Normalized attention score on the key dimension; S33. Based on attention weights, spatial features are weighted and summed to obtain a spatial context vector that is highly correlated with temporal features, thereby achieving deep coupling between spatial and temporal information; S34. The output of cross-attention is stabilized through residual connection and layer normalization to enhance the continuity of information flow during fusion and improve the convergence performance of model training. S35. Input the stabilized fused features into the fully connected prediction head to generate preliminary power prediction results for the future time range, providing basic prediction information for the subsequent introduction of physical constraints.

[0033] S4. Extract the instantaneous motion vector field (optical flow) of the cloud image to characterize the physical characteristics of cloud motion.

[0034] In this embodiment, optical flow estimation is performed on consecutive frames of sky cloud images to obtain the instantaneous velocity field, orientation field, and local deformation information of the clouds. Optical flow, as a direct observable of cloud dynamics, can describe physical processes such as cloud translation, diffusion, aggregation, and boundary changes, providing a foundation for the subsequent construction of physical loss. Specifically, it includes the following sub-steps: S41. Preprocess the cloud images of two adjacent frames by converting the input cloud images into grayscale images and smoothing them using Gaussian filtering to reduce noise interference and provide stable input for subsequent optical flow estimation. For grayscale images, the grayscale value of each pixel can be regarded as a two-dimensional variable function. And approximately a quadratic polynomial expansion: In the formula: For pixel coordinates, It is A symmetric matrix, for gradient vector, The value is the luminance constant. S42. Set the displacement between adjacent cloud maps. By calculating the displacement per unit time, we can obtain the formula for expanding the grayscale value after displacement: In the formula: ,for The velocity component in the direction, ,for The velocity component in the direction, The coefficient matrix after displacement. The vector after displacement. This is the brightness constant after displacement.

[0035] S43. Extract the instantaneous motion vector field (optical flow) of the cloud map and use it as a key input to characterize the dynamic features of cloud motion, providing constraints for the subsequent construction of physical constraints. By comparing the expansions before and after displacement and merging like terms, the following relationship can be obtained for the polynomial coefficients after the shift: The displacement vector can then be obtained by solving the problem. And the velocity vector field is derived. The specific formula is as follows: S5. Detect and classify the acceleration, divergence, and boundary physical properties of the cloud map.

[0036] In this embodiment, based on the optical flow velocity field, various physical dynamic properties of the cloud layer are further calculated, including acceleration, divergence, and cloud boundary variation characteristics. This step aims to characterize the cloud's motion trend and structural changes from a physical perspective, providing a basis for subsequent physical loss calculation. Specifically, it includes the following sub-steps: S51. In fluid dynamics, the conservation of momentum requires that acceleration be controlled by external forces, and there are usually no drastic, unfounded abrupt changes. Therefore, the instantaneous acceleration vector field of cloud maps is calculated based on the continuous optical flow field to characterize the acceleration changes of cloud motion. The calculation formula is as follows: In the formula: For the current frame optical flow, For the optical flow of the previous frame; S52. In fluid mechanics, for incompressible fluids, the mass conservation requirement necessitates a flow field divergence of 0. Therefore, the divergence of the cloud map region is calculated based on the optical flow velocity field to characterize the aggregation or diffusion behavior of cloud clusters. The calculation formula is as follows: In the formula: It is a two-dimensional velocity field, i.e., an optical flow field; S53. In fluid mechanics, the cloud flow field in actual atmospheric motion should maintain a certain continuity at the boundary of the observation area. Therefore, the edge intensity of the optical flow is calculated based on the optical flow velocity field to characterize the local drastic changes in the velocity field at the cloud boundary. The calculation formula is as follows: In the formula: It is a two-dimensional velocity field, i.e., an optical flow field; S54. Based on the statistical distribution of training data, set level thresholds to classify acceleration, divergence, and boundary physical characteristics into levels to characterize the strength of different physical quantities. Different level coefficients are assigned to different levels. The cloud condition level coefficients for acceleration, divergence, and boundary characteristics are respectively... , , In this embodiment, the acceleration levels are divided into low speed, medium speed, high speed, and extremely high speed, with corresponding level coefficients of 0.7, 1.0, 1.4, and 1.8, respectively; the divergence levels are divided into high stability, medium stability, low stability, and unstable, with corresponding level coefficients of 0.5, 1.0, 1.5, and 2.0, respectively; and the boundary feature levels are divided into simple boundary, medium boundary, complex boundary, and extremely complex boundary, with corresponding level coefficients of 0.7, 1.0, 1.3, and 1.6, respectively.

[0037] S6. Add the acceleration loss, divergence loss, and boundary loss to the original loss respectively, and train to obtain the optimal weight coefficients for each.

[0038] In this embodiment, after obtaining the acceleration, divergence, and boundary physical properties of the cloud map, corresponding physical consistency loss functions are constructed respectively. The optimal weight coefficients for each type of physical loss are determined through a phased independent training strategy, providing initial parameters for subsequent joint optimization. Specifically, the following sub-steps are included: S61. Based on the cloud image acceleration vector field, an acceleration physical consistency loss is constructed to suppress drastic abrupt changes in cloud motion, improve the smoothness of predictions and physical interpretability. The acceleration loss function formula is as follows: In the formula: and They represent the optical flow field at... and acceleration components in the direction, It is a very small positive value (such as 10). -6 To prevent numerical instability in gradient calculations, This represents the mathematical expectation of the acceleration magnitude of all pixels; S62. Based on the divergence physical quantity, a divergence physical consistency loss is constructed to constrain the passivity of cloud motion, prevent the prediction of clouds that "appear out of nowhere or dissipate out of thin air," and prevent physically unreasonable flow field divergence in the prediction results. The divergence loss function formula is as follows: In the formula: This represents the mathematical expectation of the absolute divergence values ​​over all pixels. S63. Based on the optical flow edge intensity, a boundary physical consistency loss is constructed to ensure the continuity of the cloud flow field at the boundary and prevent abnormal boundary disturbances. The boundary loss function formula is as follows: In the formula: This indicates variance calculation. This represents the set of velocities of the optical flow field at the image boundary pixels. It is a two-dimensional velocity field, i.e., an optical flow field; S64. By summing the basic prediction loss and the physical loss, a new loss function is constructed. The acceleration loss, divergence loss, and boundary loss are trained independently. By activating only one physical loss while freezing the other physical terms, the optimal weight coefficient corresponding to that physical quantity is obtained. The formula for the new loss function is as follows: In the formula: , , These are the acceleration loss weight, divergence loss weight, and boundary loss weight, respectively.

[0039] S7. Set the optimal weight coefficients as the initial values ​​of the physical loss combination weights, and use Bayesian optimization to find the optimal combination weight coefficients.

[0040] In this embodiment, based on obtaining the independent optimal weights for each physical loss, a Bayesian optimization algorithm is used to globally optimize the weight combination of multiple physical constraints to achieve a synergistic balance among different physical losses, further improving the model's prediction accuracy and physical consistency. Specifically, this includes the following sub-steps: S71. Using the optimal weights for acceleration loss, divergence loss, and boundary loss obtained respectively as the initial points of the Bayesian optimization search space, a continuous search space for the physical loss weights is constructed. Let the weight vector of the joint loss function be: In the formula: , , These are the acceleration loss weight, divergence loss weight, and boundary loss weight, respectively.

[0041] S72. Using the model's predictive performance on the validation set as the objective function of Bayesian optimization, a surrogate model is established for the objective function through Gaussian process regression, and the next set of candidate weights is selected based on the preset acquisition function. S73, For each group of candidate weights The model is trained and validated, and the validation error is fed back to the Bayesian optimization module to update the surrogate model. After multiple rounds of iterative optimization, the physical loss weight vector that optimizes the prediction performance on the validation set is output. This serves as the optimal parameter configuration for the final physical loss function.

[0042] S8. Based on the optimal combination of weight coefficients, introduce level coefficients corresponding to different cloud conditions to adaptively adjust the weights of the three types of physical loss, establish the final loss function that fully extracts the physical features of the cloud map, and perform final model training and power prediction.

[0043] In this embodiment, based on obtaining the optimal weight combination through Bayesian optimization, a cloud condition adaptive mechanism is further introduced. By dynamically adjusting the strength of physical constraints, the model can automatically optimize the loss function configuration for different weather scenarios, ultimately completing the training of a photovoltaic power prediction model with physical consistency. Specifically, this includes the following sub-steps: S81. The optimal physical loss weight coefficient is fused with the cloud condition level coefficient to construct an adaptive weight, as shown in the following formula: In the formula: , , These are the optimal physical loss weighting coefficients for acceleration, divergence, and edge strength, respectively. , , These are cloud condition level coefficients for acceleration, divergence, and edge intensity, respectively. S82. Construct the final physical loss function using adaptive weights, and combine it with the basic prediction loss to construct the final training loss function, which is used to train the complete multimodal fusion model. The joint loss function formula is as follows: In the formula: Based on the prediction of loss; S83. The entire ViT–TCN–cross-attention fusion model is trained based on the final loss function to obtain a photovoltaic power prediction model with physical consistency constraints; and the trained model is used to make the final power prediction for future time periods.

[0044] In summary, this embodiment proposes an ultra-short-term photovoltaic power prediction model that integrates Visual Transformer (ViT), Temporal Convolutional Network (TCN), and cross-attention mechanism. ViT performs block encoding and global spatial feature extraction on cloud images, overcoming the limitations of the local receptive field in traditional CNNs and achieving deep representation of cloud morphology, thickness, and boundary structure. TCN's causal and dilated convolutional structures capture multi-scale temporal dependencies of historical power while avoiding future information leakage. Finally, the cross-attention mechanism dynamically fuses spatiotemporal features, using temporal features as queries and spatial features as keys to generate physically meaningful spatiotemporal context vectors. This method significantly improves the model's ability to collaboratively perceive the spatial evolution of cloud movement and temporal fluctuations in power.

[0045] Furthermore, this embodiment introduces a mechanism for extracting optical flow field physical quantities and embedding multi-physics constraints. Based on a continuous cloud map sequence, the instantaneous motion vector field is calculated, deriving physical quantities such as acceleration, divergence, and boundary strength. Three types of physical consistency loss functions are constructed: acceleration loss constrains the smoothness of cloud motion, divergence loss reinforces the mass conservation properties of incompressible fluids, and boundary loss ensures the continuity of the cloud boundary velocity field. A Bayesian optimization and cloud condition level adaptive weighting strategy is further employed. Initial weights are determined through phased independent training, combined with Gaussian process regression to search for the optimal weight combination, and the constraint strength is dynamically adjusted according to real-time cloud conditions to achieve a precise balance between physical features and data-driven approaches.

[0046] Finally, this embodiment achieves end-to-end training of the ViT-TCN-cross-attention model by constructing a joint objective function that includes a basic prediction loss and an adaptive physical loss. This framework not only leverages the complementarity of multi-source data to improve prediction accuracy but also enhances the model's generalization ability under various weather conditions through explicit embedding of physical constraints. Experiments show that this method maintains stable physical consistency in scenarios such as sudden cloud movement and abrupt changes in irradiance, and effectively improves prediction accuracy and stability under various weather conditions.

[0047] Example 2 Building upon Example 1, this example provides an ultra-short-term photovoltaic power prediction system based on physical constraints and spatiotemporal fusion of multi-source data, used to implement the ultra-short-term photovoltaic power prediction method of Example 1. (See also...) Figure 4 The system includes: (1) Cloud image feature extraction module, used to acquire and preprocess ground panoramic or fisheye cloud image sequences, including dividing the unified image into image blocks according to the preset Patch size; flattening each image block and mapping it to a fixed-dimensional vector space through linear projection; then using Visual Transformer (ViT) to encode the projected Patch sequence in multiple layers, outputting a spatial feature sequence that represents the global and local spatial structure of the cloud image, providing spatial information input for subsequent fusion; (2) Historical power time series feature extraction module, used to read the high time resolution historical power sequence of photovoltaic power station and perform standardization and time alignment processing; input the processed time series into the temporal convolutional network (TCN) composed of causal convolution and dilated convolution, extract multi-scale time series features through residual connection, layer normalization and pooling, and output a global time series vector or time series feature sequence that can characterize short-term disturbances and long-term trends; (3) Multimodal cross-fusion module, which is used to project the spatial features of cloud map and the temporal features of historical power into a unified feature space, and realize the deep interaction between the two modes through cross-attention mechanism; cross-attention uses temporal features as queries and spatial features as keys, and obtains spatial context highly related to time steps through attention weighting, thereby generating fused spatiotemporal features and outputting preliminary future power predictions; (4) Optical flow and physical quantity extraction module, which is used to estimate the dense optical flow velocity field based on continuous frame cloud map, and calculate the physical quantities representing the instantaneous motion of the cloud layer, including velocity vector field, acceleration field, divergence field and boundary intensity index. This module ensures the robustness of optical flow estimation through preprocessing steps such as grayscale, Gaussian smoothing and local polynomial registration, and provides the calculated physical quantities to the physical constraint and cloud condition assessment module. (5) Cloud condition level assessment module, which is used to establish cloud condition level mapping rules based on acceleration, divergence and boundary intensity derived from optical flow, and discretize continuous physical quantities into several level categories, thereby quantifying and classifying the cloud condition in the current frame or short period of time; the obtained level is used to indicate the intensity of cloud movement and its importance to power prediction. (6) Physical consistency loss construction module, which is used to construct physical quantities into several physical consistency loss terms (such as acceleration consistency, divergence consistency, boundary consistency, etc.) and use them in conjunction with the basic prediction loss during training; this module supports independent activation and training of individual loss terms so as to obtain the optimal individual weight of each physical loss on the validation set. (7) Physical loss weight optimization module, which is used to use the single optimal weight as the initial point, and use Bayesian optimization or other surrogate model optimization methods to jointly search the combined weights of each physical loss on the validation set, and output the overall optimal physical loss weight vector, so as to provide reasonable initial values ​​and constraints for the final training. (8) Adaptive weight adjustment module, which is used to dynamically scale and adjust the optimal weight obtained by Bayes optimization according to the cloud condition level information to form cloud condition adaptive physical loss weight; increase the corresponding physical loss weight under severe cloud conditions, and decrease the weight of insensitive physical items under stable cloud conditions, so as to achieve differentiated physical constraints under different weather conditions. (9) Joint training and prediction module, which is used to combine multimodal fusion features, basic prediction loss and adaptive physical loss to form the final training target, and to perform end-to-end training on the complete model of ViT-TCN-cross attention; after training, the final power prediction result is output for the future short time window.

[0048] This embodiment extracts spatial features of cloud images and temporal power features through a deep learning model, and applies physical consistency constraints during training, which can significantly improve the accuracy, robustness, and interpretability of minute-level photovoltaic forecasts, providing more reliable technical support for grid dispatch and renewable energy consumption. Compared with existing technologies, this invention has advantages such as strong perception of sudden weather changes, high model prediction accuracy, and strong physical consistency and interpretability of results.

[0049] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for predicting ultra-short-term photovoltaic power based on physical constraints and spatiotemporal fusion of multi-source data, characterized in that, Using cloud map data and historical photovoltaic power data as inputs to the ultra-short-term photovoltaic power prediction model, the ultra-short-term photovoltaic power prediction results are obtained. The training process of the ultra-short-term photovoltaic power prediction model includes the following steps: Obtain cloud map data, use ViT to extract features, and obtain cloud map features; Historical photovoltaic power data is acquired, and time-series features are extracted using TCN. The predicted photovoltaic power is obtained by fusing the cloud map features and time series features based on the cross-attention mechanism, and the basic prediction loss is calculated. Extract the instantaneous motion vector field of the cloud map data to calculate multiple physical quantities of the cloud map, including acceleration, divergence and optical flow edge intensity, and match the level coefficient corresponding to each physical quantity; Under the premise of activating only the loss corresponding to a single physical quantity, training is performed separately to obtain the optimal weight coefficients for each physical loss; The optimal weight coefficients of each physical loss are used as the initial points of the search space, and the optimal combined weight coefficients are obtained through Bayesian search. The physical constraint loss is calculated based on the grade coefficient and the optimal combination weight coefficient, and the training of the ultra-short-term photovoltaic power prediction model is achieved by combining the basic prediction loss.

2. The ultra-short-term photovoltaic power prediction method based on physical constraints and spatiotemporal fusion of multi-source data as described in claim 1, characterized in that, The process of obtaining cloud map features includes the following steps: The preprocessed cloud image data is divided into multiple image blocks; For each image patch, flattening and linear projection are performed to form an embedding sequence with position encoding; The embedding sequence of each image patch is input into ViT, and the corresponding deep spatial features are extracted through a multi-head attention mechanism; For the deep spatial features of each image patch, average pooling is used to obtain the global spatial feature vector that characterizes the entire cloud map spatial structure, i.e., the cloud map features.

3. The ultra-short-term photovoltaic power prediction method based on physical constraints and spatiotemporal fusion of multi-source data as described in claim 1, characterized in that, The process of obtaining the temporal features includes the following steps: The historical photovoltaic power data is normalized and length-aligned to form a fixed-length time series. The time series is input into a causal convolutional layer, and causal padding is performed on the right side of the convolutional kernel so that the convolutional operation depends only on the current and past power values ​​at any time point. Dilated convolution is performed on causal convolution to form a dilated convolution module. The receptive field of the convolution is expanded by progressively increasing dilation coefficients to capture power change features at different time scales. The dilation coefficients grow exponentially. In each dilated convolutional module, a normalization layer, a non-linear activation function, and a Dropout regularization structure are set sequentially, and the input features and output features are added through residual connections; The time-step features extracted by multiple dilated convolutional modules are aggregated by global average pooling to obtain a global temporal feature vector that characterizes the historical power evolution pattern, i.e., the temporal features.

4. The ultra-short-term photovoltaic power prediction method based on physical constraints and spatiotemporal fusion of multi-source data according to claim 1, characterized in that, The process of obtaining the predicted photovoltaic power includes the following steps: Perform linear projection on the cloud map features and the time series features; The time-series features after linear projection are used as the Query, and the features after linear projection are used as the Key and Value of the cloud map features. Attention weights are calculated through cross-attention. The cloud map features after linear projection are weighted and summed based on the attention weights to obtain the fused features; The fused features are then subjected to residual connections and layer normalization. Based on the processed fusion characteristics, the predicted photovoltaic power is obtained using a fully connected prediction head.

5. The ultra-short-term photovoltaic power prediction method based on physical constraints and spatiotemporal fusion of multi-source data according to claim 1, characterized in that, The process of calculating acceleration, divergence, and optical flow edge intensity includes the following steps: Based on the cloud image data, the input cloud image is converted into a grayscale image and then subjected to Gaussian filtering for smoothing. Based on the preset displacement of adjacent cloud maps, the grayscale values ​​after displacement are expanded. Based on the grayscale value expansion, the instantaneous motion vector field of the cloud map is extracted; Based on the instantaneous motion vector field, the acceleration, divergence, and optical flow edge intensity of the cloud map are calculated respectively.

6. The ultra-short-term photovoltaic power prediction method based on physical constraints and spatiotemporal fusion of multi-source data according to claim 1, characterized in that, The process of obtaining the optimal weight coefficients for each physical loss includes the following steps: A new loss function is constructed by summing the basic prediction loss and the physical loss. The acceleration loss, divergence loss and boundary loss are trained independently. The optimal weight coefficients for each physical quantity are obtained by activating only a single physical loss and freezing the other physical losses.

7. The ultra-short-term photovoltaic power prediction method based on physical constraints and spatiotemporal fusion of multi-source data according to claim 1, characterized in that, The process of obtaining the optimal combination weight coefficients includes the following steps: We construct a continuous search space for physical loss weights by taking the optimal weights for acceleration loss, divergence loss, and boundary loss as the initial points of the Bayesian optimization search space. The predictive performance of the model on the validation set is used as the objective function of Bayesian optimization, and the next set of candidate weights is selected based on the preset acquisition function. For each set of candidate weights, the model is trained and validated. Feedback and updates are performed based on the validation error. After multiple rounds of iterative optimization, the physical loss weight vector that optimizes the prediction performance of the validation set is output as the optimal combination of weight coefficients for the final physical loss function.

8. The ultra-short-term photovoltaic power prediction method based on physical constraints and spatiotemporal fusion of multi-source data according to claim 1, characterized in that, The total loss function of the ultra-short-term photovoltaic power prediction model is: in, , , , , These are the total loss, basic prediction loss, acceleration loss, divergence loss, and boundary loss, respectively. , , These are the optimal weighting coefficients for acceleration loss, divergence loss, and boundary loss, respectively. , , These are the level coefficients for acceleration, divergence, and boundary loss, respectively.

9. The ultra-short-term photovoltaic power prediction method based on physical constraints and spatiotemporal fusion of multi-source data according to claim 8, characterized in that, The acceleration loss, divergence loss, and boundary loss are: in, and These represent the instantaneous motion vector field at... and acceleration components in the direction, It is a very small positive value. This represents the mathematical expectation calculated over all pixels. ,for The velocity component in the direction, ,for The velocity component in the direction, This indicates variance calculation. This represents the set of velocities of the instantaneous motion vector field at the image boundary pixels. It is the instantaneous motion vector field.

10. An ultra-short-term photovoltaic power prediction system based on physical constraints and spatiotemporal fusion of multi-source data, characterized in that, For implementing the ultra-short-term photovoltaic power prediction method as described in any one of claims 1-9, the system comprises: The cloud map feature extraction module is used to acquire cloud map data and extract features using ViT to obtain cloud map features; The historical power time-series feature extraction module is used to acquire historical photovoltaic power data, extract features using TCN, and obtain the time-series features. The multimodal cross-fusion module is used to fuse the cloud map features and temporal features based on the cross-attention mechanism to obtain the predicted photovoltaic power and calculate the basic prediction loss. The optical flow and physical quantity extraction module is used to extract the instantaneous motion vector field of the cloud image data to calculate multiple physical quantities of the cloud image, including acceleration, divergence, and optical flow edge intensity. The cloud condition level assessment module is used to match the level coefficient corresponding to each physical quantity; The physical consistency loss construction module is used to construct the loss function for each physical quantity; The physical loss weight optimization module is used to train the system separately, under the premise of activating only the loss corresponding to a single physical quantity, to obtain the optimal weight coefficients for each physical loss. The adaptive weight adjustment module is used to take the optimal weight coefficients of each physical loss as the initial points of the search space and obtain the optimal combined weight coefficients through Bayesian search. The joint training and prediction module is used to calculate the physical constraint loss based on the rank coefficients and the optimal combined weight coefficients, and to train the ultra-short-term photovoltaic power prediction model by combining the basic prediction loss.