A cross-domain short-term photovoltaic power prediction method and system based on domain generalization theory

Through a cross-domain short-term photovoltaic power prediction method based on domain generalization theory, the maximum mean difference measurement and self-attention framework are used to solve the problem of lack of historical data or data distribution differences between new stations or different stations, and high-accurate cross-domain photovoltaic power prediction is achieved.

CN119312982BActive Publication Date: 2025-06-06INNER MONGOLIA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411553399.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-02
Publication Date
2025-06-06
Estimated Expiration
2044-11-02

AI Technical Summary

Technical Problem

In the case of the lack of historical data or data distribution differences between new stations or different stations, the deep neural network photovoltaic prediction model is difficult to work effectively, resulting in a decrease in prediction accuracy.

Method used

The cross-domain short-term photovoltaic power prediction method based on domain generalization theory is adopted, and the optimal feature combination is generalized by the maximum mean difference measurement method, generalized domain data is generated, and the Informer predictor is built using the self-attention framework to realize the cross-domain prediction task.

Benefits of technology

In the absence of historical data or large data distribution differences, the cross-domain prediction capability of the photovoltaic power prediction model is significantly improved, the prediction accuracy and stability are improved, and it is suitable for cross-domain prediction tasks between new stations and different stations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119312982B_ABST
    Figure CN119312982B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-domain short-term photovoltaic power prediction method and system based on domain generalization theory, which belongs to the field of photovoltaic power generation technology. First, the generalization theory is introduced into the power prediction task of the photovoltaic system, and the maximum mean difference (MMD) is used as the difference measure of environmental sensor data in different fields and combined with adversarial generation to realize the domain generalization of data; then, the mutual information theory is used to quantify the correlation of the multi-dimensional features of the photovoltaic system, and the Informer architecture is introduced to learn the feature knowledge of the source domain and the generalization domain to predict the power of the photovoltaic system in the unknown target domain. The present invention solves the problem that the current photovoltaic prediction algorithm based on deep neural network cannot be applied in a new station without historical data, or the model failure problem between different stations with historical data but significant differences in the distribution of historical data, and has good practical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of photovoltaic power generation technology, and more specifically to a cross-domain short-term photovoltaic power prediction method and system based on domain generalization theory. Background Art

[0002] With the transformation of the global energy structure, the use of clean energy, especially solar energy, has received increasing attention. Photovoltaic power generation, as one of the main ways to utilize solar energy, has been widely used around the world due to its huge energy potential and environmental protection characteristics. However, photovoltaic power generation systems are significantly affected by weather and environmental conditions, and are characterized by randomness and intermittency, which poses a challenge to the stable operation of the power grid. Therefore, accurate power prediction of photovoltaic systems to ensure the stability and reliability of the power grid has become an important topic of current research.

[0003] The research methods of photovoltaic power prediction are mainly divided into physical methods and statistical methods. The physical method solves and simulates the physical process of photovoltaic output through mathematical derivation and simulation tools. Although it has certain advantages in theoretical accuracy, the modeling process is complex and it is difficult to respond quickly to dynamically changing environmental conditions. In contrast, the statistical method uses historical data collected by various sensors and combines statistical models to establish the connection between environmental factors and output power, which has higher flexibility and adaptability. Among them, deep learning models, as an extension of statistical models, have been widely used in power prediction due to their ability to handle high-dimensional, nonlinear and large-scale data.

[0004] Among deep learning models, the Transformer model based on the self-attention mechanism performs well in time series data prediction tasks. By combining the self-attention mechanism with a temporal convolutional network or GRU module, the accuracy and stability of photovoltaic power generation prediction can be further improved. However, deep learning models rely on sufficient historical data for training, which limits their application in actual industrial scenarios, especially in small sample cases. In addition, there are differences in the distribution of data collected by sensors between different photovoltaic systems, which makes it difficult for the trained neural network model to be universal between different systems, especially in newly built sites.

[0005] To overcome the impact of insufficient data on the model, researchers have proposed transfer learning and domain adaptation methods. Transfer learning solves the problem of insufficient data in the target domain by leveraging knowledge from similar tasks or fields. Domain adaptation is applicable to situations where target domain data is unlabeled. By training a decomposed deep domain adaptation model, time series prediction tasks on new devices can be achieved. Although these two methods effectively alleviate the problem of poor performance of deep networks caused by data shortage, they still rely on the premise that there is available data in the target domain. When there is no available data in the target domain, the application effect of these methods will be significantly reduced.

[0006] Therefore, how to provide a photovoltaic power prediction method and system that can ensure prediction accuracy when data is scarce is an urgent problem that technical personnel in this field need to solve. Summary of the invention

[0007] In view of this, the present invention provides a cross-domain short-term photovoltaic power prediction method and system based on domain generalization theory. While screening the multi-dimensional features of the photovoltaic system that are beneficial to the prediction task based on the mutual information theory, the domain generalization idea is used to generalize the data of the optimal feature combination, and an Informer predictor based on the sparse attention mechanism is constructed to learn long sequence features to achieve cross-domain prediction tasks.

[0008] In order to achieve the above object, the present invention adopts the following technical solution:

[0009] A cross-domain short-term photovoltaic power prediction method based on domain generalization theory includes the following steps:

[0010] Acquire raw domain data of multiple environmental sensors, screen features of the raw domain data, and acquire an optimal feature combination;

[0011] Using the maximum mean difference measurement method to generalize the optimal feature combination to generate generalized domain data;

[0012] A predictor is constructed based on a self-attention framework, and the original domain data and the generalized domain data are input into the predictor for cross-domain short-term photovoltaic power prediction.

[0013] Preferably, screening the features of the original domain data to obtain an optimal feature combination includes:

[0014] Preprocessing the original domain data;

[0015] Perform feature extraction on the preprocessed raw domain data;

[0016] Use mutual information to quantify the correlation between the extracted features and obtain the mutual information value;

[0017] Combining the features in order from high to low according to the mutual information values;

[0018] Use different feature combinations to make predictions, and select the feature combination that corresponds to the best prediction result as the optimal feature combination.

[0019] Preferably, the mutual information calculation formula is:

[0020]

[0021] in, Features With features The mutual information value of P(x S ) is x S The probability density function of S ) is y S The probability density function, p(x S ,y S ) is x S With y S The joint probability density function, x S Sample features of the original domain, y S is the original domain label corresponding to the original domain sample feature.

[0022] Preferably, the optimal feature combination is generalized using the maximum mean difference measurement method to generate generalized domain data, including:

[0023] The multi-dimensional feature with length n and dimension m is represented as And use F as the minimum feature unit to perform feature-by-feature generalization tasks;

[0024] Two task objectives are set according to the generalization task to increase the maximum mean difference between the original domain data and the generalized domain data, wherein the task objectives include: a generation objective and a generalization objective;

[0025] Performing adversarial training on the task objective using a generative adversarial network, and determining a loss function of the generative adversarial network according to the task objective, wherein the generative adversarial network is composed of a feature generator and a domain discriminator;

[0026] The original domain data is input into the generative adversarial network for generalization training to generate generalized domain data.

[0027] Preferably, the generalization objective is expressed as:

[0028]

[0029] in, To generalize process data; is the original domain data; is the final f generalization data; d2() is the calculation of the maximum mean difference MMD;

[0030] The generation target is expressed as:

[0031]

[0032] Among them, F' represents the generalized sample; G(·) is the feature generator.

[0033] Preferably, using the generative adversarial network to perform adversarial training on the task objective, and determining the loss function of the generative adversarial network according to the task objective, comprises:

[0034] For the feature generator, a loss function is set according to the generation target To generate the generalized domain data, the loss function of the feature generator is:

[0035]

[0036] Among them, C(·) is the domain discriminator;

[0037] A loss function is set for the domain discriminator C(·) to determine the domain of the generated data. The loss function of the domain discriminator is for:

[0038]

[0039] Setting the loss function for the generalization goal increases the original domain data With the generalized domain data Data difference, generalization loss function

[0040]

[0041] Among them, F is the original domain sample; F′ is the generalized sample;

[0042] Using the same optimizer, minimize the loss function of the feature generator At the same time, maximize the loss function of the domain discriminator With the generalization loss function The overall loss function of the generative adversarial network is:

[0043]

[0044] Among them, θ G is the model parameter of the generator; θ C are the model parameters of the domain discriminator.

[0045] Preferably, a predictor is constructed based on a self-attention framework, comprising:

[0046] Based on the probabilistic sparse self-attention operation, the encoder introduces a distillation operation to construct a feature map. The operation from layer m to layer m+1 is expressed as:

[0047]

[0048] In the formula, MaxPool(·) represents the pooling operation, ELU(·) represents the activation function, and Conv1d(·) represents the convolution operation on the feature dimension. represents sparse self-attention operation;

[0049] The decoder uses a one-shot decoding approach and adds a target placeholder filled with 0s to the input vector, which is:

[0050] X de ={X token ,X 0};

[0051] In the formula, The length is n token A sequence placeholder for is a placeholder for the target sequence;

[0052] The decoder uses a two-layer multi-head self-attention structure for forward propagation of data:

[0053]

[0054] Where FC represents the fully connected layer, LN represents the LayerNorm operation, EMA represents the interaction operation with the encoder, Mask represents the mask operation, and A(·) represents the self-attention operation;

[0055] The back propagation loss is calculated using the mean square error. The expression of the loss function is:

[0056]

[0057] Among them, θ I is the model parameter of Informer, Y i is the true value of the i-th sample sequence, Y i p is the predicted value of the i-th sample sequence.

[0058] A cross-domain short-term photovoltaic power prediction system based on domain generalization theory is used to implement the above-mentioned cross-domain short-term photovoltaic power prediction method based on domain generalization theory, including:

[0059] A special screening module is used to obtain the original domain data of various environmental sensors, screen the features of the original domain data, and obtain the optimal feature combination;

[0060] A data generalization module, used to generalize the optimal feature combination using a maximum mean difference measurement method to generate generalized domain data;

[0061] The prediction module is used to build a predictor based on a self-attention framework, and input the original domain data and the generalized domain data into the predictor to perform cross-domain short-term photovoltaic power prediction.

[0062] Through the above technical solutions, it can be known that compared with the prior art, the present invention discloses a cross-domain short-term photovoltaic power prediction method and system based on domain generalization theory, which effectively solves the problem of model failure caused by the current photovoltaic prediction algorithm based on deep neural network in the case of no historical data in the newly built site or the difference in the distribution of historical data between different sites. Specifically, by introducing the idea of ​​domain generalization, the present invention can generalize the target domain using the historical data of the source domain, realize the photovoltaic power prediction in the case of no historical data or large data distribution difference, and significantly improve the cross-domain prediction ability of the model; the multi-dimensional features of the photovoltaic system that are beneficial to the prediction task are screened by using the mutual information theory, which effectively avoids information redundancy and feature interference, and improves the accuracy and stability of the prediction model. The constructed Informer predictor based on the sparse attention mechanism can learn long sequence features more efficiently, and has better curve fitting advantages and prediction accuracy compared with other time series prediction models. Furthermore, the present invention is not only applicable to the case of no historical data in the newly built site, but also applicable to cross-domain prediction tasks where there are differences in the distribution of historical data between different sites, and has strong practicality and generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0064] Figure 1 It is the overall framework of the cross-domain short-term photovoltaic power prediction method based on domain generalization theory;

[0065] Figure 2 This is a visualization of the characteristics of site 1 and site 2;

[0066] Figure 3 This is the result diagram of the mutual information calculation between the features of site 1 and site 2;

[0067] Figure 4 This is the prediction result diagram of each feature combination of site 1 and site 2;

[0068] Figure 5 This is the generalization result diagram of each feature of Case A and Case B;

[0069] Figure 6The t-SNE dimension reduction visualization distribution results of the source domain data, target domain data, and generalized domain data of Case A and Case B;

[0070] Figure 7 This is the generalization prediction result diagram for Case A and Case B. DETAILED DESCRIPTION

[0071] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0072] The embodiment of the present invention discloses a cross-domain short-term photovoltaic power prediction method and system based on domain generalization theory. Aiming at the problem that the current photovoltaic prediction algorithm based on deep neural network fails to work on new stations without historical data or between different stations with historical data but different historical data distribution, a cross-domain short-term photovoltaic power prediction method based on environmental sensor source domain historical data and combined with domain generalization theory is proposed. The method mainly consists of three parts, and the overall structure is as follows: Figure 1 As shown in the figure. The first part is to analyze and screen the correlation of each feature vector in the environmental sensor data of the photovoltaic system. The second part is to perform generative generalization of the MMD metric for the selected optimal feature combination; the third part is to build an Informer predictor based on the ProbSparse self-attention framework to perform feature learning in the source domain and generalization domain and cross-domain power prediction.

[0073] The embodiment of the present invention discloses a cross-domain short-term photovoltaic power prediction method based on domain generalization theory, such as Figure 1 As shown, the following steps are included:

[0074] Obtain the original domain data of various environmental sensors, filter the features of the original domain data, and obtain the optimal feature combination;

[0075] The maximum mean difference measurement method is used to generalize the optimal feature combination and generate generalized domain data;

[0076] A predictor is constructed based on the self-attention framework, and the original domain data and generalized domain data are input into the predictor for cross-domain short-term photovoltaic power prediction.

[0077] Specifically, the high-dimensional feature vectors in the data collected by various environmental sensors of the photovoltaic system and the corresponding power output labels constitute a high-dimensional feature space, but the contribution of different features to power output varies significantly. In order to select the best features for subsequent tasks, the mutual information based on information theory is used to quantify the correlation between features. The features of the original domain data are screened to obtain the optimal feature combination, including:

[0078] Preprocess the original domain data; first, perform basic preprocessing on the data, normalize all data to the [0,1] interval according to the characteristics, and use the mean of adjacent data points to fill in the missing values. Visualize all the data features of the two stations on a certain day, such as Figure 2 As shown. The correlation between features is quantified using mutual information, and the mutual information between data features for 7 consecutive days is calculated. The results are shown in Figure 3 shown.

[0079] Perform feature extraction on the preprocessed raw domain data;

[0080] Use mutual information to quantify the correlation between features and obtain the mutual information value;

[0081] Combine features in order from high to low according to the mutual information value;

[0082] Use different feature combinations to make predictions, and select the feature combination that corresponds to the best prediction result as the optimal feature combination.

[0083] For the purpose of illustration, two PV arrays with different locations and different levels were selected. At the same time, in order to expand the difference in domain distribution, data from different years were selected for mutual verification.

[0084] The two photovoltaic arrays are Sails in the Desert 3-A and 3-B arrays of the Yulara Solar System on the north side of Ayers Rock, Australia. The specific parameters of the different arrays are listed in Table 1. The 3-A array uses data from 2019. The 3-B array uses data from 2018.

[0085] Table 1 PV array system parameters

[0086]

[0087] Based on the calculation results of mutual information, the correlation between the features was preliminarily analyzed, and then specifically verified. The features of the two sites were combined in order from high to low according to the mutual information value, and the combination methods are listed in Table 2 and Table 3. At the same time, different feature combinations were used for continuous day prediction tasks, and the predictor used Informer. Due to the volatility of the prediction task, the average values ​​of the results of ten rounds of experiments were selected and listed in Table 2 and Table 3. The visualization results are shown in Figure 4 shown.

[0088] Table 2 Prediction results of different feature combinations at site 1

[0089]

[0090] Table 3 Prediction results of different feature combinations at site 2

[0091]

[0092] From the evaluation index trend in the figure, it can be seen that in site 1, features 5 and 6 in the C2 combination caused RMSE / MAE to decrease by 0.1136 and 0.1123, and the improvement in accuracy was the most obvious among all feature combinations. The C3 combination with feature 7 has a significantly lower effect on improving prediction accuracy compared to C1 and C2, but the overall accuracy is still improving. The C4 and C5 feature combinations do not have obvious beneficial effects on prediction tasks, and the evaluation indicators are basically the same as the C3 combination. The C6 and C7 combinations began to reduce the accuracy of the network due to the addition of features 2 and 3 with low mutual information.

[0093] The C3 combination in site 2 achieved the best results in the prediction task, with RMSE / MAE decreasing by 0.1480 and 0.1310 respectively. However, adding feature 4 with a lower mutual information value had a significant negative impact on the prediction task, and the negative impression on the prediction task was the greatest after using the C6 feature combination. Based on the experimental results, feature 5, feature 6, and feature 7 in the C3 combination were selected for the subsequent generalization and prediction stages.

[0094] Furthermore, the mutual information calculation formula is:

[0095]

[0096] in, Features With features The mutual information value of P(x S ) is x S The probability density function of S ) is y S The probability density function, p(x S ,y S ) is x S With y S The joint probability density function, x S Sample features of the original domain, y S is the original domain label corresponding to the original domain sample feature.

[0097] After selecting the optimal features, the multi-dimensional data is generalized. According to the requirements of the generalization task, two task goals are set here, and three loss functions are designed based on the two task goals. Next, the selection of task goals and corresponding loss functions are described in detail.

[0098] In another embodiment, the optimal feature combination is generalized using the maximum mean difference measurement method to generate generalized domain data, including:

[0099] The multi-dimensional feature with length n and dimension m is represented as And use F as the minimum feature unit to perform feature-by-feature generalization tasks;

[0100] According to the generalization task, two task goals are set to increase the maximum mean difference between the original domain data and the generalized domain data. The task goals include: generation goal and generalization goal;

[0101] Use a generative adversarial network to conduct adversarial training on the task objectives, and determine the loss function of the generative adversarial network based on the task objectives. The generative adversarial network consists of a feature generator and a domain discriminator.

[0102] The original domain data is input into the generative adversarial network for generalization training to generate generalized domain data.

[0103] Furthermore, a learnable generator G(·) is used to train the source domain F in the example extracts features through downsampling convolution and reconstructs features through upsampling convolution to obtain the generalization domain. The generalization sample F' of . The generation target is expressed as:

[0104]

[0105] Among them, F' represents the generalized sample; G(·) is the feature generator.

[0106] Furthermore, in order to expand The data difference makes Covering potential so that The trained prediction task can be directly To minimize empirical risk, we will maximize and The data difference between them is used to counter the generation As our generalization goal. and The MMD empirical estimate between is set as a measure of the difference between the two data. The generalization goal is expressed as:

[0107]

[0108] in, To generalize process data; is the original domain data;

[0109] is the final f generalization data; d2() is the calculation of the maximum mean difference MMD;

[0110] Furthermore, the generative adversarial network is used to conduct adversarial training on the task objectives, and the loss function of the generative adversarial network is determined according to the task objectives, including:

[0111] For the feature generator, set the loss function according to the generation target Generate generalized domain data, the loss function of the feature generator is:

[0112]

[0113] Among them, C(·) is the domain discriminator;

[0114] The loss function is set for the domain discriminator C(·) to determine the domain of the generated data. The loss function of the domain discriminator is for:

[0115]

[0116] Setting loss function for generalization goal to increase original domain data With generalized domain data Data difference, generalization loss function

[0117]

[0118] Among them, F is the original domain sample; F′ is the generalized sample;

[0119] In order to make the training process more concise and clear, the same optimizer is used to minimize the loss function of the feature generator. At the same time, maximize the loss function of the domain discriminator And the generalization loss function To this end, and The value of is taken as the opposite number, so that these two loss functions are maximized during the gradient descent process. The overall loss function of the generated adversarial network is finally:

[0120]

[0121] Among them, θ G is the model parameter of the generator; θ C are the model parameters of the domain discriminator.

[0122] The following will explain the generalization steps with specific data. The generalization and prediction stages are divided into two parts: Case A and Case B. Case A uses site 1 as the source domain and site 2 as the unseen target domain; Case B uses site 2 as the source domain and site 1 as the unseen target domain. The target domain is set to have no available data and labels, and only the source domain and the source domain generalization data are used to train the predictor, and the prediction task is performed in the target domain.

[0123] In Case A, we selected 6 consecutive days of data in 2019 to train the CDFM framework, using the 7th day as the validation set, and the 8th day of the consecutive date in Case B as the test set. In Case B, we selected 6 consecutive days in 2018 as the training set, the 7th day as the validation set, and the 8th day of the consecutive date in Case A as the test set.

[0124] Step 1: Generalize the feature data in Case A and Case B in an adversarial generation manner, and select a deep convolutional generative adversarial network for the generalization network. At the same time, in order to cooperate with the data generation process of the generator, the convolutional network in the generator is replaced with a deconvolutional network to perform data upsampling operations. The specific model architecture and training hyperparameters of the constructed generative generalization network are shown in Table 4. The generalized data of features 5, 6, 7 and power labels in Case A and Case B after the last training are visualized as shown in Figure 5 .

[0125] Table 4 DCGGN network structure

[0126]

[0127]

[0128] Depend on Figure 5 The generalization results show that: first, the generator and the discriminator can cooperate to learn the data distribution of various features according to the adversarial learning mechanism, and fit and generate similar data distribution curves. Secondly, due to the effect of MMD, each generalized data has a certain distance difference from the source data. Due to the random input of the generator, the distance gap between the final data is different, which makes each feature data have a certain generalization interval, better covering the unknown target domain.

[0129] Then, 6 generalization samples with the same number as the training samples are uniformly extracted from the generalization data, and t-SNE dimension reduction is performed together with the source domain data and the target domain data, and the visualization is as follows: Figure 6 As shown. Figure 6From the results, we can see that different features have different cluster centers, which means that the data distributions of different features are somewhat different. The cluster centers of feature 6 and feature 7 are close, which means that their distributions are similar.

[0130] Figure 6 The circle in the middle represents the source domain data, and the triangle represents the target domain data. It can be seen that in addition to the distribution differences between different features, the source domain data and the target domain data within each feature also show a certain distribution boundary. This shows that the source domain and the target domain of the same feature have certain data distribution differences. This difference is most obvious in features 6 and 7.

[0131] Figure 6 The square in the middle represents the generalized data generated by the adversarial mechanism. The results show that the generalized data of the square is distributed between the circular source domain data and the triangular target domain data, but mostly tends to the triangular target domain distribution area. This shows that the generalized data has successfully generalized to the target domain while learning the distribution of the source domain data. The generalization effect of features 6 and 7 in Case B is the most obvious. The square generalized data almost covers the triangular target data.

[0132] In order to better perform the prediction task and overcome the quadratic computational complexity problem brought by the long sequence data of the photovoltaic system, the Informer model is used as the predictor in the overall method. The predictor is built based on the self-attention framework, including:

[0133] Based on the probabilistic sparse self-attention operation, the encoder introduces a distillation operation to construct a feature map. The operation from layer m to layer m+1 is expressed as:

[0134]

[0135] In the formula, MaxPool(·) represents the pooling operation, ELU(·) represents the activation function, and Conv1d(·) represents the convolution operation on the feature dimension. represents sparse self-attention operation;

[0136] The decoder uses a one-time decoding method to avoid the cumulative error of dynamic decoding. And adds a target placeholder filled with 0 in the input vector. The input vector is:

[0137] X de ={X token ,X 0};

[0138] In the formula, The length is n token A sequence placeholder for is a placeholder for the target sequence;

[0139] The decoder uses a two-layer multi-head self-attention structure for forward propagation of data:

[0140]

[0141] Where FC represents the fully connected layer, LN represents the LayerNorm operation, EMA represents the interaction operation with the encoder, Mask represents the mask operation, and A(·) represents the self-attention operation;

[0142] The back propagation loss is calculated using the mean square error. The expression of the loss function is:

[0143]

[0144] Among them, θ I is the model parameter of Informer, Y i is the true value of the i-th sample sequence, Y i p is the predicted value of the i-th sample sequence.

[0145] In order to obtain the optimal generalization domain data and prediction results, the present invention designs a framework training mode so that the adversarial generalization task and the cross-domain prediction task can be trained in coordination, and the generalization effect of the model is adjusted by feedback from the cross-domain prediction results, so that both the final generalization task and the prediction task can reach the optimal level.

[0146] Perform generative generalization task training on the training set to obtain the trained θ G ,θ C Then use the generalized data and training data to train the predictor, and use the validation set to tune the hyperparameters and select the best model, and finally get the trained θ I .

[0147] Specifically, first, the source domain training data Taking the generalization task loss function Perform a generalization generation task for the target. Afterwards, the generated generalization domain data With source domain training data Let's combine the predictor f Informer (·) To predict the task loss function Finally, the training hyperparameters are adjusted through the validation data, and finally the trained generator and predictor are obtained for the generalization generation of data and the power prediction of the model.

[0148] The predictor is trained using the generalized data together with the source data and verified on the validation set. The specific network architecture of the Informer predictor constructed in the embodiment of the present invention is shown in Table 5. Due to the volatility of the prediction results, the mean of the ten rounds of results is also selected as the final evaluation index and listed in Table 6. The prediction results of one of them are visualized as shown in Figure 7 shown.

[0149] Table 5 Informer predictor network architecture

[0150]

[0151]

[0152] Table 6 Generalization prediction results

[0153]

[0154] It can be seen that the curve fitting degree of both sites can reach above 0.9. The average RMSE is lower than 0.8 and the MAE is lower than 0.5. This shows that the prediction model has learned the data features in the training data and has produced a good curve fitting effect in the target domain. This shows that the CDFM framework can cover the unknown target domain data by generating generalized source domain data when the target domain is unknown, thereby achieving better task completion results in the target domain.

[0155] A cross-domain short-term photovoltaic power prediction system based on domain generalization theory is used to implement the above-mentioned cross-domain short-term photovoltaic power prediction method based on domain generalization theory, including:

[0156] A special screening module is used to obtain the original domain data of various environmental sensors, screen the features of the original domain data, and obtain the optimal feature combination;

[0157] The data generalization module is used to generalize the optimal feature combination using the maximum mean difference measurement method to generate generalized domain data;

[0158] The prediction module is used to build a predictor based on the self-attention framework and input the original domain data and generalized domain data into the predictor for cross-domain short-term photovoltaic power prediction.

[0159] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0160] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A cross-domain short-term photovoltaic power prediction method based on domain generalization theory, characterized in that: The following steps are involved: Acquire raw domain data of multiple environmental sensors, screen features of the raw domain data, and obtain an optimal feature combination, specifically including: Preprocessing the original domain data; Perform feature extraction on the preprocessed raw domain data; Use mutual information to quantify the correlation between the extracted features and obtain the mutual information value; Combining the features in order from high to low according to the mutual information values; Use different feature combinations to make predictions, and select the feature combination that corresponds to the best prediction result as the optimal feature combination; Using the maximum mean difference measurement method to generalize the optimal feature combination to generate generalized domain data; A predictor is constructed based on a self-attention framework, and the original domain data and the generalized domain data are input into the predictor for cross-domain short-term photovoltaic power prediction; the predictor is constructed based on the self-attention framework, including: Based on the probabilistic sparse self-attention operation, the encoder introduces a distillation operation to construct a feature map. The operation from layer m to layer m+1 is expressed as: In the formula, MaxPool(`) represents the pooling operation, ELU(`) represents the activation function, and Conv1d(·) represents the convolution operation on the feature dimension. represents sparse self-attention operation; The decoder uses a one-shot decoding approach and adds a target placeholder filled with 0s to the input vector, which is: X de ={X token ,X0}; In the formula, The length is n token A sequence placeholder for is a placeholder for the target sequence; The decoder uses a two-layer multi-head self-attention structure for forward propagation of data. The output of the decoder is expressed as: Where FC represents the fully connected layer, LN represents the LayerNorm operation, EMA represents the interaction operation with the encoder, Mask represents the mask operation, and A(·) represents the self-attention operation; The back propagation loss is calculated using the mean square error. The expression of the loss function is: Among them, θ I is the model parameter of Informer, Y i is the true value of the i-th sample sequence, Y i p is the predicted value of the i-th sample sequence.

2. A cross-domain short-term photovoltaic power prediction method based on domain generalization theory according to claim 1, characterized in that: The mutual information calculation formula is: in, Features With features The mutual information value of P(x S ) is x S The probability density function of S ) is y S The probability density function, p(x S ,y S ) is x S With y S The joint probability density function, x S Sample features of the original domain, y S is the original domain label corresponding to the original domain sample feature.

3. The cross-domain short-term photovoltaic power prediction method based on domain generalization theory according to claim 1 is characterized in that: The generalization processing of the optimal feature combination is performed using the maximum mean difference measurement method to generate generalization domain data, including: The multi-dimensional feature with length n and dimension m is represented as And use F as the minimum feature unit to perform feature-by-feature generalization tasks; Two task objectives are set according to the generalization task to increase the maximum mean difference between the original domain data and the generalized domain data, wherein the task objectives include: a generation objective and a generalization objective; Performing adversarial training on the task objective using a generative adversarial network, and determining a loss function of the generative adversarial network according to the task objective, wherein the generative adversarial network is composed of a feature generator and a domain discriminator; The original domain data is input into the generative adversarial network for generalization training to generate generalized domain data.

4. A cross-domain short-term photovoltaic power prediction method based on domain generalization theory according to claim 3, characterized in that: The generalization goal is expressed as: in, To generalize process data; is the original domain data; is the final f generalized data; d 2 () is the calculation of the maximum mean difference MMD; The generation target is expressed as: Among them, F' represents the generalized sample; G(·) is the feature generator.

5. A cross-domain short-term photovoltaic power prediction method based on domain generalization theory according to claim 4, characterized in that: Using the generative adversarial network to perform adversarial training on the task target, and determining the loss function of the generative adversarial network according to the task target, including: For the feature generator, a loss function L is set according to the generation target G To generate the generalized domain data, the loss function of the feature generator is: Among them, C(·) is the domain discriminator; A loss function is set for the domain discriminator C(·) to determine the domain of the generated data. The loss function of the domain discriminator is for: Setting the loss function for the generalization goal increases the original domain data With the generalized domain data Data difference, generalization loss function Among them, F is the original domain sample; F′ is the generalized sample; Using the same optimizer, minimize the loss function of the feature generator At the same time, maximize the loss function of the domain discriminator With the generalization loss function The overall loss function of the generative adversarial network is: Among them, θ G is the model parameter of the generator; θ C are the model parameters of the domain discriminator.

6. A cross-domain short-term photovoltaic power prediction system based on domain generalization theory, used to implement a cross-domain short-term photovoltaic power prediction method based on domain generalization theory as described in any one of claims 1 to 5, characterized in that: include: A special screening module is used to obtain the original domain data of various environmental sensors, screen the features of the original domain data, and obtain the optimal feature combination; A data generalization module, used to generalize the optimal feature combination using a maximum mean difference measurement method to generate generalized domain data; The prediction module is used to build a predictor based on a self-attention framework, and input the original domain data and the generalized domain data into the predictor to perform cross-domain short-term photovoltaic power prediction.