Federal learning-based cross-regional photovoltaic power generation power prediction method

Through personalized federated learning, learnable mask aggregation and nuclear norm constraints, combined with Glassman manifold optimization, data privacy protection, personalized adaptation and computing efficiency problems in cross-regional photovoltaic power prediction are solved, and efficient photovoltaic power prediction is achieved.

CN120372682AActive Publication Date: 2025-07-25NANJING UNIV OF INFORMATION SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510464486.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-25
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

There are problems in cross-regional photovoltaic power generation predictions, and the performance of traditional federated learning algorithms is degraded, the calculation volume is too large, and the optimization is difficult.

Method used

Using a personalized federated learning algorithm, through learning mask aggregation and kernel norm constraints, combined with Glassman manifold optimization, the optimization of model parameters on Glassman manifold is achieved, reducing the calculation amount and optimization difficulty.

Benefits of technology

It realizes cross-regional photovoltaic data privacy protection, personalized model adaptation, improves prediction accuracy, and maintains model performance while reducing the calculation amount, making the optimization process more efficient and stable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372682A_ABST
    Figure CN120372682A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-regional photovoltaic power generation power prediction method based on federated learning, and belongs to the technical field of photovoltaic power generation prediction, the prediction method comprises the following steps: each region participating in photovoltaic power generation power prediction corresponds to a participant, and a server is deployed for cross-regional parameter aggregation; the server distributes an initialized photovoltaic prediction global model to each participant; the participants use the local data to optimize local model parameters; locally carrying out mask aggregation on the photovoltaic prediction models of the participants; the participant optimizes a learnable mask # imgabs0 # in photovoltaic prediction, the participant sends a photovoltaic model parameter # imgabs1 # to the server, and the server calculates a parameter average value of residual models as a new global model after detecting an abnormal model; and repeating until the # imgabs2 is converged or circulated to a specified number of times. According to the method, three core objectives of cross-regional photovoltaic data privacy protection, power generation prediction accuracy improvement and effective calculation amount control can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of photovoltaic power generation prediction, and particularly relates to a cross-regional photovoltaic power generation prediction method. Background Art

[0002] In the field of cross-regional photovoltaic power generation prediction, the protection of data privacy and the construction of personalized prediction models are crucial. On the one hand, when a company deploys photovoltaic devices in multiple countries around the world, different countries, based on their own data privacy protection regulations, usually require the data generated in their own countries to be stored on local servers. At this time, the company cannot directly obtain the original data, which to a certain extent increases the difficulty of data integration and utilization. On the other hand, there are significant differences in lighting conditions in different regions. If a single model is used for power generation prediction, it is difficult to accurately adapt to the different situations in each region. In view of this, the present invention innovatively introduces a personalized federated learning algorithm. This algorithm can effectively integrate the data information of different participants while ensuring the security of the privacy data of each participant, and accurately design personalized models for different participants, so as to achieve the dual goals of cross-regional photovoltaic data privacy protection and accurate power generation prediction. However, if only the existing personalized federated learning algorithm is used, the following problems will occur:

[0003] 1) The aggregation strategy of photovoltaic prediction models in different regions is too simple, resulting in a decline in model performance. Specifically, traditional federated learning aggregation algorithms often simply perform weighted summation on the parameters of the entire model, which is obviously not conducive to the performance of the federated learning model. In response to this, the present invention uses a learnable mask for aggregation, that is, different aggregation coefficients are assigned to each model parameter to achieve more accurate model parameter aggregation.

[0004] 2) While using the learnable mask to improve the model performance, the computational complexity of photovoltaic prediction models in different regions increases sharply. Therefore, the present invention needs to consider the balance between computational complexity and accuracy. To this end, the present invention introduces the nuclear norm to constrain the computational complexity of the learnable mask. By constraining the rank of the mask matrix, the computational complexity can be significantly reduced while ensuring that the model performance is hardly affected.

[0005] 3) When optimizing the model parameters and the mask simultaneously for photovoltaic prediction models in different regions, the optimization difficulty is too high. In response to this, the present invention assumes that the model parameters are located on the Grassmann manifold, that is, complex parameters to be optimized can be represented by simple manifold parameters. At this time, the optimization problem of the model parameters is transformed into an optimization problem of the manifold parameters, and the optimization difficulty is reduced by reducing the number of parameters to be optimized.

[0006] The technical fields related to the present invention are briefly described below to better understand the core content and innovation value of the present invention. The technical fields specifically include federated learning, personalized federated learning, learnable mask aggregation, nuclear norm, and Grassmann manifold.

[0007] Regarding federated learning, it is a method for distributed machine learning at multiple participants, allowing multiple parties to jointly train a deep learning model without sharing data. In the federated learning framework, the data always remains local, and only the model updates are sent to the server for aggregation, thus protecting the data privacy of the participants. This method is particularly suitable for fields with high requirements for data privacy, such as the field of cross-regional photovoltaic power generation prediction.

[0008] Regarding personalized federated learning, it aims to solve the problem in traditional federated learning that it cannot capture the local data distributions of different participants. Personalized federated learning can generate personalized models for each participating party to adapt to its unique data distribution and requirements. This method is particularly suitable for scenarios with uneven data distributions, such as personalized cross-regional photovoltaic power generation prediction.

[0009] Regarding learnable mask aggregation, it is an advanced aggregation method in federated learning. Specifically, by assigning a learnable mask to each model parameter, it can replace the traditional simple weighted summation method. This method can aggregate the model parameters of each participating party more flexibly and accurately, thereby improving the performance of the aggregated model.

[0010] Regarding nuclear norm, it is regarded as a convex relaxation means for matrix rank constraint in matrix optimization. By constraining the nuclear norm of a matrix, the rank of the matrix can be controlled to a certain extent, and at the same time, the originally complex non-convex optimization problem is transformed into a relatively simple convex optimization problem, thereby effectively reducing the complexity of the optimization process. This method takes into account the computational feasibility and efficiency when dealing with matrix-related problems.

[0011] Regarding Grassmann manifold, it is a mathematical structure that describes all subspaces of a specific dimension in a vector space. For example, in a three-dimensional space, the Grassmann manifold can be the set of all lines or the set of all planes. By assuming that the model parameters are on the Grassmann manifold, the optimization variables can be transformed from model parameters to manifold parameters. At this time, the number of parameters to be optimized is reduced, thereby reducing the optimization difficulty. Summary of the Invention

[0012] The object of the present invention is to provide a cross-regional photovoltaic power generation prediction method based on federated learning to solve the three core problems of data privacy protection, personalized adaptation, and balance of computational efficiency and model accuracy in cross-regional photovoltaic prediction, and to break through the optimization problem of high-dimensional parameter space.

[0013] To achieve the above object, the present invention adopts the following technical solutions:

[0014] A cross-regional photovoltaic power prediction method based on federated learning, comprising the following steps:

[0015] Step 1, each region participating in photovoltaic power prediction corresponds to a participant, and a server is deployed for cross-regional parameter aggregation; each participant prepares local data including historical photovoltaic data and meteorological data, where the historical photovoltaic data includes a photovoltaic power time series, and preprocesses the local data; the server distributes an initialized global photovoltaic prediction model to each participant;

[0016] Step 2, after each participant receives the initialized photovoltaic prediction model, use it as the local model, and use the local data to optimize the local model parameters; each participant projects the local model parameters onto the Grassmann manifold. First, perform a QR decomposition on the model parameters of each layer of the local model, and then fix the orthogonal matrix Q, and use the R matrix as the manifold parameter to be optimized by the i-th participant in the t-th round of training For each layer of model parameters, and then fix the orthogonal matrix Q, and use the R matrix as the manifold parameter to be optimized by the i-th participant in the t-th round of training

[0017] Step 3, perform masked aggregation locally on the photovoltaic prediction models of each participant;

[0018] Step 4, each participant optimizes the learnable mask in the photovoltaic prediction

[0019] Step 5, each participant sends the photovoltaic model parameters To the server, and after the server detects the abnormal model, calculate the average value of the parameters of the remaining models as the new global model;

[0020] Step 6, repeat steps 2 to 5 until Convergence or loop to a specified number of times.

[0021] Further, in step 1, the local data prepared by each participant is: the local data of the i-th participant is x i And y i ,x i Is the input data for photovoltaic prediction, including historical photovoltaic data and meteorological data; among them, the historical photovoltaic data includes three hours of historical data, and the data every five minutes corresponds to a time point; the meteorological data includes temperature and solar radiation intensity; y i Is the photovoltaic power generation to be predicted in the next half hour.

[0022] Further, in step 1, the preprocessing steps for the local data are: filling the missing values based on K-nearest neighbors, where the number of neighbors is 1, and after filling the missing values, performing mean-variance normalization on the local data.

[0023] Further, in step 1, the server distributes the same initialized global photovoltaic prediction model to each participant, where t represents the number of rounds, and g indicates that the model is a global model; the global photovoltaic prediction model is a time series prediction model based on a convolutional neural network and a long short-term memory network. Among them, historical photovoltaic data obtains historical photovoltaic features through three one-dimensional convolutional layers, and meteorological data obtains meteorological features through a fully connected network layer; subsequently, the historical photovoltaic features and meteorological features are concatenated and input into the long short-term memory network, and then the time series prediction result of the network is input into a single-layer fully connected neural network to obtain a specific prediction result.

[0024] Further, the number of channels of the one-dimensional convolutional layer is 16, 32, and 64 respectively, the number of filters is 90, 20, and 50 respectively, the pooling size is 2, the convolutional kernel size is 3, the stride is 1, and the activation function selects ReLU; the output dimension of the fully connected network layer is 64; the long short-term memory network contains 100 units, sets the Dropout rate to 0.1, and the activation function is ReLU.

[0025] Further, in step 2, QR decomposition is performed on all optimizable parameters in the network, including convolutional layers and fully connected layers. When the convolutional layer is a high-dimensional tensor structure, the img2col algorithm should be used to convert it into a matrix.

[0026] Further, in step 2, for the optimization objective, it is specifically shown in formula (1):

[0027]

[0028] Satisfy the constraint conditions:

[0029] Among them, the local data of the i-th participant is x i and y i , x i is the input data for photovoltaic prediction, and y i corresponds to the photovoltaic power to be predicted; the initialized global model received by the participant is The loss of the local data on this model is In addition, the fixed parameter in the constraint condition of the Grassmann manifold is Q i ;

[0030] In formula (1) the update method is shown in formula (2)

[0031]

[0032] Among them, α represents the learning rate when optimizing the manifold parameters. In some embodiments of the present invention, α is taken as 0.01; represents the gradient of the loss with respect to In some embodiments of the present invention, is calculated through the automatic differentiation mechanism of the pytorch framework.

[0033] Furthermore, in step 3, each participant first left-multiplies the Grassmann manifold parameter by Q i to convert it back to the conventional model parameter Subsequently, the conventional model parameter and the global model parameter are subjected to masked aggregation, as specifically shown in formula (3);

[0034]

[0035] Among them, represents the learnable mask of the i-th participant during the t-th round of training, 1 represents the all-1 tensor of the same dimension as the mask, and ⊙ represents element-wise multiplication. In some embodiments of the present invention, the learnable mask is initialized as a tensor of all 1 / 2 and is gradually optimized during the training process.

[0036] Furthermore, in step 4, the nuclear norm is used to constrain the rank, and the specific optimization objective is as shown in formula (4)

[0037]

[0038] Among them, represents the nuclear norm of, β represents the correction coefficient of the nuclear norm importance, and is taken as 0.01 in the optimization of cross-region photovoltaic power prediction, represents the loss of the model after learnable mask aggregation on the local data;

[0039] For the update of in formula (4), the specific method is as shown in formula (5):

[0040]

[0041] Among them, γ represents the learning rate. In some embodiments of the present invention, γ is taken as 0.01, represents L mask with respect to gradient.

[0042] Further, in step 5, the server first calculates the mean and standard deviation of all participant models; when more than 20% of the model weights fall outside the three-standard-deviation range, the model is regarded as an abnormal model and discarded during aggregation; for the remaining normal models, the server takes the average of the parameters of these models as the new global model, as specifically shown in formula (6): where I is the number of participants;

[0043]

[0044] where I is the number of participants.

[0045] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects.

[0046] The present invention can protect the data privacy of participants in different regions. Through the federated learning algorithm, it can be ensured that when performing cross-regional power generation prediction, the data does not leave the local area, but model-level information fusion is carried out through the method of masked aggregation. This method can aggregate information from different regions and improve the model prediction performance on the premise of meeting the data privacy regulations of various countries.

[0047] The present invention can accurately adapt to personalized needs. Since there are large differences in variables such as light conditions in different regions, it is necessary to construct an exclusive model for each participant through the personalized federated learning algorithm. In this way, the present invention can effectively capture and utilize personalized information, overcoming the limitation that the global model in the traditional federated learning algorithm is difficult to adapt to personalized needs.

[0048] The present invention can improve the performance of the aggregation model through learnable masks. Specifically, by assigning different aggregation coefficients to each parameter of the photovoltaic prediction model in different regions, more accurate aggregation of model parameters can be achieved. Compared with the simple weighted summation method in the traditional federated learning aggregation algorithm, the learnable mask method of the present invention can better adapt to the characteristics of different parameters, thereby improving the overall performance of the model.

[0049] The present invention realizes the balance between computational efficiency and model accuracy. Specifically, for the learnable masks held by the photovoltaic prediction models in different regions, the present invention introduces the nuclear norm to constrain their computational amount. By constraining the rank of the mask, the present invention can significantly reduce the computational amount while ensuring that the model performance is hardly affected. This method realizes the effective balance between computational efficiency and model accuracy, thereby improving the practicability and feasibility of the method.

[0050] The present invention reduces the optimization difficulty and improves the optimization efficiency. Specifically, the parameters of the photovoltaic prediction models in different regions are located on the Grassmann manifold, and the optimization problem of the model parameters is transformed into a manifold optimization problem. By using simple manifold parameters to represent complex parameters to be optimized, the present invention can significantly reduce the optimization difficulty. This method can make the optimization process more efficient and stable, and improve the optimization effect and convergence speed of the model. Description of the Drawings

[0051] Figure 1 is a flowchart of the method for predicting the photovoltaic power generation across regions based on federated learning according to the present invention;

[0052] Figure 2 is a diagram of the model architecture of the present invention. Detailed Embodiment

[0053] The following further explains the present invention with reference to the drawings.

[0054] Taking the construction of the cross-region photovoltaic power generation prediction as an example, the present invention will analyze the photovoltaic power station dataset of Bridgestone in Changzhou as an example. This dataset contains data from January to July 2018, with each data point being every 5 minutes. The specific features include power generation, temperature, and light intensity. Through the method of the present invention, a mean square error of 0.1129 can be obtained in time series prediction. It should be noted that the present invention focuses on solving the problems of cross-region photovoltaic data privacy protection, improvement of power generation prediction accuracy, and effective control of the amount of calculation. In terms of data privacy protection, the present invention adopts federated learning technology to ensure that the data is stored locally and does not leak, thus effectively maintaining the data security of different regions. Regarding the accuracy of power generation prediction, the present invention introduces learnable mask aggregation to the photovoltaic prediction models in different regions, thereby improving the accuracy of the prediction results. In terms of calculation amount control, the present invention innovatively uses nuclear norm to constrain the rank of the learnable mask, and significantly reduces the calculation amount and improves the calculation efficiency with almost no loss of model performance. In addition, the present invention focuses on solving the problem of excessive optimization difficulty when simultaneously optimizing the mask and model parameters. To this end, the present invention proposes that the parameters of the photovoltaic prediction models in different regions are located on the Grassmann manifold. Based on this, the present invention reduces the number of parameters to be optimized, effectively reduces the complexity and difficulty of the optimization process, and thus realizes the efficient optimization of the model parameters. The following further details the technical solution of the present invention with reference to the drawings and embodiments.

[0055] Step 1, in the initialization stage, for each region participating in the photovoltaic power generation prediction, there is a corresponding participant. And, a server needs to be deployed for cross-region parameter aggregation. Specifically, x iIs the input data for photovoltaic forecasting, specifically including historical photovoltaic data of the photovoltaic power time series and meteorological data. Among them, the historical data of the photovoltaic power time series includes three hours of historical data, and each half-hour data corresponds to a time point; while the meteorological data includes temperature and solar radiation intensity. Before the data is input into the model, it is necessary to fill in the missing values based on K-nearest neighbors, where the number of neighbors can be taken as 1. After filling in the missing values, the local data is normalized by mean-variance. y i Is the photovoltaic power generation for the next half-hour to be predicted. Subsequently, the server will initialize the same global model And distribute it to each participant, where t represents the number of rounds, and g indicates that this model is a global model. Specifically, the photovoltaic forecasting global model Is a time series prediction model based on a convolutional neural network and a long short-term memory network. The specific network structure is as shown in the appendix Figure 2 As shown. For the convolutional neural network, it includes 3 one-dimensional convolutional layers. The number of channels in the one-dimensional convolutional layers are 16, 32, and 64 respectively. The convolutional kernel size is 3, the stride is 1, and the number of filters are 90, 20, and 50 respectively. And the pooling size is 2. And ReLU is used as the activation function. For the long short-term memory network, it includes 100 units, and the Dropout rate is set to 0.1. And the activation function is ReLU. Among them, the historical photovoltaic data obtains historical photovoltaic features through the convolutional neural network, and the meteorological data obtains meteorological features through a fully connected network layer. The output dimension of the fully connected network layer is 64. Subsequently, the historical photovoltaic features and meteorological features are concatenated and input into the long short-term memory network. Then, the time series prediction result of this network is input into a single-layer fully connected neural network to obtain the specific prediction result. In addition, the batch size of this model is set to 750.

[0056] Step 2, after each participant receives the initialized photovoltaic forecasting model, they use the local data to train the model and optimize the local model parameters. For the local photovoltaic forecasting model, its structure is exactly the same as that of the photovoltaic forecasting global model. It should be noted that the model parameters should be on the Grassmann manifold. The reason for choosing this manifold is that there are a large number of redundant variables in the deep neural network, and the actual effective parameters often lie in a subspace of the original parameter space. And the Grassmann manifold can describe the effective parameters in the subspace, thus reducing the difficulty of model optimization. Therefore, different from traditional deep learning that needs to optimize all parameters, in the present invention, only the parameters on the manifold need to be optimized. For this, it is necessary to first perform operations on the model parameters Perform QR decomposition to obtain the orthogonal matrix Q and the upper triangular matrix R. It should be noted that in the present invention, QR decomposition needs to be performed on all optimizable parameters in the network, including convolutional layers and fully connected layers. When the convolutional layer is a high-dimensional tensor structure, the img2col algorithm should be used to convert it into a matrix. Subsequently, the orthogonal matrix Q is fixed, and the R matrix is used as the manifold parameter of the i-th participant in the t-th round of training. The specific optimization objective is shown in Equation (1):

[0057]

[0058] Subject to the constraint:

[0059] where the local data of the i-th participant is x i and y i , and the global model received by the participant is while the loss of the local data on the model is This loss can be the mean square error in photovoltaic prediction. In addition, the fixed parameter in the constraint condition of the Grassmann manifold is Q i ;

[0060] For the update of the photovoltaic prediction model , the specific method is shown in Equation (2):

[0061]

[0062] where α represents the learning rate when optimizing the parameters, and it can be taken as 0.01 in the cross-regional photovoltaic power prediction system. represents the gradient of the mean square error L loss with respect to , which is calculated through the automatic differentiation mechanism of the pytorch framework.

[0063] Step 3, the photovoltaic prediction models of each participant are masked and aggregated locally. Specifically, each participant first multiplies the Grassmann manifold parameter on the left by Q i to convert it back to the conventional parameter Subsequently, the conventional parameter is masked and aggregated with the global model , as shown in Equation (3):

[0064]

[0065] where represents the learnable mask of the i-th participant in the t-th round of training, 1 represents the all-1 matrix with the same dimension as the mask, and ⊙ represents element-wise multiplication. It should be noted that this learnable mask is initialized as a tensor of all 1 / 2 and is gradually optimized during the training process. In the cross-regional photovoltaic power prediction system, The initialization can use a matrix with the same dimension as the global model and all elements being .

[0066] Step 4, each participant optimizes the learnable mask in the photovoltaic prediction At this time, the nuclear norm is needed to constrain the rank, and the specific optimization objective is shown in Formula (4):

[0067]

[0068] Among them, represents the nuclear norm of , and β represents the correction coefficient for the importance of the nuclear norm, which is taken as 0.01 in the optimization of cross-regional photovoltaic power prediction;

[0069] For the update of the learnable mask in the photovoltaic prediction, the specific method is shown in Formula (5):

[0070]

[0071] Among them, γ represents the learning rate, which can be taken as 0.01 in the optimization of cross-regional photovoltaic power prediction, and represents the gradient of L mask with respect to .

[0072] Step 5, each participant sends the photovoltaic prediction model parameters to the server. After the server receives the parameters of all participants, it takes the average of these parameters as the new global model, as shown in Formula (6) specifically:

[0073]

[0074] Among them, I represents the number of participants. In the cross-regional photovoltaic power prediction system, I is often between several hundred and several thousand, and the specific value is determined by the number of regions.

[0075] Step 6, repeat Step 2 to Step 5 until the global model parameters converge or loop to the specified number of times. In the cross-regional photovoltaic power prediction system, the specified number of times can be selected as 500.

[0076] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. Cross - regional photovoltaic power prediction method based on federated learning, characterized in that: It includes the following steps: Step 1: Each region participating in the photovoltaic power prediction corresponds to a participant, and a server is deployed for cross-region parameter aggregation; each participant prepares local data including historical photovoltaic data and meteorological data, where the historical photovoltaic data includes a photovoltaic power time series, and preprocesses the local data; the server distributes an initialized global photovoltaic prediction model to each participant. Step 2: After each participant receives the initialized photovoltaic prediction model, use it as the local model and optimize the local model parameters with local data; each participant projects the local model parameters onto the Grassmann manifold. First, perform a QR decomposition on the model parameters of each layer of the local model and then fix the orthogonal matrix Q and use the R matrix as the manifold parameter to be optimized by the i-th participant in the t-th round of training Step 3: The photovoltaic prediction models of each participant are masked and aggregated locally. Step 4, each participant optimizes the learnable mask in the photovoltaic prediction Step 5, each participant sends the photovoltaic model parameters to the server. After the server detects the abnormal models, it calculates the average value of the parameters of the remaining models as the new global model; Step 6, repeat Steps 2 to 5 until convergence or loop to a specified number of times.

2. The method according to claim 1, characterized in that: In step 1, the local data prepared by each participant is as follows: the local data of the i-th participant is x i and y i , x i is the input data for photovoltaic prediction, including historical photovoltaic data and meteorological data; among them, the historical photovoltaic data includes three hours of historical data, and the data every five minutes corresponds to a time point; the meteorological data includes temperature and solar radiation intensity; y i is the photovoltaic power generation for the next half hour to be predicted.

3. The method according to claim 1, characterized in that: In Step 1, the preprocessing steps for the local data are as follows: filling the missing values based on K-nearest neighbors, where the number of neighbors is 1, and after filling the missing values, performing mean-variance normalization on the local data.

4. The method according to claim 1, wherein: In step 1, the server distributes the same initialized global photovoltaic prediction model to each participant, where t represents the number of rounds and g indicates that this model is a global model; the global photovoltaic prediction model is a time series prediction model based on a convolutional neural network and a long short-term memory network. Among them, historical photovoltaic data obtains historical photovoltaic features through three one-dimensional convolutional layers, and meteorological data obtains meteorological features through a fully connected network layer. Subsequently, the historical photovoltaic features and meteorological features are concatenated and input into the long short-term memory network, and then the time series prediction result of this network is input into a single-layer fully connected neural network to obtain a specific prediction result.

5. The method according to claim 4, characterized in that: The number of channels of the one-dimensional convolutional layer is 16, 32, and 64 respectively, the number of filters is 90, 20, and 50 respectively, the pooling size is 2, the convolutional kernel size is 3, the stride is 1, and the activation function is selected as ReLU; the output dimension of the fully connected network layer is 64; the long short-term memory network contains 100 units, the Dropout rate is set to 0.1, and the activation function is ReLU.

6. The method according to claim 1, characterized in that: In Step 2, QR decomposition is performed on all optimizable parameters in the network, including the convolutional layer and the fully connected layer. When the convolutional layer is a high-dimensional tensor structure, the img2col algorithm should be used to convert it into a matrix.

7. The method according to claim 1, characterized in that: In Step 2, for the optimization objective is specifically shown in Equation (1): Meet the constraint conditions: Among them, the local data of the i-th participant is x i and y i , x i is the input data for photovoltaic prediction, and y i corresponds to the photovoltaic power generation to be predicted; the initialized global model received by the participant is The loss of the local data on this model is In addition, the fixed parameter in the constraint condition of the Grassmann manifold is Q i ; In Formula (1) is updated as shown in Formula (2): Among them, α represents the learning rate when optimizing the manifold parameters; denotes the gradient of the loss with respect to the gradient.

8. The method according to claim 1, characterized in that: In step 3, each participant first left-multiplies the Grassmann manifold parameter by Q i to convert it back to the conventional model parameter Subsequently, the conventional model parameter and the global model parameter are subjected to masked aggregation, as specifically shown in formula (3): wherein, represents the learnable mask of the i-th participant during the t-th round of training, 1 represents a tensor of all 1s with the same dimension as the mask, and ⊙ represents element-wise multiplication; the learnable mask is initialized as a tensor of all 1 / 2s and is gradually optimized during the training process.

9. The method according to claim 1, characterized in that: In step 4, the nuclear norm is used to constrain the rank, and the specific optimization objective is shown in formula (4): Among them, denotes the nuclear norm, β represents the correction coefficient of the importance of the nuclear norm, and takes 0.01 in the optimization of cross-region photovoltaic power prediction. denotes the loss of the model after learnable mask aggregation on local data; For the update of in Formula (4), the specific method is shown in Formula (5): where γ represents the learning rate, denotes the gradient of L mask with respect to .

10. The method according to claim 1, characterized in that: In Step 5, the server first calculates the mean and standard deviation of all participants' models; when more than 20% of the model weights fall outside the three-standard-deviation range, the model is regarded as an abnormal model and discarded during aggregation; for the remaining normal models, the server takes the average of the parameters of these models as the new global model, as specifically shown in formula (6): where I is the number of participants. Where I is the number of participants.

Citation Information

Patent Citations

  • High-efficiency personalized federal learning method for distribution offset robustness

    CN117787436A

  • Single-client multi-domain heterogeneous federated learning system and method based on manifold learning

    CN118734995A

  • Malware detection by distributed telemetry data analysis

    US20220114260A1