Smart grid load forecasting method based on federated learning
Patent Information
- Application Number
- CN202310661574.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-05
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-06-05
AI Technical Summary
因此现存的智能电网的电力负荷和价格精准预测存在不足之处,目前的电力的负荷预测方法单纯的在电力生产者端使用LSTM进行电厂的预测,数据上传存在一定的延迟且数据量较大,电力的负荷预测的处理效率低下
[0026]上述基于联邦学习的智能电网负荷预测方法,通过将电力数据集进行分类,分成训练集、测试集和验证集,以各个电力消耗单位为客户端,根据电线杆的编号选取100个片区作为参与联合训练的客户端,各客户端采用所述训练集构建自身的基于LSTM网络的电网负荷预测模型,获得各基于LSTM网络的电网负荷预测模型的初步网络参数,并将所述初步网络参数上传至服务器端进行联合训练,所述服务器端根据各客户端上传的初步网络参数,基于MMD的模型迁移方法对基于LSTM网络的电网负荷预测全局模型进行联合训练,再使用测试集和验证集对所述基于LSTM网络的电网负荷预测全局模型进行测试验证,获得最终网络参数,并将所述最终网络参数返回给各所述客户端,各所述客户端将自身的基于LSTM网络的电网负荷预测模型的网络参数更新为所述最终网络参数,获得各所述客户端的基于LSTM网络的电网负荷预测全局模型,各客户端获取对应片区的电力数据输入自身的基于LSTM网络的电网负荷预测全局模型进行电力负荷预测,以确定平均绝对误差,由此,处理的数据量会大大减少,可以解决数据上传的延迟且数据量较大的问题,提高了电力的负荷预测的处理效率。
Smart Images

Figure CN116706888B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart grid technology, and in particular to a smart grid load forecasting method based on federated learning. Background Technology
[0002] A smart grid is a modern power grid that efficiently manages power generation, distribution, and consumption. It involves several crucial decision-making processes based on load forecasting, such as generation planning, demand and supply management, maintenance planning, and reliability analysis. Currently, it employs different peak and off-peak pricing mechanisms. Based on these existing pricing mechanisms, reliable and accurate price estimates can enable power producers to maximize profits and power consumers to minimize costs. Since most electricity cannot be stored, a near-perfect balance needs to be maintained between power producers and consumers. Current smart grid load forecasting is performed at the power producer end, using Long Short-Term Memory (LSTM) networks to predict the overall grid load for a given region. Data collection from smart monitoring devices is distributed across various consumer terminals, resulting in data upload delays. Therefore, existing smart grid methods have shortcomings in accurately predicting power load and prices. Current load forecasting methods rely solely on LSTMs at the power producer end for power plant forecasting, leading to data upload delays, large data volumes, and low processing efficiency. Summary of the Invention
[0003] Therefore, it is necessary to provide a federated learning-based smart grid load forecasting method that can improve the processing efficiency of power load forecasting, addressing the aforementioned technical problems.
[0004] A smart grid load forecasting method based on federated learning, the method comprising:
[0005] Step S1: Classify the power dataset into training set, test set and validation set.
[0006] Step S2: Using each electricity consumption unit as a client, select 100 areas as clients U participating in the joint training based on the pole numbers. i (i = 1, 2, ..., 100);
[0007] Step S3: Each client uses the training set to construct its own LSTM-based power grid load prediction model, obtains the preliminary network parameters of each LSTM-based power grid load prediction model, and uploads the preliminary network parameters to the server for joint training.
[0008] Step S4: The server performs joint training on the global model for power grid load prediction based on LSTM network using the model transfer method based on MMD according to the preliminary network parameters uploaded by each client. Then, the server uses the test set and validation set to test and validate the global model for power grid load prediction based on LSTM network to obtain the final network parameters and return the final network parameters to each client.
[0009] Step S5: Each client updates the network parameters of its LSTM-based power grid load prediction model to the final network parameters, thereby obtaining the global LSTM-based power grid load prediction model for each client.
[0010] Step S6: Each client obtains the power data for its corresponding area and inputs it into its own LSTM-based global power grid load forecasting model to predict the power load and determine the mean absolute error. The mean absolute error is used to measure the average error between the power consumption of a single area obtained from the decomposition at a certain moment and the actual power input to that area. The expression for calculating the mean absolute error is as follows:
[0011]
[0012] Where MAE is the mean absolute error, g t To represent the actual power consumption of this region at time t, p t Let T represent the total amount of electricity generated at time t, where T represents the number of time points.
[0013] In one embodiment, the LSTM-based power grid load prediction model and the LSTM-based global power grid load prediction model have the same network structure, which includes a one-dimensional full convolution, an LSTM network, an external attention module, and a support vector machine.
[0014] After the power data is preprocessed by one-dimensional full convolution, it is input into the LSTM network to mine the correlation features between parameters and time. After the feature information is output, it is input into the external attention module for key information extraction. The extracted key information is then input into the output prediction result.
[0015] In one embodiment, the loss function of the LSTM-based power grid load forecasting model is:
[0016]
[0017] in, For the kth dataset The logarithmic loss function, To represent the k-th power dataset D kIn the i-th sample, x represents the true label value, y represents the predicted value, and n... k Let be the number of samples in the k-th dataset. for, ω is the predicted value under the current model parameters. t Let be the weight at time t. The symbol for summation is 'log()', and the symbol for log is 'logarithm'.
[0018] In one embodiment, the update formula for the network parameters when each client trains its own LSTM-based power grid load forecasting model using the training set is as follows:
[0019]
[0020]
[0021] Where, m t Let v be the first moment estimate of the gradient at time t, i.e., the mean of the gradient; let v be the second moment estimate of the gradient at time t, i.e., the biased variance of the gradient; and let g be the mean of the gradient. t Let be the gradient obtained at time t, where t represents the current learning iteration number. ⊙ is a type of multiplication involving element-wise multiplication. γ1 and γ2 are a set of hyperparameters of the LSTM-based power grid load forecasting model, where γ1, γ2 ∈ [0, 1), and γ1 = 0.9, γ2 = 0.99. and These are the mean and biased variance of the corrected gradient, where η is the learning rate. Here are the updated network parameters, θ represents the current network parameters, and m represents the current network parameters. t-1 For the first moment estimate of the gradient at time t-1, v t-1 The second moment estimate of the gradient at time t-1. Let be the first hyperparameter at time t. Let be the second hyperparameter at time t, and ∈ be a hyperparameter.
[0022] In one embodiment, the MMD-based model transfer method performs joint training on the LSTM-based global model for power grid load forecasting in the following manner:
[0023] The gradients of the initial network parameters are corrected by aligning the distribution of multiple domain MMDs based on the gradient. The correction formula is as follows:
[0024]
[0025] Where λ is the hyperparameter of gradient descent, grad i and grad jThese represent the gradients sent from client i and client j to the server, respectively. This is the corrected gradient.
[0026] The aforementioned smart grid load forecasting method based on federated learning classifies the power dataset into training, testing, and validation sets. Each power-consuming unit acts as a client, and 100 areas are selected as clients participating in joint training based on pole numbers. Each client uses the training set to construct its own LSTM-based grid load forecasting model, obtaining preliminary network parameters for each LSTM-based model. These preliminary network parameters are then uploaded to the server for joint training. The server uses the model transfer modeling method (MMD) based on the preliminary network parameters uploaded by each client to jointly train the global LSTM-based grid load forecasting model. Finally, the test set and validation set are used for further training. The validation set tests and validates the LSTM-based global power grid load forecasting model to obtain final network parameters, which are then returned to each client. Each client updates the network parameters of its own LSTM-based global power grid load forecasting model to these final network parameters, thus obtaining its own LSTM-based global power grid load forecasting model. Each client obtains the power data for its corresponding region and inputs it into its own LSTM-based global power grid load forecasting model to perform power load forecasting and determine the mean absolute error. This significantly reduces the amount of data processed, solves the problems of data upload delays and large data volumes, and improves the processing efficiency of power load forecasting. Attached Figure Description
[0027] Figure 1 This is a flowchart illustrating a smart grid load forecasting method based on federated learning in one embodiment.
[0028] Figure 2 This is a schematic diagram of a federated learning framework based on a temporal network structure in one embodiment;
[0029] Figure 3 This is a schematic diagram of a one-dimensional full convolution of power data in one embodiment;
[0030] Figure 4 This is a schematic diagram of the update gate structure of an LSTM network in one embodiment;
[0031] Figure 5 This is a schematic diagram of the forget gate structure of an LSTM network in one embodiment;
[0032] Figure 6 This is a schematic diagram of the input gate structure of an LSTM network in one embodiment;
[0033] Figure 7 This is a schematic diagram of the LSTM network performing temporal dimension feature extraction in one embodiment.
[0034] Figure 8 This is a schematic diagram of the external attention mechanism in one embodiment;
[0035] Figure 9 This is a schematic diagram of the structure of a power grid load forecasting model based on an LSTM network in one embodiment;
[0036] Figure 10 This is a schematic diagram illustrating the model transfer of a smart grid load forecasting method based on federated learning in one embodiment. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0038] In one embodiment, such as Figure 1 and Figure 2 As shown, a smart grid load forecasting method based on federated learning is provided. Taking the application of this method to a terminal as an example, the method includes the following steps:
[0039] Step S1: Classify the power dataset into training set, test set and validation set, and preprocess the samples in the training set to obtain the processed training set.
[0040] Based on publicly available power datasets, power data is collected daily by regional nodes, with data collected every 5 minutes. Each regional node receives 288 power data entries daily, which are then saved as text to create the power dataset. The regional division uses existing substations with existing power lines as regional nodes. Data acquisition devices are used on these existing substations to collect the necessary data, enabling convenient and efficient data acquisition.
[0041] The power dataset is categorized into three parts: training set Q1, test set Q2, and validation set Q3, to facilitate better training and more accurate evaluation. Training set Q1 serves as the training sequence, test set Q2 as the test sequence, and validation set Q3 as the validation set. Test set Q2 and validation set Q3 are hosted on the server and used to validate the accuracy of the global model.
[0042] In this process, each area substation is sampled independently and continuously. During the training process on the local client, the value at the first time point is used as the sample label, and the subsequent N time points are used as samples, thus constructing a labeled sample space.
[0043] Step S2: Using each electricity consumption unit as a client, select 100 areas as clients U participating in the joint training based on the pole numbers. i (i = 1, 2, ... 100).
[0044] Step S3: Each client uses the training set to build its own LSTM-based power grid load prediction model, obtains the preliminary network parameters of each LSTM-based power grid load prediction model, and uploads the preliminary network parameters to the server for joint training.
[0045] In one embodiment, the network structure of the LSTM-based power grid load prediction model includes a one-dimensional full convolution, an LSTM network, an external attention module, and a support vector machine. After the power data is preprocessed by the one-dimensional full convolution, it is input into the LSTM network to mine the correlation features between parameters and time. After the output feature information is input into the external attention module for key information extraction, the extracted key information is input into the support vector machine to output the prediction result.
[0046] The training set is divided into a new dataset D = {D1, D2, ..., D}. i}, where D i Given the power dataset for the i-th region, clients in each region use this power dataset to build a power grid load prediction model based on an LSTM network.
[0047] Step S3 includes:
[0048] Step 3-1: Input the power dataset into the LSTM-based power grid load prediction model. The one-dimensional full convolution of the LSTM-based power grid load prediction model first preprocesses the samples in the power dataset to obtain the processed training set. Specifically: A convolution process is performed on the samples in the power dataset to remove certain information for some time periods. Then, the sliding window of the sample sequence is increased by overlapping sliding. Assuming the sequence length is M, a window of length U is cut on the original data. With a sliding step size of 1, the sliding operation is performed to obtain M-U+1 training samples. The training samples are then subjected to max-min normalization to obtain the normalized training samples. The max-min normalization method is expressed as follows:
[0049]
[0050] Among them, X* Let X be the normalized training sample, and X be the normalized training sample. min X is the minimum value among all training samples. max It is the maximum value among all training samples.
[0051] Because electricity data is incomplete, neural networks require a large number of training samples for fine-tuning to achieve good performance. Furthermore, before being input into the neural network, the data needs to be transformed into a consistent input dimension, with each data point segmented into a vector of constant length—a process called windowing. Before extracting time-dimensional features using LSTM, a convolution process is performed, essentially smoothing the electricity data and removing incomplete information from certain time periods. Because electricity data has a relatively simple dimension, a one-dimensional convolution operation is used here, such as... Figure 3 As shown, a one-dimensional full convolution is used, with a kernel size K of 1*3 and a stride of 1. The input power data X is 1*n, where the input dimension "1" represents the time dimension and "n" represents the data information. Typically, the values of the one-dimensional full convolution kernel are set to η1, η2, and η3. The sliding window of the training sequence is increased by overlapping sliding. Assuming the sequence length is M, a window of length U is cut from the original data, and the sliding operation is performed with a stride of 1, resulting in M-U+1 training samples.
[0052] It should be understood that data normalization is necessary because data collection is theoretically uneven. Performing max-min normalization on the data can also facilitate the construction of power grid load forecasting models based on LSTM networks.
[0053] It should be understood that in federated learning based on temporal network structures, it is necessary to extract the correlation between parameters and time. This is done by inputting the processed training set into the LSTM network to extract the correlation between parameters and time.
[0054] Step 3-2: The preprocessed power data is input into the LSTM network. Based on the parameters in the federated learning of the temporal network structure, it is necessary to mine the correlation between the parameters and time. LSTM provides forget gates, input gates, and output gates. During joint training, parameters are selected and unnecessary parameters are removed to improve computational efficiency. In this patent, multiple LSTMs are added to the client. The more layers there are, the stronger the temporal features are trained. However, according to previous research, the prediction results of multi-layer LSTMs show that building a 3-5 layer LSTM achieves the best results. Beyond a certain number of LSTMs, the feature extraction effect no longer increases.
[0055] The LSTM network provides a forget gate, an input gate, and an output gate. During joint training, parameters are selected and unnecessary parameters are removed to improve computational efficiency. In this application, multiple LSTM networks are added to the client. The more layers there are, the stronger the temporal features are trained. Building a 3-5 layer LSTM network yields the best results. Beyond a certain number of LSTM networks, the feature extraction effect no longer increases.
[0056] It should be understood that the key to LSTM networks lies in the cell state and the selection of various gates, generally including update gates, forget gates, and input gates. In the LSTM network structure, the horizontal line is the core of the entire structure; given any input, the update gate determines the output structure. First, the update gate... Figure 4 The solid line portion is shown: where c t-1 c represents the state information from the previous moment. t This represents the current state information. In an LSTM network, the decision of what information to discard from the cell state is made through a layer called the forget gate. Where h... t-1 This represents the output of the previous cell, and its matrix is assumed to be 1*128 dimensions; x t This represents the input of the current cell, i.e., the training set x after processing by one-dimensional FULL convolution in this application. t ; σ(·) represents the sigmoid function; b f For bias. This update gate will read h. t-1 and x t Output a value between 0 and 1 for each cell in state c. t-1 The numbers in the text are represented as:
[0057] f t =σ(W f ·[h t-1 ,x t ]+b f )
[0058] Among them, f t To update the gate's output, W f This is the weight matrix.
[0059] Among them, the forget gate of LSTM network is as follows Figure 5 As shown by the solid line, less important information is discarded, while the important information is retained for feature extraction. Where h... t-1 This represents the output of the previous cell; the output matrix is manually set to 1*128 dimensions. t This represents the input of the current cell, i.e., the training set x after processing by one-dimensional FULL convolution in this application. tσ(·) represents the sigmoid function. The next step is to determine how much new information to add to the cell state. This involves two steps: first, the sigmoid(·) of the input layer determines which information needs to be updated; a tanh(·) layer generates a vector, which is the candidate content c to be updated. t Where tanh(·) is the activation function, W i and W c As the weight, b i and b c This is the bias. Combining these two parts, we update the cell's state using the following formula:
[0060] i t =σ(W i ·[h t-1 ,x t ]+b i )
[0061]
[0062] Among them, i t For updated information. Update the input state, c t-1 Updated to c t .in, To update the state, combine the input from the previous time step with f. t Multiply, discard the information that needs to be discarded, and then add. These are the new candidate values, which vary depending on the extent to which each state is updated.
[0063] The input gate of an LSTM network is as follows Figure 6 As shown by the solid line, the final output matrix is determined. This output will be based on the cell state and is a filtered output. First, a sigmoid(·) layer is run to determine which part of the cell state will be output. Next, the cell state is processed through tanh(·) and multiplied by the output of the sigmoid(·) gate to finally output the part that determines the output. The specific calculation process is as follows:
[0064] O t =σ(W o [h t-1 ,x t ]+b o )
[0065] h t =O t *tanh(c t )
[0066] The output of the LSTM network can be configured manually. In power data, it is typically set to a 1*n feature matrix F (usually n is set to 128 or 256), where the feature matrix F ∈ R. N×d N is a parameter, d is the feature dimension, and R is a set of matrices.
[0067] Step 3-3: Set the loss function of the LSTM-based power grid load forecasting model as follows:
[0068]
[0069] in, For the kth dataset The logarithmic loss function, To represent the k-th power dataset D k In the i-th sample, x represents the true label value, y represents the predicted value, and n... k Let be the number of samples in the k-th dataset. for, ω is the predicted value under the current model parameters. t Let be the weight at time t. The symbol for summation is 'log()', and the symbol for log is 'logarithm'.
[0070] The training set is used on the client side to train an LSTM-based power grid load forecasting model. The LSTM network of this model is a single recurrent neural network with 20 hidden layers. The network parameters of the output neural network are then updated. The loss function of the k-th participant can be derived from the loss function of the LSTM-based power grid load forecasting model. Finally, the network parameters θ of each LSTM-based power grid load forecasting model are updated using the Adam optimizer. The updated network parameters are given by the following formula:
[0071]
[0072]
[0073] Where, m t Let v be the first moment estimate of the gradient at time t, i.e., the mean of the gradient; let v be the second moment estimate of the gradient at time t, i.e., the biased variance of the gradient; and let g be the mean of the gradient. t Let be the gradient obtained at time t, where t represents the current learning iteration number. ⊙ is a type of multiplication involving element-wise multiplication. γ1 and γ2 are a set of hyperparameters for a power grid load forecasting model based on an LSTM network, where γ1, γ2 ∈ [0, 1), and γ1 = 0.9, γ2 = 0.99. and These are the mean and biased variance of the corrected gradient, where η is the learning rate. Here are the updated network parameters, θ represents the current network parameters, and m represents the current network parameters. t-1 For the first moment estimate of the gradient at time t-1, v t-1 The second moment estimate of the gradient at time t-1. Let be the first hyperparameter at time t. Let be the second hyperparameter at time t, and ∈ be a hyperparameter.
[0074] Steps 3-4: An external attention module is added after the LSTM network to extract important features. It calculates the attention M∈R between the input features and the external storage unit. S×d Where S and d are the hyperparameters of the external attention module, and the expression for the external attention module is:
[0075] A = (α) i,j =Norm(FM) T )
[0076] F out =AM
[0077] Among them, (α) i,j Let F be the similarity between the i-th feature and the j-th row of matrix M. out M is the feature parameter output for attention. T Let F be the transpose of matrix M, and let F be the attention feature parameters. Matrix M is a learnable parameter independent of the input, which is equivalent to the memory of the entire training set. A is the attention map inferred from prior knowledge, and the input features from M are updated through similarity in A. The external attention module implicitly learns the features of the entire input by introducing two external memory units. The two external memory units are M... k and M v M k and M v As keys and values, they are used to increase the network's capacity. The overall algorithm for the external attention module is calculated as follows:
[0078]
[0079] F out =AM v
[0080] in, External memory unit M k The transpose of .
[0081] It should be understood that, since S and d are hyperparameters of the external attention module, the overall algorithm of the external attention module is linear in terms of the number of pixels, and the external attention module allows it to be directly applied to large-scale inputs.
[0082] Steps 3-5: After adding the external attention module, a fully connected layer is used to implement the key output information using a Support Vector Machine (MLP). The MLP has a simple three-layer structure: an input layer, a hidden layer, and an output layer. The input is the output of the LSTM network, the hidden layer has 100 neurons, and the output size is the same as the input. Finally, the output layer of the MLP is the output of the LSTM-based power grid load prediction model. At this point, the LSTM-based power grid load prediction model is established and defined as w. locals w locals These are the network parameters for a power grid load forecasting model based on LSTM networks.
[0083] Step S4: Based on the preliminary network parameters uploaded by each client, the server performs joint training on the global model for power grid load forecasting based on the MMD model transfer method. Then, the server uses the test set and validation set to test and validate the global model for power grid load forecasting based on the LSTM network, obtains the final network parameters, and returns the final network parameters to each client.
[0084] It should be understood that after updating the network parameters of the LSTM-based power grid load forecasting model through step S3 and obtaining the preliminary network parameters of each LSTM-based power grid load forecasting model, these parameters are uploaded to the server for aggregation to generate a global LSTM-based power grid load forecasting model. Aggregating the preliminary network parameters of each LSTM-based power grid load forecasting model on a trusted server differs from traditional federated averaging algorithms, which only average the network parameters of each model without considering the correlation between datasets. Directly applying federated averaging algorithms can lead to negative migration when the source and target domain data are almost uncorrelated. Therefore, considering the distribution distance between the source and target domains, the Maximum Mean Discrepancy (MMD) can be used to detect power data anomalies during fault diagnosis, showing a significant improvement over conventional fault diagnosis. Therefore, this application uses an MMD-based model migration method for power load forecasting, migrating based on the differences between different domains, which is more effective than directly averaging the models. The MMD-based model transfer method effectively measures the differences between sample sets. It uses MMD to measure the differences between various LSTM-based power grid load forecasting models, specifically the distribution differences between the source and target domains. Based on the MMD value, the LSTM-based power grid load forecasting model is adjusted accordingly. The LSTM-based power grid load forecasting model is pre-trained using source domain data and jointly trained and fine-tuned using the target domain data, ultimately resulting in a global LSTM-based power grid load forecasting model with good generalization ability, thus improving the accuracy of power load forecasting.
[0085] Specifically, the transfer learning method based on MMD (maximize mean discrepancy) refers to: based on two distribution models, finding a continuous function φ(·) in the sample space, calculating the mean of the function values of samples from different distributions on φ(·), and obtaining the mean discrepancy of the two distributions corresponding to φ(·) by subtracting the two models. Finding a φ(·) that maximizes this mean discrepancy yields the MMD. Distribution alignment is performed using MMD across multiple domains based on gradients; therefore, the new aggregated gradient is equipped with information from multiple domains and better generalizes to "invisible" test data. Generalization ability can be improved by reducing domain variance and domain alignment. Domain variance can be defined by summing the MMD distances between domain pairs; the domain variance analysis formula is:
[0086]
[0087] Among them, u piand u pj Kernel embeddings representing the distributions of domains i and j, respectively. This distance u represents pi and u pj It is measured by mapping the data into the Regenerated Hilbert Space (RKHS).
[0088] The kernel mean can be expressed using empirical averaging as follows:
[0089]
[0090] Where, μ p φ(x) is the kernel mean. i Let F(ω) be the feature mapping function, so the domain variance can be calculated using the kernel mean. Based on the DNN analysis using neural tangent kernels, the objective function can be reformulated by performing a first-order Taylor expansion on the network objective F(ω):
[0091]
[0092] Where F(ω0) are the initial network parameters, and F(ω) are the final network parameters. To find the gradient, F(ω0) T As the transpose of the initial network parameters, by focusing on the parameter ω, the above approximation can be interpreted as a linear model with respect to ω, and the feature map φ(·) is the gradient initialized to ω0, given as Regarding data x, the domain variance among multiple clients can therefore be defined based on the kernel embeddings in the neural tangential kernel space, given by the following equation:
[0093]
[0094] Among them, grad i and grad j These represent the gradients sent from client i and client j to the server, respectively. Since the neural network is optimized stochastically, the final gradient used for model updates is calculated by averaging the gradients of mini-batch samples; therefore, `grad` can be used. i and grad j The nuclear mean embedding is used to represent the measurement of the distribution difference in the tangential nucleus space of the nerve.
[0095] Gradient aggregation on the centralized server side optimizes w locals Instead of performing gradient averaging to avoid potential gradient conflicts, gradient modification is performed to jointly achieve domain alignment across multiple clients.
[0096] Based on client i and client j, only when a negative transition occurs between client i and client j (i.e. And the expression for the corrected gradient of client i with respect to client j is:
[0097]
[0098] Where λ is the hyperparameter of gradient descent, grad i and grad j These represent the gradients sent from client i and client j to the server, respectively. This is the corrected gradient.
[0099] In one embodiment, the network structure of the global model for power grid load forecasting based on LSTM network includes a one-dimensional full convolution, an LSTM network, an external attention module, and a support vector machine. After the power data is preprocessed by the one-dimensional full convolution, it is input into the LSTM network to mine the correlation features between parameters and time. After the output feature information is input into the external attention module for key information extraction, the extracted key information is input into the output prediction result.
[0100] In one embodiment, the MMD-based model transfer method is used to jointly train the global model for power grid load forecasting based on LSTM networks as follows:
[0101] The gradients of the initial network parameters are corrected by aligning the distribution of multiple domain MMDs based on the gradient. The correction formula is as follows:
[0102]
[0103] Where λ is the hyperparameter of gradient descent, grad i and grad j These represent the gradients sent from client i and client j to the server, respectively. This is the corrected gradient.
[0104] In step S5, each client updates the network parameters of its LSTM-based power grid load forecasting model to the final network parameters, thereby obtaining the global LSTM-based power grid load forecasting model.
[0105] The final network parameters are determined by the server based on the preliminary network parameters uploaded by each client, using the MMD-based model transfer method to jointly train the global model of power grid load prediction based on the LSTM network, and then using the test set and validation set to test and validate the global model of power grid load prediction based on the LSTM network. The network parameters are determined after the set target prediction accuracy is achieved.
[0106] Step S6: Each client obtains the power data for its corresponding area and inputs it into its own LSTM-based global power grid load forecasting model to predict the power load and determine the mean absolute error. The mean absolute error is used to measure the average error between the power consumption of a single area obtained from the decomposition at a certain moment and the actual power input to that area. The expression for calculating the mean absolute error is:
[0107]
[0108] Where MAE is the mean absolute error, g t To represent the actual power consumption of this region at time t, p t Let T represent the total amount of electricity generated at time t, where T represents the number of time points.
[0109] It should be understood that the regional power consumption in this application can refer to the evaluation criteria of the non-intrusive load decomposition model. Here, one of the evaluation indicators, mean absolute error (MAE), is used. It is mainly used to measure the average error between the power consumption of a single region obtained by decomposition at a certain moment and the actual power input to that region. It reflects the power consumption and input of the global power grid load forecasting model based on the LSTM network of each client at a certain moment.
[0110] The aforementioned federated learning-based smart grid load forecasting method classifies the power dataset into training, testing, and validation sets. Each power-consuming unit acts as a client, and 100 areas are selected as clients participating in joint training based on pole numbers. Each client uses the training set to construct its own LSTM-based grid load forecasting model, obtaining preliminary network parameters for each LSTM-based model. These preliminary parameters are then uploaded to the server for joint training. The server, based on the preliminary network parameters uploaded by each client, uses the Model Transfer Modeling (MMD) method to jointly train the global LSTM-based grid load forecasting model. Finally, the test set is used to perform the final training. The test and validation sets are used to test and validate the global model of power grid load forecasting based on LSTM network, obtain the final network parameters, and return the final network parameters to each client. Each client updates the network parameters of its own power grid load forecasting model based on LSTM network to the final network parameters, thus obtaining the global model of power grid load forecasting based on LSTM network for each client. Each client obtains the power data of the corresponding area and inputs it into its own global model of power grid load forecasting based on LSTM network to perform power load forecasting to determine the mean absolute error. As a result, the amount of data processed is greatly reduced, which can solve the problems of data upload delay and large data volume, and improve the processing efficiency of power load forecasting.
[0111] Furthermore, the smart grid load forecasting method based on federated learning in this application utilizes an LSTM network for power load forecasting on the client side. This method can also solve the problem of user privacy information being collected in the data, thus effectively protecting user privacy information. It improves the speed of power load forecasting while also protecting user privacy.
[0112] In one embodiment, a smart grid load forecasting method based on federated learning is provided, with the following steps:
[0113] Step 1: Dataset Selection for Power Data. Daily load data from certain areas of a regional power grid in May 2020 were used to construct and backtest a power grid load forecasting model based on an LSTM network. The sampling interval for the daily load data was 5 minutes, with a total of 288 data points per day. Specific data included initial and final active power, initial and final voltage, output current, and ground susceptance, etc. Finally, the data was collected and integrated into a sequence with a length of M = 288. When constructing the LSTM-based power grid load forecasting model, the historical load data from the previous seven days was used as the training set to train the model and predict the daily load data for the eighth day. The data was divided into three sets: Q1 = 60% for training, Q2 = 20% for validation, and Q3 = 20% for testing. The Q3 test set was used on the server side to validate the performance of the jointly trained LSTM-based global power grid load forecasting model.
[0114] Step Two: The daily load data used for training may contain abrupt changes or missing data. First, one-dimensional data convolution processing is performed. Here, a portion of the data is selected for demonstration. Since all electricity data is displayed digitally, this application only shows a portion of the data undergoing one-dimensional convolution smoothing. In this implementation case, the values of the convolution kernels are η1 = 0.1, η2 = 0.2, and η3 = 0.3; see attached... Figure 3 As shown in the figure, a 1*3 convolution kernel and power data are subjected to a one-dimensional FULL convolution operation. Therefore, a window of length U=3 is cut out on the original data. According to the formula M-U+1, 286 training samples are obtained.
[0115] Step 3: Transmit the matrix-form power data after convolution processing to the LSTM network of the power grid load forecasting model. Here, it is artificially defined that the first input is the previous time step c. t-1 and h t-1Let c0 and h0 be values of 0. This application employs a two-layer LSTM network, with the first layer containing 200 hidden units and the second layer containing 100 hidden units. The final number of hidden units was determined after experimenting with different numbers of hidden units and minimizing the prediction error. During training, the network predicts the leading value at each time step. At each time step, data is learned, and the trained network is updated until each prediction from the previous time step is part of the total data for the next prediction. In this way, the LSTM network is adaptively trained. Figure 7 The diagram shown contains the internal structure of the two-layer LSTM and the input / output diagram of this implementation case, including the data content. The loss is calculated based on the loss function, and finally the output data of the LSTM network structure is obtained. The output of the LSTM network can be set to a certain size. Here, the output matrix is defined as 1*128 (for power data, the output is generally set to 1*128 or 1*256, and 1*128 is used as an example here).
[0116] Step 4: Extract weights from the LSTM-processed data in Step 3. An external attention mechanism is used here to simplify the structure of the local model (i.e., the LSTM-based power grid load forecasting model). The external attention mechanism (i.e., the external attention module) is added to the local model, such as... Figure 8 The external attention mechanism shown here extracts weights without changing the size of the matrix, only changing the information weight ratio of the key points. Therefore, the size of the output matrix is still 1*128.
[0117] Step 5: Output the processing results from Steps 1-4 to the MLP to build a local model. The flowchart for the entire local model building process is attached. Figure 9 As shown;
[0118] Step Six: Steps One through Five completed the local model setup. Step Six involves client-side migration and the establishment of the cloud-based model (i.e., the global model for power grid load forecasting based on the LSTM network). The migration process occurs entirely at the MLP layer; therefore, the preceding model structures remain the same, with model migration only performed at the final MLP layer. The migration process is detailed in the attached diagram. Figure 10 As shown, the cloud refers to the server side;
[0119] Step 7: On the server side, model processing is performed. The locally uploaded MLP model and the server-side MLP model undergo MMD-based migration. The domain difference between the models is calculated using the MMD-based migration method. After multiple iterations, a local model with better generalization performance is obtained and distributed to all participating local models. Specific implementation examples demonstrate that this invention can improve prediction performance while protecting user privacy, achieving the expected results.
[0120] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0121] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0122] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A smart grid load forecasting method based on federated learning, characterized in that, The method includes: Step S1: Classify the power dataset into training set, test set and validation set; Step S2: Using each electricity-consuming unit as a client, select 100 areas as clients participating in the joint training based on the pole numbers. ; Step S3: Each client uses the training set to construct its own LSTM-based power grid load prediction model, obtains the preliminary network parameters of each LSTM-based power grid load prediction model, and uploads the preliminary network parameters to the server for joint training. Step S4: The server performs joint training on the global model for power grid load prediction based on LSTM network using the model transfer method based on MMD according to the preliminary network parameters uploaded by each client. Then, the server uses the test set and validation set to test and validate the global model for power grid load prediction based on LSTM network to obtain the final network parameters and return the final network parameters to each client. Step S5: Each client updates the network parameters of its LSTM-based power grid load prediction model to the final network parameters, thereby obtaining the global LSTM-based power grid load prediction model for each client. Step S6: Each client obtains the power data for its corresponding area and inputs it into its own LSTM-based global power grid load forecasting model to predict the power load and determine the mean absolute error. The mean absolute error is used to measure the average error between the power consumption of a single area obtained from the decomposition at a certain moment and the actual power input to that area. The expression for calculating the mean absolute error is as follows: ; in, The mean absolute error, To represent the actual power consumption in this region at time t, The total amount of electricity generated at time t is given by T, where T represents the number of time points. The LSTM-based power grid load prediction model and the LSTM-based global power grid load prediction model have the same network structure, which includes a one-dimensional full convolution, an LSTM network, an external attention module, and a support vector machine. After the power data is preprocessed by one-dimensional full convolution, it is input into the LSTM network to mine the correlation features between parameters and time. After the feature information is output, it is input into the external attention module to extract key information. The extracted key information is then input into the module to output the prediction result. The model transfer method based on MMD performs joint training of the global model for power grid load forecasting based on LSTM networks in the following way: The gradients of the initial network parameters are corrected by aligning the distribution of multiple domain MMDs based on the gradient. The correction formula is as follows: ; Where λ is a hyperparameter of gradient descent. and These represent the client side, respectively. and client Gradients sent to the server This is the corrected gradient.
2. The method according to claim 1, characterized in that, The loss function of the LSTM-based power grid load forecasting model is: ; in, For the first In the dataset The logarithmic loss function, To represent the k-th power dataset The i-th sample, Indicates the actual value of the label. Indicates the predicted value. Let be the number of samples in the k-th dataset. for, These are the predicted values under the current model parameters. Let be the weight at time t. For the sign of the summation function, This is the logarithmic symbol.
3. The method according to claim 2, characterized in that, The update formula for the network parameters when each client trains its own LSTM-based power grid load prediction model using the training set is as follows: , ; in, Let be the first moment estimate of the gradient at time t, i.e., the mean of the gradient. Let be the second moment estimate of the gradient at time t, i.e., the biased variance of the gradient. Let be the gradient obtained at time t, where t represents the current learning iteration number. , It is a type of multiplication that multiplies elements in the same position. and These are a set of hyperparameters for the LSTM-based power grid load forecasting model. ,definition , , and These are the mean and biased variance of the corrected gradient. For learning rate, For the updated network parameters, For the current network parameters, This is the first moment estimate of the gradient at time t-1. The second moment estimate of the gradient at time t-1. Let be the first hyperparameter at time t. Let be the second hyperparameter at time t. This is a hyperparameter.
Citation Information
Patent Citations
Comprehensive evaluation method and device for energy storage power station system and readable medium
CN114024328A
Federal learning-based cloud-edge collaborative multi-residential-area load prediction method
CN114462683A