Multi-scale power short-term load prediction method based on hierarchical collaborative federated learning
Through a multi-scale short-term load prediction method of multi-scale power with hierarchical collaborative federated learning, using multi-scale training sets and model parameter exchange aggregation, the problems of low prediction accuracy and poor adaptability in the existing technology are solved, and load prediction with higher accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510289685.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-04
AI Technical Summary
The existing short-term power load prediction method based on federated learning is difficult to capture complex patterns at different time scales when training using single-scale data, resulting in low prediction accuracy and poor adaptability to emergencies. In the coordinated training of multi-user data, there are problems of data heterogeneity and model convergence.
A multi-scale short-term load prediction method for power with hierarchical collaborative federated learning is adopted. The local power load data set is divided by the client in different lengths, a multi-scale training set is constructed, and a hierarchical exchange and aggregation of load prediction model parameters is performed in the server. The local load prediction model is trained using the multi-scale training set, and the model parameters are updated and saved.
The accuracy of power load prediction and the ability to predict sudden load fluctuations is improved, the robustness and adaptability of the model are enhanced, and the problems of low prediction accuracy and poor adaptability caused by single-scale training are solved.
Smart Images

Figure CN120256952A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electric load forecasting, and more specifically, to a multi-scale short-term electric load forecasting method based on hierarchical collaborative federated learning. Background Art
[0002] Electric load forecasting, as the core central technology of modern power systems, plays an important role in ensuring the safe and stable operation of power grids, optimizing power generation plans, and power system construction. Electric load forecasting can be classified into four categories according to the time range: ultra-short-term, short-term, medium-term, and long-term. Among them, short-term load forecasting has become the core support for the intelligence of power systems due to its unique time value. On the power generation side, for hydropower generation, accurate short-term load forecasting can guide the reservoir water release period and improve the utilization rate of water energy. On the demand side, accurate short-term load forecasting enables residents to reasonably arrange the usage time of high-energy-consuming equipment, guides residents to use electricity during periods when renewable energy is sufficient, improves the consumption rate of green electricity, and reduces carbon emissions.
[0003] Short-term electric load forecasting is mainly divided into two categories: model-driven and data-driven. Among them, the model-driven method constructs a forecasting model based on physical laws, mathematical equations, or domain knowledge, and relies on expert experience to abstract the system and calibrate parameters. Its advantages are low data requirements and strong interpretability, but the considered factors are relatively single, it is difficult to capture complex non-linear relationships, and the forecasting accuracy is limited when facing complex scenarios. The data-driven method uses machine learning and deep learning to automatically mine patterns from load data and capture non-linear features without explicitly modeling the system mechanism, but it relies on a large amount of historical data, and there is a risk of overfitting when the single-user data scale is insufficient. The short-term electricity consumption of a single household is mainly affected by residents' electricity consumption behavior, and user behavior is highly random and complex. It may not be possible to comprehensively describe the overall electricity consumption pattern of residents only from recent electricity consumption behavior information, because short-term data lacks the representation of "periodic laws" and is easily interfered by accidental events. This method that only uses short-term data has poor robustness to the changing electricity consumption patterns of residents and low forecasting accuracy. Although the collaborative training of multi-user data can improve the generalization ability of the model to a certain extent, in actual operation, information security and privacy protection issues have become restrictive factors, and it is difficult to centralize the data of each user, resulting in the "data island" problem.
[0004] Federated learning, with its decentralized training mechanism and strong data privacy protection ability, provides a new solution for the field of electric load forecasting. The short-term load forecasting method based on federated learning can not only effectively protect the privacy of users' electricity consumption data, but also realize collaborative training among multiple users, while avoiding the overfitting problem caused by insufficient sample data of a single user.
[0005] The prior art discloses a power load forecasting method and system based on federated learning, which randomly selects terminal devices participating in global training based on the device selection rate, performs local training of a power load forecasting model according to the power load data recorded by the terminal devices to obtain local training network parameters, determines whether to upload the local training network parameters to the server according to the network parameter upload threshold, aggregates the local training network parameters received by the server into global parameters, distributes the global parameters to the selected terminal devices for global training of the power load forecasting model, and performs power load forecasting through the trained power load forecasting model to obtain a power load forecasting result. This solution effectively solves the problems of information security and data islands by centrally collecting and training a large amount of electricity consumption data through terminal devices, and improves the convergence speed of the federated learning forecasting model. However, this solution uses single-scale power load data to train the federated learning model, making it difficult to capture complex patterns at different time scales and resulting in the omission of important information. At the same time, it has poor adaptability to sudden events or changes in cross-scale external factors, and is prone to reducing the forecasting accuracy and robustness due to insufficient representation ability at a single time granularity. Because there are significant heterogeneities in the electricity consumption behaviors of different households, the distribution of their load data varies greatly and often shows an unbalanced characteristic. The local load forecasting models of different users often focus on different data features. This data heterogeneity will amplify the update direction conflicts during deep parameter aggregation, making it difficult for traditional federated learning methods to ensure model convergence, resulting in poor performance of the global forecasting model obtained by some clients in actual applications and limited forecasting performance for unseen load data. Summary of the Invention
[0006] To solve the problems of low forecasting accuracy and poor anti-interference ability of the current methods for short-term power load forecasting, the present invention proposes a multi-scale short-term power load forecasting method based on hierarchical collaborative federated learning, which uses a multi-scale training set to train a federated learning load forecasting model, improving the forecasting accuracy. The server exchanges and aggregates the parameters of the load forecasting model, enhancing the model's forecasting ability for sudden load fluctuations.
[0007] To achieve the above technical effects, the technical solution of the present invention is as follows:
[0008] A multi-scale short-term power load forecasting method based on hierarchical collaborative federated learning, the method comprising the following steps:
[0009] S1. The client divides the local power load data set into different time durations to construct a multi-scale training set;
[0010] S2. The client receives the load forecasting model parameters and rounds issued by the server, and updates the local load forecasting model parameters according to the rounds. The local load forecasting model includes a global aggregated load forecasting sub-model, a shallow network, and a deep network;
[0011] S3. The client uses the multi-scale training set to train the local load prediction model, updates and saves the local load prediction model parameters, where the local load prediction model parameters include shallow network parameters, deep network parameters, and global network parameters;
[0012] S4. The client uploads the shallow network parameters, deep network parameters, and global aggregation network parameters to the server. In the server, shallow network parameter aggregation, deep network parameter exchange, and global network parameter aggregation are performed, and the round is updated;
[0013] S5. Determine whether the round reaches the set round. If so, the aggregated global network parameters are sent to the client, and the client updates the local load prediction model parameters and uses the updated local load prediction model to perform short-term power load prediction; otherwise, the aggregated shallow network parameters, the exchanged deep network parameters, and the aggregated global network parameters are used as load prediction model parameters and sent to the client, and return to S2.
[0014] In this technical solution, first, the client divides the local power load data set into different time durations to construct a multi-scale training set and updates the local load prediction model parameters according to the round; uses the constructed multi-scale training set to train the local load prediction model, updates and saves the local load prediction model parameters, improving the accuracy of power load prediction; uses the server to receive the local load prediction models uploaded by different clients. In the server, the load prediction model parameters of different clients are hierarchically exchanged and aggregated and then sent to the client to predict power load data. In the present invention, the exchange and aggregation of model parameters by the server improve the prediction ability of the load prediction model for sudden load fluctuations.
[0015] Preferably, in step S1, the local power load data set is divided into different time durations to construct a multi-scale training set, and the process is as follows:
[0016] Load the local power load data set containing timestamp information into the client, where l k is the kth power load data, d k is the timestamp information of the kth power load data, and K represents the number of power load data;
[0017] Perform date encoding and holiday encoding on the local power load data, and the expression is:
[0018]
[0019] where ty represents different time scales, Ho, Da, Mo represent the hour scale, week scale, and month scale respectively, and d k,tyFor the time information of different time scales of local power load data, N ty For the size of the numerical range of different time scales, et k,ty For the date code of the k-th power load data of different time scales, eh k For the holiday code of the k-th power load data of different time scales, ET is the date code component of the local power load data of different time scales, and EH is the holiday code component of the local power load data;
[0020] Remove the outliers in the local power load data set L = {l1, l2…l K} that do not contain timestamp information;
[0021] Perform seasonal decomposition on the local power load data set L = {l1, l2…l K} that has had outliers removed to obtain the trend component LT = {l T,1 , l T,2 …l T,K} and the seasonal residual component LSR = {l SR,k};
[0022] Perform normalization on the local power load data L = {l1, l2...l K}, the trend component LT = {l r,1 , l T, …l T,K} and the seasonal residual component LSR = {l SR,k} of the local power load data to obtain the normalized local power load data set L′, the normalized trend component LT′ of the local power load data, and the normalized seasonal residual component LSR′ of the local power load data;
[0023] Integrate the date code component ET of the local power load data of different time scales, the holiday code component EH of the local power load data, the normalized local power load data set L′, the normalized trend component LT′ of the local power load data, and the normalized seasonal residual component LSR′ of the local power load data to obtain the data set Ψ o belonging to each client C Trn , and the expression is:
[0024] Ψ Tr,n = {L`, LT`, LSR`, ET, EH};
[0025] Perform data sliding on the data set Ψ Tr,n , and the expression is:
[0026]
[0027] Among them, Slide is for data sliding processing, and K Tr,n is C n is the total number of load data in the training set, W1 and W2 are set sliding windows, W1 > W2, and P is the prediction range. is the training data sample after the first window processing. is the training data sample after the second window processing. is the training label value sample.
[0028] Preferably, the load prediction model described in step S2 is expressed as Among them, represents the global aggregation load prediction sub-model. represents the shallow network. represents the deep network, t represents the round, and the round t satisfies: t ∈ [1, T2 + 1], where T2 is the set global aggregation round and T2 + 1 is the set round.
[0029] Preferably, the process of updating the parameters of the local load prediction model according to the round information in step S2 is as follows:
[0030] If the round t = 1, load the global aggregation load prediction sub-model sent by the server into the local global load prediction model n at the initial moment of the client C and define the prediction task execution flag flag = 0.
[0031] Preferably, the process of updating the parameters of the local load prediction model according to the round information in step S2 is as follows:
[0032] If the round t ∈ (1, T2 + 1] and the prediction task execution flag flag = 0, then update the local load prediction model parameters based on the load prediction model parameters sent by the server at round t. The update process satisfies the expression:
[0033]
[0034] Among them, is the global network parameter of the local load prediction model of the client C n at the current round t. is the shallow network parameter of the local load prediction model of the client C n at the current round t. is the deep network parameter of the local load prediction model of the client C n at the current round t, and β n is the client Cn The updated weight, where T1 is the parameter of the shallow training round.
[0035] Preferably, the local power load data set The power load data therein includes: active power, reactive power, voltage, and current.
[0036] Preferably, the process of training the local load prediction model using the multi-scale training set in step S3 is as follows:
[0037] Obtain the intermediate feature component using the activation function The expression is:
[0038]
[0039] where G is the global network parameter of the local load prediction model of the prediction network main body, σ is the activation function, is the training data sample;
[0040] Concatenate the intermediate feature component with the data sample to be predicted The expression is:
[0041]
[0042] where concat is the concatenation operation;
[0043] Obtain the predicted output value Y using the multi-layer perceptron MLP and the activation function o , and the expression is:
[0044] Y O = MLP(σ(G(F A )))
[0045] Update the global network parameter using the predicted output value Y o and the training label value sample The expression is:
[0046]
[0047] where y o,i is the i-th variable in Y o , is the i-th variable in is the client C n calculates the stochastic gradient under the data set ψ to be slid Tr , η is the network learning rate, SA1 is the total number of samples in the first window of data sliding, and the updated global network parameter including shallow network parameters and deep network parameters
[0048] Preferably, the client in step S4 uploads the shallow network parameters, deep network parameters, and global network parameters to the server, and the process is as follows:
[0049] Judge whether the round t is equal to one of the shallow training round parameter T1 or the global training round parameter T2. If so, the client C n uploads the global network parameters of the updated local load prediction model to the server
[0050] If not, judge whether the round t is less than the shallow training round parameter T1. If the round t is less than the shallow training round parameter T1, the client C n only uploads the shallow network parameters in the local load prediction model to the server
[0051] If the round t is greater than the shallow training round parameter T1, judge whether the round t is less than the global training round parameter T2. If the round t is less than the global training round parameter T2, the client C n only uploads the deep network parameters in the local load prediction model to the server
[0052] If the round t is greater than the global training round parameter T2, the client C n stops uploading the local load prediction model parameters.
[0053] Preferably, in step S4, the server performs aggregation of shallow network parameters, exchange of deep network parameters, and aggregation of global network parameters, and updates the round. The process is as follows:
[0054] Judge whether the round t is equal to one of the shallow training round parameter T1 or the global training round parameter T2. If so, aggregate the global network parameters of the local load prediction model The expression is:
[0055]
[0056] where Q T is the total number of samples of all clients, Q n is the number of samples of client C n and is the local load prediction model parameter on the server side, is the aggregated global network parameter;
[0057] If not, determine whether the round t is less than the shallow training round parameter T1. If the round t is less than the shallow training round parameter T1, aggregate the shallow network parameters in the local load prediction model The expression is:
[0058]
[0059] Wherein, is the aggregated shallow network parameter in the local load prediction model parameters on the server side;
[0060] If the round t is greater than the shallow training round parameter T1, determine whether the round t is less than the global training round parameter T2. If the round t is less than the global training round parameter T2, exchange the deep network parameters in the local load prediction models uploaded by different clients The expression is:
[0061]
[0062] Wherein, is the deep network parameter in the load prediction model parameter of client C n after the exchange; is the deep network parameter in the load prediction model parameter of another client after the exchange; is the deep network parameter in the load prediction model parameter of another client after the exchange;
[0063] After obtaining the local load prediction model parameters on the server side, update the round t. The expression is:
[0064] t new = t + 1
[0065] Wherein, t new is the new round;
[0066] If the round t is greater than the global training round parameter T2, end the aggregation and exchange.
[0067] Preferably, to determine whether the round reaches the set round T2 + 1 in step S5. If so, send the aggregated global network parameters on the server side to client C n Client C n updates the local load prediction model parameters and performs short-term power load prediction using the updated local load prediction model.
[0068] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0069] The present invention proposes that first, the client divides the local power load data set into different time durations to construct a multi-scale training set and updates the parameters of the local load prediction model according to the rounds; the constructed multi-scale training set is used to train the local load prediction model, and the parameters of the local load prediction model are updated and saved, improving the accuracy of power load prediction; the server receives the local load prediction models uploaded by different clients. In the server, after hierarchical exchange and aggregation of the load prediction model parameters of different clients, they are sent to the client to predict the power load data. The present invention improves the prediction ability of the load prediction model for sudden load fluctuations by using the server to exchange and aggregate the model parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 FIG. 6 shows a schematic flowchart of a multi-scale short-term power load prediction method based on hierarchical collaborative federated learning proposed in Embodiment 1 of the present invention;
[0071] Figure 2 FIG. 10 shows a schematic flowchart of step S1 of a multi-scale short-term power load prediction method based on hierarchical collaborative federated learning proposed in Embodiment 2 of the present invention;
[0072] Figure 3 FIG. 14 shows a schematic flowchart of removing outliers from the local power load data set without timestamp information proposed in Embodiment 2 of the present invention;
[0073] Figure 4 FIG. 18 shows a schematic flowchart of performing seasonal decomposition processing on the local power load data set from which outliers have been removed proposed in Embodiment 2 of the present invention;
[0074] Figure 5 FIG. 22 shows a schematic flowchart of step S2 of a multi-scale short-term power load prediction method based on hierarchical collaborative federated learning proposed in Embodiment 2 of the present invention;
[0075] Figure 6 FIG. 26 shows a schematic flowchart of training a local load prediction model using a multi-scale training set proposed in Embodiment 2 of the present invention;
[0076] Figure 7 FIG. 30 shows a schematic flowchart of the client uploading shallow network parameters, deep network parameters, and global network parameters to the server proposed in Embodiment 2 of the present invention;
[0077] Figure 8 FIG. 34 shows a schematic flowchart of performing shallow network parameter aggregation, deep network parameter exchange, and global network parameter aggregation on the server proposed in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0078] The drawings are only for illustrative purposes and should not be construed as limiting the patent;
[0079] To better illustrate this embodiment, certain parts of the drawings are omitted, enlarged or reduced, and do not represent actual sizes;
[0080] For those skilled in the art, it is understandable that some well-known content descriptions in the drawings may be omitted.
[0081] The technical solutions of the present invention will be further described below in conjunction with the drawings and embodiments.
[0082] The description of the positional relationship in the drawings is only for illustrative purposes and should not be construed as a limitation of this patent;
[0083] Embodiment 1
[0084] This embodiment proposes a multi-scale short-term power load forecasting method based on hierarchical collaborative federated learning. The schematic flow diagram of this method is shown in Figure 1 , and includes the following steps;
[0085] S1. The client divides the local dataset into different time durations to construct a multi-scale training set;
[0086] S2. The client receives the load forecasting model parameters and rounds sent by the server, and updates the local load forecasting model parameters according to the rounds. The local load forecasting model includes a global aggregated load forecasting sub-model, a shallow network, and a deep network;
[0087] S3. The client uses the multi-scale training set to train the local load forecasting model, updates and saves the local load forecasting model parameters. The local load forecasting model parameters include shallow network parameters, deep network parameters, and global network parameters;
[0088] S4. The client uploads the shallow network parameters, deep network parameters, and global aggregated network parameters to the server. In the server, shallow network parameter aggregation, deep network parameter exchange, and global network parameter aggregation are performed, and the rounds are updated;
[0089] S5. Determine whether the rounds reach the set rounds. If so, the aggregated global network parameters are sent to the client, and the client updates the local load forecasting model parameters and uses the updated local load forecasting model to perform short-term power load forecasting; otherwise, the aggregated shallow network parameters, exchanged deep network parameters, and aggregated global network parameters are used as the load forecasting model parameters and sent to the client, and return to S2.
[0090] In this embodiment, first, the client divides the local power load dataset into different time durations to construct a multi-scale training set and updates the parameters of the local load prediction model according to the rounds; the constructed multi-scale training set is used to train the local load prediction model, and the parameters of the local load prediction model are updated and saved, improving the accuracy of power load prediction; the server receives the local load prediction models uploaded by different clients. In the server, after hierarchical exchange and aggregation of the load prediction model parameters of different clients, they are sent to the client to predict power load data. The exchange and aggregation of the model parameters by the server improve the prediction ability of the federated learning prediction model for sudden load fluctuations.
[0091] Embodiment 2
[0092] In this embodiment, in step S1, the local power load dataset is divided into different time durations to construct a multi-scale training set. For the process schematic diagram of step S1, see Figure 2 , and the process is as follows:
[0093] Load the local power load dataset containing timestamp information into the client, where l k is the kth power load data, d k is the timestamp information of the kth power load data, and K represents the number of power load data;
[0094] Perform date encoding and holiday encoding on the local power load data. The expression is:
[0095]
[0096] where ty represents different time scales, Ho, Da, and Mo represent the hour scale, week scale, and month scale respectively, d k,ty is the time information of different time scales of the local power load data, N ty is the size of the numerical range of different time scales, et k,ty is the date encoding of the kth power load data of different time scales, eh k is the holiday encoding of the kth power load data of different time scales, ET is the date encoding component of different time scales of the local power load data, and EH is the holiday encoding component of the local power load data;
[0097] Specifically, N Ha = 23, N Do = 6, N Mo = 12. When k = 50, d 50 = 2022-09-10 21:30:00, and the date where d 50 is located is a weekend and it is the Mid-Autumn Festival holiday, then
[0098]
[0099]
[0100] ET
[50] = [0.413, 0.333, 0.250], EH
[50] = [1]
[0101] Remove the outliers from the local power load data set L = {l1, l2... l K}, and the flowchart for removing outliers is shown in Figure 3 , and the process is as follows:
[0102] Define the initial intermediate data component LA based on the local power load data set, and the expression is:
[0103] LA = {l A,1 , l A,2 ... l A,K} = L
[0104] where k is the index value of the corresponding power load data in the intermediate data component LA, and K represents the number of local power load data;
[0105] Set the initial value of the index value k, and determine whether the power load data indicated by the index value k satisfies the expression:
[0106]
[0107] where l k-1 is the load data of the previous sampling point, l k-60×24 / TI is the load data at the same time of the previous day, TT is the minute resolution size of the data, and H1 and H2 are the set first and second outlier ratio thresholds;
[0108] Specifically, TI = 30, H1 = 2, H2 = 2.3;
[0109] When k = 50,
[0110] L = [l1, l2, l3... l 49 , l 50 , l 51 ...] = [1.524, 1.981, 1.969... 1.963, 4.261, 2.481......]
[0111] l k = l 50 = 4.261, l k-1 = l 4z = 1.963, l K-60×24 / TI = l2 = 1.981
[0112] Then l 50 is an outlier,
[0113] When k = 51,
[0114] there is l 51 is a normal value;
[0115] If not satisfied, update the index value k, and the expression is:
[0116] k new = k + 1
[0117] Return to judge whether the power load data indicated by the updated index value k new meets the expression;
[0118] If satisfied, set the power load data in the intermediate data component LA to zero, and record the index value of the index outlier. The expression is:
[0119] l A,k = 0, Q = Q ∪ {k}
[0120] where l A,k is the k-th data variable in LA, and Q is the outlier index set;
[0121] When k = 50, l 50 is an outlier, then l A,50 = 0
[0122] Update the index value k, and the expression is:
[0123] k new = k + 1
[0124] Return to judge whether the power load data indicated by the updated index value k new meets the expression;
[0125] When k reaches the maximum value K = 1440, obtain the intermediate data component after removing the outliers, and perform interpolation processing on the intermediate data component after removing the outliers. The expression is:
[0126]
[0127] where B1 and B2 are the front and rear range sizes of l A,k ;
[0128] Specifically, B1 = B2 = 4,
[0129] When k = 50,
[0130] From the interpolation formula: get After interpolation, LA = [...1.426, 1.831, 1.873, 1.963, 2.049, 2.481, 2.584, 2.312, 1.924...]
[0131] Define the intermediate data component LA after interpolation as the local power load data set L with outliers removed;
[0132] Perform seasonal decomposition on the local power load data set L = {l1, l2...l K} with outliers removed to obtain the trend component LT = {l T,1 , l T,2 ...l T,K > and the seasonal residual component LSR = {l SR,k}. For the flow chart of the seasonal decomposition process, refer to Figure 4 , and the process is as follows:
[0133] Define the initial trend component LT based on the local power load data set. The expression is:
[0134] LT = {l T,1 , l T , 2...l T,K} = L
[0135] where k is the index value corresponding to the power load data in the trend component LT, and K represents the number of local power load data;
[0136] Set the initial value of k to 1;
[0137] Define PE = 48 as the set seasonal period;
[0138] Update the trend variable in the trend component LT based on the index value. The expression is:
[0139]
[0140] Update the index value k. The expression is:
[0141] k new = k + 1
[0142] Return and update the trend variable in the trend component LT based on the updated index value k new ;
[0143] When k = 20, L = [l1, l2, l3...l 24 ...] = [1, 412, 1.137, 0.983...2.398...]
[0144]
[0145] When k = 50,
[0146] L = [...l 48 , l 49 , l 50 , l 51 , l 52 …] = [...1.873, 1.963, 2.049, 2.481, 2.584…]
[0147]
[0148] When k = 1420,
[0149] L = [...l 1418 , l 1419 , l 1420 , l 1421 …l 1441 = [...2.371, 2798, 1783, 2.542…1.303]
[0150]
[0151] When k reaches the maximum value K = 1440, the updated trend component is obtained The seasonal residual component LSR of the local power load data is calculated, and the expression is:
[0152] LSR = L - LT.
[0153] For the local power load data L = {l1, l2…l K}, the trend component LT of the local power load data = {l T,1 , l T,2 …l T,k} and the seasonal residual component LSR of the local power load data = {l SR,k}, normalization processing is performed to obtain the normalized local power load data set L′, the normalized trend component LT′ of the local power load data, and the normalized seasonal residual component LSR′ of the local power load data;
[0154] Integrate the date coding component ET of different time scales of the local power load data, the holiday coding component EH of the local power load data, the normalized local power load data set L′, the normalized trend component LT′ of the local power load data, and the normalized seasonal residual component LSR′ of the local power load data to obtain the data set Ψ n belonging to each client C Tr,n , and the expression is:
[0155] Ψ Tr,n = {L, LT, LASR, ET, EH};
[0156] For the dataset Ψ Tr,n perform data sliding processing, and the expression is:
[0157]
[0158] where Slide is the data sliding processing, K Tr, n is the total number of C n the total number of training set load data, W1 and W2 are the set sliding windows, W1 > W2, P is the prediction range, is the training data sample after the first window processing, is the training data sample after the second window processing, is the training label value sample;
[0159] The specific process of data sliding processing is:
[0160] Define the datasets Ψ of different time scales Tr,n as the dataset Ψ to be slid I ;
[0161] Define the total number W of the first data sliding window as W I = 336, and the expression is:
[0162] W = W1
[0163] Define the total number of samples SA1 of the first data sliding window, and the expression is:
[0164] SA1 = K - P - W1 + 1
[0165] where K = 1440 represents the number of local power load data in the dataset Ψ to be slid I P = 24 is the prediction range, then SA1 = 1440 - 336 - 24 + 1 = 1081;
[0166] Extract the sub-dataset V1 without the seasonal residual component LSR' from the dataset Ψ to be slid I The expression is:
[0167]
[0168] Divide the data samples based on the sub-dataset V1 without the seasonal residual component LSR' using the data sliding window. The expression is:
[0169]
[0170] where, The data sample divided by sliding the first window for the w-th data
[0171] Use the first data sliding window to divide the label value samples based on the locally processed power load data after normalization. The expression is:
[0172]
[0173] Where is the label value sample divided by sliding the first window for the w-th data
[0174] Integrate to obtain the training data samples after processing by the first data sliding window Training label value samples The expression is:
[0175]
[0176] Define the total number W of the second data sliding window as W2 = 48. The expression is:
[0177] W = W2
[0178] Define the total number of samples SA2 of the second data sliding window as SA2 = 1369. The expression is:
[0179] SA2 = K - P - W2 + 1
[0180] Extract from the data set Ψ to be slid I The sub-data set V2 containing only the seasonal residual component LSR' and the locally processed power load data set L' after normalization is expressed as:
[0181]
[0182] Use the second data sliding window to divide the data samples based on the sub-data set V2 without the seasonal residual component LSR'. The expression is:
[0183]
[0184] Where is the data sample divided by sliding the second window for the w-th data
[0185] Integrate to obtain the data samples to be predicted after processing by the second data sliding window The expression is:
[0186]
[0187] In this embodiment, the load prediction model described in step S1 is expressed as Where represents the global aggregate load forecasting sub-model, represents a shallow network, represents a deep network, t represents a round, and round t satisfies: t∈[1, T2+1], where T2=10 is the set global aggregation round, and T2+1=11 is the set round;
[0188] In this embodiment, in step S1, the local load forecasting model parameters are updated according to the round information, and the process is as follows:
[0189] If round t = 1, the global aggregate load prediction sub-model sent by the server Load into client C n The local global load forecasting model at the initial moment And define the prediction task execution flag flag = 0.
[0190] In this embodiment, in step S1, the local load forecasting model parameters are updated according to the round information, and the process is as follows:
[0191] If the round t∈(1, 11], and the prediction task execution flag flag=0, then based on the load prediction model parameters sent by the server in round t, the local load prediction model parameters are updated, and the update process satisfies the expression:
[0192]
[0193] in, For client C n The global network parameters of the local load forecasting model at the current round t, For client C n The shallow network parameters of the local load forecasting model at the current round t, For client C n The deep network parameters of the local load forecasting model at the current round t, β n =0.6 for client C n The update weight of , T1 = 5 is the shallow training round parameter;
[0194] The flowchart of step S1 is shown in Figure 5 .
[0195] In this embodiment, the local power load data set The power load data in the system include: active power, reactive power, voltage and current.
[0196] In this embodiment, the multi-scale training set is used to train the local load forecasting model in step S3. The training flowchart is shown in Figure 6 , the process is:
[0197] Obtain the intermediate feature component using the activation function The expression is:
[0198]
[0199] where G are the global network parameters of the local load prediction model of the prediction network main body, that is, the temporal convolutional neural network TCN, σ is the activation function ReLU, is the training data sample; the kernel_size of the prediction network main body G is 3, num_chanels = [32, 64], and it has two Temporal Blocks. In the case of batch_size = 64, a sample passes through G to obtain a sub tensor of size (64, 48, 6), and the subsample tensor of the sub is of size (64, 48, 2);
[0200] Concatenate the intermediate feature component with the data sample to be predicted The expression is:
[0201]
[0202] where concat is the concatenation operation;
[0203] Use the multi-layer perceptron MLP and the activation function to obtain the predicted output value Y o , and the expression is:
[0204] Y o = MLP(σ(G(F A )))
[0205] Use the predicted output value Y o and the training label value sample to update the global network parameters The expression is;
[0206]
[0207] where y o,i is the i-th variable in Y o , is the i-th variable in is the client C n calculates the stochastic gradient under the dataset ψ to be processed by sliding, η = 0.0001 is the network learning rate, SA1 = 1081 is the total number of samples in the first window of data sliding, Tr For the mean squared error loss function MSE, the updated global network parameters include the shallow network parameters and the deep network parameters
[0208] In this embodiment, the client in step S4 uploads the shallow network parameters, deep network parameters, and global network parameters to the server. For the flowchart, see Figure 7 , and the process is as follows:
[0209] Judge whether the round t is equal to one of the shallow training round parameter T1 or the round t is equal to the global training round parameter T2. If so, the client C n uploads the global network parameters of the updated local load prediction model to the server
[0210] Specifically, θ Lc is composed of TCN and MLP. Regard the TemporalBlock of TCN as one layer of the network. MLP contains three fully connected layers. Regard it as the shallow network of the load prediction model. Regard it as the deep network of the load prediction model, D = 7, S = 4;
[0211] If not, then judge whether the round t is less than the shallow training round parameter T1. If the round t is less than the shallow training round parameter T1, then the client C n only uploads the shallow network parameters in the local load prediction model to the server
[0212] Specifically, T1 = 5;
[0213] If the round t is greater than the shallow training round parameter T1, then judge whether the round t is less than the global training round parameter T2. If the round t is less than the global training round parameter T2, then the client C n only uploads the deep network parameters in the local load prediction model to the server
[0214] Specifically, T2 = 10;
[0215] If the round t is greater than the global training round parameter T2, then the client C n stops uploading the local load prediction model parameters.
[0216] In this embodiment, in step S4, the server performs shallow network parameter aggregation, deep network parameter exchange, and global network parameter aggregation, and updates the round. For the flowchart, see Figure 8 , and the process is as follows:
[0217] Determine whether the round number t is equal to one of the shallow training round parameter T1 = 5 or the global training round parameter T2 = 10. If so, aggregate the global network parameters of the local load prediction model The expression is:
[0218]
[0219] Among them, Q T is the total number of samples of all clients, Q n is the number of samples of client C n . is the local load prediction model parameter of the server side, is the aggregated global network parameter;
[0220] Specifically, Q T = 175840, Q1 = 1081, Q 100 = 1728. When t = 5, the expression is:
[0221]
[0222] If not, determine whether the round number t is less than the shallow training round parameter T1. If the round number t is less than the shallow training round parameter T1, aggregate the shallow network parameters in the local load prediction model The expression is:
[0223]
[0224] Among them, is the aggregated shallow network parameter in the local load prediction model parameter of the server side;
[0225] If the round number t is greater than the shallow training round parameter T1, then determine whether the round number t is less than the global training round parameter T2. If the round number t is less than the global training round parameter T2, exchange the deep network parameters in the local load prediction models uploaded by different clients The expression is:
[0226]
[0227] Among them, is the load prediction model parameter of client C n after the exchange in the deep network parameter; is the deep network parameter in the load prediction model parameter of another client after the exchange;
[0228] Specifically, among them, n and m are different client numbers, and each client exchanges once; in this embodiment, when n = 1, m = 100, and t = 8, the expression is:
[0229]
[0230] After obtaining the local load prediction model parameters on the server side, update the round t, and the expression is:
[0231] t new = t + 1
[0232] where t new is the round of the new round;
[0233] If the round t is greater than the global training round parameter T2, then terminate the aggregation and exchange.
[0234] In this embodiment, to determine whether the round reaches the set round T2 + 1 in step S5, if so, the aggregated global network parameters on the server side are sent to the client C n The client C n updates the local load prediction model parameters, and uses the updated local load prediction model to perform short-term power load prediction.
[0235] Obviously, the above embodiments of the present invention are only examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A multi-scale short-term power load forecasting method based on hierarchical collaborative federated learning, characterized in that It includes the following steps: S1. The client divides the local power load dataset into different time durations to construct a multi-scale training set; S2. The client receives the load prediction model parameters and rounds sent by the server, and updates the local load prediction model parameters according to the rounds. The local load prediction model includes a global aggregated load prediction sub-model, a shallow network, and a deep network; S3. The client uses the multi-scale training set to train the local load prediction model, updates and saves the local load prediction model parameters. The local load prediction model parameters include shallow network parameters, deep network parameters, and global network parameters; S4. The client uploads the shallow network parameters, deep network parameters, and global aggregated network parameters to the server. In the server, shallow network parameter aggregation, deep network parameter exchange, and global network parameter aggregation are performed, and the rounds are updated; S5. Determine whether the rounds reach the set rounds. If so, the aggregated global network parameters are sent to the client, and the client updates the local load prediction model parameters and uses the updated local load prediction model to perform short-term power load prediction; Otherwise, the aggregated shallow network parameters, exchanged deep network parameters, and aggregated global network parameters are used as the load prediction model parameters and sent to the client, and return to S2.
2. The multi-scale short-term electric load forecasting method based on hierarchical collaborative federated learning according to claim 1, wherein, In step S1, the local power load dataset is divided into different time durations to construct a multi-scale training set. The process is as follows: Load the local power load data set containing timestamp information into the client, where l k is the k-th power load data, and d k is the timestamp information of the k-th power load data, and K represents the number of power load data; The local power load data is encoded by date and holiday. The expression is: Among them, ty represents different time scales, and Ha, Da, and Mo represent the hour scale, weekly scale, and monthly scale respectively, where d k,ty is the time information of different time scales of local power load data, and N ty is the size of the numerical range of different time scales, and et k,ty is the date code of the k-th power load data of different time scales, and eh k is the holiday code of the k-th power load data of different time scales. ET is the date code component of different time scales of local power load data, and EH is the holiday code component of local power load data; Remove the outliers from the local power load data set L = {l1, l2... l K}; The local power load data set L = {l1, l2... l K ) is subjected to seasonal decomposition to obtain the trend component LT = {l T,1 , l T,2 ... l T,K} and the seasonal residual component LSR = {l SR,k} of the local power load data with outliers removed; Normalize the local power load data \(L = \{l_1, l_2...l\ K \}\), the trend component \(LT = \{l\ T,1 , l\ T,2 ...l\ T,K \}\) and the seasonal residual component \(LSR=\{l\ SR,k \}\) of the local power load data to obtain the normalized local power load data set \(L'\), the normalized trend component \(LT'\) of the local power load data, and the normalized seasonal residual component \(LSR'\) of the local power load data; Integrate the date-coded component ET of local power load data at different time scales, the holiday-coded component EH of local power load data, the set L′ of normalized local power load data, the trend component LT′ of normalized local power load data, and the seasonal residual component LSR′ of normalized local power load data to obtain the data set Ψ belonging to each client C n as follows Tr,n , and the expression is: Ψ TTr,u = {L, LT, LSR, ET, EH}; For the data set Ψ Tr,n Perform data sliding processing, and the expression is: Among them, Slide is for data sliding processing, K Tr,n is C n is the total number of training set load data, W1 and W2 are set sliding windows, W1 > W2, p is the prediction range, is the training data sample after the first window processing, is the training data sample after the second window processing, is the training label value sample.
3. A multi-scale short-term power load forecasting method based on hierarchical collaborative federated learning according to claim 2, characterized in that, The load prediction model described in step S2 is expressed as x ∈ {G, SH, DE}, where represents the global aggregation load prediction sub-model, represents the shallow network, represents the deep network, t represents the round, and the round t satisfies: t ∈ [1, T2 + 1], where T2 is the set global aggregation round and T2 + 1 is the set round.
4. A multi-scale short-term power load forecasting method based on hierarchical collaborative federated learning according to claim 3, characterized in that, The process of updating the local load prediction model parameters according to the round information described in step S2 is as follows: If the round t = 1, load the global aggregated load prediction sub-model sent by the server into the local global load prediction model at the initial moment of the client C n , and define the prediction task execution flag flag = 0. 5. A multi-scale short-term power load forecasting method based on hierarchical collaborative federated learning according to claim 4, characterized in that The process of updating the local load prediction model parameters according to the round information described in step S2 is as follows: If the round t ∈ (1, T2 + 1] and the prediction task execution flag flag = 0, then based on the load prediction model parameters sent by the server side in round t, the local load prediction model parameters are updated, and the update process satisfies the expression: Among them, is the client C n the global network parameters of the local load prediction model at the current round t, is the client C n the shallow network parameters of the local load prediction model at the current round t, is the client C n the deep network parameters of the local load prediction model at the current round t, β n is the client C n the updated weight, and T1 is the shallow training round parameter.
6. The multi-scale short-term power load forecasting method based on hierarchical collaborative federated learning according to claim 5, characterized in that, The local power load data set The power load data therein includes: active power, reactive power, voltage, and current.
7. A multi-scale short-term power load forecasting method based on hierarchical collaborative federated learning according to claim 6, characterized in that, The process of training the local load prediction model using the multi-scale training set described in step S3 is as follows: Obtain intermediate feature components using an activation function The expression is: Among them, G is the global network parameter of the local load prediction model is the prediction network main body, σ is the activation function, is the training data sample; Concatenate the intermediate feature component with the data sample to be predicted The expression is as follows: Among them, concat is a concatenation operation; Use a multi-layer perceptron MLP and an activation function to obtain the predicted output value Y O , and the expression is: Y O = MLP(σ(G(F A ))) Using the predicted output value and the training label value samples Update the global network parameters The expression is: Among them, y O,i is the i-th variable in Y O , is the i-th variable in client C n is the stochastic gradient calculated under the dataset Ψ to be processed by sliding Tr , η is the network learning rate, SA1 is the total number of samples in the first window of data sliding, and the updated global network parameters include shallow network parameters and deep network parameters 8. The multi-scale short-term power load forecasting method based on hierarchical collaborative federated learning according to claim 7, wherein The process of the client uploading the shallow network parameters, deep network parameters, and global network parameters described in step S4 is as follows: Determine whether round t is equal to the shallow training round parameter T1 or round t is equal to the global training round parameter T2. If so, client C n uploads the global network parameters of the updated local load prediction model to the server Otherwise, it is determined whether the round t is less than the shallow training round parameter T1. If the round t is less than the shallow training round parameter T1, then the client C n only uploads the shallow network parameters in the local load prediction model to the server If the round t is greater than the shallow training round parameter T1, then determine whether the round t is less than the global training round parameter T2. If the round t is less than the global training round parameter T2, then the client C n only uploads the deep network parameters in the local load prediction model to the server If the round t is greater than the global training round parameter T2, then the client C n stops uploading the local load prediction model parameters.
9. The multi-scale short-term power load forecasting method based on hierarchical collaborative federated learning according to claim 8, characterized in that, The process of performing shallow network parameter aggregation, deep network parameter exchange, and global network parameter aggregation in the server and updating the rounds described in step S4 is as follows: Determine whether the judgment round t is equal to one of the shallow training round parameter T1 or the global training round parameter T2. If so, aggregate the global network parameters of the local load prediction model The expression is: Among them, Q r is the total sample number of all clients, and Q n is the sample number of client C n . are the local load prediction model parameters of the server side, and are the aggregated global network parameters; Otherwise, it is determined whether the round t is less than the shallow training round parameter T1. If the round t is less than the shallow training round parameter T1, the shallow network parameters in the local load prediction model are aggregated The expression is: Among them, are the aggregated shallow network parameters in the local load prediction model parameters on the server side; If the round t is greater than the shallow training round parameter T1, then it is judged whether the round t is less than the global training round parameter T2. If the round t is less than the global training round parameter T2, then the deep network parameters in the local load prediction models uploaded by different clients are exchanged The expression is: Among them, is the load prediction model parameter of the exchanged client C n in the deep network parameter; is the deep network parameter in the load prediction model parameter of another exchanged client; is the deep network parameter in the load prediction model parameter of another exchanged client; After obtaining the local load prediction model parameters on the server side, the round t is updated. The expression is: t new = t + 1 where t new is the round number of a new round; If the round t is greater than the global training round parameter T2, then the aggregation and exchange are ended.
10. A multi-scale short-term power load forecasting method based on hierarchical collaborative federated learning according to claim 9, characterized in that, Determine whether the number of judgment rounds reaches the set round T2+1 as described in step S5. If so, the aggregated global network parameters on the server side are sent to the client C n , and the client C n updates the local load forecasting model parameters and performs short-term power load forecasting using the updated local load forecasting model.
Citation Information
Cited By
Cross-network collaborative security alarm noise reduction method based on security federal learning
CN121217538A