Training method and device of power demand response prediction model, equipment, storage medium and program product
By introducing a regular term loss function in the distributed cluster of virtual power plants, the problem that the power demand response prediction model in the prior art is difficult to adapt to the power demand differences in different regions and regions, and a more efficient power demand response prediction is achieved.
Patent Information
- Application Number
- CN202510124500.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively adapt to the power demand differences in different regions and regions in virtual power plants, resulting in the power demand response prediction model not being able to achieve ideal prediction results in the local environment of some aggregators.
By receiving the global prediction model sent by the management node on each worker node in the distributed cluster, initializing the local prediction model, and iteratively training based on the local sample data, regular terms are introduced when determining the loss function to ensure that the local prediction model does not deviate too far from the global prediction model when it is updated.
The prediction effect of the power demand response prediction model is improved, the problem of overfitting or deviation from the global prediction model due to data distribution differences is reduced, and the prediction effect of the global prediction model is improved.
Smart Images

Figure CN120013007A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of electric power technology, and in particular to a training method, device, equipment, storage medium and program product for an electric power demand response prediction model. Background Art
[0002] In a virtual power plant, load aggregators can achieve flexible regulation and efficient management of the power grid by integrating multiple distributed energy resources (such as photovoltaics, wind power, energy storage equipment, etc.) and the needs of power users. Among them, power demand response (or demand response) as a key regulation mechanism can help balance the power grid load and optimize energy distribution by guiding consumers to change their electricity consumption behavior (such as reducing electricity consumption during peak hours or increasing electricity consumption during off-peak hours). Therefore, accurate demand response forecasting is crucial for effective grid regulation.
[0003] At present, researchers mainly use global models trained by standard federated learning to solve the problem of demand response forecasting. This method can jointly train the global model without exchanging original data, protect the data privacy of each load aggregator, and make full use of multi-party data resources to significantly improve the forecasting performance. However, the data characteristics of load aggregators in virtual power plants vary significantly due to factors such as geographical location, climatic conditions, and user behavior. For example, there are huge differences in electricity consumption patterns in different regions (for example, some regions may be dominated by industrial loads, while other regions are dominated by residential loads) or there are differences in seasonal electricity demand in different regions (such as heating loads in the north in winter and cooling loads in the south in summer). Due to these distribution differences, the global model trained by standard federated learning is difficult to adapt to local data, resulting in the inability to achieve ideal forecasting results in the local environment of some aggregators.
[0004] Therefore, how to improve the prediction effect of power demand response prediction model has become an urgent problem to be solved. Summary of the invention
[0005] The embodiments of the present application provide a method, apparatus, device, storage medium and program product for training an electricity demand response prediction model, which can improve the prediction effect of the electricity demand response prediction model.
[0006] In a first aspect, an embodiment of the present application provides a method for training a power demand response prediction model, which is applied to each working node in a distributed cluster; the method comprises:
[0007] receiving a global prediction model for global power demand response prediction sent from a management node in a distributed cluster;
[0008] Based on the global prediction model, a local prediction model for local power demand response prediction is initialized to obtain an initialized local prediction model;
[0009] Using the initialized local prediction model, determine the current predicted demand response value corresponding to the local sample data set related to the power demand response prediction, and determine the loss function corresponding to the initialized local prediction model based on the current predicted demand response value; the loss function includes the error between the actual demand response value corresponding to the local sample data set and the current predicted demand response value, and the regularization term between the global prediction model and the local prediction model;
[0010] With the goal of determining the minimum value of the loss function, the initialized local prediction model is iteratively trained based on the local sample data set to obtain a trained local prediction model;
[0011] The trained local prediction model is sent to the management node, so that the management node aggregates the trained local prediction models respectively sent by multiple working nodes to obtain a processed global prediction model.
[0012] In one of the embodiments, an initialized local prediction model is used to determine a current predicted demand response value corresponding to a local sample data set related to power demand response prediction, including: obtaining a local sample data set related to power demand response prediction, and an actual demand response value corresponding to the sample data set; encoding each sample data in the sample data set to obtain a data feature corresponding to each sample data; normalizing multiple data features respectively to obtain multiple normalized data features; and using the initialized local prediction model to obtain a current predicted demand response value corresponding to the sample data set based on the multiple normalized data features.
[0013] In one of the embodiments, an initialized local prediction model is used to obtain a current predicted demand response value corresponding to a sample data set based on multiple normalized data features, including: adding time series information to the multiple normalized data features respectively to obtain multiple data features with added time series information; inputting the multiple data features with added time series information into the initialized local prediction model to obtain a current predicted demand response value corresponding to the sample data set.
[0014] In one of the embodiments, the initialized local prediction model includes an encoder and a decoder; multiple data features with added timing information are input into the initialized local prediction model to obtain a current predicted demand response value corresponding to the sample data set, including: inputting multiple data features with added timing information into the encoder to obtain an encoding result; inputting the encoding result and the historical predicted demand response value into the decoder to obtain the current demand response value corresponding to the sample data set; the historical predicted demand response value is used to represent the demand response prediction result of the historical local prediction model for the local sample data set, and the historical local prediction model is the previous prediction model of the initialized local prediction model.
[0015] In one of the embodiments, multiple data features are respectively normalized to obtain multiple normalized data features, including: determining a mean and a standard deviation corresponding to each of the multiple data features; for each data feature, normalizing the targeted data feature based on the mean and the standard deviation corresponding to the targeted data feature to obtain a normalized data feature.
[0016] In one of the embodiments, the method further includes: determining an actual demand response value corresponding to the local sample data set based on the actual load corresponding to the demand response time period of all local electricity users and the load value that did not participate in the demand response in the historical time period.
[0017] In a second aspect, the present application provides a training device for a power demand response prediction model, which is applied to each working node in a distributed cluster; the device includes:
[0018] A receiving module, used for receiving a global prediction model for global power demand response prediction sent from a management node in a distributed cluster;
[0019] A processing module, used for initializing a local prediction model for local power demand response prediction based on a global prediction model to obtain an initialized local prediction model;
[0020] A determination module is used to determine the current predicted demand response value corresponding to the local sample data set related to the power demand response prediction using the initialized local prediction model, and determine the loss function corresponding to the initialized local prediction model based on the current predicted demand response value; the loss function includes the error between the actual demand response value corresponding to the local sample data set and the current predicted demand response value, and the regularization term between the global prediction model and the local prediction model;
[0021] A training module is used to iteratively train the initialized local prediction model based on the local sample data set with the goal of determining the minimum value of the loss function to obtain a trained local prediction model;
[0022] The sending module is used to send the trained local prediction model to the management node, so that the management node aggregates the trained local prediction models sent by multiple working nodes to obtain a processed global prediction model.
[0023] In a third aspect, the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0024] receiving a global prediction model for global power demand response prediction sent from a management node in a distributed cluster;
[0025] Based on the global prediction model, a local prediction model for local power demand response prediction is initialized to obtain an initialized local prediction model;
[0026] Using the initialized local prediction model, determine the current predicted demand response value corresponding to the local sample data set related to the power demand response prediction, and determine the loss function corresponding to the initialized local prediction model based on the current predicted demand response value; the loss function includes the error between the actual demand response value corresponding to the local sample data set and the current predicted demand response value, and the regularization term between the global prediction model and the local prediction model;
[0027] With the goal of determining the minimum value of the loss function, the initialized local prediction model is iteratively trained based on the local sample data set to obtain a trained local prediction model;
[0028] The trained local prediction model is sent to the management node, so that the management node aggregates the trained local prediction models respectively sent by multiple working nodes to obtain a processed global prediction model.
[0029] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the following steps are implemented:
[0030] receiving a global prediction model for global power demand response prediction sent from a management node in a distributed cluster;
[0031] Based on the global prediction model, a local prediction model for local power demand response prediction is initialized to obtain an initialized local prediction model;
[0032] Using the initialized local prediction model, determine the current predicted demand response value corresponding to the local sample data set related to the power demand response prediction, and determine the loss function corresponding to the initialized local prediction model based on the current predicted demand response value; the loss function includes the error between the actual demand response value corresponding to the local sample data set and the current predicted demand response value, and the regularization term between the global prediction model and the local prediction model;
[0033] With the goal of determining the minimum value of the loss function, the initialized local prediction model is iteratively trained based on the local sample data set to obtain a trained local prediction model;
[0034] The trained local prediction model is sent to the management node, so that the management node aggregates the trained local prediction models respectively sent by multiple working nodes to obtain a processed global prediction model.
[0035] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:
[0036] receiving a global prediction model for global power demand response prediction sent from a management node in a distributed cluster;
[0037] Based on the global prediction model, a local prediction model for local power demand response prediction is initialized to obtain an initialized local prediction model;
[0038] Using the initialized local prediction model, determine the current predicted demand response value corresponding to the local sample data set related to the power demand response prediction, and determine the loss function corresponding to the initialized local prediction model based on the current predicted demand response value; the loss function includes the error between the actual demand response value corresponding to the local sample data set and the current predicted demand response value, and the regularization term between the global prediction model and the local prediction model;
[0039] With the goal of determining the minimum value of the loss function, the initialized local prediction model is iteratively trained based on the local sample data set to obtain a trained local prediction model;
[0040] The trained local prediction model is sent to the management node, so that the management node aggregates the trained local prediction models respectively sent by multiple working nodes to obtain a processed global prediction model.
[0041] The training method, device, equipment, storage medium and program product of the above-mentioned power demand response prediction model are applied to each working node in a distributed cluster; each working node can receive a global prediction model for global power demand response prediction sent by a management node in the distributed cluster; based on the global prediction model, the local prediction model for local power demand response prediction is initialized to obtain an initialized local prediction model; using the initialized local prediction model, the current predicted demand response value corresponding to the local sample data set related to the power demand response prediction is determined, and based on the current predicted demand response value, the loss function corresponding to the initialized local prediction model is determined; the loss function includes the error between the actual demand response value corresponding to the local sample data set and the current predicted demand response value, as well as the regularization term between the global prediction model and the local prediction model; with the minimum value of the loss function as the goal, the initialized local prediction model is iteratively trained based on the local sample data set to obtain a trained local prediction model; the trained local prediction model is sent to the management node so that the management node aggregates the trained local prediction models sent by multiple working nodes to obtain a processed global prediction model. By adopting this method, a regularization term is introduced on the basis of the federated learning algorithm, so that each local prediction model will not deviate too far from the global prediction model when it is updated, thereby avoiding the problem of overfitting or excessive deviation from the global prediction model due to data distribution differences. In this way, each working node (i.e., load aggregator) can use the global prediction model as a "benchmark" when optimizing the local prediction model, so that the update of the local prediction model is kept within a reasonable range, thereby reducing the impact of data distribution differences between different load aggregators on the global prediction model, so as to improve the prediction effect of the global prediction model, and then improve the prediction effect of each local prediction model, that is, improve the prediction effect of the power demand response prediction model. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0043] Figure 1 It is a schematic diagram of an application scenario of a training method for a power demand response prediction model provided in an embodiment of the present application;
[0044] Figure 2 It is a flowchart of a method for training a power demand response prediction model provided in an embodiment of the present application;
[0045] Figure 3 It is a schematic diagram of the overall process of a training method for a power demand response prediction model provided in an embodiment of the present application;
[0046] Figure 4 It is a structural schematic diagram of a training device for a power demand response prediction model provided in an embodiment of the present application;
[0047] Figure 5 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0049] The following introduces the application scenarios of the training method of the power demand response prediction model provided in the embodiment of the present application.
[0050] See also Figure 1 , Figure 1 Schematic diagram of an application scenario of a training method for a power demand response prediction model provided in an embodiment of the present application. Figure 1 As shown, it includes a management node 101 (or referred to as a central computer device 101) and multiple working nodes (i.e., computer devices corresponding to multiple load aggregators, respectively) ( Figure 1 102 and the working node 103 are taken as examples for drawing). The management node 101 and the working node 102 and the working node 103 respectively transmit data through the network.
[0051] Among them, the working node 102 and the working node 103 can be respectively from the global prediction model for global power demand response prediction sent by the management node 101. Afterwards, the working node 102 and the working node 103 can respectively initialize the local prediction model for local power demand response prediction based on the global prediction model to obtain the initialized local prediction model; using the initialized local prediction model, determine the current predicted demand response value corresponding to the local sample data set related to the power demand response prediction, and determine the loss function corresponding to the initialized local prediction model based on the current predicted demand response value; the loss function includes the error between the actual demand response value corresponding to the local sample data set and the current predicted demand response value, as well as the regularization term between the global prediction model and the local prediction model; with the minimum value of the loss function as the goal, iteratively train the initialized local prediction model based on the local sample data set to obtain the trained local prediction model; send the trained local prediction model to the management node 101, so that the management node aggregates the trained local prediction models sent by multiple working nodes to obtain the processed global prediction model. By adopting this method, a regularization term is introduced on the basis of the federated learning algorithm, so that each local prediction model will not deviate too far from the global prediction model when it is updated, thereby avoiding the problem of overfitting or excessive deviation from the global prediction model due to data distribution differences. In this way, each working node (i.e., load aggregator) can use the global prediction model as a "benchmark" when optimizing the local prediction model, so that the update of the local prediction model is kept within a reasonable range, thereby reducing the impact of data distribution differences of different load aggregators on the global prediction model, so as to improve the prediction effect of the global prediction model, and then improve the prediction effect of each local prediction model.
[0052] Optionally, the management node 101, the working node 102, and the working node 103 may all be terminal devices or servers. The terminal devices mentioned here may include but are not limited to: smart phones, tablet computers, laptop computers, desktop computers, smart watches, smart TVs, smart car terminals, etc. The server mentioned here may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers.
[0053] See also Figure 2 , Figure 2 1 is a flow chart of a method for training a power demand response prediction model provided in an embodiment of the present application. The method can be executed by each working node in a distributed cluster (for example, the working node 101 and the working node 102 mentioned above). Figure 2 As shown, the training method of the power demand response prediction model may include but is not limited to the following steps:
[0054] S201. Receive a global prediction model for global power demand response prediction sent from a management node in a distributed cluster.
[0055] The global prediction model received by each working node can also be understood as the model parameters of the received global prediction model. For example, the model parameters of the global prediction model received by the working node are w t , where w t It represents the model parameters of the global prediction model in the tth round of training, w t →k, k=1, 2, ..., K; K represents the number of working nodes, or the number of load aggregators; k represents the order of the working nodes, or the kth load aggregator.
[0056] Optionally, when the global prediction model is initialized, that is, when federated learning starts (t=0, that is, the 0th round of training), the management node may randomly initialize the global prediction model and distribute the initialized global prediction model to each working node.
[0057] S202: Based on the global prediction model, initialize the local prediction model used for local power demand response prediction to obtain an initialized local prediction model.
[0058] That is, after receiving the global prediction model, each working node may use the model parameters of the global prediction model to update the model parameters of the local prediction model, and determine the local prediction model after the updated model parameters as the initialized local prediction model.
[0059] S203. Using the initialized local prediction model, determine the current predicted demand response value corresponding to the local sample data set related to the power demand response prediction, and based on the current predicted demand response value, determine the loss function corresponding to the initialized local prediction model; the loss function includes the error between the actual demand response value corresponding to the local sample data set and the current predicted demand response value, as well as the regularization term between the global prediction model and the local prediction model.
[0060] In an optional implementation, the loss function may be shown as the following formula (1).
[0061] (1)
[0062] In formula (1), represents the model parameters of the local prediction model corresponding to the kth working node (working node) after the tth round of training; w t It represents the model parameters of the global prediction model in the tth round of training; n k It represents the number of sample data corresponding to the working node k; xi It represents the i-th sample data in the local data; y i It represents the sample x i The label of the sample x i The actual demand response value; represents the output of function h (i.e., sample x i The predicted demand response value) and y i The mean square error loss function between ; It represents the regularization term, which is used to limit the deviation between the local prediction model and the global prediction model; μ represents the regularization coefficient, which is used to adjust the strength of the regularization term; It represents the loss function of the local prediction model corresponding to the kth working node after the tth round of training.
[0063] S204, with the goal of determining the minimum value of the loss function, iteratively train the initialized local prediction model based on the local sample data set to obtain a trained local prediction model.
[0064] S205: Send the trained local prediction model to the management node, so that the management node aggregates the trained local prediction models respectively sent by multiple working nodes to obtain a processed global prediction model.
[0065] In an optional implementation, the management node aggregates the trained local prediction models sent by multiple working nodes to obtain a processed global prediction model. The management node may aggregate the trained local prediction models sent by multiple working nodes to obtain a processed global prediction model using the following formula (2).
[0066] (2)
[0067] In formula (2), K represents the total number of working nodes (load aggregators); n k It represents the number of sample data corresponding to the kth working node; n represents the total number of sample data corresponding to K working nodes; It represents the trained local prediction model sent by the k-th working node (or the model parameters of the local prediction model corresponding to the k-th working node obtained after the t-th round of training); It represents the global prediction model obtained by aggregating the trained local prediction models sent by K working nodes, that is, the global prediction model obtained after the tth round of training.
[0068] In an embodiment of the present application, a working node may receive a global prediction model for global power demand response prediction sent from a management node in a distributed cluster; based on the global prediction model, a local prediction model for local power demand response prediction is initialized to obtain an initialized local prediction model; using the initialized local prediction model, a current predicted demand response value corresponding to a local sample data set related to power demand response prediction is determined, and based on the current predicted demand response value, a loss function corresponding to the initialized local prediction model is determined; the loss function includes the error between the actual demand response value corresponding to the local sample data set and the current predicted demand response value, as well as a regularization term between the global prediction model and the local prediction model; with the goal of determining the minimum value of the loss function, the initialized local prediction model is iteratively trained based on the local sample data set to obtain a trained local prediction model; the trained local prediction model is sent to the management node, so that the management node aggregates the trained local prediction models sent by multiple working nodes to obtain a processed global prediction model. By adopting this method, a regularization term is introduced on the basis of the federated learning algorithm, so that each local prediction model will not deviate too far from the global prediction model when it is updated, thereby avoiding the problem of overfitting or excessive deviation from the global prediction model due to data distribution differences. In this way, each working node (i.e., load aggregator) can use the global prediction model as a "benchmark" when optimizing the local prediction model, so that the update of the local prediction model is kept within a reasonable range, thereby reducing the impact of data distribution differences between different load aggregators on the global prediction model, so as to improve the prediction effect of the global prediction model, and then improve the prediction effect of each local prediction model, that is, improve the prediction effect of the power demand response prediction model.
[0069] In an optional embodiment, Figure 2 In the training method of the power demand response prediction model shown, each working node uses the initialized local prediction model to determine the current predicted demand response value corresponding to the local sample data set related to the power demand response prediction, which may include but is not limited to the following steps:
[0070] Step 1: Obtain a local sample data set related to power demand response prediction and an actual demand response value corresponding to the sample data set.
[0071] In some embodiments, the data in the local sample data set related to electricity demand response prediction may include but is not limited to humidity, temperature, weather, holiday signs, current real-time electricity price, the increment of the current real-time electricity price, the demand response value at the previous moment, etc.
[0072] In some embodiments, the actual demand response value corresponding to the sample data set can be determined by each working node in the following manner: based on the actual load corresponding to the demand response time period of all local electricity users and the load value that did not participate in the demand response in the historical time period, the actual demand response value corresponding to the local sample data set is determined.
[0073] Optionally, each working node determines the actual demand response value corresponding to the local sample data set based on the actual load corresponding to the demand response time period of all local electricity users and the load value that did not participate in the demand response in the historical time period, and the difference between the actual load corresponding to the demand response time period of all local electricity users and the load value that did not participate in the demand response in the historical time period is used as the actual demand response value corresponding to the local sample data set. The positive or negative value of the demand response value is related to the demand response. For example, if the demand response is to reduce the load (such as peak shaving), the demand response value is usually a negative value; if the demand response is to increase the load (such as valley filling), the demand response value may be a positive value.
[0074] Step 2: Encode each sample data in the sample data set to obtain the data features corresponding to each sample data.
[0075] In some embodiments, each working node may use a one-hot encoding algorithm to encode each sample data in the sample data set to obtain data features corresponding to each sample data.
[0076] Among them, one-hot encoding is also called one-bit effective encoding, which mainly uses an N-bit state register to encode N states. Each state has its own independent register bit, and only one bit is valid at any time.
[0077] Step 3: Normalize the multiple data features separately to obtain multiple normalized data features.
[0078] In some embodiments, each working node normalizes multiple data features respectively to obtain multiple normalized data features, which may include: determining the mean and standard deviation corresponding to each of the multiple data features; for each data feature, normalizing the targeted data feature based on the mean and standard deviation corresponding to the targeted data feature to obtain the normalized data feature.
[0079] Optionally, when each working node determines the mean and standard deviation corresponding to each data feature among multiple data features, the following formula (3) may be used.
[0080] (3)
[0081] In formula (3), represents the mean value corresponding to the i-th data feature; n represents the total number of characteristic components in the i-th data feature; x ij It represents the jth feature component of the i-th data feature; It represents the standard deviation corresponding to the i-th data feature.
[0082] Optionally, each working node normalizes the targeted data feature based on the mean and standard deviation corresponding to the targeted data feature, and when obtaining the normalized data feature, a z-score calculation formula can be used to normalize the targeted data feature to obtain the normalized data feature. The z-score calculation formula is shown in the following formula (4).
[0083] (4)
[0084] In formula (4), x i It represents the i-th data feature; It represents the mean value corresponding to the i-th data feature; It represents the standard deviation corresponding to the i-th data feature; z i Represents the normalized i-th data feature.
[0085] Step 4: Using the initialized local prediction model, based on multiple normalized data features, obtain the current predicted demand response value corresponding to the sample data set.
[0086] In some embodiments, each working node uses an initialized local prediction model to obtain a current predicted demand response value corresponding to the sample data set based on multiple normalized data features, which may include: adding timing information to the multiple normalized data features respectively to obtain multiple data features with added timing information; inputting the multiple data features with added timing information into the initialized local prediction model to obtain the current predicted demand response value corresponding to the sample data set.
[0087] Optionally, each working node adds timing information to a plurality of normalized data features respectively, and may add timing information to each normalized data feature respectively by means of positional encoding (PE).
[0088] Among them, the time series information added to each normalized data feature by position encoding is shown in the following formula (5).
[0089] (5)
[0090] In formula (5), pos represents the position, i represents the dimension index, and d represents the dimension of the feature matrix composed of multiple data features.
[0091] In some embodiments, the initialized local prediction model includes an encoder and a decoder; each working node inputs multiple data features with added timing information into the initialized local prediction model to obtain the current predicted demand response value corresponding to the sample data set, which may include: inputting multiple data features with added timing information into the encoder to obtain the encoding result; inputting the encoding result and the historical predicted demand response value into the decoder to obtain the current demand response value corresponding to the sample data set; the historical predicted demand response value is used to represent the demand response prediction result of the historical local prediction model for the local sample data set, and the historical local prediction model is the previous prediction model of the initialized local prediction model.
[0092] Optionally, the initialized local prediction model can be a transformer-based multi-head self-attention mechanism model. The advantage of this model is its powerful self-attention mechanism, which can capture long-range dependencies in time series data.
[0093] Optionally, each working node inputs multiple data features after adding timing information into the encoder to obtain the encoding result, which can be inputting multiple data features after adding timing information into a multi-head self-attention mechanism to obtain the encoding result.
[0094] In an optional embodiment, Figure 2 In the training method of the power demand response prediction model shown, each working node aims to determine the minimum value of the loss function, and iteratively trains the initialized local prediction model based on the local sample data set to obtain the trained local prediction model. The gradient descent method can be used to determine the minimum value of the loss function, and the initialized local prediction model is iteratively trained for E rounds based on the local sample data set to obtain the trained local prediction model.
[0095] Optionally, the model parameters of the trained local prediction model can be recorded as .
[0096] In some embodiments, It can be determined by the following formula (6).
[0097] (6)
[0098] In formula (6), It means to find the gradient of the model parameters of the local prediction model; It represents the loss gradient on local data and is used to optimize the performance of the model; It represents the regularization term, which is used to constrain the gap between the local prediction model parameters and the global prediction model and alleviate the problem of data distribution heterogeneity; It represents the learning rate, which is used to control the amplitude of each update. The specific physical meaning of each parameter can be found in the above explanation of the physical meaning of each parameter in formula (1), which will not be repeated here.
[0099] By adopting this implementation mode, each working node performs E rounds of iterative training on the initialized local prediction model based on the local sample data set by using the gradient descent method, thereby improving the prediction effect of the trained local prediction model.
[0100] In an optional embodiment, Figure 2 In the training method of the power demand response prediction model shown, after step S205, that is, the management node aggregates the trained local prediction models sent by multiple working nodes respectively, and after obtaining the processed global prediction model, it can also determine whether the processed global prediction model meets the convergence condition. If it is determined that the convergence condition is met, the training is terminated, and the processed global prediction model is sent to each working node, so that each working node updates the local prediction model based on the processed global prediction model, and uses the updated local prediction model to perform power demand response prediction; if the convergence condition is not met, the next round of training is entered, that is, the processed global prediction model is sent to each working node, so that each working node re-trains the local prediction model based on the processed global prediction model.
[0101] Optionally, the management node may determine that the processed global prediction model reaches a convergence condition when any of the following conditions are met: (1) the loss value corresponding to the processed global prediction model Less than the preset loss threshold; (2) The number of training rounds reaches T.
[0102] Among them, the loss value corresponding to the processed global prediction model is It can be determined using the following formula (7).
[0103] (7)
[0104] In formula (7), represents the processed global prediction model; K represents the total number of working nodes (load aggregators); n k It represents the number of sample data corresponding to the kth working node; n represents the total number of sample data corresponding to K working nodes; It represents the trained local prediction model sent by the k-th working node (or the model parameters of the local prediction model corresponding to the k-th working node obtained after the t-th round of training); It represents the loss value corresponding to the trained local prediction model sent by the kth working node.
[0105] By adopting this implementation mode, the management node can aggregate multiple trained local prediction models to obtain a processed global prediction model. After determining that the processed global prediction model has reached the convergence condition, the management node can stop training. In this way, a global prediction model with better prediction effect can be obtained, that is, the prediction effect of the global prediction model can be improved.
[0106] See also Figure 3 , Figure 3 Schematic diagram of the overall process of a training method for a power demand response prediction model provided in an embodiment of the present application. Figure 3 As shown, the training method of the power demand response prediction model may include but is not limited to the following steps:
[0107] S301. The central server sends a global prediction model for global power demand response prediction to a computer device corresponding to each load aggregator. Correspondingly, the computer device corresponding to each load aggregator receives the global prediction model.
[0108] S302. The computer device corresponding to each load aggregator initializes the local prediction model used for local power demand response prediction based on the global prediction model to obtain an initialized local prediction model.
[0109] In an optional implementation, the relevant description of steps S301 and S302 may refer to the description of the aforementioned steps S201 and S202, respectively, and will not be repeated here.
[0110] S303. The computer device corresponding to each load aggregator obtains a local sample data set related to power demand response prediction, and an actual demand response value corresponding to the sample data set.
[0111] S304. The computer device corresponding to each load aggregator encodes each sample data in the sample data set to obtain data features corresponding to each sample data.
[0112] S305. The computer device corresponding to each load aggregator normalizes the multiple data features respectively to obtain multiple normalized data features.
[0113] S306. The computer device corresponding to each load aggregator uses the initialized local prediction model to obtain the current predicted demand response value corresponding to the sample data set based on multiple normalized data features.
[0114] In an optional embodiment, the relevant explanation of steps S303 to S306 can refer to the aforementioned description of using the initialized local prediction model for each working node to determine the current predicted demand response value corresponding to the local sample data set related to the power demand response prediction, which will not be repeated here.
[0115] S307. The computer device corresponding to each load aggregator determines the loss function corresponding to the initialized local prediction model based on the current predicted demand response value; the loss function includes the error between the actual demand response value corresponding to the local sample data set and the current predicted demand response value, as well as the regularization term between the global prediction model and the local prediction model.
[0116] In an optional implementation, the loss function may be as shown in the above formula (1).
[0117] S308. The computer device corresponding to each load aggregator iteratively trains the initialized local prediction model based on the local sample data set with the goal of determining the minimum value of the loss function to obtain a trained local prediction model.
[0118] S309. The computer device corresponding to each load aggregator sends the trained local prediction model to the central server. Correspondingly, the central server receives the trained local prediction model sent by the computer device corresponding to each load aggregator.
[0119] S310. The central server aggregates the trained local prediction models sent by the computer devices corresponding to each load aggregator to obtain a processed global prediction model.
[0120] In an optional implementation, the central server aggregates the trained local prediction models sent by the computer device corresponding to each load aggregator to obtain the processed global prediction model, and the aforementioned formula (2) can be used.
[0121] S311. The central server determines whether the processed global prediction model has reached the convergence condition. If so, execute step S312; if not, execute step S313.
[0122] In an optional implementation, the central server may determine that the processed global prediction model reaches the convergence condition when any of the following conditions are met: (1) the loss value corresponding to the processed global prediction model Less than the preset loss threshold; (2) The number of training rounds reaches T.
[0123] S312. Send the processed global prediction model to the computer device corresponding to each load aggregator, so that the computer device corresponding to each load aggregator updates the local prediction model based on the processed global prediction model, and uses the updated local prediction model to perform power demand response prediction.
[0124] S313. Send the processed global prediction model to the computer device corresponding to each load aggregator, so that the computer device corresponding to each load aggregator retrains the local prediction model based on the processed global prediction model.
[0125] In an embodiment of the present application, a regularization term is introduced on the basis of the federated learning algorithm so that each local prediction model will not deviate too far from the global prediction model when updated, thereby avoiding the problem of overfitting or excessive deviation from the global prediction model due to differences in data distribution. In this way, each working node (i.e., load aggregator) can use the global prediction model as a "benchmark" when optimizing the local prediction model, so that the update of the local prediction model is kept within a reasonable range, thereby reducing the impact of data distribution differences between different load aggregators on the global prediction model, so as to improve the prediction effect of the global prediction model, and then improve the prediction effect of each local prediction model, that is, improve the prediction effect of the power demand response prediction model.
[0126] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0127] Based on the same inventive concept, the embodiment of the present application also provides a training device for a power demand response prediction model for implementing the training method for the power demand response prediction model involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the embodiments of the training device for one or more power demand response prediction models provided below can refer to the limitations of the training method for the power demand response prediction model above, and will not be repeated here.
[0128] See also Figure 4 , Figure 4 Schematic diagram of a training device for a power demand response prediction model provided in an embodiment of the present application. Figure 4 As shown, the training device of the power demand response prediction model may include but is not limited to:
[0129] A receiving module 401 is used to receive a global prediction model for global power demand response prediction sent from a management node in a distributed cluster;
[0130] The processing module 402 is used to initialize the local prediction model used for local power demand response prediction based on the global prediction model to obtain an initialized local prediction model;
[0131] A determination module 403 is used to determine the current predicted demand response value corresponding to the local sample data set related to the power demand response prediction using the initialized local prediction model, and determine the loss function corresponding to the initialized local prediction model based on the current predicted demand response value; the loss function includes the error between the actual demand response value corresponding to the local sample data set and the current predicted demand response value, and the regularization term between the global prediction model and the local prediction model;
[0132] A training module 404 is used to iteratively train the initialized local prediction model based on the local sample data set with the goal of determining the minimum value of the loss function to obtain a trained local prediction model;
[0133] The sending module 405 is used to send the trained local prediction model to the management node, so that the management node aggregates the trained local prediction models sent by multiple working nodes to obtain a processed global prediction model.
[0134] In one embodiment, when the determination module 403 is used to determine the current predicted demand response value corresponding to the local sample data set related to the power demand response prediction using the initialized local prediction model, it is specifically used to: obtain the local sample data set related to the power demand response prediction, and the actual demand response value corresponding to the sample data set; encode each sample data in the sample data set to obtain the data feature corresponding to each sample data; normalize the multiple data features respectively to obtain multiple normalized data features; and use the initialized local prediction model to obtain the current predicted demand response value corresponding to the sample data set based on the multiple normalized data features.
[0135] In one embodiment, when the determination module 403 is used to obtain the current predicted demand response value corresponding to the sample data set based on multiple normalized data features using the initialized local prediction model, it is specifically used to: add time series information to the multiple normalized data features respectively to obtain multiple data features with added time series information; input the multiple data features with added time series information into the initialized local prediction model to obtain the current predicted demand response value corresponding to the sample data set.
[0136] In one embodiment, the initialized local prediction model includes an encoder and a decoder; when the determination module 403 is used to input multiple data features after adding timing information into the initialized local prediction model to obtain the current predicted demand response value corresponding to the sample data set, it is specifically used to: input multiple data features after adding timing information into the encoder to obtain the encoding result; input the encoding result and the historical predicted demand response value into the decoder to obtain the current demand response value corresponding to the sample data set; the historical predicted demand response value is used to represent the demand response prediction result of the historical local prediction model for the local sample data set, and the historical local prediction model is the previous prediction model of the initialized local prediction model.
[0137] In one embodiment, when the determination module 403 is used to perform normalization processing on multiple data features respectively to obtain multiple normalized data features, it is specifically used to: determine the mean and standard deviation corresponding to each data feature in the multiple data features; for each data feature, based on the mean and standard deviation corresponding to the targeted data feature, normalize the targeted data feature to obtain the normalized data feature.
[0138] In one embodiment, the determination module 403 is further used to determine the actual demand response value corresponding to the local sample data set based on the actual load corresponding to the demand response time period of all local electricity users and the load value that did not participate in the demand response in the historical time period.
[0139] Each module in the training device of the power demand response prediction model can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the terminal device in the form of hardware, or can be stored in the memory in the terminal device in the form of software, so that the processor can call and execute the operations corresponding to each module.
[0140] In an exemplary embodiment, the present application provides a computer device, which may be a terminal device, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a training method for a power demand response prediction model is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.
[0141] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0142] In an exemplary embodiment, the present application provides a computer device including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps in the training method of the above-mentioned power demand response prediction model are implemented.
[0143] In an exemplary embodiment, the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in the training method of the above-mentioned power demand response prediction model are implemented.
[0144] In an exemplary embodiment, the present application provides a computer program product, including a computer program, which implements the steps in the above-mentioned training method of the power demand response prediction model when executed by a processor.
[0145] It should be noted that the data involved in this application (including but not limited to global prediction models, local prediction models, current predicted demand response values, actual demand response values, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0146] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0147] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0148] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A training method for a power demand response prediction model, characterized in that: Applied to each working node in a distributed cluster; the method comprises: Receiving a global prediction model for global power demand response prediction sent from a management node in the distributed cluster; Based on the global prediction model, initializing a local prediction model for local power demand response prediction to obtain an initialized local prediction model; Using the initialized local prediction model, determine the current predicted demand response value corresponding to the local sample data set related to the power demand response prediction, and based on the current predicted demand response value, determine the loss function corresponding to the initialized local prediction model; the loss function includes the error between the actual demand response value corresponding to the local sample data set and the current predicted demand response value, and the regularization term between the global prediction model and the local prediction model; With the goal of determining the minimum value of the loss function, iteratively training the initialized local prediction model based on the local sample data set to obtain a trained local prediction model; The trained local prediction model is sent to the management node, so that the management node aggregates the trained local prediction models respectively sent by the multiple working nodes to obtain a processed global prediction model.
2. The method according to claim 1, characterized in that The method of using the initialized local prediction model to determine a current predicted demand response value corresponding to a local sample data set related to power demand response prediction includes: Acquire a local sample data set related to power demand response prediction, and an actual demand response value corresponding to the sample data set; Encoding each sample data in the sample data set to obtain data features corresponding to each sample data; Normalizing the plurality of data features respectively to obtain a plurality of normalized data features; The initialized local prediction model is used to obtain a current predicted demand response value corresponding to the sample data set based on a plurality of normalized data features.
3. The method according to claim 2, characterized in that The using the initialized local prediction model to obtain the current predicted demand response value corresponding to the sample data set based on the plurality of normalized data features includes: Adding time series information to the plurality of normalized data features respectively to obtain a plurality of data features after the time series information is added; The plurality of data features to which the time series information is added are input into the initialized local prediction model to obtain a current predicted demand response value corresponding to the sample data set.
4. The method according to claim 3, characterized in that The initialized local prediction model includes an encoder and a decoder; the inputting of the plurality of data features after adding the timing information into the initialized local prediction model to obtain the current predicted demand response value corresponding to the sample data set includes: Inputting the plurality of data features after adding the timing information into the encoder to obtain an encoding result; The encoding result and the historical predicted demand response value are input into the decoder to obtain the current demand response value corresponding to the sample data set; the historical predicted demand response value is used to represent the demand response prediction result of the historical local prediction model for the local sample data set, and the historical local prediction model is the previous prediction model of the initialized local prediction model.
5. The method according to claim 2, characterized in that: The normalizing the plurality of data features respectively to obtain a plurality of normalized data features comprises: Determine a mean and a standard deviation corresponding to each of the data features; For each of the data features, based on the mean and the standard deviation corresponding to the data feature, normalization processing is performed on the data feature to obtain a normalized data feature.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Based on the actual loads of all local electricity users corresponding to the demand response time period and the load values that did not participate in the demand response in the historical time period, the actual demand response value corresponding to the local sample data set is determined.
7. A training device for a power demand response prediction model, characterized in that: Applied to each working node in a distributed cluster; the device comprises: A receiving module, configured to receive a global prediction model for global power demand response prediction sent from a management node in the distributed cluster; A processing module, configured to initialize a local prediction model for local power demand response prediction based on the global prediction model to obtain an initialized local prediction model; A determination module, configured to determine a current predicted demand response value corresponding to a local sample data set related to power demand response prediction using the initialized local prediction model, and determine a loss function corresponding to the initialized local prediction model based on the current predicted demand response value; the loss function includes an error between an actual demand response value corresponding to the local sample data set and the current predicted demand response value, and a regularization term between the global prediction model and the local prediction model; A training module, used for iteratively training the initialized local prediction model based on the local sample data set with the goal of determining the minimum value of the loss function to obtain a trained local prediction model; The sending module is used to send the trained local prediction model to the management node, so that the management node aggregates the trained local prediction models sent by multiple working nodes respectively to obtain a processed global prediction model.
8. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Electric quantity demand prediction method and system based on federal ensemble learning
CN113139341A
Personalized federal learning method, device and system based on global feature sharing
CN116777015A
Power system load prediction method and device, computer equipment and storage medium
CN118195353A
Node prediction method and system of graph structure data, storage medium and equipment
CN118708769A
Short-term photovoltaic power generation prediction method and system based on deep learning and federated learning
CN118798677A