Small sample power load prediction method and system based on transfer learning

Through the small sample power load prediction method based on transfer learning, the convolutional neural network and transfer learning technology are used to solve the problem of mismatch in the distribution of power load data under small sample events, and the adaptability and prediction performance of the model are improved.

CN120016429APending Publication Date: 2025-05-16NORTH CHINA BRANCH OF STATE GRID CORPORATION OF CHINA +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411662478.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Under small sample events, the distribution of power load data does not match the historical data, making it difficult for traditional machine learning methods to adapt and deep learning methods to be effectively deployed in high dynamic systems.

Method used

A small sample power load prediction method based on transfer learning is adopted. By obtaining historical data of the source domain and the target domain, the convolutional neural network is trained to obtain the source domain basic prediction model, and a fine-tuning model is used to construct a power load prediction model based on transfer learning.

Benefits of technology

Efficiently extract public knowledge from similar tasks and migrate it to target tasks, improving the model's performance, adaptability, and generalization capabilities on handling small sample power load prediction tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120016429A_ABST
    Figure CN120016429A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample power load prediction method and system based on transfer learning. The method comprises the following steps: acquiring historical data of a source domain and a target domain; training a convolutional neural network through historical data of a source domain to obtain a source domain basic prediction model; the source domain basic prediction model comprises a middle convolution layer and a full connection layer; keeping parameters of two middle convolutional layers and a first full connection layer of the source domain basic prediction model, performing random initialization on the parameters of other parts, and performing historical data training of a target domain to obtain a power load prediction model based on transfer learning; and performing small sample power load prediction through the power load prediction model based on transfer learning. According to the transfer learning method, public knowledge can be effectively extracted from similar tasks and transferred to the training process of the target task, so that the performance of the model on processing the target task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to power load forecasting, and in particular to a small sample power load forecasting method and system based on transfer learning. Background Art

[0002] The load forecasting of the power system has a large difference in the operation mode of society under small sample events and the same period in history, resulting in a distribution mismatch between the relevant data and the historical data, which increases the difficulty of data analysis and feature identification. Due to the difference in data distribution, the model based on traditional machine learning methods is difficult to adaptively predict the data during small sample events. Therefore, it is necessary to study the construction of relevant models that are adapted to data with abnormal distribution patterns. Due to many reasons such as the frequency of real events, sampling quality, and time scale, the current data of some high-dynamic system management and control business shows the characteristics of scarce samples, and traditional deep learning methods are difficult to effectively deploy for these highly variability and dynamic systems. Summary of the invention

[0003] Purpose of the invention: In view of the above shortcomings, the present invention provides a small sample power load forecasting method and system based on transfer learning to solve the problem of insufficient control data in high dynamic systems.

[0004] Technical solution: To solve the above problems, the present invention adopts a small sample power load forecasting method based on transfer learning, which includes the following steps:

[0005] Acquire historical data of a source domain and a target domain; the target domain is power load, and the source domain is economic factors related to power load;

[0006] A convolutional neural network is trained by using historical data of the source domain to obtain a source domain basic prediction model; the source domain basic prediction model includes an input layer, two intermediate convolutional layers, two pooling layers, two fully connected layers and an output layer;

[0007] The parameters of the two middle convolutional layers and the first fully connected layer of the source domain basic prediction model are retained, and the parameters of the rest of the source domain basic prediction model are randomly initialized. The source domain basic prediction model is trained with the historical data of the target domain to obtain the power load prediction model based on transfer learning.

[0008] Small sample power load forecasting is performed through a power load forecasting model based on transfer learning.

[0009] Furthermore, the economic factors related to power load include import amount, export amount, total import and export amount, real estate development investment, housing construction area of ​​real estate development enterprises, housing completion area of ​​real estate development enterprises, housing newly started construction area of ​​real estate development enterprises, general public budget revenue, power generation, various deposit balances of financial institutions in RMB and foreign currencies, various loan balances of financial institutions in RMB and foreign currencies, month information, and province information;

[0010] The historical data of the source domain is normalized using the Min-Max normalization method:

[0011]

[0012] Among them, x represents the original value of historical data, x * Represents the standard value of historical data after normalization, x max 、x min Indicates the maximum and minimum values ​​in the historical data of the corresponding economic factors.

[0013] Furthermore, during the training of the convolutional neural network, the parameters are learned using a loss function after forward propagation, and then updated by back propagation through gradient descent;

[0014] In the processing of the intermediate convolutional layer, the operation process of the jth convolutional layer is:

[0015]

[0016] in, represents the output of the i-th neuron in the j-th convolutional layer, represents the output of the kth neuron in the j-1st convolutional layer, n=i+s-1, s represents the convolution step size of the convolution kernel, Represents the weights and bias items corresponding to the network;

[0017] The operation process of the fully connected layer is:

[0018] y=f(W j-1 x j-1 +b j-1 )

[0019] Among them, x j-1 represents the input of the j-1th fully connected layer, W j-1 , b j-1 Represents the weights and bias terms of the j-1th fully connected layer.

[0020] Furthermore, the back propagation operation formula includes:

[0021]

[0022] Among them, y q (t) represents the expected output, represents the true value of the sample; E represents the loss error value; γ j represents the residual matrix of the ,th convolutional layer, and η represents the learning rate.

[0023] The present invention also adopts a small sample power load forecasting system based on transfer learning, comprising:

[0024] A data acquisition module, used to acquire historical data of a source domain and a target domain; the target domain is power load, and the source domain is economic factors related to power load;

[0025] A model building module, used to train a convolutional neural network through historical data of the source domain to obtain a source domain basic prediction model; the source domain basic prediction model includes an input layer, two intermediate convolutional layers, two pooling layers, two fully connected layers and an output layer;

[0026] The model fine-tuning module is used to retain the parameters of the two intermediate convolutional layers and the first fully connected layer of the source domain basic prediction model, randomly initialize the parameters of the remaining parts of the source domain basic prediction model, and obtain the power load prediction model based on transfer learning by training the source domain basic prediction model processed by the historical data of the target domain;

[0027] The prediction module is used to perform small sample power load prediction through a power load prediction model based on transfer learning.

[0028] The present invention also adopts a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0029] The present invention also adopts a computer-readable storage medium on which a computer program is stored, and the computer program implements the steps of the above method when executed by a processor.

[0030] Beneficial effect: Compared with the prior art, the significant advantage of the present invention is that the load forecasting model constructed for the situation where historical data of small sample events is insufficient and such events require rapid response is carried out based on economic power relations to perform medium-term load forecasting. The adopted transfer learning method can effectively extract common knowledge from similar tasks and transfer it to the training process of the target task, thereby improving the performance of the model in processing the target task. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 Schematic diagram of the structure and training process of CNN in the present invention.

[0032] Figure 2 Schematic diagram of the CNN model constructed in the present invention.

[0033] Figure 3 Schematic diagram of the process of transfer learning in the present invention. DETAILED DESCRIPTION

[0034] Example 1

[0035] In this embodiment, a small sample power load prediction method based on transfer learning includes the following steps:

[0036] Step 1: Obtain historical data of the source domain and the target domain; the target domain is the power load, and the source domain is the economic factors related to the power load.

[0037] In order to comprehensively consider the impact of various economic factors on electricity consumption, 13 monthly economic indicators are selected as model inputs in this embodiment, including import volume, export volume, total import and export volume, real estate development investment, housing construction area of ​​real estate development enterprises, completed housing area of ​​real estate development enterprises, newly started housing construction area of ​​real estate development enterprises, general public budget revenue, power generation, various deposit balances in RMB and foreign currencies of financial institutions, various loan balances in RMB and foreign currencies of financial institutions, month information, and province information.

[0038] In order to avoid the impact of different dimensions and magnitudes, the Min-Max normalization method is used to normalize the data set. The formula for Min-Max normalization is as follows:

[0039]

[0040] Among them, x is the original value of the data, x* is the standardized value, and x max 、x min They are the maximum and minimum values ​​in the data set consisting of the corresponding indicator data. The data values ​​after standardization transformation are in the range of [0,1].

[0041] Step 2: Train the convolutional neural network using historical data from the source domain to obtain the source domain basic prediction model.

[0042] Convolution is a linear operation that multiplies a set of weights by the input data. Similar to traditional artificial neural networks, convolutional neural networks (CNNs) are also deep learning networks with a similar structure to feedforward neural networks. However, unlike feedforward neural networks, the neuron arrangement structure of CNN has three dimensions: width, height, and depth. It can process multi-channel data simultaneously, and is therefore widely used in image processing, multi-dimensional data modeling and other fields.

[0043] like Figure 1As shown in the figure, the basic CNN is composed of multiple convolutional layers, pooling layers and fully connected layers. The model parameters are learned by using the loss function after forward propagation, and then updated by back propagation through gradient descent. The forward propagation calculation process of CNN is as follows:

[0044] For input data xi (i = 1, 2, ..., m), when passing through the i-th convolutional layer, the network operation process is as follows:

[0045]

[0046] in represents the output of the i-th neuron in the j-th convolutional layer, and Represents the output of the kth neuron in the j-1th convolutional layer. If the convolution step of the convolution kernel is s, then n = i + s-1. Since CNN uses sparse connections, the range of k is from i to n. They represent the weight and bias of the network respectively. f() represents the corresponding activation function. Common activation functions include sigmoid function and ReLU function.

[0047] After the convolution layer operation, the fully connected layer will flatten the data output by the convolution layer:

[0048] y=f(W j-1 x j-1 +b j-1 )

[0049] Among them, W j-1 , b j-1 They represent the weight and bias matrix of the network layer corresponding to the i-1th layer respectively.

[0050] After forward propagation, CNN calculates the loss function for backpropagation to update the parameters:

[0051]

[0052] Among them, t and q represent the current time and sample number respectively, and y q (t) are the expected output of the model and the true value of the sample, respectively. Afterwards, the residual between the predicted output and the true output of each layer of the network is calculated by the chain rule, and the network parameters are updated using the following formula:

[0053]

[0054] Among them, γ j is the residual matrix of the jth convolutional layer, and η is the learning rate.

[0055] The feature extractor of a convolutional neural network is generally composed of multiple stacked convolutional layers. Within the convolutional layer, each neuron is only connected to a part of the neurons in the adjacent network layer, that is, a sparse connection mode is adopted. At the same time, the neurons in each convolutional layer constitute multiple feature maps, and the related weight parameters constitute the convolution kernel. Neurons in the same feature map share the same convolution kernel. During the network training process, the convolution kernel parameters will be iterated multiple times according to the update algorithm described above until the model converges. Based on the above structure, CNN can effectively reduce model parameters and model complexity compared to traditional fully connected neural networks, avoid overfitting problems, and can also better extract data features and improve model learning efficiency.

[0056] A convolutional neural network (CNN) is used to construct a basic feature extractor to learn input features and the mapping relationship between input and output data. Convolution kernels of different sizes and numbers can be flexibly used to extract high-dimensional features in the data to improve the performance of the model. Since the amount of relevant data for small sample events is small, it is not suitable to build an overly complex network. Therefore, in this embodiment, a convolution layer composed of a 1×1 convolution kernel is used as a feature extractor. The 1×1 convolution kernel can reduce the number of output feature maps and reduce the parameters required for learning. At the same time, it is also beneficial to deepen the network depth with as few parameters as possible, and is used to realize information interaction between multi-dimensional data across channels, thereby increasing the nonlinear characteristics of the model. The CNN model finally constructed in this embodiment is as follows: Figure 2 shown.

[0057] The constructed CNN network has 7 layers. The input layer consists of a one-dimensional convolution layer with 8 convolution kernels. The two intermediate convolution layers are both composed of one-dimensional convolution layers with 4 convolution kernels. The pooling layer contains 32 neurons. The transition from the convolution layer to the fully connected layer is achieved by converting the multi-dimensional data into one dimension. Finally, there are 2 fully connected layers and an output layer, which contain 16, 8, and 1 neurons respectively. The convolution kernel size in each convolution layer is 1×1. In order to avoid the gradient vanishing problem and speed up the network convergence, the ReLU function is used as the activation function, which can be expressed as:

[0058] ReLU(x)=max{0,x}

[0059] It is worth noting that the CNN feature extraction architecture proposed in this embodiment can be flexibly adjusted or designed according to different situations to improve the model performance in a more targeted manner.

[0060] Step 3: Keep the parameters of the two middle convolutional layers and the first fully connected layer of the source domain basic prediction model, randomly initialize the parameters of the rest of the source domain basic prediction model, and train the source domain basic prediction model with historical data from the target domain to obtain the power load prediction model based on transfer learning.

[0061] The model training process is divided into two stages. The first stage is the pre-training stage, which aims to learn basic knowledge from the source domain data. First, the network is trained using the source domain data. The second stage is the fine-tuning stage, which aims to improve the performance of the model on the target domain task. Figure 3 As shown in the figure, at this stage, the parameters of the convolutional layer will be frozen to retain the sharable parameters that can be used for knowledge transfer. The parameters of the first fully connected layer will also be retained, but the subsequent fine-tuning process can still be performed. Then, the parameters of the rest of the network are randomly initialized for adaptive training on the target dataset. The network is fine-tuned using the target domain data to complete transfer learning, thereby improving the adaptability of the network to the target domain data. Since the amount of data in the target domain is small, freezing a part of the network parameters can effectively reduce the parameters that need to be trained, thereby reducing the complexity of the model. In this way, the deep neural network that originally required a large amount of data for training can also be converted into a small network with limited parameters that can be trained on small sample data, thereby effectively preventing the problem of overfitting. In the fine-tuning stage, since the parameters of the first few layers are frozen, the general characteristics between the input and output are retained. By reinitializing the remaining network parameters, the network is given a certain degree of freedom to learn new situations, so that the network can use limited target domain data for fine-tuning, thereby extracting new data features based on the learned features, achieving better generalization on the target task while maintaining common attributes, and reducing the side effects due to data distribution differences.

[0062] Step 4: Perform small sample power load forecasting through the power load forecasting model based on transfer learning.

[0063] The basic assumption of the transfer learning model is that some model parameters can be shared between the tasks to be learned in the source domain and the target domain. Since the economy and electricity consumption are closely related, even if there are large differences between economic development and electricity consumption in small sample events and normal periods, the nonlinear relationship between the two has not changed significantly. Therefore, the proposed transfer learning framework aims to learn the basic mapping relationship from economic data to electricity data from historical data and share the corresponding model parameters during small sample events. Specifically, the proposed transfer learning model task aims to use the economic data of a province in month t to predict its electricity load in month t+1. The target task can be expressed as:

[0064] EC i,t+1 =F(e 1,i,t ,e 2,i,t ,…,e n,i,t )

[0065] Among them, ECi,t+1 represents the electricity load demand of the ith province in the t+1th month, en,i,t represents the data of the nth economic indicator of the ith province in the tth month, and F() represents the electricity load forecasting model based on economic-electricity relationship transfer learning constructed in this embodiment.

[0066] Since the target task of the present invention is a small sample problem, the parameters of the convolutional layer are frozen during the model fine-tuning stage. The physical meaning is that the basic knowledge characterizing the economic-electricity relationship can remain unchanged during the fine-tuning process, so that the model can be further trained based on a small amount of data under a small sample event. At the same time, the parameters of the first fully connected layer are also retained, but training can still continue. The parameters of the remaining fully connected layers will be randomly initialized again, so that in the case of partially memorizing the parameters learned in the early stage, further personalized fine-tuning can be performed to better adapt to the target task.

[0067] Example 2

[0068] In this embodiment, a small sample power load forecasting system based on transfer learning includes:

[0069] A data acquisition module, used to acquire historical data of a source domain and a target domain; the target domain is power load, and the source domain is economic factors related to power load;

[0070] A model building module, used to train a convolutional neural network through historical data of the source domain to obtain a source domain basic prediction model; the source domain basic prediction model includes an input layer, two intermediate convolutional layers, two pooling layers, two fully connected layers and an output layer;

[0071] The model fine-tuning module is used to retain the parameters of the two intermediate convolutional layers and the first fully connected layer of the source domain basic prediction model, randomly initialize the parameters of the remaining parts of the source domain basic prediction model, and obtain the power load prediction model based on transfer learning by training the source domain basic prediction model processed by the historical data of the target domain;

[0072] The prediction module is used to perform small sample power load prediction through a power load prediction model based on transfer learning.

[0073] Example 3

[0074] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0075] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0076] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0077] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0078] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which all fall within the protection of the present invention.

Claims

1. A small sample power load forecasting method based on transfer learning, characterized in that: The following steps are involved: Acquire historical data of a source domain and a target domain; the target domain is power load, and the source domain is economic factors related to power load; A convolutional neural network is trained by using historical data of the source domain to obtain a source domain basic prediction model; the source domain basic prediction model includes an input layer, two intermediate convolutional layers, two pooling layers, two fully connected layers and an output layer; The parameters of the two middle convolutional layers and the first fully connected layer of the source domain basic prediction model are retained, and the parameters of the rest of the source domain basic prediction model are randomly initialized. The source domain basic prediction model is trained with the historical data of the target domain to obtain the power load prediction model based on transfer learning. Small sample power load forecasting is performed through a power load forecasting model based on transfer learning.

2. The small sample power load forecasting method based on transfer learning according to claim 1 is characterized in that: The economic factors related to power load include import amount, export amount, total import and export amount, real estate development investment, housing construction area of ​​real estate development enterprises, housing completion area of ​​real estate development enterprises, housing newly started construction area of ​​real estate development enterprises, general public budget revenue, power generation, various deposit balances of financial institutions in RMB and foreign currencies, various loan balances of financial institutions in RMB and foreign currencies, month information, and province information; The historical data of the source domain is normalized using the Min-Max normalization method: Among them, x represents the original value of historical data, x * Represents the standard value of historical data after normalization, x max 、x min Indicates the maximum and minimum values ​​in the historical data of the corresponding economic factors.

3. The small sample power load forecasting method based on transfer learning according to claim 2 is characterized in that: During the training of the convolutional neural network, the parameters are learned using a loss function after forward propagation, and then updated by back propagation through gradient descent; In the processing of the intermediate convolutional layer, the operation process of the jth convolutional layer is: in, represents the output of the i-th neuron in the j-th convolutional layer, represents the output of the kth neuron in the j-1th convolutional layer, n=i+s-1, s represents the convolution step size of the convolution kernel, Represents the weights and bias items corresponding to the network; The operation process of the fully connected layer is: y=f(w j-1 x j-1 +b j-1 ) Among them, x j-1 represents the input of the j-1th fully connected layer, W j-1 、b j-1 Represents the weights and bias terms of the j-1th fully connected layer.

4. The small sample power load forecasting method based on transfer learning according to claim 3 is characterized in that: The back propagation operation formula includes: Among them, y q (t) represents the expected output, represents the true value of the sample; E represents the loss error value; γ j represents the residual matrix of the jth convolutional layer, and η represents the learning rate.

5. The method for determining the output of power system units based on a double-layer model according to claim 3, characterized in that: The activation function used by the intermediate convolutional layer is the ReLU function.

6. A small sample power load forecasting system based on transfer learning, characterized in that: include: A data acquisition module, used to acquire historical data of a source domain and a target domain; the target domain is power load, and the source domain is economic factors related to power load; A model building module, used to train a convolutional neural network through historical data of the source domain to obtain a source domain basic prediction model; the source domain basic prediction model includes an input layer, two intermediate convolutional layers, two pooling layers, two fully connected layers and an output layer; The model fine-tuning module is used to retain the parameters of the two intermediate convolutional layers and the first fully connected layer of the source domain basic prediction model, randomly initialize the parameters of the remaining parts of the source domain basic prediction model, and obtain the power load prediction model based on transfer learning by training the source domain basic prediction model processed by the historical data of the target domain; The prediction module is used to perform small sample power load prediction through a power load prediction model based on transfer learning.

7. The small sample power load forecasting system based on transfer learning according to claim 6 is characterized in that: The economic factors related to power load include import amount, export amount, total import and export amount, real estate development investment, housing construction area of ​​real estate development enterprises, housing completion area of ​​real estate development enterprises, housing newly started construction area of ​​real estate development enterprises, general public budget revenue, power generation, various deposit balances of financial institutions in RMB and foreign currencies, various loan balances of financial institutions in RMB and foreign currencies, month information, and province information; The historical data of the source domain is normalized using the Min-Max normalization method: Among them, x represents the original value of historical data, x * Represents the standard value of historical data after normalization, x max 、x min Indicates the maximum and minimum values ​​in the historical data of the corresponding economic factors.

8. The small sample power load forecasting system based on transfer learning according to claim 7 is characterized in that: During the training process of the convolutional neural network, the parameters are learned using the loss function after forward propagation, and then updated by back propagation through gradient descent.

9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Cited By

  • Large-scale transformer area short-term load prediction method

    CN121390474A

  • Large-scale transformer area short-term load forecasting method

    CN121390474B