A power load forecasting method for the power Internet of Things

By decomposing the power consumption data and building a lightweight Transformer model, the existing deep learning model's problem of high computing resource consumption in power load prediction is solved, and efficient power load prediction on edge servers is achieved, which reduces cloud pressure and improves data processing efficiency.

CN117200190BActive Publication Date: 2025-06-13STATE GRID JIANGSU ELECTRIC POWER CO LTD NANJING POWER SUPPLY COMPANY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311040562.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-17
Publication Date
2025-06-13
Estimated Expiration
2043-08-17

AI Technical Summary

Technical Problem

The existing deep learning models have problems such as large computing resource consumption, large model scale, and difficulty in adapting to large-scale data processing in power load prediction, especially in environments with limited edge servers.

Method used

A lightweight power load prediction method is proposed. By decomposing trends, seasons and residuals of historical electricity consumption data, a lightweight model based on Transformer's codec structure is constructed, and the calculation complexity is reduced using DP-Attention and FDP-Attention modules are used to reduce the calculation complexity and deploy it to an edge server for prediction.

Benefits of technology

It realizes efficient power load prediction on edge servers, reduces cloud computing pressure, improves data processing and storage efficiency, and can scientifically formulate power distribution and distribution plans to ensure normal operation and life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117200190B_ABST
    Figure CN117200190B_ABST
Patent Text Reader

Abstract

A power load forecasting method for the Internet of Things in power systems. The Internet of Things in power systems has sub-servers in different areas for edge computing. A lightweight power load forecasting model is constructed, and the lightweight power load forecasting model is suitable for the computing resources of edge servers in edge computing, so as to complete the power load forecasting of the Internet of Things in power systems in edge computing. The present invention constructs a deep learning forecasting model according to historical power consumption data and its characteristics to achieve accurate load forecasting, and at the same time reduces the model scale as much as possible to lightweight the model, making it convenient to be deployed on the edge servers set by power enterprises in each area. Taking each area as a unit, the future power consumption of the area is predicted, avoiding the inconvenience caused by uploading all the power consumption data of each area to the cloud, and reducing the computing pressure on the cloud.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of smart grids, relates to the prediction of power grid power load by deep learning, and is a method for predicting power load of the Internet of Things in power systems. Background Art

[0002] With the development of technology, more and more requirements are put forward for the power system in the development of all walks of life. The implementation of smart grid construction has become the focus of the development of power grid technology. In order to promote the digital transformation of the power grid and integrate new technologies with the power system to build a new type of power system has become an inevitable trend in the development of the power industry, and then the concept of the Internet of Things in power systems has emerged. The Internet of Things in power systems refers to the full application of advanced information technologies such as 5G communication and artificial intelligence around all links of the power system to achieve interconnection and data interaction among all links of the power system, with characteristics such as comprehensive state perception, efficient information processing, and flexible and convenient application. It mainly includes four-layer structures: the perception layer, the network layer, the platform layer, and the application layer, which solve the problems of power data acquisition, transmission, management, and value creation respectively.

[0003] Power resource allocation and scheduling, as an important task of the power dispatching department, is an important link in the power system. Accurate and effective power load prediction is the basis for scientific dispatching of power backup energy and an important guarantee for the safe, stable, and economic operation of the power system. It can help relevant departments scientifically formulate power generation and distribution plans and ensure the normal operation of all walks of life and the normal life of residents. Therefore, it is necessary to achieve accurate power load prediction.

[0004] The problem of power load prediction can be classified as a time series prediction problem. In recent years, with the popularity of deep learning technology, various time series prediction models have emerged in an endless stream. The more popular ones include RNN, LSTM, and Transfomer, etc. RNN and LSTM deep neural networks have problems such as limited input data length, difficult training when the model is deeper, and difficulty in effectively analyzing the long-term dependence of sequences. They are suitable for situations with less data. However, in the 5G era, the amount of data is often relatively large, and they are difficult to handle large-scale prediction tasks when the data volume is large. Transfomer has well solved the problem of long-term dependence of sequences and has been widely used in long-time series prediction. However, Transfomer also has certain limitations in solving time series problems. Its prediction accuracy is greatly related to the input sequence. Not any time series directly fed into Transfomer can obtain satisfactory results. It is often necessary to dig deeper information based on the characteristics of the input sequence. In addition, the model scale based on Transfomer is often relatively large, and its computational complexity and required computing resources are also relatively high, which is proportional to the length of the input sequence.

[0005] Power enterprises often set up central servers (also known as cloud servers) within the enterprise and sub-servers (also known as edge servers) in each area. The cloud server is responsible for big data analysis of long-cycle data and can operate in fields such as periodic maintenance and business decision-making. The edge server focuses on the analysis of real-time and short-cycle data to better support the timely processing and execution of local services. Placing more data analysis tasks on the edge server can significantly save data transmission resources, reduce the computing pressure on the cloud, and at the same time, the storage and processing of data at the edge are more efficient and secure. However, most deep learning data analysis algorithms are often deployed in the cloud, uploading the data from each area to the cloud for processing, while the computing resources of the edge server are often very limited, and its equipment is often insufficient to run large-scale deep learning models.

[0006] In view of this, providing a lightweight power load forecasting method for the characteristics of power data has become an urgent problem to be solved in this field. Summary of the Invention

[0007] The purpose of the present invention is to propose a power load forecasting method for the Internet of Things in the power system in view of the demand for power load forecasting and the computing resource limitations of edge computing, and construct a lightweight model to facilitate its deployment on the edge servers set up by power enterprises in each area, reduce the computing pressure on the cloud, and improve the efficiency of data processing and storage.

[0008] The technical solution of the present invention is: a power load forecasting method for the Internet of Things in the power system. The Internet of Things in the power system has sub-servers in different areas for edge computing, and a lightweight power load forecasting model is constructed. The lightweight power load forecasting model is suitable for the computing resources of the edge server in edge computing, and realizes the power load forecasting of the Internet of Things in the power system in edge computing, including the following steps:

[0009] S1 Collect historical power consumption data and perform preprocessing, including outlier correction and missing value filling;

[0010] S2 Decompose the time series of historical power consumption data. The time series of historical power consumption data obtained in S1 is decomposed into three parts: a trend component, a seasonal component, and a residual component. The trend component reflects the overall trend of power consumption in the long term, the seasonal component reflects the periodic situation of power consumption in the long term, and the residual component reflects the deviation between the true value and the sum of the trend component and the seasonal component;

[0011] S3 Construct a lightweight power load forecasting model and train it.

[0012] The lightweight power load prediction model is based on the encoder-decoder structure of Transformer. Three branches are configured for prediction respectively for the three components decomposed by S2. Among them, the DP-Attention module is used to implement the attention mechanism for the prediction of the trend component and the residual component, and the FDP-Attention module is used to implement the attention mechanism for the seasonal component. The three components are respectively input into the model for prediction, and the prediction results output by the three components at the decoder end are added and reconstructed to obtain the final prediction result; a prediction model is trained and constructed with historical electricity consumption data;

[0013] S4 Model deployment: Deploy the trained lightweight power load prediction model to the sub-servers in each area, and complete the prediction of future power loads in units of areas. Input the electricity consumption data before the date to be measured into the prediction model to complete the prediction of power loads for subsequent scheduling planning.

[0014] The present invention provides a power load prediction method for the power Internet of Things. According to historical electricity consumption data and its characteristics, a deep learning prediction model is constructed to achieve accurate load prediction. At the same time, the model scale is reduced as much as possible, and the model is lightweighted, making it convenient to be deployed to the edge servers set by power enterprises in each area. The future electricity consumption of each area is predicted in units of areas, avoiding the inconvenience caused by uploading all the electricity consumption data of each area to the cloud and reducing the computing pressure on the cloud. The present invention can ultimately help relevant departments scientifically formulate power generation and distribution plans, and ensure the normal operation of all walks of life and the normal life of residents.

[0015] The present invention has the following beneficial effects:

[0016] (1) Aiming at the limitation that the traditional Transfomer model cannot accurately predict any time series, the present invention fully explores the characteristics of the historical electricity consumption data time series. According to its prominent seasonal characteristics, the time series is decomposed into trend, season, and residual components, and the prediction results are reconstructed after predicting the decomposed components respectively, which can well complete the prediction task for power data.

[0017] (2) Aiming at the obvious periodic characteristics of the decomposed seasonal component, converting the calculation of attention correlation in the time domain to the calculation in the frequency domain can make the prediction result of the seasonal component have better performance.

[0018] (3) Aiming at the disadvantages of the traditional Transformer model, such as high computational complexity, slow inference process, and large model scale, which lead to a large amount of computational resources being occupied, the DP-Attention module is designed to reduce the complexity of its main computational complexity source and accelerate the model's computational inference. At the same time, for each component after the original time series is deconstructed, its complexity is greatly reduced, and the model is easy to learn. Therefore, the stacking layers of the encoder and decoder can be reduced, and the model scale can be compressed. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is the flowchart of the power load prediction of the present invention.

[0020] Figure 2 This is the structural diagram of the prediction model constructed by the present invention.

[0021] Figure 3 This is the schematic diagram of the DP-Attention module of the present invention.

[0022] Figure 4 This is the schematic diagram of the FDP-Attention module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0024] As Figure 1 shown, this embodiment provides a power load prediction method. First, historical power consumption data is collected and the original data is preprocessed for subsequent analysis. Then, the time series is deconstructed into a trend component, a seasonal component, and a residual component, which are respectively input into the prediction model for prediction. Finally, the prediction results of each component are reconstructed to obtain the final prediction result. According to the prediction result, it helps the relevant power departments to conduct scheduling and planning of power resources.

[0025] This example specifically includes the following steps:

[0026] Step S1: Collect historical power consumption data and correct and fill in the abnormal and missing data.

[0027] According to the target prediction range, the power consumption data of a target area for a period of time is collected in units of days, months or quarters. However, due to the fact that equipment failures may occur on some days in history, resulting in gaps in the collected historical power consumption data, or emergencies may cause the power consumption to rise sharply, resulting in the daily power consumption being far from the normal range. If it is directly fed into the prediction model, it will cause serious interference to the training. Generally, it is hoped that the trained model has universality. Therefore, the input data needs to be cleaned as follows.

[0028] Correct the outliers in the original sequence data, specifically: identify the outliers in the data of a certain day through the 3σ theory, and then for the outlier points, assign and correct their values according to the data at the two adjacent moments on the same day and the data at the same moment on the two adjacent days.

[0029]

[0030]

[0031] Where x n,i is the value at the i-th moment on the n-th day, N is the judgment day interval, generally taking 7 or 14 days. σ i 2 are the mean and variance of the values at the i-th moment within the judgment day interval respectively. If the data x n,i at a certain moment on a certain day satisfies the following inequality:

[0032]

[0033] Then this value is determined as an outlier and corrected as follows:

[0034]

[0035] Where ξ is the scale threshold, usually taking 0.9 - 1.6, x n,i ′ is the corrected value at the i-th moment on the n-th day, x n+1,i and x n-1,i are the values at the i-th moment on the two adjacent days to the n-th day respectively, x n,i-1 and x n,i+1 are the values at the two adjacent moments before and after the i-th moment on the n-th day respectively.

[0036] Fill in the missing values in the original sequence data, specifically: replace each missing value with the mean of all non-missing parts of the dataset to be interpolated, that is, if the value at the i-th moment on the n-th day is missing, it is filled with the mean of the values at the i-th moment on other days within the day interval.

[0037] Specifically:

[0038]

[0039] Where, where x n,i is the value at the i-th moment on the n-th day.

[0040] Step S2 decomposes the time series of historical electricity consumption data.

[0041] In view of the unique and extremely prominent seasonal characteristics of historical electricity load - more electricity consumption in summer and winter, and relatively less electricity consumption in spring and autumn, the original time series of electricity consumption data is decomposed into three components: a trend component, a seasonal component, and a residual component. The trend component reflects the overall trend of electricity consumption over a long period of time, the seasonal component reflects the periodic pattern of electricity consumption over a long period of time, and the residual component reflects the deviation between the true value and the sum of the trend component and the seasonal component.

[0042] The present invention uses the STL algorithm, a time series decomposition method with robust locally weighted regression as the smoothing method, and decomposes the original time series based on locally weighted regression LOESS:

[0043] X original = X t + X s + X r

[0044] where, X t is the trend component, X s is the seasonal component, and X r is the residual component.

[0045] After decomposition, normalization processing is performed on each component. Taking the trend component as an example:

[0046]

[0047] where X t is the value of the original trend component, min(X t ) is the minimum value of the trend component, max(X t ) is the maximum value of the trend component, and X t ′ is the value after normalization.

[0048] Step 3: Construct a lightweight power load prediction model and train it.

[0049] Figure 2 is the prediction model constructed in the present invention. As shown in the figure, the lightweight power load prediction model is based on the encoder - decoder structure of Transformer, and three branches are respectively configured for prediction for the three components decomposed from the original time series in S2, and the sum of the prediction results of the three components is used as the final prediction result.

[0050] Specifically, for the prediction of the trend component and the residual component, a DP-Attention (Dropout Attention) module is designed as the attention mechanism module to reduce the complexity of calculating the correlation of the attention mechanism. Different from the Attention mechanism of the traditional Transformer, the DP-Attention module deactivates the input sequence vectors according to a certain proportion. Here, the deactivation is different from that of the fully connected network. It does not completely stop participating in the operation, but does not participate in the calculation of the attention mechanism. The vectors that are not deactivated still calculate the correlation through the attention mechanism. The deactivated vectors no longer calculate the correlation with other vectors, but take the mean of all input vectors to replace the correlation with other vectors. The DP-Attention module effectively reduces the computational complexity and improves the inference speed of the model.

[0051] The calculation details of DP-Attention are as follows:

[0052] As Figure 3 shown, first, the Q, K, and V of the input sequence are calculated, and the calculation formulas are as follows:

[0053] Q = W q I

[0054] K = W k I

[0055] V = W v I

[0056] Among them, I is the vector matrix after the input sequence is merged, and W q , W k , W v are the Q, K, and V parameter matrices learned by the model respectively.

[0057] Subsequently, the input sequence vectors are deactivated according to a certain proportion. The original Q, K, and V are divided into the undeactivated Q', K', V' and the deactivated Q”, K”, V”. The undeactivated Q', K', V' still perform the correlation calculation of the attention mechanism as shown in the following formula:

[0058]

[0059] Among them, d k is the dimension of the Q and K parameter matrices.

[0060] The deactivated Q”, K”, V” no longer perform the correlation calculation of the attention mechanism among themselves, but directly use the global average of all V to replace:

[0061]

[0062] Among them, M is the length of the input sequence.

[0063] For the significant periodic trend presented by the seasonal component, its correlation characteristics in the frequency domain are more prominent. Therefore, the FDP-Attention (FFT Dropout Attention) module is designed as the attention mechanism module. The input vector is Fourier-transformed to the frequency domain, and at the same time, the input vector is still subjected to random ratio inactivation to calculate the attention mechanism correlation, and then the inverse Fourier transform is performed to obtain the final correlation result.

[0064] The calculation details of FDP-Attention are roughly the same as those of DP-Attention. As Figure 4 shown, a layer of Fourier transform and inverse Fourier transform are added before and after calculating the attention mechanism correlation. For the non-inactivated Q', K', V', the following formula is used to calculate the attention mechanism correlation:

[0065]

[0066] where F represents the Fourier transform, and F -1 represents the inverse Fourier transform, and d k is the dimension of the Q, K parameter matrices.

[0067] For the inactivated Q”, K”, V”, the calculation method is the same as before.

[0068] In the encoder-decoder structure of the Transformer of the lightweight power load prediction model, for the DP-Attention module and the FDP-Attention module, V' and V” are combined together as the correlation representation of the input sequence, and then through a series of residual connections, normalizations, and feed-forward layers, the output of the encoder is formed. The decoder also uses the same Attention module as the encoder, inputs the output of the encoder and the input of the decoder into the decoder together, and through a series of residual connections, normalizations, and feed-forward layers, and finally through a linear layer to obtain the prediction output of each component.

[0069] Under the method of the present invention, the complexity of each component after the decomposition of the original time series is greatly reduced, and the model is easy to learn. Therefore, the number of stacked layers of the encoder-decoder of the prediction model of the present invention is two layers.

[0070] In the training process of the lightweight power load prediction model of the present invention, the MSE loss function is adopted. MSE measures the mean squared error between the model prediction value and the actual value. The smaller the MSE, the closer the prediction value is to the true value. The formula is as follows:

[0071]

[0072] where xl is the true value, and y l is the predicted value, and L is the length of the prediction sequence.

[0073] The final prediction result can be obtained by reconstructing the prediction results output by all components at the decoder end, that is:

[0074] Y prediction = Y t + Y s + Y r

[0075] where Y t , Y s , Y r are the predicted values of X t , X s , X r respectively.

[0076] Step S4: Model deployment. Input the electricity consumption data before the date to be measured into the model to complete the prediction of the power load for subsequent scheduling planning.

[0077] Specifically, deploy the lightweight power load prediction model to the sub-servers in each area, and complete the prediction of the future power load in units of areas. While achieving more efficient data analysis of local data, it avoids the inconvenience of uploading all a large amount of data to the cloud and alleviates the computing pressure on the cloud.

Claims

1. A power load forecasting method for the power Internet of Things, characterized in that the power Internet of Things is equipped with sub-servers in different areas for edge computing, and a lightweight power load forecasting model is constructed. The lightweight power load forecasting model is suitable for the computing resources of edge servers in edge computing, and realizes the power load forecasting of the power Internet of Things in edge computing, including the following steps: S1 Collect historical power consumption data and perform preprocessing, including outlier correction and missing value filling; S2 Decompose the time series of historical power consumption data. The time series of historical power consumption data obtained in S1 is decomposed into three parts: a trend component, a seasonal component, and a residual component. The trend component reflects the overall trend of power consumption in the long term, the seasonal component reflects the periodic situation of power consumption in the long term, and the residual component reflects the deviation between the true value and the sum of the trend component and the seasonal component; S3 Construct a lightweight power load forecasting model and train it, The lightweight power load forecasting model is based on the encoder-decoder structure of Transformer. Three branches are configured for the three components decomposed in S2 for prediction. Among them, the DP-Attention module is used to implement the attention mechanism for the prediction of the trend component and the residual component, and the FDP-Attention module is used to implement the attention mechanism for the seasonal component. The three components are respectively input into the model for prediction, and the prediction results output by the three components at the decoder end are added and reconstructed to obtain the final prediction result; the constructed prediction model is trained with historical power consumption data; In the lightweight power load forecasting model, the DP-Attention module is specifically as follows: First, calculate Q, K, and V for the input sequence, and the calculation formula is as follows: Q = W q I K = W k I V = W v I Among them, I is the vector matrix after the input sequences are merged, and W q , W k , W v are the Q, K, and V parameter matrices learned by the model respectively; subsequently, the input sequence vectors are inactivated proportionally, and the original Q, K, and V are divided into the non-inactivated Q', K', V' and the inactivated Q", K", V"; the non-inactivated Q', K', V' still perform the attention mechanism correlation calculation shown in the following formula: where d k is the dimension of the Q and K parameter matrices; the inactivated Q'', K'', and V'' no longer calculate the attention mechanism correlation among themselves, but directly use the global average of all Vs instead: where M is the length of the input sequence; The FDP-Attention module is specifically as follows: Fourier transform the input vector to the frequency domain, and at the same time still perform random proportion inactivation calculation of the input vector to calculate the attention mechanism correlation, and then perform inverse Fourier transform to obtain the final correlation result. Compared with the DP-Attention module, a layer of Fourier transform and inverse Fourier transform is added before and after calculating the attention mechanism correlation. For the non-inactivated Q', K', V', the following formula is used to calculate the attention mechanism correlation: where F represents the Fourier transform, and F -1 represents the inverse Fourier transform, and d k is the dimension of the Q, K parameter matrices; for the inactivated Q", K", V", it is the same as that of the DP-Attention module; For the DP-Attention module and the FDP-Attention module, combine V' and V” together as the correlation representation of the input sequence, and then pass through the residual connection, normalization, and feed-forward layer to form the output of the Transformer encoder; The Transformer decoder also uses the same attention modules as the encoder, that is, the DP-Attention module and the FDP-Attention module. The output of the encoder and the input of the decoder are jointly input into the decoder, and after passing through the residual connection, normalization, and feed-forward layer, finally, the prediction output of each component is obtained through a linear layer; S4 model deployment: Deploy the trained lightweight power load prediction model to the sub-servers in each area, and complete the prediction of future power loads in units of areas. Input the power consumption data before the date to be measured into the prediction model to complete the prediction of power loads for subsequent scheduling planning.

2. A power load prediction method for the power Internet of Things according to claim 1, characterized in that the preprocessing in S1 is specifically as follows: Correct the outliers in the original sequence data: Identify the outliers in the data of a certain day through the 3σ theory, and then for the outlier points, assign values and correct them according to the data at the two adjacent moments on the same day and the data at the same moment on the two adjacent days: where x n,i is the value at the i-th moment on the n-th day, N is the judgment day interval, σ i 2 are the mean and variance of the values at the i-th moment within the judgment day interval respectively. If the data x n,i satisfies the following inequality: then this value is determined as an outlier and corrected as follows: Among them, ξ is the scale threshold, taking values from 0.9 to 1.6, and x n,i ′ is the corrected value at the i-th moment on the n-th day, and x n+1,i and x n-1,i are the values at the i-th moment of the two days before and after the n-th day respectively, and x n,i-1 and x n,i+1 are the values at the two moments before and after the i-th moment on the n-th day respectively; The specific method for filling in the missing values in the original sequence data is: Replace each missing value with the mean of all non-missing parts of the dataset to be interpolated. That is, if the value at the i-th moment on the n-th day is missing, it is filled with the mean of the i-th moment on other days within the day range:

3. A power load prediction method for the power Internet of Things according to claim 1, characterized in that in S2 the specific method for decomposing the three components is: Use the STL algorithm to decompose the original time series based on the locally weighted regression LOESS: X original = X t + X s + X r Among them, X t is the trend component, X s is the seasonal component, X r is the residual component. After decomposition, each component is normalized.

Citation Information

Patent Citations

  • LS-SVM intelligent transformer area load prediction method and system based on PSO optimization

    CN112381315A

  • Transformer-based power load prediction method

    CN113592185A