An industrial electricity consumption prediction method and terminal

By screening and clustering the variable sets in the power consumption prediction and optimizing the variable sets using error-driven methods, the problem of difficult to balance prediction accuracy and computing efficiency in the prior art is solved, and more efficient and accurate power consumption prediction is achieved.

CN115310658BActive Publication Date: 2025-06-10STATE GRID FUJIAN POWER ELECTRIC CO ECONOMIC RESEARCH INSTITUTE +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210713378.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-22
Publication Date
2025-06-10
Estimated Expiration
2042-06-22

AI Technical Summary

Technical Problem

In the prior art, it is difficult to reduce the prediction operation amount and calculation time while improving the prediction accuracy.

Method used

By obtaining industry electricity consumption data samples, determining the set of variables to be selected, and filtering them based on the maximum correlation and minimum redundancy of each factor to be selected and the target variable, the initial set of variables is obtained. Then, a neural network prediction model is obtained through clustering, and a second screening set of initial screen variables is performed using error drivers to obtain the optimal variable set to output the power prediction value.

Benefits of technology

The accuracy of power consumption prediction is improved, the predicted calculation amount and calculation time are reduced, and by combining data driving and error driving methods, we can learn from each other's strengths and make up for our weaknesses, and improve computing performance and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115310658B_ABST
    Figure CN115310658B_ABST
Patent Text Reader

Abstract

The present invention discloses an industry electricity consumption prediction method and a terminal, which obtain industry electricity consumption data samples and determine a set of candidate variables. According to the maximum correlation and minimum redundancy between each candidate factor in the set of candidate variables and the target variable, a first screening of the set of candidate variables is performed to obtain a pre-screened variable set. Clustering is performed based on the industry electricity consumption data samples and the pre-screened variable set to obtain at least two neural network prediction models. In the neural network prediction model, error-driven is used to perform a second screening of the pre-screened variable set to obtain an optimal variable set, so as to obtain an electricity consumption prediction value based on the optimal variable set, that is, calculate the value of the target variable. Therefore, data-driven is first used for the preliminary screening of the variable set, and then error-driven is used for the refined selection of the variable set, combining data-driven and error-driven to make up for each other's advantages and disadvantages to improve the calculation performance and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power data prediction, and particularly relates to a method and a terminal for predicting the electricity consumption of industries. Background Art

[0002] The prediction of power demand is an important basis for signing medium- and long-term contracts, formulating power generation plans, and maintaining the balance of power and electricity in the system. Scientific prediction methods and accurate prediction results are of great significance to the economic operation of the power system. The total regional electricity is composed of the electricity consumption of each industry, and the electricity consumption of each industry is generated by the electricity consumption behaviors of individual enterprises within the industry. Due to the differences in production technologies among different enterprises, their electricity consumption characteristics are different, and different influencing factors of electricity consumption also lead to different change laws of the electricity consumption of each industry.

[0003] Currently, there are mainly two mainstream selection methods: data-driven and error-driven, and each has its own advantages. Among them, the data-driven type only measures the correlation between influencing factors and variables to be analyzed. Its calculation process can be independent of the prediction method and has a relatively fast calculation speed; the error-driven type method uses the magnitude of the prediction error as a measure for identifying the main influencing factors, generally cannot be independent of the prediction process, and requires multiple trainings of the prediction model to achieve a high prediction accuracy, but has a high calculation cost. Therefore, there is an urgent need for a more refined electricity consumption prediction model. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: to provide a method and a terminal for predicting the electricity consumption of industries, which can improve the prediction accuracy while reducing the prediction operation amount and operation time.

[0005] To solve the above technical problem, the technical solution adopted by the present invention is:

[0006] A method for predicting the electricity consumption of industries, comprising the steps of:

[0007] Obtaining a data sample of the electricity consumption of an industry and determining a set of candidate variables;

[0008] Performing a first screening on the set of candidate variables according to the maximum correlation degree and minimum redundancy degree between each candidate factor in the set of candidate variables and the target variable, to obtain a set of preliminarily screened variables;

[0009] Performing clustering according to the data sample of the electricity consumption of the industry and the set of preliminarily screened variables to obtain at least two neural network prediction models;

[0010] In the neural network prediction model, using error-driven to perform a second screening on the set of preliminarily screened variables to obtain an optimal set of variables, and outputting an electricity consumption prediction value based on the optimal set of variables.

[0011] To solve the above technical problem, another technical solution adopted by the present invention is:

[0012] An industry electricity consumption prediction terminal includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0013] Obtain industry electricity consumption data samples and determine a set of candidate variables;

[0014] According to the maximum correlation and minimum redundancy between each candidate factor in the set of candidate variables and the target variable, perform a first screening of the set of candidate variables to obtain a preliminary screening variable set;

[0015] Perform clustering based on the industry electricity consumption data samples and the preliminary screening variable set to obtain at least two neural network prediction models;

[0016] In the neural network prediction model, use error-driven to perform a second screening of the preliminary screening variable set to obtain an optimal variable set, and output an electricity consumption prediction value based on the optimal variable set.

[0017] The beneficial effects of the present invention are as follows: Obtain industry electricity consumption data samples and determine a set of candidate variables. According to the maximum correlation and minimum redundancy between each candidate factor in the set of candidate variables and the target variable, perform a first screening of the set of candidate variables to obtain a preliminary screening variable set; perform clustering based on the industry electricity consumption data samples and the preliminary screening variable set to obtain at least two neural network prediction models; in the neural network prediction model, use error-driven to perform a second screening of the preliminary screening variable set to obtain an optimal variable set, so as to obtain an electricity consumption prediction value based on the optimal variable set, that is, calculate the value of the target variable. Therefore, first use data-driven to perform a preliminary screening of the variable set, and then use error-driven to perform a refined selection of the variable set, combining data-driven and error-driven to make up for each other's strengths and weaknesses to improve calculation performance and accuracy. Description of the Drawings

[0018] Figure 1 It is a flowchart of a method for predicting industry electricity consumption according to an embodiment of the present invention;

[0019] Figure 2 It is a schematic diagram of an industry electricity consumption prediction terminal according to an embodiment of the present invention;

[0020] Figure 3 It is a flowchart of a progressive main influencing factor identification method for a method for predicting industry electricity consumption according to an embodiment of the present invention;

[0021] Figure 4 It is a flowchart of a prediction error-driven influencing factor refinement based on random forest for a method for predicting industry electricity consumption according to an embodiment of the present invention;

[0022] Figure 5It is a framework diagram of the SOM-BP prediction model for an industrial electricity consumption prediction method according to an embodiment of the present invention;

[0023] Figure 6 It is a SOM structure diagram of an industrial electricity consumption prediction method according to an embodiment of the present invention;

[0024] Figure 7 It is a relationship diagram of calendar variables and historical electricity variables in the steel industry in the second embodiment;

[0025] Figure 8 It is a comparison diagram of prediction results for typical summer months in different industries in the second embodiment;

[0026] Label description:

[0027] 1. An industrial electricity consumption prediction terminal; 2. A memory; 3. A processor. Specific implementation manner

[0028] To describe in detail the technical content, achieved purpose and effects of the present invention, the following is described in conjunction with the implementation manners and with reference to the accompanying drawings.

[0029] Please refer to Figure 1 , an embodiment of the present invention provides an industrial electricity consumption prediction method, including the steps:

[0030] Obtain industrial electricity consumption data samples and determine a set of candidate variables;

[0031] According to the maximum correlation and minimum redundancy between each candidate factor in the set of candidate variables and the target variable, perform a first screening of the set of candidate variables to obtain a pre-screened variable set;

[0032] Perform clustering according to the industrial electricity consumption data samples and the pre-screened variable set to obtain at least two neural network prediction models;

[0033] In the neural network prediction model, use error drive to perform a second screening of the pre-screened variable set to obtain an optimal variable set, and output an electricity consumption prediction value based on the optimal variable set.

[0034] As can be seen from the above description, the beneficial effects of the present invention are as follows: obtaining industry electricity consumption data samples and determining a set of candidate variables, performing a first screening of the set of candidate variables according to the maximum correlation and minimum redundancy between each candidate factor in the set of candidate variables and the target variable, to obtain a set of initially screened variables; performing clustering based on the industry electricity consumption data samples and the set of initially screened variables to obtain at least two neural network prediction models; in the neural network prediction models, performing a second screening of the set of initially screened variables using error-driven to obtain an optimal set of variables, and thus obtaining an electricity consumption prediction value based on the optimal set of variables, that is, calculating the value of the target variable. Therefore, first use data-driven to perform an initial screening of the set of variables, and then use error-driven to perform a refined selection of the set of variables, combining data-driven and error-driven to make up for each other's strengths and weaknesses to improve the calculation performance and accuracy.

[0035] Further, performing the first screening of the set of candidate variables according to the maximum correlation and minimum redundancy between each candidate factor in the set of candidate variables and the target variable, to obtain a set of initially screened variables includes:

[0036] Calculating the maximum correlation between each candidate factor in the set of candidate variables and the target variable, using an orthogonalization method to determine the information amount of the candidate factor independent of the set of candidate variables, and obtaining the redundancy of the candidate factor according to the information amount;

[0037] Combining the maximum correlation and the redundancy to obtain a maximum correlation-minimum redundancy measurement index;

[0038] According to a preset initial screening quantity, performing the first screening using the maximum correlation-minimum redundancy measurement index to obtain a set of initially screened variables.

[0039] As can be seen from the above description, after determining the maximum correlation between the candidate factor and the target variable, the redundancy is considered according to the principle of "maximum correlation - minimum redundancy" to further complete the first screening of the main influencing factors; by calculating the information amount of the candidate factor independent of the set of candidate variables, the redundancy can be indirectly measured, so as to improve the accuracy of the initial screening through the maximum correlation - minimum redundancy principle.

[0040] Further, performing the first screening according to a preset initial screening quantity using the maximum correlation-minimum redundancy measurement index to obtain a set of initially screened variables includes:

[0041] First, select the candidate factor with the highest correlation with the target variable according to the maximum correlation-minimum redundancy measurement index, and add it to the set of selected factors;

[0042] Calculate the scores of the candidate factors and the set of selected factors using the maximum correlation-minimum redundancy measurement index, and add the candidate factor corresponding to the highest score to the set of selected factors until the number in the set of selected factors reaches the preset initial screening quantity, and use the set of selected factors as the set of initially screened variables.

[0043] As can be seen from the above description, first, the candidate factor with the highest correlation with the target variable is selected according to the maximum correlation - minimum redundancy measurement index and added to the selected factor set. Subsequently, in the subsequent loop, the candidate factor is determined based on the maximum score of the candidate factors and the selected factor set and added to the selected factor set, thus ensuring the accuracy of the initial screening.

[0044] Furthermore, in the neural network prediction model, error - driven is used to perform a second screening on the initial screening variable set, and the optimal variable set obtained includes:

[0045] Initialize the back - propagation network prediction model;

[0046] Take the combination of the initial screening variable set and the target variable as the training set, perform cross - validation on the training set, and judge whether the average error is less than the minimum error. If so, output the initial screening variable set as the optimal variable set; otherwise, sort the variables in the initial screening variable set in descending order of importance according to random forest, and delete the last variable in the sorted initial screening variable set to obtain an updated initial screening variable set.

[0047] When the updated initial screening variable set is not an empty set, take the combination of the updated initial screening variable set and the target variable as the training set, and execute the step of performing cross - validation on the training set, and end when the updated initial screening variable set is an empty set.

[0048] As can be seen from the above description, by performing cross - validation on the combination of the initial screening variable set and the target variable, when the average error is less than the minimum error, the variable set is the optimal variable set; otherwise, variable screening is performed according to the importance ranking of random forest, and screening is carried out from the aspect of error - driven. Therefore, based on the initial screening results of data - driven influencing factors, adopting a prediction error - driven influencing factor refinement method based on random forest can balance the accuracy of screening and computational complexity.

[0049] Furthermore, clustering is performed according to the industry electricity consumption data sample and the initial screening variable set, and at least two neural network prediction models obtained include:

[0050] Obtain a vector composed of the candidate variable set and the target variable corresponding to each moment according to the industry electricity consumption data sample;

[0051] Use the self - organizing feature mapping network to perform clustering analysis on the vector, and form a corresponding back - propagation network prediction model according to the categories divided by the self - organizing feature mapping network clustering.

[0052] As can be seen from the above description, when using a self-organizing feature mapping network for clustering analysis and combining it with a neural network prediction model, during subsequent electricity consumption prediction, it can not only effectively capture the non-linear characteristics of the data, but also effectively reduce the computational complexity of the algorithm and shorten the computational time.

[0053] Please refer to Figure 2 , another embodiment of the present invention provides an industrial electricity consumption prediction terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0054] Obtain industrial electricity consumption data samples and determine a set of candidate variables;

[0055] According to the maximum correlation degree and minimum redundancy of each candidate factor in the set of candidate variables with the target variable, perform a first screening of the set of candidate variables to obtain a pre-screened variable set;

[0056] Perform clustering based on the industrial electricity consumption data samples and the pre-screened variable set to obtain at least two neural network prediction models;

[0057] In the neural network prediction model, use error-driven to perform a second screening of the pre-screened variable set to obtain an optimal variable set, and output an electricity consumption prediction value based on the optimal variable set.

[0058] As can be seen from the above description, obtain industrial electricity consumption data samples and determine a set of candidate variables, perform a first screening of the set of candidate variables according to the maximum correlation degree and minimum redundancy of each candidate factor in the set of candidate variables with the target variable to obtain a pre-screened variable set; perform clustering based on the industrial electricity consumption data samples and the pre-screened variable set to obtain at least two neural network prediction models; in the neural network prediction model, use error-driven to perform a second screening of the pre-screened variable set to obtain an optimal variable set, so as to obtain an electricity consumption prediction value based on the optimal variable set, that is, calculate the value of the target variable. Therefore, first use data-driven to perform a preliminary screening of the variable set, and then use error-driven to perform a refined selection of the variable set, combining data-driven and error-driven to make up for each other's strengths and weaknesses to improve computational performance and accuracy.

[0059] Further, according to the maximum correlation degree and minimum redundancy of each candidate factor in the set of candidate variables with the target variable, performing a first screening of the set of candidate variables to obtain a pre-screened variable set includes:

[0060] Calculate the maximum correlation degree of each candidate factor in the set of candidate variables with the target variable, use the orthogonalization method to determine the information amount of the candidate factor independent of the set of candidate variables, and obtain the redundancy of the candidate factor according to the information amount;

[0061] Combining the maximum correlation degree and the redundancy degree, a maximum correlation - minimum redundancy measure index is obtained;

[0062] According to the preset initial screening quantity, the maximum correlation - minimum redundancy measure index is used for the first screening to obtain an initial screening variable set.

[0063] As can be seen from the above description, after determining the maximum correlation degree between the candidate factors and the target variable, the redundancy degree is considered according to the principle of "maximum correlation - minimum redundancy" to further complete the first screening of the main influencing factors; by calculating the information amount of the candidate factors independent of the candidate variable set, the redundancy degree can be indirectly measured, thereby improving the accuracy of the initial screening through the maximum correlation - minimum redundancy principle.

[0064] Furthermore, according to the preset initial screening quantity, the maximum correlation - minimum redundancy measure index is used for the first screening, and the obtained initial screening variable set includes:

[0065] First, according to the maximum correlation - minimum redundancy measure index, the candidate factor with the highest correlation degree with the target variable is selected and added to the selected factor set;

[0066] The scores of the candidate factors and the selected factor set are calculated using the maximum correlation - minimum redundancy measure index, and the candidate factor corresponding to the highest score is added to the selected factor set until the number in the selected factor set reaches the preset initial screening quantity, and the selected factor set is used as the initial screening variable set.

[0067] As can be seen from the above description, first, according to the maximum correlation - minimum redundancy measure index, the candidate factor with the highest correlation degree with the target variable is selected and added to the selected factor set, and subsequently, the candidate factor is determined according to the maximum score of the candidate factors and the selected factor set in a loop and added to the selected factor set, thus ensuring the accuracy of the initial screening.

[0068] Furthermore, in the neural network prediction model, error - driven is used to perform a second screening on the initial screening variable set, and the obtained optimal variable set includes:

[0069] Initialize the back - propagation network prediction model;

[0070] The combination of the initial screening variable set and the target variable is used as a training set, and cross - validation is performed on the training set to determine whether the average error is less than the minimum error. If so, the initial screening variable set is output as the optimal variable set; otherwise, the variables in the initial screening variable set are sorted in descending order of importance according to random forest, and the last variable in the sorted initial screening variable set is deleted to obtain an updated initial screening variable set.

[0071] When the updated preliminary screening variable set is not an empty set, use the combination of the updated preliminary screening variable set and the target variable as the training set, and perform the step of cross-validating the training set, and end when the updated preliminary screening variable set is an empty set.

[0072] As can be seen from the above description, by cross-validating the combination of the preliminary screening variable set and the target variable, when the average error is less than the minimum error, the variable set is the optimal variable set; otherwise, variable screening is performed according to the importance ranking of the random forest, and screening is performed from the aspect of error driving. Therefore, based on the preliminary screening results of data-driven influencing factors, the method for selecting influencing factors driven by prediction error based on random forest can balance the accuracy and computational complexity of screening.

[0073] Further, clustering is performed according to the industrial electricity consumption data sample and the preliminary screening variable set, and at least two neural network prediction models are obtained, including:

[0074] Obtain a vector composed of the candidate variable set and the target variable corresponding to each moment according to the industrial electricity consumption data sample;

[0075] Use a self-organizing feature mapping network to perform clustering analysis on the vector, and form a corresponding backpropagation network prediction model according to the categories divided by the self-organizing feature mapping network clustering.

[0076] As can be seen from the above description, using a self-organizing feature mapping network for clustering analysis and combining with a neural network prediction model can not only effectively capture the non-linear characteristics of the data during subsequent electricity consumption prediction, but also effectively reduce the computational amount of the algorithm and shorten the computational time.

[0077] The above-mentioned method and terminal for predicting industrial electricity consumption of the present invention are applicable to obtaining the prediction results of typical industrial electricity consumption, while improving the prediction accuracy, reducing the prediction computational amount and computational time. The following is described through specific embodiments:

[0078] Embodiment 1

[0079] Please refer to Figure 1 and Figure 3 , a method for predicting industrial electricity consumption, including the steps of:

[0080] S1. Obtain an industrial electricity consumption data sample and determine a candidate variable set.

[0081] Specifically, after obtaining the industrial electricity consumption data, data preparation and preprocessing are performed to reduce the influence of noise data on the clustering effect and facilitate modeling. The specific steps include digital characterization of influencing factors, identification and correction of bad data, and normalization processing of some data, etc.

[0082] S2. According to the maximum correlation and minimum redundancy between each candidate factor in the candidate variable set and the target variable, perform the first screening of the candidate variable set to obtain a preliminary screening variable set.

[0083] S21. Calculate the maximum correlation between each candidate factor in the candidate variable set and the target variable, use the orthogonalization method to determine the information amount of the candidate factor independent of the candidate variable set, and obtain the redundancy of the candidate factor according to the information amount.

[0084] Specifically, the preliminary screening of the candidate variable set mainly measures the relevance between the influencing factor and the predicted target variable through the correlation degree. In this embodiment, the MIC (Maximum Information Coefficient) is used to calculate the correlation degree. The MIC can process discrete and continuous variables at the same time, can detect linear relationships and various non-linear relationships at the same time, and is more accurate, fairer and more general for the description of complex correlation relationships.

[0085] In this embodiment, let the daily electricity consumption value y of the industry be the predicted target variable, x be the candidate variable, and the selected variable will be used as the input variable of the subsequent prediction model. D(x, y) represents the finite two-dimensional data set composed of x and y. The definition of the maximum information coefficient MIC(D) is as follows:

[0086]

[0087] In the formula, x and y intervals are respectively divided on the two directions of the two-dimensional space D to form a grid G of x×y x×y , D|G represents the distribution of the data set D on the divided grid G, I(D|G) represents the mutual information of D|G, and I*(D, x, y) = max G∈Ω I(D|G) represents the maximum mutual information value on the grid set Ω composed of different division methods, and M(D) x,y represents the feature matrix composed of the maximum normalized mutual information of the data set D under different division intervals; n represents the number of data set samples, and B(n) represents the upper limit value of the grid division number, and B(n) = n 0.6 .

[0088] The standard Schmidt orthogonalization method is used to characterize the information amount GSO(x, S) of the candidate factor x independent of the selected factor set S, and indirectly measure the redundancy.

[0089] The calculation process of GSO(x, S) is as follows:

[0090]

[0091] In the formula, m is the number of times of selecting variables, S = {x1, x2,..., xm-1} is the selected variable set, v is the orthogonalized variable of x with respect to S, and u k = x k / ||xk || represents x k 's unit vector, <·,·> is the inner product of vectors; ||·|| is the norm of the vector.

[0092] S22. Combine the maximum correlation and the redundancy to obtain a maximum correlation - minimum redundancy measurement index.

[0093] Specifically, use MIC[GSO(x,S),y] as the maximum correlation - minimum redundancy measurement index for the influencing factor x and the predicted target variable y.

[0094] S23. According to the preset initial screening quantity, use the maximum correlation - minimum redundancy measurement index to perform the first screening to obtain an initial screening variable set.

[0095] S231. First, select the candidate factor with the highest correlation with the target variable according to the maximum correlation - minimum redundancy measurement index, and add it to the selected factor set;

[0096] S232. Use the maximum correlation - minimum redundancy measurement index to calculate the scores of the candidate factors and the selected factor set, and add the candidate factor corresponding to the highest score to the selected factor set until the number in the selected factor set reaches the preset initial screening quantity, and use the selected factor set as the initial screening variable set.

[0097] Specifically, let the candidate variable set be S c ={x 1 ,x 2 ,…,x K}, the selected variable set after the nth variable selection is S n , the target variable is y, the number of variables to be selected is N, and the iterative steps for screening variables using MIC[GSO(x,S),y] are as follows:

[0098] 1. When selecting variables for the first time, select the variable with the highest correlation with the output variable:

[0099]

[0100] 2. When selecting variables for the nth time (n>1), for the candidate variable x i , its score is recorded as:

[0101] Score(x i ) = MIC(x i |S n-1 ,y);

[0102] In the formula, S n-1 is the selected variable set after the (n - 1)th variable selection.

[0103] Select the variable with the highest score in the current candidate variable set:

[0104]

[0105] Add S n to the selected variable set S N-1 to form a new selected variable set Sn.

[0106] 3. Perform the next iteration until the number of selected variables n reaches a predetermined value N, and output the selected variable set S = S n

[0107] S3. Cluster according to the industry electricity consumption data sample and the pre-screened variable set to obtain at least two neural network prediction models.

[0108] S31. Obtain a vector composed of the candidate variable set and the target variable corresponding to each moment according to the industry electricity consumption data sample.

[0109] Specifically, in order to more accurately capture the fluctuation characteristics of the electricity consumption of each industry on a daily time scale, industry electricity consumption data samples are generated in units of each moment. Let the sample corresponding to moment t be a vector z(t) composed of electricity consumption and main influencing factors, and use SOM to perform clustering analysis on it. Considering the measurable data that can be obtained in practical applications, the z(t) selected in this embodiment includes the date stamp corresponding to moment t, the historical economic data and historical electricity consumption data of adjacent days. Therefore, although the electricity consumption fluctuations at different moments are the same, if the changing trends of economic indicators are quite different, they will also be classified into different categories by SOM, and then different BP prediction model parameters will be used to obtain their future electricity consumption.

[0110] S32. Use a self-organizing feature mapping network to perform clustering analysis on the vector, and form a corresponding backpropagation network prediction model according to the categories divided by the self-organizing feature mapping network clustering.

[0111] Please refer to Figure 5 , in order to refine the prediction modeling process, first divide the industry to be predicted into medium categories according to the National Economic Industry Classification (GB / T 4754—2017) as the sub-sectors. Use the SOM network to cluster the data set composed of the electricity consumption of the sub-sectors and the pre-screened industry influencing factors, and respectively construct a BP neural network prediction model composed of an input layer, an output layer and a hidden layer based on the clustering results. When building the BP neural network model, use Dropout in the fully connected layer to avoid the risk of overfitting.

[0112] Please refer to Figure 6, The self-organizing feature mapping (SOM) network has a relatively simple neural network structure and is a type of "unsupervised learning" model that can represent high-dimensional input data in a low-dimensional space and is commonly used in applications such as clustering and data visualization. The SOM structure consists of an input layer and an output layer. The input layer corresponds to the industry-specific data vectors formed by the historical electricity consumption of each industry and the influencing factor data obtained through preliminary screening. The output layer is composed of ordered nodes in a two-dimensional grid, and the two are connected by weights. SOM judges the similarity between samples through the Euclidean distance. During the learning process, the input sample finds the competitive layer unit with the shortest distance to it as the winning neuron, and updates the weights of the winning neuron and the adjacent area. This neuron represents the classification result of the input vector.

[0113] S4. In the neural network prediction model, use error-driven to perform a second screening on the preliminary screening variable set to obtain an optimal variable set, and output the electricity consumption prediction value based on the optimal variable set.

[0114] To balance accuracy and computational complexity, based on the preliminary screening results of data-driven influencing factors, a method for selecting influencing factors driven by prediction error based on random forest is adopted. It constructs multiple decision trees through random resampling technology and node random splitting technology, and obtains the prediction result by averaging the results of multiple decision trees, with characteristics such as high prediction accuracy and controllable generalization error.

[0115] S41. Initialize the Back Propagation Network (BP) prediction model.

[0116] S42. Use the combination of the preliminary screening variable set and the target variable as the training set, perform cross-validation on the training set, and judge whether the average error is less than the minimum error. If so, output the preliminary screening variable set as the optimal variable set; otherwise, sort the variables in the preliminary screening variable set in descending order of importance according to random forest, and delete the last variable in the sorted preliminary screening variable set to obtain an updated preliminary screening variable set.

[0117] Specifically, please refer to Figure 4 , The cross-validation training set is composed of the X variable set S and the target variable y, and judge whether the average error is less than the minimum error. If it is less, output S; otherwise, sort the variables in S in descending order of importance according to random forest and delete the last variable in S.

[0118] S43. When the updated preliminary screening variable set is not an empty set, use the combination of the updated preliminary screening variable set and the target variable as the training set, and execute the step of performing cross-validation on the training set, and end when the updated preliminary screening variable set is an empty set.

[0119] Specifically, please refer toFigure 4 , a new variable set S' and the corresponding training set X' are obtained. If it is determined that S' is not an empty set, the prediction model is initialized in a loop.

[0120] Therefore, in this embodiment, for the electricity consumption prediction method of typical industrial industries based on SOM-BP, through the screening of influencing factors in two stages of rough and fine, a BP network model is constructed based on the clustering results of electricity consumption of sub-industries and data samples of main influencing factors for classification prediction and integration, so as to obtain the prediction results of electricity consumption of typical industries.

[0121] Embodiment 2

[0122] This embodiment provides a specific application scenario. Specifically, a daily electricity consumption dataset of typical industries in a certain region is used as the data source. The proposed SOM-BP industry electricity consumption prediction method based on progressive identification of main influencing factors is compared with traditional prediction methods that do not perform main factor identification, progressive identification, or sub-industry clustering to verify the applicability of this method in predicting daily electricity consumption of some typical industries.

[0123] Step 1. Dataset description

[0124] The data source is a dataset of industries such as ferrous metal smelting and rolling processing (referred to as steel) and equipment manufacturing in a certain region from January 2019 to August 2020. This dataset contains daily electricity consumption information, date types, economic indicators, etc. of this region from January 2019 to September 2020. In this embodiment, the industry electricity consumption data from January 1, 2019 to May 31, 2020 is selected as the training set, and from June 1, 2020 to August 31, 2020 is the test set.

[0125] Step 2. Candidate variables

[0126] Before variable selection, a reasonable candidate variable set must first be determined. In this embodiment, in combination with the selected industry dataset, a candidate variable set is constructed from four aspects: trend variables, calendar variables, price variables, and historical electricity consumption variables. Trend variables mainly reflect the changes in electricity consumption brought about by the improvement of the national production level. Due to production and work and rest habit factors, calendar variables often have a correlation with electricity consumption. Please refer to Figure 7 , as the calendar variables change, the changes in electricity consumption show obvious periodicity; the changes in electricity consumption of continuous production industries in a year are related to the economic situation; the change trend of the load of non-continuous production industries in a year is mainly related to work and rest habits, and holidays have the greatest impact. The selected candidate variable set is shown in Table 1.

[0127] Table 1 Candidate variable set

[0128]

[0129] Among the candidate variable sets, Tre represents the changing trend of electricity consumption, which is a linearly increasing variable starting from 1 and accumulating. Among the calendar variables, M t is the month identifier, which is an integer between [1, 12]; W t is the week identifier, which is an integer between [1, 7]; Hol is the holiday identifier; WD is the weekend identifier. Among the price variables, represents the average price of the products in this industry; represents the average price of the products in the upstream industry; respectively represent the historical electricity price and the predicted electricity price; respectively represent the output of the downstream industry and the predicted output of the downstream industry. The historical electricity consumption variables include the industry electricity consumption Q t with a daily cycle and the industry electricity consumption structure R (the proportion of productive load).

[0130] Step Three: Evaluation Index

[0131] To evaluate the prediction effect, the prediction error of the test set is used as the prediction accuracy. The specific evaluation indexes are the mean absolute percentage error e MAPE and the standard deviation e SDAE of the absolute error, which are used to evaluate the average magnitude and the degree of dispersion of the prediction error respectively. The calculation formulas are as follows:

[0132]

[0133] In the formula, k is the sample size; is the actual industry electricity quantity sequence; is the industry electricity quantity prediction sequence.

[0134] Step Four: Result Analysis

[0135] 1. Comparison between different methods

[0136] The following several methods are used for the daily electricity consumption prediction of the steel industry: (1) data-driven variable selection + BP neural network; (2) error-driven variable selection + BP neural network; (3) all variables + BP neural network; (4) autoregressive moving average model; (5) the method of this embodiment;

[0137] Among them, method (4) is a traditional time series prediction method that does not consider the influencing factors other than time; method (3) is a traditional BP neural network prediction method that considers the influencing factors but does not identify the main factors first; after using different types of identification methods in methods (2) and (1), the selected main influencing factors are used to construct a BP neural network prediction model; method (5), that is, the method of this embodiment, is based on the progressive influencing factor identification idea and uses SOM-BP as the prediction model for the daily electricity consumption prediction of the steel industry.

[0138] The iron and steel industry has 4 sub - sectors, namely ironmaking, steelmaking, steel rolling processing, and ferroalloy smelting. The input layer x of SOM i (i = 1, 2, 3, 4) are 6 industry influencing factors (Tre, Hol, Q t ) and the data set composed of the electricity consumption of a certain sub - sector; the competitive layer is designed as a two - dimensional plane discrete network composed of 2×3 neurons, and is fully connected to the input layer; the number of clusters is controlled between 2 and 4, reducing the number of training and test models while maintaining good prediction accuracy. After SOM clustering, the sub - sectors of the iron and steel industry are divided into two categories. The first category is steel rolling, and the second category is ironmaking, steelmaking, and ferroalloy smelting. The characteristics are mainly reflected in: the first category mostly has a non - continuous production work system, with low electricity consumption on weekends and holidays, and the downstream industries are construction and equipment manufacturing; the second category mostly has a continuous production work system, with relatively stable daily electricity consumption fluctuations, and is more affected by the prices of upstream raw materials and the output of downstream industries.

[0139] Based on the SOM clustering results, two BP network models are constructed. The input layer of each BP neural network model is the main influencing factors after variable pre - screening; the output layer is the predicted value of the daily electricity consumption of the classified industry; the hidden layer is one layer, and the number of nodes is set to (the number of input nodes + the number of output nodes) / 2, and the activation function uses the Sigmoid function. The SOM - BP method classifies and trains, predicts, and sums up the BP network model after clustering the sub - sectors to obtain the predicted results of the industry electricity consumption. Table 2 compares the prediction accuracies of the daily electricity consumption of the iron and steel industry using various methods.

[0140] Table 2 Comparison of prediction errors of different variable selection methods

[0141]

[0142] The variable data selected by the first three methods are different. Method (1) selects 6 main influencing factors, namely Tre, Hol, Q t . Method (2) selects 9 main influencing factors, namely Tre, M t , Hol, Q t , R. Method (5) selects 6 variables after the first - step pre - screening, namely Tre, Hol, Q t ; after further refinement, 3 variables that have the greatest impact on the error are selected, namely Hol, Q t .

[0143] As can be seen from Table 2, the proposed method avoids the omission of important variables and reduces the redundancy of variable selection when considering variable selection. The prediction accuracy after prediction on this basis shows the applicability of this method in predicting the daily electricity consumption in the steel industry. The prediction accuracy based on the variable set formed by the proposed progressive main influencing factor identification method is the highest, e MAPE = 2.07%, e SDAE = 68.93 MWh. The prediction result of this method is better than that of the traditional BP neural network in method (3) that considers influencing factors. The mean absolute percentage error e MAPE decreases by 1.54%, and the standard deviation of the absolute error e SDAE decreases by 53.23 MWh; it is also significantly improved compared with the prediction result of the traditional time series prediction method in method (4). e MAPE decreases by 3.06%, e SDAE decreases by 377.35 MWh.

[0144] 2. Comparison between different types of industries

[0145] Taking July 2020 in the test set as the representative month in summer, the effects of the proposed algorithm for predicting the electricity consumption of different types of industries are compared. The electricity consumption of the steel industry (continuous production type) and the equipment manufacturing industry (non - continuous production type) on each day of this month is predicted respectively, and different evaluation indicators are used to compare the prediction accuracy of each industry. Please refer to Figure 8 , for the two prediction accuracy evaluation indicators, the prediction accuracy of the steel industry is higher than that of the equipment manufacturing industry. This is because the electricity consumption of the equipment manufacturing industry with more non - continuous production enterprises fluctuates more greatly than that of the steel industry with more continuous production enterprises during the same period, and the proportion of productive load of steel enterprises is significantly higher than that of equipment manufacturing industry enterprises. The proposed method has better prediction performance for the continuous production type steel industry, and the prediction effect for the non - continuous production type equipment manufacturing industry is also within an acceptable range.

[0146] Embodiment 3

[0147] Please refer to Figure 2 , an industry electricity consumption prediction terminal 1, including a memory 2, a processor 3, and a computer program stored on the memory 2 and executable on the processor 3. When the processor 3 executes the computer program, each step of an industry electricity consumption prediction method in Embodiment 1 or 2 is implemented.

[0148] In summary, a method and a terminal for predicting industry electricity consumption provided by the present invention obtain industry electricity consumption data samples and determine a set of candidate variables. According to the maximum correlation and minimum redundancy between each candidate factor in the set of candidate variables and the target variable, a first screening of the set of candidate variables is performed to obtain a preliminarily screened variable set. Clustering is performed based on the industry electricity consumption data samples and the preliminarily screened variable set to obtain at least two neural network prediction models. Among them, a self-organizing feature mapping network is used for clustering analysis and combined with the neural network prediction model. When predicting electricity consumption subsequently, it can not only effectively capture the non-linear characteristics of the data, but also effectively reduce the computational complexity of the algorithm and shorten the computation time. In the neural network prediction model, error-driven is used to perform a second screening of the preliminarily screened variable set to obtain an optimal variable set, and thus an electricity consumption prediction value is obtained based on the optimal variable set, that is, the value of the target variable is calculated. Among them, by performing cross-validation on the combination of the preliminarily screened variable set and the target variable, when the average error is less than the minimum error, the variable set is the optimal variable set; otherwise, variable screening is performed according to the importance ranking of the random forest, and screening is performed from the aspect of error-driven. Therefore, based on the preliminary screening results of the data-driven influencing factors, a method for selecting influencing factors by predicting error driven based on random forest can balance the accuracy of screening and the computational complexity.

[0149] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. All equivalent transformations made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in the related technical fields, are equally included in the patent protection scope of the present invention.

Claims

1. An industrial electricity consumption prediction method, characterized in that, it includes the steps of: Obtain industrial electricity consumption data samples and determine a set of candidate variables; According to the maximum correlation and minimum redundancy between each candidate factor in the set of candidate variables and the target variable, perform the first screening of the set of candidate variables to obtain a pre-screened variable set: Calculate the maximum correlation between each candidate factor in the set of candidate variables and the target variable, use the orthogonalization method to determine the information amount of the candidate factor independent of the set of candidate variables, and obtain the redundancy of the candidate factor according to the information amount; Combine the maximum correlation and the redundancy to obtain a maximum correlation-minimum redundancy measurement index; According to the preset pre-screening quantity, use the maximum correlation-minimum redundancy measurement index to perform the first screening to obtain a pre-screened variable set: First, select the candidate factor with the highest correlation with the target variable according to the maximum correlation-minimum redundancy measurement index, and add it to the set of selected factors; use the maximum correlation-minimum redundancy measurement index to calculate the scores of the candidate factors and the set of selected factors, and add the candidate factor corresponding to the highest score to the set of selected factors until the number in the set of selected factors reaches the preset pre-screening quantity, and use the set of selected factors as the pre-screened variable set; Perform clustering according to the industrial electricity consumption data samples and the pre-screened variable set to obtain at least two neural network prediction models; In the neural network prediction model, use error drive to perform the second screening on the pre-screened variable set to obtain an optimal variable set: Initialize the backpropagation network prediction model; Use the combination of the pre-screened variable set and the target variable as the training set, perform cross-validation on the training set, and judge whether the average error is less than the minimum error. If so, output the pre-screened variable set as the optimal variable set. Otherwise, sort the variables in the pre-screened variable set in descending order of importance according to the random forest, and delete the last variable in the sorted pre-screened variable set to obtain an updated pre-screened variable set; when the updated pre-screened variable set is not an empty set, use the combination of the updated pre-screened variable set and the target variable as the training set, and execute the step of performing cross-validation on the training set, and end when the updated pre-screened variable set is an empty set; Output the electricity consumption prediction value based on the optimal variable set.

2. An industrial electricity consumption prediction method according to claim 1, characterized in that, Performing clustering according to the industrial electricity consumption data samples and the pre-screened variable set to obtain at least two neural network prediction models includes: Obtain the vector composed of the set of candidate variables and the target variable corresponding to each moment according to the industrial electricity consumption data samples; Use the self-organizing feature mapping network to perform clustering analysis on the vector, and form a corresponding backpropagation network prediction model according to the categories divided by the self-organizing feature mapping network clustering.

3. An industrial electricity consumption prediction terminal, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor, characterized in that, when the processor executes the computer program, the following steps are implemented: Obtain industrial electricity consumption data samples and determine a set of candidate variables; Perform a first screening of the set of candidate variables based on the maximum correlation and minimum redundancy between each candidate factor in the set of candidate variables and the target variable, to obtain a pre-screened variable set: Calculate the maximum correlation between each candidate factor in the set of candidate variables and the target variable, use the orthogonalization method to determine the information amount of the candidate factor independent of the set of candidate variables, and obtain the redundancy of the candidate factor according to the information amount; Combine the maximum correlation and the redundancy to obtain a maximum correlation-minimum redundancy measurement index; According to a preset pre-screening quantity, perform a first screening using the maximum correlation-minimum redundancy measurement index to obtain a pre-screened variable set: First, select the candidate factor with the highest correlation with the target variable according to the maximum correlation-minimum redundancy measurement index, and add it to the set of selected factors; Use the maximum correlation-minimum redundancy measurement index to calculate the scores of the candidate factor and the set of selected factors, and add the candidate factor corresponding to the highest score to the set of selected factors until the number in the set of selected factors reaches the preset pre-screening quantity, and use the set of selected factors as the pre-screened variable set; Perform clustering based on the industrial electricity consumption data sample and the pre-screened variable set to obtain at least two neural network prediction models; In the neural network prediction model, perform a second screening on the pre-screened variable set using error drive to obtain an optimal variable set: Initialize the backpropagation network prediction model; Use the combination of the pre-screened variable set and the target variable as a training set, perform cross-validation on the training set, and judge whether the average error is less than the minimum error. If so, output the pre-screened variable set as the optimal variable set. Otherwise, sort the variables in the pre-screened variable set in descending order of importance according to random forest, and delete the last variable in the sorted pre-screened variable set to obtain an updated pre-screened variable set; When the updated pre-screened variable set is not an empty set, use the combination of the updated pre-screened variable set and the target variable as a training set, and execute the step of performing cross-validation on the training set, and end when the updated pre-screened variable set is an empty set; Output an electricity consumption prediction value based on the optimal variable set.

4. An industrial electricity consumption prediction terminal according to claim 3, wherein, Performing clustering based on the industrial electricity consumption data sample and the pre-screened variable set to obtain at least two neural network prediction models includes: Obtain a vector composed of the set of candidate variables and the target variable corresponding to each moment according to the industrial electricity consumption data sample; Use a self-organizing feature mapping network to perform clustering analysis on the vector, and form a corresponding backpropagation network prediction model according to the categories divided by the self-organizing feature mapping network clustering.

Citation Information

Patent Citations

  • Forecasting method of industrial electricity quantity based on BP-LSSVM combined optimization model

    CN109063892A

  • Photovoltaic power generation prediction method based on improved extreme learning machine

    CN114298377A