Monthly energy consumption splitting and complementing method and system

Through integrated data processing, feature extraction and deep learning models, a monthly energy consumption split completion method is provided, which solves the problem of difficult for traditional methods to adapt to complex energy consumption distribution and insufficient data processing capabilities, and achieves efficient and accurate energy consumption analysis and prediction.

CN120045854APending Publication Date: 2025-05-27STATE GRID SHANDONG ELECTRIC POWER CO +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510115775.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional energy consumption analysis methods are based on empirical rules and static statistics, and are difficult to adapt to complex energy consumption distribution and changing patterns, and cannot fully consider the complex correlation between energy consumption, especially in the energy consumption situation of multiple branches and multiple devices, and the processing ability of missing data is weak.

Method used

By integrating data collection, feature extraction, deep learning energy consumption splitting model and missing data prediction model, a monthly energy consumption splitting method and system is provided. The method includes obtaining energy consumption data, identifying and repairing abnormal data, reducing dimensionality, building a Lasso regression model for energy consumption splitting, and using the LSTM-BPNN model to predict and complete missing data.

Benefits of technology

It realizes efficient and accurate monthly energy consumption analysis and prediction, can better adapt to complex energy consumption distribution and changing modes, effectively handle multi-branch and multi-equipment situations, improve the accuracy and generalization capabilities of the splitting process, and ensure the accuracy of data completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045854A_ABST
    Figure CN120045854A_ABST
Patent Text Reader

Abstract

The invention provides a monthly energy consumption splitting and complementing method and system, and relates to the technical field of information data processing. The method includes acquiring original data; identifying and marking abnormal data in the original data and repairing the marked abnormal data; carrying out dimension reduction processing on the repaired data; building an energy consumption splitting model based on a Lasso regression model, carrying out pre-training, inputting the obtained features into the pre-trained energy consumption splitting model, carrying out year-season-month energy consumption splitting on the energy consumption, and obtaining monthly energy consumption of different branches; and building an LSTM-BPNN model as a missing data prediction model, and complementing missing data in monthly energy consumption of different branches based on the missing data prediction model. According to the method, data collection, feature extraction, deep learning of the energy consumption splitting model and the missing data prediction model are integrated, splitting, prediction and complementation of monthly energy consumption are achieved, and the method can be widely applied to the fields of industry, business, living and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information data processing, and particularly relates to a monthly energy consumption splitting and complementing method and system. Background Art

[0002] Energy consumption splitting and complementing is an important support for tasks such as optimized design, optimized control, demand-side response, and energy auditing, and is of great significance for improving energy efficiency levels and achieving energy conservation and emission reduction goals. With the rapid development of information technology and artificial intelligence technology, the continuously accumulated operation data and emerging algorithms have given greater development potential to data-driven methods.

[0003] In past energy consumption analyses, traditional statistical methods and rule models were often used for energy consumption splitting and complementing:

[0004] Split the total energy consumption into the energy consumption of each branch or energy-consuming device, providing more fine-grained energy consumption data for convenient analysis, monitoring, and optimization of different devices or regions;

[0005] Complement missing monthly energy consumption data, and infer the energy consumption value at the missing time point through time series prediction to ensure the integrity and availability of the data.

[0006] The inventors found that traditional methods are usually based on empirical rules and static statistical data, are difficult to adapt to complex energy consumption distributions and change patterns, and cannot fully consider the complex correlation relationships between energy consumptions, especially in the energy consumption scenarios of multiple branches and multiple devices. Traditional methods have weak processing capabilities for missing data and are difficult to accurately estimate missing values. Summary of the Invention

[0007] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a monthly energy consumption splitting and complementing method and system, aiming to provide an efficient and accurate monthly energy consumption analysis and prediction tool. By integrating data collection, feature extraction, a deep learning energy consumption splitting model, and a missing data prediction model, the splitting, prediction, and complementing of monthly energy consumption are realized, and it can be widely applied to fields such as industry, commerce, and residence.

[0008] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:

[0009] The first aspect of the present invention provides a monthly energy consumption splitting and complementing method.

[0010] A monthly energy consumption splitting and complementing method includes the following steps:

[0011] Obtain the energy consumption data and power consumption data of the enterprise group in the target industry as the original data;

[0012] Identify and mark the abnormal data in the original data, and repair the marked abnormal data;

[0013] Perform dimensionality reduction on the repaired data, and retain the set number of features with the most information;

[0014] Build an energy consumption splitting model based on the Lasso regression model and perform pre-training. Input the obtained features into the pre-trained energy consumption splitting model to perform annual-quarter-month energy consumption splitting on the energy consumption, and obtain the monthly energy consumption of different branches;

[0015] Build an LSTM-BPNN model as a missing data prediction model, and complete the missing data in the monthly energy consumption of different branches based on the missing data prediction model.

[0016] The second aspect of the present invention provides a monthly energy consumption splitting and completion system.

[0017] A monthly energy consumption splitting and completion system, comprising:

[0018] A data acquisition module, configured to: acquire the energy consumption data and power consumption data of the enterprise group in the target industry as the original data;

[0019] A data preprocessing module, configured to: identify and mark the abnormal data in the original data, and repair the marked abnormal data;

[0020] A data dimensionality reduction module, configured to: perform dimensionality reduction on the repaired data, and retain the set number of features with the most information;

[0021] A data splitting module, configured to: build an energy consumption splitting model based on the Lasso regression model and perform pre-training. Input the obtained features into the pre-trained energy consumption splitting model to perform annual-quarter-month energy consumption splitting on the energy consumption, and obtain the monthly energy consumption of different branches;

[0022] A data completion module, configured to: build an LSTM-BPNN model as a missing data prediction model, and complete the missing data in the monthly energy consumption of different branches based on the missing data prediction model.

[0023] The above one or more technical solutions have the following beneficial effects:

[0024] The present invention provides a monthly energy consumption splitting and completion method and system. By integrating multiple modules such as data processing, feature extraction, deep learning, and optimization algorithms, it realizes efficient and accurate monthly energy consumption analysis and prediction. By identifying and repairing the abnormal data in the original data, the accuracy of the data can be improved. When using the repaired data to train the energy consumption splitting model later, better splitting results can be obtained, and the monthly energy consumption of different branches obtained is relatively accurate;

[0025] After that, an LSTM - BPNN model is built as a missing data prediction model. For the energy consumption with missing data in the splitting results, the corresponding historical monthly energy consumption and the LSTM - BPNN model are used to complete the filling of the missing values. Finally, the filling result obtained has a high accuracy.

[0026] The present invention can better adapt to complex energy consumption distributions and change patterns, effectively handle the scenarios of multiple branches and multiple devices, and improve the accuracy and generalization ability of the splitting process.

[0027] The present invention ensures the accuracy of data filling, makes the entire monthly energy consumption analysis process more systematic and controllable, facilitates applications in different fields, and provides strong support for energy consumption optimization and management in industrial, commercial, residential and other fields.

[0028] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0030] Figure 1 It is a flowchart of the method for Embodiment 1.

[0031] Figure 2 It is a structural diagram of the LSTM - BPNN model.

[0032] Figure 3 It is a flowchart of the PSO operation. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0034] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.

[0035] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0036] Embodiment 1

[0037] As described above, traditional methods are usually based on empirical rules and static statistical data, making it difficult to adapt to complex energy consumption distributions and changing patterns, and unable to fully consider the complex correlation relationships between energy consumptions, especially in the energy consumption scenarios of multiple branches and multiple devices. Traditional methods have weak capabilities in handling missing data and are difficult to accurately estimate missing values.

[0038] Based on this, this embodiment discloses a monthly energy consumption splitting and filling method. As Figure 1 shown, the monthly energy consumption splitting and filling method proposed in this embodiment may include the following steps:

[0039] Obtain the energy consumption data and power consumption data of the enterprise group in the target industry as the original data;

[0040] Identify and mark the abnormal data in the original data, and repair the marked abnormal data;

[0041] Perform dimensionality reduction processing on the repaired data, and retain the set number of features with the most information;

[0042] Build an energy consumption splitting model based on the Lasso regression model and perform pre-training. Input the obtained features into the pre-trained energy consumption splitting model to perform annual-quarter-month energy consumption splitting on the energy consumption, and obtain the monthly energy consumption of different branches;

[0043] Build an LSTM-BPNN model as a missing data prediction model, and fill in the missing data in the monthly energy consumption of different branches based on the missing data prediction model.

[0044] In some embodiments, identifying and marking the abnormal data in the original data and repairing the marked abnormal data specifically include:

[0045] Identify and mark the abnormal data in the original data through the K-means algorithm;

[0046] Repair the marked abnormal data through the K-nearest neighbor algorithm.

[0047] In some embodiments, identifying and marking the abnormal data in the original data through the K-means algorithm specifically includes:

[0048] Define K data points as the initial clustering centers;

[0049] Calculate the Euclidean distance from the data points in the original data to the initial clustering centers, and divide the original data into K clusters based on similar data points;

[0050] Recalculate the central values of the K clusters as the new clustering centers for re-clustering, and iterate until the clustering criterion function converges to obtain the clustering results of the K clusters;

[0051] Mark the data points in the clustering results of K clusters whose distances from the corresponding cluster centers are greater than the set value as abnormal data;

[0052] Alternatively, repair the marked abnormal data through the K-nearest neighbor algorithm, specifically including:

[0053] Eliminate incomplete data from the original data and construct a complete data set matrix;

[0054] Calculate the Euclidean distance between the abnormal data and each data in the complete data set matrix;

[0055] Based on the calculated Euclidean distance, screen out the K data that are the nearest neighbors of the abnormal data;

[0056] Calculate the nearest neighbor weights of the K data that are the nearest neighbors of the abnormal data, and respectively use the nearest neighbor weights to perform weighted summation calculations on the corresponding data in the K data as the estimated abnormal data.

[0057] In some embodiments, perform dimensionality reduction processing on the repaired data through the principal component analysis method PCA, and retain the set number of features with the most information, specifically including:

[0058] Based on the repaired data, construct a data matrix represented by variables and the number of samples;

[0059] Perform Max-Min standardization transformation on the data matrix to obtain a standardized matrix;

[0060] Solve the correlation coefficient matrix of the standardized matrix;

[0061] Calculate the characteristic equation of the correlation coefficient matrix to obtain p eigenvalues;

[0062] Determine the set threshold of the cumulative contribution rate of the principal components, and select the first N principal components from the p eigenvalues that make the cumulative contribution rate of the principal components reach the set threshold of the cumulative contribution rate of the principal components as the features with the most information.

[0063] In some embodiments, the energy consumption splitting model includes:

[0064] The annual-quarter energy consumption splitting objective function is:

[0065]

[0066] where y i represents the annual energy consumption of the y i th enterprise, α j represents the energy consumption splitting coefficient of the jth quarter, and x ij represents the energy consumption of the ith enterprise in the jth quarter;

[0067] The quarter-month energy consumption splitting objective function is:

[0068]

[0069] wherein, m i represents the energy consumption of the i-th enterprise in a certain quarter, and β j represents the energy consumption splitting coefficient of the j-th month, and n ij represents the energy consumption of the i-th enterprise in the j-th month;

[0070] The calculation method of the monthly energy consumption splitting weight is as follows:

[0071] z ij = α i × β ij ;

[0072] wherein, z ij represents the energy consumption splitting coefficient of the j-th month in the i-th quarter finally, α i represents the energy consumption splitting coefficient of the i-th quarter, and β ij represents the energy consumption splitting coefficient of the j-th month in the i-th quarter.

[0073] In some embodiments, pre-training is performed on the energy consumption splitting model, specifically:

[0074] Calculate the proportion of the number of enterprises with missing energy consumption in the annual / monthly energy consumption data and the corresponding power consumption data of the enterprise group in the target industry to the total number of enterprises in the enterprise group:

[0075] When the proportion is greater than 40%, use the power consumption data of the enterprise group as the training data of the energy consumption splitting model;

[0076] When the proportion is less than 40%, use the monthly energy consumption data of enterprises with complete monthly energy consumption data as the training data of the energy consumption splitting model.

[0077] In some embodiments, an LSTM-BPNN model is built as a missing data prediction model, and the missing data in the monthly energy consumption of different branches is complemented based on the missing data prediction model, specifically including:

[0078] Build an LSTM-BPNN model by combining the long short-term memory network LSTM and the backpropagation neural network BPNN;

[0079] Train the LSTM-BPNN model through historical energy consumption data, and optimize the parameters of the LSTM-BPNN model through the particle swarm optimization algorithm PSO;

[0080] Input the historical monthly energy consumption of the branch with missing data into the trained and optimized LSTM-BPNN model to obtain the predicted missing data, and realize the complement of the missing data in the monthly energy consumption.

[0081] In some embodiments, the final hidden state h of the LSTM t is calculated as follows:

[0082]

[0083] o t = σ(W o · [h t-1 , x t + b o );

[0084]

[0085] i t = σ(W i · [h t-1 , x t + b i );

[0086] f t = σ(W f · [h t-1 , x t + b f );

[0087] where o t is the output gate, c t is the cell state at the current time, c t-1 is the cell state at the previous time, h t-1 is the hidden state at the previous time, x t is the input at the current time, is the candidate cell state at the current time, σ(·) is the Sigmoid activation function, W f , W i and W o are the weight matrices of the forget gate, input gate, and output gate respectively, W c is the weight matrix for calculating the candidate state, b f , b i and b o are the excitation threshold vectors of the forget gate, input gate, and output gate respectively, b c is the excitation threshold vector for calculating the candidate state;

[0088] Let the feature vector input to the BPNN be x = h t , where h t is the hidden state finally output by the LSTM. After the feedforward calculation of the input layer, hidden layer, and output layer, the output of the BPNN is obtained

[0089]

[0090] Among them, W out and W hidden are the connection weights of the output layer and the hidden layer respectively, and b out and b hidden are the firing thresholds of the neurons in the output layer and the hidden layer respectively. is the activation function of the hidden layer, and h is the output of the hidden layer.

[0091] In some embodiments, the parameters of the LSTM - BPNN model are optimized by the particle swarm optimization algorithm PSO, specifically as follows:

[0092] First, assign initial random positions and initial random velocities to all particles in the space;

[0093] Then, advance the positions of each particle in turn according to the velocity of each particle, the known optimal global position in the problem space, and the known optimal position of the particle;

[0094] As the calculation progresses, by exploring and exploiting the known favorable positions in the search space, the particles gather or aggregate around one or more optimal points.

[0095] Next, the technical solution of this embodiment will be further explained in conjunction with the accompanying drawings. Generally speaking, this embodiment may include the following steps:

[0096] Step S1: Identify and mark the abnormal data in the sensor data source through the K - means algorithm, and repair the marked abnormal data using the K - nearest neighbor (KNN) algorithm;

[0097] Step S2: Reduce the dimension of the data through the principal component analysis method (PCA), and retain the most informative features as the model input;

[0098] Step S3: Perform "year - quarter - month" energy consumption splitting on the energy consumption through the Lasso regression model to obtain the energy consumption of different branches;

[0099] Step S4: Combine the long short - term memory network (LSTM) and the back - propagation neural network (BPNN) to obtain the LSTM - BPNN model, and complete the missing data caused by the introduction of new devices or new branches and changes in the data source based on the LSTM - BPNN model optimized by the particle swarm algorithm (PSO).

[0100] As one or more embodiments, in step S1: Identify and mark the abnormal data in the sensor data source through the K - means algorithm, and repair the marked abnormal data using the K - nearest neighbor (KNN) algorithm. The specific steps include:

[0101] Step 101: Identify and mark the abnormal data in the sensor data source through the K-means algorithm, specifically:

[0102] The K-means algorithm divides data samples into K clusters by calculating the Euclidean distance from data points to the cluster centers, and finally divides similar data points into the same cluster. The prerequisite for the K-means algorithm to achieve effective clustering is to determine accurate cluster centers. The general method for determining them is to randomly define K data points as the initial cluster centers for clustering, then calculate the central values of the K clusters as the new cluster centers for re-clustering, and repeat this process until the clustering criterion function converges.

[0103] The convergence function is as follows:

[0104]

[0105] where E is the minimum squared error of the clusters obtained after clustering the data samples by the K-means algorithm, and u i is the mean vector of cluster C i . Its calculation method is as follows:

[0106]

[0107] When the minimum squared error E of each cluster after clustering is smaller, it indicates that the data samples within each cluster are closer to the cluster mean vector, and the similarity degree of the samples within the cluster is higher. Using the K-means algorithm can achieve a good effect of abnormal data identification. For the abnormal values marked for deletion, subsequent repair processing is required.

[0108] Specific instructions are as follows:

[0109] K-means will first randomly select k data points as the initial cluster centers. For each data point in the dataset, calculate its distances to all cluster centers, and then update by taking the mean of all points within the cluster until the change in the cluster centers is very small or reaches the predetermined number of iterations, which indicates that the algorithm has converged and the clustering is completed. Then calculate the distance from each data point to its cluster center. If the distance from a data point to its cluster center exceeds the set threshold, then this data point is considered abnormal.

[0110] Step 102: Repair the marked abnormal data using the K-Nearest Neighbor (KNN) algorithm, specifically:

[0111] The KNN algorithm is driven by historical data. By comparing the state vector of the data to be repaired with the state vectors of historical data, find the K nearest values that are most similar to the state vector of the data to be repaired, and calculate their weighted average as the estimated value of the data to be repaired.

[0112] First, eliminate incomplete data and construct a complete dataset matrix \((x 1 ,x 2 ,…,x j ,…,x m ) T ; Calculate the Euclidean distance between the data to be repaired and all the data in the complete dataset matrix. Taking a missing data \(x mis \) with a state vector of \(n\) dimensions as an example, the Euclidean distance between it and the complete data \(x j \) is:

[0113]

[0114] Then, calculate the weight \(w i \) of the nearest neighbor of the data to be repaired, as well as the replacement value \(x repair \) of the data to be repaired. The calculation methods are as follows:

[0115]

[0116]

[0117] Among them, \(x i \) is the corresponding nearest neighbor data value.

[0118] As one or more embodiments, the step S2: Perform dimensionality reduction on the data through the principal component analysis method (PCA), and retain the most informative features as the model input. The specific steps include:

[0119] Let the input data be a matrix \(X n×p \), \(p\) represents variables, and \(n\) represents the number of samples. Construct a data matrix and perform the following Max-Min standardization transformation on the matrix

[0120]

[0121] Among them, \(X ij \) represents the \(j\)-th value of the \(i\)-th sample, \(i = 1, 2, …, n\), \(j = 1, 2, …, p\).

[0122] Perform Max-Min standardization on the matrix \(X n×p \) to obtain \(Y n×p

[0123]

[0124] Among them,

[0125] Calculate the correlation coefficient matrix \(R=(r ij ) n×p \), where \(r ij \) is calculated as follows:

[0126]

[0127] Calculate the characteristic equation of the data matrix R, |R - λI p | = 0 to obtain p eigenvalues, determine the cumulative contribution rate of the principal components, and the calculation method is as follows:

[0128]

[0129] Among them, η is a constant, generally taking η = 0.85, indicating that the first p principal components can achieve the purpose of replacing multiple original indicators with fewer indicators. Through the above method, dimensionality reduction processing is performed on multiple energy consumption processing factors, and the first N principal components with a cumulative contribution rate reaching 0.85 are selected as the input of the prediction model, and η can be adjusted according to requirements.

[0130] As one or more embodiments, in step S3: perform "year-quarter-month" energy consumption splitting on the energy consumption through a Lasso regression model to obtain the energy consumption of different branches. The specific steps include:

[0131] Construct a "year-quarter-month" energy consumption splitting model based on historical energy consumption, use the annual / monthly energy consumption data of enterprise groups in the industry segment and the corresponding power consumption data as input, and calculate the proportion of the number of enterprises with missing energy consumption in the total number of enterprises in the enterprise group. When the proportion is greater than 40%, use the power consumption data of the enterprise group as the training data for the splitting model; when the proportion is less than 40%, use the monthly energy consumption data of enterprises with complete monthly energy consumption data as the training data for the splitting model.

[0132] Construct the following objective function:

[0133] The "year-quarter" energy consumption splitting objective function is:

[0134]

[0135] Among them, y i represents the annual energy consumption of the y i th enterprise, α j represents the energy consumption splitting coefficient of the jth quarter, and x ij represents the energy consumption of the ith enterprise in the jth quarter.

[0136] Through the above "year-quarter" Lasso regression objective function, α can be obtained, and thus the splitting can be completed.

[0137] The "quarter-month" energy consumption splitting objective function is:

[0138]

[0139] Among them, m irepresents the energy consumption of the i-th enterprise in a certain quarter, β j represents the energy consumption splitting coefficient of the j-th month, n ij represents the energy consumption of the i-th enterprise in the j-th month.

[0140] Through the above "quarter - month" Lasso regression objective function, β can be obtained, and then the splitting can be completed.

[0141] The calculation method of the monthly energy consumption splitting weight is as follows:

[0142] z ij = α i ×β ij

[0143] Among them, z ij represents the energy consumption splitting coefficient of the j-th month in the i-th quarter finally, α i represents the energy consumption splitting coefficient of the i-th quarter, β ij represents the energy consumption splitting coefficient of the j-th month in the i-th quarter. The monthly energy consumption splitting weights of each sub - industry are set according to the actual situation.

[0144] The Lasso regression model optimizes the "year - quarter" energy consumption splitting objective function and the "quarter - month" energy consumption splitting objective function respectively, and obtains α ij and β ij of enterprise i. Then z ij can be calculated to split the enterprise energy consumption data.

[0145] As one or more embodiments, in step S4: the missing data caused by the introduction of new equipment or new branches and the change of data sources is complemented by the LSTM - BPNN model optimized by the particle swarm optimization algorithm (PSO). The specific steps include:

[0146] Step 401: Construct an LSTM - BPNN model, and the model structure is as Figure 2 shown. Specifically:

[0147] BPNN usually consists of 1 input layer, multiple hidden layers and 1 output layer. Each neural network layer is composed of 1 or more neurons. The input layer receives the input feature vector, and each neuron corresponds to one dimension of the feature vector. The output layer outputs the result or the result vector, and each neuron corresponds to one dimension of the result or the result vector. The layer between the input layer and the output layer is the hidden layer. Neurons in adjacent layers are connected by adjustable weights. The input feature vector x becomes the output after the feed - forward calculation of the input layer, hidden layer and output layer

[0148]

[0149] Among them, W outand W hidden are the connection weights of the output layer and the hidden layer, respectively, and b out and b hidden are the firing thresholds of the neurons in the output layer and the hidden layer, respectively. is the activation function of the hidden layer, and h is the output of the hidden layer.

[0150] LSTM adds a cell state to the traditional RNN to preserve previous information, and there are three control gates (forget gate, input gate, output gate) in the cell unit to control the transmission and flow of information.

[0151] The forget gate receives the input x t at the current time step, and the hidden state h t-1 at the previous time step, and calculates the proportion f t-1 of the cell state c t from the previous time step that is retained to the current time step, and the calculation method is as follows:

[0152] f t = σ(W f · [h t-1 , x t + b f )

[0153] The input gate receives the input x t at the current time step, and the hidden state h t-1 at the previous time step, and calculates the proportion i t of the new information contained in x t-1 and h t that needs to be saved to the cell state, and the calculation method is as follows:

[0154] i t = σ(W i · [h t-1 , x t + b i )

[0155] The input x t at the current time step and the hidden state h t-1 at the previous time step generate a candidate cell state after being activated by the tanh function The calculation method is as follows:

[0156]

[0157] The cell state c t at the current time step is jointly determined by the cell state c t-1 at the previous time step, the candidate state at the current time step, f t and i t The calculation method is as follows:

[0158]

[0159] The output gate receives the input x at the current time t and the hidden state h at the previous time t-1 , and calculates the cell state c t and outputs it to the output h at the current time t in the ratio o t , and the calculation method is as follows:

[0160] o t = σ(W o ·[h t-1 , x t + b o )

[0161] The final hidden state h t is jointly determined by the output gate o t and c t , and the calculation method is as follows:

[0162]

[0163] In the above formula, σ(·) is the Sigmoid activation function, W f , W i and W o are the weight matrices of the forget gate, input gate, and output gate respectively, W c is the weight matrix for calculating the candidate state, b f , b i and b o are the excitation threshold vectors of the forget gate, input gate, and output gate respectively, b c is the excitation threshold vector for calculating the candidate state.

[0164] Step 402: Optimize the model parameters through the Particle Swarm Optimization algorithm (PSO). The PSO process is as Figure 3 shown, specifically:

[0165] The algorithm first assigns initial random positions and initial random velocities to all particles in the space. Then, according to the velocity of each particle, the known optimal global position in the problem space, and the known optimal position of the particle, the position of each particle is advanced in turn. As the calculation progresses, by exploring and exploiting the known favorable positions in the search space, the particles gather or aggregate around one or more optimal points.

[0166] The update formula for the velocity of the d-th dimension of particle i at each iteration is:

[0167]

[0168] The update formula for the position of the d-th dimension of particle i is:

[0169]

[0170] Among them, is the d-th dimensional component of the flight speed vector of particle i in the k-th iteration, is the d-th dimensional component of the position vector of particle i in the k-th iteration, c 1 , c 2 are acceleration constants, adjusting the maximum learning step size, r 1 , r 2 are two random functions, with a value range of [0, 1], to increase the search randomness. w is the inertia weight, a non-negative number, adjusting the search range of the solution space.

[0171] By using the pre-trained LSTM-BPNN and adopting a step-by-step prediction method, multiple data points are completed, ensuring that each step of the prediction makes the best use of the latest completed data points, thereby improving the prediction accuracy. Assume that the missing data in the input sequence data D = {d 1 , d 2 , …, d n} are d i and d j and i < j, f lstm-bpnn (·) is the model prediction function, then d i can be expressed as d i = f lstm-bpnn (d 1 , d 2 , …, d i-1 ), and similarly, d j = f lstm-bpnn (d 1 , d 2 , …, d j-1 ).

[0172] In this embodiment, by integrating multiple modules such as data processing, feature extraction, deep learning, and optimization algorithms, an efficient and accurate monthly energy consumption analysis and prediction tool is realized; it can better adapt to complex energy consumption distributions and change patterns, effectively handle scenarios with multiple branches and multiple devices, improve the accuracy and generalization ability of the splitting process; ensure the accuracy of data completion, make the entire monthly energy consumption analysis process more systematic and controllable, facilitate applications in different fields, and provide strong support for energy consumption optimization and management in industrial, commercial, and residential fields.

[0173] Embodiment 2

[0174] This embodiment discloses a monthly energy consumption splitting and completion system.

[0175] A monthly energy consumption splitting and completion system includes:

[0176] A data acquisition module, configured to: acquire the energy consumption data and power consumption data of an enterprise group in a target industry as original data;

[0177] A data preprocessing module, configured to: identify and mark abnormal data in the original data, and repair the marked abnormal data;

[0178] A data dimensionality reduction module, configured to: perform dimensionality reduction processing on the repaired data, and retain a set number of features with the most information;

[0179] A data splitting module, configured to: build an energy consumption splitting model based on the Lasso regression model and perform pre-training, input the obtained features into the pre-trained energy consumption splitting model, perform annual-quarter-month energy consumption splitting on the energy consumption, and obtain the monthly energy consumption of different branches;

[0180] A data completion module, configured to: build an LSTM-BPNN model as a missing data prediction model, and complete the missing data in the monthly energy consumption of different branches based on the missing data prediction model.

[0181] It can be understood that in the above data acquisition module, various sensor data sources are integrated to ensure consistent data formats; data quality control is implemented, including removing outliers and noise, correcting or removing detected outliers, and ensuring the reliability and real-time nature of data transmission;

[0182] In the above data dimensionality reduction module, through principal component analysis, retain the features with the most information as the model input;

[0183] In the above data splitting module, input the data that has been preprocessed and feature-extracted into the energy consumption splitting model, and use a linear regression model for model training and complete the splitting;

[0184] In the above data completion module, complete the training of the model through historical energy consumption data and complete the missing data caused by the introduction of new devices or new branches and changes in data sources, etc.

[0185] Those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. The present invention is not limited to any specific combination of hardware and software.

[0186] Although the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, they are not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made without creative efforts based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for splitting and completing monthly energy consumption, characterized in that: The following steps are involved: Obtain energy consumption data and electricity usage data of the target industry enterprise group as raw data; Identify and mark abnormal data in the original data, and repair the marked abnormal data; Perform dimensionality reduction on the repaired data to retain a set number of features with the most information; An energy consumption splitting model is built based on the Lasso regression model and pre-trained. The obtained features are input into the pre-trained energy consumption splitting model to split the energy consumption into year-season-month and obtain the monthly energy consumption of different branches. An LSTM-BPNN model was built as a missing data prediction model, and the missing data in the monthly energy consumption of different branches were completed based on the missing data prediction model.

2. The monthly energy consumption splitting and completion method according to claim 1, characterized in that: Identify and mark abnormal data in the original data, and repair the marked abnormal data, including: Use K-means algorithm to identify and mark abnormal data in the original data; The marked abnormal data is repaired by the K nearest neighbor algorithm.

3. The monthly energy consumption splitting and completion method according to claim 2, characterized in that: The K-means algorithm is used to identify and mark abnormal data in the original data, including: Define K data points as initial cluster centers; Calculate the Euclidean distance from the data points in the original data to the initial cluster center, and divide the original data into K clusters based on similar data points; Recalculate the center values ​​of the K clusters and use them as new cluster centers for re-clustering. Iterate until the clustering criterion function converges to obtain the clustering results of the K clusters. The data points whose distance from the corresponding cluster center in the clustering results of K clusters is greater than the set value are marked as abnormal data; Alternatively, the marked abnormal data can be repaired using the K nearest neighbor algorithm, including: Eliminate incomplete data from the original data and construct a complete data set matrix; Calculate the Euclidean distance between the abnormal data and each data in the complete data set matrix; Based on the calculated Euclidean distance, the K data closest to the abnormal data are selected; Calculate the nearest neighbor weights of the K data that are nearest neighbors to the abnormal data, and use the nearest neighbor weights to perform weighted sum calculation on the corresponding data in the K data as the estimated abnormal data.

4. The monthly energy consumption splitting and completion method according to claim 1, characterized in that: The repaired data is processed by principal component analysis (PCA) to reduce the dimensionality and retain the set number of features with the most information, including: Based on the repaired data, construct a data matrix represented by variables and sample numbers; Perform Max-Min normalization transformation on the data matrix to obtain a normalized matrix; Solve for the correlation coefficient matrix of the standardized matrix; Calculate the characteristic equation of the correlation coefficient matrix and obtain p eigenvalues; Determine the threshold value for the cumulative contribution rate of the principal component, and select the first N principal components that make the cumulative contribution rate of the principal component reach the threshold value for the cumulative contribution rate of the principal component from the p eigenvalues ​​as the most informative features.

5. The monthly energy consumption splitting and completion method according to claim 1, characterized in that: The energy consumption split model includes: The annual-seasonal energy consumption split objective function is: Among them, y i Represents the yth i Annual energy consumption of enterprises, α j represents the energy consumption split coefficient for the jth quarter, x ij represents the energy consumption of the i-th enterprise in the j-th quarter; The objective function of quarterly-monthly energy consumption splitting is: Among them, m i represents the energy consumption of the i-th enterprise in a certain quarter, β j represents the energy consumption split coefficient of the jth month, n ij represents the energy consumption of the i-th enterprise in the j-th month; The monthly energy consumption split weight calculation method is: z ij =α i ×β ij ; Among them, z ij Represents the energy consumption split coefficient of the jth month in the i-th quarter, α i represents the energy consumption split coefficient for the i-th quarter, β ij Represents the energy consumption split coefficient of the jth month in the i-th quarter.

6. The monthly energy consumption splitting and completion method according to claim 1, characterized in that: Pre-train the energy consumption splitting model, specifically: Calculate the annual / monthly energy consumption data of the target industry enterprise group and the proportion of enterprises with missing energy consumption in the corresponding electricity consumption data to the total number of enterprises in the enterprise group: When the ratio is greater than 40%, the electricity consumption data of the enterprise group is used as the training data of the energy consumption splitting model; When the ratio is less than 40%, the monthly energy consumption data of enterprises with complete monthly energy consumption data are used as training data for the energy consumption splitting model.

7. The monthly energy consumption splitting and completion method according to claim 1, characterized in that: Build an LSTM-BPNN model as a missing data prediction model, and fill in the missing data in the monthly energy consumption of different branches based on the missing data prediction model, including: By combining the long short-term memory network LSTM and the back propagation neural network BPNN, an LSTM-BPNN model is built; The LSTM-BPNN model is trained through historical energy consumption data, and the LSTM-BPNN model parameters are optimized through the particle swarm optimization algorithm PSO; The historical monthly energy consumption of branches with missing data is input into the trained and optimized LSTM-BPNN model to obtain the predicted missing data, thereby completing the missing data in the monthly energy consumption.

8. The method for splitting and completing monthly energy consumption according to claim 7, characterized in that: The final hidden state h of the LSTM t , calculated as follows: the t =σ(W o ·[h t-1 ,x t ]+b o ); i t =σ(W i ·[h t-1 ,x t ]+b i ); f t =σ(W f ·[h t-1 ,x t ]+b f ); Among them, t is the output gate, c t is the cell state at the current moment, c t-1 is the cell state at the previous moment, h t-1 is the hidden state of the previous moment, x t is the input at the current moment, is the candidate cell state at the current moment, σ(·) is the Sigmoid activation function, W f , W i and W o are the weight matrices of the forget gate, input gate, and output gate, respectively, W c To calculate the weight matrix of the candidate state, b f 、b i and b o are the activation threshold vectors of the forget gate, input gate, and output gate, respectively, and b c To calculate the excitation threshold vector of the candidate state; Assume that the feature vector input to the BPNN is x, and after feedforward calculations of the input layer, hidden layer, and output layer, the output of the BPNN is obtained. Among them, W out and W hidden are the connection weights of the output layer and the hidden layer, b out and b hidden are the firing thresholds of the output layer and hidden layer neurons, respectively. is the activation function of the hidden layer, and h is the output of the hidden layer.

9. The monthly energy consumption splitting and completion method according to claim 7, characterized in that: The LSTM-BPNN model parameters are optimized by the particle swarm optimization algorithm PSO, specifically: First, all particles in space are assigned initial random positions and initial random velocities; Then the position of each particle is advanced in turn according to its velocity, the best global position known in the problem space, and the best known position of the particle; As the computation progresses, particles aggregate or cluster around one or more optimal points by exploring and exploiting known favorable locations in the search space.

10. A monthly energy consumption splitting and completion system, characterized in that: include: The data acquisition module is configured to: acquire energy consumption data and electricity usage data of a target industry enterprise group as raw data; The data preprocessing module is configured to: identify and mark abnormal data in the original data, and repair the marked abnormal data; The data dimension reduction module is configured to: perform dimension reduction processing on the repaired data and retain a set number of features with the most information content; The data splitting module is configured to: build an energy consumption splitting model based on the Lasso regression model and perform pre-training, input the obtained features into the pre-trained energy consumption splitting model, split the energy consumption into year-season-month, and obtain the monthly energy consumption of different branches; The data completion module is configured to: build an LSTM-BPNN model as a missing data prediction model, and complete the missing data in the monthly energy consumption of different branches based on the missing data prediction model.