Enterprise monthly energy consumption prediction method and system based on subdivided industry fingerprints

By segmenting enterprises in industries and using industry indicator data to train energy consumption prediction models, the problem of limited effects of energy consumption prediction models in the existing technology is solved, and the accuracy and robustness of energy consumption prediction are improved.

CN120106295APending Publication Date: 2025-06-06STATE GRID SHANDONG ELECTRIC POWER CO +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510184738.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the enterprise energy consumption prediction of the existing technology, there is a lack of consideration of the different energy consumption of different industries, resulting in limited effect of trained energy consumption prediction models, low accuracy and low robustness.

Method used

By dividing the enterprises into sectors, using industry indicator data of the enterprise group of the same sector, the energy consumption prediction model of the segmented industry is trained, and the energy consumption of each enterprise in the segmented industry is predicted.

Benefits of technology

It improves the accuracy of enterprise energy consumption prediction and the robustness of energy consumption prediction models, increases the sample size of training data, and improves the training effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106295A_ABST
    Figure CN120106295A_ABST
Patent Text Reader

Abstract

The invention discloses an enterprise monthly energy consumption prediction method and system based on subdivided industry fingerprints. The method comprises the steps of obtaining industry index data of each enterprise in a region; according to the industry index data, subdivided industry division is carried out on enterprises in the region, and multiple groups of subdivided industry enterprise groups are obtained; complementing the missing monthly energy consumption data in the industry index data of each enterprise to obtain the complemented industry data of each enterprise; for each group of subdivided industry enterprise group, training the energy consumption prediction model corresponding to the subdivided industry by using the industry complemented data of all enterprises in the subdivided industry enterprise group, and obtaining the trained energy consumption prediction model of the subdivided industry after the training is completed; and predicting the energy consumption of each enterprise in the subdivided industry by using the trained energy consumption prediction model of the subdivided industry. The accuracy of enterprise monthly energy consumption prediction is improved, and the robustness and generalization ability of the energy consumption prediction model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing technology, and in particular to a method and system for predicting monthly energy consumption of an enterprise based on fingerprints of subdivided industries. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] During the enterprise generation process, it is necessary to predict the energy consumption of the next node of the enterprise, and then effectively allocate energy based on the predicted energy consumption.

[0004] In the related art, there is a method of using a trained energy consumption prediction model to predict the energy consumption of the next node of an enterprise. However, this method uses the enterprise's own energy consumption data or the energy consumption data of all enterprises to train the energy consumption prediction model. The training data is relatively small, and the difference in energy consumption between different industries is not considered. As a result, the trained energy consumption prediction model has limited effect, the accuracy of energy consumption prediction for the next node of the enterprise is not high, and the robustness is low. Summary of the invention

[0005] In order to solve the above problems, the present invention proposes a method and system for monthly energy consumption prediction of enterprises based on subdivided industry fingerprints. By dividing enterprises into subdivided industries, and then using the industry indicator data of all enterprises in the same subdivided industry enterprise group, the energy consumption prediction model of the subdivided industry is trained. The trained energy consumption prediction model of the subdivided industry can predict the energy consumption of each enterprise in the subdivided industry, thereby improving the accuracy of the enterprise energy consumption prediction and the robustness of the energy consumption prediction model.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] First, a method for predicting the monthly energy consumption of enterprises based on fingerprints of subdivided industries is proposed, including:

[0008] Obtain industry indicator data for each enterprise in the region;

[0009] According to the industry indicator data, the enterprises in the region are divided into sub-industries to obtain multiple groups of sub-industry enterprise groups;

[0010] Complete the missing monthly energy consumption data in the industry indicator data of each enterprise to obtain the completed industry data of each enterprise;

[0011] For each group of subdivided industry enterprise groups, the energy consumption prediction model corresponding to the subdivided industry is trained using the industry-completed data of all enterprises in the subdivided industry enterprise group. After the training is completed, the trained energy consumption prediction model for the subdivided industry is obtained;

[0012] The energy consumption prediction model trained for the subdivided industry is used to predict the energy consumption of each enterprise in the subdivided industry.

[0013] Furthermore, the process of completing the missing monthly energy consumption data in the enterprise's industry indicator data includes:

[0014] When all monthly energy consumption data of an enterprise are missing, the enterprise's monthly energy consumption splitting model is used to split the enterprise's annual energy consumption data to obtain the enterprise's monthly energy consumption data for each month;

[0015] When part of the enterprise's monthly energy consumption data is missing, determine the missing monthly energy consumption percentage, and use the missing monthly energy consumption percentage and the enterprise's annual energy consumption data to obtain the missing monthly energy consumption data of the enterprise;

[0016] The obtained monthly energy consumption data is used to complete the missing monthly energy consumption data in the enterprise's industry indicator data to obtain the enterprise's industry completed data.

[0017] Furthermore, the enterprise monthly energy consumption split model is: the enterprise's monthly energy consumption data is equal to the enterprise's monthly energy consumption split coefficient multiplied by the enterprise's annual energy consumption data;

[0018] The proportion of the enterprise's electricity consumption or output in the missing months to the annual electricity consumption or output is taken as the proportion of the enterprise's missing monthly energy consumption.

[0019] Furthermore, the monthly energy consumption splitting coefficient of the enterprise is obtained by training the monthly energy consumption splitting model of the enterprise with the goal of minimizing the energy consumption splitting error, using the energy consumption data of enterprises in the enterprise group of the subdivided industry to which the enterprise belongs that have known annual energy consumption data, quarterly energy consumption data and monthly energy consumption data.

[0020] Furthermore, when the enterprise lacks monthly energy consumption data but there is electricity or output data in the month, the proportion of the electricity or output data to the annual electricity or output is directly used as the monthly energy consumption proportion of that month; when the enterprise lacks monthly energy consumption data but there is no electricity or output data in the month, the electricity or output data of the existing months is used to predict the electricity or output data for that month, and the predicted proportion of the electricity or output data for that month to the annual electricity or output is used as the monthly energy consumption proportion of that month.

[0021] Furthermore, the energy consumption prediction model takes the industry-completed data of enterprises in the segmented industry enterprise group as input and the energy consumption prediction results as output, and is constructed using the lightGBM model.

[0022] Secondly, a monthly energy consumption forecasting system for enterprises based on subdivided industry fingerprints is proposed, including:

[0023] A data acquisition unit, used to acquire industry indicator data of each enterprise in the region;

[0024] The subdivided industry division unit is used to divide the enterprises in the region into subdivided industries according to the industry indicator data, and obtain multiple groups of subdivided industry enterprise groups;

[0025] The missing data completion unit is used to complete the missing monthly energy consumption data in the industry indicator data of each enterprise, and obtain the completed industry data of each enterprise;

[0026] The energy consumption prediction unit is used to train the energy consumption prediction model corresponding to each group of subdivided industry enterprise groups using the industry-completed data of all enterprises in the subdivided industry enterprise group. After the training is completed, the trained energy consumption prediction model of the subdivided industry is obtained; the energy consumption of each enterprise in the subdivided industry is predicted using the trained energy consumption prediction model of the subdivided industry.

[0027] In a third aspect, a computer device is provided, the device comprising:

[0028] a processor adapted to execute a computer program;

[0029] A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the method for predicting monthly energy consumption of an enterprise based on fingerprints of subdivided industries proposed in the first aspect is implemented.

[0030] In a fourth aspect, a computer-readable storage medium is proposed, wherein the computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the enterprise monthly energy consumption forecasting method based on segmented industry fingerprints proposed in the first aspect.

[0031] In a fifth aspect, a computer program product is proposed, which includes a computer program. When the computer program is executed by a processor, it implements the enterprise monthly energy consumption forecasting method based on segmented industry fingerprints proposed in the first aspect.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] The present invention provides a method and system for predicting monthly energy consumption of enterprises based on fingerprints of subdivided industries. The method divides enterprises into subdivided industries and completes the missing monthly energy consumption data in the industry indicator data. Then, for each group of subdivided industry enterprise groups, the industry-completed data of all enterprises in the subdivided industry enterprise group are used to train the energy consumption prediction model corresponding to the subdivided industry to obtain the trained energy consumption prediction model for the subdivided industry; the trained energy consumption prediction model can predict energy consumption for each enterprise in the subdivided industry; the robustness of the energy consumption prediction model is improved; and the energy consumption prediction model is trained using the data of all enterprises in the industry, which increases the sample size of the training data, further improves the training effect of the model, and improves the prediction accuracy of the enterprise's monthly energy consumption.

[0034] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The drawings in the specification, which constitute a part of the present application, are used to provide further understanding of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.

[0036] Figure 1 A flowchart of a method for predicting monthly energy consumption of an enterprise based on fingerprints of subdivided industries disclosed in an embodiment;

[0037] Figure 2 A subdivided industry tree disclosed in the embodiment;

[0038] Figure 3 A particle swarm solution flow chart disclosed in the embodiment;

[0039] Figure 4 This is a process disclosed in the embodiment of using the industry indicator data of new enterprises to update the energy consumption prediction model trained for the subdivided industry. DETAILED DESCRIPTION

[0040] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0041] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0042] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0043] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0044] Example 1

[0045] The core function of the energy consumption detection system is to monitor energy consumption data in real time and compare it with past data and databases, thereby effectively helping enterprises achieve the strategic goal of reducing costs and increasing efficiency. The widespread application of this system provides enterprises with real-time and accurate energy consumption information, enabling them to adjust their operating strategies more flexibly and further optimize resource utilization efficiency.

[0046] In recent years, the development trend of energy consumption detection systems has shown a positive integration of new technologies. The introduction of advanced technologies such as cloud computing and deep learning has significantly improved system performance. Deep learning methods such as convolutional neural networks and recurrent neural networks are widely used in model training, making the system more intelligent and able to more accurately analyze and predict energy consumption trends. Such technological upgrades not only improve the efficiency of energy consumption detection systems, but also provide better and more personalized services for different industries.

[0047] However, due to the differences in production processes and data involved in different industries, existing technologies have limitations when performing energy consumption detection tasks, and their robustness and generalization performance are relatively low. Therefore, this embodiment proposes a method for predicting monthly energy consumption of enterprises based on segmented industry fingerprints. Our model construction mainly focuses on two directions: on the one hand, it is to carefully combine the characteristics of the enterprise's production process flow to improve the prediction accuracy of the model; on the other hand, it is to reduce the human and material burden that may be caused by the over-refined production process, improve the practicality of the model, and lower the promotion threshold.

[0048] In order to strike a balance between refinement and scalability, the model of this embodiment deeply integrates the characteristics of production processes in different industries, such as production equipment, product types, and production capacity. Through expert experience, industry research, and government releases, the industry production characteristics are extracted and precipitated, and an industry index library is constructed. Based on the similarity of production process characteristics, the major industry categories are clustered layer by layer for subdivision, and combined with the business perspective, guided by the prediction results, a subdivided industry tree is innovatively constructed to extract subdivided industry fingerprints.

[0049] This example uses algorithms such as lightGBM to build an energy consumption forecasting model for subdivided industries, and at the same time builds a monthly energy consumption split model through an optimization algorithm, ultimately forming a "one enterprise group one model" with high accuracy and strong scalability. This model can be widely used in multiple scenarios, including filling gaps in enterprise energy consumption data, verification and prediction, and other needs.

[0050] In this embodiment, the disclosed method for predicting monthly energy consumption of enterprises based on subdivided industry fingerprints is as follows: Figure 1-Figure 4 As shown, including:

[0051] S1: Obtain industry indicator data for each enterprise in the region.

[0052] Among them, industry indicator data include enterprise basic information data, enterprise production equipment information data, energy consumption data, output data, product data and industry-specific attribute data.

[0053] S2: Based on the industry indicator data, the enterprises in the region are divided into sub-industries to obtain multiple groups of sub-industry enterprise groups.

[0054] S3: Complete the missing monthly energy consumption data in the industry indicator data of each enterprise to obtain the completed industry data of each enterprise.

[0055] S4: For each group of subdivided industry enterprise groups, use the industry-completed data of all enterprises in the subdivided industry enterprise group to train the energy consumption prediction model corresponding to the subdivided industry. After the training is completed, the trained energy consumption prediction model of the subdivided industry is obtained; use the trained energy consumption prediction model of the subdivided industry to predict the energy consumption of each enterprise in the subdivided industry.

[0056] The process of completing the missing monthly energy consumption data in the enterprise's industry indicator data in this embodiment is as follows:

[0057] When all monthly energy consumption data of an enterprise are missing, the enterprise's monthly energy consumption splitting model is used to split the enterprise's annual energy consumption data to obtain the enterprise's monthly energy consumption data for each month;

[0058] When part of the enterprise's monthly energy consumption data is missing, determine the missing monthly energy consumption percentage, and use the missing monthly energy consumption percentage and the enterprise's annual energy consumption data to obtain the missing monthly energy consumption data of the enterprise;

[0059] The obtained monthly energy consumption data is used to complete the missing monthly energy consumption data in the enterprise's industry indicator data to obtain the enterprise's industry completed data.

[0060] Among them, the enterprise monthly energy consumption split model is: the enterprise's monthly energy consumption data is equal to the enterprise's monthly energy consumption split coefficient multiplied by the enterprise's annual energy consumption data;

[0061] The proportion of the enterprise's electricity consumption or output in the missing months to the annual electricity consumption or output is taken as the proportion of the enterprise's missing monthly energy consumption.

[0062] The monthly energy consumption splitting coefficient of an enterprise is obtained by training the enterprise's monthly energy consumption splitting model with the goal of minimizing the quarterly energy consumption splitting error and the monthly energy consumption splitting error, using the energy consumption data of enterprises in the enterprise group of the subdivided industry to which the enterprise belongs that have known annual energy consumption data, quarterly energy consumption data and monthly energy consumption data.

[0063] When an enterprise lacks monthly energy consumption data but has electricity or output data in a month, the proportion of such electricity or output data in the annual electricity or output shall be directly used as the monthly energy consumption proportion for that month; when an enterprise lacks monthly energy consumption data but does not have electricity or output data in a month, the electricity or output data of the existing months shall be used to predict the electricity or output data for that month, and the proportion of the predicted electricity or output data in the annual electricity or output shall be used as the monthly energy consumption proportion for that month.

[0064] The energy consumption prediction model of this embodiment takes the industry-completed data of enterprises in the subdivided industry enterprise group as input and the energy consumption prediction result as output, and is constructed using the lightGBM model.

[0065] Taking a province as a region, and taking the province including 16 high-energy and high-dose industries, which can be further divided into 31 secondary sub-industries as an example, the enterprise monthly energy consumption prediction method based on sub-industry fingerprints disclosed in this embodiment is described in detail.

[0066] Combining expert experience, industry research, government releases, etc., we have established secondary subdivided industry indicator sets from the dimensions of production equipment, raw material input, and capacity output to form an industry indicator library of "Energy Consumption from Electricity". Taking the ceramic industry as an example, the ceramic industry indicator library is mainly composed of two categories (basic indicators and expanded indicators) and six indicator dimensions (production equipment, energy, output, products, and comprehensive).

[0067] Build a segmented industry tree.

[0068] With 31 secondary sub-industries as different branches, two or more enterprise groups with similar production process characteristics, production levels, and energy efficiency levels are divided out as branches, and finally a sub-industry tree is formed. The enterprise groups in the sub-industries at the bottom of the sub-industry tree can use the same energy consumption prediction model, which has the advantages of easy integrated learning and good interpretability.

[0069] The original data set of the subdivided industry tree is a summary table of the two high industries from January 2021 to May 2022. The fields include the city, county, district, address, production unit model and quantity, production capacity, output, energy consumption, coal consumption, output value, industrial added value, revenue, number of employees and industry-specific attributes, such as the water absorption rate of ceramics and whether asphalt waterproofing materials are tire-like or not.

[0070] In this embodiment, X={x 1 ,x 2 ,…,x N} represents the data of N industries, where Taking the calcium industry as an example, the input clustering model data mainly includes the following features: equipment type, equipment diameter, equipment length, equipment quantity, energy type, comprehensive energy consumption, electricity consumption, coal consumption, production capacity, output, energy efficiency level, capacity utilization rate, comprehensive energy consumption per unit product, fuel type, etc. C = {c 1 ,c 2 ,…,c K} are the expected K cluster centers. And, let Z = [z Nv ] M×K , z Nv ∈{0,1} represents the data sample x N Whether it belongs to the vth cluster, where v = 1, ..., K. Therefore, the objective function of this embodiment can be expressed as:

[0071]

[0072] To achieve this goal, we select the first center c from the dataset X in a uniformly random manner. 1 , and repeat to select the next center c i =x′∈X is:

[0073]

[0074] Until K centers are initialized. Here, D(x) is the shortest distance from x to the nearest center that has been selected. Then, the solver iteratively updates the cluster centers and memberships respectively by the following equations. Then, the solver iteratively updates the cluster centers and memberships formulated by the following equations respectively by the following equations:

[0075]

[0076] Among them, ∥x N -c v ∥ is x N and c v The Euclidean distance between .

[0077] Then, pruning optimization is performed, and grid search and other means are used to optimize the relevant clustering parameters, including the clustering parameter K and the initial centroid C, so that the verification index reaches the optimal value and the clustering result is generated. Based on the clustering results, mathematical statistics such as mean, median, mode, and similarity measurement methods such as Euclidean distance and adjusted cosine similarity are used, and the opinions of business experts are combined to judge the rationality of the industry tree constructed by the clustering results, prune the industry tree or further split it, and determine the optimal industry tree structure. Figure 1 This is a schematic diagram of the industry segmentation tree.

[0078] Among them, the industry tree structure with reasonable clustering results, small similarity between cluster centers and recommended by experts is selected as the optimal industry tree structure.

[0079] Specifically, statistical methods such as mean, median, and mode are used to analyze the central features of each cluster after clustering; if the feature differences of objects within a cluster are small, and the feature differences between clusters are large, it means that the clustering results are more reasonable.

[0080] The geometric distance between cluster centers is measured by Euclidean distance. The farther the distance, the greater the difference between clusters. A reasonable industry tree should reflect the obvious differences between industries. Clusters that are too close may need to be merged or readjusted.

[0081] The similarity between industries is measured by calculating the cosine similarity of the cluster center vectors; low similarity indicates that the industries are more independent and can be used as a reasonable basis for industry division.

[0082] Although statistics and similarity measurement methods provide a mathematical basis for the optimal industry tree structure, the rationality of industry division still needs to be combined with actual business scenarios and expert experience. Expert opinions can help correct the limitations of statistical models and ensure that the industry tree conforms to business logic.

[0083] Step (3): Build industry segmentation nodes

[0084] Each leaf node on the subdivided industry tree mainly contains "one fingerprint and two models", namely the subdivided industry fingerprint, the monthly energy consumption splitting and completion model, and the subdivided industry energy consumption forecasting model. Among them, the subdivided industry fingerprint is used to identify the subdivided industry to which the enterprise belongs, which can characterize the key production characteristics of the industry, and is concretely represented as a collection of indicators and their value ranges. The monthly energy consumption splitting and completion model converts the low-frequency annual energy consumption data into the monthly energy consumption data of the product, providing a data basis for subsequent high-frequency energy consumption forecasts. The subdivided industry energy consumption forecasting model mainly uses the correlation between historical energy consumption, electricity, production process, production capacity and current energy consumption to analyze and calculate future trends.

[0085] Generate segmented industry fingerprints: Based on the results of the industry tree construction, extract the common characteristics of the enterprise group in the segmented industry as the industry fingerprint, which is used to match the segmented industry to which the newly added enterprise samples belong, and summarize and build a segmented industry fingerprint library, in which the common characteristics of the enterprise group are extracted from the industry indicator data.

[0086] Construct a monthly energy consumption splitting and completion model: According to the missing monthly energy consumption data of the enterprise group in the subdivided industry, a monthly energy consumption splitting and completion model is constructed for each subdivided industry in two ways to complete the energy consumption data of the enterprise. First, when there is a lot of missing monthly energy consumption data, the monthly energy consumption splitting law is constructed based on electricity or output, the monthly energy consumption proportion of the enterprise is determined, and the monthly energy consumption proportion is used to obtain the missing monthly energy consumption data; second, based on the monthly energy consumption data and annual energy consumption data of enterprises in the subdivided industry, a "year-season-month" two-step energy consumption splitting model based on historical laws is constructed as a monthly energy consumption splitting model, and the monthly energy consumption splitting model is used to split the annual energy consumption data into monthly energy consumption data. Finally, the test samples are used to verify the monthly energy consumption splitting effect of subdivided industries at all levels.

[0087] Among them, the monthly energy consumption splitting coefficient of the enterprise is obtained by training the enterprise's monthly energy consumption splitting model with the goal of minimizing the quarterly energy consumption splitting error and the monthly energy consumption splitting error, using the energy consumption data of enterprises in the enterprise group of the subdivided industry to which the enterprise belongs that have known annual energy consumption data, quarterly energy consumption data and monthly energy consumption data.

[0088] Specifically, the training data for training the enterprise monthly energy consumption splitting model includes the monthly energy consumption data of enterprises under a leaf node of an industry tree of a certain industry and the electricity consumption or output data of the corresponding month.

[0089] And calculate the ratio of the number of enterprises with missing energy consumption to the total number of enterprises under the leaf node. If the ratio is greater than 30%, the electricity consumption data of the enterprises under the leaf node is used as the training data of the monthly energy consumption split model. If the ratio is less than 30%, the monthly energy consumption data of the enterprises with complete monthly energy consumption data is directly used as the training data of the monthly energy consumption split model.

[0090] The “year-season” energy consumption split model is:

[0091] y 季IN =α Il v 年 N

[0092] The “season-month” energy consumption split model is:

[0093] Z 月hIN =b hIN y 季IN

[0094] Therefore, the monthly energy consumption split model is:

[0095] Z 月hIN =b hIN α IN v 年N

[0096] Among them, v 年N represents the annual energy consumption data of the Nth enterprise, which is the known energy consumption data and is used to decompose into quarterly energy consumption; α IN represents the proportional coefficient of the Nth enterprise in the Ith quarter, I = 1, 2, 3 or 4; y 季IN represents the energy consumption data of the Nth enterprise in the Ith quarter; Z 月hIN represents the energy consumption data of the hth month in the Ith quarter of the Nth enterprise, b hIN represents the proportional coefficient of the hth month in the Ith quarter of the Nth enterprise, h = 1, 2 or 3, b hIN α IN Represents the monthly energy consumption split coefficient of the Nth enterprise.

[0097] The objective function with the goal of minimizing the monthly splitting error includes the "year-season" objective function and the "season-month" objective function, where:

[0098] The “year-season” objective function is:

[0099] min∑(|α 1N v 年N -y 1N |+|α 2Nl ν 年N -y 2N |+|α 3N ν 年N -y 3N |+|α 4N v 年N -y 4N |)

[0100] In the above formula, α 1N , α 2N , α 3N , α 4N represents the proportional coefficients of different quarters, which are optimized by particle swarm optimization to minimize the error in the objective function; v 年N represents the annual energy consumption data of the Nth enterprise, which is the known energy consumption data and is used to decompose into quarterly energy consumption; 1N ,y 2N ,y 3N and 4N Corresponding to the actual energy consumption data of the Nth enterprise in four quarters.

[0101] The “quarter-month” objective function is:

[0102] min∑(|b1IN y 季IN -Z 1IN |+|b 2IN y 季IN -Z 2IN |+|b 3IN y 季IN -Z 3IN |)

[0103] In the above formula, b 1IN 、b 2IN 、b 3IN represents the proportional coefficients of different months in a quarter, which are optimized by particle swarm optimization to minimize the error in the objective function; 季IN represents the energy consumption data of the Nth enterprise in the Ith quarter; Z 1IN , Z 2IN and Z 3IN Corresponding to the actual energy consumption data of the Nth enterprise in the three months of the Ith quarter.

[0104] Using the complete energy consumption data in the leaf nodes, the particle swarm algorithm is used to optimize the relevant parameters of the objective function. The coefficient obtained by optimizing the "year-season" objective function is multiplied by the coefficient obtained by optimizing the corresponding "season-month" objective function to obtain the monthly energy consumption split coefficient for each month. The particle swarm algorithm is introduced below.

[0105] The particle swarm algorithm is a swarm intelligence algorithm designed by simulating the predation behavior of bird flocks. The goal is to make all particles find the optimal solution in a multi-dimensional hyper-volume. First, all particles in the space are assigned initial random positions and initial random velocities. Then, the position of each particle is advanced in turn according to its velocity, the best global position known in the problem space, and the best position known to the particle. As the calculation progresses, by exploring and utilizing known favorable positions in the search space, particles gather or aggregate around one or more optimal points. The flowchart of the particle swarm algorithm is as follows: Figure 2 The formulas involved in the calculation in this process are as follows.

[0106] The update formula of the d-dimensional velocity of particle i in each iteration is:

[0107]

[0108] in, represents the velocity of particle i in the dth dimension at the kth iteration. The velocity determines how fast the particle moves in that dimension. represents the velocity of particle i in the dth dimension at the k-1th iteration, that is, the velocity of the previous iteration. w represents the inertia weight, which represents the influence of the previous velocity on the current velocity, and is usually used to control the global and local search capabilities of the algorithm. A larger w value will encourage a wider range of exploration, and a smaller w value will make the particle focus more on local search. 1 represents the self-cognition acceleration factor, which indicates how close the particle is to its own best position pbest. It controls the degree of trust the particle has in its own experience. 1 Represents a random number between [0,1] to ensure the randomness of the algorithm so that particles do not completely follow a certain direction, thereby introducing diversity. id represents the best position found by particle i in the d-th dimension history (i.e., the best solution that the particle has ever reached). 2 Represents the social cognitive acceleration factor, which indicates the degree to which the particle approaches the global optimal position gbest in the group. It reflects the particle's ability to learn from other particles. 2 Represents a random number between [0,1] to ensure the randomness of the algorithm. d It represents the global optimal position of the group in the dth dimension, that is, the best solution found in the entire particle swarm.

[0109] The update formula of the d-dimensional position of particle i is:

[0110]

[0111] in, is the d-th dimension component of the flight velocity vector of particle i in the k-th iteration, is the d-th dimension component of the position vector of particle i at the k-th iteration, c 1 ,c 2 is the acceleration constant, which adjusts the maximum learning step size, r 1 ,r 2 are two random functions with a value range of [0,1] to increase the randomness of the search. w is the inertia weight, a non-negative number, which adjusts the search range of the solution space.

[0112] Enterprise monthly energy consumption split filling: Determine the enterprise's monthly energy consumption. When all the monthly energy consumption data of the enterprise are missing, use the monthly energy consumption split coefficient optimized by the particle swarm algorithm, that is, use the enterprise monthly energy consumption split model to split the annual energy consumption data into monthly energy consumption data.

[0113] When part of an enterprise's monthly energy consumption data is missing, determine the missing monthly energy consumption percentage, and use the missing monthly energy consumption percentage and the enterprise's annual energy consumption data to obtain the enterprise's missing monthly energy consumption data.

[0114] Among them, the proportion of the enterprise's missing electricity or output in the month to the annual electricity or output is taken as the proportion of the enterprise's missing monthly energy consumption.

[0115] In the specific implementation, the quality of electricity or output data is judged, and the specific data to be used to determine the monthly energy consumption ratio is determined. The basis for judging data quality is mainly reflected in the integrity of electricity data or output data. By comparing the quality of these two types of data, the type of data with better quality is selected to determine the monthly energy consumption ratio. This ensures that the data used in the splitting process is relatively accurate and reliable.

[0116] Combine electricity or output data with energy consumption data, use calculation methods such as percentage calculation and trend calculation to build a monthly energy consumption splitting rule, combine the opinions of business experts to determine the best monthly energy consumption splitting rule, and then determine the monthly energy consumption percentage.

[0117] According to the determined monthly energy consumption proportion, the corresponding missing values ​​are calculated for filling in to form a complete monthly energy consumption curve.

[0118] The monthly energy consumption splitting rules are constructed based on the proportional relationship and trend changes of power or output data and energy consumption data, supplemented by the professional knowledge of business experts. The optimal monthly energy consumption splitting rules ensure that the energy consumption splitting is consistent with the data logic and can also reflect the actual operation of the enterprise, thereby obtaining the best monthly energy consumption splitting plan.

[0119] According to the best monthly energy consumption splitting rules, determine the enterprise's annual or quarterly energy consumption and reasonably allocate it to each month.

[0120] It is necessary to identify which months have missing energy consumption data. Missing data may be completely missing (no energy consumption data for a certain month) or partially missing (incomplete data for a certain month).

[0121] If the enterprise has other relevant data (such as electricity and production data), these data can be used as a basis. According to the proportion and trend of these data, combined with the constructed monthly energy consumption splitting rules, the energy consumption data of the missing months can be derived.

[0122] If electricity or production data are severely missing and no direct correlation can be found, it may be necessary to supplement them with annual or quarterly total energy consumption data.

[0123] According to the monthly energy consumption splitting rules, when the missing energy consumption data is completed, if there is electricity or output data in the month when the enterprise is missing monthly energy consumption data, the proportion of the electricity or output data in the annual electricity or output is directly used as the monthly energy consumption proportion of the month; when there is no electricity or output data in the month when the enterprise is missing monthly energy consumption data, the electricity or output data of the existing months is used to predict the electricity or output data of the month, and the predicted proportion of the electricity or output data of the month in the annual electricity or output is used as the monthly energy consumption proportion of the month. After that, the missing monthly energy consumption data is completed. The specific methods are as follows:

[0124] Proportional split: If the annual energy consumption data of a company is complete, but the monthly data is partially missing, the missing months can be supplemented according to the splitting rules based on the known proportion of electricity or output. For example, if the proportion of electricity in a certain month is 20%, 20% of the annual energy consumption can be allocated to that month to supplement the missing data of that month.

[0125] Completion by trend: If the company's electricity or production data shows a certain trend (such as seasonal fluctuations), the trend calculation method can be used to reasonably allocate the annual or quarterly energy consumption to the missing months according to the trend. For example, energy consumption is usually higher in summer. If data is missing in a certain month in summer, the energy consumption data for that month can be derived by referring to the trends of surrounding months and combining the opinions of business experts.

[0126] Combined with similar enterprise data: In some cases, it may be possible to use the energy consumption data of other similar enterprises as a reference to fill in the missing data according to the monthly split rules. This method can serve as a good reference, especially when enterprises in a certain industry have similar energy consumption patterns.

[0127] Constructing a sub-industry energy consumption prediction model: Based on the enterprise group data in each sub-industry, train the sub-industry energy consumption prediction model to predict the monthly energy consumption of similar enterprises. The historical monthly energy consumption, historical monthly electricity consumption, production process, production capacity and other data of the enterprise are used as the input of the prediction model. Then, a sub-industry monthly energy consumption prediction model is constructed based on lightGBM. As a mature decision tree integration algorithm, lightGBM has the advantages of high prediction accuracy, robustness and ease of use, effective processing of missing values, and rapid parameter adjustment.

[0128] The data are enterprise energy consumption data and summary data of the two high industries from January 2021 to May 2022. The fields include enterprise monthly energy consumption, monthly electricity consumption, city, county, district, address, production equipment model and quantity, production capacity, output, energy consumption, coal consumption, output value, industrial added value, revenue, number of employees, and industry-specific attributes, such as the water absorption rate of ceramics and whether asphalt waterproofing materials are tire-free.

[0129] Data input can be divided into three steps: window sliding, dynamic and static data matching, and feature engineering.

[0130] The window sliding selection length is 3, and the enterprise energy consumption data is windowed with a step size of 1 to form the basis of the data set.

[0131] Dynamic and static data matching is to match dynamic data such as electricity consumption and output, and static data such as production capacity and number of employees according to company name and time fields on the basis of window sliding.

[0132] Feature engineering is to construct the data set for the input model based on the window sliding data after dynamic and static data matching, and then go through the steps of preprocessing missing values, duplicate values, and outlier data, constructing features based on existing data, and selecting features based on correlation and mutual information. After the data set is constructed, the relevant data is input into the LightGBM model.

[0133] The learning process of decision tree is divided into two parts:

[0134] Node level: Find the optimal split point (feature value) of a leaf node

[0135] Tree structure level: choose which leaf node to split (feature)

[0136] For continuous features, the gain can be calculated using the pre-sorting method, but each split point must be tried, and the gain will be calculated many times. In order to reduce the amount of calculation, LightGBM discretizes the continuous features with equal distances, that is, histogram statistics, and only calculates the gain once at each box in the histogram. In this way, for a single feature, the time complexity of calculating the gain is reduced from the number of different feature values. In addition, LightGBM also uses histograms for differential acceleration when splitting nodes. When looking for the optimal splitting point for category features, the many-vs-many mode is used to split the nodes, and the category features are sorted according to each category. Sort them, and then construct a histogram in this order to find the optimal splitting point.

[0137]

[0138] Among them, g N is the gradient of the Nth sample, h N is the second-order derivative of the Nth sample. The sum of the gradients of all samples in the category feature represents the first-order error accumulation of the category as a whole. It represents the sum of the second-order derivatives of all samples in the category feature, and it represents the cumulative second-order derivative of the category feature as a whole. It represents the ratio of the gradient and the second-order derivative of each category feature, which is used to measure the gain of a specific interval or node of a certain category feature. When looking for the optimal split point, LightGBM uses this ratio to evaluate the splitting effect of different categories.

[0139] The tree structure is learned by the leaf growth strategy. At each split, the leaf that can bring the greatest gain is selected for splitting. Under the same number of splits, it is obvious that leaf growth can reduce the loss function more.

[0140] The overall process of lightGBM can be summarized as follows:

[0141] For a given data set: D = {(X N ,Y N ), N = 1, 2, ..., M,}, where M is the number of samples and each sample has P features. Given the loss function L(y, f(x)), output the regression tree f(x). The specific algorithm steps are as follows:

[0142] Step 1: Initialize f 0 (x), that is:

[0143]

[0144] Step 2: Calculate the negative gradient of the loss function as the residual estimate, that is:

[0145]

[0146] Step 3: Fit the residual tree and calculate the minimum value of the loss function, that is:

[0147]

[0148] Step 4: Update the regression tree, that is:

[0149]

[0150] This gives us the final f(x).

[0151] Example 2

[0152] In this embodiment, a monthly energy consumption forecasting system for an enterprise based on fingerprints of subdivided industries is disclosed, including:

[0153] A data acquisition unit, used to acquire industry indicator data of each enterprise in the region;

[0154] The subdivided industry division unit is used to divide the enterprises in the region into subdivided industries according to the industry indicator data, and obtain multiple groups of subdivided industry enterprise groups;

[0155] The missing data completion unit is used to complete the missing monthly energy consumption data in the industry indicator data of each enterprise, and obtain the completed industry data of each enterprise;

[0156] The energy consumption prediction unit is used to train the energy consumption prediction model corresponding to each group of subdivided industry enterprise groups using the industry-completed data of all enterprises in the subdivided industry enterprise group. After the training is completed, the trained energy consumption prediction model of the subdivided industry is obtained; the energy consumption of each enterprise in the subdivided industry is predicted using the trained energy consumption prediction model of the subdivided industry.

[0157] The present invention also discloses a computer device, which includes:

[0158] a processor adapted to execute a computer program;

[0159] A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the method for predicting monthly energy consumption of an enterprise based on fingerprints of subdivided industries disclosed in Example 1 is implemented.

[0160] The present invention also discloses a computer-readable storage medium storing a computer program, which is suitable for being loaded by a processor and executing the enterprise monthly energy consumption forecasting method based on subdivided industry fingerprints disclosed in Example 1.

[0161] The present invention also discloses a computer program product, which includes a computer program. When the computer program is executed by a processor, the method for predicting monthly energy consumption of an enterprise based on fingerprints of subdivided industries disclosed in Example 1 is implemented.

[0162] The method disclosed in Example 1 can be directly embodied as a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.

[0163] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0164] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. The monthly energy consumption forecasting method for enterprises based on subdivided industry fingerprints is characterized by: include: Obtain industry indicator data for each enterprise in the region; According to the industry indicator data, the enterprises in the region are divided into sub-industries to obtain multiple groups of sub-industry enterprise groups; Complete the missing monthly energy consumption data in the industry indicator data of each enterprise to obtain the completed industry data of each enterprise; For each group of subdivided industry enterprise groups, the energy consumption prediction model corresponding to the subdivided industry is trained using the industry-completed data of all enterprises in the subdivided industry enterprise group. After the training is completed, the trained energy consumption prediction model for the subdivided industry is obtained; The energy consumption prediction model trained for the subdivided industry is used to predict the energy consumption of each enterprise in the subdivided industry.

2. The method for predicting monthly energy consumption of enterprises based on subdivided industry fingerprints according to claim 1, characterized in that: The process of completing the missing monthly energy consumption data in the enterprise's industry indicator data includes: When all monthly energy consumption data of an enterprise are missing, the enterprise's monthly energy consumption splitting model is used to split the enterprise's annual energy consumption data to obtain the enterprise's monthly energy consumption data for each month; When part of the enterprise's monthly energy consumption data is missing, determine the missing monthly energy consumption percentage, and use the missing monthly energy consumption percentage and the enterprise's annual energy consumption data to obtain the missing monthly energy consumption data of the enterprise; The obtained monthly energy consumption data is used to complete the missing monthly energy consumption data in the enterprise's industry indicator data to obtain the enterprise's industry completed data.

3. The method for predicting monthly energy consumption of enterprises based on subdivided industry fingerprints according to claim 2, characterized in that: The enterprise monthly energy consumption split model is: the enterprise's monthly energy consumption data is equal to the enterprise's monthly energy consumption split coefficient multiplied by the enterprise's annual energy consumption data; The proportion of the enterprise's electricity consumption or output in the missing months to the annual electricity consumption or output is taken as the proportion of the enterprise's missing monthly energy consumption.

4. The method for predicting monthly energy consumption of enterprises based on subdivided industry fingerprints as claimed in claim 3, characterized in that: The monthly energy consumption splitting coefficient of an enterprise is obtained by training the monthly energy consumption splitting model of the enterprise with the goal of minimizing the energy consumption splitting error, using the energy consumption data of enterprises in the enterprise group of the subdivided industry to which the enterprise belongs with known annual energy consumption data, quarterly energy consumption data and monthly energy consumption data.

5. The method for predicting monthly energy consumption of enterprises based on subdivided industry fingerprints as claimed in claim 3, characterized in that: When an enterprise lacks monthly energy consumption data but has electricity or output data in a month, the proportion of such electricity or output data in the annual electricity or output shall be directly used as the monthly energy consumption proportion for that month; when an enterprise lacks monthly energy consumption data but does not have electricity or output data in a month, the electricity or output data of the existing months shall be used to predict the electricity or output data for that month, and the proportion of the predicted electricity or output data in the annual electricity or output shall be used as the monthly energy consumption proportion for that month.

6. The method for predicting monthly energy consumption of enterprises based on subdivided industry fingerprints according to claim 1, characterized in that: The energy consumption prediction model takes the industry-completed data of enterprises in the segmented industry enterprise group as input and the energy consumption prediction results as output, and is constructed using the lightGBM model.

7. The monthly energy consumption forecasting system for enterprises based on subdivided industry fingerprints is characterized by: include: A data acquisition unit, used to acquire industry indicator data of each enterprise in the region; The subdivided industry division unit is used to divide the enterprises in the region into subdivided industries according to the industry indicator data, and obtain multiple groups of subdivided industry enterprise groups; The missing data completion unit is used to complete the missing monthly energy consumption data in the industry indicator data of each enterprise, and obtain the completed industry data of each enterprise; The energy consumption prediction unit is used to train the energy consumption prediction model corresponding to each group of subdivided industry enterprise groups using the industry-completed data of all enterprises in the subdivided industry enterprise group. After the training is completed, the trained energy consumption prediction model of the subdivided industry is obtained; the energy consumption of each enterprise in the subdivided industry is predicted using the trained energy consumption prediction model of the subdivided industry.

8. An electronic device, characterized in that: The device comprises: a processor adapted to execute a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the method for predicting monthly energy consumption of an enterprise based on fingerprints of subdivided industries as described in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the enterprise monthly energy consumption forecasting method based on segmented industry fingerprints as described in any one of claims 1-6.

10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the enterprise monthly energy consumption prediction method based on segmented industry fingerprints as described in any one of claims 1 to 6.