Rapid inversion method for electricity consumption and emission of enterprise
By integrating power monitoring and machine learning technology and building a relationship model, the accuracy and efficiency of enterprise pollutant emission accounting are solved, and high-precision and low-cost pollutant emission monitoring is achieved, with strong adaptability and meeting the needs of real-time monitoring.
Patent Information
- Application Number
- CN202510334507.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology has insufficient accuracy, high calculation complexity, poor adaptability in enterprise pollutant emission accounting, and the traditional methods are inefficient, making it difficult to meet the needs of real-time monitoring and enterprise adjustment of production strategies.
By integrating power monitoring, environmental sensing and machine learning technologies, we collect enterprise electricity consumption, raw and auxiliary material usage, product output, production process parameters and pollutant concentration data, and use atmospheric diffusion model and machine learning algorithm to build relationship models to achieve rapid inversion and real-time monitoring of pollutant emissions.
High-precision and low-cost pollutant emission monitoring are achieved, and the error rate is reduced to ≤5%, meeting real-time monitoring requirements, adapting to different industries and production processes, and reducing operating costs.
Smart Images

Figure CN120280025A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of environmental monitoring and data analysis, and particularly to a method for quickly retrieving the electricity consumption and emission volume of enterprises. Background Art
[0002] In environmental protection work, accurately accounting for the pollutant emission volume of enterprises is of great importance. Traditional emission accounting methods have many defects. The material balance method is based on the law of conservation of matter, and calculates the emission volume by measuring the material input and output in the enterprise production process. However, the enterprise production process is complex, involving multiple raw materials and chemical reactions, making it difficult to accurately measure the actual consumption and conversion volume of each material. Moreover, factors such as material loss and changes in storage conditions will affect the measurement results, resulting in large errors.
[0003] The empirical coefficient method estimates the emission volume based on the industry average emission coefficient in combination with the enterprise production scale and product output. However, there are significant differences in production processes, equipment advancement, and management levels among different enterprises. Using a unified industry average emission coefficient will cause a large deviation between the accounting result and the actual emission situation.
[0004] In addition, traditional accounting methods rely on manual operations, with low efficiency, and are prone to problems such as data recording errors and omissions. The accounting cycle is long, making it difficult to meet the requirements of real-time supervision and enterprises' timely adjustment of production strategies. In contrast, enterprise electricity consumption data is easily obtained and has a low cost. However, existing emission inversion algorithms based on electricity consumption have problems such as insufficient accuracy, too high computational complexity, and poor adaptability, and cannot meet the actual requirements of efficient and accurate monitoring.
[0005] Therefore, there is an urgent need for a method for quickly retrieving the electricity consumption and emission volume of enterprises that can solve one or more of the above problems. Summary of the Invention
[0006] To solve one or more problems existing in the prior art, the present invention provides a method for quickly retrieving the electricity consumption and emission volume of enterprises, which realizes high-precision dynamic measurement of the emission volume by integrating power monitoring, environmental sensing, and machine learning technologies. The technical solutions adopted by the present invention to solve the above problems include: the collected data includes: the electricity consumption of the enterprise, the usage amount of raw and auxiliary materials, the product output, the production process parameters, the equipment operation parameters, the pollutant concentration data at the downwind position, and the background pollutant concentration data at the upwind position within a preset time period. Among them, multiple monitoring stations are respectively set at the upwind position and the downwind position of the enterprise to measure the pollutant concentration data at the upwind position, the background pollutant concentration data at the downwind position, and the meteorological data. The pollutant concentration data at the downwind position is dynamically calibrated by the background pollutant concentration data at the upwind position to eliminate the interference of the environmental background.
[0007] Data preprocessing, cleaning the collected data through data mining algorithms and statistical algorithms, and then performing normalization processing on the cleaned data;
[0008] Emission inversion, using an atmospheric diffusion model, calculating the emission per unit time of pollutants according to the pollutant concentration data at the downwind position and the corresponding meteorological data using an inversion algorithm, and obtaining the pollutant emissions by combining the emission per unit time with the enterprise production duration;
[0009] Feature selection and extraction, screening out key features strongly correlated with the pollutant emissions from the data output by the data preprocessing through a feature screening algorithm, and extracting the key features as input variables for model construction;
[0010] Model construction, constructing a relationship model between the key features and the pollutant emissions through the Lightgbm algorithm;
[0011] Model verification and optimization, verifying the relationship model through a cross-validation method, performing performance evaluation on the relationship model, and optimizing the relationship model according to the verification and evaluation results until the relationship model meets the accuracy and stability requirements;
[0012] Fast inversion calculation, inputting the real-time electricity consumption data Preal of the enterprise obtained into the optimized and verified relationship model, calculating the corresponding pollutant emissions Ereal through the relationship model, and the calculation formula of the relationship model is Ereal = f(Preal,θ), where f is the calculation function of the relationship model and θ is the parameter of the relationship model.
[0013] In some embodiments, it further includes: regional and industry analysis, performing summary analysis on the data of multiple enterprises in a predetermined region, and the enterprise data includes: the input and output data of the relationship model, and the total emissions in the predetermined region Ei is the emissions of the i-th enterprise, obtaining the total regional emissions through the summary analysis, obtaining the overall law of product production and pollutant emissions in the predetermined region through the summary analysis, and obtaining the overall law of industry production and pollutant emissions in the predetermined region through the summary analysis.
[0014] Further, it also includes: data visualization, using a data visualization module to display the pollutant emissions data of the enterprise, the overall law of product production and pollutant emissions in the predetermined region, and the overall law of industry production and pollutant emissions in the predetermined region on the monitoring platform in real time;
[0015] Real-time warning, real-time feedback of the pollutant emissions of the enterprise, and outputting a warning if the pollutant emissions exceed the set warning value.
[0016] In some embodiments, after the collected data is cleaned in the data preprocessing, outliers and missing values are removed. For missing values, corresponding interpolation algorithms need to be selected according to the time series characteristics and correlations of the data for filling. For outliers, they need to be corrected according to industry emission standards and enterprise historical data.
[0017] In some embodiments, for the emission inversion, the atmospheric diffusion model used is the Gaussian plume model. The pollutant concentration data at the downwind position includes data from multiple monitoring stations. By fitting the data from the multiple monitoring stations, the emission amount per unit time of the pollutant is solved.
[0018] The concentration expression of the Gaussian plume model is: where C(x, y, z) is the pollutant concentration at a distance x from the emission source, the lateral distance y, and the vertical distance z. Q is the emission amount per unit time of the pollutant, u is the wind speed, and σy and σz are the atmospheric diffusion coefficients in the lateral and vertical directions respectively.
[0019] The calculation formula for the pollutant emission amount is E = Q × T, where E is the pollutant emission amount, Q is the emission amount per unit time of the pollutant, and T is the production time.
[0020] In some embodiments, for the feature selection and extraction, the feature screening algorithms include the Pearson correlation coefficient analysis algorithm, the mutual information analysis algorithm, and the principal component analysis algorithm.
[0021] When using the Pearson correlation coefficient analysis algorithm, first calculate the correlation coefficients between each feature and the pollutant emission amount, and then screen out the key features according to the magnitudes of the correlation coefficients.
[0022] When using the mutual information analysis algorithm, first calculate the mutual information amounts between each feature and the pollutant emission amount, and then screen out the key features according to the magnitudes of the mutual information amounts.
[0023] When using the principal component analysis algorithm, perform dimensionality reduction on the key features screened by the Pearson correlation coefficient analysis algorithm and the mutual information analysis algorithm, extract the main components with a high cumulative contribution rate as the final feature variables, and extract the feature variables as the input variables for model construction according to requirements.
[0024] In some embodiments, for the model construction, during the process of constructing the relationship model, adjust the predetermined parameters so that the relationship model fully learns the complex mapping relationship between the enterprise electricity consumption data and the pollutant emission amount. The predetermined parameters include the learning rate, the number of trees, and the maximum depth parameter.
[0025] Further, the data obtained after the data preprocessing is divided into a training set, a validation set, and a test set according to a preset ratio and input into the training of the relationship model. The relationship model is tuned through the validation set and appropriate model parameters are selected.
[0026] In some embodiments, for the model validation and optimization, the cross-validation method includes K-fold cross-validation and leave-one-out cross-validation methods, and the performance evaluation metrics include prediction error, root mean square error, mean absolute error, and coefficient of determination.
[0027] In some embodiments, for the model validation and optimization, the optimization of the relationship model based on the validation and evaluation results includes: parameter adjustment and feature selection optimization. The parameter adjustment includes adjusting the learning rate, the number of trees, and the maximum depth parameter of the relationship model, and the feature selection optimization includes adjusting the weights of the features and re-screening the key features.
[0028] The technical effects achieved by the present invention are as follows: Through a scientific data processing process and advanced algorithms, combined with an accurate downwind concentration monitoring and emission inversion method, the internal relationship between electricity consumption and emissions is deeply explored, a highly accurate relationship model is established, the error rate of the inversion result is effectively reduced, and the advantage of high accuracy is obtained;
[0029] The fast computing performance of the algorithm and the optimized model architecture can complete the processing and analysis of a large amount of data in a short time, realize the fast inversion of emissions, and meet the strict time requirements of real-time monitoring;
[0030] The construction and use of the above relationship model fully consider the characteristics of enterprises in different industries and different production processes. Through multi-dimensional data collection and feature extraction, as well as the optimized training of the model, it can flexibly adapt to various complex and changeable production environments and data characteristics, has a wide application prospect, and obtains the advantage of strong adaptability;
[0031] The electricity consumption data is obtained by using the existing power metering system of the enterprise, without the need to install expensive online monitoring equipment on a large scale, reducing the monitoring cost. At the same time, the algorithm has relatively low requirements for computing resources, reduces the dependence on high-performance computing equipment, and further reduces the operating cost, obtaining the advantages of low-cost deployment and use; The error of the traditional metering method is about 30%, and the error of this application can be reduced to ≤5%. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a schematic block diagram of the present invention;
[0033] Figure 2 is a schematic diagram of the monitoring site setting position of the present invention;
[0034] Figure 3This is an example of the model structure of the present invention. DETAILED DESCRIPTION
[0035] In order to make the above-mentioned objects, features and advantages of the present invention more understandable, the specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from this description, and those skilled in the art can make similar improvements without violating the connotation of the present invention, so the present invention is not limited by the specific embodiments disclosed below.
[0036] like Figure 1 As shown, the present invention discloses a rapid inversion method for enterprise electricity consumption and emissions, which realizes high-precision dynamic measurement of emissions by integrating power monitoring, environmental sensing and machine learning technology, and includes: data collection, the collected data includes: enterprise electricity consumption, raw material usage, product output, production process parameters, equipment operation parameters, pollutant concentration data at the downwind position and background pollutant concentration data at the upwind position within a preset time period, wherein a plurality of monitoring stations are respectively arranged at the upwind position and the downwind position of the enterprise to measure the pollutant concentration data at the upwind position, the background pollutant concentration data at the downwind position and meteorological data, and the pollutant concentration data at the downwind position is dynamically calibrated by the background pollutant concentration data at the upwind position to eliminate local environmental interference;
[0037] The enterprise's electricity consumption refers to the electricity consumption data (hourly) collected by the enterprise in real time through the electricity metering system for each workshop and equipment, ensuring accurate acquisition of electricity consumption in the entire production process;
[0038] Data preprocessing, cleaning the collected data by using data mining algorithms and statistical algorithms, and then normalizing the cleaned data;
[0039] Emission inversion, determine the atmospheric diffusion model to be used, use the inversion algorithm to calculate the pollutant emissions per unit time based on the pollutant concentration data at the downwind location and the corresponding meteorological data, and calculate the pollutant emissions per unit time in combination with the enterprise's production duration;
[0040] Feature selection and extraction: filtering out key features that are highly correlated with the pollutant emissions from the data output by the data preprocessing through a feature screening algorithm, and extracting the key features as input variables for model construction;
[0041] Model building, combined with Figure 3 As shown, the relationship model between the key features (e.g., power consumption, raw material usage, product output) and the pollutant emissions is constructed by the Lightgbm algorithm;
[0042] Model verification and optimization: Verify the relationship model through the cross - validation method, perform performance evaluation on the relationship model, and optimize the relationship model according to the verification and evaluation results until the relationship model meets the accuracy and stability requirements;
[0043] Fast inversion calculation: The real - time electricity consumption data of enterprises Preal obtained is input into the optimized and verified relationship model (real - time data input is achieved through an efficient data transmission interface and a real - time data processing module. Among them, the real - time data processing module can adopt a streaming computing framework such as Apache Flink or Apache Storm to process and analyze real - time data). Through the relationship model, the corresponding pollutant emissions Ereal are calculated. The calculation formula of the relationship model is Ereal = f(Preal, θ), where f is the calculation function of the relationship model and θ is the parameter of the relationship model;
[0044] Regional and industry analysis: Aggregate and analyze the data of multiple enterprises in a predetermined area. The enterprise data includes: the input and output data of the relationship model, and the total emissions of the predetermined area Ei is the emissions of the i - th enterprise. Through the above - mentioned aggregate analysis, the total regional emissions are obtained. Through the aggregate analysis, the overall law of product production and pollutant emissions in the predetermined area is obtained. Through the aggregate analysis, the overall law of industry production and pollutant emissions in the predetermined area is obtained;
[0045] Data visualization: Use the data visualization module to display the pollutant emissions data of enterprises, the overall law of product production and pollutant emissions in the predetermined area, and the overall law of industry production and pollutant emissions in the monitoring platform in real - time. For example, display the data through bar charts or line charts. Among them, the data visualization module can be developed using visualization libraries such as ECharts and Highcharts;
[0046] Real - time warning: Real - time feedback on the pollutant emissions of enterprises. If the pollutant emissions exceed the set warning value, a warning is output.
[0047] In the data collection, the enterprise electricity consumption data includes: total electricity consumption, electricity consumption of each production equipment, etc. The data time range is 3 months, and the data resolution is hourly; the raw and auxiliary material usage data includes: information such as the types, usage amounts, and usage times of raw materials. The data is sourced from the enterprise's production records or procurement records; the product output data includes: information such as the output of different products and production times. The data is sourced from the enterprise's production reports or statistical records; the production process parameters include: such as chemical reaction temperature, pressure, material ratio, etc. These parameters can be obtained from the enterprise's production records or process documents; the equipment operation parameters include: such as equipment start-stop time, rotation speed, load rate, etc. These parameters can be obtained through the enterprise's equipment monitoring system.
[0048] Regarding the description of the power metering system, it adopts rail-mounted / vertical multi-loop power meters and uses split-core current transformers for monitoring. It can simultaneously access the current inputs of one to four three-phase circuits and needs to have RS485 communication and 470MHz wireless communication functions for convenient electricity consumption monitoring, centralized meter reading, and management. It can be flexibly installed in the distribution box to achieve sub-item power energy metering, statistics, and analysis for different regions and different loads.
[0049] Combined Figure 2 As shown, monitoring stations are reasonably set near the enterprise, including downwind monitoring points and upwind reference points. Nine monitoring stations 2, 3, 4, 5, 6, 7, 8, 9, and 10 are set downwind of the enterprise emission source. They are respectively set at near, medium, and long-distance points. Among them, three monitoring stations 2, 3, and 4 are about 10m away from the enterprise, three monitoring stations 5, 6, and 7 are about 50m away from the enterprise, and three monitoring stations 8, 9, and 10 are about 150m away from the enterprise. Miniature air monitoring stations are used to collect pollutant concentration data, which corresponds to the pollutant concentration data at the downwind position. One reference point (1) is set upwind, about 10m away from the enterprise, to collect background pollutant concentration data for correcting the data of the downwind monitoring points, which corresponds to the background pollutant concentration data at the upwind position.
[0050] Regarding the description of the miniature air monitoring station, it supports the detection of multiple pollutants such as particulate matter (e.g., PM2.5, PM10), nitrogen dioxide (NO2), ozone (O3), carbon monoxide (CO), nitric oxide (NO), volatile organic compounds (TVOC), and carbon dioxide (CO2), and at the same time has the monitoring ability of five meteorological parameters (wind force, wind direction, temperature, humidity, and atmospheric pressure). The equipment operating environment temperature is -10 to 50°C, the humidity is 0 to 99%, the protection level is IP65, and it has a metal outer cover for protecting against ultraviolet rays and sunlight. The equipment communicates through a 4G module, and the data is transmitted to the cloud platform, supporting remote data viewing, downloading, and system control.
[0051] The background pollutant concentration data at the upwind position dynamically calibrates the pollutant concentration data at the downwind position using the dynamic background subtraction method: the measured concentration values at each monitoring site at the downwind position are subtracted in real time from the measured background concentration values at the monitoring site at the upwind position (reference point). The purpose is to eliminate the interference caused by regional background pollution and meteorological fluctuations during the atmospheric transmission process.
[0052] According to the principle of atmospheric pollutant diffusion, a micro air monitoring station is used to collect the pollutant concentration data of each monitoring site at the downwind and the reference point at the upwind. The time range of the monitoring data is consistent with that of the data such as the electricity consumption of the enterprise, the usage of raw materials and auxiliary materials, and the product output, and the data resolution is hourly. The types of pollutants monitored by the monitoring equipment can be determined according to the emission characteristics of the enterprise, such as sulfur dioxide, nitrogen oxides, particulate matter, etc. Meteorological data during the monitoring period are collected synchronously, including wind direction, wind speed, temperature, humidity, air pressure, etc.
[0053] It should be noted that in the data preprocessing, the collected data is cleaned to remove outliers and missing values. For missing values, corresponding interpolation algorithms (such as linear interpolation, Lagrange interpolation) need to be selected according to the time series characteristics and correlations of the data for filling. For outliers, they need to be corrected according to the industry emission standards and the enterprise historical data; the normalization processing includes Min-Max standardization, Z-Score standardization, etc., and a suitable standardization method is selected according to the characteristics of the data.
[0054] Specifically, for the emission inversion, the atmospheric diffusion model used is the Gaussian plume model. The pollutant concentration data at the downwind position contains the data of multiple monitoring sites. By fitting the data of the multiple monitoring sites, the emission amount per unit time of the pollutant is solved.
[0055] The concentration expression of the Gaussian plume model is: where C(x, y, z) is the pollutant concentration at a distance x from the emission source, lateral distance y, and vertical distance z, Q is the emission amount per unit time of the pollutant, u is the wind speed, and σy and σz are the atmospheric diffusion coefficients in the lateral and vertical directions respectively.
[0056] For example: Assume that the pollutant concentration measured at the monitoring point at a distance x from the emission source at the downwind i is C i , then there is:
[0057] The calculation formula for the pollutant emission amount is E = Q × T, where E is the pollutant emission amount, Q is the emission amount per unit time of the pollutant, and T is the production time (hours, days, weeks, months, years) corresponding to the emission amount per unit time of the pollutant.
[0058] Specifically, for the feature selection and extraction, the feature screening algorithms include Pearson correlation coefficient analysis algorithm, mutual information analysis algorithm, and principal component analysis algorithm;
[0059] When using the Pearson correlation coefficient analysis algorithm, first calculate the correlation coefficients between each feature and the pollutant emissions, and then screen out the key features according to the magnitudes of the correlation coefficients;
[0060] When using the mutual information analysis algorithm, first calculate the mutual information amounts between each feature and the pollutant emissions, and then screen out the key features according to the magnitudes of the mutual information amounts;
[0061] When using the principal component analysis algorithm, perform dimensionality reduction on the key features screened by the Pearson correlation coefficient analysis algorithm and the mutual information analysis algorithm, extract the main components with high cumulative contribution rates as the final feature variables, and extract the feature variables as the input variables for model construction according to requirements;
[0062] The purpose is to reduce the data dimension, improve the model training efficiency and accuracy. For example: analyze the relationships between the electricity consumption ratios of different production equipment, the electricity consumption change trends during specific time periods, etc. and the emissions, and extract these key features as the input variables for model construction.
[0063] Specifically, combined with Figure 3 As shown, for the model construction, during the relationship model construction process, adjust the predetermined parameters to enable the relationship model to fully learn the complex mapping relationship between the enterprise electricity consumption data and the pollutant emissions. The predetermined parameters include the learning rate, the number of trees, and the maximum depth parameter. The learning rate can control the convergence speed of the model, the number of trees can affect the fitting ability of the model, and the maximum depth can control the complexity of the model;
[0064] The data obtained after data preprocessing is divided into a training set, a validation set, and a test set according to a preset ratio and input into the training of the relationship model. The relationship model is tuned through the validation set and appropriate model parameters are selected.
[0065] Specifically, for the model verification and optimization, the cross-validation methods include K-fold cross-validation and leave-one-out cross-validation methods. Select an appropriate cross-validation method according to the data volume and the complexity of the model. The performance evaluation metrics include prediction error, root mean square error, mean absolute error, and coefficient of determination. The prediction error can measure the difference between the predicted value and the actual value of the model. The root mean square error and the mean absolute error can measure the prediction accuracy of the model. The coefficient of determination can measure the fitting degree of the model;
[0066] Performing optimization on the relationship model based on the verification and evaluation results includes: parameter adjustment and feature selection optimization. The parameter adjustment includes adjusting the learning rate, the number of trees, and the maximum depth parameter of the relationship model. The feature selection optimization includes adjusting the weights of features and re-screening the key features.
[0067] An application example of this application is as follows: 1. Data collection: Establish cooperative relationships with typical enterprises in multiple industries such as chemical industry, steel industry, electronics industry, and food industry to ensure that the selected enterprises are representative and can cover different types of production processes and emission characteristics.
[0068] Obtain historical electricity consumption data for the past 3 months from the enterprise's electricity metering system, including total electricity consumption, electricity consumption of each production workshop, and electricity consumption of major production equipment. The data collection frequency is once per hour.
[0069] Collect data on the usage of raw materials and auxiliary materials by the enterprise in the past 3 months. The data collection frequency is once per day, including information such as the types and usage amounts of raw materials.
[0070] Collect data on the product output of the enterprise in the past 3 months. The data collection frequency is once per day, including information such as the output of different products.
[0071] Collect data on the pollutant emissions of the enterprise during the same period, including pollutant indicators such as chemical oxygen demand (COD), sulfur dioxide (SO2), nitrogen oxides (NOx), and particulate matter (PM). The data is sourced from the enterprise's environmental monitoring reports and on-line monitoring equipment.
[0072] Detailedly record the production process parameters of the enterprise, such as the chemical reaction process and raw material formula of chemical enterprises; the ironmaking and steelmaking process parameters of steel enterprises; the chip manufacturing process and equipment operation parameters of electronics enterprises; the processing process and production line operation time of food enterprises.
[0073] Set monitoring points, including downwind monitoring points and upwind reference points. Set 9 monitoring stations 2, 3, 4, 5, 6, 7, 8, 9, 10 downwind of the enterprise emission source. They are respectively set at near, medium, and long-distance points. Among them, 3 monitoring stations 2, 3, 4 are about 10m away from the enterprise, 3 monitoring stations 5, 6, 7 are about 50m away from the enterprise, and 3 monitoring stations 8, 9, 10 are about 150m away from the enterprise. Use micro air monitoring stations to collect pollutant concentration data. Set 1 reference point upwind, about 10m away from the enterprise, to collect background pollutant concentration data for correcting the data of the downwind monitoring points. Subtract the measured background concentration value of the upwind position monitoring station (reference point) from the measured concentration value of each monitoring station in the downwind position in real time, so as to eliminate the interference caused by regional background pollution and meteorological fluctuations during the atmospheric transmission process. Synchronously collect meteorological data during the monitoring period, including wind direction, wind speed, temperature, humidity, air pressure, etc.
[0074] 2. The data preprocessing: Clean the collected data. For the missing electricity consumption data, according to the time series characteristics of the data and the correlation of adjacent data, use the linear interpolation method to supplement it. For example, if the electricity consumption data of a certain day is missing, the electricity consumption data of the previous day and the next day can be used for linear interpolation. For abnormal emission data, conduct comprehensive analysis and correction based on industry emission standards, the enterprise's historical emission data, and the characteristics of the production process. For example, if the sulfur dioxide emission on a certain day is extremely high, it may be caused by equipment failure or production process adjustment, and it can be corrected according to historical data and the characteristics of the production process.
[0075] Normalize all the data so that its value range is strictly controlled between [0, 1]. Perform linear transformation according to the minimum and maximum values of the original data. After normalization, the values will be distributed proportionally between 0 and 1 to eliminate the dimension difference and improve the model training effect.
[0076] 3. For the emission inversion, the Gaussian plume model is used as the atmospheric diffusion model. Calculate the emissions per unit time of pollutants using the inversion algorithm based on the pollutant concentration data at the downwind position and the corresponding meteorological data. The emissions per unit time are combined with the enterprise's production time and production scale to calculate the pollutant emissions.
[0077] 4. For the feature selection and extraction, calculate the Pearson correlation coefficients between each electricity consumption, raw material and auxiliary material usage, product output feature and emissions. Select the features with a correlation coefficient greater than 0.8 as the key features for preliminary screening. For example, the correlation coefficient between the electricity consumption ratio of production equipment and sulfur dioxide emissions is 0.85, the correlation coefficient between the change trend of electricity consumption in a specific time period and sulfur dioxide emissions is 0.78, the correlation coefficient between the change trend of raw material and auxiliary material usage and sulfur dioxide emissions is 0.82, and the correlation coefficient between the change trend of product output and sulfur dioxide emissions is 0.75.
[0078] Conduct principal component analysis on the key features for preliminary screening, and extract the principal components with a cumulative contribution rate of more than 95% as the final feature variables. For example, the extracted principal components include the electricity consumption ratio of production equipment, the change trend of electricity consumption in a specific time period, etc. These principal components can effectively reduce the data dimension and improve the model training efficiency and stability.
[0079] 5. For the model construction and training, use the Lightgbm algorithm to construct a relationship model between the key features (electricity consumption, raw material and auxiliary material usage, product output) and emissions. The Lightgbm algorithm features fast training speed and low memory occupancy, and can efficiently process large-scale data.
[0080] Set the model parameters. Set the initial learning rate to 0.1 to control the learning speed of the model; set the number of trees to 100 to determine the complexity of the model; set the maximum depth to 6 to limit the depth of each tree and prevent overfitting; according to the data characteristics and model performance requirements, set other parameters such as the minimum number of data shards and the minimum number of leaf nodes.
[0081] Divide the preprocessed data according to the ratio of 70% training set, 15% validation set, and 15% test set.
[0082] Use the training set to train the Lightgbm model, continuously adjust the model parameters through the validation set, and observe the changes in indicators such as RMSE and MAE of the model on the validation set. Stop training when the indicators no longer improve significantly to obtain the optimal model performance.
[0083] 6. Model validation and optimization. Use the test set to comprehensively validate the trained model, and calculate indicators such as the prediction error, root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R 2 ) of the model. For example, the RMSE of the model is 0.05, the MAE is 0.03, and R 2 is 0.95.
[0084] According to the validation results, further optimize the model by adopting strategies such as parameter adjustment and feature selection optimization. If the RMSE exceeds the set threshold of 0.05, the MAE exceeds 0.03, and R 2 is lower than 0.95, re-adjust the model parameters, such as reducing the learning rate and increasing the number of trees, or re-select features, and train and validate again. After multiple iterations of optimization, until the model performance meets the requirements of practical applications.
[0085] 7. The fast inversion calculation and practical application mentioned above. Deploy the optimized model to the enterprise environmental monitoring system. By docking with the enterprise power metering system and production management system, real-time obtain the electricity consumption data and production process parameters of the enterprise.
[0086] After obtaining the real-time electricity consumption data of the enterprise, through an efficient data transmission interface and a real-time data processing module, quickly input the data into the optimized model, and quickly calculate the corresponding pollutant emissions.
[0087] 8. The above-mentioned regional and industrial analysis involves aggregating and analyzing the data of multiple enterprises in the region, and combining factors such as the regional industrial structure and energy structure to explore the overall production-emission patterns in the region. For example, it is found that there is a significant positive correlation between the sulfur dioxide emissions and electricity consumption in the chemical industry in this region. By comparing and clustering the data of different enterprises in the same industry, the overall production-emission patterns of the industry can be identified. For example, large chemical enterprises have relatively higher sulfur dioxide emissions, while small chemical enterprises have lower emissions.
[0088] 9. The above-mentioned data visualization shows the results by using the data visualization module to display the enterprise emissions data, regional production-emission patterns, and industrial production-emission patterns on the monitoring platform in real time. For example, the daily sulfur dioxide emissions of enterprises are shown through bar charts, and the emission trends of the region and the industry are shown through line charts.
[0089] 10. Feedback: Set an early warning threshold. When the emissions exceed the warning value, warning messages will be sent to the environmental protection department and enterprise management personnel in a timely manner, providing timely and accurate data support for the supervision and law enforcement of the environmental protection department and the environmental management of enterprises.
[0090] In summary, through a scientific data processing process and advanced algorithms, combined with accurate downwind concentration monitoring and emission inversion methods, this application deeply explores the internal relationship between electricity consumption and emissions, establishes a highly accurate relationship model, effectively reduces the error rate of the inversion results, and has the advantage of high accuracy;
[0091] The fast computing performance of the algorithm and the optimized model architecture can complete the processing and analysis of a large amount of data in a short time, realizing the rapid inversion of emissions and meeting the strict time requirements of real-time monitoring;
[0092] The construction and use of the above-mentioned relationship model fully consider the characteristics of enterprises in different industries and different production processes. Through multi-dimensional data collection and feature extraction, as well as the optimized training of the model, it can flexibly adapt to various complex and changeable production environments and data characteristics, has a wide range of application prospects, and has the advantage of strong adaptability;
[0093] The electricity consumption data is obtained by using the existing power metering system of enterprises, without the need to install a large number of expensive on-line monitoring devices on a large scale, reducing the monitoring cost. At the same time, the algorithm has relatively low requirements for computing resources, reducing the dependence on high-performance computing devices and further reducing the operating cost, having the advantages of low-cost deployment and use; the error of traditional metering methods is about 30%, and the error of this application can be reduced to ≤5%.
[0094] The above-described embodiments merely represent one or more implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the appended claims.
Claims
1. A rapid inversion method for enterprise electricity consumption and emissions, characterized in that, Including: Data collection. The collected data includes: enterprise power consumption, raw and auxiliary material usage, product output, production process parameters, equipment operation parameters, pollutant concentration data at the downwind position, and background pollutant concentration data at the upwind position within a preset time period. Among them, multiple monitoring stations are respectively set at the upwind and downwind positions of the enterprise to measure the pollutant concentration data at the upwind position, the background pollutant concentration data at the downwind position, and meteorological data. The pollutant concentration data at the downwind position is dynamically calibrated through the background pollutant concentration data at the upwind position to eliminate environmental background interference; Data preprocessing. The collected data is cleaned through data mining algorithms and statistical algorithms, and then the cleaned data is normalized; Emission inversion. Using an atmospheric diffusion model, the emission amount per unit time of pollutants is calculated using an inversion algorithm based on the pollutant concentration data at the downwind position and the corresponding meteorological data. The pollutant emission amount is obtained by combining the emission amount per unit time with the enterprise production duration; Feature selection and extraction. Key features strongly correlated with the pollutant emission amount are screened out from the data output by the data preprocessing through a feature screening algorithm, and the key features are extracted as input variables for model construction; Model construction. A relationship model between the key features and the pollutant emission amount is constructed through the Lightgbm algorithm; Model verification and optimization. The relationship model is verified through a cross-validation method, the performance of the relationship model is evaluated, and the relationship model is optimized based on the verification and evaluation results until the relationship model meets the accuracy and stability requirements; Fast inversion calculation. The real-time electricity consumption data of the enterprise obtained is input into the optimized and verified relationship model, and the corresponding pollutant emission amount is calculated through the relationship model.
2. The rapid inversion method for enterprise electricity consumption and emissions according to claim 1, characterized in that Also including: Regional and industry analysis. The data of multiple enterprises in a predetermined region is summarized and analyzed. The enterprise data includes: the input and output data of the relationship model. The total regional emissions are obtained through the summary analysis, the overall law of product production and pollutant emissions in the predetermined region is obtained through the summary analysis, and the overall law of industry production and pollutant emissions in the predetermined region is obtained through the summary analysis.
3. The rapid inversion method for enterprise power consumption and emissions according to claim 2, characterized in that Also including: Data visualization. The pollutant emission amount data of the enterprise, the overall law of product production and pollutant emissions in the predetermined region, and the overall law of industry production and pollutant emissions are displayed in real time on the monitoring platform using a data visualization module; Real-time warning. The pollutant emission amount of the enterprise is fed back in real time. If the pollutant emission amount exceeds the set warning value, a warning is output.
4. The rapid inversion method for enterprise power consumption and emissions according to claim 1, characterized in that In the data preprocessing, after the collected data is cleaned, outliers and missing values are removed. For missing values, the corresponding interpolation algorithm needs to be selected according to the time series characteristics and correlation of the data for filling. For outliers, they need to be corrected according to industry emission standards and enterprise historical data.
5. The rapid inversion method for enterprise power consumption and emissions according to claim 1, characterized in that For the emission inversion, the atmospheric diffusion model adopted is the Gaussian plume model. The pollutant concentration data at the downwind position includes data from multiple monitoring stations. By fitting the data from the multiple monitoring stations, the emission amount per unit time of the pollutant is solved.
6. The rapid inversion method for enterprise power consumption and emissions according to claim 1, characterized in that For the feature selection and extraction, the feature screening algorithms include the Pearson correlation coefficient analysis algorithm, the mutual information analysis algorithm, and the principal component analysis algorithm. When using the Pearson correlation coefficient analysis algorithm, first calculate the correlation coefficient between each feature and the pollutant emission amount, and then screen out the key features according to the magnitude of the correlation coefficient. When using the mutual information analysis algorithm, first calculate the mutual information amount between each feature and the pollutant emission amount, and then screen out the key features according to the magnitude of the mutual information amount. When using the principal component analysis algorithm, perform dimensionality reduction on the key features screened by the Pearson correlation coefficient analysis algorithm and the mutual information analysis algorithm, extract the main components with a high cumulative contribution rate as the final feature variables, and extract the feature variables as the input variables for model construction according to requirements.
7. The rapid inversion method for enterprise electricity consumption and emissions according to claim 1, characterized in that For the model construction, during the construction process of the relationship model, adjust the predetermined parameters to enable the relationship model to fully learn the complex mapping relationship between the enterprise electricity consumption data and the pollutant emission amount. The predetermined parameters include the learning rate, the number of trees, and the maximum depth parameter.
8. The rapid inversion method for enterprise electricity consumption and emissions according to claim 7, characterized in that The data obtained after data preprocessing is divided into a training set, a validation set, and a test set according to a preset ratio and input into the training of the relationship model. The relationship model is tuned through the validation set and appropriate model parameters are selected.
9. The rapid inversion method for enterprise electricity consumption and emissions according to claim 1, characterized in that For the model verification and optimization, the cross-validation methods include the K-fold cross-validation and the leave-one-out cross-validation method. The performance evaluation indicators include the prediction error, the root mean square error, the mean absolute error, and the determination coefficient.
10. The rapid inversion method for enterprise power consumption and emissions according to claim 1 or 9, characterized in that, For the model verification and optimization, the optimization of the relationship model based on the verification and evaluation results includes: parameter adjustment and feature selection optimization. The parameter adjustment includes adjusting the learning rate, the number of trees, and the maximum depth parameter of the relationship model. The feature selection optimization includes adjusting the weights of the features and re-screening the key features.