An intelligent assessment method for the risk of concealing or underreporting hazardous waste in enterprises

By constructing a multi-dimensional waste production database and using a random forest model to predict theoretical waste production, the problem of insufficient assessment of hazardous waste in the existing technology is solved, and intelligent identification and effective management of enterprise concealment and underreporting behaviors is realized, and the accuracy and efficiency of environmental supervision are improved.

CN114912787BActive Publication Date: 2025-07-01NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210486252.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-06
Publication Date
2025-07-01
Estimated Expiration
2042-05-06

AI Technical Summary

Technical Problem

The prior art has limitations in evaluating the amount of hazardous waste generated by enterprises, which makes it difficult to control the concealment and misreporting behavior in a timely and effective manner, which may lead to an increase in environmental risks.

Method used

An intelligent evaluation method is adopted to build a multi-dimensional waste production database by obtaining and matching multiple data tables, and combining manual cleaning and unsupervised anomaly detection integrated framework, identify and eliminate abnormal data, use a random forest model to predict theoretical waste production, and finally compare it with the enterprise declaration volume to identify concealment and underreporting behavior.

Benefits of technology

It improves the accuracy and efficiency of environmental supervision, ensures data reliability and model prediction accuracy, can effectively identify the company's concealment and underreporting behavior, and reduces environmental risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114912787B_ABST
    Figure CN114912787B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent assessment method for the risk of underreporting and missing reporting of enterprise hazardous waste. It obtains relevant data tables of the enterprise, completes the exact matching between the data tables, and constructs a multi-dimensional waste production database for different industries; eliminates the dirty data in the multi-dimensional database, determines the time resolution for merging, and obtains an initial sample data set; uses an unsupervised anomaly detection integration framework to identify and eliminate abnormal data in the initial sample data set, and obtains a prediction data set; uses the prediction data set to train and verify a random forest model, predicts the theoretical waste production volume and the theoretical waste production range of the enterprise during the supervision period, and calculates the probability and quantity of underreporting and missing reporting of the enterprise's hazardous waste production. Based on the basic information of the enterprise and the online monitoring data, and combining unsupervised anomaly detection and supervised machine learning methods, the present invention accurately predicts the theoretical hazardous waste production volume and the risk of underreporting and missing reporting of the enterprise, so as to realize the intelligent supervision of the source of hazardous waste.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hazardous waste production assessment, and particularly to an intelligent assessment method for the risk of underreporting and omission of hazardous waste in enterprises. Background Art

[0002] Hazardous waste refers to solid waste that is listed in the national hazardous waste catalog or identified as having hazardous characteristics (including corrosivity, toxicity, flammability, reactivity, and infectivity) according to the national hazardous waste identification standards and identification methods. In recent years, with the acceleration of the urbanization and industrialization processes, the production volume of hazardous waste in China has maintained a high growth rate. Moreover, hazardous waste is diverse in types and complex in composition, showing an overall situation of high production intensity, insufficient disposal and utilization capacity, and frequent pollution accidents, posing a huge threat to the ecological environment and human health.

[0003] The state attaches great importance to ecological civilization construction and environmental protection work. The management of solid waste, especially hazardous waste, is the key to strengthening ecological civilization construction and improving environmental quality. One of the current challenges in the management of hazardous waste is the unclear inventory of hazardous waste. In order to obtain information on the generation and flow of hazardous waste in enterprises, China currently implements a management system based on the independent declaration and registration of hazardous waste information by enterprises. However, driven by economic interests and a fluke mentality, some enterprises are prone to underreporting and omission. If the phenomenon of underreporting and omission cannot be effectively controlled in a timely manner, it may lead to a large amount of hazardous waste being outside the scope of supervision and being illegally disposed of or dumped, causing serious environmental risks.

[0004] In order to determine whether an enterprise has underreported or omitted, it is necessary to accurately master the theoretical waste production volume of the enterprise. After comparing the theoretical waste production volume with the declared value of the enterprise, it can be judged whether the enterprise has underreported or omitted. The existing methods for predicting the theoretical waste production volume of enterprises mainly include: the production and pollution discharge coefficient method, the material balance method, and the actual measurement method. The production and pollution discharge coefficient method obtains pollutant production and discharge coefficients based on various manuals such as the "Production and Pollution Discharge Accounting Methods and Coefficient Manual for Emission Source Statistical Surveys", and combines the enterprise's product output information to calculate the total emission volume of specific pollutants; the material balance method and the actual measurement method directly collect information from production facilities through on-site research and consideration of the production conditions of specific enterprises. The existing technical methods all have certain limitations: ① The waste production coefficients are considered from the average level of regions or industries, and there are limitations in the practicability and adaptability to specific enterprises; ② The material balance method and the actual measurement method require accurate mastery of the enterprise's production process and flow, with high technical difficulty and it is also difficult to implement at the national and regional levels; ③ The above methods will introduce large deviations when the process is complex and there are many interference factors.

[0005] Therefore, it is necessary to use more scientific and appropriate methods to evaluate hazardous waste emissions at the enterprise level, grasp the theoretical amount of hazardous waste generated by enterprises, and conduct verification based on self-reported data, so as to achieve intelligent identification of corporate concealment and underreporting behavior and effectively improve the level of hazardous waste management. Summary of the invention

[0006] The technical problem to be solved by the present invention is: in order to overcome the deficiencies in the prior art, the present invention provides an intelligent assessment method for the risk of underreporting or missing reporting of hazardous wastes in an enterprise, thereby improving the accuracy and efficiency of environmental supervision.

[0007] The technical solution to be adopted by the present invention to solve the technical problem is: an intelligent assessment method for the risk of underreporting or missing reporting of hazardous wastes in an enterprise, comprising the following steps:

[0008] Step 1: Obtain the enterprise basic information table, enterprise production data table, pollutant online monitoring data table, hazardous waste production declaration data table, transfer form data table, enterprise credit evaluation data table and mobile law enforcement data table, complete the precise matching between the data tables, and classify them according to industry codes to build a multi-dimensional database of waste production in different industries.

[0009] Step 2: Manually clean the data in the waste generation multidimensional database in step 1 to eliminate dirty data in the multidimensional database; determine the time resolution based on actual application needs, merge the manually cleaned data to obtain an initial sample data set; dirty data refers to data that affects the construction of the prediction model, specifically, duplicate, non-compliant, and abnormal data are collectively referred to as dirty data; time resolution refers to the time used for data collation, that is, the time for training and prediction is the amount of hazardous waste generated by the enterprise every day, month, or year.

[0010] Step 3: Use the unsupervised anomaly detection integrated framework to identify abnormal data in the initial sample data set in step 2, and then remove the abnormal data in the initial sample data set to obtain a predicted data set; among them, the unsupervised anomaly detection integrated framework is a known technology, which is widely used in current anomaly detection tasks and has a relatively complete python library.

[0011] Step 4: Using the prediction data set in step 3, with the total hazardous waste output or the output of a single type of hazardous waste as the dependent variable, train and validate the random forest model, select the best hyperparameter combination based on the average of the root mean square error RMSE and the average ratio of the regression determination coefficient R2, and predict the theoretical waste production and theoretical waste production range of the enterprise during the regulatory period. The regulatory period refers to the time period in which the enterprise's waste production needs to be predicted and the amount and probability of underreporting need to be evaluated.

[0012] Step 5: Compare the theoretical waste generation quantity obtained in Step 4 with the actual declared quantity of the enterprise, and calculate the concealment and omission probability and quantity of the enterprise's hazardous waste output.

[0013] Furthermore, Step 1 specifically includes the following steps:

[0014] Step 1-1: Obtain the enterprise-related data tables from the enterprise-level information system. Among them, the enterprise-level information system is the full life cycle monitoring system for hazardous waste, the pollutant online monitoring system, etc., which can be accessed after obtaining permission, and other information systems that meet the requirements can also be used.

[0015] The enterprise-related data tables include:

[0016] Enterprise basic information table: including but not limited to enterprise name, enterprise ID, organization code, pollution source code, industry category code, and number of enterprise employees;

[0017] Enterprise production data table: including but not limited to raw and auxiliary material names, raw and auxiliary material usage, main product names, main product output, electricity consumption, water consumption, and total enterprise output value;

[0018] Pollutant online monitoring data table: including but not limited to monitoring time, pollution source code, pollution factors (including wastewater flow, waste gas flow, total copper, total chromium, pH, ammonia nitrogen, sulfur dioxide, etc.), and pollutant emissions;

[0019] Hazardous waste output declaration data table: including but not limited to waste name, waste code, generation quantity, unit, declaration time, and name of the generating unit;

[0020] Transfer note data table: including but not limited to transfer note number, waste name, waste code, transfer-out quantity, transfer time, and name of the generating unit.

[0021] Enterprise credit evaluation data table: including but not limited to enterprise name, pollution source code, evaluation time, credit score, and credit rating;

[0022] Mobile law enforcement data table: including but not limited to enterprise name, pollution source code, inspection time, whether environmental violations are involved, and types of violations.

[0023] Step 1-2: Precisely match each data table in Step 1-1 according to the enterprise name, pollution source code, and organization code, and construct an initial waste generation multi-dimensional database.

[0024] Step 1-3: Divide the initial waste generation multi-dimensional database obtained in Step 1-2 according to the small class codes in the National Economic Industry Classification and Codes (GB / T 4754-2017), and use the historical time period data to construct waste generation multi-dimensional databases for different industries. Among them, the historical time period refers to the time period corresponding to the data set used for model construction.

[0025] Steps 1 - 4. Optionally, according to the relevant enterprise scale classification criteria (such as the "Measures for the Classification of Large, Medium, Small and Micro Enterprises in Statistics (2017)" issued by the National Bureau of Statistics), enterprises are classified into four enterprise scale levels of large, medium, small and micro according to the number of enterprise employees and total output value, and the multi-dimensional waste generation database for different industries is further divided according to the enterprise scale level, or the enterprise scale is used as one of the input variables of the subsequent prediction model.

[0026] Furthermore, Step 2 specifically includes the following steps:

[0027] Step 2 - 1: By means of manual screening, delete the data that does not meet the user-defined integrity and duplicate data in the multi-dimensional waste generation database obtained in Step 1, and delete the unavailable variables with a large number of missing values.

[0028] Step 2 - 2: Conduct compliance inspections on the waste generation enterprises in the multi-dimensional waste generation database, and preliminarily screen out the observations of enterprises with low compliance; among them, the compliance inspection is to conduct a rough screening of the data under the assumption that the worse the enterprise environmental credit, the more environmental violations, and the easier it is to falsify the declared data, so as to ensure higher data reliability for constructing the prediction model, and it also belongs to part of the manual cleaning.

[0029] Step 2 - 3: Determine the time resolution according to the actual application requirements, and merge the data in the multi-dimensional waste generation database after manual cleaning in Step 2 - 1 and Step 2 - 2 according to the specified time period to obtain the initial sample data set. Among them, the time resolution and time period are determined according to the actual requirements. For example, if you want to predict the weekly waste generation of an enterprise, you need to sum up the cleaned data by week; if you want to predict the monthly waste generation of an enterprise, you need to sum up the cleaned data by month; if you want to predict the quarterly waste generation of an enterprise, you need to sum up the cleaned data by quarter, and so on.

[0030] Specifically, the compliance inspection of enterprises in Step 2 - 2 includes the following steps:

[0031] Step 2 - 2 - 1: Obtain the enterprise compliance information table through the matching of enterprise basic information, enterprise credit evaluation data and mobile law enforcement data.

[0032] Step 2 - 2 - 2: According to the compliance information table, count the number of annual inspections of waste generation enterprises and the number of violations among them, and calculate the violation rate:

[0033]

[0034] Step 2-2-3: Calculate the annual average credit score result of the waste-producing enterprise according to the compliance information table, and determine the enterprise's environmental protection credit level; when determining the environmental protection credit level, it is determined according to relevant laws, regulations, departmental rules, etc. In this embodiment, it corresponds to the "Measures for the Environmental Protection Credit Evaluation of Enterprises and Institutions in Jiangsu Province" to determine the enterprise's environmental protection credit level.

[0035] Step 2-2-4: Regard the enterprises with non-compliant violation rates or environmental protection credit levels as low-compliance enterprises, and delete the data of these enterprises and the corresponding years.

[0036] Furthermore, in order to improve the recognition effect of abnormal data, step 3 also includes the process of optimizing and adjusting the important parameters and abnormal ratios of the abnormal detection algorithms in the unsupervised abnormal detection integration framework.

[0037] Furthermore, step 3 specifically includes the following steps:

[0038] Step 3-1: For the initial sample data set in step 2, select the monitoring values of various hazardous waste yields, various wastewater factors, and various waste gas factors as abnormal detection features, and perform standard deviation normalization operations on the abnormal detection features to obtain a standardized detection data set;

[0039] The formula for the standard deviation normalization (Z-normalization) operation is:

[0040]

[0041] where x * is the converted data, x is the original data, μ is the mean of all sample data, and δ is the standard deviation of all sample data.

[0042] Step 3-2: Construct an unsupervised abnormal detection integration framework to identify the abnormal data in the standardized detection data set.

[0043] Since the abnormal data determined by using the unsupervised abnormal detection integration framework is multi-dimensional abnormal data and cannot be plotted in two-dimensional or three-dimensional space, it is necessary to reduce the dimension of the multi-dimensional abnormal data and map it onto a two-dimensional coordinate graph to form a visual abnormal data distribution image, and optimize and adjust the important parameters and abnormal ratios of the abnormal detection algorithm. Therefore, specifically:

[0044] Step 3-3: Use a dimensionality reduction algorithm to reduce the dimensionality of the multi-dimensional abnormal data, visualize the distribution characteristics of the abnormal data after dimensionality reduction, form a distribution image of the abnormal data, and adjust the important parameters and abnormal ratios of the abnormal detection algorithms in the abnormal detection integration framework in combination with the distribution characteristics of the abnormal data in the distribution image. Preferably, select the distribution image in which the outliers in the image are all marked and there is not much overlap between the markings of the abnormal data and the normal data as the recognition result, and obtain the prediction data set after removing the outliers from the initial sample data set. Since the parameters of different abnormal detection algorithms are different, there are also differences when adjusting the parameters. However, an abnormal ratio needs to be set for each abnormal detection algorithm. Preferably, for the data mapped to a two-dimensional coordinate graph, the normal points and abnormal points are respectively marked in blue and red. When the significantly outlying observation points in the image are all marked in red and there is not much overlap between the two data distributions, the recognition effect is better.

[0045] Optionally, the dimensionality reduction algorithm used is one of the following algorithms:

[0046] Principal Component Analysis, t-SNE (t-Distributed Stochastic Neighbor Embedding), Multidimensional Scaling, etc.

[0047] Furthermore, Step 3-2 specifically includes the following steps:

[0048] Step 3-2-1: Use several abnormal detection algorithms to respectively perform abnormal recognition on the standardized detection data set described in Step 3-1 to obtain several single-dimensional abnormal score matrices.

[0049] Optionally, the commonly used abnormal detection algorithms mainly include:

[0050] Linear Model: Minimum Covariance Determinant, One-Class Support Vector Machines, etc.;

[0051] Proximity-Based: k Nearest Neighbors, Local Outlier Factor, etc.;

[0052] Probabilistic: Angle-Based Outlier Detection, etc.

[0053] Integrated detection (Outlier Ensembles): Isolation Forest, etc.;

[0054] Neural Networks: Variational AutoEncoder, etc.

[0055] Step 3-2-2: Combine the several single-dimensional outlier score matrices described in Step 3-2-1 into a multi-dimensional outlier score matrix, perform standard deviation normalization operation to obtain a normalized multi-dimensional outlier score matrix.

[0056] Step 3-2-3: Combine the normalized multi-dimensional outlier score matrix obtained in Step 3-2-2 using a combination function, and select the part of the data with the highest comprehensive outlier score according to the outlier ratio and define it as outlier data.

[0057] Optionally, the combination function used is one of the following algorithms:

[0058] Average, Weighted Average, Maximization, combination of Average and Maximization (AOM: Average of Maximum, MOA: Maximum of Average), etc.

[0059] Furthermore, Step 4 specifically includes the following steps:

[0060] Step 4-1: Determine the predicted dependent variable, using the total hazardous waste production or the production of a single type of hazardous waste as the predicted dependent variable.

[0061] Step 4-2: The training and validation of the random forest model as a whole adopt the method of k-fold cross-validation. According to the data characteristics of the predicted dependent variable, divide the prediction data set into k groups of data with consistent dependent variable data distribution. Each time, take k-1 groups of data as the training set, and the remaining 1 group of data as the validation set, for a total of k times.

[0062] Step 4-3: Determine the hyperparameters of the random forest model, set the value range and step size of each hyperparameter, generate a list of alternative hyperparameters, and use the grid search method to substitute different combinations of hyperparameters into the random forest model for training and validation; among them, the selection of the hyperparameters of the random forest model is a process of continuous debugging based on the verification results. The appropriate range of hyperparameters for different data may vary greatly. Therefore, the optimal values of the hyperparameters are obtained through training and validation. The grid search method (Grid search method) is a very common practice in parameter tuning, which is simply an exhaustive method. For example, if hyperparameter a can take [1, 2] and hyperparameter b can take [3, 4], there will be four combinations of hyperparameters: 1 and 3, 2 and 3, 1 and 4, and 2 and 4.

[0063] The main hyperparameters to be compared and selected include:

[0064] Number of decision trees (n_estimators): The number of subtrees to be built before predicting using the majority vote or average. More subtrees can make the model have better performance;

[0065] Number of features at each node (max_features): The maximum number of variables randomly selected at each node, and then the variable with the greatest influence is selected from them;

[0066] Maximum depth of the tree (max_depth): Limit the splitting height of the subtree to reduce overfitting.

[0067] Step 4-4: Compare and select the combinations of hyperparameters described in Step 4-3 according to two performance indicators: the average of the root mean square error RMSE of k-fold validation and the average of the coefficient of determination R2 of regression, and obtain the optimal combination of hyperparameters;

[0068] Root mean square error RMSE:

[0069]

[0070] where, is the true value - predicted value on the validation set, and m is the number of samples in the validation set;

[0071] Coefficient of determination R2 of regression:

[0072]

[0073] where, the numerator part represents the sum of the squared differences between the true value and the predicted value; the denominator part represents the sum of the squared differences between the true value and the mean value.

[0074] Step 4-5: Select the random forest model corresponding to the optimal hyperparameter combination according to the industry to which the target enterprise belongs as the optimal model. For the regulatory time period, organize the independent variable parameters of the enterprise and input them into the optimal model to predict the theoretical waste generation amount of the enterprise;

[0075] Step 4-6: After the independent variable parameters are input into the optimal model, according to the residual distribution of the theoretical waste generation amount of the enterprise predicted by the random forest model during the regulatory time period, determine the residual coverage range. Based on the prediction results of the prediction data set and the residual coverage range, generate the confidence interval of the prediction results, that is, the theoretical waste generation range of the enterprise. Among them, the prediction result refers to the predicted theoretical waste generation amount of the enterprise.

[0076] Specifically, the prediction of the theoretical waste generation amount range of the enterprise in Step 4-6 specifically includes the following steps:

[0077] Step 4-6-1: For the out-of-bag data set not sampled in the construction of the random forest model, use the optimal model in Step 4-5 to predict the theoretical waste generation amount, and calculate the residual ε between the predicted waste generation amount of the out-of-bag data set and the actual value Y OOB ;

[0078] Step 4-6-2: Use the independent variable parameters of the out-of-bag data, with the residual ε as the dependent variable, and reconstruct a residual prediction random forest model according to the process of Steps 4-2 to 4-4 to predict the residual of the out-of-bag data set and sum to obtain the corrected predicted waste generation value of the out-of-bag data;

[0079] Step 4-6-3: Subtract the corrected predicted waste generation value of the out-of-bag data set from the true value Y OOB to obtain the residual of the corrected out-of-bag data;

[0080] Step 4-6-4: For the newly input regulatory time period data set x new , according to the out-of-bag data set in the construction process of the optimal model in Step 4-5, the data samples that are in the same final node of the decision tree as x new constitute a new set BOP(x new ). Use the residual prediction model to calculate the corrected residuals of each data in BOP(x new ) to obtain the residual distribution of the data set; obtain the residual distribution of the data set;

[0081] Step 4-6-5: For the residual distribution obtained in Step 4-6-4, set the confidence level to α. The upper and lower limits that cover at least α% of the samples in the residual distribution are the residual coverage range;

[0082] Step 4-6-6: Add the theoretical waste production predicted in step 4-5 to the upper and lower limits of the residual coverage range to obtain the confidence interval, which is the theoretical waste production range of the enterprise.

[0083] Further, step 5 specifically includes the following steps:

[0084] Step 5-1: Obtain and calculate the target enterprise's hazardous waste production data during the forecast period as the actual reported amount, and use the enterprise's theoretical waste production obtained in step 4 as the theoretical forecast amount to calculate the amount of underreporting:

[0085]

[0086] in, is the theoretical predicted amount, and y is the actual reported amount.

[0087] Step 5-2: Under the premise that the theoretical waste generation volume conforms to the normal distribution, the cumulative distribution function curve of the theoretical waste generation volume is obtained according to the theoretical waste generation range predicted in step 4, and the probability value corresponding to the actual declared volume of the target enterprise is obtained, which is the probability of the enterprise concealing the report:

[0088] Probability of underreporting = F X (a) = P(X>a)

[0089] Among them, F X (a) is the complementary cumulative distribution function curve of theoretical waste production, P(X>a) is the probability that the theoretical waste production is greater than a, and when a is exactly the actual declared value, F X (a) It can represent the probability that the theoretical waste production volume exceeds the actual reported volume, that is, the probability of underreporting. The greater this probability is, the higher the possibility that the actual reported volume is lower.

[0090] Step 5-3: According to the actual data situation, a threshold is proposed to be selected, and enterprises with concealed quantity and concealed probability greater than the threshold are included in the list of enterprises with high concealed or underreported risk as the key targets of environmental law enforcement. As a preferred option, the threshold of concealed quantity can be selected from the average waste production of enterprises in the industry, and the threshold of concealed probability can be selected from 50%, that is, enterprises with concealed quantity greater than the average waste production of enterprises in the industry and probability greater than 50% are included in the list of enterprises with high concealed or underreported risk as the key targets of environmental law enforcement.

[0091] The beneficial effects of the present invention are:

[0092] (1) Constructing a database that integrates multi-dimensional waste production data can provide a comprehensive and reliable data basis for the accurate prediction of hazardous waste production, avoiding the shortcomings of low model accuracy, long calculation time, and small scope of application caused by improper parameter selection.

[0093] (2) By comprehensively adopting the method combining manual data cleaning and unsupervised anomaly detection integration framework, the dirty data in the multi-dimensional database can be eliminated, the problem of relatively insufficient authenticity of the currently self-declared data can be solved, the reliability of the model input data can be ensured, and the model prediction accuracy can be improved.

[0094] (3) Based on the multi-dimensional waste generation database, by using machine learning algorithms with good generalization ability, a model with small deviation and widely applicable in the industry can be constructed to solve the problems of insufficient accuracy and applicability of the existing hazardous waste accounting methods, and realize the accounting of hazardous waste emission intensity at the enterprise level.

[0095] (4) By using the whole process of the method described in the present invention, the intelligent identification of "concealing and underreporting" of hazardous waste production in waste-related enterprises can be realized, and the problems of insufficient pertinence of environmental law enforcement, relatively lagging law enforcement and limited supervision ability can be solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] The present invention will be further described below in conjunction with the drawings and embodiments.

[0097] Figure 1 It is the overall flowchart of the intelligent evaluation method of the present invention.

[0098] Figure 2 It is the flowchart of the integrated abnormal data detection method.

[0099] Figure 3 It is the flowchart of the intelligent identification method for concealing and underreporting based on the random forest model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0100] The present invention will now be described in detail with reference to the drawings. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner, so it only shows the components related to the present invention.

[0101] The present invention provides an intelligent evaluation method for the risk of concealing and underreporting of hazardous waste in enterprises. This embodiment describes the application of the method provided by the present invention to the electronic circuit manufacturing industry in Jiangsu Province (industry code C3982) to identify the behavior of concealing and underreporting of hazardous waste in enterprises.

[0102] Combined with the attached Figure 1 , an intelligent evaluation method for the risk of concealing and underreporting of hazardous waste in enterprises according to the present invention includes the following steps:

[0103] Step 1: Obtain the enterprise basic information table, enterprise production data table, pollutant online monitoring data table, hazardous waste production declaration data table, transfer note data table, enterprise credit evaluation data table and mobile law enforcement data table, complete the accurate matching between the data tables, and classify according to the industry code to construct a multi-dimensional waste generation database for different industries.

[0104] Step 2: Manually clean the data in the waste production multidimensional database in step 1 to eliminate dirty data in the multidimensional database. Specifically; determine the time resolution according to actual application needs, merge the manually cleaned data to obtain an initial sample data set; dirty data refers to data that affects the construction of the prediction model, specifically, duplicate, non-compliant, and abnormal data are collectively referred to as dirty data; time resolution refers to whether the object of training and prediction is the amount of hazardous waste generated by the enterprise every day, month, or year; period merging is to add up the daily data into monthly data, and add up the monthly data into annual data.

[0105] Step 3: Use the unsupervised anomaly detection integrated framework to identify abnormal data in the initial sample data set in step 2, and then remove the abnormal data in the initial sample data set to obtain a predicted data set; among them, the unsupervised anomaly detection integrated framework is a known technology, which is widely used in current anomaly detection tasks and has a relatively complete python library.

[0106] Step 4: Using the prediction data set in step 3, with the total hazardous waste output or the output of a single type of hazardous waste as the dependent variable, train and validate the random forest model. Select the best hyperparameter combination based on the average ratio of the root mean square error (RMSE) and the average ratio of the regression determination coefficient (R2) to predict the theoretical waste output and theoretical waste output range of the enterprise during the regulatory period.

[0107] Step 5: Compare the theoretical waste production obtained in step 4 with the actual amount reported by the enterprise, and calculate the probability and amount of underreporting of hazardous waste production by the enterprise.

[0108] Step 1 of this embodiment specifically includes:

[0109] Step 1-1: Obtain relevant data tables from enterprise-level information systems such as the hazardous waste life cycle monitoring system and the pollutant online monitoring system, including: basic enterprise information, enterprise production data, pollutant online monitoring data, hazardous waste production declaration data, transfer form data, enterprise credit evaluation data and mobile law enforcement data.

[0110] Step 1-2: Accurately match each data table according to the company name, pollution source code and organizational code to build a multi-dimensional database of waste generation.

[0111] Step 1-3: Divide the waste production multidimensional database according to the small and medium-sized category codes of the National Economic Industry Classification and Code (GB / T 4754-2017), screen out the enterprise data of 92 companies belonging to the industry C3982, and use the historical data from January 2020 to November 2021 to construct a multidimensional database of waste production for enterprises belonging to the industry C3982.

[0112] Step 2 of this embodiment specifically includes:

[0113] Step 2-1: Delete the data in the enterprise waste generation multi-dimensional database of C3982 that does not meet the user-defined integrity, duplicate data, and unavailable variables with a large number of missing values.

[0114] Step 2-2: Obtain the enterprise compliance information table by matching the enterprise basic information, enterprise credit evaluation data, and mobile law enforcement data. Count the number of inspections of waste generation enterprises each year and the number of violations among them, calculate the violation rate, and determine the enterprise environmental protection credit level according to the credit score results and the "Jiangsu Province Enterprise and Institution Environmental Protection Credit Evaluation Measures". Enterprises with a violation rate greater than 10% or an environmental protection credit level lower than the blue level are regarded as low-compliance enterprises, and the corresponding data of the enterprises are deleted.

[0115] Violation rate:

[0116]

[0117] Step 2-3, the time resolution refers to the resolution used when organizing the data, which is "month" in this embodiment. The dataset after manual cleaning is merged with a monthly resolution, that is, the data belonging to the same month are merged. The specific method is to sum up various hazardous wastes, wastewater flow, ammonia nitrogen, and COD values monthly to obtain the initial sample dataset, with a total of 608 data.

[0118] Combined with Appendix Figure 2 , Step 3 of this embodiment specifically includes:

[0119] Step 3-1: Select four features, namely the total amount of hazardous waste, wastewater flow, ammonia nitrogen, and COD in the initial sample dataset, as the anomaly detection features, and perform standard deviation normalization on the anomaly detection features to obtain the standardized detection dataset;

[0120] Standard deviation normalization (Z-normalization):

[0121]

[0122] Among them, x * is the converted data, x is the original data, μ is the mean of all sample data, and δ is the standard deviation of all sample data.

[0123] Step 3-2: Select six common anomaly detection models, namely Isolation Forest (iForest), Minimum Covariance Determinant (MCD), Local Outlier Factor (LOF), k-Nearest Neighbors (KNN), Clustering-Based Local Outlier Factor (CBLOF), and Histogram-Based Outlier Detection (HBOS), to construct an unsupervised anomaly detection integration framework. Use it to identify outliers in the standardized detection dataset and obtain six one-dimensional anomaly score matrices. Then, perform standardization processing on the six-dimensional anomaly score matrix identified by the models again, and use the combination function of AOM (Average of Maximum) to merge them. Select the part of the data with the highest comprehensive anomaly score according to the anomaly ratio and define it as the abnormal data;

[0124] Specifically, Isolation Forest (iForest) is a detection algorithm based on the integration of multiple decision trees. Its basic principle is as follows. In the isolation forest, the dataset is randomly split recursively until all sample points are isolated. By synthesizing the results of all decision trees, the ones with shorter total paths are usually outliers;

[0125] Minimum Covariance Determinant (MCD) is a detection algorithm based on Mahalanobis distance. Its basic principle is to use the minimum covariance determinant to calculate and obtain more robust estimators of the mean and covariance, and then calculate according to the Mahalanobis distance. The data points with Mahalanobis distance greater than the critical value are outliers;

[0126] Local Outlier Factor (LOF) is a density-based detection algorithm. Its basic idea is to calculate the local reachability density of each data point according to the data density around the data point, and then further calculate the outlier factor of each data point through the local reachability density. This outlier factor indicates the degree of outliers of a data point. The larger the factor value, the higher the degree of outliers;

[0127] k-Nearest Neighbors (KNN) is a distance-based detection algorithm. Its basic principle is to calculate the average distance between each sample point and its nearest k samples in turn. If the calculated average distance is greater than the threshold, it is considered an outlier;

[0128] Clustering-Based Local Outlier Factor (CBLOF) is a clustering-based detection algorithm. Its basic principle is to use clustering to determine the dense regions in the data, and then perform density estimation on each cluster;

[0129] Histogram-Based Outlier Detection (HBOS) is a statistical method-based detection algorithm. Its basic principle is to assume that each dimension is independent, divide each dimension into intervals again, and the outliers corresponding to each interval depend on the density. The higher the density, the lower the outliers;

[0130] The AOM combination function is a combination method that combines simple averaging and maximization. Specifically, the multi-dimensional anomaly score matrix is divided into several groups by dimension averaging. For each piece of data, the maximum anomaly score within the group is taken, and the average value is taken among the groups to obtain the comprehensive anomaly score.

[0131] Step 3-3: Use the t-SNE (t-Distributed Stochastic Neighbor Embedding) dimensionality reduction algorithm to visualize the distribution characteristics of multi-dimensional anomaly data, forming a distribution image of the anomaly data. The important parameters and anomaly ratio of the algorithm can be adjusted by combining the distribution characteristics of the anomaly data in the distribution image. Finally, 10% (60 pieces) of the anomaly data is selected and removed from the initial sample dataset to obtain the prediction dataset.

[0132] Specifically, the t-SNE algorithm is a non-linear dimensionality reduction technique, which can better verify the performance of the algorithm through visual visualization. The similarity between data points is converted into probability. The similarity in the high-dimensional space is represented by the Gaussian joint probability, and the similarity in the low-dimensional space is represented by the "Student's t-distribution". The dimensionality reduction of the data is completed by maximizing the similarity of the distributions in the high and low-dimensional spaces as much as possible.

[0133] Combined with Appendix Figure 3 , Steps 4 and 5 of this embodiment specifically include:

[0134] Step 4-1: Use the Random Forest algorithm, with the three characteristics of wastewater flow, ammonia nitrogen, and COD in the prediction dataset as independent variables, and the total hazardous waste production as the dependent variable, to train and verify the model;

[0135] Specifically, the random forest is an algorithm based on the integration of decision trees. When applied to regression and testing, its basic principle is to randomly and repeatedly draw k samples from the original training sample set N with replacement to generate a new training sample set, and then generate k regression trees to form a random forest. The predicted value of the new data is the average of the prediction results of all regression trees.

[0136] Step 4-2: The training and verification of the random forest model as a whole adopt the method of ten-fold cross-validation. According to the characteristics of the dependent variable data to be predicted, the prediction dataset is divided into 10 groups with consistent dependent variable data distributions. Each time, 9 groups are taken as the training set, and the remaining 1 group is taken as the verification set, for a total of 10 times;

[0137] Step 4-3: Set a certain value range and step size for the three main hyperparameters to generate a list of alternative hyperparameters. Use the grid search method to substitute different combinations of hyperparameters into the model for training and verification.

[0138] The main hyperparameters for comparison and selection include:

[0139] Number of decision trees (n_estimators): The number of subtrees to be built before predicting using the majority vote or average. A larger number of subtrees can result in better model performance.

[0140] Number of features at each node (max_features): The maximum number of variables randomly selected at each node, and then the variable with the greatest impact is selected from them.

[0141] Maximum tree depth (max_depth): Limits the splitting height of subtrees to reduce overfitting.

[0142] Step 4-4: Compare and select hyperparameter combinations based on two performance metrics, the average root mean square error (RMSE) of k-fold validation and the average coefficient of determination (R2) of regression, to obtain the optimal model. The performance metrics of the optimal model are R2 = 0.74 and RMSE = 603.21.

[0143] Root mean square error (RMSE):

[0144]

[0145] where, is the true value - predicted value on the validation set, and m is the number of samples in the validation set;

[0146] Coefficient of determination (R2) of regression:

[0147]

[0148] where, the numerator represents the sum of the squared differences between the true value and the predicted value; the denominator represents the sum of the squared differences between the true value and the mean;

[0149] Step 4-5: Taking December 2021 as the regulatory time period and Company A as an example, input the wastewater flow, ammonia nitrogen, and COD of Company A in December 2021 as independent variable feature values into the optimal model, and predict its theoretical waste production to be 381.83 tons;

[0150] Step 4-6: According to the out-of-bag dataset generated when building the random forest model, predict the residual distribution of the input dataset. Combining the prediction results of the prediction dataset and the set residual coverage range, generate a 95% confidence interval for the prediction results, and obtain the theoretical waste production range of the enterprise as [122.54, 479.93].

[0151] Step 5-1: Obtain the hazardous waste production declaration data of Company A in December 2021. The actual declared value is 137.81 tons, and the calculated underreporting quantity is 244.02 tons.

[0152] Underreporting quantity:

[0153]

[0154] Among them, is the theoretically predicted waste generation amount, and y is the actual declared amount;

[0155] Step 5-2: Under the premise assumption that the theoretical generation amount conforms to the normal distribution, obtain the complementary cumulative distribution function curve of the theoretical generation amount according to the theoretical waste generation range, and obtain that the probability value corresponding to the actual declared amount of the target enterprise is 97.5%, that is, the enterprise's concealment probability is 97.5%.

[0156] Step 5-3: Select the threshold of the concealment probability as 50%. The concealment quantity of enterprise A far exceeds 50% of the theoretical waste generation amount and the concealment probability is as high as 97.5%. It can be considered that the enterprise has a high risk of concealment and underreporting, and should be taken as the key supervision object.

[0157] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent assessment method for the risk of underreporting and missing reporting of enterprise hazardous waste, characterized in that: The following steps are involved: Step 1: Obtain enterprise basic information table, enterprise production data table, pollutant online monitoring data table, hazardous waste production declaration data table, transfer form data table, enterprise credit evaluation data table and mobile law enforcement data table, complete accurate matching between data tables, and classify them according to industry codes to build a multi-dimensional database of waste production in different industries; Step 2: Manually clean the data in the waste multidimensional database in step 1 to eliminate dirty data in the multidimensional database, determine the time resolution according to actual application requirements, merge the manually cleaned data, and obtain the initial sample data set; Step 3: Construct an unsupervised anomaly detection integrated framework, use the unsupervised anomaly detection integrated framework to identify abnormal data in the initial sample data set in step 2, and then remove the abnormal data in the initial sample data set to obtain a predicted data set; Step 4: Using the prediction data set in step 3, with the total hazardous waste output or the output of a single type of hazardous waste as the dependent variable, the random forest model is trained and validated. The best hyperparameter combination is selected based on the average of the root mean square error RMSE and the average of the regression determination coefficient R2, and the theoretical waste output and theoretical waste output range of the enterprise within the regulatory period are predicted; Step 5: Compare the theoretical waste production obtained in step 4 with the actual amount reported by the enterprise, and calculate the probability and amount of underreporting of hazardous waste production by the enterprise.

2. The intelligent assessment method for the risk of concealing or underreporting enterprise hazardous waste as claimed in claim 1, wherein: Step 1 specifically includes the following steps: Step 1-1: Obtain enterprise-related data tables from the enterprise-level information system, wherein the enterprise-related data tables include: Basic information table of the enterprise: including enterprise name, enterprise ID, organization code, pollution source code, industry category code and number of employees; Enterprise production data table: including the name of raw materials and auxiliary materials, the amount of raw materials and auxiliary materials used, the name of main products, the output of main products, electricity consumption, water consumption and total output value of the enterprise; Pollutant online monitoring data table: including monitoring time, pollution source code, pollution factor and pollutant emission; Hazardous waste production declaration data form: including waste name, waste code, production volume, unit, declaration time and name of production unit; Transfer form data sheet: including transfer form number, waste name, waste code, removal quantity, transfer time and name of the generating unit; Enterprise credit evaluation data table: including enterprise name, pollution source code, evaluation time, credit score and credit rating; Mobile law enforcement data table: including enterprise name, pollution source, inspection time, whether it involves environmental violations and violation type; Step 1-2: Accurately match the data tables in step 1-1 according to the enterprise name, pollution source code and organization code to build an initial waste generation multidimensional database; Step 1-3: Based on the initial waste generation multidimensional database obtained in step 1-2 of the national economic industry classification and the code sub-category code division, use the historical time period data to build a waste generation multidimensional database for different industries; Step 1-4: According to the relevant enterprise scale classification criteria, enterprises are classified into four enterprise scale levels of large, medium, small, and micro according to the number of enterprise employees and total output value, and the multi-dimensional waste generation databases of different industries are further divided according to the enterprise scale level, or the enterprise scale is used as one of the input variables of the subsequent prediction model.

3. The intelligent evaluation method for the risk of concealing or omitting the reporting of enterprise hazardous waste as claimed in claim 1, wherein: Step 2 specifically includes the following steps: Step 2-1: By means of manual screening, delete the data that does not meet the user-defined integrity and duplicate data in the multi-dimensional waste generation database obtained in Step 1, and delete the unavailable variables; Step 2-2: Conduct compliance inspections on the waste generation enterprises in the multi-dimensional waste generation database, and preliminarily screen out the observations of enterprises with low compliance; Step 2-3: Determine the time resolution according to the actual application requirements, and merge the data in the multi-dimensional waste generation database that has been manually cleaned in Steps 2-1 and 2-2 according to the specified time resolution to obtain the initial sample data set.

4. The intelligent evaluation method for the risk of concealing or omitting the reporting of enterprise hazardous waste as described in claim 3, wherein: The compliance inspection of enterprises in Step 2-2 includes the following steps: Step 2-2-1: Obtain the enterprise compliance information table through the matching of enterprise basic information, enterprise credit evaluation data, and mobile law enforcement data; Step 2-2-2: According to the compliance information table, count the number of annual inspections of waste generation enterprises and the number of violations among them, and calculate the violation rate: Step 2-2-3: Calculate the annual average credit score result of waste generation enterprises according to the compliance information table, and determine the enterprise environmental protection credit level; Step 2-2-4: Regard the enterprises with violation rates or environmental protection credit levels not meeting the requirements as low-compliance enterprises, and delete the data of these enterprises and the corresponding years.

5. The intelligent evaluation method for the risk of concealing or omitting the reporting of enterprise hazardous waste as claimed in claim 1, wherein: Step 3 also includes the process of optimizing and adjusting the important parameters and anomaly ratios of the anomaly detection algorithms in the unsupervised anomaly detection integration framework.

6. The intelligent assessment method for the risk of concealing or underreporting enterprise hazardous waste as claimed in claim 5, wherein: Step 3 specifically includes the following steps: Step 3-1: For the initial sample data set in Step 2, select the anomaly detection features, and perform standard deviation normalization on the anomaly detection features to obtain the standardized detection data set; The formula for the standard deviation normalization operation is: where x * is the converted data, x is the original data, μ is the mean of all sample data, and δ is the standard deviation of all sample data; Step 3-2: Construct an unsupervised anomaly detection integration framework to identify the anomaly data in the standardized detection data set; Step 3-3: Use the dimensionality reduction algorithm to reduce the dimensionality of the multi-dimensional anomaly data, and visualize the distribution characteristics of the reduced-dimensional anomaly data to form the distribution image of the anomaly data. Combine the distribution characteristics of the anomaly data in the distribution image to adjust the important parameters and anomaly ratios of the anomaly detection algorithms in the unsupervised anomaly detection integration framework, and obtain the prediction data set after removing the outliers in the initial sample data set.

7. The intelligent assessment method for the risk of concealing or underreporting enterprise hazardous waste as claimed in claim 6, wherein: Step 3-2 specifically includes the following steps: Step 3-2-1: Use several anomaly detection algorithms to respectively perform anomaly identification on the standardized detection data set described in Step 3-1 to obtain several single-dimensional anomaly score matrices; Step 3-2-2: Merge the several single-dimensional anomaly score matrices described in Step 3-2-1 into a multi-dimensional anomaly score matrix, and perform standard deviation normalization operation to obtain the standardized multi-dimensional anomaly score matrix; Step 3-2-3: Combine the standardized multi-dimensional anomaly score matrices described in Step 3-2-2, and select the part of the data with the highest comprehensive anomaly score according to the anomaly ratio and define it as the abnormal data.

8. The intelligent assessment method for the risk of concealing or underreporting enterprise hazardous waste as claimed in claim 1, wherein: Step 4 specifically includes the following steps: Step 4-1: Determine the predicted dependent variable, and use the total hazardous waste production or the production of a single type of hazardous waste as the predicted dependent variable; Step 4-2: The training and validation of the random forest model as a whole adopt the k-fold cross-validation method. According to the data characteristics of the predicted dependent variable, divide the prediction data set into k groups of data with consistent dependent variable data distributions. Each time, take k-1 groups of data as the training set, and the remaining 1 group of data as the validation set, for a total of k times; Step 4-3: Determine the hyperparameters of the random forest model, set the value range and step size of each hyperparameter, generate a list of alternative hyperparameters, and use the grid search method to substitute different hyperparameter combinations into the random forest model for training and validation respectively; Step 4-4: According to the two performance indicators of the average of the root mean square error RMSE and the average of the coefficient of determination R2 in k validations, compare and select the hyperparameter combinations described in Step 4-3 to obtain the optimal hyperparameter combination; Root mean square error RMSE: Among them, is the true value - predicted value on the validation set, and m is the number of samples in the validation set; Coefficient of determination R2: Among them, the numerator part represents the sum of the squared differences between the true value and the predicted value; the denominator part represents the sum of the squared differences between the true value and the mean; Step 4-5: Select the random forest model corresponding to the optimal hyperparameter combination according to the industry to which the target enterprise belongs as the optimal model. For the regulatory time period, organize the independent variable parameters of the enterprise and input them into the optimal model to predict the theoretical waste production of the enterprise; After the independent variable parameters are input into the optimal model, according to the residual distribution of the theoretical waste production of the enterprise predicted by the random forest model during the regulatory time period, determine the residual coverage range. Combine the prediction results of the prediction data set and the residual coverage range to generate a confidence interval for the prediction result, that is, the theoretical waste production range of the enterprise. Among them, the prediction result refers to the predicted theoretical waste production of the enterprise.

9. The intelligent evaluation method for the risk of concealing or omitting the reporting of enterprise hazardous waste as claimed in claim 8, characterized in that: The prediction of the theoretical waste production range of the enterprise in Step 4-6 specifically includes the following steps: Step 4-6-1: For the out-of-bag data set not sampled in the construction of the random forest model, use the optimal model in Step 4-5 to predict the theoretical waste generation amount, and calculate the predicted waste generation amount of the out-of-bag data set and the actual value Y OOB of the residual ε; Step 4-6-2: Using the out-of-bag data independent variable parameters, with the residual ε as the dependent variable, reconstruct a residual prediction random forest model according to the process in Steps 4-2 to 4-4 to predict the residuals of the out-of-bag dataset and sum them up to obtain the predicted value of the waste generation amount of the corrected out-of-bag data Step 4-6-3: Use the predicted value of the waste generation amount of the out-of-package dataset after calibration to subtract from the true value Y OOB to obtain the residual of the out-of-package data after calibration Step 4-6-4: For the newly input regulatory time period dataset x new , according to the out-of-bag dataset in the optimal model construction process in Step 4-5, the data samples that are in the same final decision tree node as x new form a new set BOP(x new ). Using the residual prediction model, calculate the corrected residuals of each data in BOP(x new ) to obtain the residual distribution of the dataset; Step 4-6-5: For the residual distribution obtained in Step 4-6-4, set the confidence level to α. The upper and lower limits that cover at least α% of the samples in the residual distribution are the residual coverage range; Step 4-6-6: Add the theoretical waste production predicted in Step 4-5 to both the upper and lower limits of the residual coverage range to obtain the confidence interval, which is the theoretical waste production range of the enterprise.

10. The intelligent evaluation method for the risk of concealing or underreporting enterprise hazardous waste as claimed in claim 1, characterized in that: Step 5 specifically includes the following steps: Step 5-1: Obtain and calculate the declared data of the hazardous waste production of the target enterprise during the regulatory time period as the actual declared quantity. Use the theoretical waste production of the enterprise obtained in Step 4 as the theoretical predicted quantity, and calculate the underreported quantity: Among them, is the theoretical predicted quantity, and y is the actual declared quantity; Step 5-2: Under the premise assumption that the theoretical waste production conforms to a normal distribution, obtain the cumulative distribution function curve of the theoretical waste production according to the theoretical waste production range predicted in Step 4, and obtain the probability value corresponding to the target enterprise's actual declared quantity, which is the underreporting probability of the enterprise: Concealment probability = F X (a) = P(X > a) Among them, F X (a) is the complementary cumulative distribution function curve of the theoretical waste production volume, and P(X > a) is the probability that the theoretical waste production volume is greater than a. When the value of a is exactly the actual declared value, F X (a) can represent the probability that the theoretical waste production volume exceeds the actual declared volume, that is, the probability of underreporting; Step 5-3: According to the actual situation of the data, a threshold is tentatively set, and the enterprises with underreported quantity and probability greater than the threshold are included in the list of enterprises with high risks of underreporting and omission, and are taken as the key targets for environmental protection law enforcement.

Citation Information

Patent Citations

  • Method and equipment for assessing quality of hazardous waste declaration data

    CN107194188A

  • Hazardous waste output determination method and device, equipment and storage medium

    CN113723685A