Health influence assessment method and system for coking industry
Through multi-path data collection and deep neural network models, the problem of insufficient comprehensive consideration of the contribution rate of pollution paths in the health risk assessment of the coking industry was solved, achieving more accurate carcinogenic risk assessment and supporting environmental management and health protection.
Patent Information
- Application Number
- CN202511102090.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-09-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies fail to comprehensively consider the contribution rates of different pollution pathways in health risk assessments in the coking industry, resulting in low assessment accuracy and an inability to comprehensively assess the impact of pollutants on human health.
Multi-path data collection, pollutant migration simulation, data classification processing and deep neural network model establishment are adopted. The diffusion and migration of pollutants are simulated through Gaussian model, BP neural network and hydrogeological model. Deep neural network is combined to perform risk cluster classification and carcinogenic risk calculation to form a complete health impact assessment system.
It has improved the accuracy of carcinogenic risk assessment of pollutants around coking plants, can comprehensively evaluate the health impacts of different exposure pathways, and provide scientific basis and decision-making support for environmental management and public health protection.
Smart Images

Figure CN120613148A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of health impact assessment, and in particular to a health impact assessment method and system for the coking industry. Background Art
[0002] The coking industry is a typical heavily polluting industry, involving the emission of a variety of harmful substances, such as benzene, benzene compounds, heavy metals, etc. These pollutants may have acute or chronic effects on human health through air, water, soil and other pathways.
[0003] When transmitted through the air, the emissions enter the human body and act as pro-oxidants for proteins and lipids, inducing the production of excessive reactive oxygen species and triggering oxidative stress reactions, which may eventually cause lesions in the respiratory, nervous, and cardiovascular systems, and induce serious consequences such as chronic diseases and even cancer. Pollutants from different sources not only cause serious pollution to the ecological environment, but can also directly threaten human health through the food chain.
[0004] A pollutant health risk assessment method and device with application number CN202210901793.9 obtains risk assessment results based on the source type and contribution concentration of the pollution factor and the preset health risk assessment factor. The pollutant concentration is distributed through the orthogonal matrix factor decomposition method to obtain the contribution concentration of the pollution factor corresponding to different source types, so as to facilitate subsequent risk assessment based on the source type of the pollution factor and the corresponding contribution concentration.
[0005] However, risk assessment methods often evaluate whether a site or region poses a health risk based solely on soil pollutant concentrations and ambient air pollutant concentrations. However, different pollution pathways will result in different risk levels, and the contribution rates to pollution risks vary greatly. Traditional risk assessment methods only comprehensively consider a single pollution exposure pathway and fail to comprehensively assess the impact of pollution on human health based on the pollution contribution rates brought about by different exposure pathways. Therefore, the assessment of human health is too one-sided and has low accuracy.
[0006] In summary, a health impact assessment method and system for the coking industry are provided to address the technical deficiencies mentioned in the background technology. Summary of the Invention
[0007] This application provides a health impact assessment method for the coking industry, which includes the following steps: Step 1: Acquire historical environmental monitoring data around the coking plant, wherein the historical environmental monitoring data includes meteorological data, air pollutant data, heavy metal content in soil, and water pollutant data; Step 2: Based on the meteorological data and the air pollutant data, a Gaussian model is used to simulate the diffusion pattern of air pollutants in the atmosphere. At the same time, based on the heavy metal content in the soil, a BP neural network model is used to simulate the distribution pattern of heavy metals in the soil. Based on the water pollutant data, a hydrogeological model is constructed to simulate the migration pattern of water pollutants in coking wastewater in the water body; Step 3: Based on the Gaussian model, the BP neural network model, and the hydrogeological model, pollutant data under each pollution path is collected and preprocessed to obtain multiple pollutant data packets. Outliers are detected and removed for each of the pollutant data packets to obtain a pollutant data set. The pollutant data sets are assigned to different risk clusters based on preset rules. A deep neural network model is configured to output the risk cluster type to which each pollutant data set belongs. Step 4: training the deep neural network model based on retrospective loss to improve the classification accuracy of the risk cluster type to which each pollutant data belongs, each risk cluster having a corresponding carcinogenicity proportion; Step 5: Based on the carcinogenic proportion corresponding to each risk cluster, the carcinogenic risk of pollutants from the coking plant under different pathways to the human body is established.
[0008] Specifically, in step 3, the pollutant data under each pollution path is collected and preprocessed to obtain multiple pollutant data packets, including: The pollutant data under each pollution path is defined as raw data, and the raw data is cleaned. The cleaned data is then denoised using Kalman filtering, and the data missing points are processed for the denoised data. Finally, the data under each pollution path are integrated to form a pollutant data packet corresponding to each pollution path.
[0009] Specifically, in step 3, performing outlier detection and removing each pollutant data packet to obtain a pollutant data set includes: S31, using different discretization strategies to discretize the multi-index data within the pollutant data packet under each of the pollution paths to form a plurality of discretized data groups, and then combining the plurality of discretized data groups with a plurality of candidate values of the LOF parameter K one by one to form a plurality of evaluation models corresponding one by one to the plurality of discretized data groups; S32. Based on the multiple evaluation models, respectively calculate a first score for each indicator data in each discretized data group, respectively calculate a recognition rate for each evaluation model based on the first scores, and define the evaluation model corresponding to the maximum recognition rate as an outlier evaluation model; S33. Calculate the second score of each data in the multiple pollutant data packets based on the outlier evaluation model, and compare each second score with a preset threshold. If the second score is greater than the preset threshold, it indicates that the indicator data is an outlier. After traversing all the data, eliminate the indicator data of the outlier and integrate the normal data to form the pollutant data set.
[0010] Specifically, in step 3, allocating the pollutant dataset to different risk clusters based on preset rules includes: Step 34: Perform a toxicity assessment on all the pollutant data, set a risk limit for each pollutant data, divide the pollutant data into risk clusters based on the risk limit, and assign a different cluster identifier to each pollutant data. The risk cluster is defined as acute toxicity, chronic toxicity, or carcinogenicity. Step 35: Based on the K-Means clustering algorithm, the number of clusters and the health risk level represented by each cluster are set. The K-Means clustering algorithm performs cluster analysis on the pollutant data and different risk clusters. When the preset number of clusters is reached, the cluster analysis is terminated, completing the cluster analysis of the pollutants and the risk clusters to which they belong.
[0011] Specifically, in step 4, training the deep neural network model based on retrospective loss includes: Step 41: Initialize the deep neural network model, which includes multiple neurons, assign an initial weight value to each neuron, and assign the deep neural network model a classification task of classifying pollutant data into different risk clusters; Step 42: Obtain a training data set representing the initial output and the expected output; Step 43: preheat training the deep neural network model based on the training data set; Step 44: After completing the warm-up training, enter the formal training and train the deep neural network model through multiple iterations; Step 45: When the difference between the current predicted output and the expected output is within a preset difference threshold, the training process is terminated.
[0012] Specifically, the warm-up training includes: A warm-up iterative calculation is performed on the deep neural network model, wherein the deep neural network model defines the clustering data set of the pollutant data and the risk clusters to which it belongs as the initial classification output, and defines the clustering data set of the pollutant data and the risk clusters to which it belongs based on the K algorithm as the expected classification output, and determines the first difference between the initial classification output and the expected classification output. Subsequently, the loss function of the classification task of the deep neural network model is determined based on the first-order difference of the deep neural network model, and one or more parameters of the deep neural network model are updated according to the loss function of the indication task.
[0013] Specifically, the formal training includes: causing the deep neural network model to generate the current predicted output based on the current input data, and calculating the difference between the current predicted output and the expected classification output; Calculating a loss for the classification task based on the difference, while determining a retrospective loss based on historical prediction outputs generated by historical parameter states of the deep neural network model; The value of the retrospective loss is dynamically adjusted according to the margin of the retrospective loss to complete the calculation of the loss function, the error calculated based on the loss function is back-propagated into the deep neural network model, the weight value of at least one of the multiple neurons of the deep neural network model is updated, and the updated weight value is used as part of training the deep neural network model. The deep neural network model is iteratively calculated multiple times to achieve training of the deep neural network model.
[0014] Specifically, in step 5, the steps for calculating the carcinogenic risk are: Step 501: Obtain the pollutant dataset under each path and integrate it; Step 502: Calculate exposure doses under different pathways based on human body exposure pathways and exposure parameters; Step 503: Calculate the cancer risk; Step 504: Integrate the carcinogenic risks of different exposure pathways to obtain a total carcinogenic risk.
[0015] Specifically, in step 503, the calculation formula for the carcinogenic risk is: ; in, is the cancer risk for exposure pathway v, is the exposure dose for exposure route v, is the carcinogenic intensity coefficient, is the carcinogenicity ratio of the pollutant; In step 504, the total cancer risk is calculated as follows: ; in, is the cumulative carcinogenic risk, and n is the number of exposure pathways.
[0016] The present invention provides a health impact assessment system for the coking industry, which includes: a data acquisition module, a data simulation module, a data classification module, a deep neural network model building module and a risk assessment module; A data acquisition module is used to obtain historical environmental monitoring data around the coking plant, including meteorological data, air pollutant data, heavy metal content in the soil, and water pollutant data; A data simulation module, based on the meteorological data and the air pollutant data, uses a Gaussian model to simulate the diffusion pattern of air pollutants in the atmosphere. At the same time, based on the heavy metal content in the soil, a BP neural network model is used to simulate the distribution pattern of heavy metals in the soil. A hydrogeological model is constructed based on the water pollutant data to simulate the migration pattern of water pollutants in coking wastewater in the water body. a data classification module, which collects and preprocesses pollutant data under each pollution path based on the Gaussian model, the BP neural network model, and the hydrogeological model to obtain multiple pollutant data packets, performs outlier detection and removal on each of the pollutant data packets to obtain a pollutant data set, assigns the pollutant data set to different risk clusters based on preset rules, and configures a deep neural network model to output the risk cluster type to which each pollutant data set belongs; A deep neural network model building module is used to train the deep neural network model based on retrospective loss to improve the classification accuracy of the risk cluster type to which each pollutant data belongs, each risk cluster having a corresponding carcinogenicity proportion; The risk assessment module establishes the carcinogenic risk of pollutants from the coking plant to the human body under different pathways based on the carcinogenic proportion corresponding to each risk cluster.
[0017] Compared with the prior art, the beneficial effects of the technical solution of this application are at least as follows: The present invention includes multiple steps such as multi-path data collection, pollutant migration simulation, data classification processing, deep neural network model establishment and risk assessment and prediction, forming a complete coking plant pollutant health assessment system, which can simulate the migration of pollutants around the coking plant and make risk level assessment for the migration of pollutants. At the same time, it completes the calculation of carcinogenic risks based on pollutants under different exposure paths, completes the assessment of the carcinogenic risk of pollutants around the coking plant on human health, and improves the accuracy of risk assessment. Moreover, the multi-media exposure risk assessment method can not only comprehensively and accurately assess the impact of coking industry emissions on human health, but also provide scientific basis and decision-making support for environmental management and human health protection, and is highly targeted and innovative. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 This is a schematic diagram of an embodiment of a health impact assessment method for the coking industry in an embodiment of the present application; Figure 2 Flowchart of the steps of detecting and eliminating outliers in an embodiment of the present application; Figure 3 This is a flowchart of training a deep neural network model in an embodiment of the present application; Figure 4 This is a schematic diagram of the health impact assessment system for the coking industry in an embodiment of the present application. DETAILED DESCRIPTION
[0020] The following describes the specific process of the embodiment of the present application. Figure 1 In the embodiments of the present application, an embodiment of the health impact assessment method for the coking industry includes: Step 1: Acquire historical environmental monitoring data around the coking plant, wherein the historical environmental monitoring data includes meteorological data, air pollutant data, heavy metal content in soil, and water pollutant data.
[0021] Specifically, in step 1, when detecting organic matter in the environment, online analysis technology is used to continuously monitor the target compounds in the environment, and the target compounds in the environment include halogenated hydrocarbons, aromatic hydrocarbons, oxygen-containing organic matter and nitrogen-containing organic matter. The heavy metal content in the particulate matter is collected using an atmospheric particulate matter collector. The sample is extracted by acid leaching and analyzed by atomic absorption spectrometry to obtain heavy metal element indicators such as lead, cadmium and mercury.
[0022] Biosensors were used to detect heavy metals in soil.
[0023] When monitoring the concentration of pollutants in the surrounding wastewater, online monitoring equipment is used to monitor the wastewater discharge outlet. The collected indicators include pH value, suspended solids, chemical oxygen demand, ammonia nitrogen, cyanide and volatile phenols. Groundwater monitoring wells are set up around the coking plant, and the monitoring wells are set up using a grid distribution method. The monitoring items include pH value, heavy metals and volatile organic compounds.
[0024] When biomarkers were collected from biological samples of surrounding residents, the metabolites of benzene, PAHs and their derivatives were detected using high performance liquid chromatography.
[0025] Step 2: Based on the meteorological data and the air pollutant data, a Gaussian model is used to simulate the diffusion law of air pollutants in the atmosphere. At the same time, based on the heavy metal content in the soil, a BP neural network model is used to simulate the distribution law of heavy metals in the soil. Based on the water pollutant data, a hydrogeological model is constructed to simulate the migration law of water pollutants in coking wastewater in the water body.
[0026] The Gaussian diffusion model can more accurately simulate the diffusion patterns of pollutants such as benzene and benzene compounds in the atmosphere, providing a scientific basis for environmental management and health risk assessment.
[0027] When establishing the hydrogeological model, the MODFLOW model is used to complete the construction of the groundwater model, and the SWAT model is used to complete the construction of the surface water model. Then, based on the SWAT-MODFLOW coupling model, the surface water simulation capabilities of SWAT and the groundwater simulation capabilities of MODFLOW are combined to more comprehensively consider the interaction between surface water and groundwater, construct hydrogeological models of groundwater and surface water, and simulate the migration and transformation process of pollutants in coking wastewater in water bodies, providing a scientific basis for environmental management and pollution control.
[0028] Step 3: Based on the Gaussian model, the BP neural network model and the hydrogeological model, pollutant data under each pollution path is collected and preprocessed to obtain multiple pollutant data packets.
[0029] Specifically, the pollutant data collected under each pollution path is defined as raw data, and the raw data is cleaned. The cleaned data is then denoised using Kalman filtering, and the data missing points are processed for the denoised data. Finally, the data under each pollution path are integrated to form a pollutant data packet corresponding to each pollution path.
[0030] The step of detecting and removing outliers from each pollutant data packet to obtain a pollutant data set includes: S31. Use different discretization strategies to discretize the multi-index data within the pollutant data packet under each of the pollution paths to form multiple discretized data groups, and then combine the multiple discretized data groups and multiple candidate values of the LOF parameter K one by one to form multiple evaluation models that correspond one to one with the multiple discretized data groups.
[0031] S32. Based on the multiple evaluation models, the first scores of each indicator data in each discretized data group are calculated respectively. The first scores here are used for the recognition rate calculation of the subsequent evaluation models. Through these first scores, the performance of each evaluation model in identifying outliers and normal values can be evaluated. The recognition rate of each evaluation model is calculated based on the first scores, and the evaluation model corresponding to the maximum recognition rate is defined as the outlier evaluation model.
[0032] S33. Calculate the second score of each data in the multiple pollutant data packets based on the outlier evaluation model. The second score here is calculated under the selected outlier evaluation model. Therefore, it reflects the degree of abnormality of each data point under the optimal model. This score is used for outlier detection, and each second score is compared with a preset threshold. If the second score is greater than the preset threshold, it indicates that the indicator data is an outlier. After traversing all the data, the indicator data of the outlier is eliminated, and the normal data is integrated to form the pollutant data set.
[0033] Please refer to the flowchart of the steps for outlier detection and removal. Figure 2 It is worth noting that in step S32, each discretized data group contains multiple indicator data. These indicator data are discretized and represent different features or dimensions in the pollutant data packet. The data scoring calculation module is used to calculate the first score of each data in the multiple pollutant data packets. The first score of each indicator data represents the LOF score of each indicator data point, which indicates the degree of abnormality of the data point in its local neighborhood. Therefore, the steps of calculating the LOF score of each indicator data point using the LOF algorithm are as follows: First, determine the neighborhood and select K nearest neighbors, where K is the LOF parameter. Then, calculate the local density within each neighborhood. By comparing the density difference between the data point and its neighborhood, calculate the LOF score. The higher the LOF score, the more likely the data point is an outlier. The score formula is: ; in, are the K nearest neighbors of x, is the reachable distance from x to y, is the local reachability density of y.
[0034] It is worth noting that the "first score" in the data scoring calculation stage is used to evaluate the performance of different evaluation models and help select the optimal evaluation model. The "second score" in the outlier detection stage is used to determine whether each data point is an outlier under the optimal evaluation model. Therefore, the calculation background, usage purpose and threshold setting of the two stages are significantly different.
[0035] In step 3, allocating the pollutant dataset to different risk clusters based on preset rules includes: Step 34: perform a toxicity assessment on all the pollutant data, set a risk limit for each of the pollutant data, divide the pollutant data into risk clusters based on the risk limit, and assign a different cluster identifier to each of the pollutant data. The risk clusters are defined as acute toxicity, chronic toxicity, or carcinogenicity.
[0036] Step 35: Based on the K-Means clustering algorithm, the number of clusters and the health risk level represented by each cluster are set. The K-Means clustering algorithm performs cluster analysis on the pollutant data and different risk clusters. When the preset number of clusters is reached, the cluster analysis is terminated, completing the cluster analysis of the pollutants and the risk clusters to which they belong.
[0037] At the same time, a deep neural network model is configured to output the risk cluster type to which each pollutant data belongs.
[0038] Step 4: Training the deep neural network model based on retrospective loss, including: Step 41: Initialize the deep neural network model, which includes multiple neurons, assign an initial weight value to each neuron, and give the deep neural network model a classification task of classifying pollutant data into different risk clusters.
[0039] Step 42: Obtain a training data set representing the initial output and the expected output.
[0040] Step 43: Preheat training the deep neural network model based on the training data set.
[0041] Step 44: After completing the warm-up training, enter the formal training.
[0042] Step 45: When the difference between the current predicted output and the expected output is within a preset difference threshold, the training process is terminated.
[0043] The flowchart for training a deep neural network model can be found in Figure 3It is worth noting that, during the preheating training, a warm-up iterative calculation is performed on the deep neural network model, and the deep neural network model defines the clustering data set of the pollutant data and the risk cluster to which it belongs as the initial classification output, and defines the clustering data set of the pollutant data and the risk cluster to which it belongs based on the K algorithm as the expected classification output, and determines the first difference between the initial classification output and the expected classification output. Subsequently, the loss function of the classification task of the deep neural network model is determined based on the first-order difference of the deep neural network model, and one or more parameters of the deep neural network model are updated according to the loss function of the indication task.
[0044] It is worth noting that the first-order difference refers to the rate of change of the difference between the model output and the target output of the deep neural network model, which is used to measure the gap between the predicted value of the deep neural network model and the true value. In the warm-up training stage, the first-order difference can be used to measure the degree of improvement of the deep neural network model in each iteration, that is, the difference between the predicted value of the current iteration and the predicted value of the previous iteration. In the warm-up training stage, by calculating the first-order difference, the parameters of the model can be quickly adjusted to gradually adapt to the data.
[0045] When entering formal training, the deep neural network model generates the current predicted output based on the current input data, and calculates the difference between the current predicted output and the expected classification output.
[0046] The loss of the classification task is calculated based on the difference, and the retrospective loss is determined by historical prediction outputs generated by historical parameter states of the deep neural network model.
[0047] The value of the retrospective loss is dynamically adjusted according to the margin of the retrospective loss to complete the calculation of the loss function, the error calculated based on the loss function is back-propagated into the deep neural network model, the weight value of at least one of the multiple neurons of the deep neural network model is updated, and the updated weight value is used as part of training the deep neural network model. The deep neural network model is iteratively calculated multiple times to achieve training of the deep neural network model.
[0048] The deep neural network model is subjected to convergence training. Based on this, after each pollutant sample under different paths is imported into the deep neural network model, it can be accurately and efficiently assigned to the corresponding risk cluster, thereby improving the prediction accuracy of the risk cluster to which each pollutant sample belongs, thereby completing the risk level assessment of the pollutant, and the carcinogenic proportion value W of the pollutant sample can be determined according to the risk cluster, where W represents the carcinogenic value corresponding to the risk cluster. This can effectively perform cluster analysis on pollutants and complete the diagnostic assessment of carcinogenic risks based on the risk cluster.
[0049] Step 5: Based on the carcinogenic proportion corresponding to each risk cluster, the carcinogenic risk of pollutants from the coking plant under different pathways to the human body is determined. The steps for calculating the carcinogenic risk are: Step 501: Obtain the pollutant dataset under each path and integrate it.
[0050] Step 502: Calculate exposure doses under different pathways based on the human body's exposure pathways and exposure parameters, specifically: (1) Calculation formula for respiratory exposure dose: ; in, is the concentration of pollutants in the atmosphere (mg / m 3 ), Respiratory rate (m³ / day), Frequency of exposure (days / year), Duration of exposure (years), Weight (kg), Average exposure time (days).
[0051] (2) Calculation formula for drinking water exposure dose: ; in, Concentration of pollutants in water (mg / L), Water intake (L / day).
[0052] (3) Calculation formula for skin contact exposure dose: ; in, Skin contact area (cm²).
[0053] Step 503: Calculate the carcinogenic risk based on the exposure dose and the carcinogenicity ratio of the pollutant. The carcinogenic risk calculation formula is: ; in, is the cancer risk for exposure pathway v, is the exposure dose for exposure route v (mg / kg·day), is the carcinogenic intensity coefficient (mg / kg·day), is the carcinogenicity ratio of the pollutant.
[0054] Step 504: Integrate the carcinogenic risks of different exposure pathways to obtain a total carcinogenic risk. The total carcinogenic risk is calculated as follows: ; in, is the total carcinogenic risk, and n is the number of exposure pathways.
[0055] It is worth noting that in step 501, the steps for integrating the pollutant dataset are specifically as follows: The concentration distribution data of air pollutants are obtained from the Gaussian diffusion model, and the data are organized into a time series format so that the correspondence between time and space is clear. At the same time, the second concentration distribution data of soil heavy metals are obtained from the BP neural network model, and the third concentration distribution data of surface water and groundwater pollutants are obtained from the hydrogeological model. The second concentration distribution data and the third concentration distribution data are organized into a time series format and aligned with the time and space dimensions of the air pollutant data to form a complete pollutant exposure data set.
[0056] In this example, since the exposure routes are breathing, drinking water and skin contact, the total carcinogenic risk in this example is: ; is the carcinogenicity weight under the respiratory pathway, is the carcinogenic weight under the drinking water pathway, is the carcinogenicity weight for skin contact.
[0057] based on , to determine the carcinogenic risk, if >1×10 −6 , it indicates a cancer-causing health risk.
[0058] Based on this, the carcinogenic risk of pollutants around the coking plant can be calculated according to different pollutant exposure pathways, and the carcinogenicity determination of health risks can be completed.
[0059] It is worth noting that the carcinogenic risk of pollutants calculated based on the embodiment can intuitively display the health risk levels in different regions, which can provide a basis for environmental management and public health protection. Moreover, it is convenient to identify the main pollution sources and key pollutants in the coking plant emissions and formulate corresponding emission reduction measures for different pollution sources.
[0060] The above describes the health impact assessment method for the coking industry in the embodiment of the present application. The following describes the health impact assessment system for the coking industry in the embodiment of the present application. Figure 4 In the embodiments of the present application, an embodiment of a health impact assessment system for the coking industry specifically includes: a data acquisition module, a data simulation module, a data classification module, a deep neural network model building module and a risk assessment module.
[0061] Specifically, the data acquisition module obtains historical environmental monitoring data around the coking plant, and the historical environmental monitoring data includes meteorological data, air pollutant data, heavy metal content in the soil, and water pollutant data.
[0062] The data simulation module uses a Gaussian model to simulate the diffusion pattern of air pollutants in the atmosphere based on the meteorological data and the air pollutant data. At the same time, based on the heavy metal content in the soil, a BP neural network model is used to simulate the distribution pattern of heavy metals in the soil. A hydrogeological model is constructed based on the water pollutant data to simulate the migration pattern of water pollutants in coking wastewater in the water body.
[0063] The data classification module collects pollutant data under each pollution path and preprocesses it based on the Gaussian model, the BP neural network model and the hydrogeological model to obtain multiple pollutant data packets, performs outlier detection and elimination on each of the pollutant data packets to obtain a pollutant data set, and assigns the pollutant data set to different risk clusters based on preset rules. At the same time, a deep neural network model is configured to output the risk cluster type to which each pollutant data belongs.
[0064] A deep neural network model establishment module trains the deep neural network model based on retrospective loss to improve the classification accuracy of the risk cluster type to which each pollutant data belongs, and each risk cluster has a corresponding carcinogenic proportion.
[0065] The risk assessment module establishes the carcinogenic risk of pollutants from the coking plant to the human body under different pathways based on the carcinogenic proportion corresponding to each risk cluster.
[0066] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A health impact assessment method for the coking industry, characterized in that: The specific steps include: Step 1: Acquire historical environmental monitoring data around the coking plant, wherein the historical environmental monitoring data includes meteorological data, air pollutant data, heavy metal content in soil, and water pollutant data; Step 2: Based on the meteorological data and the air pollutant data, a Gaussian model is used to simulate the diffusion pattern of air pollutants in the atmosphere. At the same time, based on the heavy metal content in the soil, a BP neural network model is used to simulate the distribution pattern of heavy metals in the soil. Based on the water pollutant data, a hydrogeological model is constructed to simulate the migration pattern of water pollutants in coking wastewater in the water body; Step 3: Based on the Gaussian model, the BP neural network model, and the hydrogeological model, pollutant data under each pollution path is collected and preprocessed to obtain multiple pollutant data packets. Outliers are detected and removed for each of the pollutant data packets to obtain a pollutant data set. The pollutant data sets are assigned to different risk clusters based on preset rules. A deep neural network model is configured to output the risk cluster type to which each pollutant data set belongs. Step 4: training the deep neural network model based on retrospective loss to improve the classification accuracy of the risk cluster type to which each pollutant data belongs, each risk cluster having a corresponding carcinogenicity proportion; Step 5: Based on the carcinogenic proportion corresponding to each risk cluster, the carcinogenic risk of pollutants from the coking plant in different pathways to the human body is established.
2. The health impact assessment method for the coking industry according to claim 1, characterized in that: In step 3, the pollutant data under each pollution path is collected and pre-processed to obtain multiple pollutant data packets, including: The pollutant data under each pollution path is defined as raw data, and the raw data is cleaned. The cleaned data is then denoised using Kalman filtering, and the data missing points are processed for the denoised data. Finally, the data under each pollution path are integrated to form a pollutant data packet corresponding to each pollution path.
3. The health impact assessment method for the coking industry according to claim 1, characterized in that: In step 3, performing outlier detection and removal on each pollutant data packet to obtain a pollutant data set includes: S31, using different discretization strategies to discretize the multi-index data within the pollutant data packet under each of the pollution paths to form a plurality of discretized data groups, and then combining the plurality of discretized data groups with a plurality of candidate values of the LOF parameter K one by one to form a plurality of evaluation models corresponding one by one to the plurality of discretized data groups; S32. Based on the multiple evaluation models, respectively calculate a first score for each indicator data in each discretized data group, respectively calculate a recognition rate for each evaluation model based on the first scores, and define the evaluation model corresponding to the maximum recognition rate as an outlier evaluation model; S33. Calculate the second score of each data in the multiple pollutant data packets based on the outlier evaluation model, and compare each second score with a preset threshold. If the second score is greater than the preset threshold, it indicates that the indicator data is an outlier. After traversing all the data, eliminate the indicator data of the outlier and integrate the normal data to form the pollutant data set.
4. The health impact assessment method for the coking industry according to claim 1, characterized in that: In step 3, allocating the pollutant dataset to different risk clusters based on preset rules includes: Step 34: Perform a toxicity assessment on all the pollutant data, set a risk limit for each pollutant data, divide the pollutant data into risk clusters based on the risk limit, and assign a different cluster identifier to each pollutant data. The risk cluster is defined as acute toxicity, chronic toxicity, or carcinogenicity. Step 35: Based on the K-Means clustering algorithm, the number of clusters and the health risk level represented by each cluster are set. The K-Means clustering algorithm performs cluster analysis on the pollutant data and different risk clusters. When the preset number of clusters is reached, the cluster analysis is terminated, completing the cluster analysis of the pollutants and the risk clusters to which they belong.
5. The health impact assessment method for the coking industry according to claim 1, characterized in that: In step 4, training the deep neural network model based on retrospective loss includes: Step 41: Initialize the deep neural network model, which includes multiple neurons, assign an initial weight value to each neuron, and assign the deep neural network model a classification task of classifying pollutant data into different risk clusters; Step 42: Obtain a training data set representing the initial output and the expected output; Step 43: preheat training the deep neural network model based on the training data set; Step 44: After completing the warm-up training, enter the formal training and train the deep neural network model through multiple iterations; Step 45: When the difference between the current predicted output and the expected output is within a preset difference threshold, the training process is terminated.
6. The health impact assessment method for the coking industry according to claim 5, characterized in that: The warm-up training includes: A warm-up iterative calculation is performed on the deep neural network model, wherein the deep neural network model defines the clustering data set of the pollutant data and the risk clusters to which it belongs as the initial classification output, and defines the clustering data set of the pollutant data and the risk clusters to which it belongs based on the K algorithm as the expected classification output, and determines the first difference between the initial classification output and the expected classification output. Subsequently, the loss function of the classification task of the deep neural network model is determined based on the first-order difference of the deep neural network model, and one or more parameters of the deep neural network model are updated according to the loss function of the indication task.
7. The health impact assessment method for the coking industry according to claim 6, characterized in that: The formal training includes: causing the deep neural network model to generate the current predicted output based on the current input data, and calculating the difference between the current predicted output and the expected classification output; Calculating a loss for the classification task based on the difference, while determining a retrospective loss based on historical prediction outputs generated by historical parameter states of the deep neural network model; The value of the retrospective loss is dynamically adjusted according to the margin of the retrospective loss to complete the calculation of the loss function, the error calculated based on the loss function is back-propagated into the deep neural network model, the weight value of at least one of the multiple neurons of the deep neural network model is updated, and the updated weight value is used as part of training the deep neural network model. The deep neural network model is iteratively calculated multiple times to achieve training of the deep neural network model.
8. The health impact assessment method for the coking industry according to claim 1, characterized in that: In step 5, the steps for calculating the carcinogenic risk are: Step 501: Obtain the pollutant dataset under each path and integrate it; Step 502: Calculate exposure doses under different pathways based on human body exposure pathways and exposure parameters; Step 503: Calculate the cancer risk; Step 504: Integrate the carcinogenic risks of different exposure pathways to obtain a total carcinogenic risk.
9. The health impact assessment method for the coking industry according to claim 8, characterized in that: In step 503, the formula for calculating the carcinogenic risk is: ; in, is the cancer risk for exposure pathway v, is the exposure dose for exposure route v, is the carcinogenic intensity coefficient, is the carcinogenicity ratio of the pollutant; In step 504, the total cancer risk is calculated as follows: ; in, is the total carcinogenic risk, and n is the number of exposure pathways.
10. A health impact assessment system for the coking industry, implementing the health impact assessment method for the coking industry as claimed in claim 1, characterized in that: The system includes: data acquisition module, data simulation module, data classification module, deep neural network model building module and risk assessment module; A data acquisition module is used to obtain historical environmental monitoring data around the coking plant, including meteorological data, air pollutant data, heavy metal content in the soil, and water pollutant data; A data simulation module, based on the meteorological data and the air pollutant data, uses a Gaussian model to simulate the diffusion pattern of air pollutants in the atmosphere. At the same time, based on the heavy metal content in the soil, a BP neural network model is used to simulate the distribution pattern of heavy metals in the soil. A hydrogeological model is constructed based on the water pollutant data to simulate the migration pattern of water pollutants in coking wastewater in the water body. a data classification module, which collects and preprocesses pollutant data under each pollution path based on the Gaussian model, the BP neural network model, and the hydrogeological model to obtain multiple pollutant data packets, performs outlier detection and removal on each of the pollutant data packets to obtain a pollutant data set, assigns the pollutant data set to different risk clusters based on preset rules, and configures a deep neural network model to output the risk cluster type to which each pollutant data set belongs; A deep neural network model building module is used to train the deep neural network model based on retrospective loss to improve the classification accuracy of the risk cluster type to which each pollutant data belongs, each risk cluster having a corresponding carcinogenicity proportion; The risk assessment module establishes the carcinogenic risk of pollutants from the coking plant to the human body under different pathways based on the carcinogenic proportion corresponding to each risk cluster.
Citation Information
Patent Citations
Pollutant health risk assessment method and device
CN115274108A