Industrial gathering area pollution risk management method and system based on multi-source data
Through the integrated learning model, the pollution characteristics and feature fusion data of industrial agglomeration areas are analyzed, and the inaccurate evaluation in the existing technology is solved, and intelligent and real-time monitoring and management of pollution risks in industrial agglomeration areas are realized.
Patent Information
- Application Number
- CN202510098987.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing pollution risk assessment methods in industrial agglomeration areas lack real-time and comprehensiveness, making it difficult to effectively deal with rapidly changing pollution situations, and the inability to fully integrate multi-source data leads to inaccurate judgments.
The integrated learning model based on multi-source data is adopted to obtain and process environmental monitoring data, hydrological geological data, industrial activity data, pollution event records and remote sensing image data, and construct feature fusion data such as comprehensive environmental indicators, pollutant migration matrix and industrial impact indicators. The integrated learning model is used for intelligent analysis to determine pollution risks.
It has achieved intelligent, targeted and real-time monitoring and management of pollution risks in industrial agglomerations, and improved the accuracy and comprehensiveness of pollution risk assessment.
Smart Images

Figure CN119990769A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of pollution risk management, and in particular to a pollution risk management method and system for industrial clusters based on multi-source data. Background Art
[0002] With the acceleration of industrialization, industrial agglomeration areas have become an important driving force for economic development, but they have also brought about serious environmental pollution problems. In industrial agglomeration areas, the concentrated development of chemical, metallurgical, mechanical and other industries has led to varying degrees of pollution of ecological environments such as soil, groundwater and air. These pollutants not only affect the health of local residents, but may also migrate and spread through soil and water bodies, posing a potential threat to a wider range of ecological environments. Therefore, timely and accurate identification and assessment of high-risk pollution areas is crucial for effective pollution control and management.
[0003] At present, the assessment methods for pollution risks in industrial agglomeration areas mainly rely on traditional monitoring and analysis methods, which often lack real-time and comprehensiveness and are difficult to effectively respond to rapidly changing pollution situations. In addition, existing monitoring systems are usually unable to fully integrate multi-source data, resulting in inaccurate judgments on pollution situations. Therefore, there is an urgent need for an intelligent auxiliary identification system to monitor and manage pollution risks in industrial agglomeration areas. Summary of the invention
[0004] The present application provides a method and system for pollution risk management of industrial agglomeration areas based on multi-source data, which can analyze the pollution situation in industrial agglomeration areas based on multi-source data, and is conducive to monitoring and managing pollution risks in industrial agglomeration areas.
[0005] In a first aspect, the present application provides a method for managing pollution risk in industrial agglomeration areas based on multi-source data. The method comprises:
[0006] Obtain multiple pollution characteristic data from multiple sources in industrial agglomeration areas;
[0007] Constructing a plurality of feature fusion data based on the pollution feature data;
[0008] Analyze the matching degree data of each intelligent sub-model in the pre-trained integrated learning model for each of the pollution feature data and feature fusion data;
[0009] Determine the input strategy of the pollution feature data and feature fusion data to the intelligent sub-model of the integrated learning model based on the matching degree data;
[0010] Based on the input strategy, the pollution feature data and feature fusion data are input into the integrated learning model to obtain the pollution risk data of the industrial agglomeration area.
[0011] By adopting the above technical solution, it is possible to analyze pollution feature data and feature fusion data separately and in a targeted manner based on the integrated learning model, so as to intelligently and reasonably determine the pollution risk status within the industrial agglomeration area, which can then help monitor and manage the pollution risks within the industrial agglomeration area.
[0012] Further, the pollution characteristic data includes one or more of environmental monitoring data, hydrogeological data, industrial activity data, pollution event records and remote sensing image data;
[0013] The feature fusion data include comprehensive environmental indicators, pollutant migration matrix, industrial impact indicators and pollution trend indicators;
[0014] The comprehensive environmental index is a quantitative value that reflects the comprehensive environmental situation and is calculated and determined by environmental monitoring data from multiple sources;
[0015] The pollutant migration matrix is a matrix that reflects the spatial migration law of pollutants and is determined by analyzing hydrogeological data and environmental monitoring data;
[0016] The industrial impact index is a quantitative value that reflects the degree of impact of industrial activities on the environment, and is determined by analyzing industrial activity data and environmental monitoring data;
[0017] The pollution trend index is a quantitative value that reflects the changing trend of each pollutant and is determined by analyzing environmental monitoring data.
[0018] Furthermore, the analysis of the matching degree data of each intelligent sub-model in the pre-trained integrated learning model for each of the pollution feature data and feature fusion data includes:
[0019] Obtain the data quality and data characteristics of pollution feature data and feature fusion data;
[0020] Analyze the data importance and data relevance of pollution feature data and feature fusion data to each intelligent sub-model in the integrated learning model based on data features;
[0021] Determine the basic matching degree of the pollution feature data and the feature fusion data for each intelligent sub-model in the integrated learning model based on the data quality, data importance and data relevance;
[0022] Combine the basic matching degree and the pre-acquired data association coefficient to determine the comprehensive matching degree of the pollution feature data and the feature fusion data for each intelligent sub-model in the integrated learning model;
[0023] The matching degree data is equal to the maximum value of the basic matching degree and the comprehensive matching degree.
[0024] Furthermore, the pollution feature data and feature fusion data are input into the integrated learning model based on the input strategy to obtain the pollution risk data of the industrial agglomeration area, including:
[0025] Determine the output of each intelligent sub-model under the action of the input as pollution risk sub-data;
[0026] Determine the model fitness of the intelligent sub-model based on the pollution feature data and / or feature fusion data actually input into the intelligent sub-model and the matching degree data of the intelligent sub-model;
[0027] Determine the sub-model weight of each intelligent sub-model according to the model fitness and the pre-acquired model reliability of each intelligent sub-model and the data characteristics of the pollution feature data and feature fusion data input into the intelligent sub-model;
[0028] The pollution risk data is determined by combining the sub-model weights of all intelligent sub-models and the pollution risk sub-data.
[0029] Furthermore, the model parameters of the intelligent sub-model in the integrated learning model are associated with pre-acquired data distribution characteristics, and the data distribution characteristics reflect the spatial and / or temporal distribution of the specified pollution characteristic data.
[0030] Furthermore, the determining of the basic matching degree of the pollution feature data and the feature fusion data for each intelligent sub-model in the integrated learning model based on the data quality, data importance and data relevance includes:
[0031] BM=w1Q(w2Imp+w3r)
[0032] Where BM is the basic matching degree for the intelligent sub-model, Q is the data quality, Imp is the data importance for the intelligent sub-model, r is the data relevance for the intelligent sub-model, and w1, w2, and w3 are the preset matching calculation weights for data quality, data importance, and data relevance, respectively.
[0033] Furthermore, the method of combining the basic matching degree and the pre-acquired data association coefficient to determine the comprehensive matching degree of the pollution feature data and the feature fusion data for each intelligent sub-model in the integrated learning model includes:
[0034] CM ij =αBM ij (1-e -x )
[0035]
[0036] In the formula, CM ij is the comprehensive matching degree of the i-th pollution feature data or feature fusion data to the j-th intelligent sub-model, BMij is the basic matching degree of the i-th pollution feature data or feature fusion data to the j-th intelligent sub-model, BM kj is the basic matching degree of the kth pollution feature data or feature fusion data to the jth intelligent sub-model, r ik is the data correlation coefficient between the i-th pollution feature data or feature fusion data and the k-th pollution feature data or feature fusion data, i≠k, α and β are respectively the preset first matching calculation coefficient and the second matching calculation coefficient.
[0037] Furthermore, the matching degree data of the intelligent sub-model based on the pollution feature data and / or feature fusion data actually input into the intelligent sub-model to determine the model adaptability of the intelligent sub-model includes:
[0038]
[0039] In the formula, A i is the model fitness of the intelligent sub-model, n i is the total number of pollution feature data and / or feature fusion data input into the i-th intelligent sub-model, MD ij It is the jth matching degree data in the pollution feature data and / or feature fusion data input into the intelligent sub-model.
[0040] Further, the determining of the sub-model weight of each intelligent sub-model according to the model fitness and the pre-acquired model reliability of each intelligent sub-model and the data features of the pollution feature data and feature fusion data input into the intelligent sub-model includes:
[0041]
[0042] γ1+γ2+γ3=1
[0043] Where n i GD is the total number of pollution feature data and / or feature fusion data input into the i-th intelligent sub-model. ij GT is the jth data attention in the contaminated feature data and / or feature fusion data in the input intelligent sub-model, l is the feature attention degree pre-acquired relative to the lth data feature, n is the number of intelligent sub-models, is the sub-model weight of the i-th intelligent sub-model, A i is the model fitness of the ith intelligent sub-model, R i is the model reliability of the i-th intelligent sub-model.
[0044] In a second aspect, the present application provides an industrial cluster pollution risk management system based on multi-source data, wherein the system is used to execute any one of the methods described in the first aspect above.
[0045] In summary, this application at least has the following beneficial effects:
[0046] A pollution risk management method and system for industrial agglomeration areas based on multi-source data are provided, which can intelligently and purposefully analyze the multi-source data of industrial agglomeration areas based on an integrated learning model, thereby determining the pollution risks within the industrial agglomeration areas.
[0047] It should be understood that the contents described in the Summary of the Invention are not intended to limit the key or important features of the embodiments of the present application, nor are they intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The above and other features, advantages and aspects of the embodiments of the present application will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0049] Figure 1 A flow chart of a method for managing pollution risk in industrial clusters based on multi-source data in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0051] In addition, the term "and / or" in this article is only a description of the association relationship between the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0052] The present application provides a pollution risk management method and system for industrial agglomeration areas based on multi-source data, which can use an integrated learning model to perform targeted processing on pollution feature data and feature fusion data, thereby comprehensively, accurately, intelligently and in real time analyzing the pollution risks within the industrial agglomeration areas, so as to facilitate the management of pollution risks within the industrial agglomeration areas.
[0053] In a first aspect, the present application discloses a method for managing pollution risk in industrial clusters based on multi-source data. The method can be applied to a server.
[0054] Figure 1 A flow chart of a method for managing pollution risk in industrial clusters based on multi-source data in an embodiment of the present application is shown.
[0055] Reference Figure 1 , the method specifically comprises the following steps:
[0056] S110: Acquire multiple pollution characteristic data from multiple sources in industrial agglomeration areas.
[0057] Pollution characteristic data are data collected in multiple ways and from multiple sources that reflect the pollution situation in industrial agglomeration areas. They can reflect the specific conditions of industrial agglomeration areas, hydrological and geographical conditions, and the presence and level of nearby pollution receptors such as water source protection areas. In an example, pollution characteristic data include environmental monitoring data, hydrogeological data, industrial activity data, pollution event records, and remote sensing image data.
[0058] Specifically, environmental monitoring data is used to reflect the real-time pollution situation in industrial agglomeration areas, which may include soil pollutant concentrations, groundwater pollutant concentrations, air pollution data, surface water quality monitoring data, etc. with timestamps; hydrogeological data is used to analyze the migration and diffusion behavior of pollutants in the site, which may include spatial geological structure and soil characteristics, groundwater level, hydraulic gradient, permeability, effective porosity, longitudinal and lateral diffusion coefficients, etc., which carry timestamps and reflect the specific situation of hydrogeological data at different times; industrial activity data is used to reflect potential pollution sources in industrial activities, which may include raw material and waste management data, underground storage tank and pipeline layout data, industrial wastewater discharge data, etc. with timestamps; pollution event records are used to reflect leakage accidents, pollution events, enterprise penalty records and other data that have occurred in industrial agglomeration areas in the past. In addition to pollution event conditions and time records, they may also include event handling plans, event handling results, pollution residual data, etc.; remote sensing image data is used to reflect important equipment and their evolution in industrial agglomeration areas, which carry timestamps.
[0059] It should be understood that in the process of surface water and groundwater pollution analysis, the impact of air pollution data is generally small or even negligible. In actual analysis tasks, it is also necessary not to analyze air pollution-related content, or to assign a smaller coefficient to air pollution-related parameters. The previous and subsequent contents in the embodiments of the present application can be based on this consideration.
[0060] The aforementioned pollution characteristic data need to undergo data cleaning, denoising and data standardization processing so that all the final pollution characteristic data have the same dimension and range. Of course, some pollution characteristic data are expressed as single-dimensional quantitative values, and some pollution data are expressed as multi-dimensional or even high-dimensional vector values.
[0061] S120: constructing a plurality of feature fusion data based on the pollution feature data.
[0062] Feature fusion data reflects the result of multi-source fusion of pollution feature data or fusion between different pollution feature data. In an example, feature fusion data may include comprehensive environmental indicators, pollutant migration matrix, industrial impact data and pollution trend indicators.
[0063] The comprehensive environmental index is a quantitative value reflecting the comprehensive environmental situation, which is calculated and determined by multi-source environmental monitoring data. Specifically, the fusion of multi-source environmental monitoring data can be realized based on convolutional neural networks. For example, if the soil pollutant concentration sequence is C soil (t) = [C soil (1,t),C soil (2,t),…,C soil (n soil ,t)](where n soil is the number of soil pollutant types), the concentration sequence of groundwater pollutants is C groundwater (t), the air pollution data sequence is A air (t), the surface water quality monitoring data sequence is C surfacewater (t). These data are constructed into a 4-dimensional spatiotemporal data cube X according to time series and spatial location (if there is spatial distribution information). env (t), where the dimensions are pollutant type, spatial location (optional), time, and environmental medium type (soil, groundwater, air, surface water). The CNN model structure includes convolutional layers, pooling layers, and fully connected layers. Assume that the convolution kernel size of the convolutional layer is k, the step size is s, and the padding is p. After the convolutional layer and the pooling layer, the feature map F is obtained. env The fully connected layer converts the feature map into a fused comprehensive environmental index I env (t). The CNN training process uses the back propagation algorithm to minimize the loss function L env , for example, the mean square error loss function:
[0064]
[0065] Where N is the number of training samples, It is the real comprehensive environmental index result.
[0066] The pollutant migration matrix is a matrix that reflects the spatial migration rules of pollutants and is determined by analyzing hydrogeological data and environmental monitoring data.
[0067] Specifically, first, define variables and matrices. Suppose there are n pollutants, m monitoring locations, and T time points. Let C ijtrepresents the concentration of the i-th pollutant at the j-th location and the t-th time point (i = 1, 2, ..., n; j = 1, 2, ..., m; t = 1, 2, ..., T). The pollutant migration matrix M is a three-dimensional matrix with dimensions n × m × T, where M ijt =C ijt .
[0068] Secondly, the influence of geological and hydrological factors on migration is considered, including the influence of groundwater flow velocity and the influence of precipitation and surface runoff. Regarding the influence of groundwater flow velocity, it is assumed that the groundwater flow velocity is v j (Unit: m / s) At the jth location, for the migration of groundwater pollutants (assuming it is the kth pollutant), the concentration change in the adjacent time interval Δt can be approximated by the following simplified form of the convection-diffusion equation (assuming one-dimensional flow and ignoring diffusion): C k,j,t+1 =C k,j-vj Δt,t, where jv j Δt represents the upstream position calculated based on the flow velocity and time interval. Regarding the impact of precipitation and surface runoff, for surface water quality and soil pollutants, precipitation P (unit: mm) and surface runoff coefficient α (dimensionless) will affect the migration of pollutants. Assuming the unit area is A (unit: m 2 ), surface runoff flow Q = P × α × A. For the lth type of pollutant in surface water, the concentration change caused by runoff between adjacent positions j and j+1 can be calculated using the law of conservation of mass: Where V j is the volume of water at position j. For soil pollutants (assuming the sth type), precipitation will lead to leaching, causing the pollutants to migrate to the lower soil layer. Assuming the soil porosity is θ (dimensionless) and the soil layer thickness is h (unit: m), the change in pollutant concentration within the time interval Δt can be expressed as:
[0069] Thirdly, considering the influence of geological structure and soil characteristics, the soil adsorption coefficient K d (Unit: L / kg) reflects the soil's adsorption capacity for pollutants. For the migration of the pth pollutant in the soil, adsorption will reduce its concentration in the soil pore water. Assume that the soil bulk density is ρ b (Unit: kg / m 3 ), the soil pore water content is θ w (dimensionless), according to the linear adsorption model, the pollutant concentration in soil pore water C w and the concentration of pollutants adsorbed on soil particles C s The relationship between them is: s =K d ×C w Total pollutant concentration Ctotal =C w ×θ w +C s ×ρ b , considering the adsorption effect, the migration equation of pollutants in the soil needs to be combined with the above relationship to correct the concentration changes.
[0070] Again, considering the exchange of air pollution data and pollutants, for gaseous pollutants (assuming the qth type), there is an exchange process at the soil-air interface and the water-air interface. Assume that the concentration of pollutants in the atmosphere is C a (Unit: mg / m 3 ), the gas exchange coefficient is k (unit: m / s), and at the interface per unit area A, according to the double membrane theory, the flux F of the pollutant at the interface (unit: mg / (m 2 ·s)) can be expressed as: F = k × (C a -C interface ), where C interface is the pollutant concentration at the interface, and its value is related to the pollutant concentration in the soil or water. For example, for the soil-atmosphere interface, the change in the pollutant concentration in the soil pore water can be expressed as:
[0071] Finally, regarding the update of the pollutant migration matrix, the influence of the above factors on the pollutant concentration is calculated. After each time step Δt, the element M in the pollutant migration matrix M is updated. ijt , to reflect the migration of pollutants in space and time.
[0072] The industrial impact index is a quantitative value that reflects the degree of impact of industrial activities on the environment and is determined by analyzing industrial activity data and environmental monitoring data.
[0073] Specifically, we first determine the change in pollutant concentration in soil, groundwater, air, and surface water. For example, regarding soil, suppose there are n types of soil pollutants. In the time interval [t1, t2], the change in the concentration of the i-th pollutant is The calculation formula is: in represents the concentration of the first soil pollutant at time t; for groundwater, for m groundwater pollutants, the change in the concentration of the jth pollutant in the same time interval is ΔC gwj The calculation formula is: in is the concentration of the jth groundwater pollutant at time t; for air, assuming there are k air pollutants, the change in the concentration of the first pollutant ΔC air, The calculation formula is: represents the concentration of the lth air pollutant at time t; for surface water, for p kinds of pollutants in surface water, the change in the concentration of the qth pollutant in the time interval [t1, t2] The calculation formula is: in is the concentration of the qth pollutant in surface water at time t.
[0074] Secondly, determine the contribution weights of potential pollution sources of industrial activities. For example, for raw material management, assuming there are r types of raw materials, the weight of the pollutant that the u-th raw material may produce is w rawu It can be determined based on its chemical properties, dosage, etc. For example, if a certain raw material will produce a large amount of heavy metal pollutants during the production process, its weight will be higher. For waste management, assuming there are s types of waste, the weight of the pollutant that the vth type of waste may produce is w wastev This can be determined based on its toxicity, treatment methods, etc. For example, hazardous waste will have a higher weight than general industrial waste. The total weight of the raw materials related to waste management data Assume that the underground storage tank and pipeline layout involves t risk points (such as tank leakage risk points, pipeline interfaces, etc.), and the weight of the w-th risk point is It can be determined based on its probability of leakage, potential leakage, etc. For example, an aging underground storage tank will have a higher leakage weight than a new one. The total weight of the underground storage tank associated with the pipeline layout data Assume that industrial wastewater discharge involves x pollutants, the weight of the yth pollutant in the wastewater is It can be determined based on its discharge volume, toxicity, etc. For example, wastewater discharge containing high concentrations of heavy metals has a higher weight. The total weight associated with industrial wastewater discharge data
[0075] The pollution trend index is a quantitative value reflecting the change trend of each pollutant, which is determined by analyzing the environmental monitoring data. For example, based on the environmental monitoring data, the slope of linear regression is used to determine the change rate of each pollutant in various environmental media such as soil, groundwater, air, and surface water, and the comprehensive change rate of each pollutant concentration is calculated based on the preset weights of different environmental media.
[0076] Of course, feature fusion data can also contain other content and adopt other specific fusion methods, which are not listed here. Feature fusion data is similar to pollution feature data, which is a single-dimensional quantitative value or a multi-dimensional vector value. It only needs to reflect the indicator data that needs to be reflected. In short, feature fusion data is determined based on pollution feature data. Compared with pollution feature data, it can better reflect certain obvious and integrated pollution conditions, which is conducive to the subsequent analysis of the pollution situation in industrial agglomeration areas.
[0077] S130: Analyze the matching degree data of each pollution feature data and feature fusion data with respect to each intelligent sub-model in the pre-trained integrated learning model.
[0078] The method of this step specifically includes: obtaining the data quality and data characteristics of the pollution feature data and the feature fusion data; analyzing the data importance and data relevance of the pollution feature data and the feature fusion data to each intelligent sub-model in the integrated learning model based on the data characteristics; determining the basic matching degree of the pollution feature data and the feature fusion data for each intelligent sub-model in the integrated learning model based on the data quality, data importance and data relevance; determining the comprehensive matching degree of the pollution feature data and the feature fusion data for each intelligent sub-model in the integrated learning model in combination with the basic matching degree and the pre-acquired data association coefficient; the matching degree data is equal to the maximum value of the basic matching degree and the comprehensive matching degree.
[0079] In the method of this step, data quality can be determined based on pre-constructed data quality assessment indicators such as completeness, accuracy, consistency, etc. Specifically, regarding completeness, let the total number of samples be N, and the number of valid samples in a certain contaminated feature data (or feature fusion data) be n, then the completeness Its value range is between 0 and 1. The closer to 1, the better the data integrity. Regarding accuracy: for data with known standard reference values (such as some calibrated environmental monitoring data with standard value comparison), let the number of comparison samples be m, the number of accurate samples be k, and the accuracy be The value range is also 0 to 1, indicating the degree of conformity between the data and the true value; Regarding consistency: it is evaluated by checking whether there are contradictions between the same data recorded from different sources (if there are multiple sources) or at different times. For example, for the same pollutant concentration data collected at the same location at different times, if the difference exceeds the reasonable fluctuation range (which can be determined by statistical analysis of the historical reasonable fluctuation range) for d times and the total number of comparisons is q, then the consistency is Its value is between 0 and 1. The higher the value, the better the data consistency. Regarding timeliness: it is measured by the interval t between the data collection time and the current analysis time (the unit can be days, hours, etc., determined according to the data characteristics) and the effective update cycle T of the data. That is, if the data is within the valid update period, the timeliness is 1, and if it exceeds the update period, it is delivered proportionally.
[0080] Data features can include numerical features and categorical features. Numerical features include statistical features such as mean, standard deviation, skewness, kurtosis, etc., and can also be distribution features presented by histograms, box plots, etc.
[0081] Regarding data importance, the feature importance of each data feature can be preset relative to each intelligent sub-model, and then the corresponding data importance can be comprehensively determined based on the feature importance of the data features carried by the contaminated feature data or feature fusion data.
[0082] Regarding data correlation, the frequency of occurrence of each data feature can be determined relative to the input record big data of each intelligent sub-model. The higher the frequency of occurrence of the data feature, the higher the correlation between the data feature and the intelligent sub-model. Accordingly, the data correlation between the contaminated feature data or feature fusion data and the intelligent sub-model can be determined based on the correlation between all data features and the intelligent sub-model.
[0083] The basic matching degree of the pollution feature data and the feature fusion data for each intelligent sub-model in the integrated learning model is determined based on the data quality, data importance and data relevance, including:
[0084] BM=w1Q(w2Imp+w3r)
[0085] Where BM is the basic matching degree for the intelligent sub-model, Q is the data quality, Imp is the data importance for the intelligent sub-model, r is the data relevance for the intelligent sub-model, and w1, w2, and w3 are the preset matching calculation weights for data quality, data importance, and data relevance, respectively.
[0086] The method of combining the basic matching degree and the pre-acquired data association coefficient to determine the comprehensive matching degree of the pollution feature data and the feature fusion data for each intelligent sub-model in the integrated learning model includes:
[0087] CM ij =αBM ij (1-e -x )
[0088]
[0089] In the formula, CM ij is the comprehensive matching degree of the i-th pollution feature data or feature fusion data to the j-th intelligent sub-model, BM ij is the basic matching degree of the i-th pollution feature data or feature fusion data to the j-th intelligent sub-model, BM kj is the basic matching degree of the kth pollution feature data or feature fusion data to the jth intelligent sub-model, r ik is the data correlation coefficient between the i-th pollution feature data or feature fusion data and the k-th pollution feature data or feature fusion data, i≠k, α and β are respectively the preset first matching calculation coefficient and the second matching calculation coefficient.
[0090] S140: Determine an input strategy of the pollution feature data and feature fusion data for an intelligent sub-model of an integrated learning model based on the matching degree data.
[0091] The integrated learning model in this step method is pre-built and trained, and can specifically include intelligent sub-models such as neural network models, LightGBM models, XGBoost models, and GBDT models.
[0092] A specific example of the method in this step is to preset a matching threshold for each intelligent sub-model, and use the pollution feature data and feature fusion data with matching data higher than the matching threshold as the input of the intelligent sub-model, so as to determine how all the pollution feature data and feature fusion data are input into the intelligent sub-model respectively.
[0093] S150: Inputting the pollution feature data and feature fusion data into the integrated learning model based on the input strategy to obtain pollution risk data of the industrial cluster area.
[0094] The method of this step specifically includes: determining that the output of each intelligent sub-model under the action of the input is pollution risk sub-data; determining the model fitness of the intelligent sub-model based on the matching data of the pollution feature data and / or feature fusion data actually input into the intelligent sub-model and oriented to the intelligent sub-model; determining the sub-model weight of each intelligent sub-model according to the model fitness and the pre-acquired model reliability of each intelligent sub-model and the data characteristics of the pollution feature data and feature fusion data input into the intelligent sub-model; and determining the pollution risk data by combining the sub-model weights and pollution risk sub-data of all intelligent sub-models.
[0095] In the method of this step, under the input strategy, the pollution feature data and the feature fusion data are respectively input into the intelligent sub-model within the integrated learning model, and the pollution risk sub-data are output under the action of the intelligent sub-model. The pollution risk sub-data is quantitative data reflecting the degree of pollution risk.
[0096] In the method of this step, the model fitness of the intelligent sub-model is determined based on all pollution feature data and / or feature fusion data actually input into the intelligent sub-model. Specifically, the model fitness of the intelligent sub-model is determined based on the matching degree data of the pollution feature data and / or feature fusion data actually input into the intelligent sub-model to the intelligent sub-model, and includes:
[0097]
[0098] In the formula, A i is the model fitness of the intelligent sub-model, n i is the total number of pollution feature data and / or feature fusion data input into the i-th intelligent sub-model, MD ijIt is the jth matching degree data in the pollution feature data and / or feature fusion data input into the intelligent sub-model.
[0099] In the method of this step, the model reliability is pre-acquired, which can be determined based on the accuracy and mean square error of the intelligent sub-model under the historical verification data set. For example, accuracy: For the classification task, let the sub-model s i The number of correctly predicted samples on the historical validation dataset (containing data samples with known true category labels) is TP i (TruePositive, true positive), the total number of samples is N i , then the accuracy About mean square error (MSE): For regression tasks, let submodel s i The predicted value on the historical validation dataset is The true value is y ij (j=1,2,…,k, k is the number of samples in the validation data set), then the mean square error Generally, MSE can be used i The reciprocal of is used to measure the reliability of the model (the larger the reciprocal, the higher the reliability, because the error is smaller), recorded as (For classification tasks, appropriate indicators such as accuracy can be used to convert them into numerical values to measure reliability).
[0100] In the method of this step, the sub-model weight of each intelligent sub-model is determined according to the model fitness and the pre-acquired model reliability of each intelligent sub-model and the data characteristics of the pollution feature data and feature fusion data input into the intelligent sub-model, including:
[0101]
[0102] γ1+γ2+γ3=1
[0103] Where n i GD is the total number of pollution feature data and / or feature fusion data input into the i-th intelligent sub-model. ij GT is the jth data attention in the contaminated feature data and / or feature fusion data in the input intelligent sub-model, l is the feature attention degree pre-acquired relative to the lth data feature, n is the number of intelligent sub-models, is the sub-model weight of the i-th intelligent sub-model, A i is the model fitness of the ith intelligent sub-model, R iis the model reliability of the i-th intelligent sub-model. Here, the feature attention of all Ken's data features is predetermined, so for each contaminated feature data or feature fusion data, the sum of the feature attention of all data features it contains is positively correlated to its data attention, and the value range of data attention is 0 to 1.
[0104] In the method of this step, the pollution risk data may specifically be equal to the weighted sum or weighted average of the pollution risk sub-data of all intelligent sub-models based on the sub-model weights.
[0105] In addition, in order to address the impact of data distribution on pollution risk, in the present method, the model parameters of the intelligent sub-model in the integrated learning model are associated with pre-acquired data distribution characteristics, and the data distribution characteristics reflect the spatial and / or temporal distribution of the specified pollution feature data.
[0106] Specifically, data distribution characteristics include spatial distribution characteristics or temporal distribution characteristics, such as the degree of spatial aggregation of possible pollution sources, the degree of aggregation of historical pollution events on the time axis, etc.
[0107] Regarding the degree of spatial aggregation of pollution sources, for example, using the nearest distance method, assuming that there are n pollution sources in the industrial agglomeration area, the coordinates of the i-th pollution source in the two-dimensional plane space are (x i ,y i ), the coordinates of the jth pollution source in the two-dimensional plane space are (x j ,y j ), calculate the distance d from each pollution source to its nearest neighbor pollution source i . It can be determined by calculating the distance between two points. For pollution sources i and j, the distance Then find the minimum d corresponding to each i ij As d i Calculate the average nearest neighbor distance of all pollution sources At the same time, assuming that these pollution sources are randomly distributed in the study area, the theoretical average nearest neighbor distance can be calculated based on the area A of the study area and the number of pollution sources n. Finally, the calculation formula of the quantitative index of spatial aggregation degree of pollution sources (NNI) is: When NNI<1, it means that the pollution sources are clustered in space; when NNI=1, it is randomly distributed; when NNI>1, it is discretely distributed.
[0108] Regarding the degree of temporal clustering of pollution events, for example, using the autocorrelation function method, assuming that there are m historical pollution events and their occurrence time is t jFirst, define the time lag τ, which can range from 1 to n-1 (n is the length of the time series, assuming that the pollution event time series is regarded as a discrete time series). j}, calculate the autocorrelation function in When τ = 1, ACF(1) reflects the correlation between adjacent time points. ACF(1) can be used as a simple quantitative indicator to measure the degree of clustering of historical pollution events. The closer ACF(1) is to 1, the higher the degree of temporal clustering; the closer it is to 0, the lower the degree of clustering.
[0109] The aforementioned data distribution characteristics are represented as a one-dimensional or multi-dimensional vector. The quantized value of each dimension of the data distribution characteristics can be used to affect the structural parameters of the intelligent sub-model, or the quantized values of different dimensions can be used to affect the structural parameters of different parts of different intelligent sub-models. For example, the data distribution characteristics can be used to affect the number of neurons and layers of the neural network model, the number of trees and leaf nodes of the LightGBM model, the number of trees and leaf node weights of the XGBoost model, and the learning rate and maximum depth of the GBDT model. In this way, the intelligent sub-model can be better adapted to the data distribution characteristics.
[0110] In summary, this method can intelligently analyze the pollution risks of industrial agglomeration areas based on multi-source pollution characteristic data within a certain period of time. The rich data sources of the analysis results make the analysis results more accurate, and the use of a variety of intelligent model structures and configurations to deal with various data situations make the final analysis results more accurate.
[0111] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.
[0112] In a second aspect, the present application discloses an industrial cluster pollution risk management system based on multi-source data. The system may include the server described in the first aspect of the present application or apply the method disclosed in the first aspect.
[0113] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0114] In summary, this application at least has the following beneficial effects:
[0115] A pollution risk management method and system for industrial agglomeration areas based on multi-source data are provided, which can intelligently and purposefully analyze the multi-source data of industrial agglomeration areas based on an integrated learning model, thereby determining the pollution risks within the industrial agglomeration areas.
[0116] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the aforementioned disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.
Claims
1. A pollution risk management method for industrial agglomeration areas based on multi-source data, characterized in that: include: Obtain multiple pollution characteristic data from multiple sources in industrial agglomeration areas; Constructing a plurality of feature fusion data based on the pollution feature data; Analyze the matching degree data of each intelligent sub-model in the pre-trained integrated learning model for each of the pollution feature data and feature fusion data; Determine the input strategy of the pollution feature data and feature fusion data to the intelligent sub-model of the integrated learning model based on the matching degree data; Based on the input strategy, the pollution feature data and feature fusion data are input into the integrated learning model to obtain the pollution risk data of the industrial agglomeration area.
2. The method according to claim 1, characterized in that The pollution characteristic data includes one or more of environmental monitoring data, hydrogeological data, industrial activity data, pollution event records and remote sensing image data; The feature fusion data include comprehensive environmental indicators, pollutant migration matrix, industrial impact indicators and pollution trend indicators; The comprehensive environmental index is a quantitative value that reflects the comprehensive environmental situation and is calculated and determined by environmental monitoring data from multiple sources; The pollutant migration matrix is a matrix that reflects the spatial migration law of pollutants and is determined by analyzing hydrogeological data and environmental monitoring data; The industrial impact index is a quantitative value that reflects the degree of impact of industrial activities on the environment, and is determined by analyzing industrial activity data and environmental monitoring data; The pollution trend index is a quantitative value that reflects the changing trend of each pollutant and is determined by analyzing environmental monitoring data.
3. The method according to claim 1, characterized in that The analysis of the matching degree data of each of the pollution feature data and feature fusion data to each intelligent sub-model in the pre-trained integrated learning model includes: Obtain the data quality and data characteristics of pollution feature data and feature fusion data; Analyze the data importance and data relevance of pollution feature data and feature fusion data to each intelligent sub-model in the integrated learning model based on data features; Determine the basic matching degree of the pollution feature data and the feature fusion data for each intelligent sub-model in the integrated learning model based on the data quality, data importance and data relevance; Combine the basic matching degree and the pre-acquired data association coefficient to determine the comprehensive matching degree of the pollution feature data and the feature fusion data for each intelligent sub-model in the integrated learning model; The matching degree data is equal to the maximum value of the basic matching degree and the comprehensive matching degree.
4. The method according to claim 1, characterized in that: The input strategy is based on which the pollution feature data and feature fusion data are input into the integrated learning model to obtain the pollution risk data of the industrial agglomeration area, including: Determine the output of each intelligent sub-model under the action of the input as pollution risk sub-data; Determine the model fitness of the intelligent sub-model based on the pollution feature data and / or feature fusion data actually input into the intelligent sub-model and the matching degree data of the intelligent sub-model; Determine the sub-model weight of each intelligent sub-model according to the model fitness and the pre-acquired model reliability of each intelligent sub-model and the data characteristics of the pollution feature data and feature fusion data input into the intelligent sub-model; The pollution risk data is determined by combining the sub-model weights of all intelligent sub-models and the pollution risk sub-data.
5. The method according to claim 1, characterized in that The model parameters of the intelligent sub-model in the integrated learning model are associated with the pre-acquired data distribution characteristics, and the data distribution characteristics reflect the distribution of the specified pollution characteristic data in space and / or time.
6. The method according to claim 3, characterized in that The basic matching degree of the pollution feature data and the feature fusion data for each intelligent sub-model in the integrated learning model based on the data quality, data importance and data relevance includes: BM=w1Q(w2Imp+w3r) Where BM is the basic matching degree for the intelligent sub-model, Q is the data quality, Imp is the data importance for the intelligent sub-model, r is the data relevance for the intelligent sub-model, and w1, w2, and w3 are the preset matching calculation weights for data quality, data importance, and data relevance, respectively.
7. The method according to claim 3, characterized in that The method of combining the basic matching degree and the pre-acquired data association coefficient to determine the comprehensive matching degree of the pollution feature data and the feature fusion data for each intelligent sub-model in the integrated learning model includes: CM ij =αBM ij (1-e -x ) In the formula, CM ij is the comprehensive matching degree of the i-th pollution feature data or feature fusion data to the j-th intelligent sub-model, BM ij is the basic matching degree of the i-th pollution feature data or feature fusion data to the j-th intelligent sub-model, BM kj is the basic matching degree of the kth pollution feature data or feature fusion data to the jth intelligent sub-model, r ik is the data correlation coefficient between the i-th pollution feature data or feature fusion data and the k-th pollution feature data or feature fusion data, i≠k, α and β are respectively the preset first matching calculation coefficient and the second matching calculation coefficient.
8. The method according to claim 4, characterized in that The method of determining the model adaptability of the intelligent sub-model based on the pollution feature data and / or feature fusion data actually input into the intelligent sub-model and the matching degree data of the intelligent sub-model includes: In the formula, A i is the model fitness of the intelligent sub-model, n i is the total number of pollution feature data and / or feature fusion data input into the i-th intelligent sub-model, MD ij It is the jth matching degree data in the pollution feature data and / or feature fusion data input into the intelligent sub-model.
9. The method according to claim 4, characterized in that The step of determining the sub-model weight of each intelligent sub-model according to the model fitness, the pre-acquired model reliability of each intelligent sub-model, and the data characteristics of the pollution feature data and feature fusion data input into the intelligent sub-model includes: Where n i is the total number of pollution feature data and / or feature fusion data input into the i-th intelligent sub-model, g ij is the jth data attention in the pollution feature data and / or feature fusion data input into the intelligent sub-model, n is the number of intelligent sub-models, is the sub-model weight of the i-th intelligent sub-model, A i is the model fitness of the ith intelligent sub-model, R i is the model reliability of the i-th intelligent sub-model.
10. A pollution risk management system for industrial clusters based on multi-source data, characterized in that: The system applies the method according to any one of claims 1-9.
Citation Information
Patent Citations
A communication loss risk prediction method and system of user equipment based on fusion processing and computer equipment
CN113592160A
Security assessment method based on feature matching degree and heterogeneous sub-model fusion
CN117436031A
Polluted site multi-source heterogeneous data fusion method
CN117593614A
Soil arsenic pollution risk assessment method, device and equipment based on optimization model
CN118982226A
Business risk prediction system and method combined with multi-source data
CN119047843A