Industrial agglomeration area pollution risk management method and system based on multi-source data

By integrating multi-source data through an ensemble learning model, the real-time and comprehensive issues of pollution risk assessment in industrial clusters have been resolved, enabling intelligent monitoring and management of pollution risks and improving the accuracy of the assessment.

CN119990769BActive Publication Date: 2026-08-25BEIJING MUNICIPAL RES INST OF ENVIRONMENT PROTECTION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510098987.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2026-08-25
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Existing pollution risk assessment methods for industrial clusters lack real-time and comprehensiveness, making it difficult to effectively respond to rapidly changing pollution situations, and the inability to fully integrate multi-source data leads to inaccurate judgments.

Method used

An ensemble learning model based on multi-source data is adopted. By acquiring and processing environmental monitoring data, hydrogeological data, industrial activity data, pollution event records and remote sensing image data, feature fusion data is constructed. Convolutional neural networks and other intelligent sub-models are used for analysis to determine pollution risks.

Benefits of technology

It enables intelligent, targeted, and real-time monitoring and management of pollution risks in industrial clusters, improving the accuracy and comprehensiveness of pollution risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990769B_ABST
    Figure CN119990769B_ABST
Patent Text Reader

Abstract

The application provides an industrial agglomeration area pollution risk management method and system based on multi-source data, belongs to the field of pollution risk management, and is used for solving the problem that the pollution risk of the industrial agglomeration area is difficult to accurately analyze in the related art.In the method and system, the pollution characteristic data and the characteristic fusion data can be respectively analyzed in a targeted manner based on an integrated learning model, so that the pollution risk in the industrial agglomeration area can be intelligently and reasonably determined, and then the pollution risk in the industrial agglomeration area can be monitored and managed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of pollution risk management, and in particular to a method and system for pollution risk management in industrial clusters based on multi-source data. Background Technology

[0002] With the acceleration of industrialization, industrial clusters have become a significant driver of economic development, but they have also brought serious environmental pollution problems. The concentrated development of industries such as chemicals, metallurgy, and machinery within these clusters has led to varying degrees of pollution of the soil, groundwater, and air. These pollutants not only affect the health of local residents but may also migrate and spread through soil and water, posing a potential threat to a wider ecological environment. Therefore, timely and accurate identification and assessment of high-risk pollution areas are crucial for effective pollution control and management.

[0003] Currently, the assessment of pollution risks in industrial clusters mainly relies on traditional monitoring and analysis methods. These methods often lack real-time and comprehensiveness, making it difficult to effectively address rapidly changing pollution situations. Furthermore, existing monitoring systems typically cannot fully integrate multi-source data, leading to inaccurate assessments of pollution conditions. Therefore, there is an urgent need for an intelligent auxiliary identification system to monitor and manage pollution risks within industrial clusters. Summary of the Invention

[0004] This application provides a method and system for pollution risk management in industrial clusters based on multi-source data. It can analyze the pollution situation in industrial clusters based on multi-source data, which is beneficial for monitoring and managing pollution risks in industrial clusters.

[0005] Firstly, this application provides a method for pollution risk management in industrial clusters based on multi-source data. The method includes:

[0006] Acquire multi-source pollution characteristic data of industrial clusters;

[0007] Multiple feature fusion data are constructed based on the pollution feature data;

[0008] The matching degree data of each pollution feature data and feature fusion data is analyzed for each intelligent sub-model in the pre-trained ensemble learning model.

[0009] Based on the matching degree data, the input strategy for the intelligent sub-model of the ensemble learning model is determined using the pollution feature data and feature fusion data.

[0010] Based on the input strategy, the pollution feature data and feature fusion data are input into the ensemble learning model to obtain pollution risk data for industrial clusters.

[0011] By adopting the above technical solutions, pollution characteristic data and feature fusion data can be analyzed separately and in a targeted manner based on the integrated learning model, thereby intelligently and reasonably determining the pollution risk status in industrial clusters, which in turn helps to monitor and manage pollution risks in industrial clusters.

[0012] Furthermore, the pollution characteristic data includes one or more of the following: environmental monitoring data, hydrogeological data, industrial activity data, pollution event records, and remote sensing image data;

[0013] The feature fusion data includes comprehensive environmental indicators, pollutant migration matrix, industrial impact indicators, and pollution trend indicators;

[0014] The comprehensive environmental index is a quantitative value reflecting the overall environmental situation, which is calculated and determined from multi-source environmental monitoring data;

[0015] The pollutant migration matrix is ​​a matrix that reflects the spatial migration pattern of pollutants and is determined by analysis of hydrogeological data and environmental monitoring data.

[0016] The industrial impact index is a quantitative value reflecting the degree of environmental impact of industrial activities, which is determined by analyzing industrial activity data and environmental monitoring data.

[0017] The pollution trend index is a quantitative value reflecting the changing trend of each pollutant, determined by environmental monitoring data analysis.

[0018] Furthermore, the analysis of the matching degree data of each pollution feature data and feature fusion data for each intelligent sub-model in the pre-trained ensemble learning model includes:

[0019] Data quality and data characteristics of pollution feature data and feature fusion data;

[0020] Based on data feature analysis, the importance and relevance of pollution feature data and feature fusion data to each intelligent sub-model within the ensemble learning model are determined.

[0021] Based on the data quality, data importance, and data relevance, the basic matching degree of the pollution feature data and feature fusion data is determined for each intelligent sub-model within the ensemble learning model;

[0022] The overall matching degree of pollution feature data and feature fusion data for each intelligent sub-model within the ensemble learning model is determined by combining the basic matching degree and the pre-acquired data correlation coefficient.

[0023] The matching score data is equal to the maximum value between the basic matching score and the comprehensive matching score.

[0024] Furthermore, the step of inputting the pollution feature data and feature fusion data into the ensemble learning model based on the input strategy to obtain pollution risk data for the industrial cluster includes:

[0025] The output of each intelligent sub-model under the influence of input is determined as pollution risk sub-data.

[0026] Based on the pollution feature data and / or feature fusion data of the actual input intelligent sub-model, the matching degree data of the intelligent sub-model is used to determine the model fit of the intelligent sub-model.

[0027] The sub-model weights of each intelligent sub-model are determined based on the model fit, the model reliability of each pre-acquired intelligent sub-model, and the data characteristics of the pollution feature data and feature fusion data input to the intelligent sub-model.

[0028] The pollution risk data is determined by combining the sub-model weights of all intelligent sub-models and the pollution risk sub-data.

[0029] Furthermore, the model parameters of the intelligent sub-model in the ensemble learning model are associated with the pre-acquired data distribution features, which reflect the spatial and / or temporal distribution of the specified pollution feature data.

[0030] Furthermore, the determination of the basic matching degree of contaminated feature data and feature fusion data for each intelligent sub-model within the ensemble learning model based on the data quality, data importance, and data relevance includes:

[0031] BM = w1Q(w2Imp + w3r)

[0032] In the formula, BM represents the basic matching degree for the intelligent sub-model, Q represents the data quality, Imp represents the data importance for the intelligent sub-model, r represents the data relevance for the intelligent sub-model, and w1, w2, and w3 are the preset matching calculation weights for data quality, data importance, and data relevance, respectively.

[0033] Furthermore, the determination of the comprehensive matching degree of pollution feature data and feature fusion data for each intelligent sub-model within the ensemble learning model, combining the basic matching degree and the pre-acquired data correlation coefficient, includes:

[0034] CM ij =αBM ij (1-e -x )

[0035]

[0036] In the formula, CM ij BM represents the comprehensive matching degree of the i-th pollution feature data or feature fusion data towards the j-th intelligent sub-model.ij BM represents the basic matching degree of the i-th pollution feature data or feature fusion data towards the j-th intelligent sub-model. kj r represents the basic matching degree of the k-th pollution feature data or feature fusion data towards the j-th intelligent sub-model. ik Let be the data correlation coefficient between the i-th pollution feature data or feature fusion data and the k-th pollution feature data or feature fusion data, where i ≠ k, and α and β are the preset first matching calculation coefficient and second matching calculation coefficient, respectively.

[0037] Furthermore, the determination of the model fit of the intelligent sub-model based on the matching degree data of pollution feature data and / or feature fusion data of the actual input intelligent sub-model includes:

[0038]

[0039] In the formula, A i For the model fit of the intelligent sub-model, n i To input the total number of pollution feature data and / or feature fusion data within the i-th intelligent sub-model, MD ij The j-th matching degree data is taken from the pollution feature data and / or feature fusion data within the input intelligent sub-model.

[0040] Furthermore, determining the sub-model weights of each intelligent sub-model based on model fit, the model reliability of each pre-acquired intelligent sub-model, and the data features of the pollution feature data and feature fusion data of the input intelligent sub-model includes:

[0041]

[0042] γ1+γ2+γ3=1

[0043] In the formula, n i To input the total number of pollution feature data and / or feature fusion data within the i-th intelligent sub-model, GD ij To input the j-th data attention from the pollution feature data and / or feature fusion data within the intelligent sub-model, GT l The feature attention value is pre-acquired relative to the l-th data feature, and n is the number of intelligent sub-models. Let A be the sub-model weight of the i-th intelligent sub-model. i R represents the model fit of the i-th intelligent sub-model. i Let represent the model reliability of the i-th intelligent sub-model.

[0044] Secondly, this application provides a pollution risk management system for industrial clusters based on multi-source data. The system is used to execute any of the methods described in the first aspect above.

[0045] In summary, this application has at least the following beneficial effects:

[0046] A method and system for pollution risk management in industrial clusters based on multi-source data are provided. This method can intelligently and purposefully analyze multi-source data of industrial clusters based on an integrated learning model, thereby determining the pollution risks within the industrial clusters.

[0047] It should be understood that the description in the Summary Section is not intended to limit the key or essential features of the embodiments of this application, nor is it intended to restrict the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0048] The above and other features, advantages, and aspects of the embodiments of this application will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0049] Figure 1 A flowchart of a pollution risk management method for industrial clusters based on multi-source data is shown in an embodiment of this application. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0051] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0052] This application provides a method and system for pollution risk management in industrial clusters based on multi-source data. It can use an integrated learning model to process pollution feature data and feature fusion data in a targeted manner, thereby conducting a comprehensive, accurate, intelligent and real-time analysis of pollution risks in industrial clusters, so as to facilitate the management of pollution risks in industrial clusters.

[0053] Firstly, embodiments of this application disclose a method for pollution risk management in industrial clusters based on multi-source data. This method can be applied to servers.

[0054] Figure 1 A flowchart of a pollution risk management method for industrial clusters based on multi-source data is shown in an embodiment of this application.

[0055] Reference Figure 1 The method specifically includes the following steps:

[0056] S110: Obtain multi-source pollution characteristic data of industrial clusters.

[0057] Pollution characteristic data are multi-source data collected through multiple methods that reflect the pollution situation within industrial clusters. They can reflect the specific conditions of industrial clusters, hydrogeological conditions, and the presence and level of nearby pollution receptors such as water source protection areas. In one example, pollution characteristic data includes environmental monitoring data, hydrogeological data, industrial activity data, pollution event records, and remote sensing image data.

[0058] Specifically, environmental monitoring data is used to reflect the real-time pollution situation in industrial clusters, including timestamped data on soil pollutant concentrations, groundwater pollutant concentrations, air pollution data, and surface water quality monitoring data. Hydrogeological data is used to analyze the migration and diffusion behavior of pollutants in the site, including spatial geological structure and soil characteristics, groundwater level, hydraulic gradient, permeability coefficient, effective porosity, and vertical and horizontal dispersion coefficients, and is timestamped to reflect the specific hydrogeological data at different times. Industrial activity data is used to reflect potential pollution sources in industrial activities, including timestamped data on raw material and waste management, underground storage tank and pipeline layout, and industrial wastewater discharge. Pollution incident records are used to reflect past leakage accidents, pollution incidents, and enterprise penalty records in industrial clusters, including incident details and time records, incident handling plans, incident handling results, and pollution residue data. Remote sensing imagery data is used to reflect important facilities and their evolution within industrial clusters, and is timestamped.

[0059] It should be understood that in the process of analyzing surface water and groundwater pollution, the impact of air pollution data is generally small or even negligible. In actual analysis tasks, air pollution-related content may not be analyzed, or air pollution-related parameters may be assigned small coefficients. The contents before and after in the embodiments of this application can be based on this consideration.

[0060] The aforementioned pollution feature data needs to undergo data cleaning, noise reduction, and data standardization to ensure that all pollution feature data have the same dimensions and range. Of course, some pollution feature data may be expressed as single-dimensional quantitative values, while some pollution data may be expressed as multi-dimensional or even high-dimensional vector values.

[0061] S120: Construct multiple feature fusion data based on the pollution feature data.

[0062] Feature fusion data reflects the result of multi-source fusion of pollution feature data or fusion of different pollution feature data. In one example, feature fusion data may include comprehensive environmental indicators, pollutant migration matrix, industrial impact data, and pollution trend indicators.

[0063] The comprehensive environmental index is a quantitative value reflecting the overall environmental situation, calculated and determined from multi-source environmental monitoring data. Specifically, the fusion of multi-source environmental monitoring data can be achieved based on convolutional neural networks. For example, let the soil pollutant concentration sequence be C. soil (t)=[C soil (1,t),C soil (2,t),…,C soil (n soil ,t)](where n soil (Number of soil pollutant species), groundwater pollutant concentration sequence is C groundwater (t), the air pollution data sequence is A air (t), the surface water quality monitoring data sequence is C surfacewater (t). These data are constructed into a 4-dimensional spatiotemporal data cube X based on time series and spatial location (if spatial distribution information is available). env (t), where the dimensions are pollutant type, spatial location (optional), time, and environmental medium type (soil, groundwater, air, surface water), respectively. The CNN model structure includes convolutional layers, pooling layers, and fully connected layers. Let the kernel size of the convolutional layer be k, the stride be s, and the padding be p. After passing through the convolutional layer and the pooling layer, the feature map F is obtained. env The fully connected layer transforms the feature maps into a fused environmental comprehensive index (I). env (t). The training process of CNN uses the backpropagation algorithm to minimize the loss function L. env For example, the mean squared error loss function:

[0064]

[0065] Where N is the number of training samples. This represents the true comprehensive environmental indicators.

[0066] The pollutant migration matrix is ​​a matrix that reflects the spatial migration patterns of pollutants and is determined by analysis of hydrogeological data and environmental monitoring data.

[0067] Specifically, first, define variables and matrices. Suppose there are n types of pollutants, m monitoring locations, and T time points. Let C... ijtLet M represent the concentration of pollutant i at position j and time point t (i = 1, 2, ..., n; j = 1, 2, ..., m; t = 1, 2, ..., T). The pollutant migration matrix M is a three-dimensional matrix with dimensions n × m × T, where M... ijt =C ijt .

[0068] Secondly, considering the influence of geological and hydrological factors on migration, including the influence of groundwater flow velocity and the influence of precipitation and surface runoff, regarding the influence of groundwater flow velocity, we assume the groundwater flow velocity is v. j (Unit: m / s) At the j-th location, for the migration of groundwater pollutants (assumed to be the k-th pollutant), the concentration change within adjacent time intervals Δt can be approximated by the simplified form of the convection-dispersion equation (assuming one-dimensional flow and neglecting dispersion): C k,j,t+1 =C k,j-vj Δt,t, where jv j Δt represents the upstream location calculated based on flow velocity and time interval. Regarding the impact of precipitation and surface runoff, for surface water quality and soil pollutants, precipitation P (unit: mm) and the surface runoff coefficient α (dimensionless) affect the migration of pollutants. Assume a unit area of ​​A (unit: m²). 2 ), Surface runoff flow Q = P × α × A. For pollutants in the l-th type of surface water, the concentration change due to runoff between adjacent locations j and j+1 can be calculated using the law of conservation of mass: Where V j Let be the volume of water at location j. For soil pollutants (assumed to be type s), precipitation leads to leaching, causing the pollutants to migrate to the lower soil layers. Assuming soil porosity is θ (dimensionless) and soil layer thickness is h (unit: m), the change in pollutant concentration over the time interval Δt can be expressed as:

[0069] Secondly, considering the influence of geological structure and soil properties, including the soil adsorption coefficient K... d (Unit: L / kg) Reflects the soil's adsorption capacity for pollutants. For the migration of pollutant p in the soil, adsorption will reduce its concentration in the soil pore water. Let the soil bulk density be ρ. b (Unit: kg / m³) 3 Soil pore water content is θ w (Dimensionless), based on the linear adsorption model, the pollutant concentration C in soil pore water... w and the concentration of pollutants C adsorbed on soil particles s The relationship between them is: C s =K d ×C w Total pollutant concentration Ctotal =C w ×θ w +C s ×ρ b When considering adsorption, the migration equation of pollutants in soil needs to be modified to account for concentration changes, taking into account the above relationships.

[0070] Secondly, considering the exchange of air pollution data with pollutants, for gaseous pollutants (let's assume it's type q), exchange processes exist at the soil-atmosphere interface and the water-atmosphere interface. Assume the concentration of pollutants in the atmosphere is C. a (Unit: mg / m³) 3 The gas exchange coefficient is k (unit: m / s). At an interface with a unit area A, according to the two-film theory, the flux of pollutants at the interface is F (unit: mg / (m²)). 2 ·s)) can be expressed as: F=k×(C a -C interface ), where C interface This refers to the pollutant concentration at the interface, and its value is related to the pollutant concentration in the soil or water body. For example, for the soil-atmosphere interface, the change in pollutant concentration in soil pore water can be expressed as:

[0071] Finally, regarding the update of the pollutant migration matrix, based on the effects of the various factors mentioned above on pollutant concentration, the elements M in the pollutant migration matrix M are updated after each time step Δt. ijt This reflects the spatial and temporal migration of pollutants.

[0072] The industrial impact index is a quantitative value reflecting the degree of environmental impact of industrial activities, determined by analysis of industrial activity data and environmental monitoring data.

[0073] Specifically, the first step is to determine the changes in pollutant concentrations in soil, groundwater, air, and surface water. For example, regarding soil, suppose there are n types of soil pollutants, and within the time interval [t1, t2], the change in concentration of the i-th pollutant... The calculation formula is: in This represents the concentration of the first soil pollutant at time t; regarding groundwater, for m types of groundwater pollutants, the change in concentration ΔC of the j-th pollutant within the same time interval. gwj The calculation formula is: in This represents the concentration of the j-th groundwater pollutant at time t; regarding air, assuming there are k air pollutants, the change in concentration ΔC of the first pollutant is... air, The calculation formula is: This represents the concentration of the l-th air pollutant at time t; regarding surface water, for p-type pollutants in surface water, it represents the change in the concentration of the q-th pollutant within the time interval [t1, t2]. The calculation formula is: in It is the concentration of the qth type of pollutant in the surface water at time t.

[0074] Secondly, determine the contribution weights of potential pollution sources from industrial activities. For example, for raw material management, assuming there are r types of raw materials, the weight w of the pollutants that the u-th raw material may generate. rawu The weight of a raw material can be determined based on its chemical properties and dosage. For example, if a certain raw material generates a large amount of heavy metal pollutants during production, its weight will be higher. For waste management, assuming there are *s* types of waste, the weight *w* of the pollutants that the *v*th type of waste might generate... wastev The weighting can be determined based on factors such as toxicity and treatment methods. For example, hazardous waste has a higher weighting than general industrial waste. The total weighting related to raw material and waste management data. Suppose that the layout of underground storage tanks and pipelines involves t risk points (such as tank leakage risk points, pipeline interfaces, etc.), and the weight of the w-th risk point is... The weighting can be determined based on factors such as leakage probability and potential leakage amount. For example, an aging underground storage tank will have a higher leakage weighting than a new one. The total weighting related to underground storage tank and pipeline layout data... Assume that industrial wastewater discharge involves x types of pollutants, and the weight of pollutant y in the wastewater is... The weight can be determined based on factors such as emission volume and toxicity. For example, wastewater containing high concentrations of heavy metals has a higher emission weight. The total weight related to industrial wastewater emission data...

[0075] The pollution trend index is a quantitative value reflecting the changing trend of each pollutant, determined by the analysis of environmental monitoring data. For example, based on environmental monitoring data, the slope of linear regression is used to determine the rate of change of each pollutant in various environmental media such as soil, groundwater, air, and surface water, and the comprehensive rate of change of the concentration of each pollutant is calculated based on the preset weights of different environmental media.

[0076] Feature fusion data can, of course, include other content and employ other specific fusion methods, which will not be listed here. Similar to pollution feature data, feature fusion data consists of single-dimensional quantitative values ​​or multi-dimensional vector values, requiring only that it reflects the indicator data to be represented. In short, feature fusion data is determined based on pollution feature data, and it is more capable of reflecting certain obvious and integrated pollution situations than pollution feature data alone, which is beneficial for subsequent analysis of pollution conditions in industrial clusters.

[0077] S130: Analyze the matching degree data of each intelligent sub-model in the pre-trained ensemble learning model based on the pollution feature data and feature fusion data.

[0078] The specific methods in this step include: acquiring the data quality and data features of pollution feature data and feature fusion data; analyzing the data importance and data relevance of pollution feature data and feature fusion data to each intelligent sub-model within the ensemble learning model based on the data features; determining the basic matching degree of pollution feature data and feature fusion data to each intelligent sub-model within the ensemble learning model based on the data quality, data importance, and data relevance; determining the comprehensive matching degree of pollution feature data and feature fusion data to each intelligent sub-model within the ensemble learning model by combining the basic matching degree and the pre-acquired data correlation coefficient; the matching degree data is equal to the maximum value of the basic matching degree and the comprehensive matching degree.

[0079] In this step of the method, data quality can be determined based on pre-built data quality assessment indicators such as completeness, accuracy, and consistency. Specifically, regarding completeness, let the total number of samples be N, and the number of valid samples in a certain contamination feature data (or feature fusion data) be n, then the completeness... Its value ranges from 0 to 1, with values ​​closer to 1 indicating better data integrity. Regarding accuracy: for data with known standard reference values ​​(e.g., some calibrated environmental monitoring data with standard value comparisons), let the number of comparison samples be m, and the number of accurate samples be k. The accuracy... The value range is also from 0 to 1, representing the degree of agreement between the data and the true value. Regarding consistency: it is assessed by checking for inconsistencies in the same data recorded from different sources (if multiple sources exist) or at different times. For example, for the same pollutant concentration data collected at the same location at different times, if the number of times the difference exceeds the reasonable fluctuation range (which can be determined through statistical analysis of historical reasonable fluctuation ranges) is d, and the total number of comparisons is q, then the consistency is considered... Its value ranges from 0 to 1, with higher values ​​indicating better data consistency. Regarding timeliness: it is measured by the interval t (which can be days, hours, etc., depending on the data characteristics) between the data collection time and the current analysis time, and the effective update cycle T of the data. That is, if the data is within the valid update cycle, the timeliness is 1; if it exceeds the update cycle, it is updated proportionally.

[0080] Data features can include numerical features and categorical features. Numerical features include statistical features such as mean, standard deviation, skewness, and kurtosis, as well as distribution features presented by histograms, box plots, etc.

[0081] Regarding data importance, the feature importance of each data feature can be preset relative to each intelligent sub-model, and then the corresponding data importance can be determined comprehensively based on the feature importance of the data features carried by the pollution feature data or feature fusion data.

[0082] Regarding data correlation, the frequency of each data feature can be determined relative to the input record big data of each intelligent sub-model. The higher the frequency of a data feature, the higher the correlation between the data feature and the intelligent sub-model. Accordingly, the data correlation between pollution feature data or feature fusion data and the intelligent sub-model can be determined based on the correlation between all data features and the intelligent sub-model.

[0083] The basic matching degree for determining contaminated feature data and feature fusion data for each intelligent sub-model within the ensemble learning model based on the data quality, data importance, and data relevance includes:

[0084] BM = w1Q(w2Imp + w3r)

[0085] In the formula, BM represents the basic matching degree for the intelligent sub-model, Q represents the data quality, Imp represents the data importance for the intelligent sub-model, r represents the data relevance for the intelligent sub-model, and w1, w2, and w3 are the preset matching calculation weights for data quality, data importance, and data relevance, respectively.

[0086] The determination of the comprehensive matching degree of pollution feature data and feature fusion data for each intelligent sub-model within the ensemble learning model, combining the basic matching degree and the pre-acquired data correlation coefficient, includes:

[0087] CM ij =αBM ij (1-e -x )

[0088]

[0089] In the formula, CM ij BM represents the comprehensive matching degree of the i-th pollution feature data or feature fusion data towards the j-th intelligent sub-model. ij BM represents the basic matching degree of the i-th pollution feature data or feature fusion data towards the j-th intelligent sub-model. kj r represents the basic matching degree of the k-th pollution feature data or feature fusion data towards the j-th intelligent sub-model. ik Let be the data correlation coefficient between the i-th pollution feature data or feature fusion data and the k-th pollution feature data or feature fusion data, where i ≠ k, and α and β are the preset first matching calculation coefficient and second matching calculation coefficient, respectively.

[0090] S140: Based on the matching degree data, determine the input strategy of the intelligent sub-model of the ensemble learning model for the pollution feature data and feature fusion data.

[0091] The ensemble learning model pre-construction and training in this step can specifically include intelligent sub-models such as neural network models, LightGBM models, XGBoost models, and GBDT models.

[0092] A specific example of this method is that a matching degree threshold is preset for each intelligent sub-model, and pollution feature data and feature fusion data with matching degree data higher than the matching degree threshold are used as inputs to the intelligent sub-model, thereby determining how all pollution feature data and feature fusion data are input into the intelligent sub-model respectively.

[0093] S150: Based on the input strategy, the pollution feature data and feature fusion data are input into the ensemble learning model to obtain pollution risk data of the industrial cluster.

[0094] The specific steps of this method include: determining that the output of each intelligent sub-model under the influence of input is pollution risk sub-data; determining the model fit of the intelligent sub-model based on the matching degree data of the pollution feature data and / or feature fusion data of the actual input intelligent sub-model to the intelligent sub-model; determining the sub-model weight of each intelligent sub-model according to the model fit, the model reliability of each pre-acquired intelligent sub-model, and the data characteristics of the pollution feature data and feature fusion data of the input intelligent sub-model; and determining the pollution risk data by combining the sub-model weights of all intelligent sub-models and the pollution risk sub-data.

[0095] In this step, under the input strategy, pollution feature data and feature fusion data are respectively input into the intelligent sub-model within the ensemble learning model, and pollution risk sub-data is output under the action of the intelligent sub-model. The pollution risk sub-data is quantitative data reflecting the degree of pollution risk.

[0096] In this step, the model fit of the intelligent sub-model is determined based on all pollution feature data and / or feature fusion data actually input into the intelligent sub-model. Specifically, determining the model fit of the intelligent sub-model based on the matching degree data of the pollution feature data and / or feature fusion data actually input into the intelligent sub-model includes:

[0097]

[0098] In the formula, A i For the model fit of the intelligent sub-model, n i To input the total number of pollution feature data and / or feature fusion data within the i-th intelligent sub-model, MD ijThe j-th matching degree data is taken from the pollution feature data and / or feature fusion data within the input intelligent sub-model.

[0099] In this step, model reliability is pre-acquired, specifically determined based on the accuracy and mean squared error of the intelligent sub-model on historical validation datasets. For example, accuracy: for a classification task, let sub-model s... i The number of correctly predicted samples on the historical validation dataset (containing data samples with known true class labels) is TP. i (TruePositive, the actual sample), the total number of samples is N. i Then the accuracy Regarding Mean Squared Error (MSE): For a regression task, let the submodel s i The predicted value on the historical validation dataset is The actual value is y ij (j = 1, 2, ..., k, where k is the number of samples in the validation dataset), then the mean squared error MSE can generally be used. i The reciprocal of the factor is used to measure model reliability (the larger the reciprocal, the higher the reliability, because the smaller the error), denoted as . (For classification tasks, appropriate metrics such as accuracy can be used to convert them into numerical values ​​that measure reliability.)

[0100] In this step, the determination of the sub-model weights of each intelligent sub-model based on model fitness, the model reliability of each pre-acquired intelligent sub-model, and the data features of the pollution feature data and feature fusion data of the input intelligent sub-model includes:

[0101]

[0102] γ1+γ2+γ3=1

[0103] In the formula, n i To input the total number of pollution feature data and / or feature fusion data within the i-th intelligent sub-model, GD ij To input the j-th data attention from the pollution feature data and / or feature fusion data within the intelligent sub-model, GT l The feature attention value is pre-acquired relative to the l-th data feature, and n is the number of intelligent sub-models. Let A be the sub-model weight of the i-th intelligent sub-model. i R represents the model fit of the i-th intelligent sub-model. iLet represent the model reliability of the i-th intelligent sub-model. Here, the feature attention of all Ken's data features is predetermined. Therefore, for each pollution feature data or feature fusion data, the sum of the feature attention of all its data features is positively correlated with its data attention, and the value of data attention ranges from 0 to 1.

[0104] In this step, the pollution risk data can specifically be equal to the weighted sum or weighted average of the pollution risk sub-data of all intelligent sub-models based on the sub-model weights.

[0105] In addition, to address the impact of data distribution on pollution risk, in this method, the model parameters of the intelligent sub-model in the ensemble learning model are associated with the pre-acquired data distribution features, which reflect the spatial and / or temporal distribution of the specified pollution feature data.

[0106] Specifically, data distribution characteristics include spatial distribution characteristics or temporal distribution characteristics, such as the degree of spatial aggregation of potential pollution sources and the degree of aggregation of historical pollution events over time.

[0107] Regarding the spatial concentration of pollution sources, for example, using the nearest distance method, suppose there are n pollution sources in an industrial cluster, and the coordinates of the i-th pollution source in a two-dimensional plane are (x... i y i The coordinates of the j-th pollution source in the two-dimensional plane space are (x, y). j y j ), calculate the distance d from each pollution source to its nearest neighbor pollution source. i The distance can be determined by calculating the formula for the distance between two points. For pollution sources i and j, the distance is... Then find the minimum d corresponding to each i. ij As d i Calculate the average nearest neighbor distance for all pollution sources. Meanwhile, assuming these pollution sources are randomly distributed within the study area, the theoretical average nearest neighbor distance can be calculated based on the area A of the study area and the number n of pollution sources. Finally, the formula for calculating the Nominal Intensity Index (NNI) of pollution source spatial concentration is as follows: When NNI < 1, it indicates that the pollution sources are spatially clustered; when NNI = 1, they are randomly distributed; when NNI > 1, they are discretely distributed.

[0108] Regarding the temporal clustering of pollution events, for example using the autocorrelation function method, suppose there are m historical pollution events, and their occurrence time is t. jFirst, define the time lag τ, which can range from 1 to n-1 (where n is the length of the time series; assume the pollution event time series is considered a discrete time series). For the time series {t} j}, calculate the autocorrelation function in When τ = 1, ACF(1) reflects the correlation between adjacent time points. ACF(1) can be used as a simple quantitative indicator to measure the degree of clustering of historical pollution events. The closer ACF(1) is to 1, the higher the degree of clustering in time; the closer it is to 0, the lower the degree of clustering.

[0109] The aforementioned data distribution characteristics are represented as a one-dimensional or multi-dimensional vector. The quantized values ​​of each dimension of the data distribution characteristics can influence the structural parameters of the intelligent sub-model, or different dimensions can be used to influence the structural parameters of different parts of different intelligent sub-models. For example, data distribution characteristics can influence the number of neurons and layers in a neural network model, the number of trees and leaf nodes in a LightGBM model, the number of trees and leaf node weights in an XGBoost model, and the learning rate and maximum depth in a GBDT model. This allows the intelligent sub-model to better adapt to the data distribution characteristics.

[0110] In summary, this method can intelligently analyze the pollution risks of industrial clusters based on multi-source pollution characteristic data over a certain period of time. The rich data sources of the analysis results make the analysis results more accurate. Furthermore, by utilizing various intelligent model structures and configuring considerations for different data situations, the final analysis results are more accurate.

[0111] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0112] Secondly, embodiments of this application disclose a pollution risk management system for industrial clusters based on multi-source data. This system may include the server described in the first aspect of the embodiments of this application or the method disclosed in the first aspect.

[0113] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system described herein can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0114] In summary, this application has at least the following beneficial effects:

[0115] A method and system for pollution risk management in industrial clusters based on multi-source data are provided. This method can intelligently and purposefully analyze multi-source data of industrial clusters based on an integrated learning model, thereby determining the pollution risks within the industrial clusters.

[0116] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A pollution risk management method for industrial clusters based on multi-source data, characterized in that, include: Acquire multi-source pollution characteristic data of industrial clusters; Multiple feature fusion data are constructed based on the pollution feature data; The matching degree data of each pollution feature data and feature fusion data is analyzed for each intelligent sub-model in the pre-trained ensemble learning model. Based on the matching degree data, the input strategy for the intelligent sub-model of the ensemble learning model is determined using the pollution feature data and feature fusion data. Based on the input strategy, the pollution feature data and feature fusion data are input into the ensemble learning model to obtain pollution risk data of industrial clusters. The pollution characteristic data includes one or more of the following: environmental monitoring data, hydrogeological data, industrial activity data, pollution event records, and remote sensing image data; The feature fusion data includes comprehensive environmental indicators, pollutant migration matrix, industrial impact indicators, and pollution trend indicators; The comprehensive environmental index is a quantitative value reflecting the overall environmental situation, which is calculated and determined from multi-source environmental monitoring data; The pollutant migration matrix is ​​a matrix that reflects the spatial migration pattern of pollutants and is determined by analysis of hydrogeological data and environmental monitoring data. The industrial impact index is a quantitative value reflecting the degree of environmental impact of industrial activities, which is determined by analyzing industrial activity data and environmental monitoring data. The pollution trend index is a quantitative value reflecting the changing trend of each pollutant, which is determined by the analysis of environmental monitoring data; The analysis of each pollution feature data and feature fusion data for matching the pre-trained ensemble learning model to each intelligent sub-model includes: Data quality and data characteristics of pollution feature data and feature fusion data; Based on data feature analysis, the importance and relevance of pollution feature data and feature fusion data to each intelligent sub-model within the ensemble learning model are determined. Based on the data quality, data importance, and data relevance, the basic matching degree of the pollution feature data and feature fusion data is determined for each intelligent sub-model within the ensemble learning model; The overall matching degree of pollution feature data and feature fusion data for each intelligent sub-model within the ensemble learning model is determined by combining the basic matching degree and the pre-acquired data correlation coefficient. The matching score data is equal to the maximum value between the basic matching score and the comprehensive matching score; The pollution risk data of the industrial cluster area obtained by inputting the pollution feature data and feature fusion data into the ensemble learning model based on the input strategy includes: The output of each intelligent sub-model under the influence of input is determined as pollution risk sub-data. Based on the pollution feature data and / or feature fusion data of the actual input intelligent sub-model, the matching degree data of the intelligent sub-model is used to determine the model fit of the intelligent sub-model. The sub-model weights of each intelligent sub-model are determined based on the model fit, the model reliability of each pre-acquired intelligent sub-model, and the data characteristics of the pollution feature data and feature fusion data input to the intelligent sub-model. The pollution risk data is determined by combining the sub-model weights of all intelligent sub-models and the pollution risk sub-data.

2. The method according to claim 1, characterized in that, The model parameters of the intelligent sub-model in the ensemble learning model are associated with the pre-acquired data distribution characteristics, which reflect the spatial and / or temporal distribution of the specified pollution feature data.

3. The method according to claim 1, characterized in that, The basic matching degree for determining contaminated feature data and feature fusion data for each intelligent sub-model within the ensemble learning model based on the data quality, data importance, and data relevance includes: In the formula, BM represents the basic matching degree for the intelligent sub-model, and Q represents the data quality. Let r represent the data importance for the intelligent sub-model, and r represent the data relevance for the intelligent sub-model. These are the preset matching calculation weights for data quality, data importance, and data relevance, respectively.

4. The method according to claim 1, characterized in that, The determination of the comprehensive matching degree of pollution feature data and feature fusion data for each intelligent sub-model within the ensemble learning model, combining the basic matching degree and the pre-acquired data correlation coefficient, includes: In the formula, The comprehensive matching degree of the i-th pollution feature data or feature fusion data with respect to the j-th intelligent sub-model. The basic matching degree of the i-th pollution feature data or feature fusion data towards the j-th intelligent sub-model. The basic matching degree of the k-th pollution feature data or feature fusion data towards the j-th intelligent sub-model. Let be the data correlation coefficient between the i-th pollution feature data or feature fusion data and the k-th pollution feature data or feature fusion data. , , These are the preset first matching calculation coefficients and second matching calculation coefficients, respectively.

5. The method according to claim 1, characterized in that, The matching degree data between the pollution feature data and / or feature fusion data based on the actual input intelligent sub-model and the intelligent sub-model, used to determine the model fit of the intelligent sub-model, includes: In the formula, Model fit for intelligent sub-models The input is the total number of pollution feature data and / or feature fusion data within the i-th intelligent sub-model. The j-th matching degree data is taken from the pollution feature data and / or feature fusion data within the input intelligent sub-model.

6. The method according to claim 1, characterized in that, The process of determining the sub-model weights of each intelligent sub-model based on model fit, the model reliability of each pre-acquired intelligent sub-model, and the data features of the pollution feature data and feature fusion data input to the intelligent sub-model includes: In the formula, The input is the total number of pollution feature data and / or feature fusion data within the i-th intelligent sub-model. The input is the attention level of the j-th data point in the pollution feature data and / or feature fusion data within the intelligent sub-model, where n is the number of intelligent sub-models. Let be the sub-model weights of the i-th intelligent sub-model. The model fit of the i-th intelligent sub-model. Let represent the model reliability of the i-th intelligent sub-model.

7. A pollution risk management system for industrial clusters based on multi-source data, characterized in that, The system uses the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Soil arsenic pollution risk assessment method, device and equipment based on optimization model

    CN118982226A

  • Landslide hazard monitoring and early warning method and system based on real 3D

    US12130401B1