Method and system for generating a regional new pollutant list based on industry atlas deduction

CN121903381BActive Publication Date: 2026-08-07SICHUAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN UNIV
Filing Date
2026-01-21
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

由于新污染物种类繁多且缺乏标准监测方法,在管控初期面临数据匮乏、难以满足模型输入要求的问题

Benefits of technology

(1)突破数据匮乏瓶颈,实现“未测先知”:针对新污染物难监测、数据稀缺的痛点,本发明改变了现有技术依赖“先监测后筛选”的滞后模式。通过引入“区域暴露指数—本地排放潜力指数—模型化预测浓度”三维互补暴露表征体系,首次实现了基于“区域产业分布图谱”推演“环境暴露”的技术路径,能够在实测数据稀缺的情况下,精准识别出与本地产业结构高度相关的潜在高风险物质。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121903381B_ABST
    Figure CN121903381B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on industry atlas deduction regional new pollutant list generation method and system, it is related to new pollutant environmental risk control and decision support technical field, the present application is with multi-source information fusion as core, constructs the hierarchical candidate library of " authority list-advanced measurement-regional industry distribution atlas ";Innovative introduction is with regional exposure index, local emission potential index (CSW) and model prediction concentration (PEC) as core complementary three-dimensional exposure characterization system, and the data blank problem of exposure evaluation is solved using multi-source data interpolation mechanism;Finally, using the differentiated empowerment strategy of "theoretical risk objective evaluation+exposure risk subjective deduction", coupled generation comprehensive classification list.The present application breaks through the dependence of traditional method to large-scale measurement, realizes from "passive screening" to "active early warning", significantly improves the accuracy and economy of local control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental risk management and decision support technology for new pollutants, specifically to a method and system for generating regional new pollutant inventories based on industry map extrapolation. Background Technology

[0002] With the expansion of chemical applications and industrial activities, a large number of "emerging pollutants" (ECs) that are not yet fully covered by current regulations but may pose risks to the environment or human health are constantly emerging. These substances are diverse, rapidly evolving, and some exhibit characteristics such as persistent degradation, bioaccumulation, and long-term toxic effects even at low concentrations. Due to their complex environmental behavior and the lack of systematic toxicological data, risk identification is highly concealed and delayed. Therefore, countries have strengthened management at the policy level; however, in practical implementation, accurately identifying high-risk targets in specific areas from thousands of potential substances remains a significant challenge.

[0003] In existing risk identification and inventory construction technologies, mainstream methods often rely heavily on large-scale on-site monitoring data or historical detection records. For example, some existing technologies collect measured concentrations of pollutants from enterprise discharge outlets or environmental media, and screen priority pollutants based on detection frequency or exceedance rate. Due to the large number of new pollutants and the lack of standardized monitoring methods, data scarcity and difficulty in meeting model input requirements are encountered in the early stages of control. Furthermore, post-hoc screening models based on actual measurements are difficult to achieve proactive early warning, and existing single fixed-weight methods cannot objectively handle conflicts between multi-dimensional indicators, making it difficult for inventory construction to balance regional industrial characteristics and the objectivity of assessment results.

[0004] Given the current situation, there is an urgent need to break through the reliance on complete measured data and develop a regionalized inventory generation method suitable for data-scarce scenarios. This method needs to address the following key technical challenges: how to effectively connect international / national universal lists with local characteristics to scientifically screen a regionally representative "preliminary database of suspected new pollutants" in the absence of large-scale local survey data; how to establish a non-measured correlation mapping mechanism between "industry characteristics and pollutants" to scientifically extrapolate and quantify "local emission potential" in scenarios with zero monitoring data, solving the problem of unknown emission source strength; and how to construct a universal high-risk new pollutant inventory generation system adapted to local watershed characteristics, thereby truly realizing a shift from a "passive measured screening" to a "proactive risk prediction" management model under limited administrative and monitoring resources. Therefore, the proposed invention has significant engineering application value in supporting local precise monitoring and hierarchical control decision-making. Summary of the Invention

[0005] In view of this, the technical problem to be solved by this invention is to propose a method and system for generating regional new pollutant inventories based on industry map extrapolation. The core of this solution lies in constructing a hierarchical non-measured correlation path of "authoritative inventory - frontier risk measurement - regional industrial distribution map"; based on the assessment of the inherent hazard attributes of substances, it further proposes a complementary three-dimensional exposure characterization system based on "regional exposure index - local emission potential index (CSW) - model-predicted concentration". This system can overcome the limitation of scarce measured data and realize the logical quantification of "industry → emission → exposure" from "regional industrial structure" to "potential environmental exposure" through non-targeted extrapolation.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method and system for generating a regional new pollutant inventory based on industry map extrapolation, comprising the following steps: S1, Construction of a hierarchical candidate pollutant database: Heterogeneous data fusion is carried out by combining authoritative control lists issued by international and national authorities and cutting-edge field measurement research results to construct a candidate set driven by both policy and academia; the candidate set is mapped to the industrial distribution map of the target region, and a preliminary screening database of candidate pollutants with regional representativeness is obtained based on the rules of material-industry correlation and environmental media adaptability. S2, Multidimensional Characterization of Theoretical Hazards: Candidate pollutants are quantitatively characterized according to factors related to biological persistence, bioaccumulation, ecotoxicity and human health risks, forming a standardized theoretical risk index matrix. S3, Local Emission Potential (CSW) Derivation: Construct an industrial distribution map of the target area, establish industry weight vectors, build a correlation scoring matrix between pollutants and industries, and derive the potential emission source strength values ​​of each candidate pollutant through matrix weighting operations; S4, Complementary 3D Exposure Characterization: Construct a 3D complementary exposure characterization system consisting of regional exposure correlation index, local emission potential index (CSW) and model-predicted concentration (PEC); establish a data availability discrimination and missing data filling mechanism, and when data in any dimension is missing or has low confidence, use data from other dimensions to fill or mutually correct, and generate a comprehensive exposure level index. S5, Objective measurement and multidimensional ranking of indicator weights: Construct an evaluation matrix of theoretical risk and exposure risk, use a differentiated weighting model based on subjective and objective integration to calculate the weight of each indicator, and use a multi-attribute decision ranking model to calculate the theoretical risk priority and exposure risk priority of each candidate pollutant. S6, Comprehensive Judgment and Dynamic Versioning Management: According to the preset fusion rules, the theoretical risk priority and the exposure risk priority are combined into a comprehensive priority, and a regionalized list of high-risk new pollutants is generated accordingly.

[0007] As a preferred option, it also includes: Establish an uncertainty propagation analysis mechanism, a traceable list version update mechanism based on trigger conditions, and a feedback-based targeted verification step; Based on the generated comprehensive priority ranking results, high-risk substances and representative locations and representative environmental media of a preset proportion or number are selected for targeted on-site monitoring; the consistency between the on-site monitoring results and the model prediction results is evaluated, and when the deviation exceeds the preset threshold, the weighting weights, model parameters or substance-industry association rules are modified, and the modification log is recorded to ensure traceability.

[0008] Preferably, the candidate pollutant database described in S1 is stored in a three-level hierarchical structure, including: a policy / authoritative list layer as basic compliance input, a measured / research aggregation layer as a supplement to cutting-edge risks, and a regional adaptation layer as localized filtering results; the database configures a unified metadata field for each candidate pollutant, the field including a unique chemical identifier, key physicochemical properties, data source evidence level, version timestamp, and environmental media adaptability identifier.

[0009] Preferably, the establishment of the pollutant-industry correlation scoring matrix in S3 specifically includes: assigning graded values ​​to the correlation strength between pollutants and industries based on authoritative lists, literature evidence, chemical registration usage information, production process inferences, or industry pollution characteristic databases; the matrix weighting operation is specifically configured as follows: the correlation scoring matrix and the industry weight vector are weighted and calculated, and an evidence confidence coefficient is introduced during the calculation process to weight and correct the reliability of data from different sources, and finally the calculation result is normalized to obtain the local emission potential index (CSW); the industry weight vector is determined based on the industry output value ratio, the number of enterprises ratio, the pollution discharge ratio, or the regional industrial policy priority.

[0010] Preferably, the regional exposure correlation index mentioned in S4 is calculated from systematic data meta-analysis, geographically similarity-weighted data, or cross-regional analogy data; the model-based predicted concentration is a relative exposure reference value derived from standardized environmental fate assumption scenario simulation, material physicochemical properties, and target watershed characteristic parameters; the three-in-one mechanism is configured with dynamic weight allocation logic to ensure that the exposure indicators are comparable across regions, across substances, and across environmental media under different circumstances, such as no measured data, only a small amount of measured data, or abundant measured data.

[0011] Preferably, the weighting process described in S5 adopts a differentiated weighting strategy: for theoretical risk indicators, an objective weighting model based on data statistical characteristics (such as the entropy method) is used to determine the weights; for exposure risk indicators, a subjective weighting model based on expert decision matrices (such as the analytic hierarchy process) is used to determine the weights, so as to strengthen the driving role of local industry characteristics in exposure assessment; all intermediate process data, weight parameter settings and ranking results generated in the weighting and ranking process are solidified and saved in the form of metadata and included in the version management module to support subsequent calculation audit, result verification and model calibration.

[0012] Preferably, the uncertainty propagation analysis mechanism described in S6 includes: performing numerical simulation, resampling analysis, or sensitivity analysis on the uncertainty distribution of key input parameters, calculating the confidence interval of the comprehensive priority, the probability of high-risk inclusion, or the robustness coefficient, and using this statistical characteristic as part of the list decision reference; the triggering conditions include the access of new high-confidence measured data, authoritative list updates, monitoring events exceeding thresholds, regional industrial structure adjustments, release of new pollutant risk research results, or changes in policy control requirements. When the triggering conditions are met, the system automatically or semi-automatically triggers the recalculation of the priority of relevant candidate pollutants and generates a new version of the list.

[0013] A regional new pollutant inventory generation system based on industry map extrapolation includes: The multi-source data access and fusion module is used to integrate policy lists, literature data, regional industrial distribution information, environmental media characteristic data, and measured monitoring data; The candidate library hierarchical management module is used to build and maintain a three-level candidate pollutant library and supports multi-environment media adaptability management. The theoretical hazard and exposure assessment module is used to calculate theoretical risk indicators and perform three-dimensional complementary exposure characterization, supporting exposure assessment across environmental media; The weighting and ranking module is used to perform objective weighting based on statistical data features, expert knowledge-driven weighting based on decision matrices, and multi-attribute decision ranking. The inventory generation and full lifecycle management module is used to generate hierarchical inventories, perform uncertainty analysis, version control, and visualize uncertainty results. The dynamic update and visualization interaction module is equipped with interconnection interfaces with external pollution discharge permit databases, monitoring systems, industry information management platforms, and policy list release platforms, supporting data-driven trigger-based updates and cross-regional list comparison analysis.

[0014] A non-transient computer-readable storage medium storing a computer program thereon, including multi-environment media adaptation, cross-regional priority comparison, and trigger-based update functions based on changes in industrial policies.

[0015] Compared with existing technologies, the present invention provides a method and system for generating regional new pollutant inventories based on industry map extrapolation, which has the following advantages: (1) Breaking through the bottleneck of data scarcity and achieving "knowing before testing": In response to the pain points of new pollutants being difficult to monitor and data being scarce, this invention changes the lagging mode of existing technologies that rely on "monitoring first and then screening". By introducing a three-dimensional complementary exposure characterization system of "regional exposure index - local emission potential index - model-predicted concentration", it has for the first time realized a technical path of "environmental exposure" based on "regional industrial distribution map", which can accurately identify potential high-risk substances that are highly related to the local industrial structure when measured data is scarce.

[0016] (2) Complementary exposure framework to enhance assessment robustness: The three-dimensional exposure system of "literature meta-analysis + industry mapping + model prediction" proposed in this invention has strong anti-interference ability. Even when a certain dimension (such as measured data) is missing, scientific assessment can still be completed through CSW deduction or model simulation, which solves the problem of traditional methods being unable to calculate or causing assessment distortion due to missing data, and ensures the comparability of cross-regional and cross-material evaluation results.

[0017] (3) Dialectical Integration of Subject and Object, Full-Process Traceability: This invention adopts a differentiated weighting strategy that combines objective weighting based on information entropy with subjective weighting based on decision matrix, and couples it with multi-attribute methods. It not only utilizes statistical laws to mine the inherent information content of the data, but also corrects the bias of multi-source heterogeneous data through confidence level classification, realizing the scientific emphasis on local high-value data. On this basis, a complete closed-loop process is constructed: "hierarchical candidate library → multi-dimensional theoretical hazard characterization → industry map deduction → complementary three-dimensional exposure characterization → differentiated weighting → multi-attribute sorting → trigger-based version management". Combined with the metadata recording and version management mechanism throughout the entire life cycle, each step of the judgment logic in the list generation is auditable and verifiable, thereby significantly improving the scientific credibility and decision reference value of the list.

[0018] (4) Dynamic evolution, less testing and high efficiency: The built-in trigger-based update and targeted verification mechanism transforms the monitoring work from "blind large-scale census" to "model-driven small sample verification". While significantly reducing the monitoring cost, it realizes the ability of the list to continuously self-calibrate and evolve as external data accumulates, which has extremely high engineering application value. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the method for generating a new regional pollutant inventory as involved in the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0021] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0022] Please refer to Example 1 Figure 1 Show: To address the problems mentioned in the technical solution, this embodiment combines the appendix... Figure 1 The core process of this invention is described below. It should be particularly noted that the specific mathematical models or software tools used in this embodiment (such as entropy method, analytic hierarchy process, TOPSIS, EPISuite, and Monte Carlo simulation) are merely exemplary preferred examples for verifying the feasibility of the technical solution of this invention, and not the sole limitation on the scope of protection. Equivalent substitutions or combinations made by those skilled in the art using other algorithmic models with equivalent functions, based on an understanding of the core concept of this invention, should all be considered to fall within the scope of protection of this invention.

[0023] The initial screening database for candidate ECs was constructed based on the principles of hierarchy, traceability, and updability, and mainly includes the following modules: (1) Authoritative Policy List Compilation Module: Constructing "List I" The existing international and domestic authoritative regulatory directories and policy lists are systematically integrated and structured. A hierarchical directory and a unified metadata structure (such as chemical identifiers, basic physicochemical information, regulatory / source indexes, and version identifiers) are established according to chemical categories. A regular synchronization and version management mechanism is set up as a set of candidate sources that conform to regulatory guidance.

[0024] (2) Academic and experimental research convergence module: constructing "Checklist II" We retrieved and integrated recent research findings on risk monitoring and academic studies related to surface water, extracted substances that were frequently detected in actual measurements and labeled as high exposure / high hazard, and stratified and archived these academic / monitoring sources. We then supplemented and integrated these findings with "List I" to expand the coverage of potentially high-risk substances not covered by policies.

[0025] (3) Regional Industry Distribution Mapping Module: Constructing "List III" We compiled industry distribution directories and spatial layout information for the target region to construct a regional industry distribution map. Based on the strength of theoretical evidence regarding the correlation between substances and industries (mainly including chemical registration and use information, pollution coefficients of production processes, and typical industry emission characteristics reported in the literature), we performed a non-measured feature mapping between the candidates in "List II" and the regional industry distribution map. Projects unrelated to the region's dominant industries (e.g., through industry code comparison) were eliminated, and substances with theoretical emission potential were retained to form a representative candidate set for the region.

[0026] (4) Environmental adaptability screening module: forming a preliminary screening database of candidate ECs; Based on List III, items are systematically screened and categorized according to their suitability in the target environmental medium (i.e., by using physicochemical parameters, such as volatility indicators, to determine whether they are suitable for long-term aqueous monitoring and risk assessment). Key judgment criteria and metadata records are retained. For items with missing parameters or insufficient predictive reliability, a process of manual review and expert judgment is established to ensure the robustness of the initial screening results in engineering applications.

[0027] The data sources, filtering rules, evidence levels, and change records used in the above modules should all be archived in a unified metadata format and incorporated into a version control system to support subsequent weighting calculations, priority sorting, and traceable dynamic updates.

[0028] This embodiment further elaborates on the specific implementation approach of the theoretical risk assessment module used to quantify the inherent hazards of candidate substances. To achieve a systematic characterization of the hazard potential of substances, this module revolves around four key factors: biopersistence, bioaccumulation, ecotoxicity, and human health risk. Authoritative measured data are prioritized for each factor; when missing data, validated prediction methods or models are used to supplement it. All inputs, inferences, and uncertainties must be recorded as traceable metadata for weighting, ranking, and subsequent verification.

[0029] (1) Bio-persistence; Biopersistence is used to characterize the ability of pollutants to resist degradation and persist in environmental media for extended periods. In practice, available degradation half-life or biodegradation potential are used as quantitative indicators of persistence. When measured values ​​are unavailable, predictive methods recognized by the academic community are employed for estimation (such as EPISuite or equivalent QSAR tools), and the reliability of model predictions is labeled and manually verified. Standardized persistence indicators serve as a crucial input for theoretical risk assessment.

[0030] (2) Bioaccumulation; Bioaccumulation reflects the potential for contaminant accumulation and amplification along the food chain in organisms. When implementing bioaccumulation, it is estimated based on available enrichment factors or physicochemical properties and empirical relationships (such as the octanol-water partition coefficient, LogKow). Measured values ​​are preferred, supplemented by model predictions when necessary. Bioaccumulation indicators should be accompanied by a confidence statement and incorporated into uncertainty propagation analysis.

[0031] (3) Human health risks; Human health risk indicators are used to characterize long-term risks to the human body through routes such as drinking water and food. Comparative health risk indicators are constructed by combining toxicity reference values ​​(such as acceptable intake or reference dose) with exposure pathway information. Exposure scenarios can be set according to regional characteristics, but the assumptions and sources must be recorded in the metadata. For substances lacking direct toxicological thresholds, authoritative databases or validated prediction methods can be used to estimate and expand the uncertainty interval to reflect the inferred risk.

[0032] (4) Ecotoxicity; Ecotoxicity indicators are used to synthesize the toxicological responses of organisms at different trophic levels to assess ecological risk. Implementation is based on chronic or acute toxicity endpoints (such as NOEC / EC). 50 / LC 50 Based on (etc.) and combined with appropriate assessment factors, the ecologically ineffective concentration (PNEC) is calculated and generated. eco When direct toxicity data is lacking, accepted prediction methods can be used for derivation, and the derivation process and uncertainties should be clearly stated in the results.

[0033] (5) Construction of a theoretical evaluation system based on objective weights; After standardization / normalization, each theoretical risk factor enters the objective weighting and multi-attribute ranking process. In practice, an objective weighting calculation method (e.g., an entropy weighting method based on information content) is used to determine the indicator weights. Then, a multi-attribute decision-making method (e.g., a TOPSIS-like framework) is used to synthesize the normalized indicators into a single theoretical risk priority index (PI). tho The evaluation process outputs the following simultaneously: raw input data, normalized values, weights, ranking results, and uncertainty assessment report. All outputs are included in version control to support auditing and traceable recalculation.

[0034] (6) Data priority and missing value handling principles; Data acquisition follows the principle of "experimental testing first, authoritative predictions second, and model / statistical imputation as a backup." A conservative imputation strategy is implemented for missing key indicators, and the imputation uncertainty is quantitatively expanded before weighting and ranking. When data is severely missing or prediction reliability is low, risk warnings or high uncertainty markers are used in the output to indicate limitations for decision-making. Multi-source value processing follows traceable synthesis rules (e.g., prioritizing the most recent measured values ​​or selecting conservative values), and all synthesis decisions must be recorded in the metadata.

[0035] This module adopts a modular and configurable design, supporting both fully automated batch calculations and manual review and expert intervention. Implementation details (such as normalization methods, specific substitution strategies for evaluation factors, interpolation rules, and uncertainty propagation methods) can be listed in the appendix as preferred embodiments for engineering implementation as needed, but do not constitute necessary limitations of this invention.

[0036] This embodiment further illustrates the exposure risk assessment module used to quantify the exposure potential of candidate substances. To achieve scientific characterization of regional exposure with limited monitoring resources, this module introduces three complementary sub-indices: the Regional Exposure Index, the Composite Source Weight (CSW), and the Modelled PEC (Predicted Concentration). These are then fused into an Exposure Risk Priority Index (PI) using a subjective weighting and uncertainty propagation method based on a decision matrix. exp Each submodule adopts a modular, configurable, and traceable data structure design to support batch calculations, manual review, and recalculation based on new evidence.

[0037] (1) Regional exposure index; Regional exposure indices are derived from systematically collecting and structurally storing available surface water monitoring and measurement data. To address the problem of data scarcity in the target area, this invention innovatively introduces a "geographic-hydrological similarity weighted algorithm," the specific implementation steps of which are as follows: Data Mining and Cleaning: The system retrieves publicly available academic databases and environmental monitoring reports to extract measured concentration data (C) of the target pollutant in Chinese surface water basins. raw Strict data cleaning rules were implemented: noisy data lacking clear sampling points and QA / QC records were removed, and data below the detection limit (LOD) were replaced with LOD / 2.

[0038] Geographic similarity weight calculation: Construct feature vectors (including but not limited to hydrological features, industrial structure, geography, and climate) of the target region and the data source region. Calculate the similarity coefficient S between the two regions using Euclidean distance or cosine similarity. k (0 < S) k ≤1).

[0039] Weighted Deduction: Based on the similarity coefficient and multi-dimensional confidence weighting, the monitoring data of the source area is weighted and corrected to estimate the theoretical background concentration C of the target area target , where k represents different data source areas.

[0040]

[0041] In the formula: C target : The theoretical background concentration of the target area after weighted deduction; k: Data source number, representing the kth valid literature or monitoring record; C raw,k : The original measured concentration value extracted from the kth record; DF k : The detection frequency of the target substance in the kth record; W time,k (Timeliness Weight): Characterizes the timeliness value of the data, and different weights are assigned according to the time span from the sampling year to the present (for example, high weight is assigned to data in the past 5 years, and low weight is assigned to old data); W rep,k (Representativeness Weight): Characterizes the representativeness of the sampling point or sample. For example, data from national control / provincial control sections is better than ordinary points, and long-term monitoring data is better than sampling during dry and rainy seasons, which is better than single sampling; W type,k (Data Type Weight): Characterizes the statistical accuracy and certainty of the data value. High weight is assigned to exact statistical data clearly reported as the median; the arithmetic mean is次之, and derivative data that needs to be converted through the concentration range is assigned a lower weight; W qa,k (Quality Control Weight): Characterizes the credibility of the data, and is assigned values according to whether there are complete QA / QC records (such as recovery rate, blank experiment, parallel samples); S k (Geographical Similarity Coefficient): Characterizes the similarity degree of the data source area and the target area in terms of hydrological characteristics, industrial structure and geographical climate (0 < Sk ≤ 1). The higher the similarity, the greater the reference value.

[0042] The present invention introduces the "dual-track risk quantification and gap interpolation algorithm": (1) When the amount of literature data N > 0: Calculate the statistical value of all valid monitoring data according to the above formula as the regional background concentration of the substance (C target,k ); (2) When N = 0: Activate the risk conservative data gap interpolation mechanism, and assign the normalized median score (such as 0.5) to this index to represent the "potential average risk" and prevent the risk omission caused by research gaps.

[0043] Standardized output: The derived C target The data is compared with and normalized to a preset benchmark value (such as the regional average background value or environmental standard), and the regional exposure index between 0 and 1 is output. All data sources, weighting rules and processing steps are recorded as metadata to ensure transparency.

[0044] (2) Local Emissions Potential Index (CSW); The Local Emission Potential Index is used to quantify the theoretical emission drivers of substances in a target area when measured emission data is unavailable. Its core logic is to establish a priori correlation matrix between "industry category" and "pollutant". The specific implementation steps are as follows: Constructing a regional industry distribution map: Extracting industry category distribution data for the target region, such as national economic industry classification codes and sewage outlet industry classification (if the target region is a river basin), and calculating the relative importance weight W of each industry. eightj (Based on the proportion of output value or the proportion of the number of enterprises); Correlation of scoring mapping and quantification rules (establishing a score) ij (Matrix): To transform qualitative industry correlations into quantitative calculation parameters, this embodiment establishes a hierarchical assignment rule. Based on authoritative databases (such as PRTR and EPA lists) and literature evidence, the system performs a semi-quantitative hierarchical assignment of the correlation strength between each candidate i and industry j: Strong correlation (Score=3.0): The substance is a characteristic pollutant, main raw material, intermediate or characteristic by-product of the industry (e.g., PFOA is used in the manufacture of fluorine-containing materials). Medium correlation (Score=2.0): This substance is widely used in the industry or as an auxiliary additive, cleaning agent, or may be generated in the process (e.g., antibiotic residues in pharmaceutical wastewater). Weak correlation (Score=1.0): The substance is used in small quantities in the industry as a general chemical, or there are only trace detection records. No association (Score=0.0): Clear process exclusion or no supporting evidence.

[0045] Confidence correction (introducing Q) ij Factor): Considering the differences in reliability among different data sources, this algorithm further introduces the evidence confidence coefficient Q. ij To address uncertainty. For corroborating evidence from national / international official lists (such as EPA), set Q. ij =1; For relevant evidence from peer-reviewed academic literature, set Q =1. ij =0.7; For associations based solely on chemical structure similarity inference (QSAR derivation), Q is set to... ij=0.4.

[0046] Source strength inference algorithm execution: This step constitutes the core data generation and preprocessing algorithm of this 3D exposure model. The comprehensive source strength weights are calculated by performing a weighted convolution operation including confidence correction using the following formula:

[0047] In the formula, CSW i Let i be the local emission potential index for the i-th substance. Weight j The normalized weight of the j-th industry (reflecting the industry structure). Score ij The strength of the theoretical correlation between materials and industries (reflecting process characteristics). Q ij This is the confidence penalty coefficient, used to weight and adjust evidence from different sources. Special note: The... Score ij The assignment is based entirely on a non-targeted industry characteristic knowledge base, without relying on the actual measurement records of sewage outlets of specific enterprises in the region, thereby achieving proactive early warning of potential risks.

[0048] (3) Theoretical predicted concentration (ModelledPEC); The model-based concentration prediction is based on a theoretical baseline derived from multi-media facultative simulations. This indicator aims to fill the gaps in monitoring data at the mechanistic level, providing a physicochemical theoretical baseline for exposure assessment that is unaffected by sampling bias. This step is achieved by constructing a "configurable environmental facultative simulation engine," the specific logic of which is as follows: Input parameter coupling: The system automatically calls the key physicochemical parameters of the candidate (such as degradation half-life, LogKow, Henry's constant, half-life) as internal variables, and combines them with the macro-hydrological characteristics of the target area (such as average flow rate, flow velocity, hydraulic residence time) as external scenario parameters.

[0049] Simulation engine kernel: A suitable multi-media mass balance model (such as Fugacity Level III or an equivalent box model) is used as the computational kernel. Under a standardized evaluation scenario, the migration, distribution, and degradation processes of various substances under steady state are simulated.

[0050] Output and Standardization: Calculate the steady-state predicted environmental concentration (PEC) in the target aqueous phase. It should be noted that the PEC here is not intended to precisely reproduce the instantaneous value at a specific monitoring section, but rather to quantify the relative fate tendency of different pollutants under the same environmental background. This PEC value, after logarithmic transformation and normalization, serves as the third dimension input to the exposure index system.

[0051] Uncertainty management: The scenario assumptions of the model, the sensitivity analysis results of the input parameters, and the simulation confidence should be recorded as metadata and participate in the subsequent weight allocation (i.e., when there are many missing input parameters of the model, the weight of that dimension should be automatically reduced).

[0052] (4) Construction of an exposure evaluation system based on expert decision matrix weighting A three-dimensional complementary system is constructed, consisting of a regional exposure index (measured domain), a local emission potential index (projected domain), and a theoretically predicted concentration (model domain). To address the significant differences in confidence levels and local correlations among multi-source heterogeneous data, this exposure assessment system introduces a subjective weighting strategy based on a decision matrix (such as the Analytic Hierarchy Process (AHP)). Then, a multi-attribute decision-making method (such as a TOPSIS-like framework) is used to synthesize the normalized indicators into a single exposure risk priority index (PI). tho The evaluation process outputs the following simultaneously: raw input data, normalized values, weights, ranking results, and uncertainty assessment report. All outputs are included in version control to support auditing and traceable recalculation.

[0053] This module design emphasizes configurability and engineering applicability: weights, geographical weighting coefficients, scoring levels, and model scenarios can be adjusted according to regional needs, and related adjustments must be based on evidence and recorded; implementation details (such as weighting strategies, standardization methods, and uncertainty propagation details) can be listed as preferred implementation methods in the appendix for engineering implementation, but do not constitute necessary limitations of this invention.

[0054] This embodiment further illustrates the module for assessing pollutant weights / priorities and generating a regional high-risk list. To achieve objective ranking and hierarchical management of candidate pollutants, this invention adopts a technical approach that combines a differentiated weighting strategy that couples subjective and objective weighting with multi-attribute decision-making.

[0055] It should be particularly noted that the core inventiveness of this invention in terms of methodology lies in constructing a full-process evaluation index system for data-scarce scenarios, namely, "hierarchical candidate library → multi-dimensional theoretical hazard characterization → industry map extrapolation → complementary three-dimensional exposure characterization → differentiated weighting → multi-attribute ranking → trigger-based version management." Regarding the mathematical processing steps after the index system is constructed (including but not limited to weighting, dimensionality reduction, ranking, and fusion), those skilled in the art, based on a full understanding of the core concept of this invention, can flexibly select various existing multi-attribute decision-making algorithms or their improved models for equivalent implementation. Therefore, the entropy method, TOPSIS, etc., listed below are merely non-limiting preferred embodiments for verifying the effectiveness of the index system of this invention and should not be considered as the sole limitation on the scope of protection of this invention.

[0056] Weighting (Example of an Objective Weighting Model): As mentioned above, this step aims to determine objective weights based on the statistical distribution characteristics of theoretical risk data. As a non-limiting preferred embodiment of the present invention, this embodiment details the calculation logic based on information entropy theory (only the core formulas are shown) for reference by those skilled in the art: (1) Construct the original evaluation matrix (Y): The standardized scores of the n pollutants to be evaluated on m indicators are used to construct a matrix, which is called Y and has a dimension of m×n.

[0057] (2) Normalization: Normalize the scores of each pollutant under each index i to obtain the probability matrix P=[P ij A correction is used to avoid zero values.

[0058] (3) Calculate the information entropy e i The representative formula is as follows:

[0059] (4) Calculate the difference coefficient and entropy weight W i Coefficient of difference d i =1-e i The entropy weight is normalized as follows:

[0060] Note: The above W i The objective weights for each indicator are calculated; the weights for the respective indicator groups of theoretical risk and exposure risk are calculated independently.

[0061] Weighting (Example of Subjective Weighting Model): As mentioned above, this step aims to introduce expert prior knowledge and multi-source data confidence assessment to determine the subjective weights of exposure risk and correct for biases in statistical data. As a non-limiting preferred embodiment of the present invention, this embodiment details the computational logic based on the Analytic Hierarchy Process (AHP) (revised version) for reference by those skilled in the art: (1) Constructing the judgment matrix (A): Based on the local relevance and reliability classification of the data source, perform pairwise comparisons of indicators at the same level. Construct the judgment matrix A=[a ij ] n×n , where a ij Assigning values ​​using the Saaty 1-9 scale indicates the importance of index i relative to index j (e.g., a). ij =5 indicates that the former is significantly more important than the latter.

[0062] (2) Weight Vector Calculation: The maximum eigenvalue of the matrix and its corresponding normalized eigenvector are calculated using the Sum-Product Method or the square root method. Taking the Sum-Product Method as an example, the column vectors of the matrix are first normalized, and then the arithmetic mean of the row vectors is calculated to obtain the normalized weight vector W=[W1,W2,...,Wn] for each indicator. T .

[0063] (3) Consistency check: To ensure logical consistency, the consistency index CI needs to be calculated as follows: CI = (λ max The weight vector is calculated as follows: -n) / (n-1) and the consistency ratio CR = CI / RI (RI is the average random consistency index). If CR < 0.1, the judgment matrix is ​​considered to have satisfactory consistency, and the weight vector is valid; otherwise, the judgment matrix needs to be modified until it passes the test.

[0064] In this embodiment, the subjective weighting method is mainly applied to the exposure risk assessment module to establish the weight hierarchy relationship between the local emission potential index (CSW), the theoretical predicted concentration (PEC), and the regional exposure index.

[0065] Priority scoring (example of a multi-attribute decision model): After obtaining the weights, the system needs to rank the candidates in multiple dimensions. As a highly robust optimization solution, this embodiment uses the Top-Optimal Solution Ranking Method (TOPSIS) to build the ranking engine. The specific calculation steps are as follows (only the core formulas are shown): (1) Constructing a weighted standard matrix: Standardize the original data (Min-Max or vector normalization can be selected) to obtain a dimensionless matrix X. ij Multiply by the corresponding weight Wi to obtain the weighted matrix V ij =W i X ij .

[0066] (2) Determine the positive ideal solution and the negative ideal solution V + and V - The calculation formula is as follows.

[0067]

[0068]

[0069] (3) Calculate the Euclidean distance of each pollutant to the ideal solution.

[0070]

[0071]

[0072] (4) Calculate the relative proximity score (TOPSIS score, denoted as C). j ):

[0073] C j The value is between 0 and 1. The closer it is to 1, the closer it is to the ideal solution and the higher its priority.

[0074] Overall Priority and Grading Rules: Overall Priority and Grading Rules: Theoretical Risk Index (PI) will be used as the basis for priority and grading rules. tho ) and Exposure Risk Index (PI) exp Multidimensional coupling is performed to generate the final comprehensive priority index (PI).

[0075] Preferred fusion rule: Nonlinear synthesis is performed using the Geometric Mean method based on risk definition, and the calculation formula is as follows. This algorithm aims to reflect the product coupling effect between hazard and exposure, strengthen the "no exposure equals no risk" veto mechanism, and avoid artificially inflating the risk of highly toxic, low-emission substances.

[0076]

[0077] Tiered and Dual Inclusion Rules: PI tiering thresholds are set based on regional management needs (exemplary divisions are as follows, which can be dynamically adjusted according to actual distribution): Tier I (extremely high risk, top 10%), Tier II (high risk, top 10% to top 30%), Tier III (medium risk, top 30% to top 50%), Tier IV (low risk, bottom 50%). This invention implements a dual-list admission mechanism: First admission: Substances with assessment results of Tier I and II are directly included in the "High-Risk Core List (High-Risk List I)"; Second admission: To avoid potential omissions of high-risk substances due to dilution of the comprehensive score, a Risk Quotient (RQ) is introduced as a supplementary judgment indicator for Tier III and Tier IV substances. Substances with an RQ exceeding a preset safety threshold (e.g., RQ>1) are included in the "High-Risk List II" (see...). Figure 1 Ultimately, the high-risk list I and high-risk list II are merged to form the aforementioned regionalized high-risk new pollutant list.

[0078] Uncertainty propagation and robustness testing: The uncertainty of input data and weighting parameters is propagated by numerical simulation or resampling to obtain the confidence interval of each substance's PI or the probability of being listed as high risk, and to provide a reference for robustness decision-making. At the same time, sensitivity analysis is carried out to identify indicators or data sources that have a significant impact on the ranking, and to provide a basis for prioritizing data collection and field verification.

[0079] Dynamic updates and traceable version management: Establish an update mechanism that combines periodic and trigger-based updates. When preset trigger conditions are met (such as updates to the authoritative list, identification of high-confidence candidates through non-targeted screening, detection of significant anomalies, or publication of important literature), the update of the candidate library and input data, quality verification, recalculation of PI, and necessary calibration of grading thresholds are performed automatically or semi-automatically. All updates, parameter adjustments, and calculation results should be versioned and archived, and the reasons for changes should be saved for traceability and review.

[0080] In the aforementioned dynamic update process, small-scale field monitoring (small sample validation) is permitted and recommended when necessary as a means of quality assurance and model calibration. This small-scale monitoring is an optional implementation measure, and its purpose is: (1) Provide local empirical evidence for regional exposure index and local emission potential (CSW) to support or revise model assumptions; (2) Evaluation of theoretical PI exp Compared with measured PI exp Or the consistency of the final rankings and trigger a correction procedure; (3) Provide on-site verification basis for changes to the list due to weight or threshold adjustments.

[0081] The method described in this section is based on the principles of modular and configurable design. Its specific implementation strategies (including normalization methods, specific algorithm selection for weighting and sorting, criteria for secondary screening, and numerical methods for uncertainty propagation) can be illustrated in the application examples as needed to promote engineering implementation. However, the above specific implementation details do not constitute a limitation on the scope of the invention.

[0082] Example 2 uses a typical watershed in Chengdu as a demonstration area. Following the approach of this invention, a total of 423 new pollutants were selected as the initial candidate set from the "Regional Candidate ECs Preliminary Screening Database". The CAS number, key physicochemical parameters, and data source of each chemical substance are recorded in the database using uniform fields to support subsequent evaluation and traceability verification.

[0083] Data and Information Source Explanation: This embodiment uses three types of data sources in combination: 1. Authoritative international / domestic lists and policy directories on high-risk areas; 2. Recent literature on measured risk assessment of surface water in China; 3. Information on sewage outlets and industrial emissions in the target area.

[0084] The data is written into the database in a structured manner and includes source, year, and quality markers, providing a robust data foundation for model operation and subsequent calibration.

[0085] Theoretical hazard dimension (weight 50%): focuses on the inherent risks of substances. Among them, human health impact (weight 0.2935) and ecotoxicity (weight 0.2934) are dominant, while persistence (0.2257) and bioaccumulation (0.1873) are secondary.

[0086] Exposure risk dimension (weight 50%): focuses on the probability of actual exposure in the region. Based on the decision matrix calculation, the local emission potential index (CSW) was assigned the highest weight (0.539), significantly higher than the model-predicted concentration (0.297) and the regional literature exposure index (0.164). This weighting strategy fully reflects the core technical idea of ​​this invention: "prioritizing the inference of local industry characteristics in data-scarce scenarios."

[0087] The comprehensive assessment results and risk grading system completed end-to-end calculations for 423 candidate substances, generating a comprehensive priority index (PI) ranging from 12.27 to 58.33. Based on the PI score distribution characteristics, the system automatically divided the list into three risk levels: (1) Level I High Risk (Top 5%, PI>48.0): There are 21 types, including perfluorooctane sulfonic acid (PFOS, Rank 1), bisphenol A (Rank 2), perfluorooctanoic acid (PFOA, Rank 3), etc.

[0088] (2) Level II Higher Risk (Top 5%-15%): A total of 42 types, covering a variety of organophosphorus flame retardants and antibiotics.

[0089] (3) Level III low to medium risk (the rest): need to be monitored or included in the watch list.

[0090] The ranking results show that perfluorooctane sulfonate (PFOS, PI=58.33), bisphenol A (BPA, PI=55.46), and perfluorooctanoic acid (PFOA, PI=51.59) consistently rank at the top. This is highly consistent with international conventions and the list of key controlled new pollutants, proving the scientific validity of this system in comprehensive hazard characterization.

[0091] Accurately identifying hidden alternatives: 6:2 fluorosulfonic acid (6:2 FTSA) ranks 8th overall in this list (PI=48.44), and its exposure priority index is as high as 63.69 (ranked 6th). This substance, as an alternative to PFOS, is currently often overlooked in routine monitoring. However, due to the significant electroplating and fine processing industries in Chengdu, the CSW module (weight 0.539) of this system keenly detected its extremely high local emission potential, thus including it in the "Level I High-Risk List". A similar finding is that another typical substance, bisphenol S (BPS), although theoretically of moderate risk, has seen its exposure priority jump to 10th place due to its wide potential application as a substitute for bisphenol A in local industries, successfully triggering an alert.

[0092] Based on the above calculation results, this system generates a targeted "Top 20 High-Risk Substance Targeted Verification Plan" to replace the traditional "blind general survey". It is recommended to conduct precise sampling for the top 20 substances in the PI ranking (such as PFOS, PFOA, 6:2FTSA, PCB-169, etc.) and high-load areas indicated by CSW (such as downstream sections of specific industrial parks).

[0093] Consistency assessment and calibration: Compare the measured concentration with the model's predicted trend (PI ranking). If the rank correlation coefficient is lower than a preset threshold (e.g., 0.6), the parameter inversion mechanism is triggered, automatically correcting the industry weights or scores in the CSW calculation formula using the following formula: =

[0094] Among them, W i old For the original entropy weight, ρ target ρ is the target relevance threshold (e.g., 0.7). obs The values ​​are the measured values, and α is an adjustable scaling factor, where (example values ​​are 0.3-0.5).

[0095] The above findings fully validate the ability of this method to identify regulatory blind spots and achieve proactive early warning. In terms of cost-effectiveness, if the traditional "large-scale field survey" model were used to comprehensively cover the initially screened 423 substances, the estimated implementation cost would be approximately RMB 1.5 million (including on-site sampling, purchase of standards, purchase of isotope internal standards, pretreatment column collection, instrument testing, and labor costs). However, this embodiment, through a "theoretical deduction + targeted verification" approach, only requires precise monitoring of 20 high-confidence substances, with a total implementation cost (including data processing and targeted sampling) of only approximately RMB 100,000. Quantitative calculations show that this method reduces monitoring costs by 93.3% and the number of species requiring on-site testing by approximately 95.3%, achieving a shift from "comprehensive blind testing" to "small-scale, high-value targeted sampling." It also outputs a 90% confidence interval for the comprehensive priority index (PI) for decision-making reference, significantly improving the economy and scientific rigor of control decisions. This embodiment only demonstrates typical calculation logic to illustrate the feasibility and typical application of the invention; the actual implementation process can be flexibly configured according to regional needs. Any adjustments, extensions, or modifications made by those skilled in the art to the parameters, data, or processes without departing from the spirit of this invention shall be deemed to fall within the protection scope of this invention.

[0096] Comparative Example 1 (Comparison of document aggregation technologies); The study on prioritizing emerging organic pollutants by Zhong et al. (2022) was selected as a representative of existing "literature-aggregated" technologies. Its core process involves constructing a candidate set of 405 unregulated new pollutants based on published literature and historical monitoring data. A two-dimensional, multi-criteria framework of "hazard potential (based on PBT properties) + exposure potential (based on reported concentration statistics and detection frequency)" is used for screening and ranking. This approach heavily relies on historically accumulated academic literature data and publicly available government monitoring reports, making it suitable for large-scale, retrospective preliminary screening of new pollutants at the national level.

[0097] In the context of refined local management, this approach has significant limitations: First, it suffers from data dependency bias, limited by literature coverage, leading to an inflated ranking of "popular substances" while "locally unique but not widely studied new substances" are directly missed due to missing literature records, exhibiting a clear "survivorship bias." Second, its risk identification model is outdated, only able to passively review known pollution, lacking foresight. The key difference between this invention and the present invention lies in the fact that, while retaining and inheriting the scientific core of "hazard + exposure" multi-indicator assessment, this invention maps "regional industrial distribution maps" with "substance-industry correlation matrices" to directly infer locally emitted substances from the source, and establishes a three-dimensional complementary system consisting of "regional exposure index (literature domain) + CSW (inference domain) + model predicted concentration (model domain)," thus enabling proactive identification of local hidden risks even in data-scarce scenarios.

[0098] Compared with the comparative schemes, this invention achieves a fundamental shift from "passive review based on historical documents" to "proactive early warning based on industry characteristics." Its beneficial effects are as follows: First, it greatly improves the sensitivity of risk identification, effectively capturing high-risk substances that are overlooked in the literature but conform to local industry characteristics, and significantly reducing the false negative rate. Second, it enhances engineering applicability; the "trigger-based update" and "small-sample targeted calibration" mechanisms proposed in this invention solve the engineering problem of static literature research results that cannot dynamically serve local refined management.

[0099] Comparative Example 2 (Comparison of large-scale experimental techniques) The large-scale field-based prioritization study on wastewater treatment plants and watersheds by Liu et al. (2024) was selected as a representative of existing "measurement-driven" technologies. This approach relies on a large-scale quantitative survey covering 46 wastewater treatment plants nationwide, employing advanced analytical methods such as LC / GC-MS / MS to obtain measured exposure levels (median concentration and detection frequency), and combining this with principal component analysis (PCA) to synthesize a priority index. This approach is highly dependent on expensive and intensive field sampling and laboratory measurement data, making it suitable for constructing high-precision, evidence-based lists of key pollutants for control.

[0100] In watershed or local-level dynamic management scenarios, this approach faces two major drawbacks: first, its cost-effectiveness is low, with survey costs often reaching millions, which is unbearable for resource-constrained areas; second, its spatiotemporal representativeness is limited, as a single large-scale sampling can only reflect a momentary projection at a specific point in time, making it difficult to capture long-term dynamic risks caused by cyclical industrial emissions. The key difference between this invention and the present invention lies in the "complementary three-dimensional exposure system," which transforms the monitoring role from "comprehensive survey" to "model-driven Top-k targeted validation." While ensuring scientific rigor, it reduces the number of species requiring actual measurement by more than 90%, achieving a highly cost-effective dynamic monitoring and early warning mechanism.

[0101] Compared with traditional methods that rely on large-scale field testing, this invention has significant advantages in terms of implementation cost control and management efficiency. Its beneficial effects are mainly reflected in the following aspects: through a closed-loop model of theoretical deduction + targeted verification, it achieves a shift from "comprehensive blind testing" to "small-scale, high-value targeted sampling," significantly reducing human and economic costs; simultaneously, this invention overcomes the "instantaneous" limitation of measured data, enabling it to provide scientific early warnings of continuous risks based on industry characteristics, making it particularly suitable for local management scenarios where monitoring resources are limited and rapid response is required.

[0102] Comparative Example 3 (Comparison of "Measurement Record Screening" Techniques) To further highlight the technical advantages of this invention in the context of "scarce monitoring data", a typical technical solution in the prior art of "screening based on actual detection records of sewage outlets or environmental media" (which is also the mainstream technical route for routine monitoring of groundwater or surface water in industrial parks) is selected as a comparison object.

[0103] (1) Simulation scenario setting: Assume that a large-scale fluorine-containing new material production enterprise has been newly introduced into the target watershed. The enterprise uses a new substitute, "6:2 fluorinated polysulfonic acid (6:2FTSA, CAS27619-97-2)". Since this substance is a new PFASs substitute, there is currently no detection record of this substance in the local historical monitoring database (there is a data gap). However, the enterprise is emitting 6:2FTSA, which objectively poses an environmental risk.

[0104] (2) Comparison of the results of the scheme (missed judgments): Traditional methods (such as the method described in patent CN121073243A) follow a post-hoc logic of "detection = risk". Since no historical test records for 6:2 FTSA can be found in the database (because it has never been measured), the comparative method determines that the "correlation" or "detection reliability" of this substance is 0. This substance is directly eliminated in the initial screening stage, leading to significant "false negatives" (high-risk substances are missed). This demonstrates that methods relying on actual measurements have inherent blind spots when facing "unknown" or "new" contaminants.

[0105] (3) Results of the operation of the method of the present invention (success warning): This invention follows the priori logic of "industry = potential emissions." The system identified a significant increase in the weight of the "fluorine-containing new materials manufacturing" industry within the region. Subsequently, it invoked the industry characteristic knowledge base and identified 6:2 FTSA as a typical substitute raw material / byproduct of this industry (Score=3.0). Despite the lack of measured data, the system calculated an extremely high exposure index and automatically triggered a "complementary assessment mode." Ultimately, 6:2 FTSA was successfully captured by the system and included in the "Level I Potential High-Risk List" (indicating the need to prioritize the development of monitoring methods for confirmation), achieving proactive early warning without prior detection. The specific risk assessment process is as follows: Theoretical risk characterization: Using the system database, the biopersistence (DHL) of 6:2 FTSA is 0.87, the bioaccumulation (LogKow) is 2.66, and the human health impact threshold (PNEC) is [not specified]. hum The concentration was 17.742 μg / L, which is below the ecological risk threshold (PNEC). eco The concentration was 865.823 μg / L. Based on multi-attribute decision assessment, the theoretical risk priority index (PI) of the 6:2 FTSA was... tho The value is 28.69.

[0106] Local Emission Potential (CSW) Projection: Although there are no local measured records, the system identified a significant weight for the fluorine-containing new materials manufacturing industry within the region. Based on the correlation matrix mapping, 6:2 FTSA, as a characteristic byproduct, is strongly correlated with the aforementioned industry (Score=3.0), resulting in a CSW index as high as 0.69, with a model-predicted concentration as high as 0.68. Although the existing measured exposure index in the region is at a moderate level (0.55), the high CSW weight directly increases the overall exposure risk.

[0107] Overall assessment conclusion: Based on the 3D exposure profile results, the system calculated the 6:2 FTSA exposure risk priority index (PI). expThe score is approximately 68.03. The final composite priority index (PI) is 48.44. According to the classification rules described in this invention, this score falls into the Level I high-risk range (top 10%), and the 6:2 FTSA is successfully identified and directly included in the "High-Risk List I", with priority monitoring recommended.

[0108] (3) Comparison of the adaptability and performance of the present invention with the comparative scheme: To demonstrate the innovative advantages of this solution, the following table, based on theoretical evaluation, summarizes the performance comparison results of the present invention and the comparative solution (traditional measured-driven type) under different monitoring resource guarantee scenarios. As shown in Table 1, the performance comparison results of the present invention and the comparative solution (traditional measured-driven type) are as follows. Table 1 shows the performance comparison results between the present invention and the comparative scheme (traditional measured drive type):

[0109] Quantitative comparison results show that the present invention, through the closed-loop model of "industry map extrapolation + targeted verification", demonstrates an absolute adaptability advantage in the extreme scenario of "data scarcity", reducing the number of species that need to be tested in the field by more than 80%, and improving the foresight of risk identification by several times compared with the comparison scheme, significantly enhancing the economy and scientific nature of watershed new pollutant management.

[0110] Note: The above three comparative examples are only intended to objectively illustrate the improvements of this invention over typical literature-driven research in terms of engineering implementation, regional adaptation, and methodological robustness. The purpose is to highlight the inventiveness and practical value of this invention, and not to negate such cutting-edge research.

[0111] Additional Implementation Notes: The method and system of this invention can be extended to other river basins or administrative regions (e.g., the upper reaches of the Yangtze River, the middle reaches of the Yellow River, the Pearl River Basin, etc.). This extension and adaptation is achieved by introducing regional characteristic data and regionally adjusting the industry weights and geographical weighting parameters in the local emission potential index. After configuring the regionalized parameters, this invention can quickly generate a candidate set and priority list of high-risk new pollutants suitable for the target region, supporting cross-regional migration and on-site secondary development to meet the management needs and data availability differences of different river basins.

[0112] In summary, this invention proposes a method and supporting system for generating a regionalized inventory of new high-risk pollutants based on multi-source data fusion, differentiated weighting strategies, and industry-emission mapping. This technology, even with limited monitoring data, enables scientific screening of candidates, theoretical and exposure-based dual-end assessment, high-confidence ranking, uncertainty quantification, and traceable dynamic updates, thus providing an actionable decision support tool for refined local supervision, monitoring resource optimization, and tiered management.

[0113] Scope of Protection: The scope of protection of this invention is not limited to the specific embodiments described above. Any equivalent substitutions, transformations, or improvements made to the method steps, parameter configurations, module implementations, or system architecture within the technical concept and principle framework disclosed in this invention, without departing from the core essence of this invention, should fall within the scope of protection of this invention. Although embodiments of this invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of this invention. The scope of this invention is defined by the appended claims and their equivalents.

Claims

1. A method for generating a regional new pollutant inventory based on industry map extrapolation, characterized in that, Includes the following steps: S1, Construction of a hierarchical candidate pollutant database: Heterogeneous data fusion is carried out by combining authoritative control lists issued by international and national authorities and cutting-edge field measurement research results to construct a candidate set driven by both policy and academia; the candidate set is mapped to the industrial distribution map of the target region, and a preliminary screening database of candidate pollutants with adaptability to the target region is obtained based on the rules of material-industry correlation and environmental media compatibility. S2, Multidimensional Characterization of Theoretical Hazards: Candidate pollutants are quantitatively characterized according to representative theoretical factors related to biological persistence, bioaccumulation, ecotoxicity and human health risks, forming a standardized theoretical risk index matrix. S3, Local Emission Potential Deduction: Construct an industrial distribution map of the target area, establish industry weight vectors, and build a correlation scoring matrix between pollutants and industries. The local emission potential evaluation results for each candidate pollutant are derived through matrix weighting operations. The establishment of the correlation scoring matrix between pollutants and industries in S3 specifically includes, but is not limited to, multi-source evidence based on authoritative lists, literature evidence, chemical registration usage information, production process inferences, or industry pollution characteristic databases, to perform graded assignment of the correlation strength between pollutants and industries. The matrix weighting operation is specifically configured as follows: the correlation scoring matrix and the industry weight vector are weighted and calculated, and an evidence confidence coefficient is introduced during the calculation process to weight and correct the reliability of data from different sources. Finally, the calculation results are normalized to obtain the local emission potential index. The industry weight vector is determined based on indicators that characterize industry features, including but not limited to industry output share, number of enterprises share, pollution discharge share, or regional industrial policy priority. S4, Complementary 3D Exposure Characterization: Constructing a 3D complementary exposure characterization system consisting of regional exposure correlation index, local emission potential index, and model-predicted concentration; establishing a data availability discrimination and missing data imputation mechanism, where data from other dimensions is used for imputation or mutual correction when any dimension is missing or has low confidence, to obtain a comprehensive exposure assessment result; the regional exposure correlation index mentioned in S4 is calculated from data mining and cleaning, geographically similar weighted data, and cross-regional analogy data; the model-predicted concentration is a relative exposure reference value derived based on preset environmental behavior simulation rules, material physicochemical properties, and geographical parameters of the target area. The three-in-one mechanism is configured with dynamic weight allocation logic to ensure that even when measured data is scarce, exposure indicators are comparable across regions, substances, and environmental media. S5, Measurement and multidimensional ranking of indicator weights: Construct an evaluation matrix of theoretical risk and exposure risk, use an objective weighting model to calculate the weight of each theoretical indicator, use a subjective weighting model to calculate the weight of each exposure indicator, and use a multi-attribute decision ranking model to calculate the theoretical risk priority and exposure risk priority of each candidate pollutant. S6, Comprehensive Judgment and Dynamic Versioning Management: According to the preset fusion rules, the theoretical risk priority and the exposure risk priority are combined into a comprehensive priority, and a regionalized list of high-risk new pollutants is generated accordingly.

2. The method for generating a regional new pollutant inventory based on industry map extrapolation according to claim 1, characterized in that, Also includes: Establish an uncertainty propagation analysis mechanism, a traceable list version update mechanism based on trigger conditions, and a feedback-based targeted verification step; Based on the generated comprehensive priority ranking results, high-risk substances and representative locations and representative environmental media of a preset proportion or number are selected for targeted on-site monitoring. The consistency between on-site monitoring results and model prediction results is evaluated. When the deviation exceeds the preset threshold, the weighting weights, model parameters or material-industry association rules are modified, and the modification log is recorded to ensure traceability.

3. The method for generating a regional new pollutant inventory based on industry map extrapolation according to claim 1, characterized in that, The candidate pollutant database described in S1 is stored in a three-level hierarchical structure, including: an authoritative policy list layer as basic compliance input, a data aggregation layer as a supplement to cutting-edge risks, and a regional adaptation layer as localized filtering results; the database configures a unified metadata field for each candidate pollutant, which includes at least a unique chemical identifier, key physicochemical properties, data source evidence level, version timestamp, and environmental media adaptability identifier.

4. The method for generating a regional new pollutant inventory based on industry map extrapolation according to claim 1, characterized in that, The weighting process described in S5 adopts a differentiated weighting strategy: the weights of theoretical risk indicators are determined using an objective weighting model based on data statistical characteristics; The weights of exposure risk indicators are determined using a subjective weighting model based on an expert decision matrix, in order to strengthen the driving role of local industry characteristics in exposure assessment. All intermediate process data, weight parameter settings, and sorting results generated during the weighting and sorting process are permanently saved in the form of metadata and incorporated into the version management module to support subsequent calculation auditing, result verification, and model calibration.

5. The method for generating a regional new pollutant inventory based on industry map extrapolation according to claim 2, characterized in that, The uncertainty propagation analysis mechanism includes, but is not limited to: numerical simulation, resampling analysis, or sensitivity analysis of the uncertainty distribution of key input parameters; calculating the confidence interval of the comprehensive priority, the statistical characteristics of the high-risk inclusion probability or robustness coefficient; and using the above statistical characteristics as part of the list decision reference. The triggering conditions of the uncertainty propagation analysis mechanism include, but are not limited to: access of new high-confidence measured data, authoritative list updates, monitoring events exceeding thresholds, regional industrial restructuring, release of research results on the risks of new pollutants, or changes in policy control requirements. When the triggering conditions are met, the system automatically or semi-automatically triggers the recalculation of the priority of relevant candidate pollutants and generates a new version of the list.

6. A system for generating a regional new pollutant inventory based on industry map extrapolation, applicable to the method for generating a regional new pollutant inventory based on industry map extrapolation as described in any one of claims 1-5, characterized in that, The core components include: The multi-source data access and fusion module is used to integrate policy lists, literature data, and regional industry distribution information; The multi-source data access and fusion module is used to integrate policy lists, literature data, regional industrial distribution information, environmental media characteristic data, and measured monitoring data; The candidate library hierarchical management module is used to build and maintain a three-level candidate pollutant library and supports multi-environment media adaptability management. The theoretical hazard and exposure assessment module is used to calculate theoretical risk indicators and perform three-dimensional complementary exposure characterization, supporting exposure assessment across environmental media; The weighting and ranking module is used to perform objective weighting based on statistical characteristics of data, subjective weighting based on decision matrices and expert experience, and multi-attribute decision ranking. The inventory generation and full lifecycle management module is used to generate hierarchical inventories, perform uncertainty analysis, version control, and visualize uncertainty results. The dynamic update and visualization interaction module is equipped with interconnection interfaces with external pollution discharge permit databases, monitoring systems, industry information management platforms, and policy list release platforms, supporting data-driven trigger-based updates and cross-regional list comparison analysis.

7. A medium for generating a regional new pollutant inventory based on industry map extrapolation, wherein a computer program is stored thereon, and the computer program, when executed by a processor, implements the steps of the method for generating a regional new pollutant inventory based on industry map extrapolation as described in any one of claims 1-5, characterized in that, It includes multi-environment media adaptation, cross-regional priority comparison, and trigger-based update functions based on changes in industrial policies.

Citation Information

Patent Citations

  • System for life cycle risk assessment

    KR100750617B1