Chao lake new pollutant composite early warning method based on knowledge graph
Patent Information
- Application Number
- CN202610291238.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-11
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-03-11
AI Technical Summary
[0008]本发明的一个目的在于提出一种基于知识图谱的巢湖新污染物复合预警方法,针对现有技术中复合污染事件早期识别能力不足、预警判定对点值预测和固定阈值依赖导致误报难控、以及缺乏可解释溯源证据链输出的问题,提出了如下技术方案:获取巢湖流域多个监测断面的新污染物监测数据、水文气象数据及排放源数据并标准化处理;基于标准化监测记录与标准化排放源记录构建包含监测断面、水系连通、排放源、污染物转化关系的知识图谱,并以监测断面为节点、水系连通关系为边生成用于预测的时空图;将时空图输入时空图Transformer预测模型对至少两种污染物浓度进行联合预测并输出不确定性量化信息;按季节类别、断面类型和水动力情景类别中的至少一种进行分层共形预测校准得到分层置信集合;依据分层置信集合与阈值关系计算超阈置信度并触发规则,执行知识图谱多跳推理实现复合污染事件识别与预警分级;提取覆盖推理路径的最小证据子图并计算贡献度与规则满足度生成证据包
[0045]1、通过时空图Transformer对至少两种新污染物浓度进行联合预测并输出不确定性量化信息,再按季节类别、断面类型和水动力情景类别中的至少一种实施分层共形预测校准,得到具有覆盖保证的分层置信集合,从而在分布随时空情景变化时仍能提高预测区间的可靠性与预警判定的一致性。
Smart Images

Figure CN122198125B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of water environment monitoring and pollution early warning, and in particular to a composite early warning method for new pollutants in Chaohu Lake based on knowledge graphs. Background Technology
[0002] With the continued presence of industrial emissions, urban domestic emissions, and agricultural non-point source pollution in the Chaohu Lake basin, the types and sources of new pollutants in the water body are becoming increasingly complex. Existing water environment monitoring systems typically deploy online or offline monitoring methods at multiple monitoring sections, combining hydrological and meteorological data to analyze water quality changes, and are gradually evolving from single-indicator exceedance judgments to comprehensive multi-indicator evaluations. In recent years, water quality prediction methods based on machine learning and deep learning have been applied, introducing time-series models, graph models, and spatiotemporal fusion models to achieve short-term predictions of water quality indicators at different sections. Simultaneously, semantic modeling technologies such as knowledge graphs are beginning to be used to integrate multi-source knowledge of pollutants, emission sources, and water system connectivity to support pollution source tracing and correlation analysis.
[0003] However, existing technologies still have the following shortcomings in the early identification and graded warning of "complex pollution events caused by new pollutants":
[0004] 1. Early warning triggers rely on point value predictions or fixed threshold over-limit judgments, making it difficult to provide early warnings of complex risks involving multiple pollutants and multiple cross-sections in the weak signal stage, and false alarms are difficult to control.
[0005] 2. Although existing prediction models can output prediction results, they are insufficient in quantifying and calibrating uncertainties. In particular, under the conditions of distribution drift such as seasonal changes, cross-sectional differences and changes in hydrodynamic scenarios, the reliability and consistency of the early warning threshold determination are poor.
[0006] 3. The source tracing and explanation are mostly based on empirical rules or post-event analysis, lacking an auditable reasoning path and evidence chain output mechanism that is consistent with the early warning triggering mechanism, making it difficult to support explainable graded early warning and regulatory decision-making for compound pollution events.
[0007] Therefore, a composite early warning method for new pollutants in Chaohu Lake that can overcome the shortcomings of the existing technology is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0008] One objective of this invention is to propose a knowledge graph-based early warning method for new pollutants in Chaohu Lake. Addressing the shortcomings of existing technologies, such as insufficient early identification capability for complex pollution events, reliance on point value prediction and fixed thresholds leading to uncontrollable false alarms, and lack of interpretable source-tracing evidence chains, the invention proposes the following technical solution: acquiring and standardizing monitoring data, hydrological and meteorological data, and emission source data from multiple monitoring sections in the Chaohu Lake basin; constructing a knowledge graph based on standardized monitoring records and standardized emission source records, encompassing the relationships between monitoring sections, water system connectivity, emission sources, and pollutant transformation; and using monitoring sections as the basis for this knowledge graph. This invention generates a spatiotemporal graph for prediction, using nodes and water system connectivity as edges. The spatiotemporal graph is then input into a spatiotemporal graph Transformer prediction model to jointly predict the concentrations of at least two pollutants and output uncertainty quantification information. Layered conformal prediction calibration is performed according to at least one of seasonal category, cross-section type, and hydrodynamic scenario category to obtain a layered confidence set. Based on the relationship between the layered confidence set and a threshold, over-threshold confidence is calculated and rules are triggered to execute multi-hop reasoning using a knowledge graph to achieve the identification and early warning classification of complex pollution events. The minimum evidence subgraph covering the reasoning path is extracted, and its contribution and rule satisfaction are calculated to generate an evidence package. This invention possesses the technical effects of achieving controllable false alarms in the weak signal stage, improving the reliability of early warnings under conditions of distribution change, and outputting an auditable evidence chain to support source tracing and interpretation.
[0009] This invention provides a knowledge graph-based method for comprehensive early warning of emerging pollutants in Chaohu Lake, comprising:
[0010] S1. Acquire new pollutant monitoring data, hydrological and meteorological data, and emission source data from multiple monitoring sections in the Chaohu Lake basin and standardize them to obtain standardized monitoring records and standardized emission source records. S2. Construct a knowledge graph for composite early warning based on the standardized monitoring records and standardized emission source records. Using monitoring sections as nodes and water system connectivity as edges in the knowledge graph, and combining pollutant concentrations and hydrological and meteorological characteristics from the standardized monitoring records, construct a spatiotemporal map for prediction. S3. Input the spatiotemporal map into the spatiotemporal map Transformer prediction model to jointly predict the concentrations of at least two pollutants within the target prediction time window, obtaining the prediction results and their corresponding confidence information parameters, forming an uncertainty quantification result. S4. Based on the prediction results, uncertainty quantification results, and corresponding historical actual concentrations, determine the context category according to at least one of seasonal category, section type, and hydrodynamic scenario category. Maintain calibration samples for each context category and update the residual distribution based on the inconsistency metric of the residuals. Perform hierarchical conformal prediction calibration on the prediction results. S5. Based on the hierarchical confidence set, combined with the preset threshold set and preset rule base, the knowledge graph rule reasoning is graded: the rule triggering is determined according to the inclusion relationship between the hierarchical confidence set and the preset threshold set, the comparison relationship between the boundary of the hierarchical confidence set and the preset threshold set, and / or the over-threshold confidence calculated by the hierarchical confidence set. When the triggering is established, multi-hop association reasoning is performed in the knowledge graph according to the entity type, relation type and maximum number of hops defined by the triggering rule to obtain the reasoning path set. Based on this, the composite pollution event and its warning level are determined, and the triggering rule identifier is generated. S6. Based on the composite pollution event, the warning level, the reasoning path set and the triggering rule identifier, the candidate subgraph covering the reasoning path set is extracted from the knowledge graph and redundant relations are deleted to obtain the minimum evidence subgraph. The contribution of nodes and edges in the minimum evidence subgraph and the rule satisfaction of the triggering rule are calculated. An evidence package containing the warning level, composite pollution event, triggering rule identifier, minimum evidence subgraph, contribution and rule satisfaction is generated and the warning result is output.
[0011] Optionally, S1 includes:
[0012] Acquire the monitoring data, hydrological and meteorological data, and emission source data of the new pollutants;
[0013] The new pollutant monitoring data and hydro-meteorological data are subjected to duplicate record removal, outlier identification and processing, missing value marking and processing, and unit of measurement standardization, so that the pollutant concentration is expressed in a preset unified unit of measurement.
[0014] Encoding mapping is performed on the monitoring section names in the new pollutant monitoring data to generate monitoring section identifiers, and encoding mapping is performed on the pollutant names in the new pollutant monitoring data to generate pollutant identifiers;
[0015] Time alignment is performed on the new pollutant monitoring data and hydro-meteorological data to ensure that the timestamp adopts a preset time granularity and corresponds to the monitoring section identifier under the same timestamp.
[0016] The standardized monitoring records are generated based on time-aligned pollutant concentrations and hydro-meteorological characteristics.
[0017] The emission source data is subjected to duplicate record removal, emission source name coding mapping to generate emission source identifiers, geographic location coordinate system one, and pollutant identifier alignment, so that the emission source identifiers are associated with geographic locations and emission association information corresponding to pollutant identifiers, and the standardized emission source records are generated accordingly.
[0018] Optionally, S2 includes:
[0019] Standardized monitoring records are used to generate monitoring section entities and pollutant entities. Monitoring section identifiers and corresponding section types are written to the monitoring section entities, and pollutant identifiers are written to the pollutant entities.
[0020] Emission source entities are generated using standardized emission source records, and emission source identifiers and geographical locations are written into the emission source entities;
[0021] Based on the water system data of Chaohu Lake Basin, entities of rivers flowing into the lake are generated, and the water system connectivity between rivers flowing into the lake and monitoring sections is established based on the spatial connectivity between the rivers flowing into the lake and the monitoring sections.
[0022] Based on the emission correlation information in the standardized emission source records, an emission correlation relationship between emission sources and pollutants is established, and an indication relationship between pollutants and emission sources is established.
[0023] Establish the transformation relationships between pollutants based on pre-defined knowledge of new pollutant transformation;
[0024] The aforementioned entities and relationships are written into the knowledge graph to complete the construction of the knowledge graph;
[0025] After the knowledge graph is constructed, the monitoring section entities in the knowledge graph are used as nodes, and the water system connectivity is used as edge relationships. The node features are generated by combining the pollutant concentrations and hydrological and meteorological characteristics corresponding to each monitoring section identifier in the standardized monitoring records, and a spatiotemporal map for prediction is generated.
[0026] Optionally, S3 includes:
[0027] In the spatiotemporal diagram, node feature sequences of each monitoring section within the historical time window are extracted according to a preset historical time window. The node feature sequences include at least two pollutant concentration sequences corresponding to the monitoring section identifier and hydro-meteorological feature sequences.
[0028] The temporal and spatial attention calculations are performed on the node feature sequence and the edge relationship of the spatiotemporal graph using the spatiotemporal graph prediction model to obtain the spatiotemporal representation of each monitoring section.
[0029] Based on the spatiotemporal representation, the predicted values of the pollutant concentrations of at least two pollutants within the target prediction time window are generated through a joint prediction output layer, forming the prediction result;
[0030] Simultaneously, confidence information parameters corresponding one-to-one with the prediction results are generated through the uncertainty output layer to form the uncertainty quantification result, wherein the confidence information parameters include the prediction variance or prediction interval boundary corresponding to the prediction value.
[0031] Optionally, S4 includes:
[0032] The corresponding seasonal category is determined based on the timestamp in the prediction results, the corresponding section type is determined based on the section type associated with the monitoring section entity in the knowledge graph, and the monitoring section corresponding to the timestamp is classified into hydrodynamic scenario category based on the hydrometeorological characteristics in the standardized monitoring records, thereby determining the context category for the prediction results.
[0033] For the context category, historical actual concentrations that match the prediction results in terms of monitoring section identifier, pollutant identifier, and timestamp are extracted from the standardized monitoring records. The non-consistency measure of the residuals is calculated, wherein the non-consistency measure of the residuals is obtained by normalizing the deviation between the historical actual concentrations and the prediction results in combination with the confidence information characterized by the uncertainty quantification results.
[0034] Write the inconsistency metric of the residual into the residual distribution corresponding to the context category, and update the residual distribution according to the preset update rule;
[0035] After the residual distribution is updated, a calibration threshold corresponding to the residual distribution is determined according to a preset significance level, and the prediction result is calibrated based on the calibration threshold to generate the hierarchical confidence set, such that the hierarchical confidence set includes the upper bound and lower bound of the confidence interval generated for the at least two pollutants respectively.
[0036] Optionally, S5 includes:
[0037] Based on the hierarchical confidence set, for each monitoring section identifier and each pollutant identifier, the upper and lower bounds of the corresponding confidence intervals are read. Then, based on the grading thresholds corresponding to the pollutant identifiers in a preset threshold set, the over-threshold confidence level of the pollutant at the monitoring section is calculated. The over-threshold confidence level is determined at least according to one of the following methods: the comparison result between the upper bound of the confidence interval and the grading threshold, the comparison result between the lower bound of the confidence interval and the grading threshold, and the inclusion relationship between the confidence interval and the grading threshold. Based on a preset rule base, rule triggering judgment is performed on the over-threshold confidence levels corresponding to at least two pollutants, ensuring that the rule triggering judgment includes at least one of the following triggering conditions: the over-threshold confidence levels of at least two pollutants under the same monitoring section identifier satisfy a preset combination condition. The system determines the following conditions: the confidence level of the same pollutant under different monitoring section identifiers meets a preset spatial aggregation condition, and the confidence level of the same pollutant under the same monitoring section identifier meets a preset persistence condition within consecutive timestamps; when a rule is triggered, the system limits the entity type and relation type in the knowledge graph according to the triggering rule, and limits the maximum number of hops for multi-hop association reasoning, generating a set of reasoning paths matching the triggering rule; a composite pollution event is generated based on the pollutant identifier set, monitoring section identifier set, and corresponding timestamp range covered by the reasoning path set, and the warning level corresponding to the composite pollution event is determined based on the corresponding hierarchical clauses in the preset rule base according to the triggering rule, combined with the confidence level of the pollutant exceeding the threshold, and a triggering rule identifier corresponding to the triggering rule is generated.
[0038] Optionally, S6 includes:
[0039] Based on the inference path set and trigger rule identifier, candidate subgraphs containing all paths in the inference path set are extracted from the knowledge graph;
[0040] Redundant relation deletion is performed on the candidate subgraph to generate a minimum evidence subgraph. This redundancy deletion includes: deleting relations not belonging to the inference path set without disrupting the connectivity of any path in the inference path set; and retaining relations associated with the trigger rule identifier and deleting other relations when multiple relations of the same type point to the same entity. A threshold confidence level is calculated based on the hierarchical confidence set, and the contribution of each node and edge in the minimum evidence subgraph is calculated based on the frequency of occurrence of each node and edge in the inference path set, the path length, and the threshold confidence level. The rule satisfaction of the trigger rule is calculated based on the rule conditions corresponding to the trigger rule identifier, combined with the upper and lower bounds of the confidence interval of the hierarchical confidence set. An evidence package is generated, comprising a warning level, a complex pollution event, a trigger rule identifier, a minimum evidence subgraph, the contribution of nodes and edges, and the rule satisfaction, used to form an interpretable chain of tracing evidence and output a warning result.
[0041] Optionally, the hierarchical confidence set, in addition to the upper and lower bounds of the confidence intervals generated for each of the at least two pollutants, also includes a joint confidence set for the at least two pollutants. The joint confidence set is generated based on the aggregated statistics of the inconsistency measures of each pollutant, so that the joint confidence set can simultaneously cover the at least two pollutants at a preset significance level.
[0042] Optionally, when limiting the maximum number of hops for multi-hop association reasoning, a propagation delay upper limit is further determined based on the hydrodynamic scenario category, and only reasoning paths that satisfy the propagation delay upper limit are retained as the set of reasoning paths.
[0043] Optionally, when generating the minimum evidence subgraph, redundant relation deletion is performed with the goal of covering the set of inference paths with the minimum cost, wherein the cost is determined by one or more of relation type weights, path lengths, and the contribution of nodes and edges.
[0044] The beneficial effects of this invention are:
[0045] 1. By using the spatiotemporal graph Transformer to jointly predict the concentrations of at least two new pollutants and output uncertainty quantification information, and then performing hierarchical conformal prediction calibration according to at least one of the seasonal category, cross section type and hydrodynamic scenario category, a hierarchical confidence set with coverage guarantee is obtained, thereby improving the reliability of the prediction interval and the consistency of the early warning judgment even when the distribution changes with spatiotemporal scenario.
[0046] 2. Based on the relationship between the hierarchical confidence set and the threshold, the over-threshold confidence level is calculated and the rule-based triggering judgment is performed. The traditional point value over-threshold triggering is upgraded to a triggering mechanism driven by confidence interval boundary comparison, inclusion relationship and over-threshold confidence level. This enables the composite risk alarm to be triggered earlier in the weak signal stage, while realizing a hierarchical early warning with controllable false alarms.
[0047] 3. After the rule is triggered, multi-hop association reasoning with the knowledge graph is performed with restrictions on entity type, relation type and maximum number of hops to obtain a set of reasoning paths. The minimum evidence subgraph covering the reasoning path is further extracted, and the node edge contribution degree and rule satisfaction degree are calculated to form an evidence package, thereby outputting an auditable and traceable evidence chain, enhancing the interpretability and traceability support of the early warning results. Attached Figure Description
[0048] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0049] Figure 1This is a flowchart of a knowledge graph-based composite early warning method for new pollutants in Chaohu Lake proposed in this invention. Detailed Implementation
[0050] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0051] refer to Figure 1 A knowledge graph-based composite early warning method for emerging pollutants in Chaohu Lake includes:
[0052] S1. Acquire new pollutant monitoring data, hydrological and meteorological data, and emission source data from multiple monitoring sections in the Chaohu Lake basin and standardize them to obtain standardized monitoring records and standardized emission source records. S2. Construct a knowledge graph for composite early warning based on the standardized monitoring records and standardized emission source records. Using monitoring sections as nodes and water system connectivity as edges in the knowledge graph, and combining pollutant concentrations and hydrological and meteorological characteristics from the standardized monitoring records, construct a spatiotemporal map for prediction. S3. Input the spatiotemporal map into the spatiotemporal map Transformer prediction model to jointly predict the concentrations of at least two pollutants within the target prediction time window, obtaining the prediction results and their corresponding confidence information parameters, forming an uncertainty quantification result. S4. Based on the prediction results, uncertainty quantification results, and corresponding historical actual concentrations, determine the context category according to at least one of seasonal category, section type, and hydrodynamic scenario category. Maintain calibration samples for each context category and update the residual distribution based on the inconsistency metric of the residuals. Perform hierarchical conformal prediction calibration on the prediction results. S5. Based on the hierarchical confidence set, combined with the preset threshold set and preset rule base, the knowledge graph rule reasoning is graded: the rule triggering is determined according to the inclusion relationship between the hierarchical confidence set and the preset threshold set, the comparison relationship between the boundary of the hierarchical confidence set and the preset threshold set, and / or the over-threshold confidence calculated by the hierarchical confidence set. When the triggering is established, multi-hop association reasoning is performed in the knowledge graph according to the entity type, relation type and maximum number of hops defined by the triggering rule to obtain the reasoning path set. Based on this, the composite pollution event and its warning level are determined, and the triggering rule identifier is generated. S6. Based on the composite pollution event, the warning level, the reasoning path set and the triggering rule identifier, the candidate subgraph covering the reasoning path set is extracted from the knowledge graph and redundant relations are deleted to obtain the minimum evidence subgraph. The contribution of nodes and edges in the minimum evidence subgraph and the rule satisfaction of the triggering rule are calculated. An evidence package containing the warning level, composite pollution event, triggering rule identifier, minimum evidence subgraph, contribution and rule satisfaction is generated and the warning result is output.
[0053] In this specific embodiment, S1 includes:
[0054] According to a unified data dictionary, new pollutant monitoring data, hydrological and meteorological data, and emission source data from multiple monitoring sections in the Chaohu Lake Basin are acquired and standardized to generate standardized monitoring records and standardized emission source records.
[0055] The new pollutant monitoring data includes a set of fields {monitoring section name, monitoring time, pollutant name, pollutant concentration, concentration measurement unit}, and a monitoring record is located by the monitoring section name and monitoring time. The hydrological and meteorological data includes a set of fields {monitoring time, monitoring section name or nearest spatial matching point name, water temperature, flow rate, water level, rainfall, wind speed, wind direction, air temperature}, and a hydrological and meteorological record is located by the monitoring time. The emission source data includes a set of fields {emission source name, unified social credit code, geographical location coordinates, latitude and longitude coordinate system, pollutant name, emission association information}, and an emission record is located by the emission source name and pollutant name.
[0056] Duplicate records are removed from new pollutant monitoring data and hydrological and meteorological data. The criteria for determining duplicate records are that the monitoring section name, monitoring time, and pollutant name are exactly the same, or the monitoring section name and monitoring time are exactly the same and the other fields are exactly the same. When removing duplicate records, the records with the highest data source credibility and the sampling method is identified as online monitoring are retained. If the credibility and sampling method are the same, the record with the earliest sampling completion time is retained.
[0057] Outlier identification and processing are performed on the new pollutant monitoring data. For each combination of monitoring section name and pollutant name, a concentration sequence in ascending order of time is constructed. A sliding window of length of 24 hours is used to calculate the median and median absolute deviation within the window. Concentration values that deviate from the threshold of 3.5 times the median absolute deviation are identified as outliers. Outliers are handled by setting them as missing and writing them into the outlier flag field.
[0058] Two levels of rules are set for missing value marking and handling. The first level rule is to use forward filling and record the filling mark for pollutant concentrations with a continuous missing duration of no more than 3 time granularities. The second level rule is to keep the missing pollutant concentrations with a continuous missing duration of more than 3 time granularities missing and record the missing mark. The missing state is retained when generating standardized monitoring records for subsequent model input masks.
[0059] The unit of measurement is set to a pre-defined unified unit of measurement. Based on a unit conversion table, unit conversion is performed on each concentration record to obtain a unified pollutant concentration. The unit conversion uses the following formula:
[0060] ;
[0061] in This indicates that the timestamp after time alignment is And the monitoring section is marked as Pollutant labeling The unified unit of measurement for pollutant concentration, Indicates the original monitoring time is And in accordance with the monitoring section markings Pollutant labeling The corresponding original concentration value, Indicates the original unit of measurement arrive Conversion factor and when hour ,when hour ,when hour Indicates the monitoring section marker. Indicates pollutant labeling, This represents the timestamp after alignment with the preset time granularity. Indicates the original monitoring timestamp. Indicates the original unit of measurement;
[0062] Encoding mapping is performed on the monitoring section names to generate monitoring section identifiers. The encoding mapping is based on the list of monitoring sections in the Chaohu Lake Basin to establish a section mapping table. Each monitoring section name is standardized into a unique key of "administrative division code + water body name + section name" and assigned an integer monitoring section identifier starting from 1 in lexicographical order. This monitoring section identifier is used to replace the section name throughout the subsequent processing.
[0063] The pollutant names are encoded and mapped to generate pollutant identifiers. The pollutant encoding mapping is based on the new pollutant list to establish a pollutant mapping table. Each pollutant name is normalized into a unique key of "Chinese common name + CAS number" and assigned an integer pollutant identifier starting from 1 in lexicographical order. This pollutant identifier is used to replace the pollutant name throughout the subsequent processing.
[0064] The monitoring data for new pollutants was aligned with the hydrological and meteorological data in terms of execution time, and a preset time granularity of 1 hour was set. The original timestamps were then used. Merged into the hourly timestamp of its respective hour. For cases where multiple concentration records appear for the same monitoring section and the same pollutant within the same hour, the median is taken as the pollutant concentration for that hour, and the aggregation method is retained. For cases where multiple records exist in the hydrological and meteorological fields within the same hour, the median is taken and the aggregation method is retained.
[0065] Standardized monitoring records are generated based on time-aligned pollutant concentrations and hydrological and meteorological characteristics. The set of fields in the standardized monitoring records is fixed as timestamps. Monitoring section markings Pollutant labeling pollutant concentration Water temperature, flow rate, water level, rainfall, wind speed, wind direction, air temperature, anomaly marker, missing marker, fill marker, aggregation method marker, and each standardized monitoring record corresponds to only one monitoring section identifier, one pollutant identifier and one timestamp;
[0066] Duplicate records are removed from emission source data. The criteria for determining duplicate records are that the unified social credit code, pollutant name, and emission outlet number in the emission-related information are exactly the same. When removing duplicate records, the record with the latest declaration time is retained.
[0067] The emission source name is encoded and mapped to generate an emission source identifier. The emission source identifier uses the unified social credit code as the primary key. For emission sources that lack a unified social credit code, a unique key of "emission source name + geographical location coordinates" is used to replace the primary key. The emission source identifier is assigned an integer starting from 1 in lexicographical order and is used to replace the emission source name throughout the subsequent processing.
[0068] For the geographic coordinate system, set it to the WGS-84 coordinate system. Convert the latitude and longitude of the non-WGS-84 coordinate system using the deterministic transformation parameters from the corresponding coordinate system to WGS-84, and retain the original coordinate system mark and the transformation mark.
[0069] The pollutant names in the emission source data are aligned with the same pollutant mapping table as those in the monitoring data to obtain pollutant identification, so that the emission source records and standardized monitoring records are consistent at the pollutant identification level.
[0070] The emission-related information is structured into a set of fields: {emission outlet number, emission method, permitted emission concentration limit, permitted emission quantity limit, declared emission quantity, declared emission concentration}. These fields are then associated with the emission source identifier, geographical location, and pollutant identifier to generate standardized emission source records. The standardized emission source record has a fixed set of fields: {emission source identifier, geographical location longitude, geographical location latitude, pollutant identifier, emission outlet number, emission method, permitted emission concentration limit, permitted emission quantity limit, declared emission quantity, declared emission concentration, coordinate system marker, transformation marker}. These fields serve as the input for constructing the knowledge graph in step S2, which establishes the emission source entity and the emission-related relationship.
[0071] In this specific embodiment, S2 includes:
[0072] A knowledge graph for composite early warning is constructed based on standardized monitoring records and standardized emission source records, and a spatiotemporal graph for prediction is generated.
[0073] The knowledge graph is implemented using an attribute graph data model and stored in a graph database. The attribute graph data model represents entities as labeled nodes and relationships as directed edges with types. In order to ensure the uniqueness of entities, unique constraints are set for monitoring section identifiers, pollutant identifiers and emission source identifiers in the graph database.
[0074] Based on standardized monitoring records, monitoring section entities and pollutant entities are generated. The node label of the monitoring section entity is set to "monitoring section" and written into the attribute set {monitoring section identifier, section type} with the monitoring section identifier as the primary key. The node label of the pollutant entity is set to "pollutant" and written into the attribute set {pollutant identifier} with the pollutant identifier as the primary key.
[0075] Emission source entities are generated based on standardized emission source records. The node label of the emission source entity is set as "emission source" and written into the attribute set {emission source identifier, geographical location longitude, geographical location latitude} with the emission source identifier as the primary key.
[0076] The entity of the river flowing into the lake is generated based on the water system data of the Chaohu Basin. The water system data of the Chaohu Basin is a vector dataset containing the geometry and topology of the river segments. Each river segment in the dataset has a unique river segment number and a definite flow direction attribute. The node label of the entity of the river flowing into the lake is set as "river flowing into the lake" and written into the attribute set {river number, river name} with the river number as the primary key.
[0077] Based on the spatial connectivity between the rivers flowing into the lake and the monitoring sections, the water system connectivity relationship between the rivers flowing into the lake and the monitoring sections is established. The process of determining the spatial connectivity is to project the section location coordinates associated with each monitoring section entity onto the river segment geometry of the Chaohu Basin water system data and select the river segment with the smallest projection distance and not exceeding 200m as the river segment to which the monitoring section belongs. Then, write the "water system connectivity relationship" edge from the river entity flowing into the lake to the monitoring section entity in the knowledge graph and write the projection distance and the river segment number to which it belongs in the edge attribute.
[0078] To obtain the inter-section connectivity edges for prediction, after completing the association between the rivers flowing into the lake and the monitoring sections, the monitoring sections under the same river flowing into the lake are sorted from smallest to largest according to the cumulative river length from their respective river segments to the lake mouth along the flow direction attribute of the water system data. For two adjacent monitoring sections, a "water system connectivity relationship" edge pointing from the upstream monitoring section to the downstream monitoring section is written. At the same time, the cumulative river length between the two sections and the sequence of river segments from upstream to downstream are written into the attributes of this edge for subsequent spatial propagation constraints.
[0079] Based on the emission association information in the standardized emission source records, emission association relationships between emission sources and pollutants are established, and indication relationships between pollutants and emission sources are also established. The emission association relationship is represented in the knowledge graph as an "emission association relationship" edge pointing from the emission source entity to the pollutant entity and written into the attribute set {emission outlet number, emission method, permitted emission concentration limit, permitted emission amount limit, declared emission amount, declared emission concentration} to solidify the emission association information. The indication relationship is represented in the knowledge graph as an "indication relationship" edge pointing from the pollutant entity to the emission source entity and written into the attribute set {indication intensity}. The indication intensity is obtained by normalizing the declared emission amount of the emission source for the pollutant in the standardized emission source records according to all emission sources and is used for path sorting in subsequent multi-hop reasoning.
[0080] Based on the pre-set new pollutant transformation knowledge, the transformation relationship between pollutants is established. The new pollutant transformation knowledge is a structured table containing a set of fields {precursor pollutant identifier, product pollutant identifier, transformation type, half-life}. In the knowledge graph, it is represented as a "transformation relationship" edge from the precursor pollutant entity to the product pollutant entity and written into the attribute set {transformation type, half-life} to support the association of the complex pollution chain during subsequent reasoning.
[0081] After writing all the aforementioned entities and relationships into the knowledge graph, a spatiotemporal graph for prediction is generated. This spatiotemporal graph uses the monitoring section entity set in the knowledge graph as the node set and the "water system connectivity" edges between sections as the edge set. Furthermore, based on the time alignment result of step S1, a node feature matrix corresponding one-to-one with the monitoring section is constructed for each timestamp, thus forming a dynamic graph sequence. This dynamic graph sequence is expressed by the formula: ;
[0082] in This represents the spatiotemporal graph sequence used for prediction. This represents a set of nodes consisting of all monitored cross-section entities, with each node uniquely identified by the monitoring cross-section identifier. This represents a set of directed edges formed by the connectivity between river sections, with each edge carrying the cumulative river length and river segment sequence attributes. This represents the complete set of timestamps determined in step S1, with a time granularity of [missing information]. express Any timestamp in, Represents timestamp The corresponding node feature matrix and by node set The node characteristics of each monitoring section are arranged in a fixed order;
[0083] Node features are obtained by aggregating standardized monitoring records along the dimensions of monitoring section identifier and timestamp, and then concatenating them using a fixed field order. The fixed field order includes the pollutant concentrations corresponding to the two target pollutant identifiers, the anomaly markers, missing markers, and filler markers for each of the two pollutants, as well as a set of hydrological and meteorological feature fields. Water temperature, flow rate, water level, rainfall, wind speed, wind direction, air temperature Two target pollutant identifiers are fixed as configuration parameters into a set during system initialization. Furthermore, it remains unchanged throughout the entire method execution to ensure consistency with the joint prediction task, and when any field is at the timestamp When a value is missing, its value is set to 0 and its missing marker is retained as 1 to ensure that subsequent models can distinguish between true zero values and missing filler values and avoid matrix dimension inconsistencies caused by missing values.
[0084] In this specific embodiment, S3 includes:
[0085] Spatiotemporal graph sequence As input, a spatiotemporal graph Transformer prediction model is constructed to jointly predict the concentrations of at least two pollutants within the target prediction time window and output the uncertainty quantification results.
[0086] The system sets the current prediction start time stamp to 1. Furthermore, the time granularity is 1 hour, and the preset historical time window length is set to... Set the target prediction time window length to Thus, each monitoring section is identified as The monitoring section in the historical time window Internal node feature sequence extraction And within the target prediction time window Internal output of two target pollutant identifier sets Corresponding concentration prediction values and confidence information parameters;
[0087] in The timestamp in step S2 is The node feature vector contains pollutant identifiers in a fixed field order. and The pollutant concentration and its anomaly markers, missing markers and filling markers, and include a set of hydro-meteorological feature fields {water temperature, flow rate, water level, rainfall, wind speed, wind direction, air temperature}, and when any field is missing, its value is set to 0 and the missing information is retained as 1 for use in the model attention mask;
[0088] The spatiotemporal graph Transformer prediction model is denoted as It consists of an input embedding layer, a temporal attention encoder, a spatial attention encoder, a joint prediction output layer, and an uncertainty output layer, and is composed of a parameter set. Determine that the input embedding layer will each Projected to dimension via linear mapping The hidden space and superimposed length is Learnable temporal location embeddings for encoding Relative position within the history window;
[0089] The temporal attention encoder identifies each monitoring section. The length is The embedded sequence is subjected to multi-head self-attention to obtain the temporal representation of the cross section within the historical window. The number of layers of the temporal attention encoder is set to 2 and the number of attention heads per layer is set to 4. In the attention weight calculation, the time step with missing flag 1 is masked to avoid missing values from interfering with attention aggregation.
[0090] The spatial attention encoder performs attention aggregation on the spatial dependencies between monitoring sections at each time step. Its spatial adjacency constraint is formed by the directed adjacency matrix of the water system connectivity relationships between monitoring sections in step S2. Determine and only allow The upstream monitoring section with connectivity contributes attention to the downstream monitoring section. At the same time, the cumulative river length of the water system connectivity edge is written into the distance matrix D, and a learnable distance embedding vector is used as the attention bias to reflect the prior that the propagation intensity decreases with the river length. The spatial attention encoder is set to 2 layers and 4 attention heads per layer, and residual connection and layer normalization are used to stabilize the training.
[0091] After completing the temporal and spatial coding, each monitoring section is identified. The corresponding spatiotemporal representation is obtained and input into the joint prediction output layer and the uncertainty output layer, where the joint prediction output layer is a two-layer feedforward network with a fixed output dimension. Simultaneously output pollutant labels as and In the future The concentration prediction values at each time step, with uncertainty, are output by a two-layer feedforward network with a fixed output dimension. The prediction variance is output one-to-one with each predicted value, and softplus activation is used to ensure that the prediction variance is positive.
[0092] The formula for joint forecasting and uncertainty output is as follows:
[0093] ;
[0094] in The monitoring section is marked as And the pollutants are labeled as In timestamp The predicted values of pollutant concentrations Indicates and The one-to-one correspondence of prediction variance is used as the confidence information parameter. This indicates the first of the two target pollutant labels. individual and Indicates the prediction step size and Indicates the current prediction start time stamp. Indicates the length of the historical time window and takes the value of This represents the timestamp index variable within the history window. The monitoring section is marked as timestamp The node feature vectors, This represents a directed adjacency matrix formed by the water system connectivity relationships between monitoring sections. This represents the cumulative river length distance matrix consistent with A. This represents the spatiotemporal graph Transformer prediction model. This represents all trainable parameters of the model;
[0095] The spatiotemporal graph Transformer prediction model is trained offline before going live in step S3 and fixed after going live. For rolling prediction, offline training employs a Gaussian negative log-likelihood loss with historical actual concentrations as the supervision signal, and simultaneously fits the data. and The training optimizer uses AdamW and the learning rate is set to 1. Weight decay is set to The batch size is set to 32, the number of training rounds is set to 50, and the training set, validation set and test set are divided in chronological order. The early stopping strategy of the validation set is used to determine the final model parameters, and the prediction results and uncertainty quantification results are output.
[0096] In this specific embodiment, S4 includes:
[0097] Based on the prediction results and uncertainty quantification results, hierarchical conformal prediction calibration is performed to generate a hierarchical confidence set and a joint confidence set at the same time.
[0098] The system sets the significance level to 1. Furthermore, independent calibration samples and residual distributions are maintained for each context category to ensure stable coverage characteristics even when the distribution changes with the context.
[0099] Context category is denoted as Furthermore, it is jointly determined by the seasonal category, cross-section type, and hydrodynamic scenario category, with the seasonal category determined by the timestamp. The corresponding months are determined and fixedly divided into spring (March to May), summer (June to August), autumn (September to November), and winter (December to February of the following year). The cross-section types are read from the cross-section type attributes of the monitored cross-section entities in the knowledge graph and fixed in the system as three categories: inflow river cross-sections, lake area cross-sections, and outflow river cross-sections. The hydrodynamic scenario categories are determined and fixedly divided into six categories based on the flow rate and wind speed in the standardized monitoring records, and are obtained by combining flow rate level and wind speed level. The flow rate level is indicated by the monitoring cross-section identifier. And the timestamp is Traffic Enter and press , The wind speed is divided into three levels: low flow, medium flow, and high flow, with the wind speed level determined by a timestamp. wind speed Enter and press and Divided into two levels: weak wind and strong wind, thus making The triplet label uniquely identifies the context category to which the prediction belongs;
[0100] For each context category Maintain separate sets of labels for the two target pollutants. Corresponding residual distribution and And maintain the joint residual distribution Each residual distribution uses a first-in-first-out queue structure to store the most recent... Inconsistency measures of calibration samples and when the queue length exceeds The earliest sample added to the queue is deleted to achieve rolling updates;
[0101] When step S3 predicts the starting point timestamp as Rolling prediction output future Prediction results at each time step With the corresponding confidence information parameters Subsequently, when the actual monitored concentration is reached, the system reads the data from the standardized monitoring records and identifies the monitoring section as... Pollutant labeling The timestamp is Consistently matched historical actual concentrations And calculate the inconsistency measure based on the prediction bias and uncertainty quantification results. And write the residual distribution of the corresponding context category, where:
[0102] ;
[0103] In the formula The monitoring section is marked as And the pollutants are labeled as timestamp Inconsistent metrics, This indicates the historical actual concentration matched with that timestamp. This represents the predicted concentration value. The prediction variance, which corresponds one-to-one with the predicted value, is used as the confidence information parameter. This represents a stable term to prevent the denominator from being zero, and its value is... Indicates the monitoring section marker. This indicates the first of the two target pollutant labels. individual and Indicates the current prediction start time stamp. Indicates the prediction step size and ;
[0104] After the residual distribution is updated, the system updates each context category. With each pollutant label From respectively Calculate calibration threshold The calculation method is to sort all inconsistent metric values in the queue in ascending order and take the first value in the sorted sequence. each element as ,in Indicates the current number of samples in the queue and when The context category is then degraded to a context category determined solely by the season category and section type, and recalculated until the condition is met. ;
[0105] The system then processes each predicted value. Calculate calibration radius It also outputs the lower bound of the confidence interval for the pollutant in the hierarchical confidence set. and the upper bound of the confidence interval ,in Indicates by timestamp Determined context category, This indicates the calibration threshold corresponding to this context category and this contaminant. Indicates the calibration radius. and These represent the lower and upper bounds of the confidence interval, respectively.
[0106] In order to obtain a joint confidence set that simultaneously covers both pollutants, the system uses the same context category. The joint inconsistency metric at each time step is defined as follows: And write And ranked according to the same sorting rules as for single pollutants. Calculate the joint calibration threshold Subsequently, for the same predicted time The joint calibration radius is calculated by replacing the single-pollutant calibration threshold with the joint calibration threshold. And output the joint confidence set accordingly. Let be a two-dimensional set consisting of the Cartesian product of the joint calibration intervals of the two pollutants, such that the hierarchical confidence set simultaneously includes the pollutant identified as... and The upper and lower bounds of the confidence interval and at the significance level The following joint confidence set that satisfies the simultaneous coverage of two pollutants is used as the input for step S5 to calculate the overthreshold confidence and triggering rule.
[0107] In this specific embodiment, S5 includes:
[0108] Based on a hierarchical confidence set, a preset threshold set, and a preset rule base, knowledge graph rule reasoning is performed to classify and determine complex pollution events and their warning levels, and to generate trigger rule identifiers.
[0109] The system at each prediction timestamp Marking each monitoring section With each target pollutant label Read the lower bound of the corresponding confidence interval from the hierarchical confidence set. and the upper bound of the confidence interval And read the pollutant identifier from the preset threshold set. Corresponding set of hierarchical thresholds And satisfy Then, for each grade threshold Calculate the overthreshold confidence level And Write the feature table used for rule triggering determination, where:
[0110] ;
[0111] In the formula Indicates the monitoring section marker. Represents two sets of target pollutant identifiers The first in Each pollutant label and This represents the prediction timestamp within the target prediction time window. Indicates the hierarchical threshold level index and The pollutant is labeled as The monitoring section is marked as And the timestamp is The lower bound of the confidence interval, This indicates the upper bound of the corresponding confidence interval. The pollutant is labeled as The Level classification threshold, Relative to the threshold The overthreshold confidence level and the range of values is ;
[0112] The system solidifies the preset rule base into a structured rule set and assigns a unique trigger rule identifier to each rule. Each rule includes pollutant combination conditions, spatial aggregation conditions, persistence conditions, warning level mapping clauses, entity type restrictions, relationship type restrictions, and maximum hop count restrictions. Among these, the pollutant combination conditions are solidified into the same monitoring section identifier. And the same timestamp The following conditions must be met simultaneously and Spatial aggregation conditions are solidified as monitoring section identifiers under the constraints of water system connectivity relationships in the knowledge graph. Within the upstream two-step range of the center, there are at least three different monitoring section markers that simultaneously meet the same pollutant marker requirements. of The continuous conditions are fixed as the same monitoring section identifier. Same pollutant label It satisfies the condition within 3 consecutive timestamps. Furthermore, the rule triggering judgment is executed sequentially by time stamp within the target prediction time window, and when any triggering condition of any rule is met, the triggering time stamp range, the triggering section identifier set, and the triggering pollutant identifier set are recorded, and the corresponding triggering rule identifier is generated.
[0113] Once the trigger is established, the system reads the entity type based on the trigger rule identifier, which is limited to [specific type]. Monitoring sections, pollutants, emission sources, and rivers flowing into the lake. And read its relation type is limited to Water system connectivity, discharge correlation, indication relationship, and transformation relationship And read its maximum number of jumps limit Subsequently, starting with each monitoring section entity in the trigger section identifier set and each pollutant entity in the trigger pollutant identifier set, multi-hop association reasoning is performed in the knowledge graph, and a breadth-first search is used to generate candidate reasoning paths. When expanding the path, it is only allowed to pass through the above-mentioned entity type restriction and relation type restriction, and the path length does not exceed the limit. Jump and remove loops from duplicate entities to ensure that the path is finite;
[0114] During multi-hop association inference, the system reads the hydrodynamic scenario category and maps it to the upper limit of propagation delay. Furthermore, this mapping is fixed in the system as a table defining the time delay upper limit for six types of hydrodynamic scenarios, with low flow and weak wind corresponding to... Low flow and strong wind response Medium flow and weak wind correspondence Medium-flow strong winds correspond High flow and weak wind correspondence High flow and strong winds correspond For each candidate inference path, the sum of its propagation delay along the water system connectivity edges is calculated as the path propagation delay, and only paths with a propagation delay not exceeding a certain threshold are retained. The candidate inference paths are selected to form an inference path set, so that the inference path set simultaneously satisfies the maximum hop count limit and the propagation delay limit;
[0115] The system generates composite pollution events based on the pollutant identifier set, monitoring section identifier set, and trigger timestamp range covered by the inference path set, and writes these composite pollution events into an event table for subsequent evidence package generation. Furthermore, it determines the warning level based on the warning level mapping clause corresponding to the trigger rule identifier, combined with the over-threshold confidence level at the trigger time. The warning level mapping clause is fixed so that when any pollutant identifier satisfies the following condition at any trigger monitoring section identifier... When it is determined to be a Level 1 warning, when any pollutant label meets the requirements. Furthermore, a Level II alert is established when the combined conditions of the two pollutants are met, and a Level II alert is established when the persistent condition is met and any pollutant identifier satisfies the alert. The alert level is determined to be a Level 3 warning, and the warning level and trigger rule identifier are output together.
[0116] In this specific embodiment, S6 includes:
[0117] Based on the complex pollution events, warning levels, inference path sets, and trigger rule identifiers, an evidence package for auditing is generated in the knowledge graph, and warning results are output.
[0118] The system records the knowledge graph as Furthermore, the set of entity types and the set of relation types limited by the trigger rule identifier are used as filtering conditions. Candidate subgraphs were obtained by performing subgraph extraction. The node set of the candidate subgraph consists of all entity nodes appearing in the inference path set and the first-order attribute nodes of these entity nodes within the same trigger timestamp range. The edge set of the candidate subgraph consists of all relation edges appearing in the inference path set and retains the relation type, direction and edge attributes written in step S2 for each edge.
[0119] The system Redundant relation deletion is performed to remove relation edges that are unrelated to the inference path set. Redundant relation deletion satisfies two constraints. The first constraint is to delete relation edges that do not belong to the inference path set without destroying the connectivity of any path in the inference path set, and simultaneously delete attribute nodes that become isolated points. The second constraint is that when there are multiple relation edges of the same type pointing to the same entity node, only the relation edge associated with the trigger rule identifier is retained and the rest of the relation edges are deleted. Here, "associated with the trigger rule identifier" is fixed as the generation rule identifier recorded in the metadata of the relation edge is equal to the trigger rule identifier.
[0120] After completing the initial deletion based on path consistency, the system further generates a minimum evidence subgraph with the goal of covering the set of inference paths with the minimum cost. The set of endpoint nodes that need to be covered is denoted as . Furthermore, the system comprises all monitoring section entities, pollutant entities, and emission source entities in the inference path set, and is therefore... Each relation edge in The cost is calculated and used as weights to perform a deterministic Steiner approximation solution process. The cost calculation and minimization objective are expressed by the following formula:
[0121] ;
[0122] in Represents the minimum evidence subgraph The total cost, express The set of relation edges, express One of the relation edges in the middle, Representing relation edges The relation type and taken from Water system connectivity, discharge correlation, indication relationship, and transformation relationship This represents the relation type weight corresponding to the relation type, and is fixed in the system as follows: 1.0 for water system connectivity relation, 1.5 for indication relation, 2.0 for emission association relation, and 0.0 for transformation relation. Representing relation edges Length factor and when When defining the water system connectivity, the value is the cumulative river length in kilometers written in step S2, with a lower limit truncated to 1 and an upper limit truncated to 50. When the value is for emission correlation, indication, or conversion relationship, it is taken as . Representing relation edges The contribution and the range of values are This represents the contribution discount factor, which is fixed at 0.5.
[0123] Contribution With node contribution The properties of the minimum evidence subgraph are calculated using reproducible statistical rules, where for each relation edge... The frequency of occurrence of a pollutant in the inference path set is calculated and divided by the maximum frequency of occurrence of the relation edge in the inference path set to obtain a frequency score. The over-threshold confidence scores of the monitoring section markers and pollutant markers adjacent to each other at both ends of the relation edge within the trigger timestamp range are then calculated. The maximum value is taken to obtain the confidence score, and the frequency score and confidence score are weighted by fixed weights of 0.6 and 0.4 respectively. For each node Calculate the frequency of its occurrence in the inference path set and combine it with the edges of its adjacent relationships. The node contribution is obtained by weighting the mean with a fixed weight of 0.5. ;
[0124] The Steiner approximation solution process employs a two-stage deterministic construction, the first stage in... Above the cost Set of endpoint nodes for weighted counterparts In the first stage, shortest path search is performed on any two endpoints to generate a complete endpoint graph. In the second stage, a minimum spanning tree is computed on the complete endpoint graph, and each endpoint edge of the minimum spanning tree is replaced with its corresponding edge in the graph. The corresponding shortest path is used to obtain a connected subgraph. Then, edge pruning is performed on this connected subgraph to delete any node that retains the set of endpoint nodes after deletion. High-cost redundant edges that are fully connected and still cover all endpoints of the inference path set are removed until no further deletions are possible. ;
[0125] The system calculates the rule satisfaction based on the rule conditions corresponding to the trigger rule identifier and the over-threshold confidence set used for triggering in step S5, and records the rule satisfaction as . The rule for calculating rule satisfaction is fixed as follows: calculate the satisfaction strength of each atomic condition contained in the triggering rule and take the minimum value as the result. Furthermore, the strength of each atomic condition satisfaction is fixed as the overthreshold confidence level used to determine that condition. The minimum value is used to ensure that the rule satisfaction decreases synchronously when any key condition becomes weaker.
[0126] Finally, the system generates an evidence package and outputs the early warning result. The evidence package uses structured data objects and includes the following fields: early warning level, compound contamination event, trigger rule identifier, trigger timestamp range, inference path set, and minimum evidence subgraph. Node contribution Contribution of the relationship edge Rule satisfaction The minimum evidence subgraph is output in the form of a node list and an edge list. The node list contains entity type and primary key identifier and retains key attributes such as monitoring section identifier, pollutant identifier and emission source identifier. The edge list contains relation type, direction, edge attribute and cost component in the evidence package to support auditable traceability of the early warning conclusion.
[0127] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0128] This invention addresses the technical problem of "early identification and graded warning of complex pollution events, and output of an interpretable chain of evidence for source tracing." It employs a combined algorithmic chain consisting of "spatiotemporal graph prediction, uncertainty quantification, hierarchical conformal prediction calibration, rule-based reasoning grading based on knowledge graphs, and evidence package output." The spatiotemporal graph Transformer performs joint predictions of at least two pollutants across multiple cross-sections under water system connectivity constraints; uncertainty quantification provides confidence information for the prediction results. After conformal prediction calibration, a hierarchical confidence set with coverage guarantees is formed, enabling warning determination to no longer rely on a single prediction point value. Instead, it allows rule triggering driven by the relationship between the confidence set and a threshold, and multi-hop reasoning within the knowledge graph. This achieves spatiotemporal aggregation alarms, complex event identification, and warning grading during weak signal phases, and outputs traceable source tracing evidence through the reasoning path.
[0129] To address the aforementioned technical issues, this invention makes scenario-oriented improvements to the algorithm structure: First, it adopts a hierarchical conformal prediction calibration method that maintains calibration samples and residual distributions hierarchically according to seasonal category, cross-section type, and hydrodynamic scenario category, mitigating the impact of spatiotemporal distribution drift on the reliability of threshold determination and improving the consistency and stability of early warnings under different scenarios; Second, it proposes a rule-triggered mechanism driven by hierarchical confidence sets, upgrading the traditional point value threshold triggering to triggering based on the comparison of upper and lower bounds of confidence intervals, inclusion relationships, and the calculated threshold confidence, making alarms earlier and false alarms controllable; Third, it combines trigger rule identifiers to extract the minimum evidence subgraph covering the reasoning path, calculates the contribution of nodes and edges with the rule satisfaction degree, and forms an evidence package, enabling the early warning results to have auditable and interpretable evidence chain output capabilities, thereby better supporting compound pollution early warning and source tracing decisions.
Claims
1. A composite early warning method for emerging pollutants in Chaohu Lake based on knowledge graphs, characterized in that, include: S1. Acquire new pollutant monitoring data, hydrological and meteorological data, and emission source data from multiple monitoring sections in the Chaohu Lake Basin and standardize them to obtain standardized monitoring records and standardized emission source records. S2. Construct a knowledge graph for composite early warning based on standardized monitoring records and standardized emission source records. Use monitoring sections in the knowledge graph as nodes and water system connectivity as edges. Combine pollutant concentrations and hydrological and meteorological characteristics in the standardized monitoring records to construct a spatiotemporal graph for prediction. S3. Input the spatiotemporal map into the spatiotemporal map Transformer prediction model, perform joint prediction of the concentrations of at least two pollutants within the target prediction time window, obtain the prediction results and their corresponding confidence information parameters, and form the uncertainty quantification results. S4. Based on the prediction results, uncertainty quantification results and corresponding historical actual concentrations, determine the context category according to at least one of the seasonal category, cross section type and hydrodynamic scenario category, maintain calibration samples for each context category and update the residual distribution based on the non-consistency measure of the residuals, perform hierarchical conformal prediction calibration on the prediction results, and obtain the hierarchical confidence set. S5. Based on the hierarchical confidence set, combined with the preset threshold set and preset rule base, the knowledge graph rule reasoning is graded: the rule triggering is determined according to the inclusion relationship between the hierarchical confidence set and the preset threshold set, the comparison relationship between the boundary of the hierarchical confidence set and the preset threshold set, and / or the over-threshold confidence calculated by the hierarchical confidence set. When the triggering is established, multi-hop association reasoning is performed in the knowledge graph according to the entity type, relation type and maximum number of hops defined by the triggering rule to obtain the reasoning path set, thereby determining the compound pollution event and its warning level, and generating the triggering rule identifier at the same time; S6. Based on the compound pollution event, warning level, inference path set and trigger rule identifier, extract candidate subgraphs covering the inference path set from the knowledge graph and remove redundant relationships to obtain the minimum evidence subgraph. Calculate the contribution of nodes and edges in the minimum evidence subgraph and the rule satisfaction of trigger rules. Generate an evidence package containing warning level, compound pollution event, trigger rule identifier, minimum evidence subgraph, contribution and rule satisfaction and output the warning result. S4 includes: The corresponding seasonal category is determined based on the timestamp in the prediction results, the corresponding section type is determined based on the section type associated with the monitoring section entity in the knowledge graph, and the monitoring section corresponding to the timestamp is classified into hydrodynamic scenario category based on the hydrometeorological characteristics in the standardized monitoring records, thereby determining the context category for the prediction results. For the context category, historical actual concentrations that match the prediction results in terms of monitoring section identifier, pollutant identifier, and timestamp are extracted from the standardized monitoring records. The non-consistency measure of the residuals is calculated, wherein the non-consistency measure of the residuals is obtained by normalizing the deviation between the historical actual concentrations and the prediction results in combination with the confidence information characterized by the uncertainty quantification results. Write the inconsistency metric of the residual into the residual distribution corresponding to the context category, and update the residual distribution according to the preset update rule; After the residual distribution is updated, a calibration threshold corresponding to the residual distribution is determined according to a preset significance level, and the prediction result is calibrated based on the calibration threshold to generate the hierarchical confidence set, such that the hierarchical confidence set includes the upper bound and lower bound of the confidence interval generated for the at least two pollutants respectively.
2. The knowledge graph-based composite early warning method for new pollutants in Chaohu Lake according to claim 1, characterized in that, S1 includes: Acquire the monitoring data, hydrological and meteorological data, and emission source data of the new pollutants; The new pollutant monitoring data and hydro-meteorological data are subjected to duplicate record removal, outlier identification and processing, missing value marking and processing, and unit of measurement standardization, so that the pollutant concentration is expressed in a preset unified unit of measurement. Encoding mapping is performed on the monitoring section names in the new pollutant monitoring data to generate monitoring section identifiers, and encoding mapping is performed on the pollutant names in the new pollutant monitoring data to generate pollutant identifiers; Time alignment is performed on the new pollutant monitoring data and hydrological and meteorological data to ensure that the timestamp adopts a preset time granularity and corresponds to the monitoring section identifier under the same timestamp. The standardized monitoring records are generated based on time-aligned pollutant concentrations and hydro-meteorological characteristics. The emission source data is subjected to duplicate record removal, emission source name coding mapping to generate emission source identifiers, geographic location coordinate system one, and pollutant identifier alignment, so that the emission source identifiers are associated with geographic locations and emission association information corresponding to pollutant identifiers, and the standardized emission source records are generated accordingly.
3. The knowledge graph-based composite early warning method for new pollutants in Chaohu Lake according to claim 1, characterized in that, S2 include: Standardized monitoring records are used to generate monitoring section entities and pollutant entities. Monitoring section identifiers and corresponding section types are written to the monitoring section entities, and pollutant identifiers are written to the pollutant entities. Emission source entities are generated using standardized emission source records, and emission source identifiers and geographical locations are written into the emission source entities; Based on the water system data of Chaohu Lake Basin, entities of rivers flowing into the lake are generated, and the water system connectivity between rivers flowing into the lake and monitoring sections is established based on the spatial connectivity between the rivers flowing into the lake and the monitoring sections. Based on the emission correlation information in the standardized emission source records, an emission correlation relationship between emission sources and pollutants is established, and an indication relationship between pollutants and emission sources is established. Establish the transformation relationships between pollutants based on pre-defined knowledge of new pollutant transformation; The aforementioned entities and relationships are written into the knowledge graph to complete the construction of the knowledge graph; After the knowledge graph is constructed, the monitoring section entities in the knowledge graph are used as nodes, and the water system connectivity is used as edge relationships. The node features are generated by combining the pollutant concentrations and hydrological and meteorological characteristics corresponding to each monitoring section identifier in the standardized monitoring records, and a spatiotemporal map for prediction is generated.
4. The knowledge graph-based composite early warning method for new pollutants in Chaohu Lake according to claim 1, characterized in that, S3 include: In the spatiotemporal diagram, node feature sequences of each monitoring section within the historical time window are extracted according to a preset historical time window. The node feature sequences include at least two pollutant concentration sequences corresponding to the monitoring section identifier and hydro-meteorological feature sequences. The temporal and spatial attention calculations are performed on the node feature sequence and the edge relationship of the spatiotemporal graph using the spatiotemporal graph prediction model to obtain the spatiotemporal representation of each monitoring section. Based on the spatiotemporal representation, the predicted values of the pollutant concentrations of at least two pollutants within the target prediction time window are generated through a joint prediction output layer, forming the prediction result; Simultaneously, confidence information parameters corresponding one-to-one with the prediction results are generated through the uncertainty output layer to form the uncertainty quantification result, wherein the confidence information parameters include the prediction variance or prediction interval boundary corresponding to the prediction value.
5. The knowledge graph-based composite early warning method for new pollutants in Chaohu Lake according to claim 1, characterized in that, S5 include: Based on the hierarchical confidence set, the upper and lower bounds of the corresponding confidence intervals are read for each monitoring section identifier and each pollutant identifier. Based on the grading thresholds corresponding to the pollutant identifiers in the preset threshold set, the over-threshold confidence level of the pollutant on the monitoring section is calculated. The over-threshold confidence level is determined at least according to one of the following methods: the comparison result between the upper bound of the confidence interval and the grading threshold, the comparison result between the lower bound of the confidence interval and the grading threshold, and the inclusion relationship between the confidence interval and the grading threshold. Based on a preset rule base, rule triggering determination is performed on the over-threshold confidence levels corresponding to at least two pollutants, such that the rule triggering determination includes at least one of the following types of triggering conditions: the over-threshold confidence levels of at least two pollutants under the same monitoring section identifier meet a preset combination condition; the over-threshold confidence levels of the same pollutant under different monitoring section identifiers meet a preset spatial aggregation condition; and the over-threshold confidence levels under the same monitoring section identifier meet a preset persistence condition within consecutive timestamps. When a rule is triggered, the entity type and relation type are limited in the knowledge graph according to the triggering rule, and the maximum number of hops for multi-hop association reasoning is limited, generating a set of reasoning paths that match the triggering rule; A composite pollution event is generated based on the pollutant identifier set, monitoring section identifier set, and corresponding timestamp range covered by the inference path set. The warning level corresponding to the composite pollution event is determined based on the corresponding hierarchical clause in the preset rule base according to the triggering rule and the over-threshold confidence level. At the same time, a triggering rule identifier corresponding to the triggering rule is generated.
6. The knowledge graph-based composite early warning method for new pollutants in Chaohu Lake according to claim 1, characterized in that, S6 include: Based on the inference path set and trigger rule identifier, candidate subgraphs containing all paths in the inference path set are extracted from the knowledge graph; Redundant relation deletion is performed on the candidate subgraph to generate a minimum evidence subgraph, wherein the redundant relation deletion includes: deleting relations that do not belong to the inference path set without disrupting the connectivity of any path in the inference path set, and retaining the relation associated with the trigger rule identifier and deleting the rest of the relations when there are multiple relations of the same type pointing to the same entity. The overthreshold confidence is calculated based on the hierarchical confidence set, and the contribution of each node and each edge in the minimum evidence subgraph is calculated based on the occurrence frequency, path length and overthreshold confidence of each node and each edge in the inference path set. Based on the rule conditions corresponding to the trigger rule identifier, and combined with the upper and lower bounds of the confidence interval of the hierarchical confidence set, the rule satisfaction degree of the trigger rule is calculated. Generate an evidence package, wherein the evidence package includes an early warning level, a complex pollution event, a trigger rule identifier, a minimum evidence subgraph, the contribution of nodes and edges, and the rule satisfaction degree, which are used to form an interpretable source tracing evidence chain and output an early warning result.
7. The knowledge graph-based composite early warning method for new pollutants in Chaohu Lake according to claim 1, characterized in that, In addition to the upper and lower bounds of the confidence intervals generated for each of the at least two pollutants, the stratified confidence set also includes a joint confidence set for the at least two pollutants. The joint confidence set is generated based on the aggregated statistics of the inconsistency measures of each pollutant, so that the joint confidence set can simultaneously cover the at least two pollutants at a preset significance level.
8. The knowledge graph-based composite early warning method for new pollutants in Chaohu Lake according to claim 5, characterized in that, When limiting the maximum number of hops for multi-hop association reasoning, the propagation time delay upper limit is further determined based on the hydrodynamic scenario category, and only the reasoning paths that satisfy the propagation time delay upper limit are retained as the inference path set.
9. The knowledge graph-based composite early warning method for new pollutants in Chaohu Lake as described in claim 6, characterized in that, When generating the minimum evidence subgraph, redundant relation deletion is performed with the goal of creating a connected subgraph with the minimum cost required to cover the set of inference paths, wherein the cost is determined by one or more of relation type weights, path lengths, and the contribution of nodes and edges.
Citation Information
Patent Citations
Circulating water treatment closed-loop optimization method and system based on water quality prediction
CN120656584A
Biopharmaceutical wastewater cooperative treatment method and system based on graph neural network
CN120943393A