Water quality analysis report generation method and system fusing multi-source data and domain knowledge

CN122797496APending Publication Date: 2026-09-22YANGTZE BASIN ECOLOGY & ENVIRONMENT MONITORING & SCIENTIFIC RESEARCH CENTER YANGTZE BASIN ECOLOGY & ENVIRONMENT ADMINISTRATION MINISTRY OF ECOLOGY & ENVIRONMENT OF THE PEOPLES REPUBLIC OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611235618.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-14
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0006]本发明提供一种融合多源数据与领域知识的水质分析报告生成方法及系统,能够深度融合领域知识与多源(天地空)观测数据、实现从数据到报告端到端智能生成的技术方案,以解决现有技术中效率低、质量不稳、分析深度不足及空间数据处理能力缺失的问题

Benefits of technology

[0021] 1. End-to-end fully automated process, significantly improving efficiency: A closed-loop workflow of "data acquisition—intelligent fusion—knowledge-driven analysis—automatic report synthesis—quality verification" has been constructed, which can complete the generation of professional reports from multi-source raw data without manual intervention. Actual testing has reduced the production time of a routine monthly water quality report from 1-2 days to less than 90 minutes, significantly improving the efficiency of emergency response and routine monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797496A_ABST
    Figure CN122797496A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for generating water quality analysis reports that integrates multi-source data and domain knowledge. The method includes: acquiring satellite / drone imagery, online monitoring data, and unstructured documents; extracting key information using an AI agent and outputting a unified data matrix or tensor using a multi-source data credibility fusion algorithm; constructing a dedicated domain knowledge base; calling standard limits for compliance judgment; performing spatiotemporal analysis and correlation analysis; calculating a comprehensive pollution risk index; matching report templates according to task type; generating professional descriptions and chart instructions using a modular prompt word-driven large language model; performing consistency verification, logic optimization, and automatically applying templates to generate a standardized water quality analysis report. This invention achieves end-to-end automated generation from multi-source data to professional reports, significantly improving efficiency and quality, and supporting precise source tracing and targeted governance decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the intersection of environmental monitoring, remote sensing technology and artificial intelligence, specifically a method and system for generating water quality analysis reports that integrates multi-source data and domain knowledge. Background Technology

[0002] Currently, the generation of water quality monitoring reports mainly relies on manual work, which has the following significant drawbacks:

[0003] 1. Inefficiency and delayed response: Environmental engineers need to manually process massive amounts of data from various instruments and sources (such as online monitoring equipment, laboratory information systems, and satellite remote sensing imagery) to perform compliance comparisons, trend calculations, spatial analysis (such as pollution distribution mapping), and chart creation. A routine monthly report often takes 1-2 days, and when remote sensing data processing and interpretation are involved, the cycle is even longer, making it difficult to meet the management needs for rapid response.

[0004] 2. Prone to errors and inconsistent results: Manual data processing is prone to calculation errors and transcription mistakes. This is especially true when fusing point-based ground monitoring data with area remote sensing data, where a lack of unified and objective processing standards leads to a high reliance on individual experience. Reports written by different personnel often vary significantly in format, analytical depth, spatial representation, and conclusion presentation, resulting in large fluctuations in quality.

[0005] 3. Difficulty in knowledge transfer and superficial analysis: Professional analytical logic (such as pollution source tracing and inference, and remote sensing image interpretation rules) and report writing experience mainly rely on individuals, making them difficult to solidify and standardize. Although some data analysis software or GIS tools exist, they can usually only perform simple statistics, generate routine charts, or display basic remote sensing images. They lack a deep understanding of professional domain knowledge (such as water quality standards, pollution spectral feature libraries, and pollutant migration models) and cross-source reasoning capabilities. They cannot automatically complete advanced analyses such as "identifying algal spatial aggregation hotspots" or "inferring non-point source pollution by combining land use," and they cannot generate complete reports with professional depth and clear conclusions. Summary of the Invention

[0006] This invention provides a method and system for generating water quality analysis reports that integrates multi-source data and domain knowledge. It can deeply integrate domain knowledge with multi-source (space, air, and ground) observation data to achieve end-to-end intelligent generation of reports from data, thereby solving the problems of low efficiency, unstable quality, insufficient analysis depth, and lack of spatial data processing capabilities in existing technologies.

[0007] A method for generating water quality analysis reports by integrating multi-source data and domain knowledge includes the following steps: S1: Collecting satellite remote sensing images, UAV remote sensing images, online monitoring data, and unstructured documents; extracting key information from the unstructured documents using an AI agent; identifying land-water boundaries and removing interference from the remote sensing images; cleaning and aligning the extracted information with structured data; outputting a unified structured data matrix or geographically labeled tensor data; weighted fusion using a multi-source data credibility fusion algorithm to obtain preprocessed point-like ground monitoring data and area remote sensing inversion data; S2: Constructing a dedicated domain knowledge base containing water quality standard limits, pollutant environmental behavior, historical cases, and an inversion model library; calling the standard limits in the knowledge base; comparing the preprocessed point-like ground monitoring data and area remote sensing inversion data obtained in step S1 with the standard limits; identifying exceedance information; and performing time series analysis. S1: Spatial evolution analysis and correlation analysis are performed to calculate the comprehensive pollution risk index, obtaining quantitative risk indicators and analysis conclusions; S2: Match a report outline template according to the task type; using the quantitative risk indicators and analysis conclusions generated in step S2, a large language model is driven by a preset prompt word template to generate professional text descriptions and chart drawing instructions; the conclusions in the generated professional text descriptions are quantitatively evaluated and graded to form a visual graphic report; S3: The visual graphic report generated in step S3 is subjected to consistency verification, checking whether the data cited in the report text is consistent with the statistical values ​​of the structured data matrix or tensor data output in step S1, and the logical coherence of the visual graphic report is optimized; based on manual review feedback, the prompt word template is iteratively optimized through an incremental learning mechanism with a forgetting factor; the verified and optimized content is automatically applied to the preset report template to generate a standardized water quality analysis report.

[0008] Furthermore, in step S1, the AI ​​agent extracts monitoring points, time, indicator values ​​and units from PDFs, Word documents or scanned documents using natural language understanding and pattern recognition technology; the water-land boundary recognition uses a deep learning model based on U-Net or Transformer architecture to extract water pixels and remove cloud layers and interfering pixels.

[0009] Furthermore, in the multi-source data credibility fusion algorithm described in step S1, the observation value of the i-th data source... fusion value for: ; ; In the formula: This represents the final merged water quality value at spatial location (x, y); For the spatial adaptive weights of the i-th data source; The real-time dynamic confidence level is calculated by weighting the i-th data source using four-dimensional quality factors; the real-time dynamic confidence level... For remote sensing data sources, the results are obtained through dynamic evaluation of image cloud cover, atmospheric correction quality, inversion model applicability, and deviation between inversion values ​​and adjacent ground monitoring values. For ground monitoring data, the results are obtained through dynamic evaluation of signal stability and deviation from historical data. The coefficient of spatial variation of water quality in a local window at coordinates (x, y) is given. λ represents the maximum coefficient of variation for the entire water body, and λ is the adjustment coefficient in the range of 0 to 1. , , , Normalize the weights for each quality dimension to satisfy + + + =1; , , , These are four quantitative quality scoring factors.

[0010] Furthermore, the formula for calculating the Comprehensive Pollution Risk Index (CPRI) mentioned in step S2 is as follows: ;in, The intensity index reflects the severity of exceeding the standard; This is a trend deterioration index, reflecting the deteriorating trend of water quality indicators; To correlate complexity indices, the combined risks of multiple pollutants occurring together are quantified; The ecotoxicity index reflects the ecological hazard level of pollutants; α, β, γ, and η are weighting coefficients; the... , , and All values ​​are calculated based on the structured data matrix or tensor data output in step S1 and the standard limits, pollutant combination risk weights, and ecotoxicity weights in the knowledge base.

[0011] Furthermore, the formula for calculating the excess strength index SI is as follows: Where N is the total number of indicators, m is the number of indicators exceeding the standard, and max(C) j ) represents the maximum concentration of index j within the time period, S j It is its standard limit. The importance weight coefficient of this indicator is obtained from the knowledge base;

[0012] The trend deterioration index TI is the mean of the sign function of the linear regression slope of the concentrations of all key indicators over time. Where k is the number of key metrics, and slope j It is the slope of the linear regression of the concentration of index j over time within a given period, where sgn is the sign function; the association complexity index CI is: Where P is the set of significantly correlated pollutant indicator pairs, r pq Let I(p,q) be the Pearson correlation coefficient between indicators p and q, and I(p,q) be the risk weight of a specific pollutant combination obtained from the knowledge base; the ecotoxicity index ETI is: Where Cj is the measured concentration of pollutant, Sj is the standard limit, and Tj is the ecotoxicity weight of pollutant stored in the knowledge base, so as to realize the risk weighting of highly toxic pollutants under the same exceedance conditions.

[0013] Furthermore, the preset prompt word templates mentioned in step S3 include trend description templates, spatial distribution feature templates, and conclusion writing templates; by driving the large language model through multi-prompt word chains, the excess list, trend statistics, pollution spatial distribution heat map metadata, and related inference conclusions generated in step S2 are input into the corresponding templates to generate professional text descriptions that logically match each chapter.

[0014] Furthermore, the consistency verification in step S4 includes: automatically checking whether the out-of-standard values, percentages, and spatial statistics cited in the report text are consistent with the statistical values ​​of the structured data matrix or tensor data output in step S1; the logical coherence optimization includes checking whether the conclusions and analysis sections correspond, and whether the spatial analysis conclusions and time change conclusions contradict each other.

[0015] Furthermore, in step S4, the prompt word template is iteratively optimized based on human review feedback using an incremental learning mechanism with a forgetting factor, as shown in the following formula: In the formula: The system's initial basic prompt word template serves as the model's default report generation rules, writing paradigm, and analysis logic baseline. This is the latest prompt word template after iterative optimization, and it serves as the final basis for subsequent report generation by the system. T represents the total number of manual feedback corrections received by the system, indicating the total number of valid correction samples within this iteration cycle. t is the feedback correction time sequence index variable, ranging from 1 to T, used to mark the chronological order of each manual feedback. t=1 corresponds to the earliest historical feedback, and t=T corresponds to the latest current feedback. Δt is the prompt word correction vector corresponding to the t-th manual feedback, representing the magnitude and direction of corrections to text expression, analysis logic, and conclusion deviations. γ is the forgetting factor, ranging from (0,1), which assigns higher weight to recent feedback and lower weight to distant historical feedback through an exponential weighting mechanism, weakening the impact of outdated and inefficient correction rules and ensuring the accuracy and timeliness of model iteration.

[0016] A water quality analysis report generation system integrating multi-source data and domain knowledge, used to execute the method described above, includes: a data acquisition module for accessing satellite remote sensing imagery, UAV remote sensing imagery, online monitoring data, and unstructured documents; a data preprocessing module for extracting key information from the unstructured documents using an AI agent, and performing land-water boundary identification and interference removal on the remote sensing imagery; cleaning and aligning the extracted information with structured data, outputting a unified structured data matrix or geographically labeled tensor data, and performing weighted fusion using a multi-source data credibility fusion algorithm to obtain preprocessed point-like ground monitoring data and area-like remote sensing inversion data; a knowledge base storage module for storing a dedicated domain knowledge base containing water quality standard limits, pollutant environmental behavior, historical cases, and an inversion model library; and an analysis and mining module for receiving the structured data matrix or tensor data, comparing it with the standard limits in the knowledge base, identifying exceedance information, and performing time-series analysis. The system employs column analysis, spatial evolution analysis, and correlation analysis to calculate a comprehensive pollution risk index, outputting quantitative risk indicators and analytical conclusions. A report synthesis module includes a report framework generation unit and a prompt word engine unit. The report framework generation unit matches a report outline template based on the task type. The prompt word engine unit utilizes the quantitative risk indicators and analytical conclusions output by the analysis and mining module, driving a large language model through preset prompt word templates to generate professional text descriptions and chart drawing instructions, outputting a visual text and graphic report. A quality verification and output module performs consistency verification on the visual text and graphic report, checking whether the data cited in the report text is consistent with the statistical values ​​of the structured data matrix or tensor data, and optimizing the logical coherence of the visual text and graphic report. Based on human review feedback, the prompt word template is iteratively optimized using an incremental learning mechanism with a forgetting factor. The verified and optimized content is automatically applied to a preset report template to generate a standardized water quality analysis report.

[0017] Furthermore, the quality verification and output module includes a consistency verification unit, a logic optimization unit, and a format output unit. The consistency verification unit is connected to the data preprocessing module and is used to verify whether the data referenced in the visualization report is consistent with the statistical values ​​of the structured data matrix or tensor data output by the data preprocessing module. The logic optimization unit is used to verify the logical coherence of the conclusions and analysis sections in the visualization report. The format output unit has built-in Word, PDF, or HTML report templates specific to various organizations, which are used to output the verified and optimized content as a standardized format report with one click.

[0018] The quality verification and output module includes a consistency verification unit, a logic optimization unit, and a format output unit. The consistency verification unit is connected to the data preprocessing module and is used to verify whether the data cited in the visualization report is consistent with the statistical values ​​of the structured data matrix or tensor data output by the data preprocessing module. The logic optimization unit is used to verify the logical coherence of the conclusions and analysis sections in the visualization report. The quality verification and output module is also used to iteratively optimize the prompt word template based on manual review feedback using an incremental learning mechanism with a forgetting factor. The optimization formula is as follows: In the formula: The system's initial basic prompt word template serves as the model's default report generation rules, writing paradigm, and analysis logic baseline. This is the latest prompt word template after iterative optimization, and it serves as the final basis for subsequent report generation by the system. T represents the total number of manual feedback corrections received by the system, indicating the total number of valid correction samples within this iteration cycle. t is the feedback correction time sequence index variable, ranging from 1 to T, used to mark the chronological order of each manual feedback. t=1 corresponds to the earliest historical feedback, and t=T corresponds to the latest current feedback. Δt is the prompt word correction vector corresponding to the t-th manual feedback, representing the magnitude and direction of corrections to text expression, analysis logic, and conclusion bias. γ is the forgetting factor, ranging from (0,1), which assigns higher weight to recent feedback and lower weight to distant historical feedback through an exponential weighting mechanism, weakening the impact of outdated and inefficient correction rules and ensuring the accuracy and timeliness of model iteration.

[0019] The format output unit has built-in multiple organization-specific Word, PDF or HTML report templates, which are used to output the verified and optimized content as a standardized format report with one click.

[0020] Compared with the prior art, the present invention has the following advantages:

[0021] 1. End-to-end fully automated process, significantly improving efficiency: A closed-loop workflow of "data acquisition—intelligent fusion—knowledge-driven analysis—automatic report synthesis—quality verification" has been constructed, which can complete the generation of professional reports from multi-source raw data without manual intervention. Actual testing has reduced the production time of a routine monthly water quality report from 1-2 days to less than 90 minutes, significantly improving the efficiency of emergency response and routine monitoring.

[0022] 2. Deeply integrate domain knowledge to generate professional and traceable intelligent analysis conclusions: By constructing a dedicated domain knowledge base including water quality standards, pollution mechanisms, and spectral feature libraries, and combining retrieval-enhanced generation (RAG) and modular prompt word engineering, static professional knowledge is dynamically injected into the large language model, enabling the system to automatically complete advanced analysis tasks such as exceeding standards, spatiotemporal trend analysis, and pollution source inference, generating reports with professional depth and traceable conclusions, eliminating the subjective bias and experience dependence of traditional manual analysis.

[0023] 3. Reliable fusion of multi-source heterogeneous data to achieve integrated "point-area" water quality perception: The original multi-source data reliability fusion algorithm (MDCFA) uses the weighted fusion of inherent device weights and real-time dynamic confidence (such as image quality and data deviation) and introduces spatial heterogeneity adaptive weights to dynamically adjust the contribution of remote sensing and ground data according to the spatial variation characteristics of water quality parameters. It effectively integrates sparse and high-precision ground point monitoring data with continuous area remote sensing inversion data to achieve the optimal balance between local accuracy and spatial continuity.

[0024] 4. Quantify the comprehensive pollution risk index to support precise early warning and targeted governance: The proposed comprehensive pollution risk index (CPRI) integrates four dimensions: exceedance intensity, deterioration trend, pollutant association complexity and ecotoxicity. It also supports probabilistic assessment based on Monte Carlo simulation, and can generate spatially continuous risk distribution maps and quantitative indicators. It supports objective ranking of pollution risks, automatic early warning and governance decisions, overcomes the limitations of traditional single indicator evaluation, and quantifies the uncertainty of assessment.

[0025] 5. Unified and standardized report quality, one-click output of standardized results: Automated consistency verification, logic optimization and format application mechanisms ensure that the report text data is strictly consistent with the source data and that the spatial logic is self-consistent. It can also output standardized reports in Word, PDF and other formats that conform to the organization's specifications with one click. At the same time, the incremental learning mechanism based on historical feedback continuously optimizes the quality of report generation, completely solving the problems of inconsistent formatting, easy errors and large quality fluctuations when writing reports manually. Attached Figure Description

[0026] Figure 1 This is a flowchart of a water quality analysis report generation method that integrates multi-source data and domain knowledge according to an embodiment of the present invention.

[0027] Figure 2 This is a schematic diagram of the spatial distribution raster layer of chlorophyll a and chemical oxygen demand in the reservoir area, generated by an empirical model based on a domain knowledge base in an embodiment of the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] Please see Figure 1 This invention provides a method for generating water quality analysis reports that integrates multi-source data and domain knowledge. Its core lies in constructing an intelligent workflow driven by a collaborative approach of "space-air-ground data-knowledge-task," integrating remote sensing observation and ground monitoring. The method comprises four steps, each working synergistically to achieve the automated generation of a professional report from raw data.

[0030] Step 1: Intelligent Acquisition and Structured Preprocessing of Multi-Source Heterogeneous Data: The aim is to solve the integration challenges caused by diverse data sources and varying formats, providing a unified, clean, and reliable spatial-temporal-attribute integrated data foundation for subsequent analysis. Specific technical methods are as follows:

[0031] 1. Full access to multi-source data: The system accesses or receives raw data from multiple sources, including satellite / UAV remote sensing imagery, online monitoring equipment, and manual record sheets. The remote sensing data includes multispectral, hyperspectral, or synthetic aperture radar imagery, specifically sourced from Landsat, Sentinel, GF series, and other satellite or UAV-borne sensors.

[0032] 2. AI Intelligent Analysis and Image Preprocessing: Utilizing a customized AI agent, the system automatically extracts key information entities such as monitoring points, time, indicator items, values, and units from unstructured documents (such as PDF reports, scanned documents, and Word documents) through Natural Language Understanding (NLU) and pattern recognition technologies. Simultaneously, for remote sensing images, the AI ​​agent calls deep learning-based object detection and semantic segmentation models (such as U-Net and Transformer architectures) to automatically identify land-water boundaries, extract water pixels, and remove interfering pixels such as clouds, shadows, and buildings.

[0033] 3. Data Standardization Preprocessing: The information extracted by AI is fused with existing structured data (such as database exported tables, CSV files, and raster data obtained from remote sensing inversion). Through data cleaning (such as removing outliers and filling in reasonable missing values) and alignment (unifying timestamps, units, spatial coordinate systems, and resolution), the output is a unified, machine-readable structured data matrix or geographically labeled tensor data.

[0034] 4. The Multi-Source ("Sky-Air-Ground") Data Credibility Fusion Algorithm (MDCFA) is used to perform weighted fusion of structured data matrices or geographically labeled tensor data to obtain preprocessed point-like ground monitoring data and area remote sensing inversion data.

[0035] Traditional water quality monitoring relies on sparse, fixed monitoring points, failing to reflect the spatial continuity and heterogeneity of water bodies. Remote sensing technology provides large-scale, periodic, and synchronous observation capabilities, but the retrieved water quality parameters (such as chlorophyll a, suspended matter, and transparency) suffer from indirectness and atmospheric interference, leading to uncertainties. This step addresses this by constructing a "sky-air-ground" data fusion framework and a dynamic confidence assessment mechanism. It utilizes high-precision ground point data to calibrate and validate remote sensing surface data, while simultaneously using remote sensing surface data to interpolate spatial blind spots between ground point data. This results in a water quality information field that possesses both local accuracy and spatial continuity, resolving the technical problem of insufficient spatiotemporal representativeness from a single data source.

[0036] To address the discrepancies in spatiotemporal representativeness and accuracy between data from different sources (especially point-based ground monitoring data and area-based remote sensing data), a weighted fusion method is employed for multiple observations of the same indicator within the same spatiotemporal point or spatial unit. Let the observation value of the i-th data source be V. i Its inherent equipment weight is W i 0, with a real-time dynamic confidence level of C. i The system incorporates spatial heterogeneity adaptive weights for basic equipment, amplifies the advantages of remote sensing surface data in heavily polluted areas, and prioritizes the use of high-precision ground point data in uniform water bodies. This achieves optimal fusion of precise ground point monitoring and continuous remote sensing surface observation, significantly improving the accuracy of water quality data across the entire region.

[0037] 1) Differentiate scoring based on error sources and quality dimensions from different data sources to achieve real-time dynamic confidence levels. The fully automated, quantifiable, and reproducible calculation is shown in the following formula:

[0038] In the formula: Let be the real-time comprehensive confidence level of the i-th data source, with a value range of [0,1]. , , , Normalize the weights for each quality dimension to satisfy + + + =1; , , , The four quantitative quality scoring factors are defined differently based on the data source.

[0039] For remote sensing data sources, the four quantization factors are defined as follows: Cloud quality score (cloud amount 0~100% linear inverse normalization) The atmospheric correction quality score (based on correction residual normalization) The inversion model fit score (water body optical type matching degree) is used to determine the model fit score. The ground verification deviation score is calculated by inversely normalizing the deviation between the remote sensing value and the nearest true value.

[0040] For ground-based monitoring data sources, the four quantitative factors are defined as follows: The equipment signal stability score, For the continuity score of time series data, For historical data consistency score, The score represents the deviation from outliers.

[0041] 2) Dynamically adjust the data source weights based on the local spatial variation characteristics of water bodies to adapt to the uneven spatial distribution of water bodies. The formula is as follows: .

[0042] In the formula: The coefficient of spatial variation of water quality in a local window at coordinates (x, y) is given. λ represents the maximum variation coefficient of the entire water body, and λ is the adjustment coefficient in the range of 0 to 1. Remote sensing surface data weights are automatically increased for areas with severe pollution gradients and complex water quality changes, while high-precision ground monitoring data are prioritized for lake centers with uniform and stable water quality.

[0043] 3) Global Adaptive Fusion Calculation Formula: .

[0044] In the formula: This represents the final merged water quality value at spatial location (x, y); For the spatial adaptive weights of the i-th data source; The real-time dynamic confidence level is obtained by weighting the i-th data source using four-dimensional quality factors. Let be the original observation value of the i-th data source. This formula relies on the real-time quality status of the data source and the dynamic weighting based on local spatial heterogeneity. Inferior data is automatically downweighted, while high-quality data dominates the fusion result. It provides dual protection for the accuracy of point-to-surface fusion from both spatial and data quality dimensions, solving the problems of poor adaptability and result distortion in traditional fusion models.

[0045] Step two, core analysis and insight mining based on Retrieval Enhanced Generation (RAG): Deeply integrating domain expertise and multi-dimensional remote sensing spatiotemporal analysis into the data analysis process to achieve professional and intelligent analysis that goes beyond simple statistics, and to generate quantitative risk indicators. The specific steps are as follows:

[0046] 1. Construct a domain-specific knowledge base: Establish and maintain a structured knowledge base in advance, covering: (a) limit tables of relevant water quality standards and regulations (such as GB 3838-2002); (b) professional knowledge such as the environmental behavior, source characteristics, and health risks of various pollutants; (c) historical typical case analysis reports; (d) professional statistical analysis and water environment models; and (e) a library of spectral characteristics of typical pollutants and a library of quantitative inversion models for water quality parameters.

[0047] 2. Automatic compliance assessment: The system calls upon the standard limits in the knowledge base and automatically compares the preprocessed point-based ground monitoring data and area remote sensing inversion data according to point location, spatial grid, time, and indicator category. It accurately identifies the areas, points, times, and multiples of exceeding the standards, and generates a structured list of exceeding standards and metadata of the spatial distribution of pollution heat map.

[0048] 3. Multi-dimensional Trend and Correlation Analysis: The AI ​​engine automatically performs time series analysis (year-on-year, month-on-month, and moving average) on the data. Simultaneously, it performs spatiotemporal analysis on remote sensing inversion parameter sequences of the same region at different time phases to identify the spatiotemporal evolution patterns of water quality parameters (such as the spatial diffusion path of algal blooms and the migration direction of pollution zones), obtaining trend statistics. Furthermore, based on pollution models in the knowledge base, it performs correlation analysis on combinations of abnormal data to obtain correlation inference conclusions. For example, by combining shoreline land use types (agricultural land, industrial areas, residential areas) and river confluence networks identified by remote sensing imagery with downstream monitored water quality anomalies, it performs spatial causal inference to propose possible pollution source locations, types, and contribution path hypotheses.

[0049] 4. Comprehensive Pollution Risk Index (CPRI): To quantitatively assess the overall risk of any spatial location or grid cell, CPRI is defined as follows: .

[0050] In the formula: α, β, γ, and η are the preset weight coefficients for each dimension.

[0051] Excess Strength Index (SI): .

[0052] This reflects the severity of exceeding the standard. Where N is the total number of indicators, m is the number of indicators exceeding the standard, and max(C) = ... j ) represents the maximum concentration of index j within the time period, S j It is its standard limit. This is the importance weighting coefficient of the indicator (obtained from the knowledge base).

[0053] Trend Deterioration Index (TI): k represents the total number of key metrics; slope jis the slope of the linear regression of the concentration of index j over time; sgn is the sign function, with positive values ​​representing an upward trend.

[0054] Association Complexity Index (CI): This is used to quantify the combined risks arising from the co-occurrence and interrelation of multiple pollutants. Here, p represents all statistically significant correlations (e.g., the absolute value of the correlation coefficient |r0|). pq | A set of pollutant indicator pairs that are greater than a certain set threshold. pq : is the Pearson correlation coefficient between indicators p and q, measuring the degree of correlation between their changes. I(p,q): is a weighted coefficient obtained from the domain knowledge base, representing the risk level of the pollution type or source indicated by a specific combination of pollutants (such as ammonia nitrogen and total phosphorus) (e.g., agricultural non-point source pollution, industrial point source pollution, etc.). The higher the risk, the larger this value.

[0055] r pq 2 The coefficient of determination reflects the strength of the association between two variables.

[0056] Ecotoxicity index (ETI): Used to differentiate the ecological hazard levels of different pollutants. .

[0057] In the formula: C j S represents the measured concentration of the pollutant. j For standard limits, T j The ecotoxicity weights of pollutants stored in the knowledge base enable risk weighting of highly toxic pollutants under the same exceedance conditions, thus meeting the needs of ecological protection.

[0058] 5. Uncertainty Assessment: This invention introduces the Monte Carlo stochastic simulation method to generate CPRI probability distribution results, thereby achieving risk assessment.

[0059] The full quantitative characterization of uncertainty is assessed, and the specific calculation process is as follows:

[0060] 1) Random perturbation sampling of data. Different perturbation intervals are set according to the error characteristics of different data sources. ±15% Gaussian random perturbation is applied to the remote sensing inversion water quality data, and ±5% Gaussian random perturbation is applied to the high-precision laboratory sampling data to generate multiple sets of input datasets with error perturbation in batches.

[0061] 2) Batch simulation iterative calculation. For each disturbance dataset, CPRI calculation is performed independently, and the simulation is repeated S times to obtain a sample set of risk values. This invention sets the number of simulations S≥10000 to ensure that the probability statistics results are convergent, stable and reproducible.

[0062] 3) Calculation of quantile risk ceiling. Based on the CPRI sample set obtained from batch simulations, the risk ceiling value at a specified percentile is calculated using the following formula: .

[0063] In the formula: To correspond to the upper limit of risk for the percentile, p is preset to values ​​of 90% and 95%. , These can be used to characterize extreme values ​​of water pollution risk at high confidence levels, thus avoiding the risk of underestimating potential risks in single-point valuations.

[0064] 4) Quantitative Calculation of Reliability Assessment. The coefficient of variation is used to quantify the dispersion and reliability of the risk assessment results, as shown in the following formula: .

[0065] In the formula: To simulate the standard deviation of CPRI samples in batches, To simulate the mean of CPRI samples in batches; The coefficient of variation is used for risk assessment. The larger the value, the stronger the impact of data disturbance on the assessment results, and the lower the reliability of the risk assessment at that location.

[0066] This invention establishes a standardized quantitative grading and adaptive response mechanism: The assessment is highly reliable, and the results are highly credible. The assessment results are subject to moderate uncertainty and may fluctuate slightly. Due to high uncertainty, risk assessment results have poor reliability. For areas with high uncertainty, the system automatically pushes monitoring optimization suggestions based on encrypted on-site sampling and supplementary remote sensing observations. This enables quantitative quality control, reliability judgment, and adaptive correction of water quality risk assessment, addressing the technical shortcomings of traditional deterministic assessments that cannot determine the credibility of the assessment and are prone to misjudgment.

[0067] Step 3: Automatic Synthesis of Structured and Modular Trustworthy Reports: This step automatically generates standardized, professional, and traceable report content based on multi-dimensional analysis results and hierarchical prompt word templates. It also quantifies and labels the credibility of conclusions, ensuring the quality of the report output. The specific implementation process is as follows:

[0068] 1. Adaptive Report Framework Matching. Based on the monitoring task type, the system automatically matches specific templates for daily, monthly, annual, and special reports, generating a complete report outline that includes a summary, monitoring overview, data analysis, risk assessment, source tracing conclusions, and governance recommendations.

[0069] 2. Modular Intelligent Content Generation. For different chapters such as trend analysis, spatial distribution, risk assessment, and pollution source tracing, a dedicated optimized prompt word chain is configured to drive the large language model to generate rigorous, objective, and professional text descriptions by combining structured analysis results and professional content from the knowledge base, avoiding colloquial and subjective expressions. For example, for an identified spatial cluster of high chlorophyll a concentration, the model will generate the following description based on the prompt words: "Based on the inversion analysis of Sentinel-2 imagery from August 15, 2023, a high chlorophyll a concentration area of ​​approximately 12 square kilometers was found in the northwest of the lake center area, indicating a risk of algal enrichment in this area, which needs further analysis in conjunction with concurrent meteorological conditions."

[0070] 3. Quantitative Labeling of Conclusion Confidence. To avoid the problems of distortion in the output of large models and unreliable conclusions, the system assigns a confidence score to each generated conclusion, using the following formula: .

[0071] Where: N evi To effectively support the number of pieces of evidence, N total Here, k represents the theoretical maximum amount of evidence, θ is the steepness coefficient, and θ is the confidence threshold. This invention establishes a standardized three-level confidence level grading and labeling mechanism, with the following specific rules: Conf(R) ≥ 0.7 indicates a high-confidence conclusion with sufficient evidence support, which is directly adopted in the report without additional labeling; 0.4 ≤ Conf(R) < 0.7 indicates a medium-confidence conclusion with moderate evidence support, and the corresponding conclusion in the report is labeled "Based on existing evidence, the reliability of the conclusion is moderate"; Conf(R) < 0.4 indicates a low-confidence conclusion with weak evidence support, and the corresponding conclusion in the report is labeled "This conclusion is based on limited evidence, further verification is recommended." Furthermore, the system automatically summarizes and aggregates all low-confidence conclusions, uniformly presenting them in the "Uncertainty Explanation" section of the report. This achieves transparent and verifiable confidence level grading of report conclusions, comprehensively avoiding the distortion of conclusions caused by the illusion of large models.

[0072] 4. Visualization Element Integration: Visualization elements are automatically generated. Based on the data analysis type, the system automatically matches chart and thematic map generation commands, outputting visualizations such as water quality spatial distribution maps, risk heat maps, time-series evolution curves, and pollution hotspot vector boundaries. It supports direct parsing and integration with third-party charts and GIS tools. For example, for the analyzed spatial distribution of water quality parameters, the system automatically outputs "Draw a thematic map of chlorophyll a concentration spatial distribution in the lake area" along with corresponding color scales, legends, and projection parameter commands; for spatiotemporal evolution, it outputs "Generate an animation of the spatiotemporal evolution of suspended solids concentration over the past three months." These commands, or the directly output data sequences and format requirements, can be parsed by chart tools (such as Matplotlib, ECharts, and GIS components) for subsequent integration into reports.

[0073] Step Four: Multi-dimensional Intelligent Quality Control and Incremental Iterative Standardized Output: This step ensures the accuracy of report data, the rigor of logic, and the standardization of format through full-dimensional verification and incremental model learning. Simultaneously, it enables the system to autonomously iterate and optimize, continuously improving output quality. The specific implementation process is as follows:

[0074] 1. Full-dimensional consistency and logical verification. The system automatically verifies the consistency between the report text data and the original data source and the integrated dataset, and verifies the standardization of units, terminology, and spatial coordinates; it also simultaneously optimizes the logic of the entire text to ensure that the time-series analysis, spatial judgment, causal tracing, and risk conclusions are consistent and without contradictions or deviations.

[0075] 2. Incremental self-learning optimization with a forgetting factor. The system relies on manual review feedback to continuously iterate and optimize the prompt word template. The prompt word template is a set of structured instructions used in this invention to drive the large language model to generate standardized and professional water quality analysis text. It has built-in fixed writing paradigms, domain reasoning logic, professional terminology specifications, and output constraints. Unlike the formatted report template, its function is to constrain the AI's analysis and reasoning process and text generation logic. Combined with historical manual correction feedback, incremental iterative optimization is achieved, ensuring that the report's analysis logic, source tracing inference, and professional expression continuously adapt to the characteristics of water quality evolution in the reservoir area. The formula is as follows:

[0076] .

[0077] In the formula: The system's initial basic prompt word template serves as the model's default report generation rules, writing paradigm, and analysis logic baseline. This is the latest prompt word template after iterative optimization, serving as the final basis for subsequent report generation by the system. T represents the total number of manual feedback corrections received by the system, signifying the total number of valid correction samples within this iteration cycle. t is the feedback correction time sequence index variable, ranging from 1 to T, used to mark the chronological order of each manual feedback; t=1 corresponds to the earliest historical feedback, and t=T corresponds to the latest current feedback. Δt is the prompt word correction vector corresponding to the t-th manual feedback, representing the magnitude and direction of corrections to textual expression, analytical logic, and conclusion biases. γ is the forgetting factor, ranging from 0 to 1, which uses an exponential weighting mechanism to assign higher weight to recent feedback and lower weight to older historical feedback, weakening the impact of outdated and inefficient correction rules and ensuring the accuracy and timeliness of model iteration. Simultaneously, the system uses self-learning maturity ML to quantify system stability and dynamically adjust the frequency of manual sampling, significantly reducing manual maintenance costs after model iteration matures.

[0078] 3. Standardized format one-click output. After verification, scoring, and optimization, the system automatically applies the organization's exclusive template, unifying the formatting, fonts, headers and footers, map scale, north arrow, etc., and outputs compliant Word and PDF standardized reports with one click, meeting the requirements for direct delivery.

[0079] Based on the aforementioned research methods, the embodiments of this invention can automate the batch production of monthly, quarterly, and annual water quality reports within a reservoir area, and are adaptable to various business scenarios such as routine water quality monitoring, periodic water quality reviews, and specialized water quality assessments. The following uses the automated generation of a monthly water quality monitoring report for a specific reservoir as an example to illustrate the complete implementation process and technical effectiveness of this invention. The specific implementation method is as follows:

[0080] 1. Data input and preprocessing (corresponding to step one)

[0081] This system automatically completes batch acquisition, intelligent parsing, and standardized preprocessing of multi-source heterogeneous data, achieving spatiotemporal alignment of multi-source monitoring data from the sky, air, and ground. It addresses the pain points of traditional data, such as disorganized formats, inconsistent spatiotemporal benchmarks, and the difficulty in reusing unstructured data. This provides a standardized, high-precision, and comprehensive data foundation for subsequent intelligent analysis. Specific data sources and processing methods are as follows:

[0082] (1) Multi-source data acquisition: The system automatically accesses three types of core monitoring data, namely online monitoring data, laboratory sampling data and satellite remote sensing image data. Among them, the online monitoring data are hourly monitoring data of the five automatic monitoring sections in the reservoir area over the past 30 days, covering conventional water quality indicators such as pH, dissolved oxygen, permanganate index, ammonia nitrogen, and total phosphorus; the laboratory sampling data are daily manual sampling laboratory analysis data of the same monitoring section in the same time period, which serve as high-precision true value reference data; the remote sensing image data are three Sentinel-2 multispectral images covering the entire reservoir area, selecting high-quality image data with 10-meter resolution and no or few clouds to ensure the effectiveness of large-scale spatial observation.

[0083] (2) Intelligent Inversion of Remote Sensing Data: The system relies on AI agents to automatically complete the preprocessing of remote sensing images, achieving precise segmentation of land-water boundaries, extraction of water body masks, and accurate removal of invalid pixels such as clouds, shoreline buildings, and shadows, avoiding interference from non-water areas. Simultaneously, it calls upon the pre-built chlorophyll a inversion model library and permanganate index inversion model library in the domain knowledge base to complete the full-domain spatial quantitative inversion of the core water quality parameters of the reservoir area, generating a 10m×10m resolution raster layer showing the spatial distribution of chlorophyll a and permanganate index. A schematic diagram of the results is shown below. Figure 2 As shown. The calculation formulas for the core water body parameter inversion bands are as follows: ; .

[0084] In the formula, bands B2, B3 and B4 correspond to the green band, red band and near-infrared band of Sentinel-2 satellite data, respectively. Through standardized band combination calculations and combined with domain experience models, accurate quantitative inversion of water quality parameters is achieved.

[0085] (3) AI Agent Information Extraction: The customized AI agent relies on natural language understanding and pattern recognition technology to automatically identify and parse various unstructured data such as laboratory test PDF reports and paper scanned documents, accurately extracting core business fields such as monitoring points, sampling time, measured values ​​of ammonia nitrogen, and measured values ​​of total phosphorus, and completing the structured translation of unstructured data. The parsed measured data is then precisely aligned with online monitoring time series data and remote sensing raster data in spatiotemporal dimensions, unifying timestamps, spatial coordinate systems, and indicator units to construct a unified and standardized fusion dataset across the entire domain.

[0086] (4) "Sky-Air-Ground" MDCFA Multi-Source Data Fusion: This invention adopts an original multi-source data credibility fusion algorithm (MDCFA) to specifically address the differences in spatiotemporal representativeness and monitoring accuracy between point ground monitoring data and area remote sensing data. It achieves precise adaptive weighted fusion in two dimensions: point location and overall spatial domain. For multi-source observation data of the same water quality indicator within the same spatiotemporal unit, the fusion calculation is completed by combining the inherent weights of the equipment with real-time dynamic confidence. At the same time, a spatial heterogeneity adaptive mechanism is introduced to amplify the spatial coverage advantage of remote sensing area data in areas with severe pollution and large water quality gradient changes. In waters with uniform and stable water quality, high-precision ground point data is given priority, thus achieving optimal fusion of point and area data.

[0087] The specific fusion process is as follows: For ground monitoring points, online ammonia nitrogen monitoring values ​​(inherent weight W=0.8, real-time dynamic confidence level C=0.9) and high-precision laboratory sampling true values ​​(inherent weight W=1.0, real-time dynamic confidence level C=0.95) are weighted and fused to effectively reduce the slight drift error caused by long-term operation of online equipment, generating time-series data of daily average values ​​with higher accuracy and stronger stability. For the entire water surface space of the reservoir area, remote sensing inversion chlorophyll a raster data and measured data from 5 ground monitoring points are cross-validated and spatially adaptively weighted and fused: For any raster unit, if a ground monitoring point is included within a 3×3 window, the remote sensing inversion weight is reduced to 0.6 and the ground measured weight is increased to 0.4, combined with Kriging interpolation to achieve point-area collaborative fusion; if there are no ground monitoring points, the remote sensing inversion data after atmospheric correction and model correction (basic weight W=0.7, excellent image quality in this case, dynamic confidence level C=0.85) is used entirely. The final output is a continuous water quality grid dataset with a resolution of 10 meters and a daily scale for the entire reservoir area. Blank dates without remote sensing observations are filled in using a spatiotemporal interpolation algorithm, achieving seamless full coverage of water quality data across the entire area and all time periods, and significantly improving the accuracy and spatial integrity of water quality monitoring data across the entire area.

[0088] 2. Core Analysis and Insight Extraction (corresponding to Step Two)

[0089] The system utilizes a built-in domain-specific knowledge base, including the GB 3838-2002 Class III standard for surface water environmental quality, a database of typical pollutant spectral characteristics in reservoir areas, a watershed land use and shoreline spatial database, and a pollution mechanism model database. Leveraging RAG retrieval enhancement technology, static domain knowledge is dynamically integrated into the data analysis process, driving AI to perform multi-dimensional intelligent analysis, quantitative risk assessment, and source tracing inference. This overcomes the limitations of traditional, purely statistical, shallow analysis. The specific analysis process is as follows:

[0090] (1) Automatic water quality compliance assessment: The system links point-based ground monitoring data with area-based remote sensing grid data, automatically compares the data with national standard limits to complete the water quality compliance screening of the entire area. Based on ground monitoring data, it was identified that there were 3 days this month where the hourly monitoring value of ammonia nitrogen exceeded the standard. The standard limit for ammonia nitrogen in Class III water quality is 1.0 mg / L, and the maximum exceedance multiple is 1.2 times. Based on the remote sensing grid data of the entire area, it accurately completed the spatial anomaly identification and determined that there is a high chlorophyll a concentration area of ​​about 2.1 square kilometers in the northwestern water inlet area of ​​the reservoir. The chlorophyll a concentration in this area is consistently greater than 30 μg / L, which is two standard deviations above the historical average for the same period in the lake area. It was determined to be an area with abnormally high algal biomass and a potential risk of algal bloom.

[0091] (2) Spatiotemporal trend and spatial pattern analysis: The spatiotemporal evolution of water quality was explored from multiple dimensions. In the point-like time series dimension, the monthly time series data of section 5 at the reservoir inlet were automatically collected. The monthly average concentration of ammonia nitrogen this month was 0.85 mg / L, an increase of 15% year-on-year. The slope of concentration change slopeNH3-N=0.02 mg / L / day was obtained by linear regression fitting, and the water quality index showed a continuous deterioration upward trend. In the area-like spatial dimension, spatial hotspot analysis was carried out on the chlorophyll a raster data of the entire reservoir area. The northern shoreline of the reservoir area was identified as a significant area of ​​high algal concentration (z-score > 2.58, p<0.01). The vector boundary and precise spatial range of the pollution hotspots were automatically extracted and output. In the spatiotemporal correlation dimension, the high chlorophyll a hotspot area was spatially overlaid with the shoreline land use type (agricultural land, forest land, and rural construction land) interpreted by remote sensing at the same time. The results showed that 86% of the high algal concentration area was distributed within the 500-meter buffer zone of agricultural land. Based on the 120mm cumulative rainfall record this month, AI automatically completed causal inference using knowledge of pollution mechanisms in the field: the heavy rainfall caused the input of nitrogen and phosphorus pollutants from agricultural non-point sources along the coast, resulting in the enrichment of nutrients in the water and ultimately causing abnormal proliferation of algae in the local area.

[0092] (3) Pollutant correlation analysis: The system automatically mines the linkage change pattern of water quality indicators before and after rainfall events. Through correlation statistical analysis, it is found that the Pearson correlation coefficient r of ammonia nitrogen and total phosphorus concentrations at ground monitoring points within 24 to 48 hours after rainfall is 0.85. The two types of pollutants show highly coordinated change characteristics, which fully corroborates the conclusion of source inference of agricultural non-point source pollution input, effectively avoids the one-sidedness of traditional single indicator analysis, and improves the scientificity and accuracy of source inference analysis.

[0093] (4) Quantitative Calculation of CPRI Comprehensive Pollution Risk: Based on the four-dimensional comprehensive pollution risk index model proposed in this invention, weighting coefficients α=0.5, β=0.3, γ=0.2, and η=0.2 were set. Monthly CPRI calculations were performed on section 5 at the reservoir inlet, and the CPRI value of the section was 0.65. According to the preset classification standards (0~0.3 for low risk, 0.3~0.7 for medium risk, and 0.7~1 for high risk), the section was determined to be of medium pollution risk. At the same time, the system completed the refined calculation of CPRI risk for the entire water area grid by grid. The exceedance intensity index and the trend deterioration index were calculated based on time-series remote sensing chlorophyll a data, and the association complexity index was weighted based on the land use knowledge base and the pollutant association model. Finally, a continuous CPRI risk spatial distribution grid for the entire reservoir area was generated. According to systematic statistics, the area of ​​high pollution risk in the reservoir area reaches 5.6 square kilometers, mainly concentrated in the northwest inlet and the northern agricultural coastal zone. This is highly consistent with the spatial distribution of water quality anomalies and the results of pollution source tracing, realizing a quantitative, spatial, and refined assessment of water quality risk.

[0094] 3. Report content synthesis and visualization output (corresponding to step three)

[0095] Based on the monitoring business scenario, the system automatically matches a standardized template for "Monthly Water Quality Report (including remote sensing thematic analysis)". Through multi-level modular prompt word chains, it accurately drives a large language model, deeply integrates structured analysis data and professional content from the domain knowledge base, and automatically generates rigorous, objective, logically closed-loop, and traceable professional report text. This avoids the subjectivity and colloquialism problems of manual writing. The core text generation example is as follows:

[0096] "Ground monitoring shows that the monthly average ammonia nitrogen level at the inlet section this month was 0.85 mg / L, up 15% year-on-year. The comprehensive pollution risk index (CPRI) was 0.65, which is classified as medium risk. The main driving factor is the trend of ammonia nitrogen increase and its strong correlation with total phosphorus (r=0.85 after rainfall), indicating an increased risk of agricultural non-point source pollution input."

[0097] "Remote sensing monitoring shows that there is significant spatial accumulation of chlorophyll a along the northern shoreline of the reservoir, covering an area of ​​approximately 2.1 square kilometers. The CPRI spatial distribution map shows that the risk level of this area is in the medium-high range (0.6~0.8). Combined with the results of shoreline land use classification, it is inferred that agricultural runoff caused by recent heavy rainfall is the main cause of the abnormal increase in algal biomass in this area."

[0098] Simultaneously, the system automatically generates standardized visualization results and adaptation commands based on multidimensional analysis results, which can be directly connected to various charts and GIS tools. Specifically, this includes: creating a pseudo-color map of the spatial distribution of chlorophyll a concentration in the reservoir area (overlaid with monitoring section locations and CPRI risk levels), with corresponding standardized color levels, legends, and projection parameters; generating monthly temporal changes in hotspot areas in the northern part of the reservoir area, and fully reconstructing the spatiotemporal evolution of algal pollution based on four periods of temporal remote sensing imagery this month. All visualization results are precisely matched with the analysis text and can be directly embedded into the report text, achieving integrated text and image output.

[0099] 4. Intelligent quality verification and standardized output (corresponding to step four)

[0100] The system performs full-dimensional automated intelligent quality control verification, strictly controlling data consistency, logical rationality, and format standardization to ensure the accuracy, rigor, and compliance of the output reports. First, data consistency verification automatically checks all core quantitative indicators in the report, such as the number of days exceeding the ammonia nitrogen standard, the area of ​​chlorophyll a hotspot, and the area of ​​risk zones, ensuring complete consistency with the preprocessed dataset and remote sensing raster statistical results. It also unifies the entire spatial coordinate system to WGS84 UTM 49N to eliminate data deviations and spatial misalignments. Second, logical rationality verification automatically verifies that the time-series analysis conclusions, spatial distribution characteristics, pollution source inferences, and risk classification results are consistent and logically self-consistent, without contradictions or deviations. Third, standardized format output: after verification, the system automatically applies the reservoir-specific official report template, standardizing the entire text layout, fonts, headers and footers, map scale, north arrow, and other standardized formats. It integrates all results, including analysis text, spatial thematic maps, time-series line graphs, and risk distribution maps, generating a standardized PDF report ready for immediate use with a single click.

[0101] Implementation effect representation:

[0102] 1. Efficiency: The water quality report production cycle has been shortened from 1-2 days to less than 90 minutes, greatly improving the efficiency of routine monitoring and emergency response.

[0103] 2. Accuracy level: The spatiotemporal adaptive fusion algorithm solves the distortion problem of traditional data fusion, takes into account both point accuracy and spatial continuity, and significantly improves the accuracy of water quality perception across the entire area.

[0104] 3. Professional level: It realizes comprehensive pollution weighted risk assessment, probabilistic uncertainty analysis, and quantitative pollution source tracing. The depth of analysis far exceeds that of traditional manual and conventional intelligent solutions, and the judgment conclusions are more in line with the actual situation of water environment governance.

[0105] 4. Quality control level: Achieve credible graded labeling of conclusions, full-dimensional quality control of reports, and autonomous system iteration, completely solving the industry pain points of uncontrollable and unoptimizable quality of intelligent reports, and ensuring that report output is standardized, regulated, and traceable.

[0106] This invention also provides a water quality analysis report generation system that integrates multi-source data and domain knowledge, used to execute the method described above, including: a data acquisition module for accessing satellite remote sensing images, UAV remote sensing images, online monitoring data, and unstructured documents; a data preprocessing module for extracting key information from the unstructured documents using an AI agent, and performing water-land boundary identification and interference removal on the remote sensing images; cleaning and aligning the extracted information with structured data, outputting a unified structured data matrix or geographically labeled tensor data, and performing weighted fusion using a multi-source data credibility fusion algorithm to obtain preprocessed point-like ground monitoring data and area-like remote sensing inversion data; a knowledge base storage module for storing a dedicated domain knowledge base containing water quality standard limits, pollutant environmental behavior, historical cases, and an inversion model library; and an analysis and mining module for receiving the structured data matrix or tensor data, calling the standard limits in the knowledge base for comparison, identifying exceedance information, and further processing. The system performs time series analysis, spatial evolution analysis, and correlation analysis to calculate a comprehensive pollution risk index and output quantitative risk indicators and analysis conclusions. The report synthesis module includes a report framework generation unit and a prompt word engine unit. The report framework generation unit matches a report outline template based on the task type. The prompt word engine unit uses the quantitative risk indicators and analysis conclusions output by the analysis and mining module to drive a large language model to generate professional text descriptions and chart drawing instructions through preset prompt word templates, outputting a visual text and graphic report. The quality verification and output module performs consistency verification on the visual text and graphic report, checking whether the data cited in the report text is consistent with the statistical values ​​of the structured data matrix or tensor data, and optimizing the logical coherence of the visual text and graphic report. Based on human review feedback, the prompt word template is iteratively optimized using an incremental learning mechanism with a forgetting factor. The verified and optimized content is automatically applied to a preset report template to generate a standardized water quality analysis report.

[0107] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for generating water quality analysis reports that integrates multi-source data and domain knowledge, characterized in that, Includes the following steps: S1: Collect satellite remote sensing images, UAV remote sensing images, online monitoring data, and unstructured documents. Use an AI agent to extract key information from the unstructured documents and perform land-water boundary identification and interference removal on the remote sensing images. After cleaning and aligning the extracted information with the structured data, output a unified structured data matrix or geographically labeled tensor data. Use a multi-source data credibility fusion algorithm for weighted fusion to obtain preprocessed point-like ground monitoring data and area remote sensing inversion data. S2: Construct a domain-specific knowledge base containing water quality standard limits, pollutant environmental behavior, historical cases, and inversion model libraries; call the standard limits in the knowledge base, compare the preprocessed point ground monitoring data and area remote sensing inversion data obtained in step S1 with the standard limits, identify the exceedance information, and perform time series analysis, spatial evolution analysis, and correlation analysis to calculate the comprehensive pollution risk index and obtain quantitative risk indicators and analysis conclusions; S3: Match the report outline template according to the task type; using the quantitative risk indicators and analysis conclusions generated in step S2, drive the large language model to generate professional text descriptions and chart drawing instructions through the preset prompt word templates; perform confidence quantification assessment and hierarchical labeling on the conclusions in the generated professional text descriptions to form a visual text and graphic report; S4: Perform consistency verification on the visualization report generated in step S3, check whether the data referenced in the report text is consistent with the statistical values ​​of the structured data matrix or tensor data output in step S1, and optimize the logical coherence of the visualization report. Based on feedback from manual review, the prompt word template is iteratively optimized using an incremental learning mechanism with a forgetting factor; The optimized content is automatically applied to a preset report template to generate a standardized water quality analysis report.

2. The method according to claim 1, characterized in that, In step S1, the AI ​​agent extracts monitoring points, time, indicator values ​​and units from PDF, Word or scanned documents using natural language understanding and pattern recognition technology; the water-land boundary recognition uses a deep learning model based on U-Net or Transformer architecture to extract water pixels and remove cloud layers and interfering pixels.

3. The method according to claim 1, characterized in that, In the multi-source data credibility fusion algorithm described in step S1, the observation value of the i-th data source fusion value for: ; ; ; In the formula: This represents the final merged water quality value at spatial location (x, y); For the spatial adaptive weights of the i-th data source; The real-time dynamic confidence level is calculated by weighting the i-th data source using four-dimensional quality factors; the real-time dynamic confidence level... For remote sensing data sources, the results are obtained through dynamic evaluation of image cloud cover, atmospheric correction quality, inversion model applicability, and deviation between inversion values ​​and adjacent ground monitoring values. For ground monitoring data, the results are obtained through dynamic evaluation of signal stability and deviation from historical data. The coefficient of spatial variation of water quality in a local window at coordinates (x, y) is given. λ represents the maximum coefficient of variation for the entire water body, and λ is the adjustment coefficient in the range of 0 to 1. , , , Normalize the weights for each quality dimension to satisfy + + + =1; , , , These are four quantitative quality scoring factors.

4. The method according to claim 1, characterized in that, The formula for calculating the Comprehensive Pollution Risk Index (CPRI) in step S2 is as follows: ; in, The intensity index reflects the severity of exceeding the standard; This is a trend deterioration index, reflecting the deteriorating trend of water quality indicators; To correlate complexity indices, the combined risks of multiple pollutants occurring together are quantified; The ecotoxicity index reflects the ecological hazard level of pollutants; α, β, γ, and η are weighting coefficients; the... , , and All values ​​are calculated based on the structured data matrix or tensor data output in step S1 and the standard limits, pollutant combination risk weights, and ecotoxicity weights in the knowledge base.

5. The method according to claim 4, characterized in that, The formula for calculating the excess strength index SI is as follows: ; Where N is the total number of indicators, m is the number of indicators exceeding the standard, and max(C j ) represents the maximum concentration of index j within the time period, S j It is its standard limit. The importance weight coefficient of this indicator is obtained from the knowledge base; The trend deterioration index TI is the mean of the sign function of the linear regression slope of the concentrations of all key indicators over time. ; Where k is the number of key metrics, slope j is the slope of the linear regression of the concentration of index j over time within a given period, and sgn is the sign function; The association complexity index CI is: ; Where P is the set of significantly correlated pollutant indicator pairs, r pq Let p be the Pearson correlation coefficient between indicators p and q, and I(p,q) be the risk weight of a specific combination of pollutants obtained from the knowledge base. The ecotoxicity index (ETI) is: ; Where Cj is the measured concentration of pollutant, Sj is the standard limit, and Tj is the ecotoxicity weight of pollutant stored in the knowledge base, so as to realize the risk weighting of highly toxic pollutants under the same exceedance conditions.

6. The method according to claim 1, characterized in that, The preset prompt word templates mentioned in step S3 include trend description templates, spatial distribution feature templates, and conclusion writing templates; By using a multi-prompt word chain to drive a large language model, the excess list, trend statistics, pollution spatial distribution heat map metadata, and related inference conclusions generated in step S2 are input into the corresponding templates to generate professional text descriptions that logically match each chapter.

7. The method according to claim 1, characterized in that, The consistency check in step S4 includes: automatically verifying whether the out-of-standard values, percentages, and spatial statistics cited in the main body of the report are consistent with the statistical values ​​of the structured data matrix or tensor data output in step S1; the logical coherence optimization includes checking whether the conclusions and analysis parts correspond, and whether the spatial analysis conclusions and time change conclusions contradict each other.

8. The method according to claim 1, characterized in that, In step S4, the prompt word template is iteratively optimized based on human review feedback using an incremental learning mechanism with a forgetting factor, as shown in the following formula: ; In the formula: The system's initial basic prompt word template serves as the model's default report generation rules, writing paradigm, and analysis logic baseline. It is the latest prompt word template after iterative optimization, and it is the final basis for the system to generate reports in the future. T is the total number of manual feedback corrections received by the system, representing the total number of all valid correction samples in this iteration cycle. t is the feedback correction time series index variable, with a value range from 1 to T, used to mark the time sequence of each manual feedback. t=1 corresponds to the earliest historical feedback, and t=T corresponds to the latest current feedback. Δt is the prompt word correction vector corresponding to the t-th manual feedback, representing the magnitude and direction of correction of text expression, analysis logic, and conclusion deviation. γ is the forgetting factor, with a value range of (0,1). Through an exponential weighting mechanism, it assigns higher weight to recent feedback and lower weight to distant historical feedback, weakening the impact of old and inefficient correction rules and ensuring the accuracy and timeliness of model iteration.

9. A water quality analysis report generation system integrating multi-source data and domain knowledge, used to perform the method according to any one of claims 1 to 8, characterized in that, include: The data acquisition module is used to access satellite remote sensing images, UAV remote sensing images, online monitoring data, and unstructured documents; The data preprocessing module is used to extract key information from the unstructured documents through an AI agent, and to identify land and water boundaries and remove interference from the remote sensing images. After cleaning and aligning the extracted information with the structured data, it outputs a unified structured data matrix or geographically labeled tensor data. A multi-source data credibility fusion algorithm is used for weighted fusion to obtain preprocessed point ground monitoring data and area remote sensing inversion data. The knowledge base storage module is used to store a dedicated domain knowledge base containing water quality standard limits, pollutant environmental behavior, historical cases, and inversion model libraries; The analysis and mining module is used to receive the structured data matrix or tensor data, call the standard limit values ​​in the knowledge base for comparison, identify the information exceeding the standard, and perform time series analysis, spatial evolution analysis and correlation analysis to calculate the comprehensive pollution risk index and output quantitative risk indicators and analysis conclusions. The report synthesis module includes a report framework generation unit and a prompt word engine unit. The report framework generation unit is used to match a report outline template according to the task type. The prompt word engine unit is used to use the quantitative risk indicators and analysis conclusions output by the analysis and mining module to drive the large language model to generate professional text descriptions and chart drawing instructions through preset prompt word templates, and output a visual text and graphic report. The quality verification and output module is used to perform consistency verification on the visualization report, check whether the data referenced in the report text is consistent with the statistical values ​​of the structured data matrix or tensor data, and optimize the logical coherence of the visualization report. Based on feedback from manual review, the prompt word template is iteratively optimized using an incremental learning mechanism with a forgetting factor; The optimized content is automatically applied to a preset report template to generate a standardized water quality analysis report.

10. The system according to claim 9, characterized in that, The quality verification and output module includes a consistency verification unit, a logic optimization unit, and a format output unit. The consistency verification unit is connected to the data preprocessing module and is used to verify whether the data referenced in the visualization report is consistent with the statistical values ​​of the structured data matrix or tensor data output by the data preprocessing module. The logic optimization unit is used to verify the logical coherence of the conclusions and analysis sections in the visualization report. The format output unit has built-in Word, PDF, or HTML report templates for various institutions, which are used to output the verified and optimized content as a standardized format report with one click. The quality verification and output module includes a consistency verification unit, a logic optimization unit, and a format output unit; The consistency verification unit is connected to the data preprocessing module and is used to verify whether the data cited in the visualization report is consistent with the statistical values ​​of the structured data matrix or tensor data output by the data preprocessing module; the logic optimization unit is used to verify the logical coherence of the conclusions and analysis sections in the visualization report; the quality verification and output module is also used to iteratively optimize the prompt word template based on manual review feedback and using an incremental learning mechanism with a forgetting factor, the optimization formula of which is: ; In the formula: The system's initial basic prompt word template serves as the model's default report generation rules, writing paradigm, and analysis logic baseline. It is the latest prompt word template after iterative optimization, and it is the final basis for the system to generate reports in the future. T is the total number of manual feedback corrections received by the system, representing the total number of all valid correction samples in this iteration cycle. t is the feedback correction time series index variable, with a value range from 1 to T, used to mark the time sequence of each manual feedback. t=1 corresponds to the earliest historical feedback, and t=T corresponds to the latest current feedback; Δt is the prompt word correction vector corresponding to the t-th manual feedback, representing the magnitude and direction of correction for text expression, analysis logic, and conclusion bias; γ is the forgetting factor, with a value range of (0,1). Through an exponential weighting mechanism, it assigns higher weight to recent feedback and lower weight to distant historical feedback, weakening the impact of old and inefficient correction rules and ensuring the accuracy and timeliness of model iteration; The format output unit has built-in multiple organization-specific Word, PDF or HTML report templates, which are used to output the verified and optimized content as a standardized format report with one click.