Water plant whole-process intelligent monitoring and control method and system based on large model

By constructing a dynamic knowledge graph and adaptively adjusting thresholds, the problem of response delay in the water plant monitoring system was solved, enabling real-time monitoring and optimization, and improving the water plant's intelligence level and anti-interference capability.

CN122151729APending Publication Date: 2026-06-05SHANDONG FENGSHI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG FENGSHI INFORMATION TECH CO LTD
Filing Date
2026-01-28
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing water plant monitoring systems lack an adaptive optimization mechanism for real-time feedback data, resulting in response delays, difficulty in adapting to water quality fluctuations and changing pollutant migration paths, and inability to identify new pollutants in a timely manner, thus affecting water quality safety.

Method used

A dynamic knowledge graph based on a large model is constructed. Through entity recognition, relation extraction, and path inference, combined with a prediction-actual path deviation feedback mechanism, the knowledge graph can be adaptively adjusted, and the threshold can be dynamically adjusted to adapt to complex working conditions.

Benefits of technology

It enables real-time response and adaptive optimization of the water plant monitoring system, improves the system's adaptability and anti-interference ability to complex operating conditions, reduces the risk of water quality accidents, and reduces the waste of operation and maintenance resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122151729A_ABST
    Figure CN122151729A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of big model-based water plant whole process intelligent monitoring control method and system, belong to water plant process intelligent control technical field.By real-time acquisition water plant pipeline node sensing data flow and outlet effect feedback data, when monitoring value deviates from preset first threshold, automatically associate space-time information to generate abnormal data flow;Further, after the abnormal data vectorization, input knowledge graph representing pollutant-process unit-reaction path, reasoning generates pollutant migration prediction path;Based on subsequent monitoring data, reconstruct the actual migration path of pollutant, by comparing prediction path and actual path to generate deviation index;If deviation exceeds second threshold, then link deviation data, real-time path and outlet feedback data, trigger the adaptive adjustment of knowledge graph.The present application is analyzed by the deviation of prediction path and actual path, combined with outlet effect feedback data, dynamically adjust the association rules in knowledge graph, realize its adaptive optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and system for intelligent monitoring and control of the entire process of a water plant based on a large model, belonging to the field of intelligent control technology for water plant processes. Background Technology

[0002] Intelligent management and control of the entire water plant process refers to the use of modern information technologies such as the Internet of Things, big data, and artificial intelligence to create a unified control system for the water plant. This system can comprehensively perceive the operational status of the entire water plant, deeply analyze massive amounts of data to obtain underlying patterns, make scientific decisions, provide optimal operating plans, and automatically control equipment to complete production tasks. Ultimately, it achieves a revolutionary transformation in water plant production from experience-driven to data-driven, and from partial automation to global intelligence.

[0003] In modern water treatment processes, real-time monitoring and source tracing of pollutants within pipelines are crucial for ensuring water quality safety. Currently, most water plants' monitoring systems rely on manual intervention for analysis and adjustment, lacking adaptive optimization mechanisms based on real-time feedback data. This leads to response delays and consequently, inefficiency. Traditional water plant monitoring systems primarily depend on fixed thresholds and static knowledge bases for anomaly detection and pollution source tracing, making it difficult to adapt to dynamic conditions such as water quality fluctuations and changing pollutant migration paths. When influent water quality changes abruptly or new pollutants appear, static rules fail to accurately reflect actual behavior, resulting in delayed system responses, frequent false alarms, and difficulty in tracking the true path of pollutants in real time. For example, when nickel-containing wastewater mixes with the influent of a wastewater treatment plant, although sensor data may be abnormal, the system, due to fixed thresholds and an outdated knowledge base, fails to identify and issue a timely warning, leading to delays in process adjustments and increasing the risk of effluent exceeding standards. Without addressing this issue, the system will be unable to respond promptly to dynamic changes, potentially delaying treatment and impacting water quality safety. Furthermore, it is prone to generating continuous false alarms during the recovery period, resulting in wasted operational resources. Summary of the Invention

[0004] The purpose of this invention is to overcome the above-mentioned shortcomings and provide a method for intelligent monitoring and control of the entire process of a water plant based on a large model. By constructing a dynamic knowledge graph that integrates multi-source process knowledge, the method performs vectorized matching and path deduction on abnormal data and introduces a "prediction-actual" path deviation feedback mechanism. The knowledge graph is dynamically adjusted based on the deviation and effluent feedback to form a fully automated control system with continuous learning capabilities.

[0005] The technical solution adopted in this invention is as follows: The intelligent monitoring and control method for the entire process of water plants based on a large model includes the following steps: S1. Based on domain text data and historical operation data, a large model is used to perform entity recognition and relationship extraction processing on the data, and a knowledge graph is constructed according to the relationship between water plant pollutants, process units and reaction paths; S2. Multiple first data acquisition nodes are set at the raw water inlet, the connecting pipes of each process unit, and the inlet of the clear water tank. A second data acquisition node is set at the pipe outlet. The second data acquisition node maintains a safe distance from the first data acquisition node. The first data acquisition node collects and monitors the data stream, and the second data acquisition node collects and feeds back the data. S3. Obtain the monitoring data stream from the previous moment, set a first threshold interval, determine whether the monitoring data stream deviates from the first threshold interval, and if so, associate the monitoring data stream with the timestamp and spatial location to generate an abnormal data stream; S4. Preprocess the abnormal data and extract key feature data, input the key feature data into the embedding model for encoding and vectorization, generate query vector data, input it into the pre-built knowledge graph for matching and reasoning, and generate prediction path data containing at least one pollutant at the current moment, as well as prediction effect feedback data associated with the prediction path data. S5. Obtain the monitoring data stream and effect feedback data at the current moment, perform spatiotemporal alignment and cleaning on the monitoring data stream, generate standardized time-series trajectory data, and use time-series association rule mining and state reasoning algorithms to deduce the actual migration sequence and transformation product sequence of pollutants between process units in a chain-like manner to obtain real-time path data. S6. Compare the real-time path data with the predicted path data to generate path deviation index data, and calculate the difference between the current effect feedback data and the predicted effect feedback data as the effect feedback data difference. S7. Set a second threshold and determine whether the path deviation index data is greater than the second threshold. If it is, trigger the adaptive adjustment of the knowledge graph based on the difference between the path deviation index data, real-time path data, and effect feedback data, and use the updated knowledge graph for path prediction in the next moment. Otherwise, do not process it.

[0006] In the above method, the domain text data mentioned in step S1 includes scientific literature data, process manual data, and chemical safety data sheets, while the historical operation data includes historical operating condition record data, manual operation log data, and laboratory test report data.

[0007] A large model is used to perform entity recognition and relation extraction on domain text data to generate initial knowledge graph data represented in tuple form; feature extraction and event reconstruction are performed on historical operation data to generate empirical verification data represented in process event chain form; the large model is used to fuse and resolve conflicts between the initial knowledge graph data and the empirical verification data to generate the basic topology structure of the knowledge graph data; attribute data and confidence data are added to the entities and relations in the basic topology structure using the large model to generate the knowledge graph.

[0008] In step S3, when the target sensor has uncollected data, the parameters collected by the associated auxiliary sensing unit are input into the trained prediction model, and the predicted monitoring data stream is used to replace it.

[0009] The adjustment method for the first threshold range is as follows: Dynamic sensitivity coefficients are generated based on the standard deviation of the monitoring data stream from the previous monitoring cycle and the types of dominant pollutants in the knowledge graph. The first threshold interval at the current moment is calculated by multiplying the first threshold interval of the previous monitoring period by the dynamic sensitivity coefficient. The dominant pollutant refers to the pollutant that poses the highest risk to achieving the effluent quality standard, calculated by a risk assessment algorithm based on predefined pollutant toxicity weight data in the knowledge graph, pollutant concentration data in the current monitoring data stream, and current process unit data.

[0010] After generating the anomalous data stream, the process also includes an anomaly confidence assessment step, which specifically includes: Obtain the historical monitoring data stream corresponding to the abnormal data stream and its corresponding effect feedback data; Calculate the mean and standard deviation of historical data streams from the same period, calculate the duration and magnitude of deviation of the current abnormal data stream from the mean and standard deviation, and combine them to generate an anomaly confidence score; If the anomaly confidence score is lower than the preset confidence threshold, the current abnormal data stream will be marked as a confidence anomaly, and logs and warnings will be issued.

[0011] The key feature data in step S4 includes pollutant type feature data, concentration gradient data, and spatial location sequence data. The pollutant prediction path data is obtained by performing similarity matching calculations between the query vector data and pollutant entity vectors in the knowledge graph data to identify one or more most likely target pollutant entities. Starting from the target pollutant entity, all relationship paths in the knowledge graph data with that entity as the head node are traversed. The occurrence probability of each traversed relationship path is weighted and evaluated to generate a comprehensive occurrence probability score for that path. These comprehensive occurrence probability scores are then sorted from high to low. This yields the probability order of various possible pollutant migration and transformation paths under the current abnormal situation, thus obtaining the prediction path data.

[0012] The process of generating path deviation index data in step S6 is as follows: extract the key node sequence and corresponding expected pollutant concentration data recorded in the predicted path data; extract the actual node sequence and corresponding measured pollutant concentration data recorded in the real-time path data; compare the consistency between the key node sequence and the actual node sequence in turn to generate the path consistency index; calculate the relative deviation between the expected pollutant concentration data and the measured pollutant concentration data at the same node in turn to generate the concentration deviation index. By integrating the path fit index and the concentration deviation index, a comprehensive path deviation index data is generated.

[0013] In step S7, based on path deviation index data and effect feedback data, the target entity or target relationship to be corrected in the knowledge graph data is located; based on real-time path data and effect feedback data, updated data for the target entity or target relationship is generated. The updated data is applied to the knowledge graph to achieve online adaptive adjustment of the knowledge graph.

[0014] The adjustment method for the second threshold is as follows: Obtain the average consistency between the predicted path data and the real-time path data at the previous moment, the standard deviation of the monitoring data stream, and the compliance rate of the effect feedback data; A dynamic weighting coefficient is assigned to each of the average fit, standard deviation, and compliance rate; The system's overall performance index is calculated by multiplying each data point by its corresponding dynamic weight coefficient and then summing the results. The system's overall performance index is input into a predefined inverse mapping function to obtain the adjusted second threshold.

[0015] The dynamic weight coefficients are configured according to the current optimization objective. If the optimization objective is to improve the compliance rate of the effect feedback data, the weight of the compliance rate is increased; if the primary objective is to improve the accuracy of the prediction model, the weight of the average fit is increased.

[0016] The process of constructing the reverse mapping function is as follows: Obtain the system's overall performance indicators from the historical operating database and the corresponding second threshold; The correlation analysis and regression fitting of the system's comprehensive performance index and the second threshold are performed, and the generated function parameters and their mapping relationship are used as the inverse mapping relationship function.

[0017] Another objective of this invention is to provide a large-scale intelligent monitoring and control system for the entire water plant process, comprising: The data acquisition module sets up multiple first acquisition nodes at the raw water inlet, the connecting pipes of each process unit, and the inlet of the clear water tank, and sets up a second acquisition node at the pipe outlet. The second acquisition node maintains a safe distance from the first acquisition node. The first acquisition node collects and monitors the data stream, and the second acquisition node collects and feeds back the data. The data acquisition module is used to acquire the monitoring data stream from the previous moment; and to acquire the monitoring data stream and effect feedback data from the current moment. The first threshold comparison and determination module is used to determine whether the monitoring data stream at the previous moment deviates from the first threshold range. If so, the monitoring data stream is associated with the timestamp and spatial location to generate an abnormal data stream. The knowledge graph matching module is used to convert the abnormal data stream into vector data and input it into a preset knowledge graph for matching and reasoning, generating prediction path data containing at least one prediction path data at the current moment, as well as prediction effect feedback data associated with the prediction path data. The knowledge graph is used to characterize the relationship between pollutants, process units and reaction paths. The real-time path generation module is used to reconstruct the actual migration path of pollutants based on the monitoring data stream at the current moment and generate real-time path data. The path deviation generation module is used to compare the real-time path data with the predicted path data, generate path deviation index data, and calculate the difference between the current effect feedback data and the predicted effect feedback data as the effect feedback data difference. The second threshold comparison and determination module is used to determine whether the path deviation index data is greater than the preset second threshold. The knowledge graph adaptive update module adjusts the knowledge graph adaptively based on the difference between path deviation index data, real-time path data, and effect feedback data for values ​​exceeding a preset second threshold.

[0018] The first acquisition node includes at least a set of auxiliary sensing units and a target sensing unit; The target sensing unit is used to collect monitoring data streams, and the auxiliary sensing unit is used to collect parameters associated with the monitoring data streams; If the target sensing unit fails to acquire data at the current moment, the parameters acquired by the auxiliary sensing unit will be input into the preset prediction model to generate the monitoring data stream for the current moment.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention analyzes the deviation between the predicted path and the actual path, combines it with feedback data on water treatment effect, and dynamically adjusts the association rules in the knowledge graph to achieve adaptive optimization. This improves the system's adaptability to complex operating conditions, automates the entire monitoring process from monitoring and reasoning to adjustment, and enhances the intelligence level and anti-interference capability of water plant operation. Attached Figure Description

[0020] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a flowchart of the operational logic for making subsequent judgments by monitoring data streams and effect feedback data in this invention. Figure 3 This is a flowchart of a method for replacing failed monitoring data streams in this invention. Figure 4 This is a structural block diagram of the intelligent monitoring and control system for the entire water plant process based on a large model, as proposed in this invention. Detailed Implementation

[0021] The present invention will be further described below with reference to specific embodiments.

[0022] Example 1: A method for intelligent monitoring and control of the entire water plant process based on a large model, including the following steps (see...). Figure 1 )as follows: S1. Based on domain text data and historical operation data, a large model is used to perform entity recognition and relationship extraction processing on the data, and a knowledge graph is constructed according to the relationship between water plant pollutants, process units and reaction paths.

[0023] The process of knowledge graph construction is as follows: Acquire domain text data from external data sources and historical operating data from within the system. Domain text data includes scientific literature data, process manual data, and chemical safety data sheets. Historical operating data includes historical operating condition records, manual operation logs, and laboratory test reports. The large model automatically identifies and labels the research topics of scientific literature data, the key points of process manual data, and the hazard classification of chemical safety data sheets based on predefined domain categories and data frameworks. The data framework is used to define and standardize the data organization method. Large models are used to perform entity recognition and relation extraction on domain text data to generate initial data of knowledge graph in the form of tuples. Feature extraction and event reconstruction are performed on historical operation data to generate empirical verification data in the form of process event chains. For example, inputting preprocessed domain text data into a large model and performing the following operations: Entity recognition and classification: Identify and extract entities of a specified type from text fragments. Entity types include, but are not limited to, chemical equipment, material composition, process parameters, and operating units. Relationship extraction: Based on a deep understanding of text semantics, the relationships between identified entities are determined and extracted. Relationship types include spatial relationships, logical relationships, input-output relationships, and conditional constraint relationships. Structured output: The extracted results are standardized and output in a preset tuple format to generate initial knowledge graph data in the format of (head entity, relation, tail entity).

[0024] Structured historical operating condition records, unstructured manual operation logs, and laboratory test reports are converted into serialized text descriptions that can be parsed by large models. By performing time-series analysis and contextual reasoning on serialized text using large models, the hidden changes in working conditions, operational intentions, and system response events in the data can be identified. Based on the inductive and deductive capabilities of large models, discrete data points are reconstructed into a chain of process events with causality and temporality, and output in this form as empirical verification data.

[0025] By using a large model, the initial data of the knowledge graph and the empirical verification data are fused and conflict-resolved to generate the basic topological structure of the knowledge graph data.

[0026] For example, initial data from text and empirical data from running data can be fed together into a large model; By leveraging the knowledge reasoning and contradiction detection capabilities of large models, we can identify points of consistency and conflict in the descriptions of the same entity or relationship in two types of data sources. For conflicting knowledge, the large model acts as an arbitrator, generating conflict resolution suggestions or outputting a revised final value based on its built-in domain prior knowledge and an assessment of the reliability of the data source. After consistency verification and conflict resolution are completed, the large model assists in completing entity alignment and relationship fusion, generating a unified and conflict-free basic topology structure for the knowledge graph data.

[0027] By utilizing the entity and relationship attribute data and confidence data in the basic topological structure of the large model, the knowledge graph data is constructed.

[0028] For example, attribute appending: Based on the generative capabilities of the large model, it automatically supplements the entities and relationships in the constructed topology with rich attribute fields. The attribute data includes quantified values ​​extracted from the original data and summary descriptive text generated by the large model. Confidence Assignment: The large model self-evaluates each piece of knowledge it participates in generating (including extracted entities, relationships, reconstructed events, and supplementary attributes), and outputs a quantitative confidence score. This score is based on the large model's internal judgment of the degree of certainty in the generation process. The final output is a knowledge graph data that includes entity, relation, attribute, and confidence data, which has undergone augmented data processing and can be directly used for downstream applications.

[0029] The process of adding more valuable information, context, or derived data to raw data to improve its quality and usability. Specific methods include: Supplementing missing attributes: Retrieve core data from one data source and then supplement it with detailed information from another data source.

[0030] Add derivative information: Generate new information through calculation or reasoning.

[0031] Connect related data: Link data scattered from different sources together to form a more complete view.

[0032] S2. Multiple first data acquisition nodes are set up at the raw water inlet, the connecting pipes of each process unit, and the inlet of the clear water tank. A second data acquisition node is set up at the pipe outlet. The second data acquisition node maintains a safe distance from the first data acquisition nodes. The first data acquisition nodes collect and monitor the data stream, and the second data acquisition node collects and provides feedback on the data. The first data acquisition node is deployed according to the pipeline configuration, and the second data acquisition node is deployed at the water outlet of the pipeline and maintains a safe distance from the first data acquisition node.

[0033] The first sampling node includes the raw water inlet, the connecting pipes and valves between various process units such as coagulation, sedimentation, and filtration, and the clear water tank inlet; the second sampling node is the outlet. When the pipe is circular, air tends to accumulate at the highest point of the pipe, forming air pockets, and sediment and impurities tend to accumulate at the lowest point of the pipe, which can easily lead to inaccurate flow and pressure measurements. In severe cases, it can cause sensor blockage, resulting in sampling failure and measurement errors. Therefore, the sampling node is usually set in the center of the pipe.

[0034] The monitoring data stream mainly reflects the real-time status of the process. For example, it can be basic water quality parameters such as pH value, water temperature, dissolved oxygen, conductivity, and turbidity; or hydraulic parameters such as flow rate and pressure; or pollutant concentration data such as the real-time concentration of target pollutants such as nickel, ammonia nitrogen, nitrate, and heavy metals. The concentration of some pollutants can be directly measured by target sensors (such as online analyzers).

[0035] The effect feedback data mainly reflects the final treatment effect of the water. It can be the effluent quality compliance rate: the actual compliance status of core assessment indicators, such as whether the total nickel concentration, COD, ammonia nitrogen, turbidity, etc. meet the discharge standards. The compliance rate is usually based on continuous sampling statistics. Or it can be the final concentration of key pollutants: the measured concentration value of target pollutants in the effluent, such as the multiple of nickel concentration exceeding the standard, residual chlorine, microbial indicators, etc. Or it can be the treatment efficiency feedback: the removal rate calculated based on the comparison of influent and effluent concentrations, used to verify the actual treatment effect of the predicted path.

[0036] S3. Obtain the monitoring data stream and effect feedback data from the previous moment, set a first threshold interval, and determine whether the monitoring data stream deviates from the first threshold interval (e.g., ...). Figure 2 Then, the monitoring data stream is associated with a timestamp and spatial location, and contextual environment data is appended to generate an abnormal data stream: The monitoring data stream and effect feedback data from the previous moment are acquired. The monitoring data stream is acquired by the sensing unit located at the first acquisition node, and the effect feedback data is acquired by the sensing unit located at the second acquisition node. For example, it continuously receives monitoring data streams generated by physical sensors installed at key nodes of the water plant pipeline. Each data point in the monitoring data stream contains timestamp data, spatial location identifiers, and monitoring values.

[0037] It continuously receives feedback data on the effectiveness of water quality monitoring points generated by physical sensors installed at the water plant's pipeline outlet, such as the actual compliance rate data of key water quality parameters.

[0038] like Figure 3 As shown, the method and process for replacing monitoring data streams that failed to be collected are introduced; The sensing unit at the first acquisition point includes at least a set of auxiliary sensing units that serve as input features of the model, and a target sensing unit that serves as the model's prediction target. The target sensing unit refers to a sensor that is fragile, expensive, or difficult to maintain directly, and the data it monitors is the monitoring data stream to be used. The auxiliary sensing unit refers to a reliable, low-cost, or easy-to-maintain sensor used to monitor parameters that have a physical or chemical correlation with the monitoring data stream. When there is uncollected data in the monitoring data stream of the target sensing unit, the parameters collected by the auxiliary sensing unit are input into the prediction model to obtain the predicted monitoring data stream. Replace the target sensing unit's monitoring data stream with the predicted monitoring data stream.

[0039] The process of constructing the prediction model used is as follows: (1) Obtain the parameters of multiple auxiliary sensing units and the corresponding monitoring data stream of the target sensing unit from the historical monitoring database; (2) Divide the preprocessed parameters and monitoring data stream into training set and validation set; The preprocessing process for parameters and monitoring data streams is as follows: Identify, remove, or repair data points in parameters and monitoring data streams that exhibit signal anomalies, communication interruptions, or are clearly outside the physically reasonable range; The parameters and monitoring data streams that have been identified, removed, or repaired will be time-aligned. The parameters and monitoring data streams after time-series alignment are normalized. (3) The parameters of the auxiliary sensing units in the training set are used as input data to the neural network model, and the monitoring data stream of the corresponding target sensing unit is used as the target output data. (4) The internal parameters of the neural network model are iteratively adjusted through the backpropagation algorithm or the gradient boosting algorithm, with the goal of reducing the value of the loss function between the model's predicted output data and the target output data. After each training iteration, the model in the current state is evaluated using the validation set, and its prediction accuracy and mean absolute error are calculated. When the accuracy no longer improves or the mean absolute error no longer decreases, the training process is stopped, and the trained model is output as the prediction data stream model. For example, suppose we have a historical monitoring dataset containing observation data at T time steps. The preprocessed parameters of the auxiliary sensing unit obtained from the historical database can be represented as matrix x, and the monitoring data stream of the target sensing unit can be represented as vector y. , ; in, Indicates at time step No. Readings from the auxiliary sensors Indicates at time step Readings from the target sensor It refers to the number of auxiliary sensing units.

[0040] The preprocessed data is divided into a training set and a validation set. The parameters of the auxiliary sensing unit in the training set and the monitoring data stream of the target sensing unit correspond one-to-one. ; ; set up This is the neural network model we need to train, and its parameters are: The training objective of the model is to learn a mapping: ; Make the predicted value As close as possible to the true value .

[0041] By minimizing the loss function To adjust model parameters Here, the loss function is the mean squared error (MSE): ; in, It is the size of the training set. The parameter θ is iteratively updated using the gradient descent algorithm, with a learning rate of η. , in, The loss function L represents the loss function with respect to parameter θ. old The gradient at that point.

[0042] After each training iteration, performance metrics are calculated using the validation set: Mean Absolute Error (MAE): ; in, This is the size of the validation set.

[0043] accuracy (Here, accuracy can be defined as the proportion of predicted values ​​falling within the allowable error range): ; in, It is an indicator function that returns 1 when the condition is true, and 0 otherwise; This is the preset error tolerance.

[0044] set up Let be the number of iterations. Training stops when the loss on the validation set no longer decreases, i.e., when the following condition is met: ; in, Indicates the first Loss on the validation set of the next iteration. It is a preset value that indicates how many consecutive rounds of observation will continue without improvement before stopping.

[0045] After training stops, the final model is output. As a prediction model: ; in, These are the parameters that perform best on the validation set during training.

[0046] Periodically obtain the model's average confidence score when processing recent data; If the average confidence score data consistently falls below a preset threshold (e.g., the preset threshold is 0.8), the model will be updated or retrained. Data with a confidence score consistently below 0.8 will be removed, and the removed data will be input into the model for retraining.

[0047] After the prediction model generates the monitoring data stream for the current moment, it performs deviation calculations on the monitoring data stream collected by the target sensing unit at the previous or next moment, and determines whether the deviation exceeds the preset deviation threshold (e.g., setting the deviation threshold to 0.85). If so, it triggers a target sensing unit fault warning and prioritizes the parameters collected by the auxiliary sensing unit to generate the required monitoring data stream through the prediction model.

[0048] When the target sensing unit's monitoring data stream is missing, the prediction model is driven by the auxiliary sensing unit parameters to generate a replacement data stream. The data stream generated by the prediction model is used to replace the missing data of the target sensing unit in real time, avoiding the interruption of the monitoring chain due to a single point of failure. For example, the loss of data at key nodes in a pipeline may lead to misjudgment of the pollution path. This reduces the absolute dependence on a single sensor and ensures the continuity of the entire process even in the event of hardware failure.

[0049] This invention sets up an auxiliary sensing unit and establishes a predictive model to automatically switch to a backup data generation mode when the target sensor fails, forming a redundant data protection mechanism. This avoids delays in control decisions caused by data interruption and improves the fault tolerance and continuity of the monitoring system.

[0050] Following the step of acquiring the monitoring data stream from the previous moment, the process also includes sensor collaborative verification and redundancy completion steps, specifically including: Acquire monitoring sub-data streams collected simultaneously by multiple sensing units of the same type deployed at the same first acquisition node; Calculate the dispersion index between each monitoring sub-data stream. If the dispersion index exceeds the preset dispersion threshold, it is determined that the sensor at that node is experiencing interference or abnormality. Based on the dispersion index, the monitoring sub-data stream with the largest deviation is selected, and the deviation data is replaced with the mean or median of the remaining monitoring sub-data streams, and the corrected monitoring data stream is output.

[0051] Specifically, the dispersion index can be the standard deviation, the range (the difference between the maximum and minimum values), or the mean absolute deviation. Since the standard deviation is sensitive to outliers and its calculation is slightly more complex, and the range is determined by only two extreme values, it is easily affected by a single faulty sensor and has poor stability, the mean absolute deviation is often chosen in engineering practice. The mean and standard deviation of the dispersion index sequence are calculated and statistically analyzed. The preset dispersion threshold can be the sum of the mean and three times the standard deviation.

[0052] The specific process for identifying abnormal data based on the acquired monitoring data stream is as follows: (a) Compare the monitoring values ​​in the real-time monitoring data stream with the dynamic operating condition baseline range data trained based on historical data in real time; When the monitoring data at a specific spatial location continuously deviates from the dynamic operating condition baseline range and reaches a set first threshold (e.g., the first threshold is set to 0.8), or when the data quality identifier indicates that the data is abnormal, an anomaly flag is triggered. Extract all monitoring data related to the specific spatial location within the specified time period, associate the timestamp and spatial location, package and attach contextual data to generate an anomaly data packet.

[0053] Specifically, the contextual environment data should include: temporal context, i.e., the time and duration of the anomaly; spatial context, i.e., the specific location of the anomaly and its position in the process chain; process status context, i.e., the operating status of the entire water plant process system at the time of the anomaly; environmental parameter context, i.e. the environmental parameters related to the anomaly at the time of the anomaly; and associated data context, i.e. other data sources that may be related to the anomaly event.

[0054] The purpose of this contextual data is to provide the complete background of the anomalous event, helping to make subsequent predictions and inferences more accurate. It is appended to the anomalous data stream to form a more complete information package for subsequent knowledge graph matching and path prediction.

[0055] (b) Anomaly confidence assessment: Obtain the historical monitoring data stream corresponding to the abnormal data stream and its corresponding effect feedback data; Calculate the mean and standard deviation of historical data streams from the same period, calculate the duration and magnitude of deviation of the current abnormal data stream from the mean and standard deviation, and combine them to generate an anomaly confidence score; If the anomaly confidence score is lower than the preset confidence threshold, the current abnormal data stream will be marked as a confidence anomaly, and logs and warnings will be issued.

[0056] Deviation : ; in, This represents the mean. It represents the standard deviation.

[0057] Deviation time : ; in, Indicates the statistical continuous time points >1.5 (i.e., the duration exceeding 1.5 standard deviations) Indicates the total duration of the abnormal data stream.

[0058] Anomaly confidence calculation: ; in, , Represents the weighting coefficient, and After obtaining the anomaly confidence level, the min-max method is used to normalize it.

[0059] The dynamic adjustment process of the first threshold specifically includes: Real-time acquisition of monitoring data streams from the previous monitoring period, and calculation of their standard deviation data. Simultaneously, the dominant pollutant entities identified in the current period are queried from the knowledge graph, and their preset toxicity weight coefficients are obtained. ; standard deviation data With toxicity weighting coefficient Input to dynamic sensitivity coefficient function A comprehensive calculation is performed to generate a dynamic sensitivity coefficient. Its mathematical expression is: ; Among them, dynamic sensitivity coefficient Compared with standard deviation data It shows a positive correlation with the toxicity weighting coefficient. There is a negative correlation; The first threshold used in the previous cycle With dynamic sensitivity coefficient Multiply, calculate and output the first threshold updated at the current time. Its mathematical expression is: ; For example, the interval (a, b) is updated to ( b ), The first threshold updated at the current moment is used for real-time data comparison to determine if there are any anomalies.

[0060] When the dynamic sensitivity coefficient is greater than 1, the new set threshold is relaxed relative to the basic threshold; when the dynamic sensitivity coefficient is less than 1, the new set threshold is tightened relative to the basic threshold.

[0061] The dominant pollutant refers to the pollutant that poses the highest risk to achieving the effluent quality standard, calculated by a risk assessment algorithm based on predefined pollutant toxicity weight data in the knowledge graph, pollutant concentration data in the current monitoring data stream, and current process unit data.

[0062] Traditional systems often rely on batch testing or manual review, which leads to delays. In contrast, this invention acquires monitoring data streams and effluent feedback data from pipeline acquisition nodes in real time. When the data deviates from the normal operating range, i.e., the preset first threshold range, it immediately identifies the anomaly. By associating timestamps and spatial locations, it shortens the response window for pollution events. Its abnormal data stream generation mechanism can trigger alarms within seconds, reducing the risk of water quality accidents, minimizing harm to public health, and saving emergency costs.

[0063] Compared to existing technologies, traditional methods use fixed threshold ranges, which cannot adapt to fluctuations in water quality and changes in pollutant types, easily leading to false alarms or missed alarms. This invention introduces a dynamic sensitivity coefficient, combining data fluctuation characteristics with the toxicity effects of pollutants, to achieve real-time dynamic adjustment of the threshold range. This avoids the response lag of fixed thresholds in sudden pollution events and prevents frequent false alarms caused by routine water quality fluctuations.

[0064] S4. Preprocess the abnormal data and extract key feature data. Input the key feature data into the embedding model for encoding and vectorization to generate query vector data. Input the query vector data into a pre-built knowledge graph for matching and reasoning to generate pollutant prediction path data containing at least one pollutant prediction path data at the current moment, as well as prediction effect feedback data associated with the prediction path data. Abnormal data packets are cleaned and standardized preprocessed; key feature data are extracted from the preprocessed data, including pollutant type feature data, concentration gradient data, and spatial location sequence data; the key feature data are input into the embedding model for encoding and vectorization to generate a machine-readable, high-dimensional query vector data.

[0065] The similarity matching calculation between the query vector data and the pollutant entity vectors in the knowledge graph data is performed to identify one or more most likely target pollutant entities. Starting with the target pollutant entity, traverse all relationship paths in the knowledge graph data that have that entity as the head node. By combining the real-time environmental condition data contained in the abnormal data packets, a weighted evaluation calculation is performed on the occurrence probability of each traversed relationship path.

[0066] The prediction performance feedback data and prediction path data are generated synchronously, as follows: Pre-defined efficiency attributes for the knowledge graph: During knowledge graph construction (step S1), the historical average removal rate (η) or theoretical conversion rate attribute is attached to each "reaction path" relationship. For example, the relationship "(nickel, in, coagulation tank)" is attached with the attribute {average removal rate: 0.45}.

[0067] Path generation and effect calculation are performed simultaneously: For each candidate predicted path generated by knowledge graph matching and reasoning, the system performs the following calculations: Input: The ordered reaction sequence of the path, and the initial contaminant concentration C0 extracted from the anomalous data stream.

[0068] Chain-based simulation calculation: Starting from the initial pollutant concentration C0, each node is processed sequentially according to the path; the removal rate corresponding to the reaction path of the current node is read from the knowledge graph; and the concentration after passing through the node is calculated.

[0069] Proceed to the next node for calculation in sequence.

[0070] Output: The predicted final concentration is obtained when the simulation reaches the end of the path. Further calculations can be performed to determine the predicted effluent compliance status.

[0071] Data Association and Output: The calculated prediction effect feedback data is appended as a set of attributes to the corresponding prediction path data object and output together. For example, a record is: {Path: Inlet → Coagulation Tank → Sedimentation Tank, Association Effect: {Predicted Final Nickel Concentration: 0.38 mg / L, Predicted Exceedance Multiple: 2.8}}.

[0072] Example: Assume abnormal data indicates an initial nickel concentration at the inlet, C0 = 1.0 mg / L. The system matching deduces two candidate paths: Path A: Inlet → Coagulation tank (η1=0.45) → Sedimentation tank (η2=0.30).

[0073] The calculated value is: A = 1.0 * (1 - 0.45) * (1 - 0.30) = 0.385 mg / L. If the standard limit is 0.1 mg / L, the predicted exceedance multiple is 2.85.

[0074] Path B: Inlet → Equalization tank → Biological treatment tank (η3=0.05).

[0075] The calculation yields: B≈0.95mg / L, with a predicted exceedance multiple of 8.5.

[0076] The system outputs the two paths mentioned above and their respective associated prediction results. The comparison shows that path A has a significantly better prediction effect, providing a direct basis for prioritizing process control to guide pollutants along path A.

[0077] Specifically, the implementation process of the probability-weighted evaluation of relationship paths is as follows: (1) Data Acquisition: Input Data I is real-time environmental condition data. The contextual information attached to the abnormal data stream is extracted to obtain a set of real-time environmental condition data characterizing the current process status, which may include, but is not limited to: influent flow rate, water temperature, pH value, dissolved oxygen concentration, and real-time dosage of key agents (such as coagulants); Input Data II is a set of candidate relationship paths, which receives the output from the knowledge graph traversal module, i.e., a set of candidate relationship paths associated with abnormal pollutants. Each path stores its typical reaction conditions, historical occurrence frequency, and path confidence in the knowledge graph.

[0078] (2) Data processing (weighted evaluation calculation): For each path in the candidate relation path set, perform the following parallel computation and fusion steps: Condition matching degree calculation: Real-time environmental condition data is compared with the typical reaction conditions pre-stored in the knowledge graph for this path. For example, if the typical reaction condition for a certain path requires "pH>7.5", the degree of conformity between the current real-time pH value and this condition is calculated. A condition matching degree score is calculated using a preset matching degree function (such as a threshold-based piecewise function or a continuous function).

[0079] Historical frequency weight extraction: The historical occurrence frequency attribute data of the path is directly read from the knowledge graph. This data is based on the statistics of historical operation data and reflects the probability of the path appearing under similar working conditions in the past, which serves as the basis for weighting.

[0080] Path confidence acquisition: The path confidence attribute data is directly read from the knowledge graph. This data combines the reliability of the knowledge source and historical verification.

[0081] Weighted fusion: The multiple indicators calculated and obtained above are weighted and fused to generate a comprehensive probability score for the occurrence of this path. The calculation formula can be expressed as: Overall probability score = F(condition matching score, historical occurrence frequency, path confidence); Where F is a preset fusion function, such as weighted summation, and the weights of each input item can be obtained from historical data.

[0082] (3) Data output: Summarize all candidate paths and their corresponding comprehensive occurrence probability scores; sort them from high to low according to the comprehensive occurrence probability scores; output the probability distribution data of the sorted paths. This data clarifies the probability order of occurrence of various possible pollutant migration and transformation paths under the current abnormal conditions.

[0083] All evaluated paths are sorted according to their occurrence probability to generate path probability distribution data; The above data is summarized and organized to obtain the predicted path data.

[0084] For example, pollutant pathway prediction: Data definition: Abnormal data stream, containing tuples ,in, For monitoring values, For timestamps, For spatial location, This is a vector representing real-time environmental conditions. Knowledge graph, where is a set of entities (pollutants, process units). For a set of relations ( (relation type) Attribute functions attached to entities and relations; The first in the knowledge graph Vector representation of each pollutant entity. .

[0085] Feature extraction and vectorization: ; in, An embedding model used to vectorize data. This is the generated query vector data.

[0086] Entity similarity matching: ; in, The cosine similarity function is used. To preset the matching threshold, For the matched set of target pollutant entities, .

[0087] Relationship path traversal: For each target entity Perform a depth-first traversal: ; in, Depth-limited (depth) The traversal function returns a sequence of elements from the beginning of the traversal. All paths from which we depart, For the set of candidate relationship paths, Represents a knowledge graph.

[0088] Path occurrence probability assessment: For each candidate path : ; in, The conditional matching degree function is used to calculate the real-time environmental conditions. With path Typical conditions The degree of conformity; For path Historical frequency of occurrence attribute; For path Path confidence attribute; , , For the preset weights, satisfy ; A comprehensive probability score is given for the occurrence of this path.

[0089] Predicted path generation: Path by Sort in descending order, then take the first k items or all items exceeding the probability threshold. The path constitutes the predicted path data. : ; Compared to existing technologies, traditional methods typically rely on manual compilation of process knowledge bases, which suffers from update lags and subjective biases. This solution, however, automatically integrates multi-source heterogeneous data through a large-scale model. For example, it rapidly merges laboratory research reports on a novel pollutant into the existing knowledge system, while dynamically adjusting the knowledge graph structure based on treatment effects observed during actual water plant operation. The separate theoretical knowledge base and empirical database in existing technologies are unified in this solution. For instance, when the actual migration path of a pollutant repeatedly deviates from theoretical predictions, the system can automatically trigger confidence downgrading and path correction of the corresponding relationships in the knowledge graph.

[0090] S5. Obtain the monitoring data stream at the current moment, perform spatiotemporal alignment and cleaning on the monitoring data stream, generate standardized time-series trajectory data, and use time-series association rule mining and state reasoning algorithms to deduce the actual migration sequence and transformation product sequence of pollutants between process units in a chain-like manner to obtain real-time path data.

[0091] Receive the monitoring data stream at the current moment; The monitoring data stream is spatiotemporally aligned and cleaned to generate standardized time-series trajectory data; Based on time-series trajectory data, the actual migration sequence and transformation product sequence of pollutants between process units are deduced through time-series association rule mining and state reasoning algorithms, thus obtaining real-time path data.

[0092] It continuously receives monitoring data streams from various "first acquisition nodes" (such as inlets, connecting pipes of each process unit, and inlets of the clear water tank). Each data entry includes a timestamp, a spatial location identifier (such as "outlet of sedimentation tank #1"), monitoring values ​​(such as pollutant concentration, pH, and flow rate), and a data quality indicator.

[0093] First, all data are time-aligned using a unified clock (e.g., resampled to 1 second / point). Then, the data is mapped to the corresponding spatial node sequence according to the process flow diagram. Data cleaning is performed: invalid values ​​are removed, noise is smoothed, and missing values ​​are filled using interpolation or process model estimation. Finally, a time-continuous, standardized concentration time-series trajectory across all process units is generated for each contaminant.

[0094] Based on the above time-series trajectory, the actual path is reconstructed through the following reasoning: Temporal association rule mining: Calculate the time-delay cross-correlation between concentration changes at upstream and downstream nodes. If the concentration change at node A is strongly correlated with the concentration change at node B after a delay time Δτ, and Δτ coincides with the hydraulic residence time, then establish the association rule "A → B".

[0095] State Reasoning and Chain Inference: Starting Point Identification: Identify the node with the earliest abnormal concentration in the time-series trajectory as the pollution source; Migration Chain Construction: Starting from the source, find the next node with an abnormal response after a reasonable delay according to association rules, and link them sequentially; Transformation Analysis: Compare the monitoring data of adjacent nodes on the chain. If the concentration decreases significantly, it is inferred that removal or precipitation has occurred in that unit; if the pollutant type changes (e.g., ammonia nitrogen decreases, nitrate increases), it is inferred that chemical or biological transformation has occurred; Iterative Output: Repeat the above process until the effluent outlet or concentration returns to normal, outputting structured real-time path data, including: migration sequence (e.g., "inlet → coagulation tank → sedimentation tank"), transformation descriptions of each stage (e.g., "adsorption and precipitation in the coagulation tank, removal rate approximately 50%), and path confidence.

[0096] Example: Suppose a nickel contamination incident occurs at a water plant, and the system executes S5: After data alignment and cleaning, the nickel concentration sequence is displayed as follows: inlet (10:00, 1.0 mg / L) → coagulation tank outlet (10:10, 0.55 mg / L) → sedimentation tank outlet (10:25, 0.4 mg / L) → outlet (10:35, 0.35 mg / L).

[0097] Association rule mining showed that the time lags for "inlet-coagulation tank outlet" and "coagulation tank outlet-sedimentation tank outlet" were 10 minutes and 15 minutes, respectively, and the correlation was significant.

[0098] The chain reaction yielded the following real-time path: Migration sequence: Inlet → Coagulation tank → Sedimentation tank → Outlet; Transformation sequence: Nickel (1.0 mg / L) at the inlet is adsorbed and precipitated in the coagulation tank (removal rate approximately 45%) → further separation in the sedimentation tank (removal rate approximately 27%) → Residual at the outlet (0.35 mg / L); Confidence level: High (due to clear temporal correlation and concentration gradient consistent with process expectations).

[0099] Compared to existing technologies, traditional methods typically only detect pollutant exceedances at the end of the system and then require significant manpower and resources for reverse investigation, resulting in a delayed response. In contrast, this invention, through chain-like deduction, can pinpoint the exact source of pollutants in real time and automatically track their migration paths and transformation processes throughout the system.

[0100] S6. Compare the real-time path data with the predicted path data to generate path deviation index data, and calculate the difference between the current effect feedback data and the predicted effect feedback data as the effect feedback data difference; Extract the key node sequences and corresponding expected pollutant concentration data recorded in the predicted path data; Extract the actual node sequence and corresponding measured pollutant concentration data recorded in the real-time path data; The consistency between the key node sequence and the actual node sequence is compared sequentially to generate a path matching index; The relative deviations between the expected pollutant concentration data and the measured pollutant concentration data at the same node are calculated sequentially to generate a concentration deviation index. By integrating the path fit index and the concentration deviation index, a comprehensive path deviation index data is generated.

[0101] Path deviation index data generation: Calculate the predicted node sequence With actual node sequence The longest common subsequence (LCS) length between them is used to calculate the fit: ; in, This indicates the degree of fit; `max` represents a function that returns the largest value in the parameter list. Represents a sequence and The length of the longest common subsequence, This represents the length function.

[0102] Concentration deviation index: For the common nodes of the predicted path and the actual path (i.e., nodes in LCS), calculate the relative deviation between the expected concentration and the measured concentration for each node, and take the average value.

[0103] Comprehensive path deviation index: The comprehensive deviation value is calculated by weighting and summing the above two indicators, namely, the degree of similarity and the concentration deviation, with the sum of the weighting coefficients being 1.

[0104] This invention reconstructs the actual migration path of pollutants based on subsequent monitoring data streams, generates real-time path data, and compares it with the predicted path to generate path deviation index data. This step ensures system self-verification: for example, if the actual flow direction of pollutants differs from the prediction (e.g., due to pipeline blockage), the deviation index quantifies the difference.

[0105] S7. Set a second threshold and determine whether the path deviation index data is greater than the second threshold. If it is, trigger the adaptive adjustment of the knowledge graph based on the difference between the path deviation index data, real-time path data, and effect feedback data, and use the updated knowledge graph for path prediction in the next moment. Otherwise, do not process it.

[0106] The specific process is as follows: Sort the path deviation index data according to the degree of deviation, and filter out the path segments whose deviation values ​​exceed the preset second threshold (e.g., set the threshold to 0.9); Map the path segment to the knowledge graph data, locate one or more relationships between the process unit entity and the pollutant entity corresponding to the segment, and mark them as target relationships to be verified; If an existing relationship cannot be mapped, a new relationship identifier is generated to be created.

[0107] The specific process for generating updated data for target entities or target relationships based on real-time path data and effect feedback data is as follows: If it is a conditional attribute update: extract the actual environmental parameter data when the target relationship occurs from the real-time path data, and use it to update or expand the reaction conditional attribute of the relationship; If it is an efficiency attribute update: Based on the final removal rate in the difference of effect feedback data and the time node information in the real-time path data, calculate the actual response efficiency data of the target relationship, and use it to update its efficiency attribute; If it is a confidence update: adjust the confidence data of the target relationship downward based on the magnitude of the path deviation index data; conversely, if the prediction is highly consistent with the actual path, adjust its confidence data upward.

[0108] The specific process of applying updated data to knowledge graph data is as follows: Incremental updates are used to submit updated data to the knowledge graph database in the form of transactions; After the update is completed, record the version log data and the reason for the adjustment. The reason for the adjustment is linked to the path deviation indicator data and effect feedback data that triggered the adjustment.

[0109] When deviations exceed the limits, the system updates the knowledge graph based on multi-source data (path deviation indicators, real-time paths, and effect feedback) to better reflect actual operating conditions. This not only solves the problem of static model aging (such as the emergence of new pollutants) but also ensures global optimization through feedback loops (such as effect feedback data indicating effluent quality).

[0110] If the path deviation index data is not greater than the preset second threshold, it indicates that the current knowledge graph's predicted path for pollutants basically matches the actual migration path, and the system's prediction accuracy is high. In this case, the system does not trigger active adjustments to the knowledge graph, but will perform the following operations: Confidence reinforcement: For the relevant entities and relational paths in the knowledge graph that were successfully matched in this prediction, appropriately increase their confidence (e.g., by increasing them by a small fixed coefficient) to reinforce the correct knowledge.

[0111] Log recording and performance statistics: Cases with a high degree of "prediction-actual" match are recorded as positive samples in the historical database to update the "historical occurrence frequency" attribute of the path and provide data for the calculation of the system's comprehensive performance index (used to dynamically adjust the second threshold).

[0112] The process continues: The system will use the current knowledge graph to directly enter the loop of the next monitoring cycle (i.e., return to step S2 / S3) and continue to perform anomaly detection and path prediction on the new monitoring data stream.

[0113] Through the above technical solutions, this invention achieves automated construction and dynamic optimization of knowledge graphs, solving the decision-making bias problem caused by the static nature of knowledge bases in traditional water plant management systems. For example, when a new pollutant not recorded in the process manual appears, the system can quickly expand the knowledge graph by integrating external literature data, while simultaneously verifying and correcting the reasoning logic based on real-time operational data, thereby improving the accuracy of path prediction under abnormal operating conditions. It transforms multi-source, heterogeneous data (from static literature to dynamic operational logs) into a machine-readable, semantically related unified knowledge network, breaking down data silos and forming a structured understanding of pollution control knowledge. It elevates traditional data processing capabilities from simple statistical analysis to complex causal reasoning and decision support, providing impetus for intelligent, precise pollution source tracing and efficient management.

[0114] Following the step of triggering adaptive adjustments to the knowledge graph, there are also steps for assessing the impact of the adjustments and rolling back, specifically including: Obtain path deviation index data and effect feedback data for multiple consecutive periods after the knowledge graph is adjusted; Calculate the adjusted average path deviation and average effect feedback index, compare them with the data before adjustment, and generate an adjustment benefit score. If the adjusted benefit score is lower than the preset benefit threshold, the knowledge graph version will be rolled back to the state before the adjustment, and the adjustment will be recorded as invalid.

[0115] Calculate the weighted average of the path deviation index data before and after the adjustment. The weights are distributed according to time decay, with more recent data having higher weights, and the sum of the weights must be 1. Calculate the weighted average of the effect feedback data before and after the adjustment, with the same weights as the path deviation index data before and after the adjustment.

[0116] Calculate the difference between the weighted average of the path deviation index data before and after adjustment, and then calculate the ratio of this difference to the weighted average of the path deviation index data before adjustment as the path deviation improvement rate. ; Calculate the difference between the weighted average of the feedback data before and after the adjustment, and then calculate the ratio of this difference to the weighted average of the feedback data before the adjustment as the improvement rate. ; Adjusting the benefit score : ; in, , Indicates weight, and It can be dynamically adjusted according to the optimization target (e.g., if the focus is on improving deviations, then increase). Focusing on feedback will improve ).

[0117] For example, the dynamic adjustment process of the second threshold includes: Real-time acquisition of system operational data, including: the average consistency between the predicted path data from the previous period and the real-time path data. Standard deviation of monitoring data stream and the rate of achievement of feedback data ; Average fit P: Reflects the average fit between the "predicted-actual" path in the previous period; obtain the K path fit indices generated by comparing all valid paths in the previous period, and calculate their mean as the average fit.

[0118] The standard deviation M of the monitoring data stream reflects the degree of fluctuation in the concentration of key pollutants in the influent during the previous period. The standard deviation of the concentration time series data of one or two dominant pollutants (determined by the knowledge graph) collected from the raw water inlet in the previous period is used as the standard deviation of the monitoring data stream.

[0119] The compliance rate E of the effect feedback data reflects the proportion of the final effluent water quality that meets the standards in the previous cycle. It involves obtaining T valid sampling values ​​for one core assessment indicator (such as total nickel concentration) at the effluent outlet in the previous cycle, along with its standard limit S. The number of samples that meet the standards is counted, and the ratio of the number of samples that meet the standards to the total number of samples is calculated as the compliance rate of the effect feedback data.

[0120] The specific process of weighted fusion processing of operational data is as follows: Average fit Standard deviation and the pass rate Each is assigned a dynamic weighting coefficient, denoted as . , , And it satisfies the weight normalization condition: ; The dynamic weighting coefficients are configured according to the current optimization objectives. If the optimization objective is to ensure stable effluent, the weight of the final compliance rate is increased; if the primary objective is to improve the accuracy of the prediction model, the weight of the average fit is increased. The overall performance index of the system is calculated by multiplying each data point by its corresponding dynamic weight coefficient and then summing the results. Its mathematical expression is: ; Comprehensive performance indicators As input, it is substituted into the predefined inverse mapping function to calculate and output the adjusted second threshold. Its mathematical expression is: ; The specific process of constructing the reverse mapping function includes: Obtain a comprehensive performance sample dataset of the system from the historical operation database. and its corresponding second threshold sample dataset ,in For the sample size, For the first A comprehensive performance sample vector, To and The corresponding second threshold; For comprehensive performance sample datasets Compared with the second threshold sample dataset Perform correlation analysis and regression fitting based on the analysis results. This process includes: Correlation analysis: Calculate the first Performance metrics , This indicates that a matrix transpose operation is performed, along with the second threshold. Correlation coefficient between The larger the absolute value, the stronger the linear relationship between the performance index and the second threshold. Its mathematical expression is the Pearson correlation coefficient: ; in, For the first The mean of a sample of performance metrics, The mean of the second threshold samples. For the first The first comprehensive performance sample vector A sample of performance metrics; Select The performance metrics are used as key performance indicators to form a new input vector. ; Regression Fitting: Establishing a vector of key performance indicators For input, second threshold The regression model for output Its general form is: ; in, The set of model parameters to be solved. For fitting residuals; Using the least squares method to evaluate the parameters The mathematical objective function for optimal estimation is: ; in, For linear regression, the loss function is... Squared loss function Solve the above optimization problem to obtain the optimal parameter set. Thus, the function is completed. The construction of.

[0121] Data output steps: Output the determined regression model and its optimal parameter set As a reverse mapping function, its mathematical expression is: ,in This is the input vector of key performance indicators. This is the predicted second threshold.

[0122] Knowledge graph adaptive adjustment: Data definition, Path deviation index data, a scalar, represents the degree of difference between the predicted path and the real-time path; Real-time path data consists of the actual observed path sequences; Performance feedback data, such as the rate of water effluent meeting standards; The knowledge graph at the current time t.

[0123] Algorithm process: Target positioning, if (Second threshold), then the path segment with the largest deviation is mapped to the knowledge graph: ; in, The target relationship to be corrected. This represents the set of all relations in the knowledge graph. If mapping fails, a new relation identifier is generated. . This means mapping real-time paths to a knowledge graph to locate the target relationships that need to be corrected.

[0124] Updated data generation: Attribute updates (for relationships) ): ; in, This indicates that the condition attributes are updated based on the actual environmental parameters. Conditional attributes representing the target relationship describe the conditions under which the relationship occurs, such as pH range, temperature, flow rate, etc. This represents actual environmental parameter data, extracted from real-time path data, used to update conditional attributes. This represents the conditional attribute of the target relationship after the update.

[0125] ; in, This represents the actual efficiency calculated based on real-time path data. and feedback data The calculated actual reaction efficiency, Efficiency attributes representing the target relation describe the processing efficiency of that relation, such as removal rate (%).

[0126] Confidence adjustment: ; in, This is an adjustment coefficient used to control the magnitude of confidence level adjustment, ranging from 0 to 1, with a deviation... The larger the value, the greater the decrease in confidence level. The confidence attribute represents the target relationship and is used to determine the credibility of the relationship. It is initially given by the large model and subsequently adjusted according to the actual situation.

[0127] Entity / relationship creation (if needed): Incremental application updates This represents a temporary knowledge graph, a temporary graph structure formed after creating new relations, used for subsequent updates: Apply the above updates to the knowledge graph to generate a new version: ; in, For the set of all updated data, For graph database transaction operations, the data set will be updated. Applying it to a knowledge graph generates a new version. This represents the updated knowledge graph, that is, the knowledge graph that has been adaptively adjusted, for prediction at the next time step.

[0128] Closed-loop feedback: Updated knowledge graph The path prediction for the next time step will be used to form an iteration: ; in, This refers to an algorithm that uses outlier data and knowledge graphs for path prediction. This represents the predicted path data for the next time step, obtained by using the updated knowledge graph and outlier data from the next time step. This represents the abnormal data stream at the next moment, generated by the first threshold comparison and judgment module.

[0129] Compared to existing technologies, traditional methods often use fixed threshold values ​​or adjust based on a single parameter, which cannot adapt to multi-objective optimization needs. For example, using the original threshold in the event of a sudden pollutant outbreak may lead to response delays. This solution, however, uses dynamic weight allocation and a reverse mapping function to automatically adjust the threshold according to real-time operational objectives. For instance, when the requirement for effluent stability increases, the weight of the effect feedback data is increased, making the threshold adjustment more inclined to suppress water quality fluctuations, thereby improving the system's adaptability to complex operating conditions.

[0130] As the average consistency and monitoring data stream change, the threshold can also adaptively adjust based on the feedback data, always keeping the system in optimal working condition and extending its effective lifespan. Automatic threshold adjustment makes the judgment criteria (such as data validity assessment and model confidence alarms) more closely aligned with the current water plant, thereby reducing false alarms and improving monitoring accuracy.

[0131] When the deviation exceeds a preset second threshold, the system triggers adaptive adjustment of the knowledge graph, forming a closed loop of "prediction-validation-self-adjustment." Existing technologies often neglect subsequent data validation, leading to model drift, such as long-term error accumulation. However, this application... Application scenario example: In a large city wastewater treatment plant: The plant has a daily processing capacity of one million tons, a complex process flow, and large fluctuations in influent water quality, which places extremely high demands on its ability to respond quickly and accurately trace the source of pollution incidents.

[0132] After the system was implemented, the physical sensor network (such as auxiliary sensing units like pH, turbidity, conductivity, and UV254 sensors, as well as target sensing units like expensive and frequently maintained online heavy metal analyzers) that are distributed throughout the plant's water pipes constituted the system's "sensory nerves," continuously collecting and uploading monitoring data streams.

[0133] When an unidentified pollutant suddenly appears in the influent, causing abnormal fluctuations in the monitoring values ​​of a certain process unit (such as a biological treatment tank), the system can compare it with the dynamic operating baseline within seconds. Once the deviation exceeds the adaptively calculated threshold, an alarm is immediately triggered, and an anomaly data packet containing contextual information such as the location of the anomaly, time, related parameters, and real-time influent flow rate is automatically generated.

[0134] Subsequently, the system's core "intelligent brain"—a knowledge graph built on a large model—began to function. Key features from anomalous data packets were extracted and transformed into a high-dimensional query vector, which was then matched against a vast array of pollutant entities within the knowledge graph. The knowledge graph not only incorporated years of historical operating logs and laboratory test reports from the plant but also deeply integrated domain knowledge from external Chemical Safety Data Sheets (MSDS) and process manuals.

[0135] Through matching, the system quickly identified the most likely target pollutant for this anomaly as "industrial nickel-containing wastewater". Then, starting with this entity, the system automatically traversed all possible relationship paths in the knowledge graph: for example, "nickel-containing wastewater" → "inhibition of nitrifying bacteria activity" → "excessive ammonia nitrogen in effluent"; or "nickel-containing wastewater" → "reaction with coagulant to form precipitate" → "removed in sedimentation tank".

[0136] The system will combine the actual influent load, water temperature and other environmental conditions at the time to perform a weighted evaluation of the probability of occurrence of these potential paths, and finally generate a predicted path distribution sorted by probability. It will also predict that if the path migrates along the main path, it will be detected at the outlet of the subsequent deep treatment unit (such as the filter bed), as well as the prediction effect feedback data associated with the predicted path data.

[0137] Simultaneously, the system's "verification and closed-loop" mechanism is activated. It continuously tracks subsequent monitoring data streams and, through time-series correlation algorithms, reconstructs the actual "travel trajectory" (i.e., real-time path data) of the pollutant throughout the plant. By comparing this actual path with the previously predicted path, the system generates a path deviation index and calculates the difference between the current effect feedback data and the predicted effect feedback data as the effect feedback data difference. For example, if the actual trajectory shows that the removal efficiency of the pollutant in the sedimentation tank is far lower than the historical experience value in the knowledge graph, this deviation will be immediately detected.

[0138] When a significant deviation occurs, the system doesn't simply issue an alarm; instead, it triggers the adaptive adjustment mechanism of the knowledge graph. It automatically locates the relationship "nickel-containing wastewater - sedimentation tank - removal efficiency" within the knowledge graph and dynamically updates its "reaction conditions" and "efficiency confidence" attributes based on real operational data (such as the pH value and coagulant dosage). This means that the knowledge graph continuously learns and improves through each practical feedback, making its next prediction more accurate.

[0139] Through the above technical solution, this invention realizes a closed-loop optimization mechanism for knowledge graphs, effectively solving the prediction inaccuracy problem caused by model rigidity in traditional systems. By dynamically correcting entity attributes and relation weights, the knowledge reasoning results are made more consistent with actual working conditions, thereby improving the accuracy of pollutant source tracing and the effectiveness of treatment plans. Simultaneously, the online update mechanism avoids the interruption problem of traditional offline training modes, ensuring the continuous and stable operation of the water plant management system.

[0140] Example 2: Intelligent Monitoring and Control System for the Entire Water Plant Process Based on a Large Model (see...) Figure 4 ),include: The data acquisition module sets up multiple first acquisition nodes at the raw water inlet, the connecting pipes of each process unit, and the inlet of the clear water tank, and sets up a second acquisition node at the pipe outlet. The second acquisition node maintains a safe distance from the first acquisition node. The first acquisition node collects and monitors the data stream, and the second acquisition node collects and feeds back the data. The data acquisition module is used to acquire the monitoring data stream and effect feedback data of the previous moment; and to acquire the monitoring data stream and effect feedback data of the current moment. The first threshold comparison and determination module is used to determine whether the monitoring data stream at the previous moment deviates from the first threshold range. If so, the monitoring data stream is associated with the timestamp and spatial location to generate an abnormal data stream. The knowledge graph matching module is used to convert the abnormal data stream into vector data and input it into a preset knowledge graph for matching and reasoning, generating prediction path data containing at least one prediction path data at the current moment, as well as prediction effect feedback data associated with the prediction path data. The knowledge graph is used to characterize the relationship between pollutants, process units and reaction paths. The real-time path generation module is used to reconstruct the actual migration path of pollutants based on the monitoring data stream at the current moment and generate real-time path data. The path deviation generation module is used to compare the real-time path data with the predicted path data, generate path deviation index data, and calculate the difference between the current effect feedback data and the predicted effect feedback data as the effect feedback data difference. The second threshold comparison and determination module is used to determine whether the path deviation index data is greater than a preset second threshold. The knowledge graph adaptive update module adjusts the knowledge graph based on path deviation index data, real-time path data, and effect feedback data for values ​​exceeding a preset second threshold.

[0141] The first acquisition node includes at least a set of auxiliary sensing units and a target sensing unit; The target sensing unit is used to collect monitoring data streams, and the auxiliary sensing unit is used to collect parameters associated with the monitoring data streams; If the target sensing unit fails to acquire data at the current moment, the parameters acquired by the auxiliary sensing unit will be input into the preset prediction model to generate the monitoring data stream for the current moment.

[0142] Specifically, the data acquisition module continuously receives multi-dimensional monitoring data from the first acquisition node and water quality feedback data from the second acquisition node. The first threshold determination module dynamically calculates the fluctuation range of the monitoring data. When abnormal data exceeding the threshold is detected, it automatically adds the acquisition time and sensor location information to form a structured abnormal record. The knowledge graph matching module converts the abnormal data into a high-dimensional vector and generates multiple possible migration path predictions and prediction effect feedback data associated with the predicted path data by searching for associated pollutant types, reaction paths, and process units in the knowledge graph. The real-time path generation module combines the sensor network data at the current moment to reverse-engineer the actual diffusion trajectory of the pollutants. The path deviation generation module generates path deviation index data by calculating the differences between the predicted path and the actual path in the time, spatial, and concentration dimensions, and calculates the difference between the effect feedback data at the current moment and the predicted effect feedback data as the effect feedback data difference. The second threshold determination module dynamically adjusts the determination threshold according to the system operating status. When the deviation exceeds the critical value, it automatically starts the incremental learning mechanism of the knowledge graph to update the model parameters.

[0143] This system achieves closed-loop control of the entire process, from data acquisition, anomaly identification, path prediction to autonomous optimization, by constructing a modular intelligent processing chain. The response delay problem caused by manual intervention in existing technologies is effectively solved by the real-time reasoning capabilities and adaptive adjustment mechanism of knowledge graphs. At the same time, through the collaborative work of multiple modules, the accuracy of pollutant source tracing and the timeliness of process control are improved.

[0144] The above is a further description of the present invention in conjunction with specific embodiments, and the scope of protection of the present invention is not limited thereto.

Claims

1. A water plant full-process intelligent monitoring and control method based on a large model, characterized by: The steps include the following: S1. Based on domain text data and historical operation data, a large model is used to perform entity recognition and relationship extraction processing on the data, and a knowledge graph is constructed according to the relationship between water plant pollutants, process units and reaction paths; S2. Multiple first data acquisition nodes are set at the raw water inlet, the connecting pipes of each process unit, and the inlet of the clear water tank. A second data acquisition node is set at the pipe outlet. The second data acquisition node maintains a safe distance from the first data acquisition node. The first data acquisition node collects and monitors the data stream, and the second data acquisition node collects and feeds back the data. S3. Obtain the monitoring data stream from the previous moment, set a first threshold interval, determine whether the monitoring data stream deviates from the first threshold interval, and if so, associate the monitoring data stream with a timestamp and spatial location to generate an abnormal data stream; S4. Preprocess the abnormal data and extract key feature data, input the key feature data into the embedding model for encoding and vectorization, generate query vector data, input it into the pre-built knowledge graph for matching and reasoning, and generate prediction path data containing at least one pollutant at the current moment, as well as prediction effect feedback data associated with the prediction path data. S5. Obtain the monitoring data stream and effect feedback data at the current moment, perform spatiotemporal alignment and cleaning on the monitoring data stream, generate standardized time-series trajectory data, and use time-series association rule mining and state reasoning algorithms to deduce the actual migration sequence and transformation product sequence of pollutants between process units in a chain-like manner to obtain real-time path data. S6. Compare the real-time path data with the predicted path data to generate path deviation index data; calculate the difference between the current effect feedback data and the predicted effect feedback data as the effect feedback data difference; S7. Set a second threshold and determine whether the path deviation index data is greater than the second threshold. If so, based on the difference between the path deviation index data, real-time path data, and effect feedback data, trigger the adaptive adjustment of the knowledge graph and use the updated knowledge graph for path prediction in the next moment.

2. The intelligent monitoring and control method for the entire process of a water plant based on a large model as described in claim 1, characterized in that, Step S1, the knowledge graph construction process, is as follows: using a large model to perform entity recognition and relation extraction on the domain text data, generating initial knowledge graph data represented in tuple form; Feature extraction and event reconstruction are performed on historical operation data to generate empirical verification data in the form of process event chains; The initial data of the knowledge graph and the empirical verification data are fused and conflict-resolved using a large model to generate the basic topology of the knowledge graph data; attribute data and confidence data are added to the entities and relationships in the basic topology using the large model to generate the knowledge graph.

3. The intelligent monitoring and control method for the entire process of a water plant based on a large model as described in claim 1, characterized in that, In step S3, when the target sensor has uncollected data, the parameters collected by the associated auxiliary sensing unit are input into the trained prediction model, and the predicted monitoring data stream is used to replace it.

4. The intelligent monitoring and control method for the entire process of a water plant based on a large model as described in claim 1, characterized in that, The adjustment method for the first threshold range is as follows: Dynamic sensitivity coefficients are generated based on the standard deviation of the monitoring data stream from the previous monitoring cycle and the types of dominant pollutants in the knowledge graph. The first threshold at the current moment is calculated by multiplying the first threshold interval of the previous monitoring period by the dynamic sensitivity coefficient; The dominant pollutant refers to the pollutant that poses the highest risk to achieving the effluent quality standard, calculated by a risk assessment algorithm based on predefined pollutant toxicity weight data in the knowledge graph, pollutant concentration data in the current monitoring data stream, and current process unit data.

5. The intelligent monitoring and control method for the entire process of a water plant based on a large model as described in claim 1, characterized in that, Step S3, after generating the abnormal data stream, also includes an anomaly confidence assessment step.

6. The intelligent monitoring and control method for the entire process of a water plant based on a large model as described in claim 1, characterized in that, The key feature data in step S4 includes pollutant type feature data, concentration gradient data, and spatial location sequence data. The pollutant prediction path data is obtained by performing similarity matching calculations between query vector data and pollutant entity vectors in knowledge graph data to identify one or more most likely target pollutant entities. Starting from the target pollutant entity, all relationship paths in the knowledge graph data with that entity as the head node are traversed. The occurrence probability of each traversed relationship path is weighted and evaluated to generate a comprehensive occurrence probability score for the path. The paths are then sorted from high to low based on the comprehensive occurrence probability scores. This yields the probability order of occurrence of various possible pollutant migration and transformation paths under the current abnormal situation, thus obtaining the prediction path data.

7. The intelligent monitoring and control method for the entire process of a water plant based on a large model as described in claim 1, characterized in that, The process of generating path deviation index data in step S6 is as follows: extract the key node sequence and corresponding expected pollutant concentration data recorded in the predicted path data; extract the actual node sequence and corresponding measured pollutant concentration data recorded in the real-time path data; compare the consistency between the key node sequence and the actual node sequence in turn to generate a path consistency index; calculate the relative deviation between the expected pollutant concentration data and the measured pollutant concentration data at the same node in turn to generate a concentration deviation index; and fuse the path consistency index and the concentration deviation index to generate comprehensive path deviation index data.

8. The intelligent monitoring and control method for the entire process of a water plant based on a large model as described in claim 1, characterized in that, The adjustment method for the second threshold is as follows: Obtain the average consistency between the predicted path data and the real-time path data at the previous moment, the standard deviation of the monitoring data stream, and the compliance rate of the effect feedback data; A dynamic weighting coefficient is assigned to each of the average fit, standard deviation, and compliance rate; The system's overall performance index is calculated by multiplying each data point by its corresponding dynamic weight coefficient and then summing the results. The system's overall performance index is input into a predefined inverse mapping function to obtain the adjusted second threshold.

9. A water plant end-to-end intelligent monitoring and control system based on a large-scale model, characterized by: include: The data acquisition module sets up multiple first acquisition nodes at the raw water inlet, the connecting pipes of each process unit, and the inlet of the clear water tank, and sets up a second acquisition node at the pipe outlet. The second acquisition node maintains a safe distance from the first acquisition node. The first acquisition node collects and monitors the data stream, and the second acquisition node collects and feeds back the data. The data acquisition module is used to acquire the monitoring data stream from the previous moment. Obtain the current monitoring data stream and effect feedback data; The first threshold comparison and determination module is used to determine whether the monitoring data stream at the previous moment deviates from the first threshold range. If so, the monitoring data stream is associated with the timestamp and spatial location to generate an abnormal data stream. The knowledge graph matching module is used to convert the abnormal data stream into vector data and input it into a preset knowledge graph for matching and reasoning, generating prediction path data containing at least one prediction path data at the current moment, as well as prediction effect feedback data associated with the prediction path data. The knowledge graph is used to characterize the relationship between pollutants, process units and reaction paths. The real-time path generation module is used to reconstruct the actual migration path of pollutants based on the monitoring data stream at the current moment and generate real-time path data. The path deviation generation module is used to compare the real-time path data with the predicted path data, generate path deviation index data, and calculate the difference between the current effect feedback data and the predicted effect feedback data as the effect feedback data difference. The second threshold comparison and determination module is used to determine whether the path deviation index data is greater than a preset second threshold. The knowledge graph adaptive update module adjusts the knowledge graph adaptively based on the difference between path deviation index data, real-time path data, and effect feedback data for values ​​exceeding a preset second threshold.

10. The intelligent monitoring and control system for the entire water plant process based on a large model as described in claim 9, characterized in that, The first acquisition node includes at least a set of auxiliary sensing units and a target sensing unit; The target sensing unit is used to collect monitoring data streams, and the auxiliary sensing unit is used to collect parameters associated with the monitoring data streams; If the target sensing unit fails to acquire data at the current moment, the parameters acquired by the auxiliary sensing unit will be input into the preset prediction model to generate the monitoring data stream for the current moment.