A Method and System for Predicting Biochemical Treatment Efficiency Based on Dynamic Evolution of Water Quality Fingerprint
By collecting and analyzing water quality fingerprint time-series data in real time, combined with deep learning models and expert knowledge bases, the problem of predicting the impact of dynamic changes in water quality on the biochemical system was solved, enabling advanced prediction and control of biochemical treatment efficiency, and improving the stability and efficiency of the wastewater treatment system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS
- Filing Date
- 2026-04-24
- Publication Date
- 2026-06-30
AI Technical Summary
Existing water fingerprinting technology lacks the ability to mine the temporal characteristics of dynamic changes in water quality, thus failing to capture its impact on the biochemical system. Furthermore, it fails to establish a predictive relationship between the dynamic characteristics of water fingerprints and key performance indicators of biochemical treatment, resulting in the system being unable to take control measures in advance, which affects the stability of effluent water quality and treatment efficiency.
By collecting and analyzing influent water quality fingerprint time-series data in real time, a mapping relationship between the data and the efficiency of biochemical treatment is established. Deep learning models such as long short-term memory networks, bidirectional long short-term memory networks, or Transformer models are used for prediction. Combined with expert knowledge base, process control suggestions are generated to achieve advanced prediction and early warning of future biochemical treatment efficiency.
It enables accurate and proactive prediction and targeted control of biochemical treatment efficiency, improves the stability and treatment efficiency of wastewater treatment systems, breaks through the traditional passive management mode of lagging monitoring, and provides technical support for intelligent and low-energy operation of wastewater treatment plants.
Smart Images

Figure CN122089171B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wastewater treatment technology, specifically to a method and system for predicting biochemical treatment efficiency based on the dynamic evolution of water quality fingerprints, which is used to achieve advanced prediction and early warning control of the biochemical treatment efficiency of wastewater treatment systems. Background Technology
[0002] Performance monitoring and prediction of wastewater treatment biochemical systems is a core technology area in the water treatment industry, and it is of great significance for ensuring the stability of effluent quality and improving system operating efficiency. With the increasing demand for intelligent water treatment, related technology research is becoming increasingly in-depth.
[0003] Traditional wastewater biochemical treatment monitoring technologies mainly rely on discrete sampling analysis and conventional online monitoring equipment. For example, parameters such as COD, ammonia nitrogen, and total nitrogen are determined through laboratory analysis of periodically collected water samples, or key indicators are monitored in real time using online sensors for ammonia nitrogen, pH, and dissolved oxygen. These methods can reflect the current operating status of the system, but they lack predictability, and the long sampling frequency and analysis cycle lead to data delays.
[0004] In recent years, spectroscopic analysis techniques have been applied in the field of water quality monitoring. In particular, ultraviolet-visible spectroscopy (UV-Vis) and fluorescence excitation-emission matrix (EEM) techniques have been used to construct "water fingerprints." This method can quickly and without reagents obtain overall characteristic information of organic matter and certain inorganic matter in water bodies. Existing technologies can already use water fingerprints to identify toxic substances in water bodies or estimate routine water quality parameters.
[0005] However, existing water fingerprinting technology has two main technical problems: First, it lacks in-depth analysis of the temporal characteristics of water fingerprint data, making it impossible to capture the impact of dynamic changes in water quality on the biological system; second, it fails to establish a predictive relationship between the dynamic characteristics of water fingerprints and key performance indicators of biological treatment (such as nitrification rate and denitrification effect), resulting in the system being unable to take control measures in advance for impending changes in treatment efficiency, affecting the stability of effluent water quality and treatment efficiency. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for predicting the efficiency of biochemical treatment based on the dynamic evolution of water quality fingerprints. By collecting and analyzing the time series data of influent water quality fingerprints in real time, a mapping relationship between the data and the efficiency of biochemical treatment is established, enabling advanced prediction of future biochemical treatment efficiency. Based on the prediction results, early warning information and process control suggestions are generated to improve the stability and treatment efficiency of the wastewater treatment system.
[0007] To achieve the above objectives, this invention provides a method for predicting the effectiveness of biochemical treatment based on the dynamic evolution of water quality fingerprints, comprising:
[0008] Spectral data of the influent to the wastewater treatment system are acquired and preprocessed to obtain a standardized spectral matrix. Spectral feature parameters are extracted based on the standardized spectral matrix to form a water quality fingerprint feature vector. The continuously acquired water quality fingerprint feature vectors are organized according to time to construct a water quality fingerprint time series dataset.
[0009] The efficiency indicators of the effluent from the biochemical treatment system are collected and standardized to obtain standardized efficiency indicator data. A time correspondence is established based on the water quality fingerprint time series dataset and the standardized efficiency indicator data, and a training dataset is constructed. Feature engineering is performed on the training dataset to obtain model training data.
[0010] The model training data is used to train and optimize the prediction model to obtain an optimized prediction model. Real-time water quality fingerprint time series data is received to construct real-time prediction data. The optimized prediction model is used to predict the real-time prediction data and output the predicted value of the biochemical treatment efficiency index.
[0011] Based on historical performance index data, a graded early warning threshold is determined. The predicted value of the biochemical treatment performance index is compared with the graded early warning threshold to generate early warning information. Based on the early warning information, an expert knowledge base is queried and an experience learning model is applied to generate process control suggestions. The process control suggestions are then optimized to form a control scheme.
[0012] Further, the acquisition and preprocessing of spectral data of the wastewater influent to obtain a standardized spectral matrix, the extraction of spectral feature parameters based on the standardized spectral matrix to form a water quality fingerprint feature vector, and the organization of the continuously acquired water quality fingerprint feature vectors according to time to construct a water quality fingerprint time-series dataset, including:
[0013] An ultraviolet-visible spectrometer and a three-dimensional fluorescence spectrometer were installed at the inlet of the wastewater treatment system. An automatic cleaning device and a calibration system were configured, and the sampling frequency was set to once every 30 minutes to acquire the raw spectral data stream.
[0014] The raw spectral data stream is subjected to scattered light correction, internal filter effect correction, baseline drift correction and instrument response normalization to obtain a standardized ultraviolet-visible absorption spectral matrix and a three-dimensional fluorescence spectral matrix.
[0015] Based on the standardized UV-Vis absorption spectral matrix and the three-dimensional fluorescence spectral matrix, the characteristic wavelength absorbance, slope index, and spectral integral area of the UV-Vis spectrum, as well as the fluorescence region volume, fluorescence peak intensity ratio, and parallel factor analysis component score of the three-dimensional fluorescence spectrum, are extracted to form a multidimensional water quality fingerprint feature vector.
[0016] The continuously collected multidimensional water quality fingerprint feature vectors are organized by timestamps to construct a time-series feature matrix with a sliding time window, thus obtaining the water quality fingerprint time-series dataset.
[0017] Furthermore, based on the standardized UV-Vis absorption spectral matrix and the three-dimensional fluorescence spectral matrix, the characteristic wavelength absorbance, slope index, and spectral integral area of the UV-Vis spectrum, as well as the fluorescence region volume, fluorescence peak intensity ratio, and parallel factor analysis component scores of the three-dimensional fluorescence spectrum, are extracted to form a multidimensional water quality fingerprint feature vector, including:
[0018] The response values of each wavelength or wavelength pair in the standardized UV-Vis absorption spectral matrix and the three-dimensional fluorescence spectral matrix are used as the dimensions of the high-dimensional feature space. The high-dimensional features are projected onto the two-dimensional or three-dimensional space through principal component analysis or t-SNE dimensionality reduction technology to form a set of geometric feature points.
[0019] Based on the principles of water chemistry, in the set of geometric feature points, when the water quality characteristics represented by two feature points are correlated or causally related during the biochemical treatment process, a connection is established between the two feature points to form an intersection diagram.
[0020] Based on the intersection diagram, feature points are selected from the set of geometric feature points through correlation analysis as the terminal nodes of the Steiner tree. The parameter λ is introduced to adjust the balance between the complexity of the tree and the fitting effect. Dynamic programming and divide-and-conquer strategies are applied, and the optimal value of parameter λ is found through gradient descent or simulated annealing to obtain the optimal Steiner tree structure.
[0021] Using the optimal Steiner tree structure, the depth, width, and branching factor of the tree are calculated as topological features, the path length and path strength between terminal nodes are extracted as path features, the centrality index of each node in the tree is calculated as node centrality features, and dynamic evolution parameters are extracted by comparing the changes in the Steiner tree structure at different time points.
[0022] The multidimensional water quality fingerprint feature vector is obtained by combining the topological features, path features, node centrality features, dynamic evolution parameters, characteristic wavelength absorbance, slope exponent, spectral integral area, fluorescence region volume, fluorescence peak intensity ratio, and parallel factor analysis component scores.
[0023] Furthermore, the application of dynamic programming and divide-and-conquer strategies includes:
[0024] The intersection graph is decomposed into multiple subgraphs, and the local optimal Steiner tree is solved independently for each subgraph.
[0025] The intermediate results of solving each subgraph are recorded using a dynamic programming method. These intermediate results include the optimal connection method within the subgraph and the corresponding objective function value.
[0026] By merging the locally optimal Steiner trees of each subgraph using a divide-and-conquer strategy, calculating the connection cost between subgraphs based on the intermediate results, selecting the connection scheme that optimizes the global objective function, and constructing a global Steiner tree structure.
[0027] Furthermore, the efficiency indicators of the collected biochemical treatment system effluent are standardized to obtain standardized efficiency indicator data. A time-series correlation is established based on the water quality fingerprint dataset and the standardized efficiency indicator data, and a training dataset is constructed. Feature engineering is performed on the training dataset to obtain model training data, including:
[0028] The concentrations of ammonia nitrogen, nitrate nitrogen, total nitrogen, chemical oxygen demand (COD), sludge settling ratio, specific oxygen consumption rate (SOH), and nitrification rate at the effluent of the biochemical treatment system are collected. The units of the ammonia nitrogen concentration, nitrate nitrogen concentration, total nitrogen concentration, COD, sludge settling ratio, specific SOH, and nitrification rate are standardized to obtain standardized performance index data.
[0029] Based on the hydraulic residence time, an influent timestamp is added to each performance index data point in the standardized performance index data. The influent timestamp is associated with the timestamp in the water quality fingerprint time series dataset to establish an influent-effluent correspondence.
[0030] Based on the water quality fingerprint time series dataset and the standardized performance index data with the added water inflow timestamp, a data pair of correspondences between the water quality fingerprint sequence at time t and n hours before it and the performance index at time t+m is constructed by timestamp matching, where m represents the prediction lead time.
[0031] Outlier identification and missing value processing are performed on the corresponding data pairs. Outlier data are processed using moving median filtering, interpolation, or statistical methods to obtain a quality-controlled training dataset.
[0032] A sliding window design is applied to the quality-controlled training dataset, dividing the water quality fingerprint sequence into multiple time windows of fixed length. Temporal difference feature extraction is performed on the water quality fingerprint feature vectors within the time windows, the difference between feature values at adjacent time points is calculated, and periodic feature encoding and time encoding are added to the time windows to obtain structured model training data.
[0033] Furthermore, the outlier identification and missing value processing of the corresponding relationship data pairs includes:
[0034] The corresponding data pairs are subjected to quality review to identify outliers that exceed the normal range and missing values in the data records;
[0035] The outliers in the time series are smoothed by using a moving median filter, and missing values are filled by interpolation or by using statistical methods to calculate the estimated values of missing data points, thus obtaining the quality-controlled training dataset.
[0036] Further, the step of training and optimizing the prediction model using the model training data to obtain an optimized prediction model, receiving real-time water quality fingerprint time-series data to construct real-time prediction data, and using the optimized prediction model to predict and output predicted values of biochemical treatment efficiency indicators from the real-time prediction data includes:
[0037] The model training data is used to train a long short-term memory network model, a bidirectional long short-term memory network model, or a Transformer model. The model training is carried out using learning rate scheduling, gradient pruning, and early stopping strategies. The hyperparameters are adjusted through cross-validation to obtain the model parameters that have been trained and converged.
[0038] The converged model parameters are loaded into the Long Short-Term Memory Network model, the Bidirectional Long Short-Term Memory Network model, or the Transformer model. The prediction accuracy, mean absolute error, and root mean square error are calculated on the validation set to evaluate the model performance under different prediction lead. Based on the prediction accuracy, mean absolute error, and root mean square error, the model is integrated or its structure is adjusted to obtain an optimized prediction model.
[0039] Receive real-time water quality fingerprint time-series data, and apply feature engineering methods consistent with the sliding window design, time-series differential feature extraction, periodic feature encoding, and time encoding to construct real-time prediction data that meets the model input requirements;
[0040] The optimized prediction model is used to make multi-time-scale predictions on the real-time prediction data, outputting the predicted values of biochemical treatment efficiency indicators at different future time points, and calculating the prediction confidence interval based on the optimized prediction model.
[0041] Further, the training of the Long Short-Term Memory (LSTM) network model, Bidirectional LSM network model, or Transformer model using the model training data employs learning rate scheduling, gradient pruning, and early stopping strategies for model training, and adjusts hyperparameters through cross-validation to obtain the converged model parameters, including:
[0042] The computational graph structure of the Long Short-Term Memory Network model, the Bidirectional Long Short-Term Memory Network model, or the Transformer model is represented as a directed tree, where nodes represent computational units in the model, and edges represent data flow and dependencies.
[0043] Based on the forward propagation and gradient backpropagation characteristics of the model, the gradient norm, feature expression entropy and prediction error contribution rate of each node in the directed tree are calculated. Nodes with gradient norm lower than a preset gradient threshold, feature expression entropy lower than a preset entropy threshold, or prediction error contribution rate higher than a preset error threshold are identified as information bottleneck positions. An enhanced edge candidate set is generated for each of the information bottleneck positions, and each enhanced edge connects two nodes in the directed tree that are not originally directly connected.
[0044] Define an augmentation edge utility function to quantify the improvement in model performance after adding each candidate edge in the augmentation edge candidate set. Construct a linear programming relaxation problem, use a rounding strategy better than 2 approximations to convert the fractional solution of the linear programming into an integer solution, and iteratively optimize to select the optimal subset of augmentation edges to add to the directed tree.
[0045] Based on the optimal subset of enhanced edges, the model architecture is modified by adding residual connections, cross-layer connections, or attention mechanisms to form an enhanced connection structure. A parameter initialization strategy adapted to the enhanced connection structure is designed, and a phased training strategy is adopted. First, the original structural parameters of the model are frozen and only the connection weight parameters of the enhanced connection structure are trained. Then, the original structural parameters and the connection weight parameters are jointly optimized. Regularization techniques are applied to control the numerical range of the connection weight parameters to obtain the model parameters that have been trained and converged.
[0046] Furthermore, the defined enhancement edge utility function quantifies the improvement in model performance after adding each candidate edge from the enhancement edge candidate set, including:
[0047] Based on the two nodes connected by each candidate edge in the enhanced edge candidate set, each candidate edge is temporarily added to the directed tree to form a temporarily modified directed tree. According to the temporarily modified directed tree, a corresponding temporary connection structure is added to the model. The model training data is used to perform forward propagation calculation on the model after adding the temporary connection structure. The difference in prediction error before and after adding each candidate edge is calculated on the validation set to obtain the change in prediction error.
[0048] The number of new parameters and the number of new floating-point operations in the model caused by adding temporary connection structures corresponding to each candidate edge are counted to obtain the increase in model parameters and the increase in computational cost.
[0049] Based on the change in prediction error, the increase in model parameters, and the increase in computational cost, the ratio of the change in prediction error to the weighted sum of the increase in model parameters and the increase in computational cost is calculated. This ratio is defined as the utility function of the enhanced edge, and the utility score of each candidate edge in the candidate enhanced edge set is obtained.
[0050] Furthermore, the phased training strategy involves first freezing the original structural parameters of the model and training only the connection weight parameters of the enhanced connection structure, and then jointly optimizing the original structural parameters and the connection weight parameters, including:
[0051] Based on the optimal subset of enhanced edges, the model architecture is modified by adding residual connections, cross-layer connections, or attention mechanisms to the node positions corresponding to the optimal subset of enhanced edges to form an enhanced connection structure, thus obtaining the model after adding the enhanced connection structure.
[0052] Based on the model with the added enhanced connection structure, the original structural parameters of the Long Short-Term Memory Network model, the Bidirectional Long Short-Term Memory Network model, or the Transformer model and the connection weight parameters of the enhanced connection structure are labeled. The original structural parameters are set to an untrainable state, and the connection weight parameters are set to a trainable state. The connection weight parameters are optimized by gradient descent using the model training data. The convergence of the loss function is monitored on the validation set to obtain the connection weight parameters after the first stage of optimization.
[0053] Remove the untrainable state of the original structural parameters, set the original structural parameters and the connection weight parameters optimized in the first stage to trainable state simultaneously, and use the model training data to perform joint gradient descent optimization on the original structural parameters and the connection weight parameters optimized in the first stage to obtain the original structural parameters and connection weight parameters optimized in the second stage.
[0054] Based on the connection weight parameters optimized in the second stage, the L1 norm and L2 norm of the connection weight parameters are calculated, and a regularization loss term is constructed as a weighted combination of the L1 norm and the L2 norm. The regularization loss term is added to the total loss function of the model after adding the enhanced connection structure. The sparsity and numerical value of the connection weight parameters are controlled by adjusting the regularization coefficient, and the model parameters of the training convergence are obtained.
[0055] Further, the process of determining a graded early warning threshold based on historical performance index data, comparing the predicted values of the biochemical treatment performance indicators with the graded early warning threshold to generate early warning information, querying an expert knowledge base based on the early warning information and applying an experience learning model to generate process control suggestions, and optimizing the process control suggestions to form a control scheme includes:
[0056] Based on the statistical analysis and processing requirements of historical performance index data, combined with expert knowledge, the early warning thresholds for different performance indicators are determined, and the thresholds are dynamically adjusted considering seasonal factors and changes in working conditions to obtain graded early warning thresholds.
[0057] The predicted value of the biochemical treatment efficiency index is compared with the graded early warning threshold. When the predicted value of the biochemical treatment efficiency index exceeds the graded early warning threshold and the prediction confidence interval meets the preset requirements, early warning information including early warning time, early warning type, severity and credibility is generated.
[0058] Integrate the experience of experts and process operation procedures in the field of wastewater treatment to establish an expert knowledge base that includes a three-element correlation rule of problem-cause-measure;
[0059] Collect historical process control records and corresponding effect data, and use machine learning methods to mine the relationship patterns between control parameters and effects to obtain an experience learning model;
[0060] Based on the warning information, combined with the current system operating status parameters, the expert knowledge base is queried and the experience learning model is applied to generate process control suggestions that include the adjustment direction, magnitude and timing of aeration volume, reflux ratio and reagent dosage.
[0061] The proposed process control recommendations are optimized through multi-objective optimization to balance treatment effect, energy consumption, and operating cost, forming an optimized control scheme. The optimized control scheme is then transformed into execution guidelines that include execution time, operation sequence, and precautions.
[0062] Furthermore, the statistical analysis and processing techniques based on historical performance indicator data, combined with expert knowledge, determine the early warning thresholds for different performance indicators, and dynamically adjust the thresholds considering seasonal factors and changes in operating conditions, resulting in tiered early warning thresholds, including:
[0063] Based on the historical performance index data, the historical value ranges of ammonia nitrogen concentration, nitrate nitrogen concentration, total nitrogen concentration, chemical oxygen demand (COD), sludge settling ratio, specific oxygen consumption rate (SOH), and nitrification rate are statistically analyzed. Using these values as coordinate axes and their respective historical value ranges as the value ranges for each axis, a multi-dimensional early warning space is constructed. This space is then divided into a grid structure according to a preset discretization step size. The intersections of the grid structure are defined as nodes. Directed edges are established between adjacent nodes to represent early warning state transition relationships. Weights are assigned to each directed edge based on the frequency and impact of state transitions in the historical performance index data. Based on statistical analysis of the historical performance index data, the historical occurrence probability and impact severity score are calculated for each node, resulting in a weighted multi-dimensional early warning space graph structure.
[0064] For each performance index among ammonia nitrogen concentration, nitrate nitrogen concentration, total nitrogen concentration, chemical oxygen demand, sludge settling ratio, specific oxygen consumption rate, and nitrification rate, a corresponding parameterized threshold function is constructed. The parameter set of the parameterized threshold function is defined as a parameter vector θ. The parameter vector θ includes a time-sensitive parameter controlling the influence of prediction lead time on the threshold strictness, a seasonal adjustment parameter controlling the influence of seasonal factors and historical data distribution on the baseline threshold, a dynamic adjustment parameter controlling the influence of current treatment load and process parameters on the allowable fluctuation range, and a cross-constraint parameter controlling the mutual influence between different performance indices. The parameterized threshold functions are combined according to the performance index type to form a threshold function set.
[0065] An objective function is constructed, which takes the parameter vector θ as input variable and comprehensively calculates the false negative rate, false positive rate, and alarm timeliness index generated based on the threshold function set. Cost weights are assigned to the false negative rate, the false positive rate, and the alarm timeliness index respectively. The optimization problem of the objective function is mapped to a graph search problem of finding the optimal path in the weighted multidimensional early warning space graph structure. An adaptive A-star algorithm or Monte Carlo tree search is applied to search in the parameter space of the parameter vector θ. During the search process, dynamic programming is applied to store intermediate results and pruning techniques are used to eliminate low-value search branches to obtain the optimal parameter θ star that makes the objective function reach the optimal value.
[0066] The optimal parameter θ is substituted into the threshold function set, which is then deployed for real-time threshold calculation. An online learning mechanism is established, and comparative data between actual warning results and real anomalies are collected as feedback on the warning effect. Based on the warning effect feedback, the parameter vector θ is incrementally updated using gradient descent or Bayesian optimization. Three severity levels—mild, moderate, and severe—are set based on the thresholds output by the threshold function set. For each severity level, two confidence levels—high certainty and low certainty—are set, constructing a multi-level warning system with six levels to obtain the graded warning thresholds.
[0067] This invention also provides a biochemical treatment efficiency prediction system based on the dynamic evolution of water quality fingerprints, comprising:
[0068] The water quality fingerprint acquisition and feature extraction module is used to acquire spectral data of the influent of the sewage treatment system and preprocess it to obtain a standardized spectral matrix. Based on the standardized spectral matrix, spectral feature parameters are extracted to form a water quality fingerprint feature vector. The continuously acquired water quality fingerprint feature vectors are organized according to time to construct a water quality fingerprint time series dataset.
[0069] The efficiency index collection and training data construction module is used to collect the efficiency index of the effluent from the biochemical treatment system and perform standardization processing to obtain standardized efficiency index data. Based on the water quality fingerprint time series dataset and the standardized efficiency index data, a time correspondence is established and a training dataset is constructed. Feature engineering processing is performed on the training dataset to obtain model training data.
[0070] The prediction model training and prediction module is used to train and optimize the prediction model using the model training data to obtain an optimized prediction model, receive real-time water quality fingerprint time series data to construct real-time prediction data, and use the optimized prediction model to predict the real-time prediction data and output the predicted value of the biochemical treatment efficiency index.
[0071] The early warning and control decision-making module is used to determine the graded early warning threshold based on historical performance index data, compare the predicted value of the biochemical treatment performance index with the graded early warning threshold to generate early warning information, query the expert knowledge base based on the early warning information and apply the experience learning model to generate process control suggestions, and optimize the process control suggestions to form a control scheme.
[0072] The beneficial effects of this invention include:
[0073] This invention establishes a mapping relationship between the dynamic evolution of water quality fingerprints and the effectiveness of biochemical treatment by continuously collecting and analyzing time-series data of influent water quality fingerprints and combining deep learning technology. This enables advanced prediction and targeted control of future biochemical treatment effectiveness, and has the advantages of predictability, accuracy and practicality. Attached Figure Description
[0074] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0075] Figure 1 This is a flowchart illustrating the biochemical treatment efficiency prediction method based on the dynamic evolution of water quality fingerprints according to the present invention.
[0076] Figure 2 This is a schematic diagram of the training and prediction method for the biochemical treatment efficiency index prediction model of the present invention.
[0077] Figure 3 This is a schematic diagram of the structure of the biochemical treatment efficiency prediction system based on the dynamic evolution of water quality fingerprints according to the present invention. Detailed Implementation
[0078] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0079] Example 1
[0080] like Figure 1 As shown, this invention provides a method for predicting the effectiveness of biochemical treatment based on the dynamic evolution of water quality fingerprints, comprising:
[0081] Step S1: Obtain spectral data of the influent of the sewage treatment system and preprocess it to obtain a standardized spectral matrix. Extract spectral feature parameters based on the standardized spectral matrix to form a water quality fingerprint feature vector. Organize the continuously collected water quality fingerprint feature vectors according to time to construct a water quality fingerprint time series dataset.
[0082] Step S2: Collect the performance indicators of the effluent from the biochemical treatment system and perform standardization processing to obtain standardized performance indicator data. Based on the water quality fingerprint time series dataset and the standardized performance indicator data, establish a time correspondence and construct a training dataset. Perform feature engineering processing on the training dataset to obtain model training data.
[0083] Step S3: Train the prediction model using the model training data and optimize it to obtain an optimized prediction model. Receive real-time water quality fingerprint time series data to construct real-time prediction data. Use the optimized prediction model to predict the real-time prediction data and output the predicted value of the biochemical treatment efficiency index.
[0084] Step S4: Determine the graded early warning threshold based on historical performance index data, compare the predicted value of the biochemical treatment performance index with the graded early warning threshold to generate early warning information, query the expert knowledge base based on the early warning information and apply the experience learning model to generate process control suggestions, and optimize the process control suggestions to form a control scheme.
[0085] Specifically, a water quality fingerprint time-series dataset is established to provide input features for the prediction model. Multispectral or hyperspectral online analyzers are installed at the inlet of the wastewater treatment plant to continuously collect spectral data of the influent, typically at a frequency of once per hour or higher. The collected raw spectral data undergoes a series of preprocessing steps, including noise removal, baseline drift correction, and elimination of scattering effects, to obtain a standardized spectral matrix. Based on the standardized spectral matrix, key spectral feature parameters are extracted, such as absorbance at characteristic wavelengths, reflectance, slope of specific bands, area ratio, and spectral derivative features, to construct a water quality fingerprint feature vector representing the water quality characteristics. These feature vectors can effectively capture information on pollutants such as organic matter, nitrogen compounds, and phosphorus compounds in the water without requiring complex physicochemical analysis. The system organizes the continuously collected water quality fingerprint feature vectors in chronological order to form a time-series dataset reflecting the dynamic changes in influent water quality. This dataset not only contains single-point water quality information but, more importantly, records the evolution patterns of water quality parameters over time. This is crucial for predicting the response of the biochemical system, as the effectiveness of biochemical treatment often depends on the trend and historical state of water quality changes, rather than solely on the current state.
[0086] Simultaneously, a performance indicator monitoring and data processing workflow was established to provide training targets for the predictive model. Key performance indicator data reflecting treatment efficiency were periodically collected at the effluent outlet of the biological treatment system, such as effluent water quality indicators like ammonia nitrogen concentration, nitrate nitrogen concentration, total nitrogen concentration, and chemical oxygen demand (COD), as well as indicators reflecting system biological activity like sludge settling ratio, specific oxygen consumption rate, and nitrification rate. These raw indicator data were standardized, including unit unification and numerical normalization, to eliminate dimensional differences between different indicators and facilitate subsequent model processing. Then, based on the hydraulic retention time characteristics of the biological treatment system, a time-series correlation was established between influent water quality fingerprint data and effluent performance indicators. This typically requires considering an 8-24 hour delay effect, meaning that the current effluent performance is influenced by the influent characteristics 8-24 hours prior. Based on this time-series correlation, the water quality fingerprint time-series data was paired with subsequent performance indicator data at corresponding time points to construct a complete training dataset. Feature engineering is performed on the constructed training dataset, including sliding window design, temporal feature extraction, outlier handling, and missing value imputation, to enhance the representativeness and learnability of the data and ultimately obtain structured model training data.
[0087] Next, we move into the model training and prediction phase to map water quality fingerprints to future performance. We select suitable deep learning models for time-series prediction tasks, such as Long Short-Term Memory (LSTM), Bidirectional Long Short-Term Memory (BiLSTM), or Transformer models, which can effectively capture temporal dependencies in long-sequence data. Using the prepared model training data, we train the selected network architecture, employing techniques such as learning rate scheduling, gradient pruning, and early stopping to improve training effectiveness. We also adjust network hyperparameters through cross-validation. We evaluate model performance on independent validation datasets and optimize the model based on the evaluation results, including network structure adjustment, ensemble learning, and parameter regularization, ultimately obtaining an optimized prediction model. In the practical application phase, we continuously receive real-time collected water quality fingerprint data and apply the same preprocessing and feature extraction procedures as in the training phase to construct the input data required for real-time prediction. We input this real-time data into the optimized prediction model for forward computation, outputting predicted values of biochemical treatment performance indicators at specific future time points (e.g., 4 hours later, 8 hours later), along with the confidence interval of the prediction, quantifying the uncertainty of the prediction. These forecasts provide crucial information for subsequent early warning and control decisions.
[0088] Finally, intelligent early warning and control are implemented based on the prediction results to achieve proactive prevention and optimized operation. By analyzing the distribution characteristics of historical performance index data, combined with treatment process requirements and expert experience, graded early warning thresholds for each performance index are determined, typically including three warning levels: mild, moderate, and severe. Considering the impact of seasonal factors and changes in operating conditions, these thresholds are dynamically adjusted to adapt to early warning needs under different conditions. The predicted values of performance indicators output by the prediction model are compared with the currently applicable graded early warning thresholds. When the predicted value exceeds the threshold and the prediction confidence interval meets the reliability requirements, structured early warning information is generated, including elements such as early warning time, type, severity, and credibility. Based on the generated early warning information, a pre-established wastewater treatment expert knowledge base is automatically queried. This knowledge base contains a large number of problem-cause-response ternary association rules. At the same time, an experience-based learning model learned from historical control records is applied to analyze the possible effects of adjusting different control parameters under the current conditions. Combining the knowledge base query results and the experience model analysis results, preliminary process control suggestions are generated, including the adjustment direction, magnitude, and timing of key parameters (such as aeration rate, reflux ratio, and reagent dosage). Finally, these preliminary suggestions are optimized for multiple objectives, balancing energy consumption and operating costs while ensuring treatment effectiveness, to form the final optimized control plan, which is then transformed into detailed operational guidelines for easy implementation by operators.
[0089] This invention establishes a mapping relationship between the dynamic evolution of water quality fingerprints and the efficiency of biochemical treatment, enabling accurate prediction and early warning of the future state of the treatment system. This breaks through the traditional passive management model that relies on delayed monitoring, providing technical support for the intelligent and low-energy operation of wastewater treatment plants. The core advantage of the system lies in its integration of technologies from multiple fields, including spectral analysis, time-series deep learning, and knowledge engineering, constructing a complete closed loop from data acquisition and pattern recognition to prediction, early warning, and intelligent control. This allows it to cope with complex and ever-changing influent water quality challenges, ensuring the stable and efficient operation of the biochemical treatment system. Through continuous data accumulation and model optimization, the system's predictive accuracy and control effectiveness will continuously improve, ultimately achieving comprehensive intelligent management of the wastewater treatment process.
[0090] Example 2
[0091] In this embodiment, the process of acquiring and preprocessing the spectral data of the wastewater influent to obtain a standardized spectral matrix, extracting spectral feature parameters based on the standardized spectral matrix to form a water quality fingerprint feature vector, and organizing the continuously acquired water quality fingerprint feature vectors according to time to construct a water quality fingerprint time-series dataset includes:
[0092] An ultraviolet-visible spectrometer and a three-dimensional fluorescence spectrometer were installed at the inlet of the wastewater treatment system. An automatic cleaning device and a calibration system were configured, and the sampling frequency was set to once every 30 minutes to acquire the raw spectral data stream.
[0093] The raw spectral data stream is subjected to scattered light correction, internal filter effect correction, baseline drift correction and instrument response normalization to obtain a standardized ultraviolet-visible absorption spectral matrix and a three-dimensional fluorescence spectral matrix.
[0094] Based on the standardized UV-Vis absorption spectral matrix and the three-dimensional fluorescence spectral matrix, the characteristic wavelength absorbance, slope index, and spectral integral area of the UV-Vis spectrum, as well as the fluorescence region volume, fluorescence peak intensity ratio, and parallel factor analysis component score of the three-dimensional fluorescence spectrum, are extracted to form a multidimensional water quality fingerprint feature vector.
[0095] The continuously collected multidimensional water quality fingerprint feature vectors are organized by timestamps to construct a time-series feature matrix with a sliding time window, thus obtaining the water quality fingerprint time-series dataset.
[0096] Specifically, high-precision ultraviolet-visible spectrometers and three-dimensional fluorescence spectrometers are installed at the inlet of the wastewater treatment system as online monitoring devices. The ultraviolet-visible spectrometer primarily monitors the absorption spectrum in the wavelength range of 200-800 nm, effectively capturing the characteristic absorption of dissolved organic matter, nitrates, nitrites, and other substances in the water. The three-dimensional fluorescence spectrometer, through excitation-emission matrix measurements, monitors the three-dimensional characteristics of fluorescent substances in the water, exhibiting particularly good response characteristics to humic substances, proteins, and microbial metabolites. To ensure the reliability and continuity of monitoring data, an automatic cleaning device is installed on the monitoring system to remove any microbial films or suspended solids that may adhere to the probe surface through periodic rinsing or mechanical wiping. Simultaneously, an automatic calibration system is configured to periodically calibrate the instrument's zero point and range using standard solutions, eliminating the effects of instrument drift. The monitoring system is set to a sampling frequency of once every 30 minutes to ensure sufficient data acquisition density to capture dynamic changes in water quality, while avoiding data redundancy and storage pressure caused by overly frequent sampling. Through this monitoring system, the raw spectral data stream of the influent is continuously acquired, providing a fundamental data source for subsequent analysis.
[0097] The acquired raw spectral data stream undergoes systematic preprocessing to eliminate interference from various physical and instrumental factors, resulting in a standardized spectral matrix. First, scattered light correction is performed, identifying and removing Rayleigh and Raman scattering in the three-dimensional fluorescence spectrum. Rayleigh scattering refers to scattered light of the same wavelength as the incident light produced when the incident light is elastically scattered by the sample, typically appearing as a diagonal band in the fluorescence matrix; while Raman scattering is inelastic scattering caused by water molecule vibrations, appearing as a secondary scattering band parallel to the diagonal. These scattering bands are removed using interpolation or spectral difference methods to avoid interference with the true fluorescence signal. Next, internal filtering effect correction is performed. Internal filtering effect refers to the phenomenon where the absorption of excitation or emission light by certain substances in the sample leads to an underestimation of fluorescence intensity. Based on the absorbance values of the sample at the corresponding wavelengths, a mathematical model is used to compensate and adjust the fluorescence intensity. Finally, baseline drift correction is performed, eliminating baseline drift by subtracting the spectral signal of a blank water sample or applying a polynomial fitting method, ensuring a consistent baseline for spectral data acquired at different time points. Finally, instrument response normalization is performed. Based on the instrument sensitivity curve or the response of the standard reference material, the spectral intensity is normalized to ensure data comparability and consistency. Through this series of preprocessing steps, standardized UV-Vis absorption spectral matrix and three-dimensional fluorescence spectral matrix are obtained, laying the foundation for subsequent feature extraction.
[0098] Based on the preprocessed standardized spectral matrix, a series of spectral characteristic parameters that can characterize water quality properties are extracted. For UV-Vis spectroscopy, absorbance values at characteristic wavelengths are extracted, such as the absorbance value at 254 nm (UV254), which is closely related to the content of dissolved organic matter in the water. The slope index is calculated, which is the ratio of the spectral slope within a specific wavelength range (such as 275-295 nm and 350-400 nm), reflecting the molecular weight distribution and aromaticity of organic matter. The spectral integral area is calculated, and a comprehensive index reflecting the total organic load is obtained by integrating the entire spectral curve or a specific interval. For three-dimensional fluorescence spectroscopy, the fluorescence region volume is calculated, and the three-dimensional fluorescence matrix is divided into multiple regions (such as protein-like regions, humic substance-like regions, etc.). The fluorescence intensity integral value of each region is calculated to characterize the relative content of different types of organic matter. The fluorescence peak intensity ratio is calculated, such as the intensity ratio of protein-like fluorescence peaks to humic substance-like fluorescence peaks, reflecting the changes in organic matter composition. Parallel factor analysis (PARAFAC) is applied to decompose the three-dimensional fluorescence data to obtain component scores, i.e., the relative content of each characteristic component in the sample. PARAFAC is a multivariate statistical method that decomposes complex three-dimensional fluorescence data into several independent fluorescent components, each with specific spectral characteristics and relative concentrations. This helps in identifying and quantifying different types of fluorescent substances in water. By extracting and calculating the aforementioned feature parameters, high-dimensional spectral data is transformed into multidimensional water quality fingerprint feature vectors with clear physical and chemical meanings, facilitating subsequent pattern recognition and predictive analysis.
[0099] The continuously collected multidimensional water quality fingerprint feature vectors are organized according to timestamps to construct a time-series feature matrix with a sliding time window. Specifically, each collected water quality fingerprint feature vector serves as a time point in the time-series data. As monitoring continues, these feature vectors are arranged chronologically to form time-series data. A sliding time window technique is used to process this time-series data; a fixed-length time window (e.g., data from the past 12 hours) is selected, and the window slides and updates over time, ensuring that it always contains the most recent n time points. This sliding window method can capture the dynamic trends of water quality changes while limiting the amount of data analyzed in a single session, thus improving computational efficiency. In the constructed time-series feature matrix, each row represents a time point, and each column represents a feature parameter. The time span of the matrix is determined by the size of the sliding window. In this way, a water quality fingerprint time-series dataset reflecting the dynamic evolution characteristics of water quality is obtained, providing a structured data foundation for subsequent time-series feature analysis and predictive model training.
[0100] Water fingerprints, in this context, refer to a set of multidimensional features acquired through methods such as spectral analysis that characterize water properties, similar to a human fingerprint used for identification. In this invention, water fingerprints primarily consist of ultraviolet-visible spectral and three-dimensional fluorescence spectral features, reflecting the types, quantities, and structural characteristics of organic and certain inorganic substances in the water. The formation principle of water fingerprints is based on the different absorption or fluorescence emission characteristics of different substances at specific wavelengths of light. By analyzing these characteristics, a "fingerprint" of the composition of substances in the water can be obtained. The steps for constructing a water fingerprint include: acquiring raw spectral data, performing spectral preprocessing, extracting feature parameters, and forming feature vectors. The dynamic evolution of water fingerprints refers to the changing patterns of these features over time, reflecting the dynamic changes in water quality.
[0101] Example 3
[0102] In this embodiment, based on the standardized UV-Vis absorption spectral matrix and the three-dimensional fluorescence spectral matrix, the characteristic wavelength absorbance, slope index, and spectral integral area of the UV-Vis spectrum, as well as the fluorescence region volume, fluorescence peak intensity ratio, and parallel factor analysis component scores of the three-dimensional fluorescence spectrum, are extracted to form a multidimensional water quality fingerprint feature vector, including:
[0103] The response values of each wavelength or wavelength pair in the standardized UV-Vis absorption spectral matrix and the three-dimensional fluorescence spectral matrix are used as the dimensions of the high-dimensional feature space. The high-dimensional features are projected onto the two-dimensional or three-dimensional space through principal component analysis or t-SNE dimensionality reduction technology to form a set of geometric feature points.
[0104] Based on the principles of water chemistry, in the set of geometric feature points, when the water quality characteristics represented by two feature points are correlated or causally related during the biochemical treatment process, a connection is established between the two feature points to form an intersection diagram.
[0105] Based on the intersection diagram, feature points are selected from the set of geometric feature points through correlation analysis as the terminal nodes of the Steiner tree. The parameter λ is introduced to adjust the balance between the complexity of the tree and the fitting effect. Dynamic programming and divide-and-conquer strategies are applied, and the optimal value of parameter λ is found through gradient descent or simulated annealing to obtain the optimal Steiner tree structure.
[0106] Using the optimal Steiner tree structure, the depth, width, and branching factor of the tree are calculated as topological features, the path length and path strength between terminal nodes are extracted as path features, the centrality index of each node in the tree is calculated as node centrality features, and dynamic evolution parameters are extracted by comparing the changes in the Steiner tree structure at different time points.
[0107] The multidimensional water quality fingerprint feature vector is obtained by combining the topological features, path features, node centrality features, dynamic evolution parameters, characteristic wavelength absorbance, slope exponent, spectral integral area, fluorescence region volume, fluorescence peak intensity ratio, and parallel factor analysis component scores.
[0108] Specifically, the standardized UV-Vis absorption spectral matrix and three-dimensional fluorescence spectral matrix data are represented in a high-dimensional feature space and then subjected to dimensionality reduction processing. The original spectral data has extremely high dimensionality; for example, UV-Vis spectra may contain absorbance values at hundreds of wavelengths, while three-dimensional fluorescence spectra may contain intensity values for tens of thousands of excitation-emission wavelength pairs. The response value of each wavelength point or wavelength pair is considered as one dimension in the high-dimensional feature space, and the entire spectral data constitutes a point in that space. Because this representation is too high-dimensional, it is not conducive to direct analysis and visualization. Therefore, dimensionality reduction techniques are used to project the data into a lower-dimensional space. Principal Component Analysis (PCA) is a linear dimensionality reduction method that projects high-dimensional data into a lower-dimensional space composed of several principal components by identifying the main directions of variation (principal components). Meanwhile, t-distributed random neighborhood embedding (t-SNE) is a non-linear dimensionality reduction technique, particularly suitable for preserving local structural relationships in high-dimensional data and better showcasing clustering characteristics. These dimensionality reduction techniques project complex spectral data into two-dimensional or three-dimensional space, forming a more intuitive set of geometric feature points. Each point represents the spectral characteristics of a water sample, and the distance between points reflects the similarity of water quality characteristics.
[0109] Based on the principles of water chemistry, connections are established between points in the geometric feature point set formed after dimensionality reduction, forming an intersection graph. An intersection graph is a special graph structure where nodes represent feature points and edges represent relationships between nodes. In this embodiment, a connection is established between two feature points when the water quality characteristics they represent have a correlation or causal relationship during the biochemical treatment process. For example, the presence of certain organic matter may affect the conversion efficiency of ammonia nitrogen, or the intensity of certain fluorescent substances may be related to microbial activity; these chemical or biological correlations are represented as edges in the intersection graph. Methods for determining whether a correlation exists between feature points include: calculating the Pearson correlation coefficient or Spearman's rank correlation coefficient between feature points, establishing a connection when the correlation coefficient exceeds a preset threshold; manually defining connections between specific feature points based on domain expert knowledge and known chemical reaction relationships or biodegradation pathways between water quality parameters; and analyzing the temporal patterns of feature point changes in historical data, establishing a connection when the changing trends of two feature points show a significant time dependency. The intersection map constructed in this way not only contains spatial distribution information of water quality characteristics, but also incorporates the chemical and biological relationships between water quality parameters, providing a theoretical basis for the subsequent construction of Steiner trees.
[0110] Based on the constructed intersection graph, the Steiner tree algorithm is applied to extract the most representative water quality feature correlation structure. A Steiner tree is an algorithm that finds the minimum connection tree in a given set of points, connecting all key nodes at minimal cost. In this embodiment, firstly, correlation analysis is used to select feature points highly correlated with target biochemical performance indicators (such as ammonia nitrogen removal rate, total nitrogen removal rate, etc.) from the geometric feature point set as the terminal nodes of the Steiner tree. These terminal nodes represent water quality features that have a significant impact on the prediction target. To balance the complexity of the tree with the fitting effect, a parameter λ is introduced as an adjustment factor. A larger λ value results in a simpler tree structure but may reduce fitting accuracy; a smaller λ value results in a more complex tree structure but may more accurately reflect data relationships. Determining the optimal λ value is an optimization problem. It can be achieved by continuously adjusting the λ value on the training data using gradient descent, evaluating the performance of the corresponding Steiner tree on the prediction task, and selecting the λ value that minimizes the prediction error. Alternatively, simulated annealing can be used to perform a random search in the parameter space, accepting suboptimal solutions with a certain probability to escape local optima, and eventually converging to the globally optimal or near-optimal λ value. By applying dynamic programming and a divide-and-conquer strategy to optimize the computation process, the optimal Steiner tree structure that best represents the relationship between water quality fingerprint features is finally obtained.
[0111] Using the obtained optimal Steiner tree structure, a series of structured feature parameters that characterize the relationships between water quality features are extracted. First, the topological features of the tree are calculated, including the tree depth (the maximum distance from the root node to the farthest leaf node), width (the number of nodes in the layer with the most nodes), and branching factor (the average number of child nodes in non-leaf nodes). These parameters reflect the complexity and hierarchy of the relationship structure between water quality features. Then, path features between terminal nodes are extracted, including path length (the minimum number of edges required to connect two terminal nodes) and path strength (the geometric mean of the weights of all edges on the path). These features characterize the strength and complexity of the relationship between key water quality parameters. Next, the centrality indices of each node in the tree are calculated, such as degree centrality (the number of edges directly connected to the node), proximity centrality (the reciprocal of the average shortest path length from the node to all other nodes), and betweenness centrality (the proportion of shortest paths passing through the node to the total number of shortest paths). These centrality indices reflect the importance and influence of different water quality features in the entire system. Finally, by comparing the changes in the Steiner tree structure at different time points, dynamic evolution parameters were extracted, such as tree structure similarity (the degree of matching between tree structures at two time points), key edge change rate (stability of key connection relationships), and topological feature change rate (time derivative of tree topological features). These dynamic parameters capture the patterns of water quality feature association structure evolution over time.
[0112] By combining the structured features extracted based on Steiner trees with traditional spectral feature parameters, an enhanced multidimensional water quality fingerprint feature vector is formed. Specifically, graph-based features such as topological features (e.g., tree depth, width, branching factor), path features (e.g., path length and path strength between terminal nodes), node centrality features (e.g., degree centrality, proximity centrality, betweenness centrality), and dynamic evolution parameters (e.g., tree structure similarity, critical edge change rate, topological feature change rate) are combined with traditional UV-Vis spectral features (e.g., characteristic wavelength absorbance, slope exponent, spectral integral area) and three-dimensional fluorescence spectral features (e.g., fluorescence region volume, fluorescence peak intensity ratio, parallel factor analysis component score). This combination not only retains the water quality chemical composition information contained in traditional spectral features but also introduces Steiner tree-based structural features, enabling the characterization of complex relationships and dynamic evolution patterns among water quality components. The resulting multidimensional water quality fingerprint feature vector has stronger characterization capabilities and predictive value, comprehensively capturing the static composition and dynamic changes of water quality, providing a more comprehensive and in-depth data foundation for subsequent prediction of biochemical treatment efficiency.
[0113] Steiner trees, a concept in graph theory, refer to the minimum-weighted tree connecting a specified set of terminal nodes in a given graph. Unlike minimum spanning trees, Steiner trees allow the introduction of non-terminal nodes (called Steiner points) as connection points to reduce the overall connection cost. In this invention, Steiner trees are used to model the association structure between water quality features. Terminal nodes represent key water quality features, and the tree structure reflects the optimal connection relationships between these features. The principle of constructing a Steiner tree is to determine the optimal connection structure by minimizing the connection cost (such as the distance between feature points or the inverse of the association strength). The construction steps include: determining the set of terminal nodes, setting the connection cost function, applying the Steiner tree algorithm to solve for the optimal tree structure, and extracting the structural features of the tree.
[0114] An intersection graph is a graph structure built upon a set of geometric feature points, based on the relationships between those points. Nodes in the graph represent water quality feature points, and edges represent correlations or causal relationships between features. The construction principle of an intersection graph is to transform abstract water quality feature relationships into a visual graph structure for analysis using graph theory algorithms. The construction steps include: determining the nodes (feature point set) of the graph, defining edge connection conditions (such as correlation coefficient thresholds or expert knowledge rules), establishing connections between relevant nodes based on the conditions, and assigning weights to the edges (such as correlation coefficient strength or expert scores). The intersection graph provides the basic graph structure for subsequent Steiner tree construction, reflecting the potential network of relationships between water quality features.
[0115] Example 4
[0116] In this embodiment, the application of dynamic programming and divide-and-conquer strategy includes:
[0117] The intersection graph is decomposed into multiple subgraphs, and the local optimal Steiner tree is solved independently for each subgraph.
[0118] The intermediate results of solving each subgraph are recorded using a dynamic programming method. These intermediate results include the optimal connection method within the subgraph and the corresponding objective function value.
[0119] By merging the locally optimal Steiner trees of each subgraph using a divide-and-conquer strategy, calculating the connection cost between subgraphs based on the intermediate results, selecting the connection scheme that optimizes the global objective function, and constructing a global Steiner tree structure.
[0120] Specifically, the intersection graph is decomposed into multiple smaller subgraphs to reduce problem complexity, and the locally optimal Steiner tree is solved independently for each subgraph. Intersection graph decomposition is a key step in solving large-scale Steiner tree problems; it breaks down a large problem whose computational complexity may grow exponentially into multiple smaller problems that can be solved in a reasonable amount of time. The decomposition process is based on the structural characteristics of the graph and mainly employs the following strategies: Boundary point decomposition, which identifies cut vertices (points whose removal would split the graph into multiple connected components) or cut edges (edges whose removal would split the graph) in the intersection graph, and decomposes the graph into subgraphs using these cut vertices or cut edges as boundaries; Community detection, which uses community detection methods such as the Louvain algorithm or the Girvan-Newman algorithm to identify tightly connected groups (communities) in the intersection graph, and treats each community as a subgraph; Spectral clustering, which maps the nodes in the graph to a low-dimensional space and clusters them based on the eigenvectors of the graph's Laplacian matrix, with each cluster forming a subgraph; and Domain-driven decomposition, which divides the subgraphs according to the chemical properties or biological functions of water quality characteristics, such as assigning organic matter indicators, nutrient indicators, and microbial activity indicators to different subgraphs. After decomposition, Steiner tree algorithms, such as Prim's algorithm or variants of Kruskal's algorithm, are applied independently to each subgraph to solve for the locally optimal Steiner tree connecting the terminal nodes in the subgraph. Since the subgraph is significantly smaller than the original graph, the computational complexity is greatly reduced, making it possible to find a high-quality solution in a finite amount of time.
[0121] Dynamic programming is employed to record intermediate results during the solution process of each subgraph, avoiding redundant calculations and improving algorithm efficiency. Dynamic programming is an effective method for solving problems with overlapping subproblems and optimal substructure characteristics, and it is particularly applicable in the Steiner tree problem. When applying the Steiner tree algorithm to each subgraph, the following intermediate results are recorded and stored: the optimal connection method within the subgraph, i.e., for each possible subset of terminal nodes in the subgraph, recording the optimal tree structure connecting these nodes; the corresponding objective function values, including evaluation metrics such as the total weight of the tree, the average edge weight of the tree, and the maximum edge weight of the tree; subtree characteristic parameters, such as the tree's depth, width, node degree distribution, and other topological features; and boundary node connection information, recording the connection relationships between the boundary nodes of the subgraph and internal nodes, providing a basis for subsequent subgraph merging. These intermediate results are stored in a hash table or cache matrix. When the algorithm needs to repeatedly calculate a subproblem, it can directly look up the result in the table without recalculating. The core idea of dynamic programming is to decompose complex problems into simpler subproblems and store the solutions to the subproblems to avoid redundant calculations, significantly improving algorithm efficiency, especially when dealing with large-scale Steiner tree problems.
[0122] By merging the locally optimal Steiner trees of each subgraph using a divide-and-conquer strategy, a globally optimal or near-optimal Steiner tree structure can be formed. The divide-and-conquer strategy is a classic method for handling complex problems, its core idea being "divide and conquer": breaking down the problem into multiple subproblems, solving them separately, and then merging the results. In the Steiner tree problem, subgraph merging is a key step in the divide-and-conquer strategy. The merging process includes the following stages: First, construct a subgraph connection graph, treating each subgraph as a node and the potential connections between subgraphs as edges, with edge weights determined based on the connection costs of the boundary nodes between subgraphs; then, based on the connection costs between subgraphs, calculate the minimum spanning tree or the minimum weight perfect match to determine the optimal connection scheme between subgraphs; next, according to the determined connection scheme, connect the locally optimal Steiner trees of each subgraph through boundary nodes to form a preliminary global tree structure; finally, optimize and adjust the merged global tree, such as pruning (removing unnecessary edges or nodes connecting terminal nodes) or reconnecting (adjusting the connection method to reduce the total cost), to obtain the final global Steiner tree structure. During the subgraph merging process, the previously recorded intermediate results, especially the connection information of boundary nodes, are fully utilized to avoid redundant calculations and ensure algorithm efficiency. Through this bottom-up merging strategy, a Steiner tree structure with optimal or near-optimal global objective function is finally constructed, effectively balancing computational complexity and solution quality.
[0123] In practical applications, the above decomposition-solution-merging process may require multiple iterations to obtain a better solution. After each iteration, the quality of the global Steiner tree is evaluated. If it does not meet the preset quality criteria or convergence conditions, the decomposition strategy, optimization parameters, or merging scheme are adjusted, and the next iteration is performed. The iterative process usually converges to a stable solution, which approximates the global optimum under computational resource constraints. By combining dynamic programming and divide-and-conquer strategies, the exponential complexity of the Steiner tree problem is successfully transformed into an approximate problem solvable in polynomial time, making it possible to process large-scale spectral data in real time in practical water quality monitoring systems. The constructed global Steiner tree structure effectively captures the core correlations between water quality fingerprint features, providing strong data support for water quality dynamic change analysis and biochemical treatment efficiency prediction.
[0124] Dynamic programming is an algorithmic strategy that decomposes a complex problem into overlapping subproblems and stores the solutions to these subproblems to avoid redundant computation. Its core ideas are "optimal substructure" and "overlapping subproblems." In the Steiner tree problem, the application principle of dynamic programming is to decompose the large Steiner tree problem into smaller Steiner tree problems on subgraphs and store intermediate results to improve computational efficiency. The implementation steps include: defining the state (e.g., the subgraph and its set of terminal nodes), establishing state transition equations (e.g., the relationship between the optimal solution after adding a node and the previous state), designing a memoized storage structure (e.g., a hash table), solving the subproblems in a specific order and recording the results, and constructing the final solution based on the subproblem solutions. Dynamic programming significantly reduces the computational complexity of the Steiner tree problem, making it possible to solve large-scale problems.
[0125] The divide-and-conquer strategy is an algorithmic design paradigm that decomposes a large problem into multiple independent smaller problems, solves each problem individually, and then combines the results. In this invention, the principle of the divide-and-conquer strategy is to decompose a large-scale intersection graph into multiple subgraphs, solve the locally optimal Steiner tree on each subgraph, and then construct a global tree structure through a merging strategy. The implementation steps include: problem decomposition (decomposing the graph into subgraphs), independent solving (solving the Steiner tree problem on each subgraph), result merging (merging the subgraph solutions into a global solution), and solution optimization (adjusting the merged results to improve global optimality). The combination of the divide-and-conquer strategy and dynamic programming effectively balances computational complexity and solution quality, making it suitable for processing complex correlation structures in large-scale water quality fingerprint data.
[0126] Example 5
[0127] In this embodiment, the efficiency indicators of the effluent from the biochemical treatment system are collected and standardized to obtain standardized efficiency indicator data. A time correspondence is established based on the water quality fingerprint time-series dataset and the standardized efficiency indicator data, and a training dataset is constructed. Feature engineering is performed on the training dataset to obtain model training data, including:
[0128] The concentrations of ammonia nitrogen, nitrate nitrogen, total nitrogen, chemical oxygen demand (COD), sludge settling ratio, specific oxygen consumption rate (SOH), and nitrification rate at the effluent of the biochemical treatment system are collected. The units of the ammonia nitrogen concentration, nitrate nitrogen concentration, total nitrogen concentration, COD, sludge settling ratio, specific SOH, and nitrification rate are standardized to obtain standardized performance index data.
[0129] Based on the hydraulic residence time, an influent timestamp is added to each performance index data point in the standardized performance index data. The influent timestamp is associated with the timestamp in the water quality fingerprint time series dataset to establish an influent-effluent correspondence.
[0130] Based on the water quality fingerprint time series dataset and the standardized performance index data with the added water inflow timestamp, a data pair of correspondences between the water quality fingerprint sequence at time t and n hours before it and the performance index at time t+m is constructed by timestamp matching, where m represents the prediction lead time.
[0131] Outlier identification and missing value processing are performed on the corresponding data pairs. Outlier data are processed using moving median filtering, interpolation, or statistical methods to obtain a quality-controlled training dataset.
[0132] A sliding window design is applied to the quality-controlled training dataset, dividing the water quality fingerprint sequence into multiple time windows of fixed length. Temporal difference feature extraction is performed on the water quality fingerprint feature vectors within the time windows, the difference between feature values at adjacent time points is calculated, and periodic feature encoding and time encoding are added to the time windows to obtain structured model training data.
[0133] Specifically, at the effluent outlet of the biological treatment system, several key indicators characterizing treatment efficiency were systematically collected. These included core parameters representing water quality such as ammonia nitrogen concentration, nitrate nitrogen concentration, total nitrogen concentration, and chemical oxygen demand (COD). These parameters directly reflect the effluent's compliance with treatment standards and the biological system's removal efficiency for specific pollutants. Simultaneously, parameters characterizing activated sludge status, such as sludge settling ratio, specific oxygen consumption rate, and nitrification rate, were collected. These parameters reflect the activity and treatment capacity of the microbial community. Ammonia nitrogen concentration characterizes the ammonia nitrogen content in the water and is a direct indicator for evaluating the nitrification effect of the biological treatment system. Nitrate nitrogen concentration characterizes the accumulation of nitrification products. Total nitrogen concentration comprehensively characterizes the total content of various forms of nitrogen in the water, reflecting the overall denitrification effect. Chemical oxygen demand characterizes the organic matter content in the water and is an important indicator for evaluating organic matter removal efficiency. Sludge settling ratio is a parameter characterizing the settling performance of activated sludge, affecting the system's solid-liquid separation effect. Specific oxygen consumption rate characterizes the respiration intensity and metabolic activity of activated sludge. Nitrification rate directly characterizes the activity of the nitrifying bacterial community and the speed of the nitrification reaction. These indicators are collected through a combination of automated monitoring equipment and periodic sampling and analysis, establishing a complete performance indicator monitoring system to provide comprehensive target parameters for subsequent performance prediction.
[0134] The collected performance index data underwent unit unification and numerical standardization to eliminate dimensional differences between different parameters and make the data suitable for model training. First, unit unification was performed, converting parameters with different units into standardized units. For example, ammonia nitrogen, nitrate nitrogen, total nitrogen, and chemical oxygen demand were unified to milligrams per liter (mg / L); sludge settling ratio was expressed as a percentage (%); specific oxygen consumption rate was expressed as milligrams of oxygen per gram of mixed liquor volatile suspended solids per hour (mg O2 / gMLVSS / h); and nitrification rate was expressed as milligrams of ammonia nitrogen per gram of mixed liquor volatile suspended solids per hour (mg NH4⁺-N / gMLVSS / h). After unifying units, numerical standardization is performed. Common standardization methods include: Min-Max standardization, which linearly maps the data to the [0,1] or [-1,1] interval; Z-score standardization, which transforms the data into a distribution with a mean of 0 and a standard deviation of 1; logarithmic transformation, which takes the logarithm of the data to compress the numerical range and is suitable for data with an exponential distribution; and quantile transformation, which performs a non-linear transformation based on the quantile distribution of the data and is insensitive to outliers. The standardization process needs to consider the data distribution characteristics of each indicator and select an appropriate standardization method to ensure that the transformed data distribution is conducive to model training. Through unit unification and numerical standardization, a standardized performance indicator dataset is obtained, eliminating differences in the units and numerical ranges between different indicators, giving each indicator an equal chance of weighting in model training.
[0135] Based on the hydraulic retention time (HRT) of the biochemical treatment system, a time-correlation relationship is established for standardized performance index data, linking effluent indicators with their corresponding influent characteristics. Hydraulic retention time (HRT) refers to the average time water resides in the biochemical treatment system, determining the time delay between influent and effluent. Based on the measured HRT or the actual flow time determined using tracer experiments, the corresponding influent time is calculated for each effluent performance index data point, and an influent timestamp is added. For example, if the average HRT of the system is 8 hours, then the influent time corresponding to the effluent sample collected at time t should be approximately t-8 hours. In actual operation, considering the non-ideal flow regime of the system and the differences in residence time among different process units, it may be necessary to establish more complex time-correlation models, such as dynamic HRT calculations considering flow rate changes, or residence time distribution functions based on tracer experiment results. By adding accurate influent timestamps, the effluent performance index data is correlated with the corresponding time point data in the water quality fingerprint time series dataset, establishing a clear influent-effluent correspondence and providing a clear data foundation for subsequent model training based on causal relationships.
[0136] Based on the established time correspondence, suitable data pairs for time-series prediction tasks are constructed to precisely match the input features (water quality fingerprint sequences) with the prediction targets (future performance indicators). For each performance indicator data point marked with an inflow timestamp, the corresponding time point (denoted as t) and the water quality fingerprint sequence n hours prior are extracted from the water quality fingerprint time-series dataset as input features; simultaneously, the performance indicator at time t+m is used as the prediction target, where m represents the prediction lead time, i.e., how far in advance the model is expected to predict the future performance indicator. For example, if n=12 and m=4, it means using the water quality fingerprint sequence from t-12 hours to time t to predict the performance indicator at t+4 hours, achieving a 4-hour advance warning. The choice of prediction lead time m needs to balance prediction accuracy and warning lead time: a smaller m value usually achieves higher prediction accuracy, but the warning lead time is limited; a larger m value provides a longer lead time, but prediction accuracy may decrease. In practical applications, multiple prediction lead times (such as 2 hours, 4 hours, 6 hours, etc.) can be set to construct multiple corresponding datasets and train models for different prediction time ranges. In this way, a large number of data pairs are generated that correspond to the water quality fingerprint sequence at time t and the preceding n hours → the performance index at time t+m, providing a foundation for the model to learn the mapping relationship between the water quality fingerprint sequence and future performance changes.
[0137] Quality control is performed on the constructed correspondence data pairs to handle outliers and missing values, ensuring data reliability and integrity. First, outlier identification is performed using statistical methods (such as the 3σ rule and box plots), machine learning methods (such as Isolation Forest and Single-Class SVM), or domain rules (such as values exceeding the physically possible range) to identify anomalies in the data. Then, missing value detection is performed, including both complete missing values (no record at a certain time point) and partial missing values (missing some features at a certain time point). For identified outliers, a moving median filter is used for smoothing, replacing the outlier with the median of adjacent data points within a time window. This method has good robustness against sudden outliers. For missing values, an appropriate imputation method is selected based on the temporal characteristics of the data. For example, linear interpolation and spline interpolation are suitable for short-term missing data with stable changes; time-series model interpolation (such as ARIMA) is suitable for data with obvious temporal patterns; or imputation based on similar daily patterns is suitable for data with periodic characteristics. The quality control process must maintain the original characteristics of the data and avoid over-smoothing or introducing spurious patterns. Therefore, the processing intensity and method selection should be based on data characteristics and professional judgment. These processes yield a quality-controlled training dataset, providing a high-quality data foundation for subsequent feature engineering and model training.
[0138] Temporal feature engineering is performed on the quality-controlled training dataset to extract and enhance pattern features in the time-series data, thereby improving the model's ability to learn temporal dependencies. First, a sliding window design is applied, dividing each water quality fingerprint sequence into a uniform time window of a fixed length (e.g., 12 hours). The choice of window length needs to balance capturing short-term fluctuations with understanding long-term trends; the optimal window size is usually determined through cross-validation. Then, temporal difference feature extraction is performed on the water quality fingerprint feature vectors within the time window, calculating the difference (first-order difference Δx) between feature values at adjacent time points. t = x t - x {t-1} The process involves several steps: First, capturing the changing trends and rates of features, these differential features typically reflect the dynamic characteristics of the system. Next, periodic feature encoding is added to identify and encode periodic patterns in the data, such as decomposing time information into hourly, daily, and weekly periods, using sine and cosine functions to preserve the continuous characteristics of the periodicity. Finally, time encoding is added to directly encode time information into features usable by the model, enabling the model to learn time-related patterns, such as using positional encoding or learnable time embedding vectors to represent different time points. Through this series of feature engineering processes, the original time-series data is transformed into structured model training data, which includes both the original water quality fingerprint features and derived features such as time-series differences, periodicity, and time encoding. This comprehensively characterizes the static features and dynamic change patterns of the water quality fingerprint, providing rich information input for the deep learning model.
[0139] Hydraulic retention time (HRT) refers to the average residence time of water in a wastewater treatment system or specific reaction unit, calculated as the ratio of reactor volume to flow rate. In this invention, HRT is a key parameter connecting influent water quality characteristics and effluent efficiency indicators, used to establish a time-to-effect relationship. The principle of HRT is based on the fact that the residence time of fluid in the reactor determines the reaction time and efficiency of the treatment process. The steps to determine the actual HRT include: theoretical calculation based on the design volume and flow rate; tracer experiment, adding tracers (such as fluorescent dyes) to the influent and monitoring the change of tracer concentration in the effluent over time to obtain the residence time distribution; and fluid dynamics simulation, establishing a computational fluid dynamics model of the system to simulate the water flow path and residence time under different operating conditions. Accurate HRT estimation is crucial for establishing the influent-effluent correspondence.
[0140] Forecast lead time refers to the length of time a forecasting model can predict future states, i.e., the time difference between the target prediction time and the current time. In this invention, forecast lead time (m) represents the model's ability to predict biochemical treatment efficiency indicators m hours later using current and historical water quality fingerprint data. The principle behind setting forecast lead time is based on the need to balance forecast accuracy and timely warning: a larger lead time provides operators with more response time, but forecast accuracy usually decreases as the lead time increases. The steps to determine the optimal forecast lead time include: multi-leader testing to evaluate the model's predictive performance under different lead times; system response time analysis, considering the minimum time required for process adjustments to show results; and application scenario requirement assessment, determining the necessary early warning time based on actual operational needs (such as control decisions, reagent preparation time, etc.). A multi-leader strategy is typically adopted, providing short-term, medium-term, and long-term forecasts simultaneously to meet the needs of different application scenarios.
[0141] Temporal difference features are features generated by calculating the differences between data points at adjacent time points in a time series, used to capture the trends and rates of data change. In this invention, temporal difference features are generated by calculating the difference (Δx) between water quality fingerprint feature values at adjacent time points. t = x t - x {t-1} The principle of time-series differencing is to transform the original non-stationary time series into a form that is closer to a stationary series, highlighting the changes in the data rather than absolute values, which helps the model learn the patterns of change. The steps for generating time-series differencing features include: calculating the first-order difference, the difference between adjacent time points; calculating higher-order differences, such as the second-order difference (Δ²x). t = Δx t - Δx {t-1} The time-series difference features are used to capture changes in the rate of change; they are used to calculate differences over specific time intervals, such as the difference at the same time point as the previous day, to eliminate the influence of periodicity; and the difference features are standardized to make the difference features of different variables comparable. Time-series difference features can effectively capture the dynamic changes in water quality fingerprints and are an important input for prediction models.
[0142] Periodic feature encoding is a technique that converts periodic time information into feature vectors usable by models. In this invention, periodic feature encoding is used to characterize regular changes such as hourly, daily, and weekly cycles that may exist in water quality and biochemical treatment systems. The principle of periodic feature encoding is to convert cyclical time features into continuous feature representations through mathematical transformations, avoiding numerical jumps at cycle boundaries (such as from 11 PM to midnight). The encoding steps include: cycle identification, determining the main cycles present in the data (such as 24 hours, 7 days, etc.); sine-cosine transformation, converting the periodic variable x into (sin(2πx / P), cos(2πx / P)), where P is the cycle length; multi-cycle combination, simultaneously encoding multiple possible cycles, such as hourly, daily, weekly, and monthly cycles; and cycle intensity adjustment, adjusting the encoding amplitude according to the importance of different cycles. Through periodic feature encoding, the model can better learn and utilize the periodic patterns in water quality data, improving prediction accuracy.
[0143] Temporal encoding is a technique that converts temporal information into a numerical representation that can be directly used by deep learning models. In this invention, temporal encoding enables models to learn patterns related to absolute time points, such as seasonal variations and weekday patterns. The principle of temporal encoding is to create numerical vectors that can characterize time positions and relative relationships, allowing models to identify and utilize time-dependent patterns. The implementation steps include: positional encoding, such as the position vectors generated by the sine-cosine function used in Transformers, which can express the relative positional relationships in a sequence; temporal embedding, mapping timestamps to low-dimensional dense vectors, which can be pre-defined or learned by the model; temporal feature engineering, decomposing time into multiple features, such as year, month, day, hour, minute, day of the week, and whether it is a holiday; and time-aware attention, introducing time interval information into the attention mechanism, enabling the model to consider the time distance between events. Through temporal encoding, models can better understand and utilize time-related patterns in data, improving their ability to model time-series dependencies.
[0144] Example 6
[0145] In this embodiment, the outlier identification and missing value processing of the corresponding relationship data pairs includes:
[0146] The corresponding data pairs are subjected to quality review to identify outliers that exceed the normal range and missing values in the data records;
[0147] The outliers in the time series are smoothed by using a moving median filter, and missing values are filled by interpolation or by using statistical methods to calculate the estimated values of missing data points, thus obtaining the quality-controlled training dataset.
[0148] Specifically, a systematic quality review is conducted on the established data pairs corresponding to water quality fingerprint sequences and performance indicators to comprehensively identify outliers and missing values in the dataset. Outlier identification is a crucial step in data quality control, primarily employing the following methods: statistical analysis, such as calculating the mean and standard deviation of each parameter to identify data points exceeding the μ±3σ range, or using box plots to identify outliers exceeding 1.5 times the interquartile range; cluster analysis, clustering data points in the feature space to identify points far from their cluster centers; density estimation, calculating the density of the region where each data point resides to identify points located in low-density areas; and neighborhood rule analysis, setting reasonable value ranges based on water treatment processes and the physicochemical characteristics of water quality parameters to identify outliers exceeding these ranges, such as negative COD or ion concentrations exceeding solubility. Missing value identification includes: explicit missing values, i.e., items explicitly marked as null or with special symbols in the data records; implicit missing values, such as duplicate values, zero values, or certain specific values that may actually represent missing data; and temporal discontinuity, identifying missing time points by checking the integrity of the timestamp sequence. During the quality review process, it is also necessary to analyze the distribution patterns of outliers and missing values, such as whether they exhibit systematicity (e.g., continuous anomalies caused by specific sensor malfunctions) or randomness. This helps in selecting appropriate processing strategies. Through quality review, a complete list of data quality issues is established, providing guidance for subsequent targeted data processing.
[0149] For identified outliers, a moving median filter is used for smoothing, preserving the overall trend of the data while reducing the impact of abnormal fluctuations. Moving median filtering is a non-linear filtering technique particularly effective at removing spike noise (such as momentary sensor malfunctions) from time-series data, while also preserving data edges and rapidly changing characteristics. The specific implementation steps include: first, determining the filter window size w (usually an odd number, such as 3, 5, or 7). The choice of window size needs to balance filtering effect and signal fidelity; a larger window can filter noise more effectively but may over-smooth the signal. Then, for each time point t, data from w time points before and after it are taken; these w data points are sorted, and the median is selected as the filtering result for the current time point; this sliding window process is repeated for the entire time series. Moving median filtering is particularly sensitive to outliers, effectively removing impulse noise without introducing new frequency components, making it an ideal method for handling sudden anomalies in sensor data. However, for consecutive outliers (more than half of the outliers within the window), the effect of moving median filtering is limited, in which case it may be necessary to increase the window size or combine it with other methods.
[0150] For identified missing values, interpolation methods are used to impute them, restoring the continuity and integrity of the data. Interpolation impute is a common method for handling missing values in time-series data. Different interpolation strategies can be selected based on the data characteristics: linear interpolation, suitable for short-term missing values and relatively stable data, imputes missing values by connecting known points on both sides of the missing value to form a straight line; polynomial interpolation, such as cubic spline interpolation, can generate a smoother impute curve, suitable for situations where data smoothness needs to be maintained; time-series model interpolation, such as ARIMA (Autoregressive Integral Moving Average) models, can capture the time-series characteristics of the data for predictive impute, suitable for data with obvious time dependencies; k-nearest neighbor interpolation, based on similar samples in the feature space, is suitable for multivariate data where there are correlations between variables; machine learning interpolation, such as random forests or neural networks, can learn complex relationships between variables for predictive impute, suitable for large-scale datasets with complex relationships between variables. The choice of interpolation method should consider the characteristics of the data, the missing value pattern, and computational resource constraints. For long-term or systematic missing values, special processing may be required, combining domain knowledge or additional monitoring data, or using algorithms that can handle missing values during model training.
[0151] For complex situations that are difficult to handle using conventional methods, such as large missing segments or systematic anomalies, statistical methods are used to calculate estimates of missing data points to restore the statistical properties of the data as much as possible. These statistical methods include: multiple imputation, which generates multiple possible imputation values and their probability distributions to better characterize the uncertainty of missing data; the Expectation-Maximization (EM) algorithm, which iteratively optimizes the imputation values while preserving the data distribution characteristics; imputation based on similar time patterns, which uses similar daily or weekly patterns in historical data to estimate missing values, particularly suitable for parameters with obvious periodicity; collaborative filtering, which uses the correlation between multiple parameters to predict missing parameters based on existing parameters, similar to techniques in recommendation systems; and Bayesian inference, which combines prior knowledge and observed data to estimate the posterior distribution of missing values. These methods are generally computationally complex, but they better preserve the statistical properties and uncertainty information of the data, making them suitable for situations with high data quality requirements. When applying these methods, the reasonableness of the imputation results should be evaluated to avoid introducing unrealistic data patterns or biases.
[0152] Through the outlier and missing value processing described above, a quality-controlled training dataset is obtained. This dataset is continuous, consistent, and representative, effectively supporting subsequent model training and predictive analysis. It is worth noting that the degree of data processing should balance "data cleanliness" and "information preservation." Over-processing may remove valuable variation information contained in the data, while under-processing may cause the model to learn noise rather than true patterns. Therefore, outlier and missing value processing should be combined with domain knowledge, data characteristics, and modeling objectives, adopting appropriate processing strategies, and the effectiveness of the processing should be verified in subsequent model evaluation. High-quality data processing improves the reliability and representativeness of the model training data, laying the foundation for building accurate and stable predictive models.
[0153] Example 7
[0154] like Figure 2 As shown, in this embodiment, the process of training and optimizing a prediction model using the model training data to obtain an optimized prediction model, receiving real-time water quality fingerprint time-series data to construct real-time prediction data, and using the optimized prediction model to predict and output predicted values of biochemical treatment efficiency indicators from the real-time prediction data includes:
[0155] Step S31: Use the model training data to train a long short-term memory network model, a bidirectional long short-term memory network model, or a Transformer model. Use learning rate scheduling, gradient pruning, and early stopping strategies to train the model, and adjust the hyperparameters through cross-validation to obtain the model parameters that have been trained and converged.
[0156] Step S32: Load the converged training model parameters into the Long Short-Term Memory Network model, the Bidirectional Long Short-Term Memory Network model, or the Transformer model; calculate the prediction accuracy, mean absolute error, and root mean square error on the validation set; evaluate the model performance under different prediction lead amounts; and perform model integration or structural adjustment based on the prediction accuracy, mean absolute error, and root mean square error to obtain an optimized prediction model.
[0157] Step S33: Receive real-time water quality fingerprint time series data, and apply the feature engineering method consistent with the sliding window design, the time series differential feature extraction, the periodic feature encoding and the time encoding to construct real-time prediction data that meets the model input requirements;
[0158] Step S34: Use the optimized prediction model to perform multi-time-scale prediction on the real-time prediction data, output the predicted values of biochemical treatment efficiency indicators at different future time points, and calculate the prediction confidence interval based on the optimized prediction model.
[0159] The predicted value of biochemical treatment efficiency indicators refers to the point estimate of efficiency indicators (ammonia nitrogen concentration, total nitrogen concentration, chemical oxygen demand, nitrification rate, etc.) output by the optimized prediction model for a specific future time (such as t+2h, t+4h, t+8h, etc.), denoted as ,in For the current moment, To predict lead time. Specifically, the model's receiving length is... After the real-time water quality fingerprint is input, it performs forward calculations based on the time-series mapping relationship learned internally, and the output result is the predicted value of the biochemical treatment efficiency index.
[0160] The prediction confidence interval quantifies the uncertainty of the above-mentioned point prediction, representing the range of values within which the actual performance index value falls with a certain probability (e.g., 95%). This invention employs the Monte Carlo Dropout method to calculate the prediction confidence interval: maintaining the activation of the Dropout layer during the model inference phase, and performing calculations on the same input... Sub-independent forward computation ( ), to obtain the sampled prediction set ,in For the number of samples, For the first Second sample output value. Predicted mean. With the predicted standard deviation They are respectively:
[0161]
[0162]
[0163] The 95% confidence interval for prediction is The system only triggers an alert when the out-of-range value of the confidence interval still exceeds the corresponding warning threshold, in order to ensure the reliability of the alert.
[0164] Specifically, a deep learning time-series prediction model was trained using model training data that had undergone feature engineering. A suitable network architecture for capturing the time-series features of water quality fingerprints was selected, and advanced training optimization strategies were employed. The model selection primarily considered the following network architectures: Long Short-Term Memory (LSTM) network models, a special type of recurrent neural network with input, forget, and output gates, which effectively learn long-term dependencies in long-sequence data and are suitable for capturing long-term trends in water quality parameters; Bidirectional Long Short-Term Memory (BiLSTM) network models, which introduce two independent forward and backward LSTM layers on top of LSTM, simultaneously considering past and future contextual information, improving the accuracy of predicting intermediate states in the sequence; and Transformer models, an architecture based on self-attention mechanisms, which can process the entire sequence in parallel, capturing dependencies between any positions in the sequence, and are particularly effective for modeling long-distance dependencies. Several optimization strategies were employed during model training: learning rate scheduling, such as cosine annealing, which uses a large learning rate for rapid convergence in the early stages of training and a small learning rate for fine-tuning in the later stages, or gradually decays the learning rate after warm-up to avoid instability in the early stages of training; gradient clipping, which sets a gradient norm threshold and scales the gradient when it exceeds the threshold to prevent gradient explosion; and early stopping, which monitors the performance on the validation set and stops training when the metrics fail to improve for several consecutive rounds to prevent overfitting. Furthermore, K-fold cross-validation was used to adjust model hyperparameters, such as hidden layer size, number of layers, and dropout rate, to find the optimal model configuration. Finally, through this series of training and optimization steps, a time-series prediction model with convergent parameters was obtained, which can effectively capture the mapping relationship between water quality fingerprint sequences and future performance metrics.
[0165] The converged model parameters are loaded into the selected network architecture, and a comprehensive performance evaluation is performed on independent validation datasets. The final optimized prediction model is selected or constructed based on a comprehensive comparison of multiple metrics. Performance evaluation employs various metrics: prediction accuracy, such as accuracy rate (the proportion of predictions within a predetermined error range) and R² coefficient of determination (the degree to which predicted values explain the variation in actual values); mean absolute error (MAE), which calculates the average absolute difference between predicted and actual values, reflecting the average magnitude of the prediction error; and root mean square error (RMSE), which calculates the square root of the sum of the squares of the differences between predicted and actual values, and is more sensitive to larger errors. These metrics reflect the model's predictive performance from different perspectives and must be considered comprehensively. Particularly important is evaluating the model's performance at different prediction lead times, such as evaluating the performance of predictions 2 hours, 4 hours, and 6 hours in advance, to determine the model's effective prediction range. Based on the evaluation results, further model optimization measures are taken: model ensemble, such as combining the prediction results of multiple base models through averaging, weighted averaging, or stacking, to improve the stability and accuracy of predictions; structural adjustment, such as adjusting the network depth and width according to different prediction tasks, or designing dedicated prediction heads for specific performance indicators; and feature importance analysis to identify and enhance the most critical feature inputs for the prediction task. Through these evaluation and optimization steps, an optimized prediction model with good prediction performance under various conditions is finally obtained, preparing it for real-time prediction applications.
[0166] The system receives real-time water quality fingerprint time-series data from a water quality monitoring system and applies the same feature engineering methods as in the training phase to construct real-time prediction data that meets the model's input requirements. This process requires ensuring the consistency of the processing pipeline, meaning that the same preprocessing steps and feature extraction methods are applied to the real-time data as to the training data. First, the same data preprocessing methods are applied, including outlier handling and missing value imputation. Then, a sliding time window consistent with the training phase is constructed, typically holding water quality fingerprint data from the most recent n hours, with the window size the same as that used in the training model. Next, temporal difference feature extraction is performed, calculating the difference in feature values between adjacent time points in the real-time data to capture parameter change trends. Then, periodic feature encoding is added, generating periodic codes based on the current time information (hour, day, week, etc.). Finally, time encoding is added, converting the current moment into a time representation that the model can understand. These processing steps need to be executed in real-time to ensure that newly acquired data can be promptly converted into the model input format. To ensure data processing consistency, feature engineering parameters used in the training phase (such as standardization coefficients and encoding methods) are typically saved and directly applied to real-time data processing. In this way, it is ensured that the predicted data constructed in real time has the same feature distribution and structure as the training data, avoiding the degradation of model performance due to inconsistent data processing.
[0167] An optimized prediction model is used to perform forward predictions on real-time data across multiple time scales, outputting predicted values and confidence intervals for biochemical treatment efficiency indicators at different future time points. Multi-time scale prediction refers to simultaneously predicting different time ranges, typically including short-term (e.g., 2 hours later), medium-term (e.g., 4-6 hours later), and long-term (e.g., 8-12 hours later) predictions. During the prediction process, prepared real-time prediction data is input into the optimized prediction model. The model generates predicted values for the corresponding time points based on the learned mapping relationship between water quality fingerprint sequences and future efficiency indicators. For different prediction time ranges, the following strategies can be adopted: direct multi-output prediction, where the model outputs predicted values for multiple time points simultaneously; iterative prediction, first predicting recent values and then incorporating the prediction results into the input sequence to predict further future time points; and dedicated model prediction, training specialized models for different prediction ranges. In addition to point predictions, it is also necessary to calculate prediction confidence intervals to characterize the uncertainty of the prediction results. Common methods include: Monte Carlo Dropout, which maintains dropout activation during the inference phase and obtains the distribution through multiple predictions; ensemble variability, which estimates uncertainty based on the dispersion of predictions from each base model in the ensemble model; Bayesian neural networks, which directly model the model parameters as probability distributions; and quantized regression, which predicts quantiles rather than point estimates. Prediction confidence intervals are crucial for subsequent early warning decisions, and high-confidence anomaly predictions usually require a higher level of response. Prediction results should include multiple key performance indicators, such as concentrations of ammonia nitrogen, total nitrogen, and COD, as well as activated sludge state parameters, to comprehensively characterize the future operating status of the system.
[0168] The presentation and analysis of forecast results should consider user needs and practical application scenarios, transforming complex forecast results into understandable and actionable information. Typical result presentations include: time series visualization, displaying historical data, current values, and future forecast trends to intuitively reflect parameter changes; multi-parameter dashboards, simultaneously displaying current and forecast values of multiple key indicators for easy overall status assessment; warning status indicators, using color coding or icons to identify different levels of warning status based on the comparison between forecast values and thresholds; and confidence interval bands, displaying the range of confidence intervals around the forecast line to reflect the uncertainty of the forecast. For key decision-makers, in-depth analysis functions should also be provided, such as parameter sensitivity analysis to identify the input features that have the greatest impact on the forecast results; anomaly cause analysis, inferring factors that may lead to forecast anomalies based on model interpretability techniques; and multi-scenario simulation to evaluate changes in forecast effects under different intervention measures. These functions enable the forecasting system not only to provide numerical estimates of future performance but also to support operators in understanding the reasons behind the forecasts and making more informed regulatory decisions. By providing comprehensive forecast result presentation and analysis functions, the practical value of the forecasting model is maximized, transforming advanced algorithms into practically usable decision support tools.
[0169] In this invention, the prediction confidence interval is an interval estimate that quantifies and represents the uncertainty of the prediction result, giving the range and probability of the predicted value falling into it. The prediction confidence interval is used to assess the reliability of the predicted value of biochemical treatment efficacy, supporting risk-based decision-making. The principle of the confidence interval is based on probability statistics, extending point estimation to interval estimation, reflecting the sources of uncertainty in the model prediction (such as data noise, model approximation error, etc.). The steps for constructing the prediction confidence interval include: Monte Carlo Dropout sampling, maintaining the activation of the Dropout layer of the neural network during the inference phase, and generating a prediction distribution through multiple samplings; ensemble methods, using multiple different models or different parameter versions of the same model for prediction, and determining the confidence interval based on the distribution of the predicted values; quantile regression, training the model to directly predict different quantiles of the distribution (such as 10%, 50%, 90%), rather than a single expected value; and Bayesian methods, performing Bayesian inference on the model parameters or the prediction distribution to obtain the confidence interval of the posterior distribution. Accurate prediction confidence intervals are crucial for risk assessment and early warning decisions, helping operators determine the credibility of prediction anomalies and take appropriate response measures.
[0170] Example 8
[0171] In this embodiment, the training of a Long Short-Term Memory (LSTM) network model, a Bidirectional LSM network model, or a Transformer model using the model training data employs learning rate scheduling, gradient pruning, and early stopping strategies for model training. Hyperparameters are adjusted through cross-validation to obtain converged model parameters, including:
[0172] The computational graph structure of the Long Short-Term Memory Network model, the Bidirectional Long Short-Term Memory Network model, or the Transformer model is represented as a directed tree, where nodes represent computational units in the model, and edges represent data flow and dependencies.
[0173] Based on the forward propagation and gradient backpropagation characteristics of the model, the gradient norm, feature expression entropy and prediction error contribution rate of each node in the directed tree are calculated. Nodes with gradient norm lower than a preset gradient threshold, feature expression entropy lower than a preset entropy threshold, or prediction error contribution rate higher than a preset error threshold are identified as information bottleneck positions. An enhanced edge candidate set is generated for each of the information bottleneck positions, and each enhanced edge connects two nodes in the directed tree that are not originally directly connected.
[0174] Define an augmentation edge utility function to quantify the improvement in model performance after adding each candidate edge in the augmentation edge candidate set. Construct a linear programming relaxation problem, use a rounding strategy better than 2 approximations to convert the fractional solution of the linear programming into an integer solution, and iteratively optimize to select the optimal subset of augmentation edges to add to the directed tree.
[0175] Based on the optimal subset of enhanced edges, the model architecture is modified by adding residual connections, cross-layer connections, or attention mechanisms to form an enhanced connection structure. A parameter initialization strategy adapted to the enhanced connection structure is designed, and a phased training strategy is adopted. First, the original structural parameters of the model are frozen and only the connection weight parameters of the enhanced connection structure are trained. Then, the original structural parameters and the connection weight parameters are jointly optimized. Regularization techniques are applied to control the numerical range of the connection weight parameters to obtain the model parameters that have been trained and converged.
[0176] Specifically, the computational process of deep learning models is abstracted into a directed tree structure for structured analysis and optimization. Deep neural network models can be represented as data flow graphs, containing various computational units (such as LSTM units, self-attention layers, fully connected layers, etc.) and the data transfer relationships between them. This data flow graph can be further simplified into a directed tree structure. Although, strictly speaking, some models (such as networks with residual connections) form directed acyclic graphs rather than trees, they can still be transformed into equivalent tree structures for analysis through appropriate node duplication or the introduction of virtual nodes. In this directed tree, each node represents a computational unit, such as the memory unit, input gate, forget gate, and output gate in an LSTM model, or the multi-head self-attention layer and feedforward network layer in a Transformer model; while edges represent data flow and dependencies, indicating how information is transferred from one computational unit to another, including information flow in forward propagation and gradient flow in backpropagation. The root node of the tree typically corresponds to the final output layer of the model, and the leaf nodes correspond to the input features. This tree-like representation makes the computational structure and information flow of the model intuitive and visible, providing a foundation for subsequent bottleneck analysis and structural optimization. The process of constructing a directed tree requires analyzing the model's computational graph definition, identifying the input-output relationships of each computational unit, and mapping them to nodes and edges in the tree, while ensuring the integrity and consistency of the information flow.
[0177] Based on the constructed directed tree structure, key bottleneck locations are identified through in-depth analysis of the model's information flow characteristics, providing precise targets for structural optimization. Specifically, the model is trained and inferred, and the following three key indicators are recorded and analyzed: gradient norm, which calculates the L2 norm of the gradient vector received by each node during backpropagation; a low gradient norm (below a preset gradient threshold) indicates that the node receives weak signals during parameter updates, potentially indicating a vanishing gradient problem; and feature representation entropy, which measures the diversity of the node's output activation value distribution, i.e., the richness of information expressed by the node when processing different input samples. "Information richness" specifically means: the more uniform the activation value distribution, the stronger the node's ability to respond to various input patterns and the stronger its information expression capability; the more concentrated the activation value distribution (a large number of values approaching zero or a fixed value), the weaker the node's ability to distinguish inputs, indicating a risk of information loss. Let the node... For batch The activation values of each sample are output as follows Normalize it into a probability distribution Then the node Feature representation entropy for:
[0178]
[0179] in, For nodes For the The activation output value of each sample. Its normalized probability, This represents the sample batch size. The theoretical maximum value is (Obtained when uniformly distributed). When Below the preset entropy threshold If the node outputs insufficient information, there is a risk of information loss or node inactivation. It is then marked as an information bottleneck and its structure is reinforced.
[0180] The prediction error contribution rate is quantitatively assessed through perturbation analysis to evaluate the magnitude of each node's contribution to the final prediction error of the model. Specifically, the "degree of influence" is defined as: during the forward propagation process, the impact on the target node... The sum of the output activation values has a mean of zero and a standard deviation of . Use Gaussian random noise (10% of the historical standard deviation of the node's activation values) to observe the relative change in the root mean square error of the model on the validation set before and after applying the perturbation. Prediction error contribution rate Calculate using the following formula:
[0181]
[0182] in, This represents the baseline root mean square error of the model on the validation set without any perturbation. For nodes The root mean square error after applying perturbation noise. The higher the value, the greater the contribution of that node to the final prediction error. When Higher than the preset error threshold When this node is identified as an information bottleneck, structural enhancement is needed to improve its information transmission efficiency in order to improve the overall predictive performance of the model.
[0183] Nodes meeting any of the above conditions (low gradient norm, low entropy, or high error contribution rate) are marked as information bottlenecks. These locations either have poor information transmission (e.g., intermediate layers in deep networks) or low information processing efficiency (e.g., simple structures in complex tasks), requiring structural enhancement to improve them. For each identified bottleneck location, a set of potential enhancement connection candidates is generated based on network connection patterns and information flow analysis. These candidate connections (i.e., enhancement edges) connect node pairs that are not directly connected in the computation graph, aiming to create "shortcuts" for information flow and alleviate information pressure at the bottleneck. The generation of enhancement edge candidates considers the model's hierarchical structure, functional modules, and information flow direction, ensuring that the added connections are functionally meaningful and feasible in implementation.
[0184] The generated candidate set of enhancement edges is systematically evaluated and optimized to determine the optimal structural enhancement scheme. This process first requires defining a utility function for the enhancement edges, quantifying the improvement in model performance after adding each candidate edge. After defining the utility function, a linear programming relaxation problem is constructed, formalizing the enhancement edge selection problem into an optimization problem. The goal of this linear programming problem is to maximize the total utility value of the selected enhancement edges under the constraint of controlling the total number of enhancement edges (or the total number of newly added parameters). Since linear programming may produce fractional solutions (i.e., some edges are partially selected), a rounding strategy is needed to convert fractional solutions into integer solutions (i.e., deterministically selecting or not selecting each edge). This embodiment uses a rounding strategy better than 2 approximations, a special random rounding method that provides better approximation guarantees than simple rounding while maintaining the quality of the expected solution. During the rounding process, random decisions are made based on the selection probability of each edge obtained from the linear programming, but the probability distribution is adjusted to ensure that the total utility of the final set of selected edges is not less than half of the optimal linear programming solution. Through an iterative optimization process, the constraints and objective function of the linear programming problem are continuously adjusted. An alternating minimization strategy is employed (first fixing edge weights and then fixing edge weights and optimizing edge selection), ultimately determining the optimal subset of augmented edges. This process requires repeated evaluation of the effects of different combinations of augmented edges on the validation dataset to ensure that the selected edge set truly improves model performance and does not merely overfit the training data.
[0185] Based on the determined optimal subset of enhanced edges, the original model architecture is modified and enhanced, and a specialized training strategy is designed. Specifically, for each node pair connected by an enhanced edge, corresponding structural enhancement elements are added to the model: residual connections (ResNet style), providing shortcuts for information flow to bypass certain layers, helping to alleviate the gradient vanishing problem in deep networks; cross-layer connections (DenseNet style), directly passing features from previous layers to subsequent layers, enhancing feature reuse and gradient flow; and attention mechanisms, using the correlation between connected nodes to guide information fusion, improving the selectivity and effectiveness of information transmission. For these newly added structural elements, a specialized parameter initialization strategy is designed: small initial values close to zero are used for residual connections to ensure that the original information flow is not excessively interfered with in the early stages of training; attention weights are initialized with a uniform distribution, giving each information source equal initial importance; and the projection matrix of cross-layer connections is orthogonally initialized to maintain the structural properties of the feature space. Then, a phased training strategy is adopted, first freezing the original structural parameters of the model (such as the weight matrix of LSTM units, multi-head attention parameters in Transformer, etc.), and training only the newly added enhanced connection structural parameters. This strategy allows new connections to find the optimal combination of information without perturbing the previously learned features. After the first phase of training, the original parameters are unfrozen, and joint optimization of all parameters is performed, enabling the augmented connections to work collaboratively with the original structure. During this process, regularization techniques are applied to control the numerical range of the augmented connection weights, preventing the new connections from excessively dominating the information flow and ensuring the model's generalization ability. Through this refined structural augmentation and training strategy, a high-performance model with optimized structure and convergent parameters is ultimately obtained.
[0186] In this invention, information bottleneck locations refer to nodes in the computational graph of a deep learning model where information transmission is obstructed or inefficient. Information bottleneck locations are identified using three key indicators: excessively low gradient norm (vanishing gradient), excessively low feature representation entropy (information loss), and excessively high prediction error contribution rate (critical nodes). The formation principle of information bottleneck locations is that in deep neural networks, due to problems such as activation function saturation, gradient vanishing / exploding, and limited representational capacity, certain nodes become "bottlenecks" in information flow, limiting the overall performance of the model. The steps for identifying information bottleneck locations include: calculating the gradient norm of each node during training to quantify the gradient signal strength; calculating the activation value entropy of the node output to assess the information content; quantifying the influence of nodes on the final prediction through perturbation analysis; and identifying nodes that meet the bottleneck conditions based on preset thresholds. The identification of information bottleneck locations provides precise targets for targeted model structure enhancement.
[0187] Linear programming relaxation is an optimization modeling method that relaxes discrete decision problems (such as whether to add a reinforcing edge) into continuous variable problems, making them easier to solve using linear programming algorithms. In this invention, reinforcing edge selection is modeled as a 0-1 integer programming problem, and then relaxed to become a linear programming problem. The principle of relaxation is to relax the integer constraints on variables, allowing them to take any value in the interval [0,1], making the problem easier to solve, and the optimal solution of the relaxation problem provides an upper bound for the original integer programming problem. The steps for constructing and solving the relaxation problem include: defining the selection variable xi∈[0,1] for each candidate edge; setting the objective function to maximize total utility; adding resource constraints (such as an upper limit on the total number of edges or the total number of new parameters); using a linear programming algorithm (such as the simplex method or interior point method) to solve the relaxation problem; and applying a rounding strategy to convert the continuous solution back to an integer solution. This method can effectively handle large-scale edge selection optimization problems and find solutions close to the global optimum.
[0188] The better-than-2 rounding strategy is a special random rounding algorithm used to convert fractional solutions of linear programming relaxation problems into integer solutions while ensuring the quality of the solutions. In this invention, this strategy is used to convert continuous edge selection probabilities into deterministic 0-1 selection decisions. The principle of this rounding strategy is based on probability theory and approximation algorithm analysis. Through a carefully designed randomization process, it ensures that the expected target value of the final integer solution is not less than half of the optimal solution of the linear programming problem (i.e., a 2 approximation), while in practice, a better approximation ratio (better than 2) is usually obtained. The implementation steps include: obtaining the optimal solution of the linear programming problem, where each variable xi represents the probability of selecting the i-th edge; constructing a special probability distribution based on these probability values, possibly adjusting the original probabilities to improve the quality of the final solution; randomly sampling based on the adjusted probabilities to determine whether each edge is selected (1) or not selected (0); and, if necessary, performing post-processing steps to ensure that all constraints are satisfied. This rounding strategy has performance guarantees in theory and can usually produce high-quality integer solutions close to the optimal solution in practice.
[0189] Example 9
[0190] In this embodiment, the defined enhancement edge utility function quantifies the improvement in model performance after adding each candidate edge from the enhancement edge candidate set, including:
[0191] Based on the two nodes connected by each candidate edge in the enhanced edge candidate set, each candidate edge is temporarily added to the directed tree to form a temporarily modified directed tree. According to the temporarily modified directed tree, a corresponding temporary connection structure is added to the model. The model training data is used to perform forward propagation calculation on the model after adding the temporary connection structure. The difference in prediction error before and after adding each candidate edge is calculated on the validation set to obtain the change in prediction error.
[0192] The number of new parameters and the number of new floating-point operations in the model caused by adding temporary connection structures corresponding to each candidate edge are counted to obtain the increase in model parameters and the increase in computational cost.
[0193] Based on the change in prediction error, the increase in model parameters, and the increase in computational cost, the ratio of the change in prediction error to the weighted sum of the increase in model parameters and the increase in computational cost is calculated. This ratio is defined as the utility function of the enhanced edge, and the utility score of each candidate edge in the candidate enhanced edge set is obtained.
[0194] Specifically, each candidate edge in the enhancement edge candidate set is evaluated independently, and its actual impact on model performance is quantified experimentally. For each candidate enhancement edge, the two nodes it connects are first identified. These two nodes are not directly connected in the original directed tree structure of the model, but adding a connection between them may create a shortcut or supplementary path for information flow, alleviating information bottlenecks. The evaluation process uses a temporary addition method, that is, without changing other parts of the original model, only this candidate edge is added to the directed tree structure, forming a temporarily modified directed tree structure. Then, based on this temporarily modified tree structure, the corresponding connection structure is added to the actual neural network model. The specific implementation depends on the type of nodes connected by the candidate edge: if it connects two hidden layer nodes, a direct feature transfer path may be added; if it connects layers with large spans, a dimension-matching mapping layer may be needed; if it connects different types of functional units, a specific adaptation mechanism such as an attention layer may be needed. After adding temporary connections, forward propagation is performed on the modified model using a subset (usually a mini-batch) of the model's training dataset. The model's performance is then evaluated on an independent validation set, and key metrics such as prediction error (mean squared error, mean absolute error, etc.) are calculated. By comparing the difference in prediction error on the validation set before and after adding temporary connections, the performance change brought about by each candidate edge is determined. This change directly reflects the impact of the enhanced edges on the model's predictive ability and is a core component of the utility function.
[0195] Beyond performance improvements, the increased computational cost of augmented edges must also be considered, requiring a comprehensive cost-benefit analysis. Adding new connection structures inevitably increases model complexity, primarily in two ways: the number of model parameters increases and the computational cost (number of floating-point operations). For each candidate augmented edge, these two cost metrics due to the addition of the corresponding temporary connection structure need to be statistically analyzed. The increase in the number of parameters depends on the specific implementation of the connection: a simple identity mapping may not increase the number of parameters; a linear projection layer increases the number of parameters by multiplying the input dimension by the output dimension; and attention mechanisms may add more complex parameter structures. The increase in computational cost refers to the additional floating-point operations in forward and backward propagation, including matrix multiplication, activation function calculation, gradient calculation, etc. This statistic needs to consider the actual batch size and sequence length, as the computational complexity of some operations varies with these factors. For example, if the augmented edge connects two LSTM layers with a hidden state size of h and is implemented through a linear projection, the number of new parameters is approximately h², while the additional computational cost is also related to the batch size and sequence length. By accurately calculating these cost metrics, the resource requirements of each enhanced edge can be comprehensively assessed, providing a cost basis for subsequent utility function calculations.
[0196] Based on collected data on performance improvements and resource costs, a utility function balancing these two factors is constructed to provide a quantitative basis for edge enhancement selection. The core design of the utility function is to evaluate the performance improvement per unit cost, i.e., "cost-effectiveness." Specifically, the utility of an enhanced edge is defined as the ratio of the change in prediction error to the weighted sum of resource costs. The change in prediction error is the reduction in prediction error on the validation set before and after adding the edge; a larger change indicates a greater contribution from the edge. Resource costs are a weighted combination of the increase in model parameters and the increase in computational cost, with weighting coefficients reflecting the relative importance of these two resources in a specific application scenario. For example, on resource-constrained edge devices, the number of parameters may be more important; while in applications requiring real-time response, computational cost may be more critical. The utility function formula can be expressed as: Utility = Reduction in Prediction Error / (α × Increase in Parameters + β × Increase in Computational Cost), where α and β are weighting coefficients set according to application requirements. Through this utility function, the value of each candidate enhanced edge is quantified into a utility score; a high score indicates that the edge brings significant performance improvement at a relatively low cost. These scores provide clear quantitative indicators for subsequent edge selection optimization, allowing the edge selection problem to be transformed into a utility-maximization-based optimization problem. It is important to note that the utility scores are calculated by independently evaluating each edge while keeping other structural elements constant. However, in practice, multiple edges may interact (co-enhance or cancel each other out) when used in combination, a factor that needs to be considered in subsequent edge selection optimization.
[0197] To more comprehensively evaluate the value of augmenting edges, multiple dimensions are introduced beyond the basic utility function. First, edge complementarity is considered, i.e., the degree of complementarity between a given edge and the existing set of edges. This is measured by conditional utility increment, calculating the additional performance improvement gained by adding the current edge to an existing set of edges. Highly complementary edge combinations can produce a "1+1>2" effect, which is particularly important in complex model optimization. Second, robustness is considered, testing the stability of augmented edges under different data distributions (e.g., different batches, different noise levels) to prioritize edges that provide stable improvements under various conditions. Third, generalization is considered, evaluating the difference in performance improvement between the training and validation sets to avoid connections that are only effective on specific data but have poor generalization ability. Finally, long-term utility evaluation is introduced, i.e., the impact of adding edges on the long-term training process of the model, prioritizing edges that not only improve short-term performance but also enhance long-term convergence characteristics. These multi-dimensional evaluations are integrated into the final utility score through weighted combination, forming a more comprehensive and reliable quantitative indicator of edge value. This multi-faceted utility evaluation ensures that the selected reinforcing edges are not only effective under specific conditions, but also have good generalization, stability, and long-term value, providing more reliable guidance for model structure optimization.
[0198] The enhanced edge utility function is a quantitative evaluation metric used to measure the performance improvement of adding new connections to the computational graph of a model. In this invention, the enhanced edge utility function is defined as the ratio of the change in prediction error to the resource cost (a weighted combination of the increase in parameters and the increase in computational cost). The design principle of the utility function is based on cost-benefit analysis, aiming for the maximum performance improvement per unit of resource input. The steps for calculating the enhanced edge utility include: temporarily adding candidate edges to the model for forward computation; measuring the change in prediction error before and after the addition on the validation set; statistically analyzing the increase in parameters and computational cost caused by adding the edge; and using the ratio of the improvement in computational error to the resource cost as the utility score. Edges with high utility scores can bring significant performance improvements with relatively small resource overhead and should be given priority for addition to the model.
[0199] Example 10
[0200] In this embodiment, the phased training strategy involves first freezing the original structural parameters of the model and training only the connection weight parameters of the enhanced connection structure, and then jointly optimizing the original structural parameters and the connection weight parameters, including:
[0201] Based on the optimal subset of enhanced edges, the model architecture is modified by adding residual connections, cross-layer connections, or attention mechanisms to the node positions corresponding to the optimal subset of enhanced edges to form an enhanced connection structure, thus obtaining the model after adding the enhanced connection structure.
[0202] Based on the model with the added enhanced connection structure, the original structural parameters of the Long Short-Term Memory Network model, the Bidirectional Long Short-Term Memory Network model, or the Transformer model and the connection weight parameters of the enhanced connection structure are labeled. The original structural parameters are set to an untrainable state, and the connection weight parameters are set to a trainable state. The connection weight parameters are optimized by gradient descent using the model training data. The convergence of the loss function is monitored on the validation set to obtain the connection weight parameters after the first stage of optimization.
[0203] Remove the untrainable state of the original structural parameters, set the original structural parameters and the connection weight parameters optimized in the first stage to trainable state simultaneously, and use the model training data to perform joint gradient descent optimization on the original structural parameters and the connection weight parameters optimized in the first stage to obtain the original structural parameters and connection weight parameters optimized in the second stage.
[0204] Based on the connection weight parameters optimized in the second stage, the L1 norm and L2 norm of the connection weight parameters are calculated, and a regularization loss term is constructed as a weighted combination of the L1 norm and the L2 norm. The regularization loss term is added to the total loss function of the model after adding the enhanced connection structure. The sparsity and numerical value of the connection weight parameters are controlled by adjusting the regularization coefficient, and the model parameters of the training convergence are obtained.
[0205] Specifically, based on the optimal subset of augmented edges obtained from previous analysis and optimization, the architecture of the original deep learning model is systematically modified, and augmented connection structures are added at specified locations. For each selected augmented edge, the most suitable connection implementation method is chosen according to the type and function of the node pairs it connects: residual connection is the most direct form of augmentation, suitable for connecting nodes with similar layers and the same feature dimensions. It achieves direct addition of features through identity mapping or simple linear transformation, in the form y = F(x) + x, where F(x) is the original computation path; cross-layer connection is suitable for information transmission across multiple layers, usually requiring dimension matching processing, which can be achieved through projection matrix or adaptive pooling, in the form y = F(x) + W·G(z), where z is the output of the distant layer, G is the feature processing function, and W is the projection matrix; attention mechanism is suitable for scenarios that require selective fusion of information, especially when two nodes represent information of different types or different levels of abstraction. It guides information fusion by calculating relevance weights, in the form y = F(x) + Attention(x,z)·z, where the Attention function calculates the relevance weights between x and z. In practical implementation, the specific architectural characteristics of the model also need to be considered: for LSTM models, it may be necessary to distinguish between connections between different gating units and connections between hidden states; for Transformer models, it may be necessary to consider enhanced connections within or across layers of the multi-head attention mechanism. After adding these connection structures, the forward propagation logic of the model needs to be adjusted to ensure that the new data flow path is correctly implemented. Through this step, the original model is reconstructed into a new model with enhanced connection structures, ready for parameter optimization training.
[0206] On the model with the added enhanced connections, the first phase of targeted training is implemented, focusing on optimizing the weight parameters of the newly added connection structures. First, it's crucial to clearly distinguish between two types of parameters in the model: original structural parameters, including the input weights, recursive weights, and biases of the LSTM, or the attention projection matrix and feedforward network weights in the Transformer; and enhanced connection weight parameters, i.e., the trainable parameters in the newly added connection structures, such as the projection matrix and attention weights. In the first phase of training, the original structural parameters are set to a non-trainable state (usually achieved through `freeze` or `requires` in the framework). grad=False) means that these parameters will not be updated during backpropagation. Simultaneously, the weight parameters of the enhanced connection structure are set to a trainable state, allowing them to be optimized via gradient descent. The core idea of this strategy is to allow the newly added connection structure to learn how to optimally combine and propagate these features without altering the existing, initially trained feature extraction capabilities. The training process uses the same loss function and optimizer as regular training, but only updates the parameters marked as trainable. During training, the convergence of the loss function needs to be monitored on the validation set, including the decreasing trend and stability of the validation loss. When the validation loss no longer decreases significantly over multiple consecutive training epochs, or reaches the preset upper limit of training epochs, the first stage of training is considered to have converged, yielding the optimized connection weight parameters. This stage of training typically converges quickly because the number of parameters to be optimized is relatively small, and the existing feature extraction structure provides a good foundation.
[0207] After completing the first stage of training, the second stage, joint optimization of all parameters, is initiated, enabling the enhanced connections to co-evolve with the original structure. At this point, the untrainable state of the original structure parameters is removed, and all parameters in the model (including the original structure parameters and the connection weights optimized in the first stage) are simultaneously set to a trainable state. This means that during backpropagation, gradients will flow to all parameters, allowing for end-to-end joint optimization of the entire model. The goal of joint optimization is to enable the original feature extraction structure to adapt to the new information flow path, while the enhanced connections can further optimize their combination based on the adjusted features. This stage uses the same loss function, but may require adjustments to the learning rate strategy: typically, a smaller learning rate is used for the original structure parameters to avoid drastically altering the learned useful features; a slightly larger learning rate can be used for the enhanced connection parameters to allow them to continue to adapt actively. Performance is also monitored on the validation set during training. When the jointly optimized model reaches a stable state on the validation set—that is, the validation loss no longer decreases or begins to increase (signs of overfitting)—the second stage of training ends. This stage typically requires a longer training time because it involves the joint optimization of all parameters, and the optimization space is more complex. This two-stage training strategy ensures that the enhanced connections can play a full role while maintaining the effective features of the original model structure, achieving the dual goals of structural enhancement and parameter optimization.
[0208] To further improve the model's generalization ability and prevent overfitting, a special regularization strategy for enhancing connection weights is introduced based on the second-stage optimization. First, the L1 and L2 norms of the connection weight parameters after the second-stage optimization are calculated. The L1 norm is the sum of the absolute values of the weights, promoting parameter sparsity and causing most weights to approach zero, thus achieving implicit feature selection. The L2 norm is the square root of the sum of the squares of the weights, limiting the overall size of the weights and preventing overfitting caused by excessively large individual weight values. Then, a regularization loss term is constructed in the form λ1×L1 norm + λ2×L2 norm, where λ1 and λ2 are hyperparameters controlling the strength of the two regularization methods, and their optimal values need to be determined through cross-validation. This hybrid regularization combines the sparsity promotion of L1 and the smoothing constraint of L2, making it particularly suitable for controlling the behavior of enhancing connections. The constructed regularization loss term is added to the model's total loss function, participating in the optimization along with the original task losses (such as mean squared error, cross-entropy, etc.). By adjusting the regularization coefficients, the characteristics of augmenting connection weights can be precisely controlled: a larger L1 coefficient encourages more weights to approach zero, effectively "closing" some augmenting connections and achieving automatic structural simplification; an appropriate L2 coefficient prevents excessively large remaining non-zero weights, ensuring that augmenting connections do not overly dominate the information flow. The essence of this regularization strategy is to guide the model to find the minimum necessary set of augmenting connections, achieving performance improvement while avoiding unnecessary complexity increases. Through this optimization step, a training convergence model that is both structurally sound and parameter-stable is ultimately obtained, possessing both high performance and good generalization ability.
[0209] After the entire training process is completed, a final model evaluation and analysis are performed to confirm the effectiveness of the augmented connections and prepare for practical applications. First, model performance is evaluated on independent test sets, comparing differences in various metrics before and after augmentation, such as prediction accuracy, mean squared error, and mean absolute error. Particular attention is paid to the model's stability under different conditions (e.g., different pollutant loads, different seasonal operating conditions) to verify whether the augmented connections improve the model's robustness. Second, contribution analysis of the augmented connections is conducted by visualizing the weight distribution and activation patterns of each augmented connection to identify the most critical connection paths and information flows. This helps in understanding the model's working mechanism and provides guidance for future model iterations. Then, the model's computational efficiency is evaluated by measuring changes in inference time, memory usage, and energy consumption before and after augmentation to ensure the augmented model can still operate efficiently in resource-constrained environments. Finally, a model deployment package is generated, including optimized model parameters, inference code, and usage documentation, facilitating integration into practical water quality monitoring systems. Through these comprehensive evaluation and preparation steps, we ensure that the models trained with structural enhancement and optimization not only have excellent theoretical performance but also provide stable and reliable predictive capabilities in practical applications, thus providing strong support for water quality monitoring and treatment efficiency early warning.
[0210] The phased training strategy is a neural network optimization method that divides the model training process into multiple continuous but distinct phases. In this invention, this strategy specifically refers to a two-phase method: first training the parameters of the augmented connections, and then jointly optimizing all parameters. The principle of phased training is based on transfer learning and incremental learning concepts. By controlling the range and order of parameter updates, more refined model optimization can be achieved. The implementation steps include: In the first phase, the original model parameters are frozen, and only the parameters of the newly added augmented connection structure are trained, allowing the new connections to learn how to optimally combine existing features; in the second phase, the parameter freeze is lifted, and all parameters are optimized simultaneously, possibly using different learning rates, allowing the original structure and augmented connections to co-evolve. The advantage of this strategy is that it avoids the interference of unstable early-stage augmented connection training on the original features, while ensuring the overall consistency of the final model, making it an effective method for optimizing complex model structures.
[0211] The regularization loss term is an additional term added to the original loss function of the model to control model complexity and parameter characteristics. In this invention, the regularization loss term specifically refers to the weighted combination of the L1 norm and L2 norm used to control the characteristics of the augmented connection weight parameters. The principle of the regularization loss term is to guide the model to learn a simpler and more generalized parameter distribution by adding a penalty term to the parameters in the optimization objective. The steps to construct the regularization loss term include: calculating the L1 norm (sum of absolute values) and L2 norm (square root of the sum of squares) of the target parameters; determining the weight coefficients of L1 and L2, reflecting the importance attached to parameter sparsity and size; adding the weighted combination of regularization terms to the original loss function; and considering both task loss and regularization loss during the optimization process. By adjusting the regularization coefficients, the characteristics of augmented connections can be precisely controlled: L1 regularization promotes weight sparsity, achieving feature selection and model simplification; L2 regularization limits the weight size, preventing overfitting and numerical instability. A reasonable regularization strategy is a key technology for improving the generalization ability of the model.
[0212] Example 11
[0213] In this embodiment, the process of determining a graded early warning threshold based on historical performance index data, comparing the predicted value of the biochemical treatment performance index with the graded early warning threshold to generate early warning information, querying an expert knowledge base based on the early warning information and applying an experience learning model to generate process control suggestions, and optimizing the process control suggestions to form a control scheme includes:
[0214] Based on the statistical analysis and processing requirements of historical performance index data, combined with expert knowledge, the early warning thresholds for different performance indicators are determined, and the thresholds are dynamically adjusted considering seasonal factors and changes in working conditions to obtain graded early warning thresholds.
[0215] The predicted value of the biochemical treatment efficiency index is compared with the graded early warning threshold. When the predicted value of the biochemical treatment efficiency index exceeds the graded early warning threshold and the prediction confidence interval meets the preset requirements, early warning information including early warning time, early warning type, severity and credibility is generated.
[0216] Integrate the experience of experts and process operation procedures in the field of wastewater treatment to establish an expert knowledge base that includes a three-element correlation rule of problem-cause-measure;
[0217] Collect historical process control records and corresponding effect data, and use machine learning methods to mine the relationship patterns between control parameters and effects to obtain an experience learning model;
[0218] Based on the warning information, combined with the current system operating status parameters, the expert knowledge base is queried and the experience learning model is applied to generate process control suggestions that include the adjustment direction, magnitude and timing of aeration volume, reflux ratio and reagent dosage.
[0219] The proposed process control recommendations are optimized through multi-objective optimization to balance treatment effect, energy consumption, and operating cost, forming an optimized control scheme. The optimized control scheme is then transformed into execution guidelines that include execution time, operation sequence, and precautions.
[0220] Specifically, a dynamically adaptive early warning threshold system is first constructed based on historical performance data and treatment process requirements, combined with expert knowledge. This process begins by collecting a large amount of historical operational data and conducting statistical analysis on key performance indicators (such as ammonia nitrogen concentration, nitrate nitrogen concentration, total nitrogen concentration, chemical oxygen demand, sludge settling ratio, specific oxygen consumption rate, nitrification rate, etc.). Their mean, standard deviation, quantiles, and other statistical characteristics are calculated, and frequency distribution histograms and cumulative distribution curves are plotted to understand the basic distribution characteristics of the data. Based on this, the starting point of the baseline early warning threshold is determined with reference to national and industry emission standards. Then, the statistical analysis results are interpreted and adjusted in conjunction with the opinions of process experts with extensive operational experience, transforming expert experience into quantitative guidance for threshold setting. Particularly important is considering the impact of seasonal factors on the biological treatment system, such as the significant effect of temperature changes on microbial activity. Through seasonal or temperature-range-specific subset analysis of historical data, specific threshold models for different seasons are constructed. Furthermore, considering factors such as fluctuations in influent water quality and quantity, a dynamic threshold adjustment mechanism is designed. For example, during the rainy season or under special conditions, the system can automatically adjust the stringency of the thresholds. Finally, a tiered early warning system is established, typically including three levels: mild warning (alert level), moderate warning (intervention level), and severe warning (emergency level). Each level has different thresholds and corresponding response measures. This multi-dimensional and dynamic early warning threshold design ensures the sensitivity and adaptability of the early warning system to changes in the state of the wastewater treatment system.
[0221] Based on the established tiered early warning threshold system, the system monitors and evaluates the predicted values of biochemical treatment efficiency indicators in real time, and generates structured early warning information promptly. Specifically, when the system obtains the predicted values of efficiency indicators for a future time point (e.g., 2 hours later, 4 hours later) from the optimized prediction model, it immediately compares them with the applicable tiered early warning thresholds under the current conditions. The comparison process considers not only the predicted values themselves but also fully utilizes the prediction confidence interval information output by the prediction model to assess the reliability of the prediction. The system will trigger an early warning only if both of the following conditions are met simultaneously: first, the predicted value exceeds (is worse than) the currently applicable early warning threshold, indicating a possible decline in treatment effectiveness; second, the prediction confidence interval meets preset requirements, such as the lower bound of the 95% confidence interval also exceeding the early warning threshold, or the confidence interval width being within an acceptable range, indicating that the prediction has sufficient credibility. When the triggering conditions are met, the system generates structured early warning information, including the following core elements: early warning time, including the predicted issuance time and the predicted occurrence time of the anomaly; early warning type, clearly indicating which performance indicators may have problems, such as excessive ammonia nitrogen or decreased sludge activity; severity, assessing the severity of the potential impact based on the degree and duration of the predicted value exceeding the threshold, typically categorized into three levels: mild, moderate, and severe; and credibility rating, assessing the credibility of the early warning based on the historical accuracy of the prediction model, the confidence interval of the current prediction, and the stability of the most recent prediction, typically categorized into three levels: high, medium, and low. This structured early warning information design provides both key information about the abnormal situation and an assessment of the uncertainty of the early warning, facilitating subsequent decision analysis and the formulation of intervention measures.
[0222] After the early warning information is generated, the system needs to provide professional problem diagnosis and solutions, which primarily relies on a systematically constructed wastewater treatment expert knowledge base. The core of the knowledge base is a set of ternary association rules for problem-cause-solution, which systematically represents the diagnostic and treatment logic for common problems in the wastewater treatment process. The knowledge base construction process includes: expert knowledge collection, gathering experience and knowledge from senior process engineers, operations managers, and researchers through structured interviews, case studies, and workshops; literature review, systematically sorting out process problem solutions in industry standards, technical manuals, and research literature; case analysis, retrospectively analyzing abnormal events in historical operation records to extract problem handling patterns; knowledge standardization, transforming the collected knowledge into rule entries in a unified format, with each rule including a problem description (e.g., abnormal specific indicators), possible cause analysis (e.g., insufficient dissolved oxygen, microbial inhibition), and recommended measures (e.g., adjusting aeration rate, adding carbon sources); knowledge organization, organizing rule entries according to process units (pretreatment, biological tanks, secondary sedimentation tanks, etc.), problem types (nitrification problems, sludge bulking, etc.), and severity to construct a multi-level index structure; knowledge verification, verifying the correctness and applicability of the rules through expert review and historical case testing; and a knowledge update mechanism, establishing a regular review and dynamic update process to ensure that the knowledge base keeps pace with the latest technological developments and practical experience. Through this systematic knowledge base construction process, an expert knowledge system with both theoretical foundation and practical support is formed, providing a reliable basis for subsequent problem diagnosis and solution formulation.
[0223] To overcome the limitations of purely rule-based expert knowledge bases in quantitative analysis, the system also needs to establish an experience-based learning model using a data-driven approach to capture the complex relationship between process control and its effects. This process begins with the collection and organization of historical data, including process control records (such as aeration rate adjustments, reflux ratio changes, and reagent dosing) and corresponding effect data (changes in various performance indicators). The data preprocessing stage addresses time alignment issues to ensure accurate causal correspondence between control actions and effect observations; it also addresses data quality issues, including outlier detection and missing value handling; and it employs feature engineering to construct effective features reflecting changes in the system state before and after control. The model building stage utilizes various machine learning methods, selecting appropriate algorithms based on problem characteristics and data availability: regression models (such as multiple linear regression and support vector regression) are suitable for predicting the quantitative impact of control parameter changes on performance indicators; time-series models (such as ARIMA and LSTM) are suitable for capturing the dynamic changes in performance indicators after control; reinforcement learning models are suitable for optimizing multi-step control sequences to maximize long-term effects; and causal inference models (such as Bayesian networks) are suitable for identifying the causal relationship network between control parameters and performance indicators. During model training, cross-validation is used to evaluate model performance and avoid overfitting. Sensitivity analysis is used to identify the most critical control parameters. An online learning mechanism is designed during deployment to enable the model to continuously learn and improve from new control records. Through this data-driven, experience-based learning model, the system can quantitatively evaluate the expected effects of different control schemes based on historical experience, providing more precise guidance for process control decisions.
[0224] Once a potential problem is detected and an early warning message is generated, a comprehensive diagnosis and suggestion generation process will be initiated. This process integrates knowledge base queries and empirical model reasoning to generate targeted process control recommendations. First, based on the type and severity of the early warning message, and combined with current system operating parameters (such as dissolved oxygen, pH, MLSS concentration, etc.), relevant problem-cause-response rules are retrieved from the expert knowledge base. The retrieval process employs fuzzy matching and similarity ranking to ensure that the most relevant rule entries are found. Simultaneously, the current system status and early warning message are input into an empirical learning model to predict the potential effects of changes in different control parameters, quantifying the effectiveness of various control options. Then, the knowledge base query results and model prediction results are merged, comprehensively considering the professional rationality of the rule recommendations and the quantitative accuracy of the model predictions to generate preliminary process control recommendations. These recommendations mainly focus on key operating parameters, including: aeration rate adjustment, which directly affects dissolved oxygen levels and microbial activity; changes in the return ratio, which affect sludge concentration distribution and retention time; and adjustments to reagent dosage, such as the addition of external carbon sources, coagulants, and pH adjusters. For each control parameter, it is recommended to clearly specify three elements: the direction of adjustment (increase, decrease, or remain unchanged); the magnitude of adjustment (specific numerical range or percentage change); and the timing of the operation (the time point and duration, considering system response lag). This comprehensive diagnostic method, combining expert knowledge and data models, has both theoretical support and practical basis, providing operators with more comprehensive and reliable decision support.
[0225] Based on the initial process control recommendations, further consideration is given to the multiple constraints and objectives of actual operation, leading to multi-objective optimization and the final control scheme. The optimization process first clarifies several potentially conflicting objectives: treatment effectiveness, ensuring effluent quality meets standards, especially key indicators in the early warning system; energy consumption control, reducing operational energy consumption, particularly for high-energy-consuming equipment such as aeration systems; operating costs, including direct costs such as reagent consumption and equipment maintenance; and system stability, avoiding system oscillations caused by drastic adjustments. Then, a multi-objective optimization model is constructed, parameterizing each control recommendation and setting quantitative evaluation indicators and constraints for each objective. The Pareto optimality principle is used for optimization, generating a series of non-dominated solutions, each achieving a different balance among different objectives. Based on the current system state and operating strategy, the solution that best meets the requirements is selected as the final scheme. Once the plan is finalized, the abstract control parameters are transformed into specific operational guidelines, including: an execution schedule, specifying the exact time points for each operational step; an operational sequence, especially for interdependent parameters, determining the optimal adjustment order; the scope of operation, breaking down large adjustments into multiple smaller steps for gradual implementation; precautions, highlighting key points and potential risks during operation; and monitoring requirements, indicating which parameter changes require close monitoring during the adjustment process. This detailed operational guidance ensures the control plan is accurately understood and implemented, while maximizing the balance between multiple objectives such as treatment effectiveness, energy consumption, and cost, thereby achieving optimized operation of the wastewater treatment system.
[0226] After the implementation of the plan, it is necessary to track its effectiveness and continuously optimize it to form a closed-loop feedback mechanism. First, monitoring points are established to record changes in key parameters in real time and compare them with the expected results. Then, the effectiveness of early warnings is assessed, including whether early warning indicators have returned to normal ranges and whether stable operation has been restored. Next, the entire early warning-diagnosis-control process is reviewed and analyzed to evaluate the accuracy of early warnings, the rationality of diagnosis, and the effectiveness of control, identifying areas for improvement. Finally, the difference between the actual and predicted results is fed back to the experience-learning model to update the reliability rating of relevant rules in the knowledge base, enabling the system to learn and continuously optimize itself. Through this closed-loop feedback mechanism, the early warning and control system can continuously accumulate experience, improve accuracy, and gradually develop into an adaptive, self-learning intelligent process management system, providing strong support for the stable operation and efficiency improvement of wastewater treatment plants.
[0227] Example 12
[0228] In this embodiment, the statistical analysis and processing technology based on historical performance index data, combined with expert knowledge, determines the early warning thresholds for different performance indicators, and dynamically adjusts the thresholds considering seasonal factors and changes in operating conditions to obtain tiered early warning thresholds, including:
[0229] Based on the historical performance index data, the historical value ranges of ammonia nitrogen concentration, nitrate nitrogen concentration, total nitrogen concentration, chemical oxygen demand (COD), sludge settling ratio, specific oxygen consumption rate (SOH), and nitrification rate are statistically analyzed. Using these values as coordinate axes and their respective historical value ranges as the value ranges for each axis, a multi-dimensional early warning space is constructed. This space is then divided into a grid structure according to a preset discretization step size. The intersections of the grid structure are defined as nodes. Directed edges are established between adjacent nodes to represent early warning state transition relationships. Weights are assigned to each directed edge based on the frequency and impact of state transitions in the historical performance index data. Based on statistical analysis of the historical performance index data, the historical occurrence probability and impact severity score are calculated for each node, resulting in a weighted multi-dimensional early warning space graph structure.
[0230] Among them, the historical occurrence probability of the node This refers to the frequency of a specific performance indicator combination state node appearing in the total historical data sample, reflecting the likelihood of that state occurring in actual operation. Impact Severity Score This is an indicator that comprehensively quantifies the negative impact of a certain node state on the wastewater treatment system. The specific calculation method is as follows: First, based on historical data, the degree to which each performance indicator exceeds the corresponding emission standard threshold under that node state is statistically analyzed. The degree of exceeding the standard of each indicator Defined as:
[0231]
[0232] in, For the state of this node, the first Historical average of each performance indicator These correspond to the emission standard thresholds. The degree of exceedance of each indicator constitutes a vector. A weight vector is preset based on the importance of each indicator to the compliance of effluent and the stability of the system. (If the sum of all components is 1), then the severity score is:
[0233]
[0234] The higher the value, the greater the potential harm that the node's state poses to the system, requiring a higher level of early warning response and more urgent regulatory intervention.
[0235] For each performance index among ammonia nitrogen concentration, nitrate nitrogen concentration, total nitrogen concentration, chemical oxygen demand, sludge settling ratio, specific oxygen consumption rate, and nitrification rate, a corresponding parameterized threshold function is constructed. The parameter set of the parameterized threshold function is defined as a parameter vector θ. The parameter vector θ includes a time-sensitive parameter controlling the influence of prediction lead time on the threshold strictness, a seasonal adjustment parameter controlling the influence of seasonal factors and historical data distribution on the baseline threshold, a dynamic adjustment parameter controlling the influence of current treatment load and process parameters on the allowable fluctuation range, and a cross-constraint parameter controlling the mutual influence between different performance indices. The parameterized threshold functions are combined according to the performance index type to form a threshold function set.
[0236] An objective function is constructed, which takes the parameter vector θ as input variable and comprehensively calculates the false negative rate, false positive rate, and alarm timeliness index generated based on the threshold function set. Cost weights are assigned to the false negative rate, the false positive rate, and the alarm timeliness index respectively. The optimization problem of the objective function is mapped to a graph search problem of finding the optimal path in the weighted multidimensional early warning space graph structure. An adaptive A-star algorithm or Monte Carlo tree search is applied to search in the parameter space of the parameter vector θ. During the search process, dynamic programming is applied to store intermediate results and pruning techniques are used to eliminate low-value search branches to obtain the optimal parameter θ star that makes the objective function reach the optimal value.
[0237] The optimal parameter θ is substituted into the threshold function set, which is then deployed for real-time threshold calculation. An online learning mechanism is established, and comparative data between actual warning results and real anomalies are collected as feedback on the warning effect. Based on the warning effect feedback, the parameter vector θ is incrementally updated using gradient descent or Bayesian optimization. Three severity levels—mild, moderate, and severe—are set based on the thresholds output by the threshold function set. For each severity level, two confidence levels—high certainty and low certainty—are set, constructing a multi-level warning system with six levels to obtain the graded warning thresholds.
[0238] Specifically, the process begins by utilizing abundant historical performance indicator data to construct a multidimensional early warning space and graph structure, providing a data foundation and computational framework for subsequent threshold optimization. This process starts with in-depth analysis of historical data for various key performance indicators (ammonia nitrogen concentration, nitrate nitrogen concentration, total nitrogen concentration, chemical oxygen demand, sludge settling ratio, specific oxygen consumption rate, nitrification rate, etc.), including calculating basic statistics (mean, median, standard deviation, quantiles, etc.) and distribution characteristics (normality test, skewness, kurtosis, etc.). Based on the analysis results, the historical numerical range for each indicator is determined, typically using the range from minimum to maximum value, or an effective range after removing extreme anomalies (such as the 3σ range or the 5%~95% quantile range). Next, these performance indicators are used as coordinate axes in the multidimensional space, with the value range of each axis representing the historical numerical range of that indicator, thus constructing an n-dimensional early warning space (where n is the number of performance indicators). To facilitate computation, this continuous space needs to be discretized into a grid structure. Choosing an appropriate discretization step size is crucial; too large a step size leads to a loss of accuracy, while too small a step size results in a significant increase in computational cost. After discretization, the intersection of each grid is defined as a node, representing a specific combination of performance indicators. The next step is to construct the connections between these nodes, establishing directed edges between adjacent nodes to represent possible state transition paths. By analyzing the continuous changes in performance indicators in historical data, we can identify which state transitions occur frequently in practice and which occur infrequently. Based on the above analysis of the continuous changes in performance indicators in historical data, we quantify and assign weights to each directed edge between adjacent nodes, specifically as follows: Let there be a directed edge... Indicates from node To the node To determine the state transition, first count the number of times the transition occurred in historical data. Divide it by the number of nodes Total number of state transitions starting from the start The normalized transition frequency is obtained. This reflects the historical probability of the transfer; secondly, it uses the target node... Severity of impact score (The calculation method is shown in section [0218-1]) After max-min normalization, the result is obtained. This serves as a quantitative indicator of the impact of the state transition on the system. (Directed edge) weight Defined as:
[0239]
[0240] in, This is a balancing coefficient used to adjust the proportion of historical occurrence frequency and impact severity in the weighting. The default value is [value missing]. In practical applications, the system's preference for false alarms and false negatives can be adjusted. The higher the value, the more frequently the state transition has occurred historically and the greater its potential impact on the system. Paths passing through such high-weight directed edges should be given priority in graph search and threshold optimization.
[0241] Simultaneously, two key indicators are calculated for each node (i.e., each state combination): historical occurrence probability, indicating the frequency of the state's occurrence in the past; and impact severity score, comprehensively considering the degree of impact of the state on effluent quality, treatment costs, and system stability. Through this series of processes, a weighted multidimensional early warning space graph structure is finally obtained. It not only represents the possible value space of the effectiveness indicators but also includes historical patterns of state transitions and impact assessments, providing a rich information structure for subsequent early warning threshold optimization.
[0242] Based on the constructed multidimensional early warning space map structure, parameterized threshold functions are further designed, transforming the setting of early warning thresholds from static rules to a dynamically adaptive mathematical model. For each performance indicator, a specific parameterized threshold function is constructed. These functions are no longer simple fixed values, but complex functions dynamically calculated based on multiple factors. The design of the threshold functions needs to consider the following key elements: baseline threshold, a basic reference value set according to emission standards and process requirements; time-sensitive adjustment, adjusting the stringency of the threshold as the prediction lead time increases, considering the increase in prediction uncertainty over time; seasonal adjustment, adjusting the threshold according to temperature and biological activity changes in different seasons, such as allowing slightly lower nitrification efficiency in winter; operating load adjustment, dynamically adjusting the acceptable fluctuation range according to the current influent load and treatment capacity; and cross-constraint adjustment, considering the mutual influence between different indicators, such as the correlation between nitrate and total nitrogen. All these adjustment factors are controlled by a parameter vector θ, which contains multiple sub-parameter groups: a time-sensitive parameter, controlling the functional relationship between prediction lead time and threshold strictness; a seasonal adjustment parameter, controlling the influence coefficient of seasonal factors and historical data distribution on the baseline threshold; a dynamic adjustment parameter, controlling the influence factors of current processing load and process parameters on the allowable fluctuation range; and a cross-constraint parameter, controlling the mutual influence weights between different performance indicators. These parameterized threshold functions are combined according to the type of performance indicator to form a function set, which together constitute the core calculation module of the early warning system. Through this parameterized design, the early warning threshold is no longer a static judgment standard, but an intelligent criterion that can be dynamically adjusted according to the system status and operating environment, greatly improving the adaptability and accuracy of the early warning system.
[0243] To determine the optimal parameter combination in the parameterized threshold function, an objective function needs to be constructed and an efficient optimization algorithm designed. The core of the objective function design is to balance multiple performance indicators of the early warning system: the false negative rate (the proportion of actual anomalies not detected by the system, reflecting the sensitivity of the early warning system); the false positive rate (the proportion of warnings issued by the system but no actual anomaly occurring, reflecting the specificity of the early warning system); and the alarm timeliness (the amount of time the warning precedes the actual anomaly, reflecting the foresight of the early warning system). These three indicators are usually interdependent and need to be assigned different weights in the objective function to reflect their relative importance in specific application scenarios. The objective function accepts a parameter vector θ as input, calculates the early warning effect on historical data using a set of threshold functions, comprehensively evaluates the three key indicators, and finally outputs an overall performance score. The optimization goal is to find the parameter vector θ that optimizes this performance score. To improve optimization efficiency, this parameter optimization problem is transformed into a graph search problem: finding the optimal path from the initial state to the target state in a weighted multidimensional early warning space graph structure, where the quality of the path is defined by the objective function. This transformation utilizes the graph structure characteristics of the early warning space, allowing the application of efficient graph search algorithms. The specific implementation employs two advanced algorithms: the Adaptive A* algorithm, combined with a heuristic function-guided search, significantly reduces the search space and is particularly suitable for situations where the target state is clearly known; and Monte Carlo tree search, which expands the search tree through random sampling and evaluation, excelling at handling high-dimensional spaces and complex objective functions. To further improve search efficiency, dynamic programming is applied to store and reuse intermediate computation results, avoiding redundant calculations; simultaneously, pruning techniques are used to eliminate obviously suboptimal search branches based on upper bound estimates. Through these optimization algorithms and techniques, the optimal parameter vector θ* that maximizes the objective function is obtained, defining the optimal parameter combination for the threshold function set, achieving an optimal balance of early warning performance on given historical data.
[0244] Finally, the optimized threshold function model is deployed to the actual early warning system, and a continuous improvement online learning mechanism is established. First, the optimal parameter θ* is substituted into the threshold function set to form a complete dynamic threshold calculation system. When deploying this system, an efficient calculation process needs to be designed to ensure rapid threshold judgment on real-time data streams; simultaneously, a suitable data caching mechanism should be established to store intermediate calculation results to improve efficiency. The output thresholds are graded, typically with three severity levels: mild warning (indicating performance indicators begin to deviate from the normal range and require attention); moderate warning (indicating the deviation is increasing and may require intervention); and severe warning (indicating significant anomalies and requiring immediate action). For each severity level, two confidence levels are further set: high certainty (indicating high reliability of the warning judgment, usually based on high-quality predictions and clear threshold exceedances); and low certainty (indicating some uncertainty in the warning judgment, possibly based on marginal predictions or slight threshold exceedances). This constructs a multi-level early warning system with six levels (three severity levels multiplied by two confidence levels), providing operators with more detailed risk assessments. After deployment, the system is not static but continuously optimized through an online learning mechanism: it collects feedback on the actual situation after each warning, records whether the warning is accurate and timely, and forms evaluation data on the warning effect; based on this evaluation data, the parameter vector θ is updated periodically or triggered, using methods such as gradient descent (suitable for cases where the parameter space is continuous and the objective function is smooth) or Bayesian optimization (suitable for cases where the parameter space is complex or the objective function is non-convex). Through this closed-loop feedback and continuous learning, the warning system can adapt to process changes and seasonal variations, maintaining long-term high performance.
[0245] In actual operation, the system's early warning results are integrated with the expert knowledge base and experience learning model to form a complete intelligent decision support system. When an early warning is triggered, the system not only provides warning information but also automatically queries the knowledge base for relevant problem diagnoses and handling suggestions based on the current operating conditions. Simultaneously, the experience learning model provides quantitative control suggestions based on the handling effects of similar historical situations. These suggestions, after multi-objective optimization, form an optimal control scheme that balances treatment effect, energy consumption, and cost. Finally, the optimized control scheme is transformed into specific operational guidance, including detailed information such as execution time, operation sequence, and precautions, for operators to refer to and implement. This complete closed loop from early warning to specific operational guidance greatly improves the intelligence level and operational efficiency of the wastewater treatment system. By continuously collecting operational feedback and effect evaluation, the system can continuously learn and improve, gradually increasing the accuracy of early warnings and the effectiveness of control suggestions, ultimately achieving intelligent management and optimized operation of the wastewater treatment system.
[0246] The tiered early warning threshold is a multi-level, dynamically adjustable judgment standard system used to assess whether predicted biochemical treatment efficiency indicators are in an abnormal or potentially risky state. In this invention, the tiered early warning threshold sets multiple severity levels and confidence levels for key indicators such as ammonia nitrogen, total nitrogen, and COD. The design principle of the tiered early warning threshold is based on risk stratification and dynamic adaptation, optimizing the early warning effect by balancing the detection rate (sensitivity) and false alarm rate (specificity). The steps for constructing the tiered early warning threshold include: statistically analyzing historical data to establish a baseline threshold; introducing a multi-factor dynamic adjustment mechanism, such as seasonal factors and load changes; setting severity levels and confidence levels; applying optimization algorithms to determine the optimal threshold parameters; and establishing an online learning mechanism to continuously optimize the threshold settings. Through this tiered and dynamic threshold system, the early warning system can identify potential problems in a timely and accurate manner, while reducing false alarms and missed alarms.
[0247] The multidimensional early warning spatial graph structure is a mathematical representation method that constructs a multidimensional space from multiple performance indicators, and then builds a graph structure composed of nodes and directed edges within this space to characterize the system state and its change patterns. In this invention, the structure uses each performance indicator as a coordinate axis to construct a discretized grid, and adds node weights and edge weights based on historical data. The construction principle of the multidimensional early warning spatial graph structure is to discretize a high-dimensional continuous space into a computable graph model, while retaining historical patterns of state transitions and impact assessment information. The construction steps include: determining the coordinate axis range of each indicator; performing spatial discretization according to a preset step size; defining grid intersections as state nodes; establishing directed edges between nodes based on historical data; calculating and assigning edge weights; and calculating the historical occurrence probability and impact severity for each node. This graph structure provides rich background information and a computational framework for threshold optimization, supporting the transformation of the threshold optimization problem into a graph search problem, thus improving computational efficiency.
[0248] A parameterized threshold function is a set of mathematical functions that transforms a static threshold into a dynamically adjustable calculation model based on various factors through an adjustable parameter vector. In this invention, the parameterized threshold function accepts factors such as time, season, and operating conditions as input and outputs a dynamically adjusted warning threshold. The design principle of the parameterized threshold function is to capture the relationship between the threshold and various influencing factors through mathematical modeling, enabling the threshold to adapt to the dynamic changes of the system. The construction steps include: determining the baseline threshold; identifying key influencing factors (time, season, operating conditions, etc.); designing adjustment functions for each factor; defining parameter vectors to control the adjustment intensity; combining the adjustment functions to form a complete threshold calculation model; and verifying the model's performance on historical data. Parameterized design transforms threshold judgment from static rules to dynamically adaptive intelligent calculation, significantly improving the accuracy and adaptability of the warning system.
[0249] A multi-level early warning system is a structured early warning classification system that provides fine-grained risk assessment and response guidance by combining different severity levels and confidence levels. In this invention, the multi-level early warning system includes three severity levels: mild, moderate, and severe. Each level is further divided into two confidence levels: high certainty and low certainty, forming a total of six early warning levels. The design principle of the multi-level early warning system is to separate the severity and certainty of early warning information, enabling response measures to more accurately match risk characteristics. The construction steps include: defining severity grading standards; designing confidence assessment methods; combining to form a multi-level matrix; configuring corresponding response strategies for each level; establishing escalation and demotion mechanisms between levels; and designing the presentation method of early warning information. This hierarchical early warning system enables operators to take appropriate measures based on the severity and certainty of the early warning, avoiding resource waste and risk neglect.
[0250] Experience-based learning (EPL) models are data-driven machine learning systems that analyze historical records of process control and their effects to capture the complex relationship between control parameters and processing efficiency, providing guidance for control decisions in new situations. In this invention, EPL models utilize supervised learning and reinforcement learning techniques to mine successful experiences from historical data. The construction principle of EPL models is based on pattern recognition and causal inference, establishing a mapping relationship between control behaviors and effects by identifying successful patterns in historical data. Development steps include: collecting historical control records and effect data; data preprocessing and feature engineering; selecting suitable learning algorithms (such as regression models, time series models, reinforcement learning models, etc.); model training and validation; and model deployment and online updates. EPL models compensate for the shortcomings of pure rule-based knowledge bases in quantitative analysis, providing accurate control suggestions based on historical data, and are a core component of intelligent decision-making systems.
[0251] Example 13
[0252] like Figure 3 As shown, the present invention also provides a biochemical treatment efficiency prediction system based on the dynamic evolution of water quality fingerprints, comprising:
[0253] The water quality fingerprint acquisition and feature extraction module 10 is used to acquire spectral data of the influent of the sewage treatment system and preprocess it to obtain a standardized spectral matrix. Based on the standardized spectral matrix, spectral feature parameters are extracted to form a water quality fingerprint feature vector. The continuously acquired water quality fingerprint feature vectors are organized according to time to construct a water quality fingerprint time series dataset.
[0254] The efficiency index collection and training data construction module 20 is used to collect the efficiency index of the effluent from the biochemical treatment system and perform standardization processing to obtain standardized efficiency index data. Based on the water quality fingerprint time series dataset and the standardized efficiency index data, a time correspondence is established and a training dataset is constructed. Feature engineering processing is performed on the training dataset to obtain model training data.
[0255] The prediction model training and prediction module 30 is used to train and optimize the prediction model using the model training data to obtain an optimized prediction model, receive real-time water quality fingerprint time series data to construct real-time prediction data, and use the optimized prediction model to predict the real-time prediction data and output the predicted value of the biochemical treatment efficiency index.
[0256] The early warning and control decision module 40 is used to determine the graded early warning threshold based on historical performance index data, compare the predicted value of the biochemical treatment performance index with the graded early warning threshold to generate early warning information, query the expert knowledge base based on the early warning information and apply the experience learning model to generate process control suggestions, and optimize the process control suggestions to form a control scheme.
[0257] In this embodiment of the invention, an ultraviolet-visible spectrometer and a three-dimensional fluorescence spectrometer are installed at the inlet of the wastewater treatment system to acquire raw spectral data streams. These data streams are then preprocessed and feature extracted to construct a time-series dataset reflecting dynamic changes in water quality. Simultaneously, various performance indicators of the effluent from the biochemical treatment system are collected, and a time correspondence is established with the influent water quality fingerprint data to construct a training dataset. This training dataset is used to train a deep learning model (such as LSTM, BiLSTM, or Transformer) to realize the mapping relationship from water quality fingerprint time-series data to future biochemical treatment performance. Real-time water quality fingerprint data is received, and the trained model is used to predict future biochemical treatment performance. Based on the prediction results, early warning information and process control suggestions are generated.
[0258] The key technological innovations in this invention include: a time-series characteristic analysis and feature extraction method based on water quality fingerprints to capture the dynamic evolution of water quality; establishing a mapping relationship between water quality fingerprint time-series data and future biochemical treatment efficiency through deep learning technology to achieve advanced prediction; and combining an expert knowledge base and an experience learning model to transform the prediction results into actionable process control suggestions, forming a closed-loop control system.
[0259] This invention has the following advantages: it can predict the performance changes of the biochemical treatment system several hours in advance, providing sufficient response time for process control; by analyzing the dynamic evolution characteristics of water quality fingerprints, it improves the accuracy and reliability of prediction; it integrates expert experience and machine learning models, and the generated process control suggestions are highly practical and targeted; the entire system can achieve adaptive optimization, and the accuracy of prediction and control is continuously improved with the accumulation of data.
[0260] The main application areas of this invention include: operation and management of biochemical treatment systems in urban wastewater treatment plants; intelligent monitoring and control of industrial wastewater treatment facilities; and early warning and handling of water quality anomalies in water environment management. By implementing this invention, the treatment efficiency and effluent quality stability of wastewater treatment systems can be significantly improved, while energy consumption and operating costs can be reduced, resulting in significant environmental and economic benefits.
[0261] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for predicting the effectiveness of biochemical treatment based on the dynamic evolution of water quality fingerprints, characterized in that, include: Spectral data of the influent to the wastewater treatment system is acquired and preprocessed to obtain a standardized spectral matrix. Based on this standardized spectral matrix, spectral feature parameters are extracted to form a water quality fingerprint feature vector. The continuously acquired water quality fingerprint feature vectors are organized chronologically to construct a water quality fingerprint time-series dataset. This includes: acquiring the raw spectral data stream; performing scattering correction, internal filtration effect correction, baseline drift correction, and instrument response normalization on the raw spectral data stream to obtain a standardized UV-Vis absorption spectral matrix and a three-dimensional fluorescence spectral matrix; using the response values of each wavelength or wavelength pair in the standardized UV-Vis absorption spectral matrix and the three-dimensional fluorescence spectral matrix as the dimension of a high-dimensional feature space; projecting the high-dimensional features onto a two-dimensional or three-dimensional space using principal component analysis or t-SNE dimensionality reduction techniques to form a set of geometric feature points; based on water chemistry principles, when two feature points represent water quality features that are correlated or causally related during the biochemical treatment process, a connection is established between the two feature points to form an intersection graph; based on the intersection graph, phase... The correlation analysis selects feature points from the set of geometric feature points as terminal nodes of the Steiner tree. A parameter λ is introduced to balance the complexity of the tree with the fitting effect. Dynamic programming and divide-and-conquer strategies are applied, and the optimal parameter λ value is found through gradient descent or simulated annealing to obtain the optimal Steiner tree structure. Using the optimal Steiner tree structure, the tree's depth, width, and branching factor are calculated as topological features. The path length and path strength between terminal nodes are extracted as path features. The centrality index of each node in the tree is calculated as node centrality features. Dynamic evolution parameters are extracted by comparing the changes in the Steiner tree structure at different time points. The topological features, path features, node centrality features, and dynamic evolution parameters are combined with characteristic wavelength absorbance, slope exponent, spectral integral area, fluorescence region volume, fluorescence peak intensity ratio, and parallel factor analysis component scores to obtain a multidimensional water quality fingerprint feature vector. The continuously collected multidimensional water quality fingerprint feature vectors are organized by timestamps to construct a time-series feature matrix with a sliding time window, resulting in the water quality fingerprint time-series dataset. The efficiency indicators of the effluent from the biochemical treatment system are collected and standardized to obtain standardized efficiency indicator data. A time correspondence is established based on the water quality fingerprint time series dataset and the standardized efficiency indicator data, and a training dataset is constructed. Feature engineering is performed on the training dataset to obtain model training data. The model training data is used to train and optimize the prediction model to obtain an optimized prediction model. Real-time water quality fingerprint time series data is received to construct real-time prediction data. The optimized prediction model is used to predict the real-time prediction data and output the predicted value of the biochemical treatment efficiency index. Based on historical performance index data, a graded early warning threshold is determined. The predicted value of the biochemical treatment performance index is compared with the graded early warning threshold to generate early warning information. Based on the early warning information, an expert knowledge base is queried and an experience learning model is applied to generate process control suggestions. The process control suggestions are then optimized to form a control scheme.
2. The method according to claim 1, characterized in that, The application of dynamic programming and divide-and-conquer strategies includes: The intersection graph is decomposed into multiple subgraphs, and the local optimal Steiner tree is solved independently for each subgraph. The intermediate results of solving each subgraph are recorded using a dynamic programming method. These intermediate results include the optimal connection method within the subgraph and the corresponding objective function value. By merging the locally optimal Steiner trees of each subgraph using a divide-and-conquer strategy, calculating the connection cost between subgraphs based on the intermediate results, selecting the connection scheme that optimizes the global objective function, and constructing a global Steiner tree structure.
3. The method according to claim 1, characterized in that, The efficiency indicators of the effluent from the biochemical treatment system are collected and standardized to obtain standardized efficiency indicator data. A time-series correlation is established based on the water quality fingerprint dataset and the standardized efficiency indicator data, and a training dataset is constructed. Feature engineering is performed on the training dataset to obtain model training data, including: The concentrations of ammonia nitrogen, nitrate nitrogen, total nitrogen, chemical oxygen demand (COD), sludge settling ratio, specific oxygen consumption rate (SOH), and nitrification rate at the effluent of the biochemical treatment system are collected. The units of the ammonia nitrogen concentration, nitrate nitrogen concentration, total nitrogen concentration, COD, sludge settling ratio, specific SOH, and nitrification rate are standardized to obtain standardized performance index data. Based on the hydraulic residence time, an influent timestamp is added to each performance index data point in the standardized performance index data. The influent timestamp is associated with the timestamp in the water quality fingerprint time series dataset to establish an influent-effluent correspondence. Based on the water quality fingerprint time series dataset and the standardized performance index data with the added water inflow timestamp, a data pair of correspondences between the water quality fingerprint sequence at time t and n hours before it and the performance index at time t+m is constructed by timestamp matching, where m represents the prediction lead time. Outlier identification and missing value processing are performed on the corresponding data pairs. Outlier data are processed using moving median filtering, interpolation, or statistical methods to obtain a quality-controlled training dataset. A sliding window design is applied to the quality-controlled training dataset, dividing the water quality fingerprint sequence into multiple time windows of fixed length. Temporal difference feature extraction is performed on the water quality fingerprint feature vectors within the time windows, the difference between feature values at adjacent time points is calculated, and periodic feature encoding and time encoding are added to the time windows to obtain structured model training data.
4. The method according to claim 3, characterized in that, The outlier identification and missing value handling for the corresponding data pairs include: The corresponding data pairs are subjected to quality review to identify outliers that exceed the normal range and missing values in the data records; The outliers in the time series are smoothed by using a moving median filter, and missing values are filled by interpolation or by using statistical methods to calculate the estimated values of missing data points, thus obtaining the quality-controlled training dataset.
5. The method according to claim 3, characterized in that, The process of training and optimizing a prediction model using the model training data to obtain an optimized prediction model, receiving real-time water quality fingerprint time-series data to construct real-time prediction data, and using the optimized prediction model to predict and output predicted values of biochemical treatment efficiency indicators from the real-time prediction data includes: The model training data is used to train a long short-term memory network model, a bidirectional long short-term memory network model, or a Transformer model. The model training is carried out using learning rate scheduling, gradient pruning, and early stopping strategies. The hyperparameters are adjusted through cross-validation to obtain the model parameters that have been trained and converged. The converged model parameters are loaded into the Long Short-Term Memory Network model, the Bidirectional Long Short-Term Memory Network model, or the Transformer model. The prediction accuracy, mean absolute error, and root mean square error are calculated on the validation set to evaluate the model performance under different prediction lead. Based on the prediction accuracy, mean absolute error, and root mean square error, the model is integrated or its structure is adjusted to obtain an optimized prediction model. Receive real-time water quality fingerprint time-series data, and apply feature engineering methods consistent with the sliding window design, time-series differential feature extraction, periodic feature encoding, and time encoding to construct real-time prediction data that meets the model input requirements; The optimized prediction model is used to make multi-time-scale predictions on the real-time prediction data, outputting the predicted values of biochemical treatment efficiency indicators at different future time points, and calculating the prediction confidence interval based on the optimized prediction model.
6. The method according to claim 5, characterized in that, The training of a Long Short-Term Memory (LSTM) network model, a Bidirectional LSM network model, or a Transformer model using the model training data employs learning rate scheduling, gradient pruning, and early stopping strategies. Hyperparameters are adjusted through cross-validation to obtain converged model parameters, including: The computational graph structure of the Long Short-Term Memory Network model, the Bidirectional Long Short-Term Memory Network model, or the Transformer model is represented as a directed tree, where nodes represent computational units in the model, and edges represent data flow and dependencies. Based on the forward propagation and gradient backpropagation characteristics of the model, the gradient norm, feature expression entropy and prediction error contribution rate of each node in the directed tree are calculated. Nodes with gradient norm lower than a preset gradient threshold, feature expression entropy lower than a preset entropy threshold, or prediction error contribution rate higher than a preset error threshold are identified as information bottleneck positions. An enhanced edge candidate set is generated for each of the information bottleneck positions, and each enhanced edge connects two nodes in the directed tree that are not originally directly connected. Define an augmentation edge utility function to quantify the improvement in model performance after adding each candidate edge in the augmentation edge candidate set. Construct a linear programming relaxation problem, use a rounding strategy better than 2 approximations to convert the fractional solution of the linear programming into an integer solution, and iteratively optimize to select the optimal subset of augmentation edges to add to the directed tree. Based on the optimal subset of enhanced edges, the model architecture is modified by adding residual connections, cross-layer connections, or attention mechanisms to form an enhanced connection structure. A parameter initialization strategy adapted to the enhanced connection structure is designed, and a phased training strategy is adopted. First, the original structural parameters of the model are frozen and only the connection weight parameters of the enhanced connection structure are trained. Then, the original structural parameters and the connection weight parameters are jointly optimized. Regularization techniques are applied to control the numerical range of the connection weight parameters to obtain the model parameters that have been trained and converged.
7. The method according to claim 6, characterized in that, The defined augmentation edge utility function quantifies the improvement in model performance after adding each candidate edge from the augmentation edge candidate set, including: Based on the two nodes connected by each candidate edge in the enhanced edge candidate set, each candidate edge is temporarily added to the directed tree to form a temporarily modified directed tree. According to the temporarily modified directed tree, a corresponding temporary connection structure is added to the model. The model training data is used to perform forward propagation calculation on the model after adding the temporary connection structure. The difference in prediction error before and after adding each candidate edge is calculated on the validation set to obtain the change in prediction error. The number of new parameters and the number of new floating-point operations in the model caused by adding temporary connection structures corresponding to each candidate edge are counted to obtain the increase in model parameters and the increase in computational cost. Based on the change in prediction error, the increase in model parameters, and the increase in computational cost, the ratio of the change in prediction error to the weighted sum of the increase in model parameters and the increase in computational cost is calculated. This ratio is defined as the utility function of the enhanced edge, and the utility score of each candidate edge in the candidate enhanced edge set is obtained.
8. The method according to claim 6, characterized in that, The phased training strategy first freezes the original structural parameters of the model and trains only the connection weight parameters of the enhanced connection structure, then jointly optimizes the original structural parameters and the connection weight parameters, including: Based on the optimal subset of enhanced edges, the model architecture is modified by adding residual connections, cross-layer connections, or attention mechanisms to the node positions corresponding to the optimal subset of enhanced edges to form an enhanced connection structure, thus obtaining the model after adding the enhanced connection structure. Based on the model with the added enhanced connection structure, the original structural parameters of the Long Short-Term Memory Network model, the Bidirectional Long Short-Term Memory Network model, or the Transformer model and the connection weight parameters of the enhanced connection structure are labeled. The original structural parameters are set to an untrainable state, and the connection weight parameters are set to a trainable state. The connection weight parameters are optimized by gradient descent using the model training data. The convergence of the loss function is monitored on the validation set to obtain the connection weight parameters after the first stage of optimization. Remove the untrainable state of the original structural parameters, set the original structural parameters and the connection weight parameters optimized in the first stage to trainable state simultaneously, and use the model training data to perform joint gradient descent optimization on the original structural parameters and the connection weight parameters optimized in the first stage to obtain the original structural parameters and connection weight parameters optimized in the second stage. Based on the connection weight parameters optimized in the second stage, the L1 norm and L2 norm of the connection weight parameters are calculated, and a regularization loss term is constructed as a weighted combination of the L1 norm and the L2 norm. The regularization loss term is added to the total loss function of the model after adding the enhanced connection structure. The sparsity and numerical value of the connection weight parameters are controlled by adjusting the regularization coefficient, and the model parameters of the training convergence are obtained.
9. The method according to claim 5, characterized in that, The process involves determining a tiered early warning threshold based on historical performance index data, comparing the predicted values of the biochemical treatment performance indicators with the tiered early warning threshold to generate early warning information, querying an expert knowledge base based on the early warning information and applying an experience learning model to generate process control suggestions, and optimizing the process control suggestions to form a control scheme, including: Based on the statistical analysis and processing requirements of historical performance index data, combined with expert knowledge, the early warning thresholds for different performance indicators are determined, and the thresholds are dynamically adjusted considering seasonal factors and changes in working conditions to obtain graded early warning thresholds. The predicted value of the biochemical treatment efficiency index is compared with the graded early warning threshold. When the predicted value of the biochemical treatment efficiency index exceeds the graded early warning threshold and the prediction confidence interval meets the preset requirements, early warning information including early warning time, early warning type, severity and credibility is generated. Integrate the experience of experts and process operation procedures in the field of wastewater treatment to establish an expert knowledge base that includes a three-element correlation rule of problem-cause-measure; Collect historical process control records and corresponding effect data, and use machine learning methods to mine the relationship patterns between control parameters and effects to obtain an experience learning model; Based on the warning information, combined with the current system operating status parameters, the expert knowledge base is queried and the experience learning model is applied to generate process control suggestions that include the adjustment direction, magnitude and timing of aeration volume, reflux ratio and reagent dosage. The proposed process control recommendations are optimized through multi-objective optimization to balance treatment effect, energy consumption, and operating cost, forming an optimized control scheme. The optimized control scheme is then transformed into execution guidelines that include execution time, operation sequence, and precautions.
10. The method according to claim 9, characterized in that, The statistical analysis and processing techniques based on historical performance index data, combined with expert knowledge, determine early warning thresholds for different performance indicators. These thresholds are then dynamically adjusted, taking into account seasonal factors and changes in operating conditions, resulting in tiered early warning thresholds, including: Based on the historical performance index data, the historical value ranges of ammonia nitrogen concentration, nitrate nitrogen concentration, total nitrogen concentration, chemical oxygen demand (COD), sludge settling ratio, specific oxygen consumption rate (SOH), and nitrification rate are statistically analyzed. Using these values as coordinate axes and their respective historical value ranges as the value ranges for each axis, a multi-dimensional early warning space is constructed. This space is then divided into a grid structure according to a preset discretization step size. The intersections of the grid structure are defined as nodes. Directed edges are established between adjacent nodes to represent early warning state transition relationships. Weights are assigned to each directed edge based on the frequency and impact of state transitions in the historical performance index data. Based on statistical analysis of the historical performance index data, the historical occurrence probability and impact severity score are calculated for each node, resulting in a weighted multi-dimensional early warning space graph structure. For each performance index among ammonia nitrogen concentration, nitrate nitrogen concentration, total nitrogen concentration, chemical oxygen demand, sludge settling ratio, specific oxygen consumption rate, and nitrification rate, a corresponding parameterized threshold function is constructed. The parameter set of the parameterized threshold function is defined as a parameter vector θ. The parameter vector θ includes a time-sensitive parameter controlling the influence of prediction lead time on the threshold strictness, a seasonal adjustment parameter controlling the influence of seasonal factors and historical data distribution on the baseline threshold, a dynamic adjustment parameter controlling the influence of current treatment load and process parameters on the allowable fluctuation range, and a cross-constraint parameter controlling the mutual influence between different performance indices. The parameterized threshold functions are combined according to the performance index type to form a threshold function set. An objective function is constructed, which takes the parameter vector θ as input variable and comprehensively calculates the false negative rate, false positive rate, and alarm timeliness index generated based on the threshold function set. Cost weights are assigned to the false negative rate, the false positive rate, and the alarm timeliness index respectively. The optimization problem of the objective function is mapped to a graph search problem of finding the optimal path in the weighted multidimensional early warning space graph structure. An adaptive A-star algorithm or Monte Carlo tree search is applied to search in the parameter space of the parameter vector θ. During the search process, dynamic programming is applied to store intermediate results and pruning techniques are used to eliminate low-value search branches to obtain the optimal parameter θ star that makes the objective function reach the optimal value. The optimal parameter θ is substituted into the threshold function set, which is then deployed for real-time threshold calculation. An online learning mechanism is established, and comparative data between actual warning results and real anomalies are collected as feedback on the warning effect. Based on the warning effect feedback, the parameter vector θ is incrementally updated using gradient descent or Bayesian optimization. Three severity levels—mild, moderate, and severe—are set based on the thresholds output by the threshold function set. For each severity level, two confidence levels—high certainty and low certainty—are set, constructing a multi-level warning system with six levels to obtain the graded warning thresholds.
11. A biochemical treatment efficiency prediction system based on the dynamic evolution of water quality fingerprints, characterized in that, include: The water quality fingerprint acquisition and feature extraction module is used to acquire spectral data of the influent of the wastewater treatment system and preprocess it to obtain a standardized spectral matrix. Based on the standardized spectral matrix, spectral feature parameters are extracted to form a water quality fingerprint feature vector. The continuously acquired water quality fingerprint feature vectors are organized by time to construct a water quality fingerprint time-series dataset, including: acquiring the original spectral data stream; performing scattering light correction, internal filtration effect correction, baseline drift correction, and instrument response normalization on the original spectral data stream to obtain a standardized ultraviolet-visible absorption spectral matrix and a three-dimensional fluorescence spectral matrix; using the response value of each wavelength or wavelength pair in the standardized ultraviolet-visible absorption spectral matrix and the three-dimensional fluorescence spectral matrix as the dimension of the high-dimensional feature space; projecting the high-dimensional features onto a two-dimensional or three-dimensional space through principal component analysis or t-SNE dimensionality reduction technology to form a set of geometric feature points; based on the principles of water chemistry, when the water quality features represented by two feature points in the set of geometric feature points have a correlation or causal relationship in the biochemical treatment process, a connection is established between the two feature points to form an intersection map; based on the... The intersection point diagram is used to select feature points from the set of geometric feature points as terminal nodes of the Steiner tree through correlation analysis. A parameter λ is introduced to balance the complexity and fitting effect of the tree. Dynamic programming and divide-and-conquer strategies are applied, and the optimal parameter λ value is found through gradient descent or simulated annealing to obtain the optimal Steiner tree structure. Using the optimal Steiner tree structure, the depth, width, and branching factor of the tree are calculated as topological features. The path length and path strength between terminal nodes are extracted as path features. The centrality index of each node in the tree is calculated as node centrality features. Dynamic evolution parameters are extracted by comparing the changes in the Steiner tree structure at different time points. The topological features, path features, node centrality features, and dynamic evolution parameters are combined with characteristic wavelength absorbance, slope exponent, spectral integral area, fluorescence region volume, fluorescence peak intensity ratio, and parallel factor analysis component scores to obtain a multidimensional water quality fingerprint feature vector. The continuously collected multidimensional water quality fingerprint feature vectors are organized by timestamps to construct a time-series feature matrix with a sliding time window, resulting in the water quality fingerprint time-series dataset. The efficiency index collection and training data construction module is used to collect the efficiency index of the effluent from the biochemical treatment system and perform standardization processing to obtain standardized efficiency index data. Based on the water quality fingerprint time series dataset and the standardized efficiency index data, a time correspondence is established and a training dataset is constructed. Feature engineering processing is performed on the training dataset to obtain model training data. The prediction model training and prediction module is used to train and optimize the prediction model using the model training data to obtain an optimized prediction model, receive real-time water quality fingerprint time series data to construct real-time prediction data, and use the optimized prediction model to predict the real-time prediction data and output the predicted value of the biochemical treatment efficiency index. The early warning and control decision-making module is used to determine the graded early warning threshold based on historical performance index data, compare the predicted value of the biochemical treatment performance index with the graded early warning threshold to generate early warning information, query the expert knowledge base based on the early warning information and apply the experience learning model to generate process control suggestions, and optimize the process control suggestions to form a control scheme.
Citation Information
Patent Citations
Method and system for determining formation pressure of underground gas storage, electronic equipment and storage medium
CN117473846A
A multi-parameter environmental quality intelligent monitoring system and control method thereof
CN119756489A