Water quality fingerprint identification and feedforward dosing control method and system
By acquiring the UV254 absorbance and three-dimensional fluorescence spectral characteristics of the influent and combining them with a model for carbon source dosing control, the problems of lag and waste in carbon source dosing in wastewater treatment systems have been solved, and rapid and precise carbon source dosing control has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-15
AI Technical Summary
Existing wastewater treatment systems lack rapid and accurate methods for identifying the characteristics of influent organic matter, resulting in delays and waste in carbon source addition and making it difficult to achieve precise control.
By acquiring the UV254 absorbance and three-dimensional fluorescence spectral characteristics of the influent, water quality fingerprint data is generated. Combined with the UV254-BOD/COD prediction model and the fluorescence characteristic ratio correction model, a feedforward-feedback composite control strategy is adopted to achieve precise carbon source addition.
It enables rapid identification and accurate assessment of the characteristics of organic matter in influent, shortens response time, avoids excessive or insufficient carbon sources, and improves control accuracy and system stability.
Smart Images

Figure CN121596813B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wastewater treatment, and in particular to a water quality fingerprint identification and feedforward dosing control method and system for biochemical treatment systems, applied to carbon source dosing control in wastewater treatment processes. Background Technology
[0002] In the field of wastewater treatment, the denitrification efficiency of biological treatment systems directly affects effluent quality and treatment costs. During biological denitrification, the carbon-to-nitrogen ratio (C / N) is a key parameter influencing denitrification efficiency. Because many industrial and municipal wastewater sources are insufficient in carbon, external carbon sources are required to ensure the smooth progress of the denitrification process.
[0003] Currently, common carbon source addition control methods mainly include fixed dosage mode and simple feedback control based on water quality parameters. Fixed dosage mode sets a constant amount of carbon source to be added according to design parameters, which cannot cope with fluctuations in influent water quality; while simple feedback control usually depends on the total nitrogen or nitrate nitrogen concentration in the effluent, and only adjusts the carbon source dosage when the nitrogen content in the effluent exceeds the standard, which has obvious lag.
[0004] In recent years, advanced control strategies based on online monitoring technology have been increasingly applied to wastewater treatment processes. These technologies typically use conventional water quality parameters such as COD and BOD as control criteria, but the measurement of these parameters often takes a long time, making rapid response difficult. In particular, BOD measurement usually takes 5 days, and even rapid BOD measurement methods take several hours, failing to meet the needs of real-time control. Furthermore, existing control strategies generally lack the ability to assess the biodegradability of influent organic matter, and relying solely on total quantity indicators makes it difficult to accurately calculate the required carbon source.
[0005] The main shortcomings of existing technologies are: on the one hand, there is a lack of rapid and accurate means to identify the characteristics of organic matter in the influent, making it difficult to obtain information on the biodegradability of the influent in a timely manner; on the other hand, the control system mostly adopts a single feedback control mode, which has a significant lag in response to sudden changes in the quality of the influent, and cannot predictably adjust the amount of carbon source added, resulting in either excessive addition causing waste or insufficient addition affecting the treatment effect. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides a water quality fingerprinting and feedforward dosing control method and system. This method rapidly detects the water quality characteristics of the influent, assesses its biodegradability, calculates the carbon source demand, and employs a feedforward-feedback composite control strategy to achieve precise carbon source dosing.
[0007] This invention provides a water quality fingerprint recognition and feedforward dosing control method, comprising:
[0008] The UV254 absorbance value of the influent and the signal intensity of the tyrosine-like fluorescence peak and the tryptophan-like fluorescence peak in the three-dimensional fluorescence spectrum are obtained. The UV254 absorbance value, the signal intensity of the tyrosine-like fluorescence peak and the signal intensity of the tryptophan-like fluorescence peak are fused and preprocessed to generate influent water quality characteristic fingerprint data.
[0009] Based on the influent water quality fingerprint data, the predicted BOD / COD ratio and available carbon content are calculated using a pre-established UV254-BOD / COD prediction model and a fluorescence feature ratio correction model.
[0010] Based on the predicted BOD / COD ratio and effective carbon content, combined with the real-time monitored influent flow rate and total nitrogen concentration, the theoretically required carbon source amount is calculated according to the denitrification theoretical model, and the effective carbon content is deducted to obtain the external carbon source demand. After time-series adjustment of the external carbon source demand, the predicted carbon source dosage is output.
[0011] Based on the predicted carbon source dosage as a feedforward control signal and combined with the total nitrogen monitoring data of the effluent as a feedback control signal, the final control quantity is calculated through a composite control algorithm, and the final control quantity is converted into an operation command for the carbon source dosing device and executed.
[0012] Further, the process of fusing and preprocessing the UV254 absorbance, the signal intensity of the tyrosine-like fluorescence peak, and the signal intensity of the tryptophan-like fluorescence peak to generate influent water quality characteristic fingerprint data includes:
[0013] Based on the UV254 absorbance, the signal intensity of the tyrosine-like fluorescence peak, and the signal intensity of the tryptophan-like fluorescence peak, the UV254 absorbance and the signal intensity are fused to generate preliminary fused data.
[0014] Based on the preliminary fused data, noise is eliminated using the sliding window averaging method to generate denoised data.
[0015] Based on the noise-reduced data, time tags are added to generate the influent water quality characteristic fingerprint data.
[0016] Further, the acquisition of the UV254 absorbance value of the influent and the signal intensity of the tyrosine-like fluorescence peak and the tryptophan-like fluorescence peak in the three-dimensional fluorescence spectrum includes:
[0017] Based on the inlet pipe of the sewage treatment system, an online UV-visible spectrometer and a three-dimensional fluorescence spectrometer are installed on the inlet pipe, and a sampling frequency of 5-15 minutes is set to establish a monitoring point.
[0018] Based on the online UV-visible spectrometer, the absorbance at a wavelength of 254 nm is continuously monitored and data is calibrated according to a preset calibration curve to obtain the calibrated UV254 absorbance value.
[0019] Based on the aforementioned three-dimensional fluorescence spectrometer, a three-dimensional fluorescence spectrum of the incoming water is acquired. A peak recognition algorithm is used to identify and quantify the signal intensity of the tyrosine-like fluorescence peak with an excitation wavelength of 230-275 nm and an emission wavelength of 300-320 nm, as well as the signal intensity of the tryptophan-like fluorescence peak with an excitation wavelength of 270-280 nm and an emission wavelength of 340-380 nm. The signal intensity of the tyrosine-like fluorescence peak and the signal intensity of the tryptophan-like fluorescence peak are then output.
[0020] Furthermore, a UV254-BOD / COD prediction model is established, including:
[0021] Based on at least three months of historical monitoring data of the influent of the wastewater treatment system, the influent UV254 absorbance, laboratory-measured BOD and COD values were collected to construct a training dataset.
[0022] Based on the training dataset, the first functional relationship between UV254 absorbance and BOD and the second functional relationship between UV254 absorbance and COD are established using multiple linear regression or support vector machine regression algorithms, respectively, to generate the UV254-BOD / COD prediction model.
[0023] Furthermore, a fluorescence characteristic ratio correction model is established, including:
[0024] Based on the signal intensity of the tyrosine-like fluorescence peak and the signal intensity of the tryptophan-like fluorescence peak, the intensity ratio of the tyrosine-like fluorescence peak to the tryptophan-like fluorescence peak is calculated.
[0025] Using the intensity ratio and the UV254 absorption value as input variables, and combining the laboratory-measured BOD / COD ratio as the output variable, a functional relationship between the input variables and the output variables is established to generate the fluorescence characteristic ratio correction model.
[0026] Furthermore, it also includes adaptive learning steps based on the deep actor criticism framework, including:
[0027] Based on historical influent water quality characteristic fingerprint data and corresponding measured BOD / COD values, carbon source dosage and treatment effect data, an expert database is constructed. The expert database contains data groups of water quality status, evaluation results, rewards and subsequent status.
[0028] Based on the historical influent water quality characteristic fingerprint data, an actor network is designed. The input of the actor network is the historical influent water quality characteristic fingerprint data, and the output of the actor network is the BOD / COD value and effective carbon content to be trained.
[0029] Based on the historical influent water quality characteristic fingerprint data and the BOD / COD value to be trained, a critic network is designed. The input of the critic network is the historical influent water quality characteristic fingerprint data and the BOD / COD value to be trained, and the output of the critic network is the carbon source dosage.
[0030] Based on the expert database, a preset number of data groups are randomly selected from the expert database to form batch training data. The actor network and the critic network are used to perform offline policy learning on the batch training data. The parameters of the actor network are updated by calculating the policy gradient and the parameters of the critic network are updated by minimizing the temporal difference error. The trained actor network model is then output.
[0031] Based on the influent water quality characteristic fingerprint data, the influent water quality characteristic fingerprint data is input into the trained actor network model, and the predicted BOD / COD ratio and available carbon content are output.
[0032] Further, the offline policy learning using the actor network and the critic network on the batch training data includes:
[0033] Based on the data set in the batch training data, the water quality status in the data set is input into the actor network, and the current policy action is output. The current policy action is the BOD / COD value and effective carbon content to be trained. The water quality status and the current policy action are input into the critic network, and the Q value of the current state-action pair is output.
[0034] Based on the Q value, the policy gradient is calculated and the actor network parameters are updated to generate the updated actor network.
[0035] Based on the rewards and subsequent states of the data set, calculate the target value of the temporal difference error, update the parameters of the critic network to minimize the squared difference between the Q value output by the critic network and the target value of the temporal difference error, and generate the critic network with updated parameters.
[0036] Based on the updated actor network with the parameters, the water quality state is input into the updated actor network, and the updated policy action is output. The water quality state and the updated policy action are input into the updated critic network, and the updated Q value is output. Based on the updated Q value, a policy constraint term is added to the policy gradient calculation to ensure that the updated policy does not deviate from the expert policy. The policy constraint term is the expected value of the square of the difference between the updated policy action and the expert action in the data set, and the trained actor network model is generated.
[0037] Furthermore, it also includes personalized computational steps based on multimodal dynamic agent learning, including:
[0038] Based on multimodal data including the influent water quality characteristic fingerprint data, the predicted BOD / COD ratio and effective carbon content, influent flow rate, pH value, temperature, total nitrogen concentration, ammonia nitrogen concentration, historical carbon source dosage and treatment effect data, the multimodal data is collected and integrated, and the multimodal data is standardized and denoised to generate preprocessed multimodal data.
[0039] Based on the preprocessed multimodal data, feature extraction and cross-modal fusion are performed on data from different sources through a gated cross-modal fusion network to generate fused feature representations;
[0040] Based on the fusion feature representation, water inflow type clustering is performed through a dual-constraint proxy optimization mechanism and a dynamic candidate management mechanism. The current water inflow sample is assigned to the identified water inflow type and the membership degree to each water inflow type is calculated. The clustering results and membership degrees are then output.
[0041] Based on the clustering results, a dedicated carbon source demand calculation model is constructed for each identified influent type. The dedicated carbon source demand calculation model maps the fusion feature representation to the carbon source dosage.
[0042] Based on the membership degree and the dedicated carbon source demand calculation model, the carbon source addition amount output by the dedicated carbon source demand calculation model for each influent type is weighted and combined to output the external carbon source demand amount.
[0043] Furthermore, the step of extracting features and fusing cross-modal data from different sources through a gated cross-modal fusion network to generate a fused feature representation includes:
[0044] Based on the preprocessed multimodal data, modality-specific encoders are used to encode the spectral data, conventional water quality parameter data, and process operation parameter data respectively, generating modality-specific features;
[0045] Based on the modality-specific features, the correlation strength between different modalities is calculated using an attention mechanism to generate an attention weight matrix;
[0046] Based on the modality-specific features and the attention weight matrix, information flow is controlled by a gating unit to perform weighted fusion of different modality features and generate the fused feature representation.
[0047] Furthermore, the influent type clustering through the dual-constraint proxy optimization mechanism and dynamic candidate management mechanism assigns the current influent sample to the identified influent type and calculates the membership degree to each influent type, outputting the clustering results and membership degrees, including:
[0048] Based on the fusion feature representation, initialize multiple inlet type proxy vectors;
[0049] Based on the fusion feature representation, the similarity between the current inflow sample and the proxy vector of each inflow type is calculated. The current inflow sample is assigned to each inflow type using a fuzzy clustering method, and the membership degree is calculated to generate preliminary clustering results.
[0050] The total nitrogen compliance rate and carbon source utilization efficiency of the effluent are obtained as treatment effect feedback data. Based on the preliminary clustering results and the treatment effect feedback data, the influent type proxy vector is updated periodically. The number of categories is dynamically adjusted according to the intra-class sample density and inter-class distance. The clustering results and membership degree are output.
[0051] Further, the step of adjusting the external carbon source demand based on time series and outputting the predicted carbon source dosage includes:
[0052] Based on the external carbon source demand, combined with the bioreaction kinetics and system hydraulic residence time, the external carbon source demand is adjusted in a time sequence, a phased addition strategy is formulated, and the predicted value of the carbon source addition is output.
[0053] Furthermore, the calculation of the final control quantity using a composite control algorithm, based on the predicted carbon source dosage as a feedforward control signal and combined with the effluent total nitrogen monitoring data as a feedback control signal, includes:
[0054] Based on the predicted carbon source dosage, the feedforward control output is calculated by the feedforward controller.
[0055] The total nitrogen concentration in the effluent is monitored in real time. The deviation between the total nitrogen concentration in the effluent and the preset target value is calculated. The feedback control output is calculated through a proportional-integral-derivative control algorithm.
[0056] Based on the feedforward control output and the feedback control output, the final control quantity is generated by weighting and combining the set feedforward weight coefficient and feedback weight coefficient.
[0057] Furthermore, after converting the final control quantity into an operation instruction for the carbon source dosing device and executing it, the process further includes:
[0058] Based on the final control quantity, the final control quantity is converted into an operation command to control the flow rate or switching frequency of the carbon source pump and sent to the carbon source dosing device;
[0059] Based on the execution process of the carbon source dosing device, the dosing flow rate, dosing time, and equipment status parameters are recorded to generate process data records;
[0060] Based on the process data records and the effluent total nitrogen monitoring data, the deviation between the predicted BOD / COD ratio and the measured value is calculated as a prediction accuracy index, and the effluent total nitrogen compliance rate is calculated as a control effect index. When the prediction accuracy index or the control effect index is lower than a preset threshold, the control parameters in the UV254-BOD / COD prediction model, the fluorescence feature ratio correction model, and the composite control algorithm are updated, and system optimization suggestions are output.
[0061] Furthermore, the final control quantity is calculated using a composite control algorithm, which further includes a robust optimization step based on hypercube half-space learning:
[0062] Based on the UV254 absorbance, the intensity ratio of the tyrosine-like fluorescence peak to the tryptophan-like fluorescence peak, the predicted BOD / COD ratio, the influent flow rate, and the total nitrogen concentration in the influent, a control feature space is defined, and each feature in the control feature space is normalized and mapped to a unit hypercube.
[0063] Based on historical control data, a labeled sample set is constructed, where the labels in the sample set represent the effectiveness of control decisions;
[0064] Based on the labeled sample set, a hypercube half-space model is learned using a full polynomial time learning algorithm. The hypercube half-space model includes hyperplane parameters. The hyperplane parameters are updated using a projective gradient descent method until the empirical risk is less than a preset accuracy threshold. The trained hypercube half-space model is then output.
[0065] Based on the trained hypercube half-space model, a feedforward control quantity is calculated. Combined with a feedback control quantity based on the total nitrogen monitoring data of the effluent, the feedforward weight coefficient and feedback weight coefficient are dynamically calculated according to the current position in the unit hypercube. The feedforward control quantity and the feedback control quantity are then weighted and combined to generate the final control quantity.
[0066] Further, the hypercube half-space model is learned using a full polynomial-time learning algorithm. This hypercube half-space model includes hyperplane parameters. The hyperplane parameters are updated using a projective gradient descent method until the empirical risk is less than a preset accuracy threshold. The trained hypercube half-space model is then output, including:
[0067] Based on the labeled sample set, a noise model is defined, which includes the mixing ratio of the real data distribution and the noise data distribution, and the maximum tolerable noise level is calculated.
[0068] Based on the labeled sample set, the hyperplane parameters are initialized. Through iterative processes, batch samples are randomly sampled, the gradient of the loss function is calculated, and the hyperplane parameters are updated by applying projective gradient descent. When the empirical risk is less than a preset accuracy threshold, the iteration stops, and the initially trained hypercube half-space model is output.
[0069] Based on the initially trained hypercube half-space model, it is extended into a multi-objective optimization framework. A multi-objective function is defined, which includes denitrification efficiency, carbon source utilization efficiency, and energy consumption. The Pareto front is calculated, and a trade-off point is selected on the Pareto front according to historical operating preferences. The trained hypercube half-space model is then output.
[0070] The present invention also provides a water quality fingerprint recognition and feedforward dosing control system, comprising:
[0071] The water quality fingerprint generation module is used to acquire the UV254 absorption value of the influent and the signal intensity of the tyrosine-like fluorescence peak and the tryptophan-like fluorescence peak in the three-dimensional fluorescence spectrum. The module performs data fusion and preprocessing on the UV254 absorption value, the signal intensity of the tyrosine-like fluorescence peak and the signal intensity of the tryptophan-like fluorescence peak to generate influent water quality characteristic fingerprint data.
[0072] The prediction calculation module is used to calculate the predicted BOD / COD ratio and available carbon content based on the influent water quality characteristic fingerprint data, using a pre-established UV254-BOD / COD prediction model and a fluorescence characteristic ratio correction model.
[0073] The carbon source demand calculation module is used to calculate the theoretically required amount of carbon source based on the predicted BOD / COD ratio and effective carbon content, combined with the real-time monitored influent flow rate and total nitrogen concentration in the influent, and subtract the effective carbon content to obtain the external carbon source demand. After time-series adjustment of the external carbon source demand, the predicted carbon source dosage value is output.
[0074] The composite control module is used to calculate the final control quantity based on the predicted carbon source dosage as a feedforward control signal, combined with the total nitrogen monitoring data of the effluent as a feedback control signal, through a composite control algorithm, and convert the final control quantity into an operation command for the carbon source dosing device and execute it.
[0075] The present invention has the following advantages over the prior art:
[0076] 1. This invention achieves rapid identification of the characteristics of organic matter in influent by acquiring the UV254 absorbance value and three-dimensional fluorescence spectral characteristics of the influent, and can provide biodegradability assessment results in a short time (minutes), which significantly shortens the response time compared with traditional BOD / COD determination methods (several hours to several days);
[0077] 2. The present invention adopts a feedforward-feedback composite control strategy, which can make a predictive response to changes in influent water quality, avoid the lag of simple feedback control, and improve control accuracy and system stability.
[0078] 3. This invention achieves accurate assessment of the biodegradability of organic matter in influent through data fusion and multi-model calculation, providing a reliable basis for calculating carbon source dosage and avoiding the problem of poor treatment effect caused by excessive or insufficient carbon source.
[0079] 4. The system structure of this invention is reasonable and the control logic is clear. It is easy to implement and promote in existing sewage treatment facilities and has good application prospects. Attached Figure Description
[0080] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0081] Figure 1 This is a flowchart illustrating the water quality fingerprinting and feedforward dosing control method of the present invention.
[0082] Figure 2 This is a flowchart illustrating the method for generating influent water quality characteristic fingerprint data according to the present invention.
[0083] Figure 3 This is a flowchart illustrating the method for establishing the UV254-BOD / COD prediction model of the present invention.
[0084] Figure 4 This is a schematic diagram of the water quality fingerprint recognition and feedforward dosing control system of the present invention. Detailed Implementation
[0085] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0086] Example 1
[0087] like Figure 1 As shown, the present invention provides a water quality fingerprint recognition and feedforward dosing control method, comprising:
[0088] Step S1: Obtain the UV254 absorption value of the influent and the signal intensity of the tyrosine-like fluorescence peak and the tryptophan-like fluorescence peak in the three-dimensional fluorescence spectrum. Perform data fusion and preprocessing on the UV254 absorption value, the signal intensity of the tyrosine-like fluorescence peak and the signal intensity of the tryptophan-like fluorescence peak to generate influent water quality characteristic fingerprint data.
[0089] Step S2: Based on the influent water quality characteristic fingerprint data, calculate using the pre-established UV254-BOD / COD prediction model and fluorescence characteristic ratio correction model, and output the predicted BOD / COD ratio and effective carbon content.
[0090] Step S3: Based on the predicted BOD / COD ratio and effective carbon content, combined with the real-time monitored influent flow rate and total nitrogen concentration, calculate the theoretically required carbon source amount according to the denitrification theoretical model and subtract the effective carbon content to obtain the external carbon source demand. After time-series adjustment of the external carbon source demand, output the predicted carbon source dosage value.
[0091] Step S4: Based on the predicted carbon source dosage as a feedforward control signal, and combined with the total nitrogen monitoring data of the effluent as a feedback control signal, calculate the final control quantity through a composite control algorithm, convert the final control quantity into an operation command for the carbon source dosing device, and execute it.
[0092] Specifically, in step S1, the generation of water quality characteristic fingerprint data is a fundamental step in achieving precise carbon source dosing. By acquiring and processing UV254 absorbance values and three-dimensional fluorescence peaks, a multidimensional dataset characterizing the influent water properties is established. This step includes three sub-steps: optical parameter acquisition, signal preprocessing and calibration, and multi-parameter data fusion.
[0093] First, real-time acquisition of optical parameters is fundamental to obtaining raw data. UV254 absorbance refers to the absorbance of a water sample at a 254 nm ultraviolet wavelength, primarily reflecting the content of aromatic compounds and organic matter containing conjugated double bonds in the water, and is positively correlated with dissolved organic matter (especially recalcitrant components). Three-dimensional fluorescence spectroscopy refers to the measurement of fluorescence intensity of samples under different excitation-emission wavelength combinations, forming a three-dimensional data matrix that can identify different types of fluorophores in the water. In the implementation of this method, an online UV254 analyzer (sampling frequency 15 minutes / time) and a three-dimensional fluorescence spectrometer (sampling frequency 30 minutes / time) were used for continuous monitoring of the incoming water. Two key characteristic peaks were observed in the three-dimensional fluorescence spectroscopy: the tyrosine-like fluorescence peak (excitation wavelength approximately 270-280 nm, emission wavelength approximately 300-320 nm, abbreviated as T peak), which mainly reflects the content of proteins, peptides, and tyrosine-containing substances in the water and is associated with easily biodegradable organic matter; and the tryptophan-like fluorescence peak (excitation wavelength approximately 270-280 nm, emission wavelength approximately 330-350 nm, abbreviated as W peak), which mainly characterizes components such as proteins and humic substances and is associated with the content of organic nitrogen in the water. The acquisition equipment was installed in the mixing area after the influent pump room to ensure sample representativeness, and was equipped with an automatic cleaning and calibration system to reduce the impact of biofilm and dirt on the measurements. The collected raw data was transmitted in real time to the central control system via an industrial-grade data acquisition unit, forming the basic data stream.
[0094] Secondly, preprocessing and calibration of the acquired signals are crucial for ensuring data quality. Raw optical signals typically contain various interferences and noises, requiring a series of preprocessing steps to improve signal quality. For UV254 data, baseline correction is first performed to eliminate instrument drift; then, medium-range filtering (window size 5) is applied to eliminate short-term spike interference; finally, temperature compensation is performed, adjusting the readings based on the experimentally determined temperature coefficient (approximately 0.026 / ℃) to eliminate the influence of water temperature variations. For three-dimensional fluorescence data, preprocessing is more complex: first, scattered light correction is performed to eliminate interference from Rayleigh and Raman scattering; then, internal filtering effect correction is performed, using synchronously measured absorption spectrum data to correct for fluorescence intensity attenuation; next, the PARAFAC (parallel factor analysis) algorithm is applied to decompose the fluorescence matrix and extract the main fluorescent components; finally, quantitative calibration is performed, using standards (such as proteins and humic acid) to establish a calibration curve for fluorescence intensity versus concentration. All calibration parameters are verified weekly using standard samples to ensure long-term measurement stability. In addition, an automatic outlier detection function has been implemented: when the rate of change of three consecutive measurements exceeds the preset threshold (25% for UV254 and 35% for fluorescence peak), a manual confirmation process is triggered to prevent abnormal data from entering the subsequent analysis stage.
[0095] Finally, multi-parameter data fusion is the core step in generating water quality characteristic fingerprints. Water quality characteristic fingerprint data refers to a multi-dimensional vector formed by combining multiple optical parameters, capable of comprehensively characterizing the type, content, and biodegradability of organic matter in the influent. In practical implementation, due to the different sampling frequencies of the UV254 analyzer and the fluorescence spectrometer, time alignment is first required. Cubic spline interpolation is used to map all parameters onto a unified 15-minute time grid. Then, characteristic ratios are calculated, especially the intensity ratio (T / W ratio) of the tyrosine-like fluorescence peak to the tryptophan-like fluorescence peak. This ratio is an important indicator of the biodegradability of organic matter. Next, parameter standardization is performed, using the Z-score method to convert each parameter into a standard form with a mean of 0 and a standard deviation of 1, eliminating the influence of differences in the dimensions and ranges of different parameters. To improve fingerprint recognition, principal component analysis (PCA) is also applied for dimensionality reduction and feature extraction, typically retaining the principal components (usually 3-4) needed to explain 95% of the variance. The final generated water quality fingerprint data is a multi-dimensional vector containing original parameters, feature ratios, and principal component scores, comprehensively characterizing the organic properties of the influent. A historical database of fingerprint data was also established; by comparing the current fingerprint with historical patterns, trends and anomalies in the influent characteristics can be quickly identified.
[0096] In step S2, the predictive analysis of organic matter characteristics is the foundation for calculating carbon source demand. By establishing a model relating UV254 absorbance to the BOD / COD ratio and correcting it using fluorescence characteristics, accurate predictions of the biodegradability of influent organic matter and available carbon content can be achieved. This step includes three sub-steps: application of the UV254-BOD / COD prediction model, correction of the fluorescence characteristic ratio, and calculation of available carbon content.
[0097] First, applying the UV254-BOD / COD prediction model is fundamental to rapidly assessing the biodegradability of organic matter. The BOD / COD ratio is a crucial indicator of wastewater biodegradability; traditional measurements require 5 days (BOD5), which is insufficient for real-time control. By establishing a prediction model for UV254 and the BOD / COD ratio, this parameter can be rapidly estimated. In practical applications, a machine learning model pre-trained with a large amount of historical data is used for prediction. This model employs an ensemble learning approach, combining the advantages of random forests and support vector regression (SVR), capable of handling the nonlinear relationship between UV254 and BOD / COD. Specifically, the random forest component constructs 100 decision trees, each trained using a randomly selected subset of features, effectively capturing major patterns in the data; the SVR component uses radial basis function (RBF) kernels, capable of finely handling local features and boundary conditions. The outputs of the two sub-models are weighted and averaged (weights of 0.7 and 0.3, respectively) to obtain the final predicted value. The model takes standardized UV254 values and their historical trends (differences between the first three time points) as input and outputs the predicted BOD / COD ratio. The model is recalibrated quarterly using the latest laboratory measurements to ensure it adapts to seasonal water quality changes. In real-time operation, the model receives the latest UV254 data and outputs a preliminary predicted BOD / COD ratio as the basis for subsequent corrections.
[0098] Secondly, applying fluorescence characteristic ratio correction is key to improving prediction accuracy. Relying solely on UV254 to predict BOD / COD has limitations, especially when the influent contains industrial wastewater or special components. The fluorescence characteristic ratio correction model improves prediction accuracy by incorporating fluorescence spectral information to correct the initial prediction results. The core of this correction model is the intensity ratio of tyrosine-like peaks to tryptophan-like peaks (T / W ratio), which is highly correlated with the biodegradability of organic matter: a high T / W ratio indicates a relatively high content of easily degradable proteins and good biodegradability; conversely, a low T / W ratio indicates poor biodegradability. The correction model uses a piecewise linear function, applying different correction coefficients according to different ranges of the T / W ratio: when the T / W ratio is within the normal range (0.8-1.2), the correction is small; when the T / W ratio deviates significantly from the normal range, the correction magnitude increases. Specifically, the model calculates a correction factor K based on the T / W ratio, and then multiplies the initial BOD / COD value predicted by UV254 by K to obtain the corrected value. The correction formula was derived from fitting a large amount of experimental data and took into account the effects of factors such as season and temperature. Furthermore, the correction model incorporates a dynamic weighting mechanism based on historical data: when the influent characteristics are continuously monitored as stable, the correction magnitude is reduced; when a sudden change in influent characteristics is detected, the correction magnitude is increased. This adaptive correction strategy enables the prediction system to better adapt to various influent conditions, including special cases such as industrial wastewater inrush and rainy season dilution.
[0099] Finally, calculating the available carbon content is a crucial step in assessing the amount of carbon source available for denitrification in the influent. Available carbon content refers to the amount of carbon source in the influent that can be utilized by denitrifying bacteria, and it forms the basis for determining the amount of additional carbon source needed. Based on the corrected BOD / COD ratio and the influent COD concentration, the available carbon content can be estimated. The calculation involves two steps: first, the readily biodegradable COD (RBCOD) content in the influent is estimated using an empirical formula: RBCOD = COD × f Z × (BOD / COD), where f Z A conversion factor (typically 1.5-1.8, adjusted according to wastewater characteristics) is used. Then, the effective carbon content is calculated based on RBCOD. Considering the carbon utilization efficiency during denitrification, a correction factor (typically 0.7-0.8) is applied to convert it into the actual amount of carbon available for nitrogen removal. This calculation also considers the distribution of carbon sources in the system, including biosynthetic consumption, aerobic leakage, and anaerobic fermentation. A calibration function based on historical data is also implemented: by comparing the predicted effective carbon content with the observed nitrogen removal effect during actual operation, the conversion and correction factors are dynamically adjusted to improve calculation accuracy. Furthermore, a temperature correction mechanism is introduced to adjust the effective carbon calculation according to seasonal temperature changes, reflecting the impact of temperature on microbial activity. The final output effective carbon content data provides crucial input for subsequent carbon source demand calculations.
[0100] In step S3, the calculation and adjustment of carbon source demand is the core step in determining the precise dosage. By integrating influent characteristic data, theoretical model calculations, and time series optimization, the optimal carbon source dosage strategy is determined. This step includes three sub-steps: theoretical carbon source demand calculation, effective carbon source deduction and determination of net demand, and time series adjustment based on system characteristics.
[0101] First, calculating the theoretical carbon source requirement based on a denitrification theoretical model is the scientific basis for determining the basic dosage. The denitrification theoretical model is a mathematical model describing the relationship between carbon source requirement and nitrogen removal during denitrification, based on the principles of microbial kinetics and stoichiometry. In practical applications, real-time data on influent flow rate and total nitrogen concentration (usually from online ammonia nitrogen and nitrate nitrogen analyzers and flow meters) are acquired and combined with the process parameters of the wastewater treatment system (such as the volume of the denitrification zone and the mixed liquor recirculation ratio) to calculate the theoretically required carbon source. The calculation considers several key factors: First, the target nitrogen removal rate is determined, typically based on the influent total nitrogen concentration, effluent total nitrogen standard, and the nitrogen removal pathways in the system (including assimilation, denitrification, etc.). Then, the theoretical carbon required to remove one unit of nitrogen is determined using stoichiometry. The classic denitrification stoichiometry C / N ratio is approximately 3.5-4.5 (based on COD), but in practical applications, this is adjusted based on factors such as carbon source type and sludge characteristics. Next, denitrification efficiency factors are considered, including temperature effects (corrected using a temperature coefficient θ=1.06), DO inhibition (significantly affecting denitrification when DO>0.5mg / L), and pH effects (optimal pH is 7-8.5). Finally, a safety factor (typically 1.1-1.2) is introduced to ensure stable nitrogen removal under fluctuating conditions. An adaptive optimization function is also implemented: by continuously monitoring the difference between the actual and predicted nitrogen removal effects, model parameters, such as the C / N ratio and efficiency coefficient, are dynamically adjusted to gradually bring the theoretical model closer to the actual system characteristics. This calculation method, which combines theory and practice, provides a scientific basis for carbon source addition.
[0102] Second, deducting the effective carbon content and determining the net demand for additional carbon source is a crucial step in optimizing carbon source usage. The additional carbon source demand refers to the amount of carbon source that needs to be added after considering the existing effective carbon in the influent, directly affecting carbon source utilization efficiency and cost. The calculation process first compares the theoretically required total carbon source with the effective carbon content of the influent calculated in the previous step: when the theoretical demand is greater than the effective carbon content, the difference is the amount of additional carbon source needed; when the theoretical demand is less than or equal to the effective carbon content, in principle, no additional carbon source is needed, but considering system stability, a minimum addition amount (such as 10% of the theoretical demand) is usually set as a baseline. In practical applications, this calculation also considers several practical factors: First, the problem of uneven carbon source distribution. The effective carbon in the influent may be unevenly distributed in the system, leading to local carbon source insufficiency. Therefore, a distribution coefficient (usually 0.8-0.9) is used to reduce the effective carbon. Second, the problem of carbon source utilization rate. Different types of added carbon sources (such as methanol, sodium acetate, etc.) have different utilization efficiencies, requiring a conversion coefficient to convert the theoretical demand into the actual dosage. Finally, the problem of system response time. Considering the lag between changes in influent characteristics and system response, predictive correction is introduced into the calculation, adjusting the dosage strategy in advance based on the changing trend of influent characteristics. An upper and lower limit control mechanism for carbon source dosage is also established: the upper limit is set based on the maximum pump capacity and system safety operation requirements to prevent over-dosing; the lower limit ensures minimum system stability requirements. This net demand calculation that considers multiple factors ensures denitrification effect while avoiding carbon source waste, achieving the basic goal of precise control.
[0103] Third, timing adjustments based on system characteristics are an advanced strategy for optimizing carbon source utilization efficiency. Simply calculating the total demand is insufficient for optimal control; the timing and distribution of carbon source addition must also be considered. Timing adjustments refer to the process of configuring the total dosage according to the optimal time distribution based on the bioreaction kinetics and system hydraulic retention time. In practical applications, the system designs specific timing strategies based on the characteristics of the wastewater treatment process (such as A² / O, MBBR, or SBR). For continuous flow processes (such as A² / O), a "front-increase + balanced" addition mode is typically adopted: 60-70% of the total demand is rapidly added at the front end to meet the initial high denitrification rate requirement; the remaining 30-40% is evenly distributed over subsequent time to maintain continuous denitrification. This distribution considers the denitrification kinetics (the initial reaction rate is the highest, then gradually decreases) and hydraulic characteristics (the flow distribution of wastewater in the system). For intermittent processes (such as SBR), timing adjustments are more refined, dynamically adjusting the addition ratio according to different operating stages (such as the influent period, reaction period, etc.). It also implements a dynamic adjustment function based on real-time feedback: by monitoring key parameters online (such as ORP, nitrate concentration, etc.), when a deviation between the actual denitrification effect and the expected result is detected, the dosage in subsequent periods is automatically fine-tuned, forming a closed-loop optimization. Furthermore, the timing strategy also considers seasonal factors: increasing the dosage ratio in the early stages during low-temperature seasons to compensate for the reduced reaction rate caused by low temperatures; and appropriately balancing the distribution during high-temperature seasons to improve overall utilization efficiency. This timing optimization based on system characteristics and kinetics ensures that carbon source dosage is not only precise in total amount but also optimal in temporal distribution, significantly improving carbon source utilization efficiency.
[0104] In step S4, the execution of the feedforward-feedback composite control is a key step in achieving the integration of theoretical calculations and actual control. By combining model-predictive feedforward control and effluent quality-based feedback control, robust and precise carbon source dosing control is achieved. This step includes three sub-steps: feedforward control signal generation, feedback control signal processing and fusion, and control command conversion and execution.
[0105] First, generating feedforward control signals based on predicted values is fundamental to achieving predictive control. Feedforward control is a control method that takes control measures before the disturbance affects the controlled system, based on the measurement and prediction of disturbances. In carbon source dosing control, the feedforward control signal originates from the predicted carbon source dosing value calculated in the aforementioned steps, representing an estimate of the optimal carbon source dosing based on the current influent characteristics and system state. In practical applications, the generation of feedforward control signals involves several technical measures: First, the predicted value is smoothed using an exponentially weighted moving average method to eliminate short-term fluctuations and prevent frequent changes in the control signal; then, a feasibility check is performed to ensure that the predicted value is within the system's physical constraints (such as pumping capacity, pipeline flow limits, etc.); next, a rate of change limit is applied to prevent excessive changes in the dosing amount within adjacent control cycles, typically limited to ±15% / hour, unless a significant change in influent characteristics is detected; finally, unit conversion is performed, converting the predicted carbon source demand (usually expressed as kg COD / h or mg COD / L) into actual control quantities (such as pump speed, flow rate, etc.). An anomaly handling mechanism was also designed: when a potential failure of the prediction model is detected (such as sensor malfunction, abnormal fluctuations in predicted values, etc.), the system automatically switches to a safe default value (usually the recent average dosage) to ensure stable system operation. The feedforward control signal is updated every 15-30 minutes, consistent with the water quality monitoring frequency, ensuring real-time control. This prediction-based feedforward control is proactive and can make appropriate adjustments before water quality changes affect the system, greatly improving the timeliness and accuracy of control.
[0106] Secondly, generating feedback control signals and fusing them with effluent total nitrogen monitoring data is crucial for ensuring long-term control stability. Feedback control is a method based on the deviation between the system output and the desired output, possessing self-correcting capabilities. In practical applications, the effluent total nitrogen concentration (usually the sum of ammonia nitrogen, nitrate nitrogen, and nitrite nitrogen) is continuously monitored using an online total nitrogen analyzer. The measured value is compared with the preset target value (usually 80-90% of the discharge standard, with a safety margin) to calculate the deviation. Based on this deviation, a PID (proportional-integral-derivative) control algorithm is used to calculate the feedback control output. The parameters (Kp, Ki, Kd) of the PID controller are determined through system identification and actual debugging, with typical values of Kp=0.8, Ki=0.05, and Kd=0.1, but adjustments are made based on system characteristics and operating conditions. Due to the certain lag in total nitrogen measurement (analysis cycle is typically 30 minutes to 1 hour), feedback control is mainly used for long-term correction, compensating for model prediction bias and system drift. The fusion of feedforward and feedback control signals adopts a weighted combination method:
[0107] U=α f ·U ff +β b ·U fb
[0108] Where α f and β b These are the feedforward and feedback weight coefficients, respectively, satisfying α f +β b = 1. During the stable operation phase, α f Typically set to 0.7-0.8, it emphasizes feedforward control, U ff For feedforward control output, U fb For feedback control output;
[0109] When a system anomaly or a decrease in model accuracy is detected, it will be dynamically adjusted to α. f =0.5-0.6, enhancing the feedback correction effect. A smooth transition mechanism was also designed, using gradual rather than abrupt changes when weights need adjustment to avoid sudden shocks to the system caused by abrupt changes in control quantities. This feedforward-feedback composite control strategy combines model-based predictive capabilities with measurement-based correction, enabling both rapid response to changes in influent and long-term maintenance of stable effluent, significantly improving control performance.
[0110] Third, converting the final control quantity into equipment operation commands and executing them is a crucial step in translating theoretical control into practical operation. "Operation commands" refer to specific instructions that directly control the actions of field equipment, such as pump start / stop, frequency setting, and valve opening. In carbon source dosing control, the final control quantity typically needs to be converted into specific operating parameters of the carbon source dosing pump. The conversion process considers several factors: first, the pump's performance curve, establishing a correspondence between the control quantity (e.g., kg / h) and the pump frequency or flow rate; second, the carbon source characteristics, including concentration, density, and viscosity, which affect the actual amount of effective component added; and third, the selection of the dosing point, which may require adjusting the carbon source distribution ratio in different reaction zones based on process characteristics and current operating conditions. The converted operation commands are then sent to the field controller (PLC or DCS) via an industrial communication network (e.g., Modbus, PROFIBUS, or Industrial Ethernet) to control the relevant equipment. A comprehensive monitoring mechanism was also designed: real-time recording of equipment operating status, dosage, and key parameters to form a complete operation log; periodic (e.g., daily or per shift) carbon source consumption statistics and dosage effect evaluations to provide a basis for management decisions; and a multi-level alarm mechanism that automatically triggers alarms and takes safety measures when equipment abnormalities, communication interruptions, or excessive control deviations are detected. To improve system reliability, redundancy design was implemented, including backup pumps and dual-path communication, ensuring basic functionality is maintained in the event of a single point of failure. Furthermore, a manual intervention interface is provided, allowing operators to temporarily adjust or take over control in special circumstances, enhancing the system's flexibility and safety.
[0111] Example 2
[0112] In this embodiment, obtaining the UV254 absorbance value of the influent and the signal intensity of the tyrosine-like fluorescence peak and the tryptophan-like fluorescence peak in the three-dimensional fluorescence spectrum includes: installing an online UV-Vis spectrometer and a three-dimensional fluorescence spectrometer on the influent pipe of the wastewater treatment system, setting the sampling frequency to 5-15 minutes, and establishing monitoring points; continuously monitoring the absorbance at a wavelength of 254 nm using the online UV-Vis spectrometer and calibrating the data according to a preset calibration curve to obtain the calibrated UV254 absorbance value; obtaining the three-dimensional fluorescence spectrum of the influent using the three-dimensional fluorescence spectrometer, identifying and quantifying the signal intensity of the tyrosine-like fluorescence peak with an excitation wavelength of 230-275 nm and an emission wavelength of 300-320 nm, and the signal intensity of the tryptophan-like fluorescence peak with an excitation wavelength of 270-280 nm and an emission wavelength of 340-380 nm using a peak recognition algorithm, and outputting the signal intensity of the tyrosine-like fluorescence peak and the signal intensity of the tryptophan-like fluorescence peak.
[0113] Specifically, monitoring equipment was first installed on the influent pipe of the wastewater treatment system. The selection and placement of the equipment are crucial, requiring consideration of water flow characteristics, water quality representativeness, and ease of maintenance. In this embodiment, the main influent pipe of Wastewater Treatment Plant A was selected as the primary monitoring point, as this location represents the influent characteristics of the entire plant area. During installation, sampling ports were first opened on the pipe, avoiding sedimentation areas and dead zones to ensure representativeness of the collected samples. Then, an automatic sampling system was installed, including a sampling pump, a filtration unit, and a sample dispensing system. The sampling pump was a peristaltic pump made of corrosion-resistant material, with a flow rate set to 500 mL / min; the filtration unit used a primary filter with a 100 μm pore size to remove large suspended particles and prevent clogging of subsequent monitoring equipment; the sample dispensing system used a solenoid valve to distribute water samples to different analytical instruments at regular intervals. Subsequently, an imported HACH DR3900 UV-Vis spectrometer and a domestically produced FP-8500 three-dimensional fluorescence spectrometer were installed. Both devices are secured with stainless steel brackets and equipped with a temperature control system to maintain the ambient temperature within 23±2℃, thus eliminating the impact of temperature fluctuations on the measurement results. After installation, the sampling frequency was set to once every 10 minutes. This frequency captures short-term fluctuations in water quality without generating excessive redundant data, while also taking into account the equipment's lifespan and maintenance costs.
[0114] Secondly, based on the installed UV-Vis spectrometer, the absorbance at a wavelength of 254 nm was monitored. The UV254 absorbance value refers to the absorbance of a water sample at a wavelength of 254 nm in the ultraviolet light, primarily reflecting the content of dissolved organic matter (especially aromatic organic matter) in the water. In practice, baseline calibration of the equipment is first required, using ultrapure water as a reference, with its absorbance set to zero. Then, automatic instrument calibration is performed every 72 hours to eliminate systematic errors caused by light source intensity attenuation and detector drift. During routine monitoring, a dedicated calibration curve was established to eliminate the influence of interference factors such as turbidity. This calibration curve was obtained by fitting the relationship between the absorbance of standard solutions with different turbidities (0-500 NTU range, in 50 NTU intervals) at 254 nm and the actual organic matter concentration. The specific calibration process includes: simultaneously measuring the absorbance and turbidity values at 254 nm for each sample, and then correcting the original absorbance according to the pre-established calibration curve to obtain the calibrated UV254 absorbance value. The calibration algorithm uses a polynomial correction method: UV254 calibration = UV254 original - (a × turbidity) 2 The formula is: + b × turbidity), where a and b are calibration coefficients, determined experimentally to be a = 0.00025 and b = 0.015. This method effectively eliminates the interference of turbidity on UV254 measurements, improving the accuracy and reliability of the data.
[0115] Finally, the three-dimensional fluorescence spectral characteristics of the influent were obtained using a three-dimensional fluorescence spectrometer. Three-dimensional fluorescence spectroscopy (EEM) is a three-dimensional spectrum characterizing the distribution of fluorescent substances in water. The horizontal axis represents the emission wavelength, the vertical axis represents the excitation wavelength, and the color intensity indicates fluorescence intensity. In wastewater, protein-like substances (such as tyrosine-like and tryptophan-like substances) are important fluorescent components, and their fluorescence characteristics are closely related to the biodegradability of organic matter. The FP-8500 three-dimensional fluorescence spectrometer was set with an excitation wavelength range of 200-400 nm, with 5 nm intervals; an emission wavelength range of 250-500 nm, with 2 nm intervals; a scan speed of 12000 nm / min; and excitation and emission slit widths of 5 nm. Before each measurement, the instrument automatically performed Raman scattering correction to eliminate interference from Raman scattering in the water. The acquired raw fluorescence spectral data underwent four processing steps: First, internal filtering effect correction was performed using absorbance for first and second-order internal filtering effect correction; then, Rayleigh scattering removal was performed using the tangent method to remove first and second-order Rayleigh scattering bands; next, spectral normalization was performed by dividing the spectral intensity by the fluorescence intensity of the quinine standard solution measured daily; finally, peak identification was performed using a self-developed "watershed-local maximum" composite algorithm to automatically identify characteristic fluorescence peaks. Specifically, the tyrosine-like fluorescence peak (T peak) was identified as the strongest signal within the excitation wavelength range of 230-275 nm and the emission wavelength range of 300-320 nm; the tryptophan-like fluorescence peak (W peak) was identified as the strongest signal within the excitation wavelength range of 270-280 nm and the emission wavelength range of 340-380 nm. The key steps of the peak identification algorithm include image filtering and denoising, fluorescence region segmentation, local maximum search, and peak feature extraction, enabling accurate identification and quantification of the signal intensity of specific fluorescence peaks against complex backgrounds. Finally, the signal intensities of the tyrosine-like fluorescence peak and the tryptophan-like fluorescence peak are output as important indicators for characterizing the properties of organic matter in the influent.
[0116] In the actual application at Wastewater Treatment Plant A, the installation and commissioning of the equipment took two weeks. After the system stabilized, continuous and automatic monitoring of water quality characteristics was achieved. The HACH DR3900 UV-Vis spectrometer has a measurement accuracy of ±0.005 absorbance units and a linear range of 0-3.5 absorbance units, fully meeting the requirements for wastewater influent monitoring. The FP-8500 three-dimensional fluorescence spectrometer has a signal-to-noise ratio greater than 1000:1, a wavelength accuracy of ±1.5 nm, and a fluorescence intensity repeatability error of less than 3%. Data from both devices is transmitted to the central control system via the MODBUS RTU protocol, enabling real-time data acquisition and storage. After three months of operation, the equipment reliability reached 99.5%, with only a small amount of data loss due to maintenance and calibration, which was supplemented using data interpolation methods from nearby time points. This high-frequency, continuous monitoring of water quality characteristics laid a solid data foundation for subsequent water quality fingerprinting and carbon source dosing control.
[0117] Example 3
[0118] like Figure 2 As shown, in this embodiment, the data fusion and preprocessing of the UV254 absorbance value, the signal intensity of the tyrosine-like fluorescence peak, and the signal intensity of the tryptophan-like fluorescence peak to generate influent water quality characteristic fingerprint data includes:
[0119] S11. Based on the UV254 absorption value, the signal intensity of the tyrosine-like fluorescence peak and the signal intensity of the tryptophan-like fluorescence peak, the UV254 absorption value and the signal intensity are fused to generate preliminary fused data.
[0120] S12. Based on the preliminary fused data, noise is eliminated using the sliding window averaging method to generate noise-reduced data;
[0121] S13. Based on the noise-reduced data, add time tags to generate the influent water quality characteristic fingerprint data.
[0122] Specifically, the fusion of multi-source spectral data is fundamental to constructing a water quality characteristic fingerprint. Feature-level fusion refers to combining data from different sources at the feature level to form a new feature representation, unlike original data-level fusion or decision-level fusion. In this embodiment, a feature-level fusion method is used to combine the UV254 absorbance value with the signal intensities of the tyrosine-like fluorescence peak (T peak) and the tryptophan-like fluorescence peak (W peak) into a three-dimensional feature vector [UV254, T peak, W peak]. This fusion method preserves the original information of each parameter and facilitates subsequent processing and model building. The fusion process first performs data standardization, unifying the three parameters to the same numerical range to avoid weight bias caused by different units. Standardization uses the Z-score method, i.e., Z = (X-μ) / σ, where X is the original value, μ is the historical average, and σ is the historical standard deviation. After standardization, all three parameters are converted to standard normal distribution data with a mean of 0 and a standard deviation of 1. Then, the three standardized parameters are aligned by timestamp to ensure that data from the same time point are combined together. To address potential data asynchrony issues (such as a missing parameter at a certain point in time), a linear interpolation method is used to fill in the missing values. Finally, the three aligned parameters are combined into a feature vector to form preliminary fused data. In the actual application at Wastewater Treatment Plant A, the data fusion process is executed every 10 minutes, consistent with the sampling frequency, ensuring the timeliness and continuity of the data.
[0123] Secondly, noise reduction of the initially fused data is a crucial step in improving data quality. In actual monitoring, sensor errors, environmental interference, and random fluctuations can cause the raw data to contain a large amount of noise, affecting the accuracy of subsequent analysis and decision-making. This embodiment uses the sliding window averaging method for noise reduction, a common time-series data smoothing technique. Specifically, it selects data from five sampling points (approximately a 50-minute time span) before and after the current time point, and calculates their arithmetic mean as the smoothed result for the current time point. The mathematical expression is:
[0124] X-smooth(t) = (1 / 5) × Σ[Xoriginal(t-2), Xoriginal(t-1), Xoriginal(t), Xoriginal(t+1), Xoriginal(t+2)], where t represents a time point and X represents a dimension of the feature vector. To handle the data at the beginning and end of the time series, a mirror-fill strategy was adopted, mirroring two points at each end of the sequence. This method is simple, effective, and computationally inefficient, suitable for online real-time processing, and effectively eliminates the impact of short-term random fluctuations while preserving the long-term trend of the data. After sliding window averaging, the signal-to-noise ratio of the data is significantly improved, and the average noise level is reduced. For outlier handling, a threshold detection mechanism is also implemented. When the data at a certain time point deviates from the average of the five points before and after it by more than three times the standard deviation, it is considered an outlier and replaced with linear interpolation of the preceding and following data. This method effectively identifies and handles data anomalies caused by equipment failure, sampling anomalies, or external interference, ensuring the quality of the denoised data.
[0125] Finally, time tags are added to the denoised data to generate the final influent water quality characteristic fingerprint data. Time tags are not only an index of the data but also an important dimension for analyzing water quality change patterns. In this embodiment, the time tags contain three levels of information: absolute timestamp (UTC time accurate to the second), relative time information (time since system startup), and time features (such as time of day, day of the week, month, etc.). The absolute timestamp adopts the ISO 8601 standard format (YYYY-MM-DDThh:mm:ssZ) for easy data exchange with other systems; the relative time information is in seconds for easy calculation of time intervals; and the time features extract periodic factors that may affect water quality, such as dividing a 24-hour day into six time periods (early morning, morning, forenoon, noon, afternoon, and evening) and marking a week as weekdays and weekends. These time features serve as auxiliary information for the water quality characteristic fingerprint, helping to discover the periodic change patterns of the influent water quality. For example, in the actual operation of Wastewater Treatment Plant A, it was found that the influent organic matter concentration peaked between 8-10 AM and 6-8 PM on weekdays, while weekends showed a different pattern. The correlation analysis between these temporal characteristics and water quality parameters provided important basis for subsequent predictive models. Finally, the feature vector [UV254, T peak, W peak] was combined with time labels to form complete influent water quality characteristic fingerprint data, which was stored in a time-series database, providing fundamental data support for subsequent biodegradability assessment and carbon source demand calculation.
[0126] Through the three steps described above, the original discrete monitoring data is transformed into structured, continuous water quality fingerprint data. This processed data is not only smooth and stable, reducing the influence of random factors, but also incorporates more dimensions of information through time stamps, enabling a more comprehensive and accurate reflection of the true characteristics and changing patterns of the influent water quality. In practical application at Wastewater Treatment Plant A, this method has made the trends of water quality changes clearer, effectively supporting subsequent decision analysis. The system generates 144 sets of water quality fingerprint data daily (once every 10 minutes), which are displayed in real-time through a visual interface and regularly generate trend reports to provide reference for process operation. After three months of operational verification, this data processing method has demonstrated excellent performance in capturing sudden changes in water quality, identifying typical patterns, and predicting short-term trends, providing a reliable data foundation for carbon source dosing control.
[0127] Example 4
[0128] like Figure 3 As shown, in this embodiment, the UV254-BOD / COD prediction model is established, including:
[0129] S21. Based on at least 3 months of historical monitoring data of the influent of the wastewater treatment system, collect the influent UV254 absorbance, laboratory-measured BOD and COD values from the historical monitoring data, and construct a training dataset.
[0130] S22. Based on the training dataset, use multiple linear regression or support vector machine regression algorithms to establish the first functional relationship between UV254 absorbance and BOD and the second functional relationship between UV254 absorbance and COD, respectively, to generate the UV254-BOD / COD prediction model.
[0131] Specifically, firstly, collecting historical monitoring data of the wastewater influent to the wastewater treatment system and constructing a training dataset is the foundation for model building. In this embodiment, historical data for the past six months (from January 2022 to June 2022) of wastewater treatment plant A were collected, including daily influent UV254 monitoring values and laboratory-measured BOD5 and COD values. The data collection time span covered water quality changes under different seasons and weather conditions, ensuring the representativeness and diversity of the data. The collected raw data included 180 complete data points, each containing the sampling date, UV254 value (unit: cm^-1), BOD5 value (unit: mg / L), COD value (unit: mg / L), and BOD5 / COD ratio. During data collection, special attention was paid to the consistency of sampling time, ensuring that the online UV254 monitoring value matched the laboratory sample collection time, typically selecting a fixed time period of 10:00 AM daily for comparison sampling. For the collected raw data, data cleaning was first performed to remove obvious outliers (such as data exceeding the normal range by 3 times the standard deviation) and incomplete records. The data was then standardized, converting UV254, BOD5, and COD values into Z-scores. Next, exploratory analyses were performed, including correlation analysis, scatter plot analysis, and distribution characteristic analysis, to preliminarily determine the relationships between variables. Finally, the processed dataset was randomly divided into a training set (126 groups) and a validation set (54 groups) in a 7:3 ratio. The training set was used for model training and parameter optimization, while the validation set was used to evaluate the model's generalization performance.
[0132] Secondly, based on the constructed training dataset, a prediction model was established between UV254 and BOD and COD. This embodiment attempted various regression algorithms, including multiple linear regression (MLR), support vector machine regression (SVR), and random forest regression (RFR). After comparison, the SVR algorithm performed best on this dataset, and therefore was ultimately selected to build the prediction model. Support vector machine regression is a nonlinear regression method based on statistical learning theory. Its core idea is to map the input space to a high-dimensional feature space and then construct a linear regression function in the feature space. The advantage of SVR is its ability to handle nonlinear relationships, its insensitivity to outliers, and its good performance even with small sample sizes. In the specific implementation, two SVR models were established: one for predicting BOD (first functional relationship) and the other for predicting COD (second functional relationship). Key parameters of the SVR model include kernel function type, regularization parameter C, epsilon bandwidth, and kernel function parameter gamma. The optimal parameter combination was determined using grid search combined with 5-fold cross-validation: for the BOD prediction model, a radial basis function (RBF) kernel function was selected, with C=10, epsilon=0.1, and gamma=0.01; for the COD prediction model, the RBF kernel function was also selected, but the parameters were adjusted to C=15, epsilon=0.05, and gamma=0.05. After model training, the model performance was evaluated on the validation set. The average relative error for BOD prediction was 8.5%, with a coefficient of determination (R²) of 0.86; the average relative error for COD prediction was 7.2%, with an R² of 0.86. 2 The value is 0.89. Compared with the traditional multiple linear regression method, the SVR model improves the prediction accuracy by about 20%, and is particularly robust when dealing with extreme water quality conditions.
[0133] To further enhance the model's application value, several key optimization measures were implemented. First, a model update mechanism was introduced, incrementally updating the model monthly using newly collected data to adapt to long-term changes in water quality characteristics. Second, prediction uncertainty estimation was implemented by generating multiple model replicas using the Bootstrap method and calculating the confidence intervals of predicted values, providing uncertainty information for decision-making. Third, segmented prediction models were established for different water quality conditions, selecting appropriate sub-models based on different ranges of UV254 values to further improve prediction accuracy. Finally, the established model was encapsulated as a web service, accessible through a RESTful API interface for other systems, enabling convenient model application.
[0134] In the actual operation of Wastewater Treatment Plant A, the UV254-BOD / COD prediction model receives new UV254 monitoring values every 10 minutes and outputs predicted BOD and COD values in real time. Compared with the traditional method of requiring 5 days for BOD5 measurement and several hours for rapid BOD, this significantly shortens the response time and provides timely decision-making basis for carbon source dosing control. Over the past three months of model operation, the average deviation from laboratory measurements has remained within 10%, fully meeting the accuracy requirements of engineering applications. Furthermore, through continuous comparison of predicted and measured values, a model drift detection mechanism has been established. When the prediction deviation exceeds 15% for five consecutive times, the model retraining process is automatically triggered to ensure the long-term effectiveness of the model. This rapid BOD / COD prediction method based on UV254 provides solid technical support for assessing the biodegradability of influent organic matter and is a key link in realizing water quality fingerprinting and feedforward control.
[0135] Example 5
[0136] In this embodiment, establishing a fluorescence characteristic ratio correction model includes: calculating the intensity ratio of the tyrosine-like fluorescence peak to the tryptophan-like fluorescence peak based on the signal intensity of the tyrosine-like fluorescence peak and the signal intensity of the tryptophan-like fluorescence peak; establishing a functional relationship between the input variables and the output variables based on the intensity ratio and the UV254 absorbance value as input variables, and combining the laboratory-measured BOD / COD ratio as the output variable, thereby generating the fluorescence characteristic ratio correction model.
[0137] Specifically, firstly, calculating the intensity ratio of the tyrosine-like fluorescence peak to the tryptophan-like fluorescence peak is the basis for the correction model. The tyrosine-like fluorescence peak (T peak) and the tryptophan-like fluorescence peak (W peak) are two major protein-like fluorescent substances in wastewater, and their relative content ratio is closely related to the composition and biodegradability of organic matter. In this embodiment, based on the three-dimensional fluorescence spectral data obtained in Example 1, the signal intensities of the T peak (excitation wavelength 230-275 nm, emission wavelength 300-320 nm region) and the W peak (excitation wavelength 270-280 nm, emission wavelength 340-380 nm region) for each sample were extracted, and then their ratio T / W was calculated. To ensure the accuracy and comparability of the calculation results, the original fluorescence signal was first subjected to multiple corrections. These included: removing the effects of Rayleigh scattering and Raman scattering, eliminating the internal filtration effect (fluorescence signal attenuation due to sample absorbance), performing instrument response correction, and standardization. The ratio of the corrected T peak and W peak signal intensities was calculated using a simple numerical division: T / W = I T / I W , where I T and I WThe values represent the corrected T-peak and W-peak signal intensities, respectively. During a 6-month monitoring period at Wastewater Treatment Plant A, the daily T / W ratio was recorded, and laboratory BOD / COD measurements were performed on water samples collected during the same period. Data analysis showed that the T / W ratio varied between 0.5 and 2.5, exhibiting a significant positive correlation with the BOD / COD ratio, with a correlation coefficient of 0.76. This finding is significant because the T / W ratio reflects the relative content of readily degradable proteins (corresponding to the T-peak) and recalcitrant proteins (corresponding to the W-peak), directly related to the biodegradability of organic matter. Particularly on days with a high proportion of industrial wastewater, the correlation between the T / W ratio and the traditional BOD / COD ratio was even more pronounced, providing a new approach for real-time assessment of influent biodegradability.
[0138] Secondly, based on the calculated T / W ratio and UV254 absorbance, a fluorescence characteristic ratio correction model was established. This model uses the T / W ratio and UV254 absorbance as input variables and the laboratory-measured BOD / COD ratio as the output variable, establishing a functional relationship between them. During model establishment, a correlation analysis was first performed to confirm a significant nonlinear relationship between the input and output variables. Then, various regression models were explored, including multiple linear regression, multinomial regression, support vector regression, and artificial neural networks. After comparison, a bivariate nonlinear regression model was selected, specifically in the following form:
[0139] BOD / COD = a + b (T / W) + c (A) UV254 +d·(T / W)·A UV254 +e·(T / W) 2 +f·A UV254 2 ,
[0140] Where a, b, c, d, e, and f are regression coefficients, which are optimized using the least squares method. A UV254 The UV254 absorbance is used. The model was trained using the training set (70% of the total data) and evaluated using the validation set (30%). The final model parameters were: a=0.12, b=0.18, c= -0.05, d=0.07, e= -0.03, f=0.01, and the coefficient of determination R on the validation set was... 2The model achieved a correlation coefficient of 0.85 and a mean relative error of 6.8%. To further improve the model's generalization performance, K-fold cross-validation (K=5) was implemented, and the results showed that the model performed stably on different subsets of data. Furthermore, sensitivity analysis was conducted to assess the importance of the input variables. The results showed that the T / W ratio accounted for approximately 65% of the model output, while the UV254 value accounted for approximately 35%, indicating that the fluorescence characteristic ratio is indeed an important indicator for assessing the biodegradability of organic compounds. After model training, a practical application test was performed. The real-time monitored T / W ratio and UV254 value were input into the model, and the predicted BOD / COD ratio was output. Compared with the laboratory measured values, the correlation coefficient was as high as 0.92, significantly better than the prediction model using only UV254 (correlation coefficient of 0.83).
[0141] In practical applications, the establishment of the fluorescence characteristic ratio correction model also includes several key optimization measures. First, in the data preprocessing stage, outliers are identified and processed using the 3σ criterion to ensure the model is not affected by extreme values. Second, feature engineering was implemented; in addition to the original T / W ratio and UV254 value, more feature combinations were explored, such as the ratio of the T peak to the total fluorescence intensity and the relative position change of the W peak. However, these additional features ultimately proved to have limited impact on model performance and were not adopted to maintain model simplicity. Third, seasonal adjustment of the model was implemented; by introducing time features (such as season and temperature) as auxiliary variables, a seasonal adjustment coefficient was established to enable the model to adapt to changes in wastewater characteristics under different seasons. Finally, an online update mechanism was implemented, using newly collected data to calibrate the model monthly, maintaining its timeliness.
[0142] In the actual operation of Wastewater Treatment Plant A, the biodegradability assessment system based on this modified model updates its predictions every 10 minutes, providing accurate and timely data support for carbon source dosing control. Especially when there are sudden changes in influent composition (such as a sharp increase in the proportion of industrial wastewater), the system can quickly detect changes in biodegradability and adjust the carbon source dosing strategy accordingly, avoiding carbon source waste or poor treatment results caused by inaccurate biodegradability assessments. Over the past three months of operation, the average deviation between this modified model and laboratory-measured BOD / COD ratios has remained within 7%, outperforming similar systems in the industry (typically 10-15%). Furthermore, correlation analysis with treatment effects confirms a high correlation between the biodegradability assessment results based on this model and actual nitrogen removal efficiency, validating the model's application value in practical engineering. This modified model, combining UV254 absorbance and fluorescence characteristic ratios, provides an innovative solution for rapid biodegradability assessment in wastewater treatment, effectively overcoming the time-consuming nature of traditional BOD / COD measurements and achieving real-time, accurate monitoring of influent organic matter characteristics.
[0143] Example 6
[0144] In this embodiment, the adaptive learning steps based on the deep actor criticism framework include: constructing an expert database based on historical influent water quality characteristic fingerprint data and corresponding measured BOD / COD values, carbon source dosage, and treatment effect data. The expert database contains data sets on water quality status, evaluation results, rewards, and subsequent status. Based on the historical influent water quality characteristic fingerprint data, an actor network is designed. The input of the actor network is the historical influent water quality characteristic fingerprint data, and the output of the actor network is the BOD / COD value and available carbon content to be trained. Based on the historical influent water quality characteristic fingerprint data and the BOD / COD value to be trained, a critic network is designed. The input of the critic network is the historical influent water quality characteristic fingerprint data. The influent water quality fingerprint data and the BOD / COD value to be trained are used, and the output of the critic network is the carbon source dosage. Based on the expert database, a preset number of data groups are randomly selected from the expert database to form batch training data. The actor network and the critic network are used to perform offline policy learning on the batch training data. The parameters of the actor network are updated by calculating the policy gradient and by minimizing the temporal difference error, and the trained actor network model is output. Based on the influent water quality fingerprint data, the influent water quality fingerprint data is input into the trained actor network model, and the predicted BOD / COD ratio and available carbon content are output.
[0145] Specifically, firstly, constructing a high-quality expert database is the foundation of the adaptive learning system. The "expert database" refers to a collection of data on historically optimal operational decisions and their effects, including information such as water quality status, assessment results, operational decisions, and their effects. In this embodiment, 5000 sets of high-quality data were selected from the operational records of Wastewater Treatment Plant A over the past two years (July 2020 to June 2022) to construct the expert database. The data selection criteria included: data completeness (no missing values), representativeness (covering different seasons and influent conditions), and operational effectiveness (achieving effluent total nitrogen standards and high carbon source utilization). Each set of data contains four parts of information: water quality status s (a water quality fingerprint composed of UV254 values, tyrosine-like fluorescence peak intensity, and tryptophan-like fluorescence peak intensity), assessment result a (measured BOD / COD ratio and available carbon content), reward r (a score considering both effluent total nitrogen compliance rate and carbon source utilization efficiency, with a maximum score of 100), and subsequent status s' (the water quality status at the next time point). These four parts of information constitute a complete state-action-reward-nextstate (SARS) transition sample, which is the standard data format for reinforcement learning. To ensure data quality, all data were normalized, mapping each parameter to the [0,1] interval; data balancing was also performed to ensure a relatively balanced number of samples under different water quality conditions, preventing the model from being biased towards specific conditions. Furthermore, to enhance data diversity, data augmentation techniques were employed, such as adding slight random noise and performing time-series shifting, expanding the original 5000 sets of data to approximately 8000 sets. Finally, the complete expert database was stored in a distributed database system for easy batch retrieval and updates.
[0146] Secondly, designing a neural network structure suitable for water quality assessment tasks is the core of the system's adaptive learning. This embodiment designs an Actor-Critic network architecture based on the Deep Deterministic Policy Gradient (DDPG) algorithm. The Actor network is responsible for policy generation, i.e., predicting the BOD / COD ratio and available carbon content based on water quality status; the Critic network is responsible for policy evaluation, calculating the value function under a given state and action. Specifically, the Actor network adopts a 4-layer fully connected neural network structure: the input layer receives 3D water quality feature fingerprint data (UV254 value, tyrosine-like fluorescence peak intensity, and tryptophan-like fluorescence peak intensity); hidden layer 1 contains 64 neurons, using the ReLU activation function; hidden layer 2 contains 32 neurons, also using the ReLU activation function; the output layer contains 2 neurons, corresponding to the predicted BOD / COD ratio and available carbon content, respectively. The output is restricted to the range [0,1] using the Sigmoid activation function, and then mapped to the actual value range through inverse normalization. The critic network structure is similar to the actor network, but with two key differences: First, the input layer receives 5-dimensional data, including 3-dimensional water quality features and 2-dimensional actor network output (i.e., BOD / COD ratio and predicted effective carbon content); second, the output layer has only one neuron, outputting a Q-value representing the expected effect score of the current state-action pair, using a linear activation function. To improve the network's generalization ability, a Dropout layer (dropout rate 0.2) and a batch normalization layer are added after each hidden layer. Furthermore, L2 regularization (regularization coefficient 0.001) is introduced to prevent overfitting. The network parameters are initialized using the He initialization method, which helps the training of deep networks converge.
[0147] Then, based on the constructed expert database and the designed network structure, implementing offline policy learning is a key process for the system to acquire knowledge. Offline policy learning refers to a method where the system does not directly interact with the environment but learns policies from collected historical data. It is suitable for scenarios where real-time control systems cannot frequently conduct trial and error. In this embodiment, offline policy learning based on the DDPG algorithm is implemented. The specific process includes: first, setting learning parameters, including learning rate (0.0001 for actor network, 0.001 for critic network), batch size (64), discount factor γ (0.99), and target network soft update coefficient τ (0.001), etc. Then, initializing the actor network, critic network, and target network (network copies used to stabilize the learning process). Next, the iterative training phase begins. In each iteration, 64 sets of data are randomly selected from the expert database to form a batch, and the network parameters are updated through backpropagation. The training process focused on addressing two key issues in offline reinforcement learning: distribution shift, which was mitigated by introducing policy constraints to limit policy changes and ensure the model doesn't deviate excessively from expert policies; and overestimation bias, which was reduced through double-Q learning and delayed updates to the target network to minimize optimism in value estimation. Loss function values, policy entropy, and validation performance were recorded during training. An early stopping mechanism was triggered when validation performance showed no significant improvement after 10 consecutive iterations. After approximately 2000 iterations, the model finally converged, with the actor network loss function dropping below 0.005 and the critic network loss function dropping below 0.01. The training time was approximately 4 hours (using a server equipped with a specific GPU model).
[0148] Finally, the trained model was applied to real-time water quality assessment and its performance was evaluated. During the application phase, the system received new water quality characteristic fingerprint data (UV254 value and fluorescence peak intensity) every 10 minutes, inputting it into the trained actor network model. The model then outputs the predicted BOD / COD ratio and available carbon content in real time, providing a basis for subsequent carbon source dosing control decisions. To evaluate model performance, the prediction results for 30 consecutive days were compared with laboratory measurements, and the mean relative error and coefficient of determination R0 were calculated. 2 Indices such as consistency index were used. Results showed that the trained model had an average relative error of 5.8% in BOD / COD prediction on the test set, and R... 2 The R-squared value reached 0.91, and the consistency index reached 0.94, which is significantly better than the traditional statistical model (mean relative error 8.5%, R-squared value 0.94). 2(The value is 0.86). Especially when dealing with extreme water quality conditions (such as industrial wastewater impact or early rainy season scouring), the deep actor critique model shows stronger robustness and adaptability. In addition, a continuous optimization mechanism for the model is implemented: new operational data is collected weekly, high-performing samples (operational decisions with good treatment effects) are added to the expert database, and the model is retrained regularly (monthly) to ensure that the system can continuously adapt to changes in water quality characteristics and continuously improve the evaluation accuracy.
[0149] In the actual operation of Wastewater Treatment Plant A, the adaptive learning system based on the deep actor criticism framework significantly improved the accuracy and stability of water quality assessment. In particular, it demonstrated adaptability not found in traditional models when dealing with complex situations such as seasonal changes and industrial wastewater impacts.
[0150] Example 7
[0151] In this embodiment, the offline policy learning using the actor network and the critic network on the batch training data includes: based on the data sets in the batch training data, inputting the water quality status of the data sets into the actor network, outputting the current policy action, where the current policy action is the BOD / COD value and available carbon content to be trained; inputting the water quality status and the current policy action into the critic network, outputting the Q value of the current state-action pair; calculating the policy gradient based on the Q value and updating the parameters of the actor network to generate a parameter-updated actor network; calculating the temporal difference error target value according to the reward and subsequent states of the data sets, and updating the parameters of the critic network. To generate a parameter-updated critic network, the method minimizes the squared difference between the Q-value output by the critic network and the target value of the temporal difference error. Based on the parameter-updated actor network, the water quality state is input into the parameter-updated actor network, and the updated policy action is output. The water quality state and the updated policy action are input into the parameter-updated critic network, and the updated Q-value is output. Based on the updated Q-value, a policy constraint term is added to the policy gradient calculation to ensure that the updated policy does not deviate from the expert policy. The policy constraint term is the expected value of the squared difference between the updated policy action and the expert action in the data set. This generates the trained actor network model.
[0152] Specifically, the first step in offline learning is to evaluate the current policy of the agent network based on batch training data. In this embodiment, each training round randomly selects 64 sets of data from the expert database to form a batch. Each set of data includes water quality state s, expert action aE (measured BOD / COD value and available carbon content), reward r (treatment effect score), and subsequent state s'. The water quality state s in the batch is input into the current agent network to obtain the action a (i.e., the predicted values of BOD / COD value and available carbon content) under the current policy. This process is represented as follows: ,in Represents the policy function. The parameters of the current actor network are represented. Then, the state-action pair (s,a) is input into the critic network, and the output is the Q(s,a) value. This value represents the expected long-term cumulative reward under the current strategy and is a quantitative representation of the expected effect of the evaluation result. In the actual implementation, the water quality state s is represented by a 3D vector, including the normalized UV254 value, tyrosine-like fluorescence peak intensity, and tryptophan-like fluorescence peak intensity; action a is represented by a 2D vector, including the normalized BOD / COD prediction value and the effective carbon content prediction value. The Q value is calculated in batches, processing the entire batch of data at once to improve computational efficiency. Furthermore, to enhance the stability of the evaluation, an integrated evaluation mechanism is introduced, maintaining multiple critic networks with different initialization parameters, and taking the average of their Q values as the final evaluation result, effectively reducing the variance and uncertainty of a single network evaluation.
[0153] Secondly, updating the critic network parameters based on the calculated Q-value is a crucial step in improving evaluation accuracy. In standard reinforcement learning algorithms, critic network parameter updates typically employ Temporal Difference (TD) learning, aiming to minimize the mean squared error between the predicted Q-value and the TD target value. In the offline learning environment of this embodiment, the temporal difference target value is calculated as follows: ,in For instant rewards, The discount factor is set to 0.99. This indicates that the target policy will be executed in the subsequent state s'. The expected Q-value (generated by the target actor network) (evaluated by the target critic network). Parameters of the target network. Obtained through soft updates of the original network parameters: ,in The soft update coefficient (set to 0.001) helps stabilize the learning process. Critics network parameters. The update uses stochastic gradient descent, with the goal of minimizing the loss function. Where N is the batch size. In the actual implementation, the Adam optimizer is used with a learning rate of 0.001, and L2 regularization (coefficient 0.0001) is introduced to prevent overfitting. To address the overestimation problem in offline learning, a double-Q learning technique is implemented, maintaining two critic networks and using their minimum output as the basis for calculating the target value, effectively reducing the optimism bias of the value function. Furthermore, a prioritized experience replay technique is implemented, adjusting the sample sampling probability according to the magnitude of the TD error, making training more focused on difficult samples and accelerating the learning process.
[0154] Then, updating the actor network parameters based on the updated critic network is a core step in optimizing the strategy. In the DDPG algorithm, the actor network parameters... The update objective is to maximize the expected Q-value, which is achieved through the policy gradient method. The policy gradient calculation formula is as follows: ,in This represents the gradient of the Q-value with respect to action a. This indicates the policy function with respect to the parameters. The gradient is calculated using the Adam optimizer. Intuitively, this formula means adjusting policy parameters along the action direction that increases the Q-value, so that the actions generated by the policy achieve higher evaluation scores. In practice, the gradient is calculated using backpropagation, and the parameters are updated using the Adam optimizer with a learning rate of 0.0001. To enhance learning stability, gradient pruning is introduced, limiting the gradient norm to no more than 1 to prevent excessive parameter updates from causing training instability. Furthermore, a parameter noise exploration mechanism is implemented. By adding noise to the parameter space rather than the action space, a more structured and consistent exploration is achieved, which helps discover better policies. The standard deviation of the parameter noise is initially set to 0.1 and dynamically adjusted according to training progress, being larger in the early stages to promote exploration and gradually decreasing in the later stages to promote utilization. In this way, the actor network can gradually learn better policies, generating more accurate BOD / COD predictions and effective carbon content estimates.
[0155] Finally, to ensure the stability of offline learning, adding a policy constraint term in the policy gradient calculation is a crucial step. One of the main challenges of offline reinforcement learning is the distribution shift problem, where the action distribution generated by the learned policy may differ from the action distribution of the data collection policy (expert policy), leading to inaccurate value function estimation. To address this issue, this embodiment introduces a policy constraint mechanism, limiting the deviation between the learned policy and the expert policy by adding a constraint term to the policy gradient. Specifically, the policy constraint term is defined as... ,in This is the weighting coefficient (set to 0.01). The action generated for the current policy, This is the expert action (measured BOD / COD value). This constraint penalizes the deviation between the policy-generated action and the expert action, making the policy more closely resemble expert experience. The final actor network update objective becomes:
[0156] ,
[0157] That is, to maximize the Q-value while minimizing the deviation from the expert's actions. In practical implementation, the constraint weights are dynamically adjusted for different types of water quality conditions. For water quality conditions that appear frequently in expert data, use a smaller [size / weight] This value allows for more strategy optimization; for rare water quality conditions, a larger value can be used. The forced strategy more closely resembles expert behavior. Furthermore, a progressive constraint relaxation mechanism is implemented, gradually reducing the constraint as training progresses. The value was initially set to 0.05 and eventually reduced to 0.01, allowing for greater optimization potential while maintaining policy stability. This constrained optimization method enables the system to learn safely and effectively on offline data, avoiding performance crashes caused by policy deviations while preserving the ability to improve the policy.
[0158] In practical application at Wastewater Treatment Plant A, the offline policy learning algorithm successfully converged after 2000 iterations. During training, the prediction error of the actor network gradually decreased from an initial 15% to 5.8%, and the TD error of the critic network decreased from an initial 0.23 to 0.008, demonstrating the effectiveness of the learning process. To verify the learning results, during a 30-day testing period, the BOD / COD values predicted by the model were compared with laboratory measurements. The results showed that 92% of the predicted values deviated within ±10%, 78% deviated within ±5%, and the average relative error was 5.8%, significantly better than the traditional statistical model and the initially trained neural network model. Furthermore, by applying the prediction results to carbon source dosing control, the system reduced carbon source usage while maintaining effluent quality standards, demonstrating the important value of accurate BOD / COD assessment for optimal resource utilization. Through offline policy learning, the system not only gained more accurate assessment capabilities but also achieved the ability to extract knowledge from expert experience and continuously optimize, providing strong support for intelligent control of the wastewater treatment process. The successful implementation of this learning algorithm demonstrates the application potential of reinforcement learning technology in the optimization of traditional industrial processes and opens up new paths for the application of intelligent technologies in the environmental protection field.
[0159] Example 8
[0160] In this embodiment, the personalized computation steps based on multimodal dynamic agent learning include: collecting and integrating multimodal data based on the influent water quality characteristic fingerprint data, the predicted BOD / COD ratio and available carbon content, influent flow rate, pH value, temperature, total nitrogen concentration, ammonia nitrogen concentration, historical carbon source dosage, and treatment effect data; standardizing and denoising the multimodal data to generate preprocessed multimodal data; and using a gated cross-modal fusion network to extract features and fuse data from different sources based on the preprocessed multimodal data to generate fused feature representations. Based on the fusion feature representation, influent types are clustered using a dual-constraint proxy optimization mechanism and a dynamic candidate management mechanism. The current influent sample is assigned to the identified influent type, and the membership degree to each influent type is calculated. The clustering results and membership degrees are output. Based on the clustering results, a dedicated carbon source demand calculation model is constructed for each identified influent type. The dedicated carbon source demand calculation model maps the fusion feature representation to the carbon source dosage. Based on the membership degree and the dedicated carbon source demand calculation model, the carbon source dosage output by the dedicated carbon source demand calculation model for each influent type is weighted and synthesized to output the external carbon source demand.
[0161] Specifically, firstly, the collection and preprocessing of multimodal data forms the data foundation for personalized computing. "Multimodal data" refers to data from different sources with varying characteristics and representations; in this invention, it specifically refers to diverse heterogeneous data describing wastewater characteristics and treatment processes. In this embodiment, Wastewater Treatment Plant B constructed a comprehensive multimodal dataset, including three main categories: spectral data (including UV254 absorbance values, three-dimensional fluorescence spectral EEM features, etc.), conventional water quality parameters (including influent flow rate, pH value, temperature, total nitrogen concentration, ammonia nitrogen concentration, etc.), and operational parameters (including historical carbon source dosage, effluent water quality, treatment efficiency, etc.). During data acquisition, spectral data was collected every 10 minutes using an online spectral analyzer; conventional water quality parameters were collected every 5-15 minutes using various online sensors, with varying collection frequencies for different parameters; operational parameters were recorded using a combination of automated system data and manual data recording, with time granularities ranging from minutes to hours. To address the asynchronicity issue of different data sources, the system employs a time window alignment technique, using a 30-minute time window to map all data onto a unified time grid based on timestamps. Representative values are calculated from multiple samples within the window using the mean. For the collected raw multimodal data, missing value processing is first performed using multiple imputation to fill in a small number of missing data points. This method generates multiple possible imputed values to reflect the uncertainty of missing values, outperforming simple mean or linear interpolation. Next, z-score standardization is performed, converting parameters of different dimensions into standardized variables with a mean of 0 and a standard deviation of 1. The formula is z = (x - μ) / σ, where x is the original value, μ is the mean, and σ is the standard deviation. After standardization, wavelet transform denoising is applied, a powerful denoising technique suitable for processing non-stationary signals. Specifically, Daubechies wavelet (db4) is used for multi-scale decomposition, typically to level 4. Then, a soft thresholding method is applied to shrink the wavelet coefficients, with the threshold set to λ = σ. Where σ is the noise standard deviation estimate and N is the signal length. This method can effectively remove random noise while preserving the abrupt changes and trend information of the data. After preprocessing, the signal-to-noise ratio of various types of data is significantly improved, laying the foundation for subsequent modality fusion and feature extraction.
[0162] Secondly, feature extraction and fusion of preprocessed multimodal data using a gated cross-modal fusion network is a key step in achieving effective information integration. A gated cross-modal fusion network is a specially designed neural network structure capable of adaptively learning the correlations between different modalities and selectively fusing the most relevant information. In its implementation at Wastewater Treatment Plant B, a network structure with three parallel branches was designed to process the three modalities respectively. Each branch first extracts the latent feature representation of that modality through a modality-specific encoder. The spectral data encoder employs a one-dimensional convolutional neural network (1D-CNN) structure, containing three convolutional layers (kernel sizes of 5, 3, and 3, and filter counts of 32, 64, and 128, respectively) and two pooling layers, effectively capturing local features and global patterns in spectral data. The conventional water quality parameter encoder uses a multilayer perceptron structure, containing three fully connected layers (with 64, 96, and 128 neurons, respectively), using Leaky ReLU as the activation function, suitable for processing numerical features. The runtime parameter encoder also uses a multilayer perceptron structure, but considering temporal dependencies, it adds an LSTM (Long Short-Term Memory) layer, enabling it to capture the temporal dynamics of runtime parameters. These three encoders map the raw data from different modalities to a unified 128-dimensional feature space, achieving preliminary modality alignment. Then, a multi-head self-attention mechanism (8 heads) is used to calculate the correlation strength between modalities, forming an attention weight matrix. The core idea of this mechanism is the "Query-Key-Value" computation paradigm. Each modality feature can serve as both a query vector and a key-value pair. Attention weights are obtained by calculating the similarity between the query and the key and performing softmax normalization. Then, the value vectors are weighted and summed. This approach can dynamically capture the correlation between different modalities, such as the special variation patterns of spectral features under high pH values. Finally, a sigmoid function-based gating unit is designed to control the information flow between different modalities. The input of the gating unit is the modality feature and attention weights, and the output is a control coefficient between 0 and 1, determining how much information can pass through. This gating mechanism allows the network to selectively combine modal information under different conditions. For example, when sensor malfunctions or data quality deteriorates, the weights of the corresponding modalities are automatically reduced, enhancing the robustness of the system. The entire network is trained and optimized end-to-end using the cross-entropy loss function and the Adam optimizer on a training set containing three years of historical data. After convergence, the model achieves a classification accuracy of 91.8% on the test set, demonstrating its powerful feature extraction and fusion capabilities.
[0163] Next, clustering influent types based on fused feature representations is the foundation for personalized processing. The "dual-constraint surrogate optimization mechanism" is a clustering optimization method that combines sample distribution and processing effect feedback, while the "dynamic candidate management mechanism" enables the system to adaptively adjust the number of clusters and surrogate vectors. In the implementation at Wastewater Treatment Plant B, six surrogate vectors for each influent type were first initialized based on historical data distribution and expert experience. Each vector is a 128-dimensional feature vector representing a potential influent type feature center. An improved K-means++ algorithm was used for initialization to make the initial surrogate vector distribution more uniform. For each new influent sample, its cosine similarity with each surrogate vector was calculated. The cosine similarity calculation formula is cos(θ) = (A·B) / (||A||·||B||), where A and B are the sample feature vector and the surrogate vector, respectively. Then, an improved fuzzy C-means (FCM) clustering algorithm was used to assign samples to each category, while simultaneously calculating membership values. The standard FCM algorithm minimizes the objective function as follows:
[0164]
[0165] Where u ij Let d(x) represent the membership degree of sample i to category j, m be the fuzziness factor (usually set to 2), and d(x) be the membership degree of sample i to category j. i , c jThe distance between the sample and the class center is denoted as . In this embodiment, two improvements were made to the FCM algorithm: first, sample weights were introduced, assigning different weights based on the representativeness and recentity of the samples, making clustering focus more on important samples; second, a regularization term was added to prevent any one class from becoming overly dominant. The system collects data on the effluent total nitrogen compliance rate and carbon source utilization efficiency weekly as feedback on the processing effect, and evaluates the current clustering effect based on this, forming a processing effect score. A soft update strategy is adopted for surrogate vector updates. The new surrogate vector is the weighted sum of the old surrogate vector and the average vector of samples within the class, with the weight coefficient α (learning rate) set to 0.1 to ensure the smooth evolution of the surrogate vector. The dual-constraint optimization is reflected in: on the one hand, the sample distribution guides the surrogate vector update, making the surrogate vector reflect the natural clustering of the data; on the other hand, the processing effect feedback adjusts the update direction, prompting the surrogate vector to evolve in a direction that is conducive to improving the processing effect. The dynamic candidate management mechanism periodically (monthly) evaluates clustering quality and dynamically adjusts the number of categories based on three indicators: intra-cluster density (average distance between samples within a category), inter-cluster distance (distance between surrogate vectors of different categories), and treatment effect difference (difference in treatment effect scores between different categories). When two surrogate vectors are detected to be too close (distance less than the threshold of 0.2) and have similar treatment strategies, the system automatically merges these two categories. When the intra-cluster sample distribution is too scattered (intra-cluster variance greater than the threshold of 1.5) and the treatment effect is unstable, the system automatically splits the category into two sub-categories. Through this dynamic adjustment mechanism, the influent classification system of Plant B, after three months of operation, dynamically adjusted from the initial 6 categories to 5 categories: high-nitrogen and low-carbon type (mainly industrial wastewater), medium-load balanced type (typical urban domestic sewage), high-carbon and low-nitrogen type (containing a large amount of easily degradable organic matter), high-salt inhibition type (containing high concentration of salt), and complex impact type (large fluctuations in water quality parameters). This automatically identified influent type classification is more accurate than the traditional fixed threshold classification, and can capture subtle differences in influent characteristics, providing a foundation for subsequent personalized treatment.
[0166] Finally, the ultimate goal of this step is to construct a dedicated carbon source demand calculation model based on the clustering results and achieve personalized dosing calculations. One of the innovations of this invention is to construct a dedicated carbon source demand calculation model for each identified influent type, rather than using a single general model. In the implementation at Wastewater Treatment Plant B, dedicated carbon source demand calculation models were constructed for each of the five identified influent types. Each dedicated model uses the same network structure: the input layer receives the fused feature representation (128-dimensional vector), passes through three fully connected layers (with 64, 32, and 16 neurons respectively), and one output layer, outputting the predicted carbon source dosing amount (unit: mg / L or kg / h). Model training employs supervised learning, using the historically optimal carbon source dosing amount for that type of influent (the dosing amount that results in effluent total nitrogen compliance and the highest carbon source utilization rate) as the label. Although the models for different influent types have the same structure, the learned parameters and feature focus points differ significantly. For example, models for high-nitrogen, low-carbon influent focus more on the difference between nitrogen content and available carbon content, while models for high-salinity-inhibiting influent focus more on the relationship between salinity and microbial activity. Each specialized model performed excellently on the test data for its corresponding type, with an average prediction error controlled within 7%, significantly better than the general model (error approximately 12%). For newly entered influent samples, the system calculates their membership degree to each influent type based on the aforementioned clustering results. The membership degree value reflects the degree to which a sample belongs to a certain category, ranging from 0 to 1, with the sum of the membership degree values for all categories being 1. Then, the sample is input into each specialized model to obtain the carbon source addition amount predicted by different models. Finally, a weighted average of the outputs of each model is calculated based on the membership degree value to determine the final external carbon source demand. This soft allocation method is more robust than hard classification (using only the model corresponding to the highest membership degree) and can smoothly handle boundary conditions and mixed-type influents. The calculation formula is: Final carbon source demand = ∑(u i ·M i (x)), where u i M represents the membership degree of a sample to the i-th class. i (x) is the predicted output of the i-th specialized model for input x.
[0167] Example 9
[0168] In this embodiment, the step of extracting features and fusing cross-modal data from different sources through a gated cross-modal fusion network to generate a fused feature representation includes: based on the preprocessed multimodal data, using a modality-specific encoder to encode spectral data, conventional water quality parameter data, and process operation parameter data respectively to generate modality-specific features; calculating the correlation strength between different modalities based on the modality-specific features using an attention mechanism to generate an attention weight matrix; and controlling information flow through a gating unit to perform weighted fusion of different modality features to generate the fused feature representation.
[0169] Specifically, firstly, using modality-specific encoders to encode data of different modalities is fundamental to effectively extracting features from various types of data. A modality-specific encoder is a neural network structure specifically designed to process specific types of data, effectively capturing the feature patterns of that type of data. In the implementation at Wastewater Treatment Plant B, three dedicated encoders were designed to process spectral data, conventional water quality parameters, and process operating parameters, respectively. The spectral encoder (Es) employs a one-dimensional convolutional neural network (1D-CNN) structure, specifically designed to process spectral data with continuous wavelengths or time-series characteristics. This encoder contains three convolutional blocks, each consisting of a convolutional layer, a batch normalization layer, an activation function layer, and a pooling layer. The first convolutional block uses 32 kernels of size 5 with a stride of 1 to capture local spectral features; the second convolutional block uses 64 kernels of size 3 to capture medium-scale features; and the third convolutional block uses 128 kernels of size 3 to extract high-level abstract features. All convolutional layers employ zero-padding to preserve feature map size, and the activation function used is PReLU (parameterized ReLU), which has adaptive learning capabilities. Max pooling is used with a window size of 2, effectively reducing feature dimensionality. After convolution, global average pooling and fully connected layers map the features to a 128-dimensional vector space. This structural design fully considers the continuity and local correlation of spectral data, effectively extracting key patterns from UV254 and fluorescence spectra. The conventional water quality encoder (Eq) uses a multilayer perceptron structure, suitable for processing discrete numerical parameters. This encoder contains three fully connected layers with 64, 96, and 128 neurons respectively, each followed by batch normalization and a Leaky ReLU activation function (with a negative half-axis slope of 0.2). To enhance robustness, a Dropout layer with a dropout rate of 0.3 is added after the first and second fully connected layers to effectively prevent overfitting. Furthermore, considering the potential nonlinear relationships between conventional water quality parameters, a self-attention module is added before the last layer, enabling the model to adaptively learn the interactions between parameters. The runtime parameter encoder (Ep) is also based on a multilayer perceptron, but it specifically considers the temporal dependence of the parameters. This encoder first processes 24 hours of continuous historical data through an LSTM (Long Short-Term Memory) layer with 64 LSTM units. The output sequence is then converted into a 96-dimensional vector via a fully connected layer, and then concatenated with the parameters at the current time step. Both vectors are then fed into two fully connected layers (96 and 128 neurons respectively), ultimately mapping to a 128-dimensional feature space. This structural design allows the encoder to simultaneously consider both current runtime parameters and historical trends, providing more comprehensive information. The three encoders are optimized for the data types they process. After individual pre-training, they can transform the raw data of their respective modalities into semantically rich feature representations, laying the foundation for subsequent cross-modal fusion.
[0170] Secondly, calculating the correlation strength between different modalities through an attention mechanism is key to achieving effective cross-modal fusion. The "multi-head attention mechanism" is a technique that can simultaneously learn correlation patterns from multiple representation subspaces. Derived from the Transformer model, it performs exceptionally well in cross-modal learning. In its implementation at Wastewater Treatment Plant B, an 8-head self-attention mechanism was designed to calculate the correlation strength between modalities based on the modality-specific features output by three encoders. Specifically, three transformations—Query, Key, and Value—are defined to map each modal feature to Q, K, and V representations, respectively. For the features Ei of modality i and Ej of modality j, a linear transformation is used to obtain...
[0171] Qi = Ei·WQ, Kj = Ej·WK, Vj = Ej·WV,
[0172] Where WQ, WK, and WV are learnable transformation matrices with a dimension of 128×16 (single-head size). Then, attention weights are obtained by calculating the dot product of the query and the key, followed by scaling and softmax normalization.
[0173] A(i,j) = softmax(Qi·Kj Q / ),
[0174] Where d=16 is the dimension of the attention head, a scaling factor. Used to stabilize gradients. Finally, the attention weights are multiplied by the values to obtain the weighted result: O(i,j) = A(i,j)·Vj. This process is computed in parallel in 8 independent attention heads, each focusing on a different feature subspace. The outputs of the 8 heads are then concatenated and mapped back to the original dimension through a linear transformation to obtain the final attention output. This multi-head design enhances the model's expressive power, enabling it to simultaneously capture complex correlation patterns between multiple modalities. In practical applications, the attention weight matrix intuitively reflects the degree of attention between different modalities. For example, the system learns to focus more on fluorescence features under high-temperature conditions and more on conventional water quality parameters when pH is abnormal. To further improve the effectiveness of the attention mechanism, positional encoding and residual connections are also implemented: positional encoding adds relative positional information to the features, helping the model understand the sequential relationships in the parameter sequence; residual connections retain the original feature information through skip connections, alleviating the gradient vanishing problem in deep networks. Through the attention mechanism, the interaction between different modal data can be adaptively learned, and the importance of each modality can be dynamically adjusted according to the current situation, improving the expressive and discriminative power of the fused features.
[0175] Finally, controlling information flow through gating units and performing weighted fusion of different modal features is a crucial step in ensuring fusion quality. A gating unit is a mechanism for controlling information flow, similar to the gating structure in LSTM, selectively allowing information to pass through while filtering out irrelevant or redundant information. In the implementation at Wastewater Treatment Plant B, a sigmoid function-based gating unit was designed to control information exchange between different modes. For the information flow between modes i and j, the gating value is calculated as follows:
[0176] G(i,j)=σ(W g ·[Ei; A(i,j);Ej] + b g ),
[0177] Where [;] denotes vector concatenation, W g and b g σ is a learnable parameter, and σ is the sigmoid activation function that maps the output to the range of 0-1. A gating value close to 1 allows information to pass through completely, while a value close to 0 blocks information flow. This design enables the system to adaptively select useful modal combinations under different conditions. For example, when the data quality of a certain sensor deteriorates, the corresponding modal information can be partially or completely blocked, reducing its impact on the fusion result and enhancing the robustness of the system. The final fused feature representation is calculated as follows:
[0178] F=∑ i,j G(i,j)·A(i,j)·Transform(Ei, Ej),
[0179] The `Transform` method is a linear transformation used to adjust the feature dimensions. This fusion approach comprehensively considers intra-modal features, inter-modal attention, and gating control, generating a unified feature representation that incorporates multi-source information. To further enhance the fusion effect, two technical improvements were implemented: feature calibration and contrastive learning. Feature calibration dynamically adjusts features from different modalities using feature statistics (mean and variance) to mitigate distributional differences between modalities; contrastive learning enhances the discriminative power of the fused features by maximizing the consistency of different representations of the same sample and minimizing the similarity of different sample representations. These techniques work together to enable the gated cross-modal fusion network to extract and integrate the most valuable information from heterogeneous data, providing high-quality feature representations for subsequent water inflow classification and decision-making.
[0180] In the practical application at Wastewater Treatment Plant B, the gated cross-modal fusion network was trained end-to-end using the Adam optimizer with an initial learning rate of 0.001. A cosine annealing scheduling strategy was employed to dynamically adjust the learning rate. The training data included three years of historical records, approximately 100,000 samples, divided into training, validation, and test sets in an 8:1:1 ratio. To address data imbalance, a weighted sampling strategy was used to increase the sampling probability of rare class samples. An early stopping strategy was employed during training; training was halted when the validation set performance showed no improvement for five consecutive cycles to prevent overfitting. Ultimately, the model achieved a classification accuracy of 91.8% and an F1 score of 0.897 on the test set, demonstrating its powerful feature extraction and fusion capabilities. Visual analysis revealed that the fused features exhibited a clear clustering structure after t-SNE dimensionality reduction, with samples from different influent types forming distinct clusters, validating the high discriminative power of the fused features. Furthermore, ablation experiments demonstrated the significant contributions of gating mechanisms and multi-head attention to system performance: removing the gating mechanism resulted in a 4.2 percentage point decrease in accuracy, while removing multi-head attention led to a 3.7 percentage point decrease. After deployment, through continuous incremental learning using newly collected data, the model performance gradually improved, reaching a test accuracy of 93.5% after three months. This adaptive cross-modal fusion method not only improved the accuracy of carbon source dosing at Plant B but also provided a new approach for processing multi-source heterogeneous data in the water treatment field.
[0181] Example 10
[0182] In this embodiment, the influent type clustering through a dual-constraint proxy optimization mechanism and a dynamic candidate management mechanism, assigning the current influent sample to the identified influent type and calculating the membership degree to each influent type, and outputting the clustering result and membership degree, includes: initializing multiple influent type proxy vectors according to the fused feature representation; calculating the similarity between the current influent sample and each of the influent type proxy vectors based on the fused feature representation, assigning the current influent sample to each influent type using a fuzzy clustering method and calculating the membership degree to generate a preliminary clustering result; obtaining the effluent total nitrogen compliance rate and carbon source utilization efficiency as treatment effect feedback data; periodically updating the influent type proxy vectors based on the preliminary clustering result and the treatment effect feedback data; dynamically adjusting the number of categories according to the intra-class sample density and inter-class distance; and outputting the clustering result and membership degree.
[0183] Specifically, the initialization of multiple influent type surrogate vectors is the starting point of the clustering process. An influent type surrogate vector is a centroid vector representing a specific influent characteristic in the feature space, and is a core concept in cluster analysis. In the implementation at Wastewater Treatment Plant B, based on historical data and expert experience, six influent type surrogate vectors were initialized. Each surrogate vector is a 128-dimensional feature vector, corresponding to a point in the fused feature space. The initialization employed an improved K-means++ algorithm, which selects initial centroids through a distance-weighted probability distribution, ensuring a more uniform distribution of the initial surrogate vectors and avoiding the local optima problem that may occur with traditional K-means random initialization. Specifically, a sample point is first randomly selected as the first surrogate vector. Then, the shortest distance from all samples to the selected surrogate vector is calculated, and probability sampling is performed using the square of the distance as a weight to select the next surrogate vector. This process is repeated until all six surrogate vectors have been selected. This initialization method ensures that the initial surrogate vectors are distributed as evenly as possible throughout the sample space, providing a good starting point for subsequent clustering. To further optimize the initial surrogate vectors, expert knowledge was incorporated to fine-tune them, making them closer to the characteristics of typical influent types identified by experts. For example, based on engineers' experience, the surrogate vector representing high-nitrogen, low-carbon influent was adjusted, resulting in higher values for nitrogen-related dimensions and lower values for carbon-related dimensions. This initialization method, combining data-driven approaches and expert knowledge, preserves the natural distribution characteristics of the data while incorporating prior knowledge from domain experts, making the initial clustering more reasonable. After initialization, the six surrogate vectors represent different potential influent types, providing a foundation for subsequent dynamic optimization.
[0184] Secondly, calculating the similarity between the current influent sample and each surrogate vector based on the fused feature representation, and then assigning samples using fuzzy clustering, is the core step in implementing soft clustering. Fuzzy clustering differs from traditional hard clustering methods (such as K-means) in that it does not absolutely assign samples to a particular category, but rather calculates the membership degree of a sample to each category, making it more suitable for handling situations where category boundaries are ambiguous or samples have mixed characteristics. In the implementation at Wastewater Treatment Plant B, for each new influent sample, its cosine similarity with each surrogate vector is first calculated. Cosine similarity is a commonly used similarity measure in vector space, and its calculation formula is:
[0185] ,
[0186] Where A and B are the sample feature vector and the surrogate vector, respectively, and · denotes the vector dot product. The vector norm is represented. The cosine similarity value ranges from [-1, 1], with a larger value indicating that the two vectors are closer in direction, meaning the sample has a higher similarity to that type. Then, an improved fuzzy C-means (FCM) clustering algorithm is used to calculate the membership degree of samples to each category. Standard FCM minimizes the weighted sum of squared distances by iteratively optimizing the membership matrix and class centers. In this embodiment, the FCM algorithm has three important improvements: first, it uses cosine distance instead of Euclidean distance, which is more suitable for high-dimensional feature spaces; second, it introduces sample weights, assigning different weights based on the representativeness and recentity of the samples, making clustering focus more on important samples; and third, it adds a regularization term to prevent one category from becoming overly dominant. The improved membership degree calculation formula is:
[0187] ,
[0188] in Let be the membership degree of sample i to category j. Let m be the cosine similarity between the sample and the class center, and m be the fuzzy factor (set to 2.2 in this embodiment). The calculated membership values range from [0,1], and for each sample, the sum of the membership values of all classes is 1. This soft assignment method reflects the complexity of the influent better than hard clustering. For example, a sample may simultaneously possess the characteristics of "high nitrogen and low carbon" and "high salt suppression," and this mixed attribute can be quantified through membership degree. To improve the clustering quality, the system also implements anomaly detection based on membership degree. When the maximum membership degree of a sample to all classes is lower than the threshold (set to 0.4), it is marked as a potential new type, triggering a manual review process to promptly identify possible new influent types.
[0189] Third, regularly updating the surrogate vector and dynamically adjusting the number of categories based on processing effect feedback data is a key step in maintaining clustering effectiveness. "Dual-constraint surrogate optimization" refers to optimizing the surrogate vector while simultaneously considering data distribution characteristics and processing effect feedback, while "dynamic candidate management" is a mechanism for dynamically adjusting the number of categories based on clustering quality indicators. In the implementation at Wastewater Treatment Plant B, the system collects data on effluent total nitrogen compliance rate and carbon source utilization efficiency weekly as processing effect feedback, and evaluates the current clustering effect based on this. The processing effect score is calculated as follows:
[0190] ,
[0191] The pass rate and utilization efficiency were both normalized to the 0-100 range. Then, the processing results were summarized according to category, and the average score and variance for each category were calculated. The surrogate vector update adopted a dual-constraint soft update strategy, combining the constraints of sample distribution and processing results. The update formula is:
[0192] ,
[0193] in The learning rate is set to 0.1. The sample distribution weights are set to 0.8. For the sample set assigned to the k-th class, For sample weights, For sample features, This is an adjustment vector driven by processing effects. Wherein, Calculated based on the processing effect of samples within the category, the sample pairs with good results are... It contributes more significantly, guiding the surrogate vector to evolve in a direction that improves processing performance. This dual-constraint update mechanism is one of the innovations of this invention, enabling clustering to both reflect the natural distribution of data and serve the practical goal of improving processing performance.
[0194] To achieve dynamic adjustment of the number of categories, a comprehensive evaluation mechanism was designed to periodically (usually monthly) assess clustering quality. Evaluation metrics include: intra-cluster clustering (average distance between samples within a category), inter-cluster separation (distance between surrogate vectors of different categories), processing performance difference (difference in processing performance scores between different categories), and category usage frequency (distribution of sample numbers for each category). Based on these metrics, two dynamic adjustment operations are implemented: merging and splitting. When two surrogate vectors are detected to be too close (cosine similarity greater than 0.8) and have similar processing strategies (carbon source addition strategy difference less than 10%), the two categories are automatically merged. The merging operation includes generating a new surrogate vector (a weighted average of the two original surrogate vectors), redistributing samples, and updating the dedicated model. When the intra-category sample distribution is too scattered (average intra-category similarity below the threshold of 0.6) and the processing performance is unstable (score variance greater than the threshold of 15), the category is automatically split into two sub-categories. The splitting uses a binary K-means algorithm to re-cluster the samples in the original category, generating two new surrogate vectors, and training a dedicated model for each new category. Through this dynamic adjustment mechanism, the influent classification system of Plant B evolved from an initial six categories to five categories after three months of operation: the two original similar categories ("Medium Load Balance Type I" and "Medium Load Balance Type II" were merged into a single "Medium Load Balance Type"), while the original "Complex Mixed Type" was split into two more refined categories: "High Carbon Low Nitrogen Type" and "Complex Impact Type". This adaptive category management allows the system to dynamically adjust according to the actual data distribution and processing needs, avoiding the limitations of a preset fixed number of categories.
[0195] In the actual operation of Wastewater Treatment Plant B, the dual-constraint proxy optimization and dynamic candidate management mechanism significantly improved the accuracy and adaptability of influent classification. Compared with traditional fixed-category classification methods, this method achieved continuous improvement in cluster quality over a three-month operation period: the average intra-cluster similarity increased from 0.71 to 0.83, and the average inter-cluster distance increased from 0.42 to 0.58, indicating a more compact and separated cluster structure. More importantly, the treatment effect based on optimized clustering was significantly improved, with the average treatment effect score for each category increasing from 78.5 to 89.3, and the standard deviation decreasing from 16.2 to 9.7, indicating that the treatment effect was both improved and more stable. The system can accurately identify and respond to different types of influent. For example, during the maintenance of the industrial area's drainage system, the system successfully identified "high-salt inhibition" influent characteristics and automatically adjusted to a more conservative carbon source addition strategy, avoiding potential decline in treatment effect. In addition, the dynamic category management mechanism enables the system to have self-evolution capabilities, adapting to long-term changes in influent characteristics, such as seasonal changes or water quality changes caused by drainage system renovations. This clustering method is not only used for carbon source dosing control, but also provides decision support for other process controls in wastewater treatment plants, such as aeration control and chemical dosing, further improving overall treatment efficiency. Evaluation based on actual operational data shows that the optimized five-category influent classification is more reasonable than the initial six categories or the traditional four categories, more accurately reflecting the natural distribution characteristics of Plant B's influent and providing a solid foundation for personalized treatment. This dynamic and adaptive clustering method surpasses traditional static classification methods, opening up new technical pathways for intelligent decision-making in wastewater treatment processes.
[0196] Example 11
[0197] In this embodiment, the step of adjusting the demand for external carbon sources in a time series and then outputting the predicted value of the carbon source dosage includes: adjusting the time series allocation of the demand for external carbon sources based on the demand for external carbon sources, taking into account the bioreaction kinetics characteristics and the hydraulic residence time of the system, formulating a phased dosage strategy, and outputting the predicted value of the carbon source dosage.
[0198] Specifically, firstly, analyzing the bioreaction kinetics is the theoretical basis for formulating a timing-based dosing strategy. "Bioreaction kinetics" refers to the rate characteristics of the absorption, utilization, and transformation of substrates (such as carbon sources) by microorganisms in a biological treatment system, directly affecting denitrification efficiency and carbon source utilization. In the implementation at Wastewater Treatment Plant C, extensive experimental research revealed that the denitrification process exhibits a three-stage kinetic pattern of "fast-slow-stable." The first stage is the rapid reaction period (approximately 0-2 hours), during which the denitrification rate is highest, reaching 0.15-0.25 mgNO3-N / gMLSS·h. Microorganisms rapidly absorb and utilize carbon sources, but insufficient carbon supply at this stage will directly lead to a decrease in denitrification efficiency. The second stage is the deceleration period (approximately 2-6 hours), where the denitrification rate gradually decreases to 0.08-0.15 mgNO3-N / gMLSS·h, and the microorganisms' demand for carbon sources is relatively stable. The third stage is the steady-state period (after 6 hours), where the denitrification rate further decreases to 0.05-0.08 mgNO3-N / gMLSS·h. This stage primarily involves the removal of residual nitrogen, and the demand for carbon sources is relatively low. To accurately grasp this kinetic characteristic, Plant C conducted multiple batch experiments, measuring the time-varying curves of the denitrification rate under different carbon-to-nitrogen ratios (C / N ratios). The results show that at the optimal C / N ratio (approximately 5.5-6.5), the denitrification rate curve exhibits a clear exponential decay characteristic, which can be described by a modified first-order kinetic equation.
[0199] r(t) = r0·e (-k·t) + r min ,
[0200] Where r(t) is the denitrification rate at time t, r0 is the initial denitrification rate, and k is the decay coefficient (determined to be approximately 0.3 h in experiments at plant C). -1 ), r min The minimum denitrification rate was determined. Based on this kinetic model, the distribution of microbial carbon source demand at different stages of the reaction process can be calculated, providing a theoretical basis for sequential addition. Of particular note is that the study also found that different types of carbon sources (such as methanol, sodium acetate, and industrial waste molasses) have different kinetic parameters. The modified starch waste liquid used by Plant C as a carbon source had its kinetic parameters specifically calibrated to ensure the accuracy of the model. Furthermore, temperature has a significant impact on denitrification kinetics. Plant C established a temperature correction coefficient θ' = 1.08 to adjust the kinetic parameters under different temperature conditions, adapting the model to seasonal variations.
[0201] Secondly, analyzing the system's hydraulic characteristics is crucial to ensuring the proper distribution of carbon sources within the reactor. "Hydraulic retention time (HRT)" refers to the average time water remains in the reactor, directly impacting pollutant treatment efficiency and system stability. In the implementation at Wastewater Treatment Plant C, A... 2 The / O (anaerobic-anoxic-aerobic) process has a total HRT of approximately 12 hours, with 2 hours in the anaerobic zone, 4 hours in the anoxic zone, and 6 hours in the aerobic zone. To accurately grasp the hydraulic characteristics of the system, Plant C conducted tracer experiments using lithium chloride as a tracer. After addition at the inlet, the concentration changes at various points in the system were monitored to obtain the actual hydraulic distribution characteristics. The experimental results showed that the hydraulic characteristics of the system deviated from the ideal completely mixed model, exhibiting certain short-circuiting and dead zone characteristics, with the actual effective HRT being approximately 85% of the design value. Simultaneously, computational fluid dynamics (CFD) simulations revealed uneven water flow distribution within the anoxic zone due to factors such as the inlet pipe layout and the location of the mixing equipment. This resulted in higher carbon source utilization near the front end, while denitrification efficiency decreased at the back end due to insufficient carbon sources. Based on these findings, the hydraulic characteristics were quantitatively described, and residence time distribution (RTD) curves for various points in the system were established, providing a spatial reference for the temporal allocation of carbon sources. Furthermore, continuous monitoring of nitrogen concentration changes at different locations revealed that the denitrification load in the first third of the anoxic zone accounted for approximately 50% of the total load, the middle third approximately 30%, and the last third approximately 20%. This load distribution characteristic is also an important factor to consider when formulating the dosing strategy. The system hydraulic characteristic analysis also considered the impact of seasonal variations and flow fluctuations, establishing a dynamic HRT adjustment model based on influent flow rate to ensure a reasonable dosing strategy is maintained under different flow conditions.
[0202] Finally, based on the bioreaction kinetics and system hydraulic characteristics, designing a pre-emptive + balanced sequential dosing scheme is the core of optimizing carbon source utilization. In the implementation at Wastewater Treatment Plant C, an innovative sequential allocation adjustment strategy was designed based on the aforementioned analysis results. The specific implementation plan is as follows: the theoretically calculated total carbon source demand is allocated sequentially according to the principle of pre-emptive + balanced dosing. Specifically, 60% of the total demand is added within the first 2 hours of the system (corresponding to the first 1 / 3 of the anoxic zone), and the remaining 40% is added evenly over the subsequent 4-6 hours (corresponding to the middle and later stages of the anoxic zone and the system circulation flow). The high proportion of addition in the first 2 hours meets the high carbon source demand of microorganisms in the initial rapid reaction stage, avoiding limitations in the denitrification process due to insufficient carbon source; while the balanced dosing in the subsequent periods ensures the continuity of carbon source supply throughout the entire reaction process, preventing excessive waste of carbon source. In practical operation, the 60% dosage for the first two hours is further divided into two periods: 40% in the first hour and 20% in the second hour, forming a decreasing gradient that better reflects the exponential decay characteristic of denitrification. The remaining 40% dosage for the next 4-6 hours is evenly distributed, approximately 6.7-10% of the total dosage per hour. To achieve this precise timing control, the system employs a combination of a programmable logic controller (PLC) and a variable frequency carbon source dosing pump. The pump's operating frequency is automatically adjusted based on the calculated timing scheme, achieving precise control of the carbon source dosage. Notably, the system also incorporates a dynamic adjustment mechanism based on real-time feedback. By monitoring the ORP (oxidation-reduction potential) and nitrate nitrogen concentration at different locations in the anoxic zone online, when poor denitrification is detected in a certain section (e.g., excessively high ORP or abnormal nitrate nitrogen concentration), the system fine-tunes the original timing scheme, increasing the carbon source dosage ratio for the corresponding period to ensure overall denitrification effectiveness.
[0203] This timing adjustment strategy has yielded significant results after implementation at Wastewater Treatment Plant C. Compared to the traditional uniform dosing method (which distributes the total demand evenly throughout the HRT), the forward-leaning + equalization strategy improved carbon source utilization efficiency by approximately 15%, reduced the total carbon source dosage by 12.5%, and simultaneously reduced the effluent total nitrogen concentration from an average of 9.8 mg / L to 8.2 mg / L, increasing the compliance rate from 92% to 98.5%.
[0204] Example 12
[0205] In this embodiment, the step of using the predicted carbon source dosage as a feedforward control signal and combining it with the total nitrogen monitoring data of the effluent as a feedback control signal to calculate the final control quantity through a composite control algorithm includes: calculating the feedforward control output based on the predicted carbon source dosage using a feedforward controller; monitoring the total nitrogen concentration of the effluent in real time, comparing the total nitrogen concentration of the effluent with a preset target value to calculate the deviation, and calculating the feedback control output using a proportional-integral-derivative control algorithm; and generating the final control quantity by weighting the feedforward control output and the feedback control output using set feedforward weighting coefficients and feedback weighting coefficients.
[0206] Specifically, firstly, designing a feedforward controller based on the predicted carbon source dosage is key to leveraging the model's predictive capabilities. "Feedforward control" is a control method that measures control before the disturbance affects the controlled system, based on the measurement and prediction of disturbances, and is characterized by its predictive and rapid response features. In this embodiment, the feedforward controller designed for Wastewater Treatment Plant C takes the influent water quality characteristics and the predicted BOD / COD ratio as input, and outputs the predicted carbon source dosage as the feedforward control output U. ff The core of the feedforward controller is the water quality fingerprint identification and carbon source demand calculation model established based on Examples 1-9. This model has already considered various factors such as influent characteristics, biological reaction requirements, and system characteristics. To further improve the accuracy of feedforward control, Plant C made two key optimizations to the original model output: First, a temperature correction coefficient was introduced to adjust the carbon source demand according to the wastewater temperature. The correction formula is as follows:
[0207] U ff = U ff_base ×θ (T-20) ,
[0208] Where θ is the temperature coefficient (experimentally determined to be 1.06), and T is the current wastewater temperature (°C); secondly, a dynamic response mechanism for changes in influent flow rate is incorporated. When a sudden change in influent flow rate is detected (the rate of change exceeds 20% / hour), the carbon source dosage is adjusted in advance based on the flow rate change trend to cope with the upcoming load change. This prediction-based feedforward control has strong predictive ability, enabling corresponding measures to be taken before water quality changes affect the system, greatly reducing water quality fluctuations caused by delayed response in traditional feedback control. In specific implementation, the feedforward controller updates its output value every 10 minutes, consistent with the data update frequency of the online monitoring equipment, ensuring the timeliness of control. In addition, to enhance the reliability of feedforward control, the system sets reasonable constraints to limit the variation range of the feedforward output and avoid over-adjustment due to prediction errors. Specifically, the rate of change between two adjacent outputs must not exceed ±15%, unless a significant change in influent characteristics is detected (such as industrial wastewater impact). Through this design, the feedforward controller can calculate a reasonable carbon source dosage in advance based on an accurate grasp of influent characteristics, providing a benchmark value for the entire control system.
[0209] Secondly, designing a feedback controller based on effluent total nitrogen monitoring is crucial for ensuring long-term stable operation. Feedback control is a method of control based on the deviation between the system output and the desired output, possessing self-correcting and stability characteristics. In the implementation at Wastewater Treatment Plant C, a feedback controller based on the Proportional-Integral-Derivative (PID) algorithm was designed, using the deviation of the effluent total nitrogen concentration from the target value (10 mg / L) as input to calculate the feedback control output U. fb The mathematical expression for a PID controller is:
[0210] U fb = Kp×e + Ki×∫e·dt + Kd×de / dt,
[0211] Where e represents the deviation (target value minus measured value), and Kp, Ki, and Kd are the proportional, integral, and derivative parameters, respectively. The proper setting of these three parameters is crucial to the control effect. Plant C determined the optimal parameter combination through system identification and parameter optimization methods. First, a pseudo-random binary sequence (PRBS) signal was used to excite the system, collecting input (carbon source dosage) and output (total nitrogen in effluent) data to establish a dynamic system model. Then, based on the established model, the MATLAB optimization toolbox was used to optimize the parameters and initially determine the parameter range. Finally, through on-site closed-loop testing and manual fine-tuning, the final parameters were determined: Kp=0.8, Ki=0.05, Kd=0.1. This set of parameters ensures control stability while possessing sufficient response speed and anti-interference capability. Considering the noise problem of the online total nitrogen monitoring data, a data preprocessing stage was also designed in the feedback loop, using median filtering and exponentially weighted moving average (EWMA) methods to remove outliers and smooth data, enhancing the reliability of the feedback signal. Furthermore, to adapt to system load changes and seasonal effects, an adaptive adjustment function for PID parameters was implemented: increasing the Kp value during high-load periods (such as the rainy season) to accelerate the response; increasing the Ki value during low-temperature seasons to enhance the integral action; and appropriately increasing the Kd value to improve predictability when water quality fluctuates significantly. The feedback controller's update frequency is set to once every 30 minutes, a result that considers both the system response time and the measurement cycle of the online total nitrogen analyzer, ensuring both timely control and avoiding overly frequent adjustments.
[0212] Third, designing a fusion strategy for feedforward and feedback signals is the core of realizing the advantages of composite control. "Composite control" refers to the organic combination of multiple control strategies to leverage their respective advantages and compensate for the shortcomings of a single control method; it is an important application of modern control theory. In the implementation at Wastewater Treatment Plant C, a weighted fusion method was used to integrate the feedforward control output U... ff and feedback control output U fb Combined into the final control quantity U.
[0213] Based on long-term operational experience and system performance analysis, Plant C determined the optimal weight configuration for different operating conditions. During the stable operation phase, α=0.7 and β=0.3 are set, favoring model-based predictive feedforward control, which helps maintain control stability and predictability. When a sudden change in influent water quality is detected (e.g., UV254 value change rate exceeds 30% / hour or a new influent type is detected), the system automatically switches to "fast response mode," adjusting the weights to α=0.8 and β=0.2 to further enhance the role of feedforward control and improve the system's response speed to changes in influent. Conversely, when a decline in system performance is detected (e.g., total nitrogen in effluent exceeds or approaches the limit for three consecutive days) or the accuracy of the feedforward prediction model decreases (the deviation between predicted and measured values continues to exceed 15%), the system switches to "correction priority mode," adjusting the weights to α=0.5 and β=0.5 to enhance the feedback correction effect and ensure system stability. This dynamic weight adjustment mechanism enables the composite control system to adaptively adjust the control strategy according to different situations, balancing response speed and control stability.
[0214] In addition, to further enhance the performance of the composite control system, Plant C implemented three innovative designs: First, a "disturbance suppressor" was introduced. When an external disturbance that may affect system stability (such as a rainstorm or equipment failure) is detected, the system temporarily increases the α value and decreases the control gain to reduce control fluctuations. Second, a "model credibility evaluation module" was designed. By continuously comparing predicted and measured values, the model's credibility index is calculated, and α is dynamically adjusted accordingly. f The system increases the feedforward control weights when the model's reliability is high, and increases the feedback control weights when the model's reliability is low. Finally, it implements a "smooth transition mechanism" so that when the weight coefficients need to be adjusted, the system adopts a gradual change rather than an abrupt change to avoid sudden changes in the control quantity from impacting the system.
[0215] In the actual operation of Wastewater Treatment Plant C, the feedforward-feedback composite control system performed exceptionally well. Compared to using feedback control alone, the composite control reduced the standard deviation of the effluent total nitrogen concentration by 42%, significantly improving control accuracy. Compared to using feedforward control alone, the composite control eliminated long-term cumulative deviations, ensuring long-term stable system operation. Especially when dealing with sudden changes in influent water quality, the response time of the composite control system was shortened by approximately 65% compared to traditional PID control, greatly reducing the risk of exceeding standards. After one year of continuous operation and evaluation, the control system reduced the average effluent total nitrogen concentration of Plant C from 9.1 mg / L to 7.8 mg / L, stably meeting the Class A discharge standard. Simultaneously, the accuracy of carbon source dosing improved, and the total usage decreased by approximately 14%, achieving a win-win situation for both environmental and economic benefits. This composite control method, combining feedforward predictive and feedback corrective capabilities, provides a new technical path for the intelligent control of wastewater treatment plants and has broad application prospects.
[0216] Example 13
[0217] In this embodiment, after converting the final control quantity into an operation instruction for the carbon source dosing device and executing it, the method further includes: based on the final control quantity, converting the final control quantity into an operation instruction for controlling the flow rate or switching frequency of the carbon source pump, and sending it to the carbon source dosing device; according to the execution process of the carbon source dosing device, recording the dosing flow rate, dosing time, and equipment status parameters, and generating process data records; based on the process data records and the effluent total nitrogen monitoring data, calculating the deviation between the predicted BOD / COD ratio and the measured value as a prediction accuracy index and the effluent total nitrogen compliance rate as a control effect index; when the prediction accuracy index or the control effect index is lower than a preset threshold, triggering the update of the control parameters in the UV254-BOD / COD prediction model, the fluorescence feature ratio correction model, and the composite control algorithm, and outputting system optimization suggestions.
[0218] Specifically, the key step in translating theoretical control into practical execution is converting the calculated final control quantity into executable operating instructions for the equipment. "Operating instructions" refer to specific commands that directly control the actions of on-site equipment, such as pump start / stop, frequency setting, and valve opening. In the implementation at Wastewater Treatment Plant D, the carbon source dosing system uses a variable frequency speed-regulating pump as the actuator, requiring the conversion of the abstract control quantity U (unit: kg / h or mg / L) into a specific pump frequency setting value (unit: Hz). The conversion process considers multiple on-site factors, including the pump's performance curve, piping system characteristics, and the density and concentration of the carbon source liquid. The conversion formula is: f = f min + (f max -f min ) × (U - U min ) / (U max - U min ), where f is the pump frequency setpoint, f min and f max These represent the pump's minimum and maximum operating frequencies (set to 10Hz and 50Hz by factory D), and U is the current control variable. min and U maxThese are the lower and upper limits of the control quantity (determined according to the process design of Plant D). To ensure the accuracy of the conversion, the system periodically calibrates the pump flow rate, establishes a frequency-flow calibration curve, and dynamically updates the conversion parameters based on the calibration results. Furthermore, considering potential fluctuations in carbon source concentration, an automatic compensation mechanism is designed: by periodically measuring the effective concentration of the carbon source solution (such as COD or TOC values), when a concentration change exceeding 5% is detected, the conversion coefficient is automatically adjusted to ensure that the actual effective carbon dosage is consistent with the control requirements. Operation commands are sent via industrial Ethernet communication. The control system (central PLC or DCS) sends the frequency setpoint to the frequency converter via the Modbus TCP / IP protocol. After receiving the command, the frequency converter controls the pump motor speed to achieve precise adjustment of the carbon source flow rate. To enhance system reliability, the communication system employs a redundant design and includes an automatic communication fault detection and emergency handling mechanism: when a communication interruption is detected for more than 30 seconds, the system automatically switches to local control mode, maintaining the last valid dosage setting, and simultaneously triggers an alarm to prompt operator intervention. Furthermore, an N+1 redundancy configuration was implemented to address equipment failures, meaning a backup pump is provided. When the operating pump fails, the system automatically switches to the backup pump, ensuring the continuity of the control system. This precise control command conversion and reliable execution mechanism provide a solid hardware foundation for accurate carbon source dosing.
[0219] Secondly, establishing a comprehensive process data recording and monitoring system is the data foundation for evaluating control effectiveness and continuous optimization. In the implementation at Wastewater Treatment Plant D, a comprehensive data acquisition and storage architecture was designed to record various key parameters of the carbon source addition process. Data acquisition is divided into three levels: basic operating parameters, control process parameters, and system status parameters. Basic operating parameters include the real-time flow rate of the carbon source pump (measured using an electromagnetic flowmeter with an accuracy of ±0.5%), cumulative addition amount, pump operating frequency, current, power, etc., with a sampling frequency of once per minute; control process parameters include the intermediate calculation results of each level of control algorithm, such as feedforward control output, feedback control output, weighting coefficients, and various calculated values of PID control, etc., with a sampling frequency of once every 10 minutes; system status parameters include equipment operating status, alarm information, communication quality indicators, etc., recorded using event-triggered methods. All collected data, after preprocessing (including outlier detection, data completion, and standardization), is stored in a time-series database (using InfluxDB), and an automatic archiving and backup mechanism is set up to ensure data integrity and traceability. Based on the collected process data, the system implements multi-level monitoring functions. Real-time monitoring at the operational level displays the current dosing status, equipment operating parameters, and key process indicators through an HMI (Human-Machine Interface), supporting daily monitoring and operation by operators. At the management level, trend analysis provides historical data queries, trend charts, and statistical reports through a web application, supporting engineers in performance evaluation and optimization analysis. At the system level, anomaly monitoring automatically identifies potential problems, such as equipment performance degradation, increased control deviation, or reduced model prediction accuracy, through statistical models and rule engines, generating early warning information. Particularly noteworthy is the system's establishment of a correlation analysis function between the carbon source dosing process and effluent water quality. Through time-series correlation analysis, lag analysis, and causal inference, the impact of carbon source dosing strategies on total nitrogen in the effluent is quantified, providing a scientific basis for subsequent optimization. This comprehensive data recording and monitoring system not only supports daily operation management but also provides rich historical data for continuous system optimization.
[0220] Third, continuous system optimization based on performance evaluation is a key mechanism to ensure long-term efficient operation. In the implementation at Wastewater Treatment Plant D, a self-optimization system driven by performance indicators was established, regularly evaluating system performance and triggering corresponding optimization actions. Core performance indicators fall into two main categories: prediction accuracy indicators and control effectiveness indicators. Prediction accuracy indicators primarily assess the consistency between model predictions and actual conditions, including BOD / COD prediction deviation (compared to laboratory measurements, calculated weekly) and carbon source demand prediction deviation (compared to the actual optimal dosage, evaluated monthly). Control effectiveness indicators assess the system's final treatment effect, including effluent total nitrogen compliance rate (calculated daily), carbon source utilization efficiency (calculated weekly), and system stability (characterized by the degree of fluctuation in control parameters, evaluated monthly). The system automatically calculates these indicators daily and compares them with preset performance thresholds: BOD / COD prediction deviation should not exceed 10%, effluent total nitrogen compliance rate should not be lower than 95%, and carbon source utilization efficiency should not be lower than 90% of the industry benchmark. When any indicator falls below the threshold, a corresponding optimization process is triggered. The optimization process consists of three levels: parameter fine-tuning, model updating, and system refactoring. Parameter fine-tuning is the most basic optimization level, triggered when indicators deviate slightly from the target. It mainly adjusts control parameters such as PID control parameters and weighting coefficients, without involving changes to the model structure. Model updating is a medium-level optimization, triggered when prediction accuracy continues to decline. It collects new data from the last 30 days, retrains the UV254-BOD / COD prediction model and the fluorescence feature ratio correction model, and updates model parameters to adapt to new water quality characteristics. Reconstruction is the highest level of optimization, triggered when multiple indicators fail to meet standards for an extended period or when significant changes in the system environment are detected (such as significant changes in influent characteristics or process adjustments). At this point, the entire control strategy needs to be re-evaluated, which may involve deep-seated changes such as model structure adjustments and control algorithm modifications, usually requiring the involvement of professional engineers. To make the optimization process more intelligent, a machine learning-based root cause analysis function is also integrated. When performance degradation is detected, it automatically analyzes possible causes, such as sensor drift, changes in influent characteristics, seasonal effects, or equipment failure, and provides targeted solutions in the optimization suggestion report.
[0221] Example 14
[0222] In this embodiment, the final control quantity is calculated using a composite control algorithm, further including a robust optimization step based on hypercube half-space learning: A control feature space is defined based on the UV254 absorbance, the intensity ratio of the tyrosine-like fluorescence peak to the tryptophan-like fluorescence peak, the predicted BOD / COD ratio, the influent flow rate, and the influent total nitrogen concentration; each feature in the control feature space is normalized and mapped to a unit hypercube; a labeled sample set is constructed based on historical control data, where the labels in the sample set represent the effectiveness of the control decision; and a full multinomial algorithm is used based on the labeled sample set. A time-based learning algorithm is used to learn a hypercube half-space model, which includes hyperplane parameters. The hyperplane parameters are updated using a projective gradient descent method until the empirical risk is less than a preset accuracy threshold, and the trained hypercube half-space model is output. Based on the trained hypercube half-space model, a feedforward control quantity is calculated. Combined with a feedback control quantity based on the total nitrogen monitoring data of the effluent, the feedforward weight coefficient and the feedback weight coefficient are dynamically calculated according to the current position in the unit hypercube. The feedforward control quantity and the feedback control quantity are then weighted and combined to generate the final control quantity.
[0223] Specifically, robust optimization based on hypercube half-space learning is an innovative control optimization method that transforms the control problem into a half-space learning problem in a high-dimensional space, thereby achieving a more accurate and disturbance-resistant control strategy. This process includes four sub-steps: defining and mapping the control feature space, constructing a labeled sample set, training the hypercube half-space model, and calculating position-aware dynamic weights.
[0224] First, defining and normalizing the control feature space is fundamental to constructing the hypercube space model. The "control feature space" refers to the multidimensional feature vector space describing the system state and control decisions, while the "unit hypercube" refers to a regularized feature space where all feature dimensions are normalized to the [0,1] interval. In the implementation at Wastewater Treatment Plant E, based on system analysis and expert experience, a 5-dimensional control feature space was defined, containing five key parameters: UV254 absorbance, the intensity ratio of tyrosine-like fluorescence peaks to tryptophan-like fluorescence peaks (T / W ratio), the predicted BOD / COD ratio, influent flow rate, and influent total nitrogen concentration. These five parameters collectively describe the core factors influencing carbon source dosing control decisions: the first three parameters characterize the properties and biodegradability of organic matter, while the latter two characterize the system load status. To allow features of different dimensions and ranges to be processed within the same framework, the system performed min-max normalization on all features; that is, for feature x, its normalized value is x' = (x - x) / x'. min ) / (x max - x min ), where xmin and x max These are the historical minimum and maximum values of the feature, respectively. Through this mapping, all feature values are constrained within the interval [0,1], forming a unit hypercube space. In practical implementation, to avoid the influence of extreme values on the mapping, factory E adopts a robust normalization scheme: x... min and x max The values were set to the 1st and 99th percentiles of the historical data for this feature, rather than absolute extrema, and normalized results outside the [0,1] interval were truncated. Furthermore, to handle correlations and redundancy among features, the system performed principal component analysis (PCA) preprocessing, retaining the principal components (typically 3-4) needed to explain 95% of the variance, and then performing subsequent modeling in the principal component space. This definition and mapping of the feature space lays the foundation for the mathematical processing of complex control problems, transforming the originally complex control decision problem into a classification problem in a regularized geometric space.
[0225] Secondly, constructing a labeled sample set is the data foundation for supervised learning. A labeled sample set refers to a dataset where each sample point, in addition to its feature vector, is accompanied by a label value representing the quality or effectiveness of the decision. In the implementation at Wastewater Treatment Plant E, a training set containing 2000 samples was constructed based on historical control data from the past 12 months. Each sample consists of two parts: a 5-dimensional feature vector x (mapped to a point in a unit hypercube) and a binary label y (y=1 indicates effective control, y=0 indicates ineffective control). The label was determined based on a comprehensive evaluation of historical control results, primarily considering three indicators: effluent total nitrogen compliance, carbon source utilization efficiency, and system stability. Specifically, when a control decision results in an effluent total nitrogen concentration below 9 mg / L (far below the 10 mg / L limit), a carbon source utilization rate above 85%, and system parameter fluctuations less than a preset threshold, the control is marked as "effective" (y=1). Conversely, if it results in effluent total nitrogen approaching or exceeding the limit, low carbon source utilization, or significant system fluctuations, it is marked as "ineffective" (y=0). To ensure the representativeness and balance of the samples, Plant E adopted a stratified sampling strategy: first, historical data was divided into multiple strata according to factors such as season, influent type, and load level; then, random sampling was performed within each stratum to ensure that the final sample set covered various operating conditions. In addition, to improve sample quality, outlier detection and data cleaning were implemented: the Local Outlier Factor (LOF) algorithm was used to identify outliers in the feature space, and combined with expert review, obviously erroneous samples were removed; for samples with high noise but whose error status could not be determined, local weighted regression (LOWESS) was used for smoothing. The final sample set contains approximately 70% "valid" samples and 30% "invalid" samples. This moderate class imbalance reflects the success rate distribution in actual control and helps the model learn more practical decision boundaries.
[0226] Next, training the hypercube half-space model using a full polynomial-time learning algorithm is the core of learning the effective control decision boundary. A "hypercube half-space model" refers to a mathematical model in which a unit hypercube space is divided into two parts (half-spaces) by a hyperplane; one part corresponds to the effective control region, and the other part corresponds to the ineffective control region. The "full polynomial-time learning algorithm" is a class of efficient learning algorithms whose computational complexity increases polynomially with the problem size. In the implementation at Wastewater Treatment Plant E, a half-space learning algorithm based on the Support Vector Machine (SVM) principle was adopted. This algorithm essentially seeks a hyperplane h(x) = w TThe algorithm assigns x + b to the hyperplane, ensuring that samples labeled "valid" are positioned on one side of the hyperplane and samples labeled "invalid" on the other. The goal is to find the optimal hyperplane parameters w and b that minimize the classification error. The training process uses projective gradient descent, an optimization algorithm combining gradient descent and parameter constraints. First, the hyperplane parameters w and b are randomly initialized. Then, the following steps are iteratively executed: randomly sample batches of data from the sample set (batch size set to 64); calculate the gradient of the loss function of the current model on this batch of data; update the parameters along the gradient direction; and project the updated parameters back to the constraint set (usually projecting w onto the L2 unit sphere to control model complexity). The loss function uses a combination of hinge loss and L2 regularization.
[0227] L(w,b) = (1 / n)·∑max(0, 1-y i ·(w T ·x i +b)) +λ Z ·||w|| 2 ,
[0228] Where λ Z The regularization coefficient is set to 0.01 (Efactor is set to 0.01). This loss function both penalizes classification errors and encourages finding the hyperplane with the largest margin, enhancing the model's generalization ability. The iterative process continues until the empirical risk (average loss on the training set) is less than the preset accuracy threshold (Efactor is set to 0.05) or the maximum number of iterations (5000 times) is reached. To improve training efficiency and model quality, Efactor implements three technical improvements: adaptive learning rate adjustment (initial value 0.01, dynamically adjusted according to loss changes); early stopping strategy (stopping training when validation set performance shows no improvement for 10 consecutive rounds); and ensemble learning (training multiple models and merging prediction results through voting or weighted averaging). After training, the resulting hypercube semi-space model can not only determine whether a control state is effective, but also calculate the "effectiveness probability" p(x) by the distance from a point to the hyperplane, providing a basis for subsequent dynamic weight calculations.
[0229] Finally, calculating the dynamic weight coefficients based on the model and generating the final control quantity is crucial for achieving adaptive control. In the implementation at Wastewater Treatment Plant E, the feedforward control quantity U_ff is generated by the water quality assessment model based on Examples 1-9, and the feedback control quantity U_fb is generated by the PID controller based on the total nitrogen in the effluent. These two basic control quantities need to be combined using weight coefficients to obtain the final control quantity U. Unlike the traditional fixed weight method, Plant E designed a dynamic weight calculation strategy based on the current state position, that is, dynamically determining the feedforward weight coefficient α(x) and the feedback weight coefficient β(x) according to the position of the current state x in the unit hypercube. Specifically, through the trained half-space model, the "effectiveness probability" p(x) of the current state x can be calculated. This probability reflects the reliability of the control decision in this state. The closer state x is to the center of the historical "effective" region, the closer p(x) is to 1; conversely, if x is located in the "ineffective" region or near the boundary between two regions, p(x) is smaller. Based on this probability, the dynamic weight calculation formula is:
[0230] α(x) = α min + (α max - α min )·p(x),
[0231] β(x) = 1 - α(x),
[0232] Where α min and α max These represent the minimum and maximum values of the feedforward weights (Efactor is set to 0.5 and 0.9, respectively). The core idea of this scheme is: when the system state is in a region where historical control performance has been good, more reliance is placed on model-based feedforward control to leverage its predictive advantages; when the system state is in an unfamiliar region or where historical control performance has been poor, more reliance is placed on feedback control to strengthen its error correction capabilities. The final control quantity is calculated using the following formula:
[0233] U = α(x)·U ff + β(x)·U fb .
[0234] In addition, to further improve the robustness of control, Factory E designed two enhancement mechanisms: one is the "trust region constraint," which applies when the calculated final control quantity U is compared with the control quantity U at the previous moment. prev When the difference is too large (more than 20%), the change range is automatically limited to prevent drastic fluctuations in control; secondly, "safety boundary monitoring" continuously detects whether the system state is close to the half-space boundary. When the distance is less than the preset threshold, the feedback control weight is automatically increased to improve the system's conservatism and safety.
[0235] Example 15
[0236] In this embodiment, the process of learning a hypercube half-space model using a full polynomial-time learning algorithm, wherein the hypercube half-space model includes hyperplane parameters, and updating the hyperplane parameters using a projective gradient descent method until the empirical risk is less than a preset accuracy threshold, and outputting the trained hypercube half-space model, includes: defining a noise model based on the labeled sample set, wherein the noise model includes a mixture ratio of real data distribution and noise data distribution, and calculating the maximum tolerable noise level; initializing the hyperplane parameters according to the labeled sample set, randomly sampling batch samples through an iterative process, calculating the gradient of the loss function, and updating the hyperplane parameters by applying projective gradient descent, stopping the iteration when the empirical risk is less than the preset accuracy threshold, and outputting the initially trained hypercube half-space model; and expanding the initially trained hypercube half-space model into a multi-objective optimization framework, defining a multi-objective function including denitrification efficiency, carbon source utilization efficiency, and energy consumption, calculating the Pareto front, selecting a tradeoff point on the Pareto front according to historical operating preferences, and outputting the trained hypercube half-space model.
[0237] Specifically, firstly, defining a noise model and analyzing the maximum tolerable noise level is fundamental to ensuring the robustness of the learning algorithm. A "noise model" refers to a mathematical representation of the noise contained in the training data, used to evaluate the performance stability of the algorithm in the presence of noise. In this embodiment, the F wastewater treatment plant adopts a mixed-distribution noise model, representing the actual data distribution D as the true data distribution D0. clean and noise data distribution D noise Weighted sum:
[0238] D = (1-η)·D clean + η·D noise ,
[0239] Where η is the noise ratio, representing the proportion of noisy samples in the dataset. This noise model considers two main sources of noise: measurement noise, which arises from sensor inaccuracies, calibration errors, etc., and typically manifests as random fluctuations in the data; and label noise, which is the inaccurate labeling of the effectiveness of control decisions, usually caused by factors such as operational record errors and subjective effect evaluation. To quantify the algorithm's tolerance to noise, F plant conducted systematic theoretical analysis and numerical simulations. The theoretical analysis is based on computational learning theory. Through VC dimension and sample complexity analysis, it is proved that under certain conditions, the learning algorithm has a PAC (Probably Approximately Correct) learning guarantee. Specifically, when the sample size n satisfies n ≥ O((d / ε) 2When η = log(1 / δ), the algorithm outputs the hypothesis that the error does not exceed ε with a probability of at least 1-δ, where d is the feature dimension. Numerical simulations test the performance trend of the algorithm by injecting different proportions of artificial noise into the training data. Simulation results show that when the noise proportion η does not exceed 15%, the algorithm's test accuracy decreases by no more than 5 percentage points, remaining within an acceptable range; when η exceeds 20%, the performance begins to decline significantly. Based on the combined theoretical analysis and simulation results, the maximum tolerable noise level η of the system is determined. max The threshold is 15%. This threshold provides an important reference for subsequent data preprocessing and algorithm design: on the one hand, data cleaning is needed to ensure that the noise level of the training data is below this threshold; on the other hand, the algorithm design needs to consider stability under this noise level. Specific measures include applying a double verification mechanism to the original data (e.g., only marking it as valid or invalid when the processing results are consistent over several consecutive days), and introducing a noise-robust loss function (such as a truncated loss function to limit abnormally large loss values). Through these methods, F factory ensures the reliability of the learning algorithm in real-world noisy environments.
[0240] Secondly, training the hypercube half-space model through an iterative process is the core of learning the effective control decision boundary. In the implementation at Wastewater Treatment Plant F, the initialization of the hyperplane parameters w and b adopted a heuristic method based on historical data distribution, rather than simple random initialization. Specifically, by calculating the centroids of effective and invalid samples, the orientation and position of the initial hyperplane were determined, giving the initial model a reasonable prior structure and accelerating the convergence process. Training employed batch stochastic gradient descent, randomly selecting 64 samples from the training set each time to form a batch, calculating the gradient of the loss function on that batch, and updating the parameters accordingly. The design of the loss function is a crucial step in the algorithm; Plant F adopted a combination of hinge loss and L2 regularization:
[0241] L(w,b) = (1 / n)·∑max(0, 1-y i ·(w T ·x i +b)) + λz·||w|| 2
[0242] Where λ ZThe regularization parameter is set to 0.01. The advantage of this loss function is that the hinge loss still penalizes correctly classified but marginally insufficient samples, encouraging the search for the decision boundary with the largest margin; the L2 regularization term controls model complexity, prevents overfitting, and improves generalization ability. Parameter updates use projective gradient descent, an optimization algorithm that combines gradient descent and parameter constraints. After each gradient update, the weight vector w is projected back to the L2 unit sphere. This projection step serves two purposes: first, it prevents the norm of the weight vector from becoming too large, leading to numerical instability; second, it enhances the model's generalization ability by constraining the hypothesis space. The learning rate uses an adaptive adjustment strategy, initially set to 0.01, and then dynamically adjusted according to changes in the loss function: when the loss decreases for several consecutive rounds, the learning rate is slightly increased to accelerate convergence; when the loss fluctuates or increases, the learning rate is decreased to stabilize the training process. Iterative training continues until one of two stopping conditions is met: the empirical risk (average loss on the training set) is less than a preset accuracy threshold of 0.05, or the maximum number of iterations of 5000 is reached. To prevent overfitting, F-factory adopted an early stopping strategy: the original dataset was divided into an 80% training set and a 20% validation set, and training was stopped when the performance on the validation set showed no improvement for 15 consecutive iterations. Furthermore, to improve the model's generalization ability and robustness, 5-fold cross-validation and ensemble learning were implemented: five models were trained, each using a different data partition, and the prediction results were merged through a voting mechanism. This comprehensive and optimized training process ensured that the final model not only performed well on the training data but also possessed predictive capabilities on new data.
[0243] Then, extending the initially trained model into a multi-objective optimization framework is key to balancing multiple performance indicators. "Multi-objective optimization" refers to an optimization method that simultaneously considers multiple potentially conflicting objective functions to find the optimal or satisfactory solution. In the implementation at Wastewater Treatment Plant F, control decisions need to balance three key objectives: denitrification efficiency (f1, the reciprocal of the total nitrogen concentration in the effluent), carbon source utilization efficiency (f2, the reciprocal of the amount of carbon source consumed to remove a unit of nitrogen), and energy consumption (f3, mainly aeration and pumping energy consumption). These three objectives often have an inverse relationship: for example, improving denitrification efficiency usually requires increasing the amount of carbon source added, leading to a decrease in carbon source utilization efficiency; reducing aeration can reduce energy consumption, but may affect the nitrification process, indirectly reducing denitrification efficiency. To find the optimal balance point among these three objectives, Plant F extended the initially trained hypercube half-space model into a multi-objective optimization framework. First, an objective function integrating multiple objectives was defined: F(w,b) = (f1,f2,f3), where each component corresponds to a performance indicator. Since these objectives may conflict with each other, there is no single solution that simultaneously optimizes all objectives. Therefore, it is necessary to find a Pareto optimal set of solutions (Pareto front). The Pareto front refers to the set of solutions that cannot improve any objective without compromising at least one objective; it represents the optimal trade-off in multi-objective problems. Factory F uses an improved NSGA-II (Non-dominated sorting genetic algorithm II) to calculate the Pareto front. This algorithm generates a series of candidate solutions by simulating the evolutionary process and selects them based on non-dominated sorting and crowding calculations, approaching the true Pareto front generation by generation. The improvements are: the fitness function incorporates the classification results of the hypercube half-space model, ensuring that the generated solutions not only meet the multi-objective optimization requirements but also lie within the "effective control" region; a local search strategy is introduced, performing local optimization on elite individuals every few generations to accelerate convergence; and an adaptive population size is implemented, dynamically adjusting the population size according to the front complexity. After about 100 generations of evolution, the algorithm converges to a Pareto front containing about 50 non-dominated solutions, each solution corresponding to a set of hyperplane parameters (w,b) and corresponding multi-objective performance indices.
[0244] Selecting the final solution on the Pareto front is a crucial step in multi-objective optimization. Traditional methods typically rely on weighted summation or priority settings, but these methods often fail to accurately reflect the decision-makers' true preferences. To address this, Plant F innovatively designed a solution selection method based on historical operational preferences. First, the decision-making patterns of operators were extracted from historical operational data, including the relative importance placed on the three objectives under different conditions. Specifically, control decisions made by five experienced operators under over 200 typical operating conditions were collected, establishing a decision preference database. Then, based on this data, a preference learning model was trained, capable of predicting operators' possible decision preferences under given conditions. The model employs inverse reinforcement learning to infer the underlying reward function from observed decisions, reflecting the operators' implicit preferences. During operation, the system first uses the preference model to predict appropriate objective weights based on the current operating conditions; then, based on these weights, the best-matching point on the Pareto front is selected as the final solution. This approach combines mathematical optimality with human experience and wisdom, ensuring Pareto optimality while also considering preferences and constraints in practical operation, making the control system more practical and reliable. Notably, the system also implements an online learning function for the preference model: when operators manually adjust system parameters, these adjustments are recorded as new preference data points, and the preference model is updated periodically, allowing the system to gradually adapt to the decision-making style of the operating team.
[0245] In practical applications at Wastewater Treatment Plant F, the control system based on full polynomial time learning and multi-objective optimization has achieved significant results. Compared to traditional single-model and fixed-weight methods, the new system has improved in three key performance indicators: the average total nitrogen concentration in the effluent decreased from 8.2 mg / L to 7.5 mg / L, carbon source utilization efficiency increased by 17%, and energy consumption per unit treated volume decreased by 8%. Particularly under conditions of large system load fluctuations, the multi-objective optimization framework exhibits excellent adaptability, dynamically adjusting the control strategy according to actual conditions to maximize resource utilization efficiency while ensuring effluent compliance. Long-term operational data shows that it not only improves treatment efficiency but also reduces the need for manual intervention, decreasing the frequency of manual adjustments by operators by 65%, significantly alleviating operational burden. Furthermore, through continuous data accumulation and model updates, the system's performance continues to improve, demonstrating good adaptability and development potential. This method, combining computational learning theory and multi-objective optimization, provides a new technical path for intelligent control of wastewater treatment, achieving a balance between efficiency, economy, and stability, and has broad application prospects and promotional value.
[0246] Example 16
[0247] like Figure 4As shown, the present invention also provides a water quality fingerprint recognition and feedforward dosing control system, comprising:
[0248] The water quality fingerprint generation module 10 is used to acquire the UV254 absorption value of the influent and the signal intensity of the tyrosine-like fluorescence peak and the tryptophan-like fluorescence peak in the three-dimensional fluorescence spectrum. The module performs data fusion and preprocessing on the UV254 absorption value, the signal intensity of the tyrosine-like fluorescence peak and the signal intensity of the tryptophan-like fluorescence peak to generate influent water quality characteristic fingerprint data.
[0249] The prediction calculation module 20 is used to calculate the predicted BOD / COD ratio and effective carbon content based on the influent water quality characteristic fingerprint data, using a pre-established UV254-BOD / COD prediction model and a fluorescence characteristic ratio correction model.
[0250] The carbon source demand calculation module 30 is used to calculate the theoretically required amount of carbon source based on the predicted BOD / COD ratio and effective carbon content, combined with the real-time monitored influent flow rate and total nitrogen concentration in the influent, and subtract the effective carbon content to obtain the external carbon source demand. After time-series adjustment of the external carbon source demand, the predicted carbon source dosage value is output.
[0251] The composite control module 40 is used to calculate the final control quantity based on the predicted value of the carbon source dosage as a feedforward control signal, combined with the total nitrogen monitoring data of the effluent as a feedback control signal, through a composite control algorithm, and convert the final control quantity into an operation command for the carbon source dosing device and execute it.
[0252] This water quality fingerprinting and feedforward dosing control system achieves precise carbon source dosing control through four core modules. The water quality fingerprint generation module collects the UV254 absorbance and specific fluorescence spectral peak information of the influent, generating a water quality characteristic fingerprint through data fusion. The prediction and calculation module uses these fingerprint data and a specialized model to predict the BOD / COD ratio and available carbon content. The carbon source demand calculation module combines the prediction results with real-time monitored flow and total nitrogen data, calculates the external carbon source demand based on a denitrification theoretical model, and performs time-series optimization. The composite control module uses the predicted values as a feedforward signal, combines them with effluent total nitrogen monitoring data as a feedback signal, calculates the final control quantity through a composite control algorithm, and converts it into operational commands for execution. This integrated scheme organically combines spectral analysis, predictive modeling, and intelligent control, achieving precise control of carbon source dosing and optimized resource utilization.
[0253] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A water quality fingerprint recognition and feedforward dosing control method, characterized in that, include: The UV254 absorbance value of the influent and the signal intensity of the tyrosine-like fluorescence peak and the tryptophan-like fluorescence peak in the three-dimensional fluorescence spectrum are obtained. The UV254 absorbance value, the signal intensity of the tyrosine-like fluorescence peak and the signal intensity of the tryptophan-like fluorescence peak are fused and preprocessed to generate influent water quality characteristic fingerprint data. Based on the influent water quality fingerprint data, the predicted BOD / COD ratio and available carbon content are calculated using a pre-established UV254-BOD / COD prediction model and a fluorescence feature ratio correction model. Based on the predicted BOD / COD ratio and effective carbon content, combined with the real-time monitored influent flow rate and total nitrogen concentration, the theoretically required carbon source amount is calculated according to the denitrification theoretical model, and the effective carbon content is deducted to obtain the external carbon source demand. After time-series adjustment of the external carbon source demand, the predicted carbon source dosage is output. Based on the predicted carbon source dosage as a feedforward control signal, and combined with the total nitrogen monitoring data of the effluent as a feedback control signal, the final control quantity is calculated through a composite control algorithm, and the final control quantity is converted into an operation command for the carbon source dosing device and executed. The method also includes an adaptive learning step based on a deep actor criticism framework, including: Based on historical influent water quality fingerprint data and corresponding measured BOD / COD values, carbon source dosage, and treatment effect data, an expert database is constructed. This database includes data sets on water quality status, evaluation results, rewards, and subsequent status. Based on the historical influent water quality fingerprint data, an actor network is designed. The input to the actor network is the historical influent water quality fingerprint data, and the output is the BOD / COD value and available carbon content to be trained. Based on the historical influent water quality fingerprint data and the BOD / COD value to be trained, a critic network is designed. The input to the critic network is the historical influent water quality fingerprint data and the BOD / COD value to be trained. The BOD / COD value is calculated, and the output of the critic network is the carbon source addition amount. Based on the expert database, a preset number of data groups are randomly selected from the expert database to form batch training data. The actor network and the critic network are used to perform offline policy learning on the batch training data. The parameters of the actor network are updated by calculating the policy gradient and by minimizing the temporal difference error, and the trained actor network model is output. Based on the influent water quality characteristic fingerprint data, the influent water quality characteristic fingerprint data is input into the trained actor network model, and the predicted BOD / COD ratio and available carbon content are output.
2. The method according to claim 1, characterized in that, The process of fusing and preprocessing the UV254 absorbance, the signal intensity of the tyrosine-like fluorescence peak, and the signal intensity of the tryptophan-like fluorescence peak to generate influent water quality characteristic fingerprint data includes: Based on the UV254 absorbance, the signal intensity of the tyrosine-like fluorescence peak, and the signal intensity of the tryptophan-like fluorescence peak, the UV254 absorbance and the signal intensity are fused to generate preliminary fused data. Based on the preliminary fused data, noise is eliminated using the sliding window averaging method to generate denoised data. Based on the noise-reduced data, time tags are added to generate the influent water quality characteristic fingerprint data.
3. The method according to claim 1, characterized in that, The acquisition of the UV254 absorbance value of the influent and the signal intensity of the tyrosine-like fluorescence peak and tryptophan-like fluorescence peak in the three-dimensional fluorescence spectrum includes: Based on the inlet pipe of the sewage treatment system, an online UV-visible spectrometer and a three-dimensional fluorescence spectrometer are installed on the inlet pipe, and a sampling frequency of 5-15 minutes is set to establish a monitoring point. Based on the online UV-visible spectrometer, the absorbance at a wavelength of 254 nm is continuously monitored and data is calibrated according to a preset calibration curve to obtain the calibrated UV254 absorbance value. Based on the aforementioned three-dimensional fluorescence spectrometer, a three-dimensional fluorescence spectrum of the incoming water is acquired. A peak recognition algorithm is used to identify and quantify the signal intensity of the tyrosine-like fluorescence peak with an excitation wavelength of 230-275 nm and an emission wavelength of 300-320 nm, as well as the signal intensity of the tryptophan-like fluorescence peak with an excitation wavelength of 270-280 nm and an emission wavelength of 340-380 nm. The signal intensity of the tyrosine-like fluorescence peak and the signal intensity of the tryptophan-like fluorescence peak are then output.
4. The method according to claim 1, characterized in that, Establish a UV254-BOD / COD prediction model, including: Based on at least three months of historical monitoring data of the influent of the wastewater treatment system, the influent UV254 absorbance, laboratory-measured BOD and COD values were collected to construct a training dataset. Based on the training dataset, the first functional relationship between UV254 absorbance and BOD and the second functional relationship between UV254 absorbance and COD are established using multiple linear regression or support vector machine regression algorithms, respectively, to generate the UV254-BOD / COD prediction model.
5. The method according to claim 1, characterized in that, Establish a fluorescence characteristic ratio correction model, including: Based on the signal intensity of the tyrosine-like fluorescence peak and the signal intensity of the tryptophan-like fluorescence peak, the intensity ratio of the tyrosine-like fluorescence peak to the tryptophan-like fluorescence peak is calculated. Using the intensity ratio and the UV254 absorption value as input variables, and combining the laboratory-measured BOD / COD ratio as the output variable, a functional relationship between the input variables and the output variables is established to generate the fluorescence characteristic ratio correction model.
6. The method according to claim 1, characterized in that, The offline policy learning using the actor network and the critic network on the batch training data includes: Based on the data set in the batch training data, the water quality status in the data set is input into the actor network, and the current policy action is output. The current policy action is the BOD / COD value and effective carbon content to be trained. The water quality status and the current policy action are input into the critic network, and the Q value of the current state-action pair is output. Based on the Q value, the policy gradient is calculated and the actor network parameters are updated to generate the updated actor network. Based on the rewards and subsequent states of the data set, calculate the target value of the temporal difference error, update the parameters of the critic network to minimize the squared difference between the Q value output by the critic network and the target value of the temporal difference error, and generate the critic network with updated parameters. Based on the updated actor network with the parameters, the water quality state is input into the updated actor network, and the updated policy action is output. The water quality state and the updated policy action are input into the updated critic network, and the updated Q value is output. Based on the updated Q value, a policy constraint term is added to the policy gradient calculation to ensure that the updated policy does not deviate from the expert policy. The policy constraint term is the expected value of the square of the difference between the updated policy action and the expert action in the data set, and the trained actor network model is generated.
7. The method according to claim 1, characterized in that, It also includes personalized computation steps based on multimodal dynamic agent learning, including: Based on multimodal data including the influent water quality characteristic fingerprint data, the predicted BOD / COD ratio and effective carbon content, influent flow rate, pH value, temperature, total nitrogen concentration, ammonia nitrogen concentration, historical carbon source dosage and treatment effect data, the multimodal data is collected and integrated, and the multimodal data is standardized and denoised to generate preprocessed multimodal data. Based on the preprocessed multimodal data, feature extraction and cross-modal fusion are performed on data from different sources through a gated cross-modal fusion network to generate fused feature representations; Based on the fusion feature representation, water inflow type clustering is performed through a dual-constraint proxy optimization mechanism and a dynamic candidate management mechanism. The current water inflow sample is assigned to the identified water inflow type and the membership degree to each water inflow type is calculated. The clustering results and membership degrees are then output. Based on the clustering results, a dedicated carbon source demand calculation model is constructed for each identified influent type. The dedicated carbon source demand calculation model maps the fusion feature representation to the carbon source dosage. Based on the membership degree and the dedicated carbon source demand calculation model, the carbon source addition amount output by the dedicated carbon source demand calculation model for each influent type is weighted and combined to output the external carbon source demand amount.
8. The method according to claim 7, characterized in that, The process of extracting features from and fusing cross-modal data from different sources using a gated cross-modal fusion network to generate a fused feature representation includes: Based on the preprocessed multimodal data, modality-specific encoders are used to encode the spectral data, conventional water quality parameter data, and process operation parameter data respectively, generating modality-specific features; Based on the modality-specific features, the correlation strength between different modalities is calculated using an attention mechanism to generate an attention weight matrix; Based on the modality-specific features and the attention weight matrix, information flow is controlled by a gating unit to perform weighted fusion of different modality features and generate the fused feature representation.
9. The method according to claim 7, characterized in that, The process involves clustering influent types using a dual-constraint proxy optimization mechanism and a dynamic candidate management mechanism. This process assigns the current influent sample to the identified influent type and calculates the membership degree for each type, outputting the clustering results and membership degrees. Based on the fusion feature representation, initialize multiple inlet type proxy vectors; Based on the fusion feature representation, the similarity between the current inflow sample and the proxy vector of each inflow type is calculated. The current inflow sample is assigned to each inflow type using a fuzzy clustering method, and the membership degree is calculated to generate preliminary clustering results. The total nitrogen compliance rate and carbon source utilization efficiency of the effluent are obtained as treatment effect feedback data. Based on the preliminary clustering results and the treatment effect feedback data, the influent type proxy vector is updated periodically. The number of categories is dynamically adjusted according to the intra-class sample density and inter-class distance. The clustering results and membership degree are output.
10. The method according to claim 1, characterized in that, The step of adjusting the external carbon source demand based on time series and outputting the predicted carbon source dosage includes: Based on the external carbon source demand, combined with the bioreaction kinetics and system hydraulic residence time, the external carbon source demand is adjusted in a time sequence, a phased addition strategy is formulated, and the predicted value of the carbon source addition is output.
11. The method according to claim 1, characterized in that, The process of using the predicted carbon source dosage as a feedforward control signal, combined with the total nitrogen monitoring data of the effluent as a feedback control signal, and calculating the final control quantity through a composite control algorithm includes: Based on the predicted carbon source dosage, the feedforward control output is calculated by the feedforward controller. The total nitrogen concentration in the effluent is monitored in real time. The deviation between the total nitrogen concentration in the effluent and the preset target value is calculated. The feedback control output is calculated through a proportional-integral-derivative control algorithm. Based on the feedforward control output and the feedback control output, the final control quantity is generated by weighting and combining the set feedforward weight coefficient and feedback weight coefficient.
12. The method according to claim 11, characterized in that, After converting the final control quantity into an operation instruction for the carbon source dosing device and executing it, the process further includes: Based on the final control quantity, the final control quantity is converted into an operation command to control the flow rate or switching frequency of the carbon source pump and sent to the carbon source dosing device; Based on the execution process of the carbon source dosing device, the dosing flow rate, dosing time, and equipment status parameters are recorded to generate process data records; Based on the process data records and the effluent total nitrogen monitoring data, the deviation between the predicted BOD / COD ratio and the measured value is calculated as a prediction accuracy index, and the effluent total nitrogen compliance rate is calculated as a control effect index. When the prediction accuracy index or the control effect index is lower than a preset threshold, the control parameters in the UV254-BOD / COD prediction model, the fluorescence feature ratio correction model, and the composite control algorithm are updated, and system optimization suggestions are output.
13. The method according to claim 5, characterized in that, The final control quantity is calculated using a composite control algorithm, further including a robust optimization step based on hypercube half-space learning: Based on the UV254 absorbance, the intensity ratio of the tyrosine-like fluorescence peak to the tryptophan-like fluorescence peak, the predicted BOD / COD ratio, the influent flow rate, and the total nitrogen concentration in the influent, a control feature space is defined, and each feature in the control feature space is normalized and mapped to a unit hypercube. Based on historical control data, a labeled sample set is constructed, where the labels in the sample set represent the effectiveness of control decisions; Based on the labeled sample set, a hypercube half-space model is learned using a full polynomial time learning algorithm. The hypercube half-space model includes hyperplane parameters. The hyperplane parameters are updated using a projective gradient descent method until the empirical risk is less than a preset accuracy threshold. The trained hypercube half-space model is then output. Based on the trained hypercube half-space model, a feedforward control quantity is calculated. Combined with a feedback control quantity based on the total nitrogen monitoring data of the effluent, the feedforward weight coefficient and feedback weight coefficient are dynamically calculated according to the current position in the unit hypercube. The feedforward control quantity and the feedback control quantity are then weighted and combined to generate the final control quantity.
14. The method according to claim 13, characterized in that, The hypercube half-space model is learned using a full polynomial-time learning algorithm. This model includes hyperplane parameters. The hyperplane parameters are updated using a projective gradient descent method until the empirical risk is less than a preset accuracy threshold. The trained hypercube half-space model is then output, including: Based on the labeled sample set, a noise model is defined, which includes the mixing ratio of the real data distribution and the noise data distribution, and the maximum tolerable noise level is calculated. Based on the labeled sample set, the hyperplane parameters are initialized. Through iterative processes, batch samples are randomly sampled, the gradient of the loss function is calculated, and the hyperplane parameters are updated by applying projective gradient descent. When the empirical risk is less than a preset accuracy threshold, the iteration stops, and the initially trained hypercube half-space model is output. Based on the initially trained hypercube half-space model, it is extended into a multi-objective optimization framework. A multi-objective function is defined, which includes denitrification efficiency, carbon source utilization efficiency, and energy consumption. The Pareto front is calculated, and a trade-off point is selected on the Pareto front according to historical operating preferences. The trained hypercube half-space model is then output.
15. A water quality fingerprint recognition and feedforward dosing control system, characterized in that, include: The water quality fingerprint generation module is used to acquire the UV254 absorption value of the influent and the signal intensity of the tyrosine-like fluorescence peak and the tryptophan-like fluorescence peak in the three-dimensional fluorescence spectrum. The module performs data fusion and preprocessing on the UV254 absorption value, the signal intensity of the tyrosine-like fluorescence peak and the signal intensity of the tryptophan-like fluorescence peak to generate influent water quality characteristic fingerprint data. The prediction calculation module is used to calculate the predicted BOD / COD ratio and available carbon content based on the influent water quality characteristic fingerprint data, using a pre-established UV254-BOD / COD prediction model and a fluorescence characteristic ratio correction model. The carbon source demand calculation module is used to calculate the theoretically required amount of carbon source based on the predicted BOD / COD ratio and effective carbon content, combined with the real-time monitored influent flow rate and total nitrogen concentration in the influent, and subtract the effective carbon content to obtain the external carbon source demand. After time-series adjustment of the external carbon source demand, the predicted carbon source dosage value is output. The composite control module is used to calculate the final control quantity based on the predicted value of carbon source dosage as a feedforward control signal, combined with the total nitrogen monitoring data of the effluent as a feedback control signal, through a composite control algorithm, and convert the final control quantity into an operation command for the carbon source dosing device and execute it. The water quality fingerprinting and feedforward dosing control system is also used for: constructing an expert database based on historical influent water quality characteristic fingerprint data and corresponding measured BOD / COD values, carbon source dosage, and treatment effect data. The expert database includes data sets on water quality status, evaluation results, rewards, and subsequent status. Based on the historical influent water quality characteristic fingerprint data, an actor network is designed, with the historical influent water quality characteristic fingerprint data as input and the BOD / COD value and available carbon content to be trained as output. Based on the historical influent water quality characteristic fingerprint data and the BOD / COD value to be trained, a critic network is designed, with the historical influent water quality characteristic fingerprint data as input. The system uses characteristic fingerprint data and the BOD / COD value to be trained. The output of the critic network is the carbon source addition amount. Based on the expert database, a preset number of data groups are randomly selected from the expert database to form batch training data. The actor network and the critic network are used to perform offline policy learning on the batch training data. The parameters of the actor network are updated by calculating the policy gradient and by minimizing the temporal difference error. The trained actor network model is then output. Based on the influent water quality characteristic fingerprint data, the influent water quality characteristic fingerprint data is input into the trained actor network model, and the predicted BOD / COD ratio and available carbon content are output.