Method for predicting total nitrogen concentration of effluent of sewage treatment plant based on attention mechanism

Through the total nitrogen concentration prediction method of sewage treatment plant effluent based on attention mechanism, the long-term delay and variable correlation problems during carbon source injection process are solved, and the accurate prediction of total nitrogen concentration of effluent is achieved, which improves the carbon source utilization efficiency and sewage treatment effect.

CN120496657APending Publication Date: 2025-08-15CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510681684.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing wastewater treatment carbon source application process has long delays and complex variable correlations, which makes it difficult to accurately predict the total nitrogen concentration of the effluent, affecting the carbon source utilization efficiency and treatment cost.

Method used

The total nitrogen concentration prediction method for effluent in sewage treatment plants based on attention mechanism is adopted, and the total nitrogen concentration prediction of effluent in sewage treatment plants is extracted through time scale and variable dimension feature, combined with a multi-layer perceptron, a M-DMCAN model is constructed, and the model training and optimization is used for PyTorch framework.

Benefits of technology

Accurate prediction of the total nitrogen concentration of the effluent water is achieved, the efficiency of carbon source utilization is improved, the treatment cost is reduced, and the water quality of the effluent is stable and meets the standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496657A_ABST
    Figure CN120496657A_ABST
Patent Text Reader

Abstract

The invention relates to a sewage treatment plant effluent total nitrogen concentration prediction method based on an attention mechanism, and belongs to the technical field of sewage treatment. The method comprises the following steps: a time scale feature extraction module for performing down-sampling, decomposition and period and trend feature mixing on a single target variable time sequence; the variable dimension feature extraction module captures cross correlation among multi-dimensional variables through a Patch-based segmentation method and a multi-head attention mechanism; and the multi-feature prediction and polymerization module maps the features extracted by the two modules into effluent total nitrogen concentration prediction output through a multi-layer perceptron, and performs weighted polymerization to obtain a final prediction result. The model can effectively capture the long-term dependence and complex interaction relation between the effluent total nitrogen concentration and the carbon source adding amount, is high in prediction precision, provides an accurate basis for sewage treatment carbon source adding control, can improve the operation effect of the sewage treatment process and the carbon source utilization efficiency, reduces the treatment cost, and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of sewage treatment and relates to a method for predicting the total nitrogen concentration in effluent from a sewage treatment plant based on an attention mechanism. Background Art

[0002] Carbon source addition is crucial in wastewater treatment, directly impacting key operational indicators such as denitrification efficiency and nitrogen emissions. However, current carbon source addition processes in wastewater treatment face numerous challenges. Firstly, the mechanisms are extremely complex and the dynamics are protracted, making it difficult to accurately describe the dynamic characteristics of the carbon source addition process. For example, microbial carbon source utilization is influenced by numerous factors, including temperature, dissolved oxygen, and the carbon-nitrogen ratio. These intertwined factors increase the complexity of the process. Secondly, the actual carbon source addition control process in wastewater treatment suffers from long reaction times, nonlinearity, and time variability. Traditional carbon source addition control methods, such as those relying on manual experience or simple formulas to calculate carbon source dosage, are ill-suited to complex and changing operating conditions and cannot accurately control carbon source dosage. This not only leads to low carbon source utilization efficiency and increased treatment costs, but can also cause excessive effluent total nitrogen concentrations, negatively impacting the environment. Furthermore, because carbon source addition occurs within the denitrification tank, there is a significant time lag between carbon source addition and its impact on effluent total nitrogen concentration, typically requiring 4-5 hours or even longer. Moreover, the sewage treatment process involves numerous variables, and the correlations between these variables are complex, further increasing the difficulty of accurately predicting the total nitrogen concentration in the effluent. Existing sewage treatment water quality prediction models, whether mechanism-based or data-driven, are unable to effectively address these issues. For example, mechanism-based models require a precise understanding of system dynamics and complex parameter calibration, and have poor generalization capabilities. While data-driven models can, to a certain extent, uncover potential patterns in the data, they still suffer from insufficient accuracy and adaptability when dealing with long time lags and complex variable correlations.

[0003] Therefore, developing a model and method that can accurately predict the total nitrogen concentration in sewage treatment effluent is of great practical significance for realizing intelligent control of carbon source addition, improving sewage treatment effects and reducing costs. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a method for predicting total nitrogen concentration in sewage treatment plant effluent based on an attention mechanism, addressing the difficulty in accurately predicting effluent total nitrogen concentration due to the long time lag and complex variable correlations in the carbon source addition process. By accurately predicting effluent total nitrogen concentration, a reliable basis is provided for intelligent control of carbon source addition, thereby achieving a synergistic improvement in carbon source utilization efficiency and denitrification efficiency, reducing sewage treatment costs, and ensuring that effluent water quality remains stable and meets standards.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] A method for predicting total nitrogen concentration in effluent from a sewage treatment plant based on an attention mechanism includes the following steps:

[0007] S1: Data acquisition and preprocessing: Obtain sewage treatment data, including indicators such as carbon source addition flow rate, effluent total nitrogen concentration, influent total nitrogen concentration, influent flow rate, influent chemical oxygen demand, and influent nitrate nitrogen. Perform preprocessing on the data, such as missing value filling, time alignment, and outlier processing, to obtain a data set, which is divided into a training set and a validation set.

[0008] S2: Model construction and training: Construct a prediction model for effluent total nitrogen concentration, including a time-scale feature extraction module, a variable-dimensional feature extraction module, and a multi-feature prediction and aggregation module; train the model using the training set, evaluate model performance using the validation set, and save the optimal model parameters;

[0009] The time scale feature extraction module is used to extract the time scale features of the time series itself, including downsampling the input data to construct a multi-time scale series, embedding it into a unified feature space, and then learning and mixing the information of the two features of cycle and trend by stacking the decomposition module and the mixing module of the time series;

[0010] The variable dimension feature extraction module is used to extract variable dimension features between variables, including using a patch-based segmentation method to segment all variables of the input data by time step, linearly projecting the data segments and adding position information, and then using multi-head attention to extract cross-correlation features and residual connections between variables;

[0011] The multi-feature prediction and aggregation module maps the features processed by the time scale feature extraction module and the variable dimension feature extraction module into the prediction output of the total nitrogen concentration in the effluent through a multi-layer perceptron (MLP), and then performs weighted aggregation to obtain the final prediction result;

[0012] S3: Model prediction: Input the test data into the trained effluent total nitrogen concentration prediction model to obtain the predicted value of effluent total nitrogen concentration in the future.

[0013] Furthermore, in step S2, the specific implementation steps of the time scale feature extraction module are:

[0014] S201: Downsample the input data, halving the time step each time, to obtain a set of multi-time-scale sequences, where the bottom-level sequence is equivalent to the original sequence, and the highest-level pooling layer sequence provides the most macroscopic feature perspective; then, embed the multi-time-scale sequences into a unified feature space for subsequent processing.

[0015] S202: By stacking multiple layers of time series decomposition modules and mixing modules, time dimension features containing the periodicity and trend characteristics of the data itself are obtained, specifically including:

[0016] In the Cycle and Trend Decomposition module, each scale series is decomposed into a cyclical component and a trend component to avoid confusion in the overall analysis. The cyclical component is calculated using the Avgpool function, and the trend component is obtained by subtracting the cyclical component from the original series.

[0017] In the feature component mixing module, the periodic information is mixed in a bottom-up order through residual connections, and the trend information is mixed in a top-down order through residual connections.

[0018] Through this mixture of cycles and trends in different directions, more detailed cycle and trend modeling can be achieved, effectively capturing long-term fluctuation patterns and cyclical changes in time series.

[0019] Furthermore, in step S202, the time dimension features including the periodicity and trend features of the data itself are obtained through the decomposition module and the mixing module of the L-layer time series.

[0020]

[0021] Among them, UM(·) represents the mixed operation of the period, which includes two linear layers and the GELU activation function; VM(·) represents the mixed operation of the trend, which includes two linear layers and the GELU activation function; φ(·) includes two linear layers and the GELU activation function of the middle layer; M represents the maximum number of sampling layers, m represents the number of downsampling layers, l represents the number of trend decomposition layers, and L represents the period of the sequence. represents the mth periodic component of the lth pooling layer; Represents the mth trend component of the lth pooling layer.

[0022] Furthermore, in step S2, the specific implementation steps of the variable dimension feature extraction module are:

[0023] S211: Based on the Patch segmentation method, all variables of the input data are segmented according to the time step according to a specific length to obtain a set of time segments;

[0024] S212: Use linear projection to encode the data segments and incorporate positional information. The mapping matrix used is shared with the time-scale feature extraction module. This processing not only includes information about the variables themselves, but also incorporates positional information, helping to better capture correlations between variables.

[0025] S213: Multi-head attention is used to extract cross-correlation features between variables, and through residual connection, variable dimension cross-correlation features containing cross-dependencies between different variables are obtained, thereby effectively capturing the complex interactive relationships between multidimensional variables.

[0026] Furthermore, in step S213, the expression of the variable dimension cross-correlation feature E is obtained as follows:

[0027] E=Norm(E+MAS(Q,K,V))

[0028] Among them, Norm(·) is the normalization operation; MAS(Q,K,V) represents the multi-head attention layer, Q, K, V represent the query, key and value respectively, and all variables share the same MAS layer; E represents the embedding set, Among them, e i,d is the embedding of the i-th time segment corresponding to the d-th variable sequence, D is the input variable, L seg is the time segment length; T is the sequence window size.

[0029] Furthermore, in step S2, the multi-feature prediction and aggregation module maps the features obtained after time scale feature extraction and variable dimension feature extraction processing into the predicted output of total nitrogen concentration in effluent water through a multi-layer perceptron. The specific implementation steps are as follows:

[0030] S221: During the time scale feature prediction and aggregation process, the prediction results from different scales are normalized and processed by a multi-layer perceptron before being aggregated.

[0031] S222: During the variable dimension feature prediction and aggregation process, hierarchical cross-attention learning is performed according to the different lengths of the time series segments, and the prediction outputs at different levels are aggregated to obtain the final prediction output;

[0032] S223: The time scale feature prediction results and the variable dimension feature prediction results are weighted and aggregated to obtain the final effluent total nitrogen concentration prediction value. The weight coefficient is adjusted in the range of [0,1] according to the actual situation to optimize the prediction effect.

[0033] Furthermore, in step S223, the final predicted value Y of the effluent total nitrogen concentration is expressed as:

[0034] Y=(1-α)·Y S-pred +αY M-pred

[0035] Among them, α∈[0,1] is the cross-dimensional prediction weight that controls cross attention; Y S-pred is the predicted value obtained by time scale feature extraction, and the expression is in, represents the mth time dimension feature with a sequence period of L, MLP represents a multi-layer perceptron, and Norm(·) represents a normalization operation;

[0036] Y M-pred is the predicted value obtained by extracting variable dimension features, and the expression is Among them, c is the number of layers of cross attention learning, Y c is the hierarchical prediction output, Among them, E is the cross-correlation feature of variable dimension, y i is the predicted value corresponding to the i-th time segment, H is the window size of the prediction sequence, L seg is the length of the time segment, d is the d-th variable sequence, and D is the input variable.

[0037] Furthermore, in step S2, model training specifically includes: using the training set to build a water outlet total nitrogen concentration prediction model based on the PyTorch framework, setting model hyperparameters, using the mean square error as the loss function, and using the Adam algorithm to update the model parameters. During the training process, the validation set is used to evaluate the model performance and save the optimal model parameters.

[0038] The beneficial effects of the present invention are:

[0039] (1) Accurate prediction performance: The prediction model of the present invention can effectively capture the long-term dependence and complex interactive relationship between the total nitrogen concentration in the effluent and the amount of carbon source added by integrating time scale feature extraction and variable dimension feature extraction. Compared with traditional prediction models, it has significant advantages in dealing with long time lags and complex variable correlation problems. The experimental results on real sewage data sets show that the model has excellent performance in various long-term and short-term prediction tasks, and the prediction error is better than the comparison model in multiple indicators (such as mean square error, mean absolute error and root mean square error), and can more accurately predict the changing trend of the total nitrogen concentration in the effluent.

[0040] (2) Improve sewage treatment efficiency: Accurate prediction of effluent total nitrogen concentration provides a reliable basis for intelligent control of carbon source addition. By accurately adjusting the carbon source dosage based on the prediction results, it is possible to avoid excessive or insufficient carbon source addition, improve carbon source utilization efficiency, reduce reagent waste, and reduce treatment costs. At the same time, it helps maintain the stable operation of the sewage treatment system, ensure that the effluent total nitrogen concentration remains stable and meets the standard, and improve the overall operating efficiency and effluent water quality of the sewage treatment plant.

[0041] (3) Strong adaptability: The model structure is flexible and can adapt to the operating data characteristics of different sewage treatment plants. By learning features at multiple time scales and variable dimensions, the model can better handle noise and uncertainty in the data, has strong adaptability to complex and changing sewage treatment conditions, and provides a universal prediction solution for sewage treatment plants of different scales and processes.

[0042] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0044] Figure 1 This is the overall framework diagram of the effluent total nitrogen concentration prediction model (M-DMCAN) proposed in the present invention;

[0045] Figure 2 It is the time scale feature extraction module;

[0046] Figure 3 It is a variable dimension feature extraction module;

[0047] Figure 4 The prediction effect of total nitrogen concentration in effluent before data preprocessing;

[0048] Figure 5 The prediction effect of total nitrogen concentration in effluent after data preprocessing;

[0049] Figure 6 Comparison of experimental results between M-DMCAN and variant models. DETAILED DESCRIPTION

[0050] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0051] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0052] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0053] See also Figures 1 to 6 The present invention provides a method for predicting the total nitrogen concentration in the effluent of a sewage treatment plant based on an attention mechanism, which specifically includes the following steps:

[0054] S1: Data Acquisition and Preprocessing: Obtain real data from the sewage treatment plant, including indicators such as carbon source addition flow rate, effluent total nitrogen concentration, influent total nitrogen concentration, influent flow rate, influent chemical oxygen demand, and influent nitrate nitrogen. Because real data may contain missing values, outliers, and inconsistent time scales, data preprocessing is required. Missing sampling points are filled using methods such as linear interpolation, the time scale of the dataset is aligned to ensure consistent sampling intervals, and outliers are processed to improve data quality and provide reliable data support for subsequent model training.

[0055] S2: M-DMCAN model construction and training: The prediction model of the present invention is built based on the PyTorch framework, and appropriate hyperparameters are set, such as the number of pooling layers, the number of attention heads, the number of cross-dimensional cross-attention scale layers, the learning rate, etc. The mean square error is used as the loss function, and the Adam algorithm is used to update the model parameters. During the training process, the data set is divided into a training set, a validation set, and a test set according to a certain ratio. The model is trained using the training set. In each iteration, the model calculates the predicted value based on the current parameters and compares it with the true value. The error is calculated through the loss function, and then the Adam algorithm is used to adjust the model parameters according to the error so that the loss function gradually decreases. At the same time, the model performance is evaluated using the validation set. If the evaluation indicators on the validation set (such as mean square error, mean absolute error, etc.) do not improve within multiple consecutive epochs, the training is stopped, the model parameters at this time are saved, and the optimal model is obtained.

[0056] S3. Model prediction: The pre-processed test data is input into the trained model. The model predicts the total nitrogen concentration of the effluent in the future based on the learned characteristics and rules, and outputs the predicted value to provide a decision-making basis for the carbon source addition control in the sewage treatment process.

[0057] Figure 1 This is the overall framework diagram of the M-DMCAN model, as shown Figure 1 As shown in Figure 1, the M-DMCAN model mainly consists of two parts: (1) univariate time scale feature extraction based on trends and cycles; (2) variable dimension cross-attention feature extraction. Specifically, the time feature extraction based on trends and cycles is used to extract the time scale characteristics of the time series itself, and the multivariate cross-attention is used to extract the variable dimension characteristics between variables. Finally, the extracted features are mapped to the predicted output of the total nitrogen concentration in the effluent through a multi-layer perceptron (MLP). The two parts work together to improve the prediction ability of the model. The symbols and definitions used in this invention are shown in Table 1.

[0058] Table 1

[0059] symbol definition <![CDATA[X,y t ]]> Historical key indicator series, predicted value corresponding to time t T,t Window size and sequence sampling points of historical key indicator series H The window size for predicting the sequence M,m Maximum number of downsampling layers and number of downsampling layers L,l The period and trend decomposition level of the sequence S,s A set of multi-scale sequences after sequence downsampling, a sequence of a certain scale after downsampling u,v Period and trend subsequences decomposed from the sequence <![CDATA[x i.d ]]> The i-th time segment after the time series of variable d is segmented <![CDATA[e i,d ]]> The embedding corresponding to the i-th time segment of variable d <![CDATA[E :,d ]]> The embeddings corresponding to all time segments of variable d <![CDATA[E i,: ]]> The embeddings corresponding to the i-th time segment of all variables <![CDATA[L seg ]]> Time segment length c Number of layers for crisscross attention learning α Cross-dimensional multi-scale cross-attention weight coefficient

[0060] In the prediction of time series data x, the core idea is to use the time window of past length T to achieve the prediction of the future length H. This paper focuses on the feature extraction of historical time scale in a single time variable series (CS_TN), and achieves this goal by decomposing and fusing the time series. The specific implementation process is detailed in Figure 2 .

[0061] First, a multi-time scale representation of the single target variable time series is constructed. :,y ∈R 1×TDownsampling is performed, and each downsampling time step is halved to obtain a set of multi-time scale series X y ={x0,x1,...,x M},in For multi-time scale sequences, the bottom sequence x0 is essentially equivalent to the original sequence x :,y , which contains the most detailed features; the highest level pooling layer sequence x M , providing the most macroscopic feature perspective of the time series. Then, the multi-time scale series X y Embedded into a unified feature space to obtain deep features

[0062] S 0 =W e X y (1)

[0063] in, is the mapping matrix.

[0064] Then, we stack L layers of time series decomposition modules and mixing modules to learn and mix the information of both period and trend features. is the l-th layer feature representation, where The specific operation of extracting the l layer is as follows:

[0065] (1) Cycle and trend feature decomposition module

[0066] Research has shown that data series have significant periodic and trend characteristics. In view of this, a decomposition module is used to process each scale series, breaking them down into periodic and trend components one by one, in order to avoid interference and errors that may arise when analyzing the series as a whole.

[0067] Specifically, for the lth pooling layer, the multi-scale time series is first decomposed into periodic components and trend components Expressed as:

[0068]

[0069] (2) Feature component mixing module

[0070] When exploring the periodic characteristics, it was found that a larger period can be regarded as a collection of multiple smaller periods, and periodic information is crucial for predicting future periodic fluctuations. Based on this concept, when performing periodic fusion operations, a hierarchical order from low to high is adopted, and the sequence details of the low-scale level are gradually integrated into the higher-scale level. In this way, richer and more detailed information is injected into the periodic modeling of the more macroscopic sequence at the high-scale level, thereby improving the completeness and accuracy of the modeling. Specifically, for the multi-scale periodic sequence collection For the lth scale, a bottom-up mixing layer (denoted as UM) is used to realize bottom-up periodic information interaction in the form of residual connection, which can be expressed as follows:

[0071]

[0072] Among them, UM(·) represents a periodic mixed operation, which includes two linear layers and GELU activation function, and its input dimension is T / 2 m-1 , the output dimension is T / 2 m .

[0073] When analyzing trend characteristics, it is found that the coarse-grained time series at the high-scale level can more intuitively present the macro trend direction. In contrast, the time series at the low-scale level is easily affected by factors such as noise, resulting in blurred trend characteristics. In view of this, unlike the periodic fusion strategy, this paper adopts a high-to-low fusion method, using the macro trend information at the high-scale level to provide guidance for the trend modeling at the low-scale level, so as to enhance the model's ability to capture the real trend. Similarly, for the multi-scale trend series set For the lth scale, a top-down mixing layer (denoted as VM) is used to realize top-down trend information interaction in the form of residual connection, which can be expressed as follows:

[0074]

[0075] Similarly, VM(·) represents a trend mixing operation, which consists of two linear layers and GELU activation function, and its input dimension is T / 2 m+1 , the output dimension is T / 2 m .

[0076] Finally, after the L-layer time series decomposition module and mixing module, the time dimension features containing the periodic characteristics and trend characteristics of the data itself are obtained.

[0077]

[0078] Among them, φ(·) contains two linear layers and the GELU activation function in the middle layer, which is used for feature enhancement. The GELU activation function is selected to increase the nonlinear expression of features and alleviate the problem of gradient disappearance.

[0079] The existing multi-dimensional time series prediction model based on the attention mechanism generally integrates the data points of each variable at the same time step into a vector, and mines the correlation of the time dimension through the self-attention mechanism. However, this method is difficult to directly reflect the intrinsic connection between variables, and has limitations when dealing with complex scenarios such as sewage treatment where variables have strong coupling. The present invention focuses on the correlation characteristics of multi-dimensional variables in the process of carbon source addition in sewage treatment. To this end, before extracting the variable dimension features, the multi-variable time series data is encoded using a patch-based segmentation strategy, and then the variable dimension features are extracted with the help of a multi-head attention mechanism, thereby effectively capturing the cross-correlation between variables. For details of the specific implementation process, see Figure 3 .

[0080] First, according to the length L seg The length of , all variables of the input data X are segmented according to the time step, so as to obtain a set of time segments, which is formalized as:

[0081]

[0082] Among them, x i,d Indicates that the length of the i-th variable in the d-th variable sequence is L seg data segment; for the task of predicting the total nitrogen concentration in effluent, the number of input variables D = 6.

[0083] The data segments are then encoded. This section uses a simple linear projection to encode the data segments and adds position information to facilitate subsequent attention calculation. The process is expressed as:

[0084]

[0085] Among them, the same mapping matrix W is used for time scale feature extraction e , is the learnable position information embedding corresponding to the position (i, d). Therefore, for the set of time series segments corresponding to Equation (7), the corresponding embedding set is obtained:

[0086]

[0087] Among them, e i,d is the embedding corresponding to the i-th time segment of the d-th variable sequence.

[0088] Then, multi-head attention is used to extract cross-correlation features between variables and residual connections to obtain variable dimension cross-correlation features, which can be specifically expressed as:

[0089]

[0090] E=Norm(E+MAS(Q,K,V)) (11)

[0091] Among them, Norm(·) is a common normalization operation; MAS(Q,K,V) represents a multi-head attention layer, Q, K, and V represent query, key, and value respectively, and all variables share the same MAS layer.

[0092] The E obtained through the above operations contains the cross-dependencies among different variables. Not only from the original e i,d , which also implies that any x i,d information.

[0093] Multi-feature prediction and aggregation, the time-scale feature sequence obtained after the input variable is processed by time-scale feature extraction The variable dimension feature sequence E obtained after the variable dimension feature extraction process is used as input to obtain the output through the multi-layer perceptron (MLP).

[0094] Specifically, the prediction results Ym from different scales m are aggregated to obtain the prediction result Y about the variable y S -pred , the time scale feature prediction and aggregation process is expressed as:

[0095]

[0096] Among them, Y m ∈R 1×H represents the prediction sequence of variable y at scale m; MLP(·) is a multi-layer perceptron for the prediction task.

[0097] The variable dimension feature sequence E obtained after the input variable is extracted is used to obtain the prediction output through the multi-layer perceptron (MLP). In the cross attention learning, according to the time series segment length L seg The different levels of attention learning can be divided into different levels. Generally, the length of the time series segments of adjacent scales is 2 times. First, the hierarchical prediction output Y is obtained. c Then, the prediction outputs of different levels are aggregated to obtain the final prediction output The variable dimension feature prediction and aggregation process is expressed as:

[0098]

[0099] Where c is the number of layers of cross-attention learning.

[0100] The prediction obtained by time scale feature extraction is Y S-pred , the prediction obtained by variable dimension feature extraction is Y M-pred . The final prediction is therefore given by:

[0101] Y=(1-α)·Y S-pred +αY M-pred (16)

[0102] Among them, α∈[0,1] is the cross-dimensional prediction weight that controls the cross attention.

[0103] In the prediction task of the present invention, the mean square error is used as the training loss function of the prediction model, that is:

[0104]

[0105] Among them, H is the prediction window size, Y true is the true value of the target variable.

[0106] Model training is performed by minimizing the error between the true and predicted values (arg minL(θ)). Among all possible parameters θ, the one that minimizes the loss function L(θ) is found. The Adam algorithm is used to update model parameters. The Adam optimizer combines momentum and adaptive learning rates to automatically adjust the learning rate of each parameter, making it particularly well-suited for the diverse gradient dynamics of multi-head attention and MLP parameters.

[0107] Experimental design of prediction model for total nitrogen concentration in effluent from sewage treatment process.

[0108] Within the framework of control theory, a carbon source dosing control system equipped with a predictive function offers significant advantages in improving efficiency. Its innovation lies in its ability to break through traditional control modes and make intelligent decisions based on the future dynamic evolution of the controlled object, effectively avoiding the lag inherent in traditional feedback control and achieving a qualitative leap in control accuracy. Therefore, the accuracy of the predictive model, like the system's "core engine," directly determines the performance of the entire control system and is crucial to the accuracy and reliability of carbon source dosing control. This study employed deep data analysis methods to comprehensively explore the trend characteristics, cyclical variation patterns, and coupling relationships between multiple variables to construct a prediction model for effluent total nitrogen concentration. This model is capable of real-time dynamic monitoring and accurate prediction of key parameters throughout the entire wastewater treatment process. By establishing an efficient and accurate feedback mechanism, it provides detailed data support and a solid theoretical basis for the subsequent design and optimization of the carbon source dosing predictive controller, driving wastewater treatment systems toward intelligent and refined operation models and helping the industry achieve the dual goals of green, low-carbon, and efficient treatment.

[0109] The data set used in this paper is from a real sewage treatment plant in Chongqing. An introduction to the relevant indicators of the data set is provided, and the data is analyzed. Based on this, a prediction model is designed to predict the effluent total nitrogen concentration in the next 15 minutes, 30 minutes, and 60 minutes using the historical 200-minute carbon source addition flow rate, effluent total nitrogen concentration, influent total nitrogen concentration, influent flow rate, influent chemical oxygen demand, and influent nitrate nitrogen.

[0110] To better reflect the predictive model's performance for wastewater treatment carbon source addition scenarios with long time delays, the dataset was linearly interpolated to fill in missing sampling points and the time scale of the dataset was aligned to ensure consistent sampling intervals. Training, validation, and test sets were then generated in a 7:1:2 ratio.

[0111] Evaluation indicators:

[0112] Three common evaluation indicators are used in this scheme: mean square error (MSE), mean absolute error (MAE) and root mean square error (RMSE) to comprehensively evaluate the prediction accuracy of the model to measure the performance of the prediction model.

[0113] Mean Squared Error (MSE) is one of the most commonly used loss functions for prediction tasks. It measures the average of the squares of the differences between the model's predictions and the true values. By calculating the square of the difference between each predicted value and the true value, MSE quantifies the overall bias of the model. Its calculation formula is as follows:

[0114]

[0115] Among them, y i is the model prediction value, y i is the true value, N represents the number of sample points in the prediction window, and the subsequent evaluation index symbols have the same meaning, so they are not introduced again.

[0116] Mean Absolute Error (MAE) measures the average absolute difference between the predicted value and the true value. Unlike MSE, MAE gives equal weight to the error of each sample and is more robust to outliers. Its formula is as follows:

[0117]

[0118] The root mean squared error (RMSE) is the square root of the mean squared error, which normalizes the MSE error metric to its original units. RMSE is slightly less sensitive to error than MSE and can directly reflect the degree of error in the model's predictions. Therefore, it is often used to evaluate the overall performance of the model. Its calculation formula is as follows:

[0119]

[0120] In general, MSE, MAE, and RMSE are commonly used metrics for evaluating predictive model performance. MSE has a stronger penalty for large errors, while MAE provides a more robust evaluation method for outliers. RMSE is a standardized version of MSE, using the same units as the original data, making it easier to intuitively interpret the magnitude of the error. Furthermore, MSE, MAE, and RMSE are all negatively correlated with the predictive performance of the model; that is, smaller values indicate better predictions.

[0121] Baseline method introduction: In order to observe the effectiveness of the model as a whole, it will be compared with related models. This section will introduce the control models involved in detail, including:

[0122] (1) TimeMixer, a time series prediction model based on a multi-scale fusion architecture, achieves time series prediction by decoupling the past information and future predictions of multi-scale time series. It uses a full MLP architecture and efficiently analyzes the periodicity and trend components, achieving state-of-the-art prediction results.

[0123] (2) CrossFormer, a deep learning model based on the Transformer architecture, can effectively capture the long-term and short-term dependencies in time series by encoding the position of time series and introducing a cross-scale attention mechanism, thereby achieving prediction of time series.

[0124] (3) Linear, a simple single-layer linear model, directly performs a linear transformation on the input time series data to generate forecast results. This method shares weights between different variables and does not model spatial correlation.

[0125] (4) DLinear, based on Linear, decomposes the time series into trend and residual series, uses two single-layer linear networks to model them respectively, and finally adds the prediction results of the two to generate the final prediction.

[0126] (5) NLinear takes data distribution shift into account based on Linear, and uses normalization and linear changes to enhance the robustness of the prediction model to distribution shift.

[0127] (6) Transformer, which applies the Transformer model to the field of time series prediction. It captures long-distance dependencies through the attention mechanism, thereby effectively processing complex patterns in time series data. It also shows great potential with its global modeling capabilities and flexible architecture.

[0128] (7) Autoformer, through deep decomposition architecture, effectively separates the trend and periodic components of time series. It uses the autocorrelation mechanism based on the periodicity of time series and discovers the dependency between subsequences through sequence-level connections, which effectively improves the prediction effect of time series with periodicity and trend.

[0129] (8) Informer, based on Transform, introduces the ProbSparse self-attention mechanism and halves the length of the input sequence of the cascade layer, making the model more suitable for the prediction of long time series. It also uses encoders of different time scales to better capture the long-term and short-term dependencies in the time series.

[0130] (9) LSTM, a special recurrent neural network, mainly solves the gradient vanishing and gradient exploding problems faced by recurrent neural networks when processing long sequence data. It controls the flow of information by introducing a gating mechanism, thereby effectively capturing long-term dependencies.

[0131] Model parameter details. The model was implemented in PyTorch 1.12.0 and experiments were conducted on the following hardware platforms: an Intel Xeon E5-2678 CPU and an NVIDIA GeForce RTX 3090 GPU. To facilitate comparison and enhance the reliability of the experiments, the model configuration remained the same as the baseline model, except for necessary hyperparameters. The prediction input length was 40 (200 minutes), and the prediction output lengths were 3, 6, and 12 (15 minutes, 30 minutes, and 60 minutes). Average pooling was used with 3 pooling layers. The prediction layer had 1 MLP hidden layer and 64 dimensions. The number of attention heads was 2, and the number of cross-dimensional attention scales was 2. The model was trained using a mini-batch method with a uniform batch size of 128 and a learning rate of 0.01. The cross-dimensional multi-scale cross-attention weight coefficient α was set to 0.4. It is worth noting that the experiment used an early stopping strategy during training; if the MAE on the validation set did not improve within 5 epochs, training was terminated.

[0132] Experimental results analysis:

[0133] Due to problems such as sensors being blocked by impurities in sewage during real sewage treatment, data may be missing or abnormal. Therefore, in order to further evaluate the impact of data preprocessing on model performance, this paper compares the model performance with and without preprocessing on a real sewage dataset (WWTP). Data preprocessing includes steps such as linear interpolation to fill missing values, time alignment, and outlier processing. Table 2 and Figure 4 、 Figure 5 The experimental results and the visual comparison between the predicted values and the true values are shown.

[0134] The results in Table 2 show that across all prediction tasks, the model with data preprocessing significantly outperformed the unpreprocessed model in terms of MAE, MSE, and RMSE. Data preprocessing enabled the M-DMCAN model to more accurately capture short-term dynamics and reduce prediction errors. This demonstrates that data preprocessing plays a crucial role in capturing long-term dependencies and trends in real-world wastewater treatment processes.

[0135] Table 2

[0136]

[0137] from Figure 4 The analysis results show that when the data is not preprocessed, the M-DMCAN model has difficulty accurately matching the dramatic fluctuations in the data in multiple data subintervals (such as the marked areas in the figure). In particular, there is a significant deviation between the model's predicted values and the actual values at the data extremes. A deeper investigation reveals that this is mainly due to the presence of missing values and outliers in the original data. These data defects interfere with the model's learning of the data's intrinsic distribution characteristics, resulting in a decrease in the model's prediction accuracy and an inability to effectively capture data trends.

[0138] Figure 5 Experimental results show that after data preprocessing, the M-DMCAN model's ability to fit the data is significantly improved. In particular, at extreme data points and in areas of significant fluctuation (as marked in the figure), the consistency between the model's predicted values and the true values is significantly improved. This phenomenon clearly demonstrates that data preprocessing effectively improves data quality by filling in missing data and removing outliers. The processed data provides the model with higher-quality learning samples, enabling it to more accurately capture and learn the time series characteristics of the data, thereby improving forecasting performance.

[0139] Experimental research results confirm that data preprocessing is indispensable for enhancing the overall performance of the M-DMCAN model. Specifically, by using data completion technology to fill missing values, applying anomaly detection algorithms to remove outliers, and standardizing the time scale, the reliability and consistency of the data are significantly improved. These optimization measures have built a better training foundation for the model, enabling the M-DMCAN model to more accurately capture the temporal correlation characteristics between data when processing actual sewage treatment data. Whether it is dealing with short-term rapid fluctuation predictions or long-term trend evolution analysis, the model has demonstrated better prediction accuracy and stability, effectively improving the reliability of key parameter predictions in the sewage treatment process.

[0140] To verify the effectiveness of the effluent total nitrogen concentration prediction model, an overall performance comparison was conducted with relevant SOTA models on a real sewage dataset. The experimental results are shown in Table 3.

[0141] Table 3

[0142]

[0143] The experimental data in Table 3 clearly demonstrates that the M-DMCAN model demonstrates outstanding performance across various forecasting tasks, maintaining a leading edge or approaching optimal levels in core evaluation metrics such as mean squared error (MSE), mean absolute error (MAE), and root mean squared error (RMSE). Notably, in the 60-minute forecasting task, M-DMCAN achieves an MSE as low as 0.1935, significantly outperforming other compared models and highlighting its strong forecasting accuracy. While models such as TimeMixer, CrossFormer, DLinear, and NLinear only narrowly surpass M-DMCAN in the 15-minute short-term forecasting task, their performance gradually declines as the forecast duration increases. In contrast, M-DMCAN demonstrates superior stability and accuracy in long-term forecasting scenarios, fully demonstrating its superiority in handling complex time series relationships. In comparison, traditional Linear, Transformer, and LSTM models significantly lag behind in forecasting performance, particularly in long-term forecasting tasks, where their error values significantly exceed those of advanced models such as M-DMCAN. This series of experimental results strongly confirms that M-DMCAN can deeply explore the complex dynamic characteristics in long-time lag data and accurately capture the underlying laws of time series, providing a more reliable and practical prediction solution for practical application scenarios such as carbon source addition control in sewage treatment processes.

[0144] To further explore the specific contributions of each model component to performance, this study constructed three variant models and conducted multi-dimensional prediction experiments. By systematically removing core model modules, we quantitatively evaluated the impact of each component on overall performance. The designed variant models were: removing the trend and cycle feature extraction module (without DM), removing the cross-attention module (without CAN), and removing the multi-scale feature learning module (without M). Figure 6 The visualization results of the ablation experiment are presented intuitively. Table 4 provides a detailed comparative analysis of the performance indicators of each variant model from the perspectives of mean square error (MSE), mean absolute error (MAE), and root mean square error (RMSE). The specific conclusions are as follows:

[0145] When the trend and period feature extraction module (without DM) is removed, model performance declines significantly. Experimental data shows that across all forecasting tasks, both the MSE and MAE metrics show a significant increase. In particular, in the 60-minute long-term forecasting task, the MSE value climbs to 0.2878 and the MAE reaches 0.3764, significantly reducing forecast accuracy. This phenomenon demonstrates that the trend and period feature extraction module plays a key role in capturing the periodic patterns and long-term evolution trends of time series. Without this module, the model struggles to effectively learn the dynamic patterns of change in time series, resulting in a significant increase in forecast error.

[0146] In experiments with the Cross-Attention Module removed (without CAN), model performance degraded even more dramatically, particularly in long-term prediction scenarios. For a 60-minute prediction task, for example, the MSE value soared to 0.3146, and the MAE reached 0.3888. A deeper analysis reveals that the Cross-Attention Module, as a core component for modeling complex dependencies between multidimensional variables, effectively captures high-order interactions between features from different dimensions. However, removing this module deprives the model of its ability to exploit deep connections within multidimensional data, leading to a sharp decline in prediction performance.

[0147] In contrast, removing the multi-scale feature learning module (without M) has a relatively mild impact on model performance. Despite a decrease in overall performance, this variant model performs relatively close to the full model in short-term forecasting tasks. This phenomenon is attributed to the fact that the multi-scale feature learning module is mainly used to integrate feature information of different time scales and variable dimensions to capture the dynamic changes of time series at different time granularities. In short-term forecasting scenarios, feature information at a single scale can already meet the model's forecasting needs, while in long-term forecasting tasks, multi-scale feature learning helps the model capture data features more comprehensively, thereby significantly improving forecasting performance.

[0148] Table 4

[0149]

[0150] Ablation experiments clearly reveal that the trend and period feature extraction module and the cross-attention module are key determinants of model performance. Removing these two key modules significantly increases the model's mean squared error (MSE) and mean absolute error (MAE), particularly in long-term forecasting tasks, where the prediction error jumps significantly. This phenomenon demonstrates that the trend and period feature processing module is the cornerstone for accurately capturing the long-term evolution of time series, while the cross-attention module is the core hub for constructing complex dependencies between multidimensional variables. Together, they play a decisive role in the model's prediction accuracy and stability. In stark contrast, the multi-scale feature learning module exhibits significant scenario-specific impact on model performance. In short-term forecasting tasks, even without this module, the model maintains good prediction performance thanks to its other components. This demonstrates that the core value of the multi-scale feature learning module lies particularly in long-term forecasting scenarios. By integrating feature information from different time granularities and variable dimensions, it provides the model with richer contextual awareness, effectively improving the accuracy of processing complex time series data.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for predicting total nitrogen concentration in effluent from a sewage treatment plant based on an attention mechanism, characterized in that: The method specifically comprises the following steps: S1: Data acquisition and preprocessing: Obtain sewage treatment data, including carbon source addition flow rate, effluent total nitrogen concentration, influent total nitrogen concentration, influent flow rate, influent chemical oxygen demand, and influent nitrate. Preprocess the data to obtain a data set, which is divided into a training set and a validation set. S2: Model construction and training: Construct a prediction model for effluent total nitrogen concentration, including a time-scale feature extraction module, a variable-dimensional feature extraction module, and a multi-feature prediction and aggregation module; train the model using the training set, evaluate model performance using the validation set, and save the optimal model parameters; The time scale feature extraction module is used to extract the time scale features of the time series itself, including downsampling the input data to construct a multi-time scale series, embedding it into a unified feature space, and then learning and mixing the information of the two features of cycle and trend by stacking the decomposition module and the mixing module of the time series; The variable dimension feature extraction module is used to extract variable dimension features between variables, including using a patch-based segmentation method to segment all variables of the input data by time step, linearly projecting the data segments and adding position information, and then using multi-head attention to extract cross-correlation features and residual connections between variables; The multi-feature prediction and aggregation module maps the features processed by the time scale feature extraction module and the variable dimension feature extraction module into the prediction output of the total nitrogen concentration in the effluent through a multi-layer perceptron, and then performs weighted aggregation to obtain the final prediction result; S3: Model prediction: Input the test data into the trained effluent total nitrogen concentration prediction model to obtain the predicted value of effluent total nitrogen concentration in the future.

2. The method for predicting total nitrogen concentration in effluent from a sewage treatment plant based on an attention mechanism according to claim 1, characterized in that: In step S2, the specific implementation steps of the time scale feature extraction module are: S201: Downsample the input data, halving the time step each time, to obtain a set of multi-time-scale sequences, where the bottom-level sequence is equivalent to the original sequence, and the highest-level pooling layer sequence provides the most macroscopic feature perspective; then, embed the multi-time-scale sequences into a unified feature space; S202: By stacking multiple layers of time series decomposition modules and mixing modules, time dimension features containing the periodicity and trend characteristics of the data itself are obtained, specifically including: In the cycle and trend feature decomposition module, each scale series is decomposed into a periodic component and a trend component. The periodic component is calculated by the Avgpool function, and the trend component is obtained by subtracting the periodic component from the original series. In the feature component mixing module, the periodic information is mixed in a bottom-up order through residual connections, and the trend information is mixed in a top-down order through residual connections.

3. The method for predicting total nitrogen concentration in effluent from a sewage treatment plant based on an attention mechanism according to claim 2, characterized in that: In step S202, the time dimension features including the periodicity and trend features of the data itself are obtained through the decomposition module and the mixing module of the L-layer time series. Among them, UM(·) represents the mixed operation of the period, which includes two linear layers and the GELU activation function; VM(·) represents the mixed operation of the trend, which includes two linear layers and the GELU activation function; φ(·) includes two linear layers and the GELU activation function of the middle layer; M represents the maximum number of sampling layers, m represents the number of downsampling layers, l represents the number of trend decomposition layers, and L represents the period of the sequence. represents the mth periodic component of the lth pooling layer; Represents the mth trend component of the lth pooling layer.

4. The method for predicting total nitrogen concentration in effluent from a sewage treatment plant based on an attention mechanism according to claim 1, characterized in that: In step S2, the specific implementation steps of the variable dimension feature extraction module are: S211: Based on the Patch segmentation method, all variables of the input data are segmented according to the time step according to a specific length to obtain a set of time segments; S212: Encode the data segment using linear projection and add position information thereto, wherein the mapping matrix used is shared with the time scale feature extraction module; S213: Multi-head attention is used to extract the cross-correlation features between variables, and through residual connection, the variable dimension cross-correlation features containing the cross-dependency relationship between different variables are obtained.

5. The method for predicting total nitrogen concentration in effluent from a sewage treatment plant based on an attention mechanism according to claim 4, characterized in that: In step S213, the expression of the variable dimension cross-correlation feature E is obtained as follows: E=Norm(E+MAS(Q,K,V)) Among them, Norm(·) is the normalization operation; MAS(Q,K,V) represents the multi-head attention layer, Q, K, V represent the query, key and value respectively, and all variables share the same MAS layer; E represents the embedding set, Among them, e i,d is the embedding of the i-th time segment corresponding to the d-th variable sequence, D is the input variable, L seg is the time segment length; T is the sequence window size.

6. The method for predicting total nitrogen concentration in effluent from a sewage treatment plant based on an attention mechanism according to claim 1, characterized in that: In step S2, the specific implementation steps of the multi-feature prediction and aggregation module are: S221: During the time scale feature prediction and aggregation process, the prediction results from different scales are normalized and processed by a multi-layer perceptron before being aggregated. S222: During the variable dimension feature prediction and aggregation process, hierarchical cross-attention learning is performed according to the different lengths of the time series segments, and the prediction outputs at different levels are aggregated to obtain the final prediction output; S223: The time scale feature prediction results and the variable dimension feature prediction results are weighted and aggregated to obtain the final predicted value of the effluent total nitrogen concentration. The weight coefficient is adjusted in the range of [0,1] according to the actual situation.

7. The method for predicting total nitrogen concentration in effluent from a sewage treatment plant based on an attention mechanism according to claim 6, characterized in that: In step S223, the final predicted value Y of the effluent total nitrogen concentration is expressed as: Y=(1-α)·Y S-pred +αY M-pred Among them, α∈[0,1] is the cross-dimensional prediction weight that controls cross attention; Y S-pred is the predicted value obtained by time scale feature extraction, and the expression is in, represents the mth time dimension feature with a sequence period of L, MLP represents a multi-layer perceptron, and Norm(·) represents a normalization operation; Y M-pred is the predicted value obtained by extracting variable dimension features, and the expression is Among them, c is the number of layers of cross attention learning, Y c is the hierarchical prediction output, Among them, E is the cross-correlation feature of variable dimension, y i is the predicted value corresponding to the i-th time segment, H is the window size of the prediction sequence, L seg is the length of the time segment, d is the d-th variable sequence, and D is the input variable.

8. The method for predicting total nitrogen concentration in effluent from a sewage treatment plant based on an attention mechanism according to claim 1, characterized in that: In step S2, model training specifically includes: using the training set to build a water outlet total nitrogen concentration prediction model based on the PyTorch framework, setting model hyperparameters, using mean square error as the loss function, and using the Adam algorithm to update the model parameters. During the training process, the validation set is used to evaluate the model performance and save the optimal model parameters.