Small sample sewage soft measurement method based on optimized minimum count summary

By combining the CMS algorithm optimized by attenuation factor and the LSTM-Attention model, the frequency and time domain characteristics in the sewage water quality data are extracted and utilized, and the problem of low prediction accuracy in the small sample data environment is solved, achieving higher robustness and adaptability.

CN120123733APending Publication Date: 2025-06-10BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510198630.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-23
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing sewage water quality monitoring methods have low prediction accuracy in small sample data environments, and traditional models are difficult to effectively utilize the frequency domain characteristics in the data, resulting in unstable prediction results in dynamically changing environments.

Method used

A method combining the minimum count summary (CMS) algorithm based on attenuation factor optimization and a deep learning model (LSTM-Attention) is used to extract frequency domain features and combine time domain features to enhance the data expression ability and improve the robustness and adaptability of the model in a small sample environment.

Benefits of technology

It significantly improves the accuracy and robustness of wastewater quality prediction, and can provide reliable prediction results in small samples and dynamic data flows, suitable for complex and high noise environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123733A_ABST
    Figure CN120123733A_ABST
Patent Text Reader

Abstract

The invention relates to a sewage treatment water quality prediction method based on an optimized minimum count summary algorithm and a deep learning model, and aims to improve the accuracy of water quality monitoring and reduce the monitoring cost. Along with the acceleration of industrialization and urbanization processes, the problems of water pollution and water resource shortage are more and more serious. In order to improve the utilization efficiency of water resources, sewage recovery and treatment are particularly critical. In the sewage treatment process, water quality monitoring data provides an objective evaluation standard for the treatment process, but due to monitoring period and cost differences of different water quality indexes, a traditional direct measurement method faces great challenges. According to the method, the minimum count summary (Count Min Sketch, CMS) algorithm and the deep learning model are combined, the optimized minimum count summary algorithm is used for extracting frequency domain information, and the frequency domain information serves as features to be input into the long short-term memory network (LSTM)-Attention model, so that the precision of water quality prediction is remarkably improved. The method not only improves the prediction precision of the model, but also is efficient and low in cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of environmental monitoring and data modeling, and specifically to a small-sample sewage data feature enhancement method based on an optimized Count-Min Sketch (CMS) algorithm and deep learning, which is used for soft-sensor modeling and prediction of sewage water quality. Background Art

[0002] With the exacerbation of environmental problems globally, the shortage of water resources and water pollution problems are becoming increasingly severe. Sewage treatment has become one of the key means to ensure the sustainable utilization of water resources. Sewage treatment not only helps to protect water resources, improve environmental quality, but also plays a crucial role in maintaining ecological balance and promoting human health. To cope with the increasingly severe water pollution challenges, improve sewage treatment efficiency and reduce operating costs, it is particularly important to accurately predict key water quality indicators (such as biochemical oxygen demand BOD, chemical oxygen demand COD, etc.). Indicators such as BOD and COD are usually used to evaluate the concentration of organic pollutants in water bodies and can provide reliable control basis for the sewage treatment process. Therefore, real-time and accurate prediction of these key indicators can not only optimize the sewage treatment process, but also improve the economy and efficiency of water quality management.

[0003] However, most traditional sewage water quality monitoring methods rely on laboratory tests and expensive sensing devices. Although these methods can provide relatively accurate water quality data, due to the complex test process and long time period, the real-time performance is poor, and they cannot meet the needs of rapid decision-making in a dynamically changing environment. In addition, the high cost of traditional measurement devices, the complexity of equipment maintenance, and the instability of monitoring accuracy also limit their popularization in practical applications. Especially in scenarios with tight resources and high real-time requirements, the limitations of traditional methods are more obvious.

[0004] In recent years, with the rapid development of machine learning and deep learning technologies, data-driven soft-sensor methods have been widely applied in the field of sewage treatment. Soft-sensor technology predicts key water quality indicators that are difficult to directly measure by establishing a mathematical model with the help of easily measurable auxiliary variables (such as flow rate, temperature, pH value, etc.). Compared with traditional methods, soft-sensor methods can significantly reduce monitoring costs and improve the ability of real-time monitoring. Especially in dealing with complex and dynamically changing environments, they have great application potential. This makes water quality monitoring more flexible and efficient in practical applications, especially suitable for sewage treatment systems that require rapid decision-making and adjustment.

[0005] However, existing data-driven soft sensing methods usually rely on large-scale, high-quality training data, while in actual sewage treatment applications, water quality data often presents characteristics such as sparse samples, high noise and data imbalance. These characteristics make the prediction accuracy of traditional soft sensing models low in small sample data environments, especially when data is scarce or there are abnormal fluctuations. The model is prone to overfitting or deviation, making it difficult to obtain reliable prediction results. Therefore, how to improve the robustness of soft sensing technology in small sample data environments by optimizing algorithms and model design in the case of insufficient samples has become a technical problem that needs to be solved urgently in the field of sewage treatment.

[0006] In addition, existing soft sensing methods usually focus on the temporal characteristics in time series data, but make insufficient use of the potential frequency domain characteristics in the data. In sewage quality prediction, water quality parameters are not only affected by time changes, but may also be significantly affected by periodic fluctuations or frequency characteristics. For example, changes in auxiliary variables such as temperature and pH value may show a certain periodic pattern, which is difficult to reflect in time domain analysis, but can be effectively extracted and expressed through frequency domain analysis. Existing models mainly rely on traditional time series data processing methods, which often ignore the potential frequency domain characteristics in the data, thereby limiting the performance and prediction accuracy of the model. Therefore, how to develop a small sample data enhancement method that can combine time domain and frequency domain characteristics at the same time to improve the prediction accuracy and robustness of the model under small sample conditions has become a technical problem that needs to be solved urgently in the current sewage treatment field. Summary of the invention

[0007] The present invention proposes a sewage water quality prediction method based on a minimum count summary optimized by a decay factor combined with a deep learning model, aiming to solve the problem of predicting key water quality indicators in the sewage treatment process. The small sample problem and high noise interference faced by the prior art in water quality prediction usually lead to low prediction accuracy, especially in an environment with large dynamic changes in time series data. The present invention not only improves the prediction accuracy of the model by introducing the optimized CMS algorithm, namely Decay CMS, and the LSTM-Attention model in deep learning, but also improves the robustness and adaptability in the case of small samples and dynamic data streams.

[0008] The uniqueness of the present invention lies in: First, it optimizes the frequency-domain features by combining the Decay CMS technology, and dynamically adjusts the influence of historical data on the model using the decay factor; Second, it adopts a data enhancement method that combines chi-square binning and one-hot encoding, effectively enhancing the expression ability of the frequency-domain features; Finally, through the LSTM-Attention network, by combining time-domain and frequency-domain features, it significantly improves the prediction performance of the model in a small-sample environment. Compared with traditional soft sensing methods, the present invention provides a more accurate, low-cost, and adaptable sewage water quality prediction solution. The technical solution of the present invention includes the following steps:

[0009] The technical solution of the present invention includes the following core parts: the Decay CMS module, the frequency-domain data processing module, and the water quality prediction module based on LSTM-Attention.

[0010] S1: Decay CMS module

[0011] The main function of this module is to extract the frequency-domain features of sewage water quality data, helping the model capture the periodic features and dynamic change trends of the data. Through the optimized Decay CMS algorithm, while ensuring the real-time performance and adaptability of the model, it can effectively reduce the interference of outdated data. The introduced decay factor can dynamically adjust the influence of historical data, ensuring the efficient processing of real-time data by the model.

[0012] S2: Frequency-domain data processing module

[0013] The function of the frequency-domain data processing module is to enhance the expression ability of the frequency-domain features through data processing techniques such as chi-square binning and one-hot encoding. Chi-square binning helps divide the data into different intervals, maximizing the feature differences between intervals, and converting the frequency-domain data into a sparse matrix through one-hot encoding, facilitating model learning. On this basis, combined with the features of time-domain data, the final enhanced feature set is formed, enabling the model to fully utilize frequency-domain and time-domain information for prediction.

[0014] S3: Water quality prediction module based on LSTM-Attention

[0015] The function of this module is to process the time-series data through the LSTM network, capture the long-term dependencies and dynamic changes in the sewage water quality data, and combine the Attention mechanism to improve the attention to key time-step data. LSTM can effectively avoid the problem of gradient disappearance in traditional neural networks when processing time-series data, while the Attention mechanism helps the model focus on important time nodes, improving the prediction accuracy. The design of this module ensures that even in a small-sample and high-noise environment, the model can provide accurate and stable prediction results.

[0016] The present invention constructs an enhanced dataset based on the frequency-domain features extracted from the optimized Decay CMS and the original time-series features, and uses this dataset to predict and model the biochemical oxygen demand (BOD) of the target water quality index. Specifically, the method proposed in the present invention is compared with the model based only on time-series features, and the Adam algorithm is used for all models to optimize the parameters to ensure the efficiency and stability of the training process.

[0017] Furthermore, the dataset used is sourced from the open-source sewage water quality dataset on the UCI Machine Learning Repository website. This dataset records the monitoring data of different treatment processes in a sewage treatment plant from January 1990 to August 1991, covering a total of 527 days of data for different treatment processes, and each sample represents the observation results of one day. The dataset contains 38 attributes, and all attributes are continuous numerical values. To ensure the relevance of model training, the present invention selects the biochemical oxygen demand, i.e., BOD, in the second treatment tank as the target variable for soft sensing.

[0018] In the process of selecting auxiliary variables, the present invention first adopts the Pearson correlation coefficient. By analyzing the correlation between different variables and the target variable BOD, six auxiliary variables that are most relevant to the target variable are screened out. Specifically, through this method, it is ensured that the selected auxiliary variables have a strong correlation with the prediction task, thus providing more reliable data support for the subsequent training of the model.

[0019] In the data preprocessing process, the present invention first conducts outlier detection on the dataset and deletes the abnormal data that does not conform to the actual situation. Further, for the missing values in the dataset, the median interpolation method is used for filling to ensure the integrity of the data. To ensure the elimination of the influence of different attribute dimensions, the dataset is also normalized, mapping all data values to the interval [0, 1]. Finally, the dataset is divided into a training set and a test set in the ratio of 8:2 in chronological order from front to back, where 80% of the data is used for model training and 20% of the data is used for model performance evaluation.

[0020] In terms of model evaluation, the present invention uses the root mean square error (RMSE) as the evaluation index. Specifically, RMSE reflects the average value of the squared prediction errors. Further, the smaller the value of RMSE, the higher the prediction accuracy of the model, which helps to comprehensively evaluate the performance of the model in the water quality prediction task.

[0021] In the experimental comparison, the model based on temporal features in the present invention was compared with the optimized Decay CMS model. Specifically, the model based on temporal features was trained only using the original temporal data, while the optimized Decay CMS model was trained by combining temporal and frequency domain features. The experimental results show that the optimized Decay CMS model has significantly improved in evaluation metrics such as RMSE and MAE compared with the model that only relies on temporal features. By introducing the Decay CMS algorithm, the model can effectively extract frequency domain features, and through the LSTM-Attention mechanism, it combines temporal and frequency domain information, further improving the accuracy of water quality prediction.

[0022] Further analysis of the experimental results reveals that the optimized Decay CMS model shows higher robustness and adaptability in dealing with small sample data and dynamic changes. Specifically, the optimized Decay CMS model performs more excellently on the test set compared with the traditional model based on temporal features, with a significantly reduced RMSE, verifying the effectiveness of the method of the present invention. Compared with the model that only relies on temporal features, the optimized Decay CMS model can capture the trend of water quality changes more accurately. Especially in a complex and dynamically changing environment, it can better adapt to the changes in water quality data, thus providing more accurate predictions.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] (1) Improved the accuracy of water quality prediction: By combining the Decay CMS optimization algorithm with the LSTM-Attention model, the present invention can utilize both temporal and frequency domain features, significantly improving the accuracy of water quality prediction. Compared with the traditional model that only relies on temporal data, the optimized model can capture the changing trend of water quality data more comprehensively. Especially when dealing with temporal data with long-term dependence relationships, it shows higher accuracy.

[0025] (2) Effectively solved the small sample problem: By introducing the Decay CMS optimization algorithm, the present invention can still provide reliable prediction results in a small sample data environment. Decay CMS uses the decay factor to dynamically adjust the influence of historical data, enabling the model to pay more attention to the features of the current temporal data, reducing the interference of historical data, and enhancing the robustness in the case of small samples.

[0026] (3)Enhanced model adaptability and robustness: The optimized Decay CMS can adapt to data changes under different environmental conditions by extracting frequency-domain features and combining time-domain data. Especially in the sewage water quality monitoring with large dynamic changes, it can accurately capture the changing trends of water quality indicators and provide more stable and accurate predictions. This method makes the model of the present invention more adaptable and robust than traditional models when dealing with complex and high-noise environments.

[0027] (4)Efficient feature extraction and processing methods: The present invention enhances the expression ability of frequency-domain features through data processing methods such as chi-square binning and one-hot encoding, enabling the model to utilize frequency-domain information more effectively. This feature enhancement method not only improves the learning ability of the model but also reduces the computational complexity, showing high efficiency in practical applications.

[0028] (5)Applicable to dynamic and variable data environments: By introducing the LSTM-Attention mechanism, the present invention can handle long-term dependencies in time-series data. At the same time, the Attention mechanism focuses on data at key time steps, further enhancing the modeling ability for dynamically changing data. Especially when processing real-time water quality monitoring data, it can adapt to changes in different water quality conditions in real time and provide efficient and accurate predictions. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is the overall process framework diagram of the present invention;

[0030] Wherein:

[0031] 1 Data preprocessing: Normalize the original sewage water quality data to eliminate the dimensional differences between different variables;

[0032] 2 Optimize the CMS algorithm: Dynamically update the counter by introducing a decay factor to enhance the frequency-domain feature extraction ability;

[0033] Feature discretization and encoding: Discretize the frequency-domain features using the 3 chi-square binning algorithm and generate a sparse feature matrix using 4 one-hot encoding;

[0034] 5 Feature fusion: Concatenate the frequency-domain features and time-domain features to form an enhanced feature set;

[0035] 6 Prediction model construction: Input the enhanced features into a deep learning model composed of LSTM and Attention to obtain 7 model prediction values.

[0036] Figure 2 is the processing flow chart for optimizing the minimum count sketch algorithm;

[0037] Wherein:

[0038] A two-dimensional array of 1w*d

[0039] 2 Time-domain data input at the current moment

[0040] 3 Time-domain data input at the next moment

[0041] 4 Decay factor

[0042] Figure 3 is the structural diagram of the LSTM-Attention model;

[0043] Among them:

[0044] 1 The combined dataset including time domain and frequency domain input

[0045] 2 LSTM layer

[0046] 3 Attention mechanism layer

[0047] 4 Fully connected layer

[0048] 5 Target water quality index Specific implementation manner

[0049] The technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] Based on the restrictive problem of the small sample dataset, the present invention proposes an optimized soft sensor framework, which combines the frequency domain feature extraction ability of the Count-Min Sketch (CMS) algorithm and the time series modeling ability of the LSTM-Attention deep learning model, and can effectively improve the accuracy and robustness of sewage water quality prediction.

[0051] As Figure 1 shown, the soft sensor framework of the present invention mainly includes the following steps:

[0052] 1. Select auxiliary variables and data preprocessing: Select the auxiliary variable with the largest correlation with BOD using the pearson coefficient, normalize the original sewage water quality data, map the data values to the interval [0, 1], and eliminate the dimensional differences between different variables. At the same time, use the linear interpolation method to supplement the missing values to ensure the integrity of the input data.

[0053] 2. Optimize the CMS algorithm: The present invention introduces a decay factor to optimize the CMS algorithm, as Figure 2As shown, by dynamically adjusting the counter value, the weights of historical data and current data are better balanced, enhancing the ability to capture the frequency-domain characteristics of time-series data.

[0054] 3. Frequency-domain feature extraction and processing: The extracted frequency-domain features are discretized by the chi-square binning algorithm, and the binned features are further transformed into a sparse matrix form through one-hot encoding, providing high-quality input for the deep learning model.

[0055] 4. Feature fusion: The frequency-domain features are concatenated with the original time-domain features to form an enhanced feature set, providing more comprehensive data input for the model.

[0056] 5. Construction of the LSTM-attention prediction model: As Figure 3 shown, the enhanced feature set is input into the LSTM-Attention model for prediction.

[0057] The LSTM module captures the long-term dependence characteristics of the data, and the Attention module dynamically adjusts the feature weights to strengthen the expression of key features. Finally, the predicted value is output through the fully connected layer.

[0058] (1) Selection of auxiliary variables and data preprocessing

[0059] The target of this water quality prediction is BOD. Using the Pearson correlation coefficient method, water quality parameters (auxiliary variables) that are correlated with the prediction target are selected from the original dataset.

[0060] The formula for the Pearson correlation coefficient is:

[0061]

[0062] where: r is the Pearson correlation coefficient, representing the strength of the linear relationship between two variables X and Y. X i and Y i are the i-th observations of variables X and Y, respectively. and are the means of variables X and Y, respectively. n is the number of samples.

[0063] The relevant information of the finally selected auxiliary variables is shown in Table 1

[0064]

[0065] The values of the selected auxiliary variables are normalized, and the formula is as follows:

[0066]

[0067] where, x max and x minThey are the minimum and maximum values of water quality variables respectively. x(t) is the actual value of water quality data. The normalized water quality data forms the time domain part X of the data set. time .

[0068] (2) Optimize the Count-Min Sketch algorithm (CMS)

[0069] The Count-Min Sketch is a space-efficient data structure based on hashing technology to approximately estimate the frequency of elements in the data. Multiple hash functions are used to map each element in the data to different positions in a matrix. Whenever an element appears, the counter value is updated in different hash table rows through the corresponding hash function. Finally, for each element, the estimated value of the frequency is obtained by querying the counter values of all rows corresponding to the element and taking the minimum value.

[0070] The size of the matrix is w*d, where w represents the length of the array. The value of w is determined by the required error ∈. If the error is ∈, then w is usually set to:

[0071]

[0072] where e is the base of the natural logarithm.

[0073] d represents the width of the array, and its value is usually related to the required hash collision probability δ. To make the probability of hash collision small enough, the selection of d is based on:

[0074]

[0075] where δ is the allowed hash collision probability, usually a very small value, such as 0.01.

[0076] The value of w is 8 and the value of d is 5.

[0077] When optimizing CMS to update the frequency estimation, a decay factor α is introduced. For each counter update, the decay factor is applied to the weight of the historical data, so that the frequency contribution of the older data gradually weakens, thereby reducing the error caused by the change of the data stream. The update formula of the optimized CMS is:

[0078] C i (t + 1) = α·C i (t) + Δ i (t)

[0079] where, C i (t + 1) is the value of the i-th counter at time t + 1, C i (t) is the value of the i-th counter at time t, and Δ i(t) is the new increment of the current time-series data, obtained through the update of the hash function. α is the decay factor. As an optimization method, its value range is (0, 1). The selection of α should be adjusted according to the characteristics of the data stream and the requirements of the application scenario. Specifically, if the data stream is relatively static, a smaller value of the decay factor can be set; if the data stream changes frequently, the decay factor should be set to a larger value to ensure that the influence of new data on the estimation result is increased.

[0080] For the specific value of α, different values of α can be tested through experiments to observe the performance of the model under different decay factors. When using this dataset for the experiment, the values of α are set to 1, 0.85, and 0.7 respectively for model training. When α = 0.85, the RMSE and MAE values of the model are the smallest, and the accuracy of the model is the highest.

[0081] (3) Frequency-domain feature extraction and processing

[0082] ① Use chi-square binning to divide the frequency-domain features into m intervals. Chi-square binning maximizes the differences between intervals and minimizes the differences within intervals, thereby extracting important distribution patterns in the data. This operation converts the features from the original continuous values into intervals with more obvious distinguishability. The formula for chi-square binning is:

[0083]

[0084] where O i and E i are the observed frequency and the expected frequency of the i-th interval respectively. The value of m is determined by experiments or cross-validation. χ 2 is to maximize the chi-square statistic to guide the value of the number of intervals m. The m when χ 2 is the largest is the optimal number of intervals.

[0085] ② One-hot encode the chi-square binning results to generate a sparse vector X freq of the frequency-domain features for subsequent model input;

[0086] X freq = [x 1 , x 2 , …, x i , … x m

[0087] where m is the total number of bins (i.e., the number of intervals); x i represents the i-th position, and x m∈ {0, 1}; In the one-hot encoded vector, only one position is 1 and the rest are 0; After chi-square binning, the data is processed by one-hot encoding, and each interval is transformed into an m-dimensional sparse vector. This kind of sparse vector has higher discriminability, can capture more details of the data in subsequent model training, and prevent possible order relationships between different features.

[0088] (4) Feature fusion

[0089] Concatenate the enhanced frequency-domain features and the original time-domain features column-wise to form the feature set F as the model input:

[0090] F = [X time , X freq

[0091] (5) Prediction model based on LSTM-attention network

[0092] Input the feature set into a deep learning model composed of a long short-term memory network and an attention mechanism to predict the sewage water quality index; The data set is first input into the LSTM model. The number of samples BS for each training is 32, indicating that each sample has 3 time points in the sequence, that is, the time step T, and the number of data features IS at each time step is 30. The LSTM model has 3 layers, each layer includes 256 hidden units, and the activation function uses the tanh function; The output expression is as follows:

[0093] f t = σ(W f · [x t , h t-1 + b f )

[0094] i t = σ(W i · [x t , h t-1 + b i )

[0095]

[0096] o t = σ(W o · [x t , h t-1 + b o )

[0097] h t = o t * tanh(C t )

[0098] Among them, f t , i t and ot represent the input gate, forget gate, and output gate; σ represents the Sigmoid function; Wf and bf represent the weights and biases of the forget gate respectively, Wi and bi represent the weights and biases of the input gate respectively, Wo and bo represent the weights and biases of the output gate respectively; h t-1 represents the hidden layer state, x t represents the input at the current time step. They pass through a tanh layer in the forget gate to obtain a new candidate value W c and b c represent the weights and biases of the candidate layer respectively; C t represents the output after the cell state is updated; C t - 1 represents the cell state at the previous time step, h t represents the final output. Before the start of training, all weight and bias terms are initialized by Xavier initialization.

[0099] The Attention mechanism performs a weighted sum of the hidden states output by the LSTM to generate a weighted context vector. The specific calculation process is as follows:

[0100] a t = Attention(h) = softmax(W a h t )

[0101] where a is the Attention weight, representing the weighting coefficient for each time step, W a is a trainable parameter matrix, which is optimized through backpropagation during training. W a maps the hidden state at each time step to a new space, enabling the model to calculate the attention weights for each time step. The softmax function normalizes W a h t . The output hidden state of the LSTM is weighted and summed using a to generate the final context vector c:

[0102]

[0103] where T is the sequence length, a t is the attention weight at the t-th time step, and h t is the hidden state at the t-th time step.

[0104] Finally, the weighted context vector c is input into the fully connected layer for the final prediction. The expression is:

[0105]

[0106] Among them, W is the weight matrix of the fully connected layer, b is the bias term, and the final output is the predicted value of the model, which is expressed as:

[0107]

[0108] Among them, y 1 , y 2 , …, y k are the respective dimensions of the predicted value.

[0109] During the training process, an early stopping strategy is used, with the tolerance set to 20. When the RMSE of the loss function on the validation set has not improved in 20 consecutive iterations, the training is stopped.

[0110]

[0111] where n is the number of predicted values, is the predicted value of the i-th sample, and y i is the true value of the sample.

[0112] To verify the effectiveness of the optimized CMS method and the LSTM-Attention model proposed in the present invention, the following three deep learning models were selected for experimental comparison: Temporal Convolutional Network (TCN), LSTM, and the LSTM-Attention model proposed in the present invention. These models were trained and tested under two conditions: without using the CMS method and using the CMS method. The test results are shown in Table 1.

[0113] Table 1 Test Set Results

[0114]

[0115] As shown in Table 1, among the three experimental models, using the CMS method significantly improved the test performance of the models. This indicates that the CMS method proposed in the present invention can effectively enhance the expression ability of frequency domain features, thereby improving the prediction accuracy. Among all the models, the LSTM-Attention model combined with the CMS method has the best prediction performance, and its RMSE (Root Mean Square Error) and MAE (Mean Absolute Error) are significantly lower than those of other models, further verifying the effectiveness and superiority of the soft sensor framework of the present invention.

Claims

1. A small sample sewage soft sensing method based on optimizing the minimum counting summary is characterized by: The following steps are involved: 1) Selection of auxiliary variables and data preprocessing: The Pearson coefficient is used to select the auxiliary variable with the greatest correlation with BOD, and the raw sewage water quality data is normalized to map the data values ​​to the [0,1] interval to eliminate the dimensional differences between different variables; at the same time, linear interpolation is used to supplement the missing values ​​to ensure the integrity of the input data; 2) Optimize the CMS algorithm: Introduce a decay factor to optimize the CMS algorithm; balance the weights of historical data and current data by dynamically adjusting the counter value; 3) Frequency domain feature extraction and processing: The extracted frequency domain features are discretized using the chi-square binning algorithm, and the binned features are further converted into a sparse matrix form through one-hot encoding; 4) Feature fusion: concatenate frequency domain features with original time domain features to form an enhanced feature set; 5) Construction of prediction model based on LSTM-attention; The enhanced feature set is input into the LSTM-Attention model for prediction; the LSTM module captures the long-term dependency characteristics of the data, and the Attention module dynamically adjusts the feature weights to strengthen the expression of key features, and finally outputs the predicted value through the fully connected layer; The details are as follows (1) Selection of auxiliary variables and data preprocessing The target of water quality prediction is BOD. The Pearson correlation coefficient method is used to select water quality parameters that are correlated with the prediction target from the original data set, namely auxiliary variables; The formula for the Pearson correlation coefficient is: Where: r is the Pearson correlation coefficient, which indicates the strength of the linear relationship between two variables X and Y; X i and Y i are the i-th observation values ​​of variables X and Y respectively; and are the means of variables X and Y respectively; n is the sample size; The auxiliary variables finally selected are as follows The values ​​of the selected auxiliary variables are normalized, and the formula is as follows: Among them, x max and x min are the minimum and maximum values ​​of the water quality variables respectively; x(t) is the actual value of the water quality data; the normalized water quality data forms the time domain part of the data set X time ; (2) Optimizing the Minimum Count Summary Algorithm (CMS) Use multiple hash functions to map each element in the data to different positions in a matrix; whenever an element appears, update the counter value in different hash table rows through the corresponding hash function; finally, for each element, the frequency estimate is obtained by querying the counter values ​​of all rows corresponding to the element and taking the minimum value; The size of the matrix is ​​w*d, where w represents the length of the array; the value of w is determined by the required error ∈; if the error is ∈, then w is usually set to: Where e is the base of natural logarithms; d represents the width of the array, and its value is usually related to the required hash collision probability δ; in order to make the probability of hash collision small enough, d is selected based on: Among them, δ is the allowed hash collision probability; w is 8, d is 5; When optimizing CMS, an attenuation factor α is introduced when updating the frequency estimation; the updating formula of optimizing CMS is: C i (t+1)=α·C i (t)+Δ i (t) Among them, C i (t+1) is the value of the i-th counter at time t+1, C i (t) is the value of the i-th counter at time t, Δ i (t) is the increment of the current time series data, which is obtained by updating the hash function; α is the attenuation factor, which is used as an optimization method and has a value range of (0, 1); Frequency domain feature extraction and processing ① Use chi-square binning to divide the frequency domain features into m intervals. Chi-square binning maximizes the differences between intervals and minimizes the differences within intervals, thereby extracting important distribution patterns in the data; this operation converts the features from the original continuous values ​​to intervals with more obvious distinction; the formula for chi-square binning is: Among them, O i and E i are the observed frequency and expected frequency of the ith interval respectively; the value of m is determined by experiment or cross-validation, and χ 2 To maximize the chi-square statistic, the value of the number of guidance intervals m, χ 2 The maximum value of m is the optimal number of intervals; ② Perform unique hot encoding on the chi-square binning results to generate a sparse vector X of frequency domain features freq , used for subsequent model input; X freq =[x1,x2,…,x i ,…x m ] Where m is the total number of bins (i.e. the number of intervals); x i represents the i-th position, x m ∈{0, 1}; only one position in the one-hot encoded vector is 1, and the rest are all 0; the data after chi-square binning is processed by one-hot encoding, and each interval is converted into an m-dimensional sparse vector; this sparse vector has higher discrimination, can capture more details of the data in subsequent model training, and prevent possible sequential relationships between different features; (3) Feature Fusion The enhanced frequency domain features are concatenated with the original time domain features to form a feature set F as the model input: F=[X time ,X freq ] (4) Prediction model based on LSTM-attention network The feature set is input into a deep learning model composed of a long short-term memory network and an attention mechanism to predict sewage quality indicators. The data set is first input into the LSTM model. The number of samples BS in each training is 32, which means that the sequence of time points in each sample is 3, that is, the time step T, and the number of data features in each time step IS is 30. The LSTM model has 3 layers, each layer includes 256 hidden units, and the activation function uses the tanh function. The output expression is as follows: f t =σ(W f ·[x t ,h t-1 ]+b f ) i t =σ(W i ·[x t ,h t-1 ]+b i ) the t =σ(W o ·[x t ,h t-1 ]+b o ) h t =o t *tanh(C t ) Among them, f t ,i t and t represents the input gate, forget gate and output gate; σ represents the Sigmoid function; W f and b f Denote the weight and bias of the forget gate, Wi and bi denote the weight and bias of the input gate, Wo and bo denote the weight and bias of the input gate, respectively; h t-1 represents the hidden layer state, x t Represents the input at the current moment. They pass through a tanh layer in the forget gate to obtain a new candidate value W c and b c Represent the weight and bias of the candidate layer respectively; C t Represents the output after the cell state is updated; C t-1 Indicates the cell state at the last moment, h t Represents the final output; before training begins, all weights and bias terms are initialized by Xavier; The Attention mechanism performs weighted summation on the hidden states output by LSTM to generate a weighted context vector. The specific calculation process is as follows: a t =Attention(h)=softmax(W a h t ) Among them, a is the Attention weight, which represents the weighted coefficient of each time step, W a is a trainable parameter matrix, which is optimized by back propagation during training. a The hidden state of each time step is mapped to a new space so that the model can calculate the attention weight of each time step. The softmax function converts W a h t Normalize; use a to perform weighted summation on the output hidden state of LSTM to generate the final context vector c: Among them, T is the sequence length, a t is the attention weight at the tth time step, h t is the hidden state at the tth time step; Finally, the weighted context vector c is input into the fully connected layer for final prediction, expressed as: Among them, W is the weight matrix of the fully connected layer, b is the bias term, and the final output is the predicted value of the model, expressed as: Among them, y1,y2,…,y k For each dimension of the predicted value; Use the early stopping strategy during training, set the tolerance to 20, and stop training when the RMSE loss function of the validation set does not improve in 20 consecutive iterations; Where n is the number of predicted values, is the predicted value of the i-th sample, y i is the true value of the sample.

Citation Information

Cited By

  • Method and system for predicting sewage treatment capacity of environmental protection industry based on big data

    CN120852124A