Regenerated Water Quality Prediction Method and System Based on Deep Learning Algorithm
By constructing a convolutional neural network model and using deep learning algorithms to process sewage water quality and water quality purification processing data, it solves the problem that traditional methods are difficult to accurately predict complex nonlinear water quality changes, and achieves higher accuracy and robust water quality prediction.
Patent Information
- Application Number
- CN202411114101.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-08-14
AI Technical Summary
Traditional water quality prediction methods are difficult to accurately predict water quality changes in complex nonlinear systems, and cannot fully capture the dynamic characteristics of water quality changes.
The water quality prediction method for regenerated water based on deep learning algorithm is adopted, and the water quality prediction method is used to construct a convolutional neural network model, and the feature selection and pre-processing of wastewater water quality raw data, water quality purification processing data and regenerated water quality detection data are used for feature selection and pre-processing, and the deep learning model is trained for water quality prediction.
It improves the accuracy of water quality prediction, can fully capture the complex relationship between various factors affecting water quality, and enhances the robustness and adaptability of the model.
Smart Images

Figure CN118861531B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of water quality monitoring, and particularly relates to a reclaimed water quality prediction method and system based on a deep learning algorithm. Background Art
[0002] Reclaimed water (water that has been treated to meet specific water quality standards and can be reused within a certain range) is an important water resource, and its water quality safety is directly related to the ecological environment and human health. However, the water quality of reclaimed water is affected by various factors, including sewage water quality, treatment processes, equipment operating conditions, etc. The relationships between these factors are complex and difficult to accurately predict using traditional methods. Therefore, it is of great significance to develop an efficient and accurate reclaimed water quality prediction method.
[0003] Most traditional water quality prediction methods are based on statistical analysis or physical models. These methods have limitations in dealing with complex nonlinear systems and are difficult to comprehensively capture the dynamic characteristics of water quality changes. In recent years, the rapid development of deep learning technology has provided new ideas for water quality prediction. Deep learning models have powerful feature learning capabilities and nonlinear modeling capabilities, and can automatically extract useful information from massive data, providing a new solution for reclaimed water quality prediction. Summary of the Invention
[0004] The present invention provides a reclaimed water quality prediction method and system based on a deep learning algorithm, aiming to use a neural network to build a model and apply it to reclaimed water quality prediction to solve the problem that traditional methods cannot give accurate predictions.
[0005] On the one hand, the present invention provides a reclaimed water quality prediction method based on a deep learning algorithm, and the method includes the following steps:
[0006] Step S1, respectively obtain the original sewage water quality data, water quality purification treatment data, and reclaimed water quality detection data of reclaimed water from the factory sewage outlet, reclaimed water treatment equipment, and the reclaimed water samples after treatment, form a multi-dimensional data group, and select a characteristic data group that affects water quality prediction from the multi-dimensional array.
[0007] Step S2, preprocess the data of the characteristic data group, the preprocessing includes outlier processing and normalization processing, and use the preprocessed characteristic data group to train a deep learning model to obtain model parameters.
[0008] Step S3, use the trained model to obtain the water quality prediction result of reclaimed water.
[0009] Furthermore, the original sewage water quality data includes 9 dimensions: sewage chemical oxygen demand, total nitrogen in sewage, total phosphorus in sewage, sewage suspended solids, biochemical oxygen demand in sewage, sewage temperature, sewage pH value, sewage turbidity, and sewage conductivity.
[0010] Further, the water quality purification treatment data includes 4 dimensions: sedimentation time, aeration volume, filter rate, and backwashing frequency.
[0011] Further, the reclaimed water quality detection data includes 9 dimensions: chemical oxygen demand of reclaimed water, total nitrogen of reclaimed water, total phosphorus of reclaimed water, suspended solids of reclaimed water, biological oxygen demand of reclaimed water, temperature of reclaimed water, pH value of reclaimed water, turbidity of reclaimed water, and conductivity of reclaimed water.
[0012] Further, the multi-dimensional data set is represented as matrix X n : X n =(X w , X c , X z ), where each row represents the data at a time point, each column represents a specific water quality index or treatment parameter, X w represents the original sewage water quality data, X c represents the water quality purification treatment data, X z represents the reclaimed water quality detection data, n represents the number of time windows, then the dimension of matrix X n is equal to 22n.
[0013] Further, the feature data set is represented as matrix X nt : X nt =(X wt , X ct , X zt ), X wt represents the feature data of the original sewage water quality data, X ct represents the feature data of the water quality purification treatment data, X zt represents the feature data of the reclaimed water quality detection data, the total number of feature data is k, 6≥k≥3, then the dimension of matrix X nt is equal to kn.
[0014] Further, by calculating the mutual information between the original sewage water quality data or the water quality purification treatment data and the reclaimed water quality detection data, the feature data set affecting water quality prediction is selected from the multi-dimensional array:
[0015]
[0016] Among them, I(x; y) represents the mutual information value between two data, x∈X, X represents the original sewage water quality data or the water quality purification treatment data, y∈Y, Y represents the reclaimed water quality detection data; P(x,y) is the joint probability distribution, and P(x) and P(y) are the marginal probability distributions.
[0017] Select k feature data from high to low according to the magnitude of the comparative mutual information value to form a feature data group.
[0018] Furthermore, the outlier processing is determined by calculating the standard score of each data point:
[0019] Among them, q is the data point, μ is the mean of the data, and σ is the standard deviation of the data; when |z| is greater than the set threshold, the data point is considered an outlier, and the detected outlier is processed, and the outlier is replaced with the mean or median.
[0020] The min-max method is used for data normalization: Among them, q is the original data, q′ is the normalized data, q min and q max are the minimum and maximum values of the data, respectively.
[0021] Furthermore, the deep learning model includes a convolutional neural network, and the neural network includes: an input layer, a graph convolutional layer, a feature aggregation layer, a fully connected layer, and an output layer.
[0022] The input layer receives the preprocessed feature data group and initializes it as node features, and the initialization is expressed as: Among them, represents the initial feature of node i, and x i represents the node feature of node i.
[0023] The graph convolutional layer uses heterogeneous graph convolution operations to aggregate and update node features of different types, and the function formula of the graph convolutional layer is:
[0024]
[0025] Among them, represents the feature of node i at the l-th layer, represents the set of neighbor nodes of node i, and c ij is the normalization coefficient, and W (l) and are the weight matrices of the l-th layer, and ReLU is the activation function of the graph convolutional layer;
[0026] The feature aggregation layer aggregates the features of each node to generate a new node feature representation, and the function formula of the feature aggregation layer is:
[0027]
[0028] Among them, represents the feature of node i at the last layer, and Aggregate represents the feature aggregation function;
[0029] The fully connected layer inputs the aggregated node features into the fully connected layer for feature transformation and non-linear transformation. The functional formula of the fully connected layer is as follows:
[0030]
[0031] where y i represents the predicted output delay value of node i, W fc and b fc are the weights and biases of the fully connected layer, and softmax is the activation function of the fully connected layer;
[0032] The output layer outputs the predicted value y i of the reclaimed water quality detection data of the node.
[0033] Furthermore, the training process of the deep learning model includes the following steps:
[0034] Step S21: Divide the preprocessed feature data set into a training set and a validation set. The training set is used for model training, and the validation set is used for model validation.
[0035] Step S22: Evaluate the model using the cross-validation method and adjust the model parameters according to the evaluation results. Specifically, it includes: dividing the training set into several subsets for k-fold cross-validation; in each fold, select one subset as the validation set, and the remaining subsets as the training set, train the model and evaluate its performance on the validation set; calculate the average validation error of all folds as the evaluation index of the model.
[0036] Step S23: Evaluate the performance of the model on the validation set, select the best model parameters, and retrain the model using the entire training set to obtain the parameters of the final deep learning model.
[0037] On the other hand, the present invention provides a reclaimed water quality prediction system based on a deep learning algorithm. The system includes: a data acquisition module, a model training module, and a water quality prediction module.
[0038] Furthermore, the data acquisition module is used to respectively obtain the original sewage water quality data, water quality purification treatment data, and reclaimed water quality detection data of reclaimed water from the factory sewage outlet, reclaimed water treatment equipment, and the treated reclaimed water samples, form a multi-dimensional data set, and select the feature data set that affects water quality prediction from the multi-dimensional data set.
[0039] Furthermore, the model training module is used to preprocess the data of the feature data set. The preprocessing includes outlier processing and normalization processing, and uses the preprocessed feature data set to train the deep learning model to obtain model parameters.
[0040] Further, the water quality prediction module is used to obtain the water quality prediction result of reclaimed water by using the trained model.
[0041] Further, the original sewage water quality data includes 9 dimensions: sewage chemical oxygen demand, total nitrogen in sewage, total phosphorus in sewage, sewage suspended solids, sewage biological oxygen demand, sewage temperature, sewage pH value, sewage turbidity, and sewage conductivity.
[0042] The water quality purification treatment data includes 4 dimensions: sedimentation time, aeration volume, filter rate, and backwashing frequency;
[0043] The reclaimed water quality detection data includes 9 dimensions: reclaimed water chemical oxygen demand, total nitrogen in reclaimed water, total phosphorus in reclaimed water, reclaimed water suspended solids, reclaimed water biological oxygen demand, reclaimed water temperature, reclaimed water pH value, reclaimed water turbidity, and reclaimed water conductivity.
[0044] The multi-dimensional data group is represented as matrix X n : X n =(X w , X c , X z ), where each row represents the data at a time point, each column represents a specific water quality index or treatment parameter, X w represents the original sewage water quality data, X c represents the water quality purification treatment data, X z represents the reclaimed water quality detection data, and n represents the number of time windows. Then the dimension of matrix X n is equal to 22n;
[0045] The feature data group is represented as matrix X nt : X nt =(X wt , X ct , X zt ), X wt represents the feature data of the original sewage water quality data, X ct represents the feature data of the water quality purification treatment data, X zt represents the feature data of the reclaimed water quality detection data. The total number of feature data is k, 6≥k≥3. Then the dimension of matrix X nt is equal to kn.
[0046] By calculating the mutual information between the original sewage water quality data or the water quality purification treatment data and the reclaimed water quality detection data, the feature data group affecting water quality prediction is selected from the multi-dimensional array:
[0047]
[0048] Among them, I(x; y) represents the mutual information value between two data, where x ∈ X, X represents the original sewage water quality data or the water quality purification treatment data, y ∈ Y, and Y represents the reclaimed water quality detection data; P(x, y) is the joint probability distribution, and P(x) and P(y) are the marginal probability distributions.
[0049] According to the comparison of the magnitudes of the mutual information values, k characteristic data are selected from high to low to form a characteristic data group.
[0050] Furthermore, the outlier processing is determined by calculating the standard score of each data point:
[0051] Among them, q is the data point, μ is the mean of the data, and σ is the standard deviation of the data. When |z| is greater than the set threshold, the data point is considered an outlier, and the detected outlier is processed by replacing it with the mean or median.
[0052] The min-max method is used for data normalization: Among them, q is the original data, q′ is the normalized data, q min and q max are the minimum and maximum values of the data, respectively.
[0053] Furthermore, the deep learning model includes a convolutional neural network, and the neural network includes: an input layer, a graph convolutional layer, a feature aggregation layer, a fully connected layer, and an output layer;
[0054] The input layer receives the preprocessed characteristic data group and initializes it as node features, and the initialization is expressed as: Among them, represents the initial feature of node i, and x i represents the node feature of node i;
[0055] The graph convolutional layer uses heterogeneous graph convolution operations to aggregate and update node features of different types, and the function formula of the graph convolutional layer is:
[0056]
[0057] Among them, represents the feature of node i at the l-th layer, represents the set of neighbor nodes of node i, c ij is the normalization coefficient, W (l) and are the weight matrices of the l-th layer, and ReLU is the activation function of the graph convolutional layer;
[0058] The feature aggregation layer aggregates the features of each node to generate a new node feature representation, and the function formula of the feature aggregation layer is:
[0059]
[0060] Among them, represents the feature of node i in the last layer, and Aggregate represents the feature aggregation function;
[0061] The fully connected layer inputs the aggregated node features into the fully connected layer for feature transformation and non-linear transformation. The function formula of the fully connected layer is:
[0062]
[0063] where y i represents the predicted output delay value of node i, W fc and b fc are the weights and biases of the fully connected layer, and softmax is the activation function of the fully connected layer; the output layer outputs the predicted value y i of the reclaimed water quality detection data of the node.
[0064] Furthermore, the training process of the deep learning model includes the following steps:
[0065] Step S21: Divide the preprocessed feature data set into a training set and a validation set. The training set is used for model training, and the validation set is used for model validation.
[0066] Step S22: Evaluate the model using the cross-validation method and adjust the model parameters according to the evaluation results. Specifically, it includes: dividing the training set into several subsets for k-fold cross-validation; in each fold, select one subset as the validation set, and the remaining subsets as the training set, train the model and evaluate its performance on the validation set; calculate the average validation error of all folds as the evaluation index of the model.
[0067] Step S23: Evaluate the performance of the model on the validation set, select the best model parameters, and retrain the model using the entire training set to obtain the parameters of the final deep learning model.
[0068] Compared with the prior art, the beneficial effects of the present invention are:
[0069] The present invention learns and models the multi-dimensional data set through a deep learning model, can comprehensively capture the complex relationships between various factors affecting water quality, and improve the accuracy of water quality prediction; through feature selection and outlier processing, further improve the robustness and adaptability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 is a flowchart of the reclaimed water quality prediction method based on the deep learning algorithm of the present invention;
[0071] Figure 2 It is a schematic diagram of the module composition of a reclaimed water quality prediction system based on a deep learning algorithm. Specific implementation manners
[0072] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0073] Embodiment 1
[0074] As Figure 1 shown, it is a flowchart of a reclaimed water quality prediction method based on a deep learning algorithm, and the method includes the following steps:
[0075] Step S1, respectively obtain the original sewage water quality data, water quality purification treatment data, and reclaimed water quality detection data of reclaimed water from the factory sewage outlet, reclaimed water treatment equipment, and treated reclaimed water samples, form a multi-dimensional data group, and select a characteristic data group affecting water quality prediction from the multi-dimensional array.
[0076] The original sewage water quality data includes 9 dimensions: sewage chemical oxygen demand, total nitrogen in sewage, total phosphorus in sewage, sewage suspended solids, biological oxygen demand in sewage, sewage temperature, sewage pH value, sewage turbidity, and sewage conductivity.
[0077] The water quality purification treatment data includes 4 dimensions: sedimentation time, aeration volume, filter rate, and backwashing frequency.
[0078] The reclaimed water quality detection data includes 9 dimensions: reclaimed water chemical oxygen demand, total nitrogen in reclaimed water, total phosphorus in reclaimed water, reclaimed water suspended solids, biological oxygen demand in reclaimed water, reclaimed water temperature, reclaimed water pH value, reclaimed water turbidity, and reclaimed water conductivity.
[0079] The multi-dimensional data group is represented as matrix X n : X n =(X w , X c , X z ), where each row represents the data at a time point, each column represents a specific water quality index or treatment parameter, X w represents the original sewage water quality data, X c represents the water quality purification treatment data, X z represents the reclaimed water quality detection data, n represents the number of time windows, and the dimension of matrix X n is equal to 22n.
[0080] The characteristic data group is represented as a matrix X nt : X nt =(X wt , X ct , X zt ), where X wt represents the characteristic data of the original sewage water quality data, X ct represents the characteristic data of the water quality purification treatment data, X zt represents the characteristic data of the reclaimed water quality detection data. The total number of characteristic data is k, and 6≥k≥3. Then the dimension of the matrix X nt is equal to kn.
[0081] By calculating the mutual information between the original sewage water quality data or the water quality purification treatment data and the reclaimed water quality detection data, a characteristic data group affecting water quality prediction is selected from the multi-dimensional array:
[0082]
[0083] Among them, I(x; y) represents the mutual information value between two data, x∈X, X represents the original sewage water quality data or the water quality purification treatment data, y∈Y, Y represents the reclaimed water quality detection data; P(x, y) is the joint probability distribution, and P(x) and P(y) are the marginal probability distributions.
[0084] According to the comparison of the magnitudes of the mutual information values, k characteristic data are selected from high to low to form a characteristic data group.
[0085] Step S2: Preprocess the data of the characteristic data group. The preprocessing includes outlier processing and normalization processing, and use the preprocessed characteristic data group to train a deep learning model to obtain model parameters.
[0086] The outlier processing is judged by calculating the standard score of each data point:
[0087] Among them, q is the data point, μ is the mean of the data, and σ is the standard deviation of the data; when |z| is greater than the set threshold, the data point is considered an outlier, and the detected outlier is processed, and the outlier is replaced by the mean or median.
[0088] The minimum-maximum method is used for data normalization: Among them, q is the original data, q′ is the normalized data, q min and q max are the minimum and maximum values of the data respectively.
[0089] The deep learning model includes a convolutional neural network, and the neural network includes: an input layer, a graph convolutional layer, a feature aggregation layer, a fully connected layer, and an output layer.
[0090] The input layer receives the preprocessed feature data group and initializes it as node features, and the initialization is expressed as: Where, represents the initial feature of node i, and x i represents the node feature of node i.
[0091] The graph convolutional layer uses heterogeneous graph convolutional operations to aggregate and update node features of different types. The function formula of the graph convolutional layer is:
[0092]
[0093] Where, represents the feature of node i at the l-th layer, represents the set of neighbor nodes of node i, and c ij is the normalization coefficient, and W (l) and are the weight matrices of the l-th layer, and ReLU is the activation function of the graph convolutional layer;
[0094] The feature aggregation layer aggregates the features of each node to generate a new node feature representation. The function formula of the feature aggregation layer is:
[0095]
[0096] Where, represents the feature of node i at the last layer, and Aggregate represents the feature aggregation function;
[0097] The fully connected layer inputs the aggregated node features into the fully connected layer for feature transformation and non-linear transformation. The function formula of the fully connected layer is:
[0098]
[0099] Where, y i represents the predicted value of the output delay of node i, and W fc and b fc are the weights and biases of the fully connected layer, and softmax is the activation function of the fully connected layer;
[0100] The output layer outputs the predicted value y i .
[0101] The training process of the deep learning model includes the following steps:
[0102] Step S21: Divide the preprocessed feature data set into a training set and a validation set. The training set is used for model training, and the validation set is used for model validation.
[0103] Step S22: Evaluate the model using the cross-validation method and adjust the model parameters according to the evaluation results. Specifically, it includes: dividing the training set into several subsets for k-fold cross-validation; in each fold, select one subset as the validation set and the remaining subsets as the training set, train the model and evaluate its performance on the validation set; calculate the average validation error of all folds as the evaluation metric of the model.
[0104] Step S23: Evaluate the performance of the model on the validation set, select the best model parameters, and retrain the model using the entire training set to obtain the parameters of the final deep learning model.
[0105] Step S3: Use the trained model to obtain the water quality prediction results of the reclaimed water.
[0106] For example, for a sewage treatment plant, the goal is to predict the water quality indicators of the treated reclaimed water to ensure that it meets the environmental protection standards. The following data is collected: raw sewage water quality data (recorded hourly, including 9 dimensions): sewage chemical oxygen demand (COD), sewage total nitrogen (TN), sewage total phosphorus (TP), sewage suspended solids (SS), sewage biological oxygen demand (BOD), sewage temperature (Temp), sewage pH value (pH), sewage turbidity (Turbidity), sewage conductivity (Conductivity); water quality purification treatment data (recorded hourly, including 4 dimensions): settling time (Settling Time), aeration volume (Aeration), filter rate (Filter Rate), backwash frequency (Backwash Frequency); reclaimed water quality detection data (recorded hourly, including 9 dimensions): reclaimed water chemical oxygen demand (COD); reclaimed water total nitrogen (TN); reclaimed water total phosphorus (TP); reclaimed water suspended solids (SS); reclaimed water biological oxygen demand (BOD); reclaimed water temperature (Temp); reclaimed water pH value (pH); reclaimed water turbidity (Turbidity); reclaimed water conductivity (Conductivity).
[0107] The above data is obtained from the sewage outlet of the sewage treatment plant, the reclaimed water treatment equipment, and the treated reclaimed water samples, and a multi-dimensional data group matrix is formed. Each row represents the data at a time point, and each column represents a specific water quality index or treatment parameter; by calculating the mutual information between the original sewage water quality data or water quality purification treatment data and the reclaimed water quality detection data, the following 6 characteristics are selected from the multi-dimensional data group as the characteristic data groups that affect water quality prediction: sewage chemical oxygen demand (COD), total phosphorus in sewage (TP), suspended solids in sewage (SS), aeration volume (Aeration), filter rate (Filter Rate), total phosphorus in reclaimed water (TP).
[0108] After training the neural network, the following data is obtained at a certain moment: sewage chemical oxygen demand (COD): 300 mg / L, total phosphorus in sewage (TP): 5 mg / L, suspended solids in sewage (SS): 100 mg / L, aeration volume (Aeration): 50 m 3 / h, filter rate (Filter Rate): 10 m 3 / h, total phosphorus in reclaimed water (TP): 1 mg / L; these data are input into the trained deep learning model, and the model outputs the predicted reclaimed water quality index. Through this method, the water quality of reclaimed water can be predicted in advance to ensure that it meets the environmental protection standards, so as to make corresponding adjustments and optimizations.
[0109] Example 2
[0110] As Figure 2 shown, it is a schematic diagram of the module composition of a reclaimed water quality prediction system based on a deep learning algorithm. The system includes: a data acquisition module, a model training module, and a water quality prediction module.
[0111] The data acquisition module is used to respectively obtain the original sewage water quality data, water quality purification treatment data, and reclaimed water quality detection data of reclaimed water from the factory sewage outlet, reclaimed water treatment equipment, and treated reclaimed water samples, form a multi-dimensional data group, and select the characteristic data groups that affect water quality prediction from the multi-dimensional array.
[0112] The model training module is used to preprocess the data of the characteristic data group. The preprocessing includes outlier processing and normalization processing, and uses the preprocessed characteristic data group to train a deep learning model to obtain model parameters.
[0113] The water quality prediction module is used to obtain the water quality prediction result of reclaimed water using the trained model.
[0114] The original sewage water quality data includes 9 dimensions: sewage chemical oxygen demand, sewage total nitrogen, sewage total phosphorus, sewage suspended solids, sewage biological oxygen demand, sewage temperature, sewage pH value, sewage turbidity, and sewage conductivity.
[0115] The water quality purification treatment data includes 4 dimensions: sedimentation time, aeration volume, filter rate, and backwashing frequency;
[0116] The reclaimed water quality detection data includes 9 dimensions: reclaimed water chemical oxygen demand, reclaimed water total nitrogen, reclaimed water total phosphorus, reclaimed water suspended solids, reclaimed water biological oxygen demand, reclaimed water temperature, reclaimed water pH value, reclaimed water turbidity, and reclaimed water conductivity.
[0117] The multi-dimensional data set is represented as matrix X n : X n =(X w , X c , X z ), where each row represents the data at a time point, each column represents a specific water quality index or treatment parameter, X w represents the original sewage water quality data, X c represents the water quality purification treatment data, X z represents the reclaimed water quality detection data, and n represents the number of time windows. Then the dimension of matrix X n is equal to 22n;
[0118] The characteristic data set is represented as matrix X nt : X nt =(X wt , X ct , X zt ), X wt represents the characteristic data of the original sewage water quality data, X ct represents the characteristic data of the water quality purification treatment data, X zt represents the characteristic data of the reclaimed water quality detection data. The total number of characteristic data is k, 6≥k≥3. Then the dimension of matrix X nt is equal to kn.
[0119] By calculating the mutual information between the original sewage water quality data or the water quality purification treatment data and the reclaimed water quality detection data, the characteristic data set affecting water quality prediction is selected from the multi-dimensional array:
[0120]
[0121] Among them, I(x; y) represents the mutual information value between two data, where x ∈ X, X represents the original sewage water quality data or the water quality purification treatment data, y ∈ Y, and Y represents the reclaimed water quality detection data; P(x, y) is the joint probability distribution, and P(x) and P(y) are the marginal probability distributions.
[0122] According to the magnitude comparison of the mutual information values, k characteristic data are selected from high to low to form a characteristic data group.
[0123] The outlier processing is determined by calculating the standard score of each data point:
[0124] Among them, q is the data point, μ is the mean of the data, and σ is the standard deviation of the data. When |z| is greater than the set threshold, the data point is considered an outlier, and the detected outlier is processed by replacing it with the mean or median.
[0125] The min-max method is used for data normalization: Among them, q is the original data, q′ is the normalized data, q min and q max are the minimum and maximum values of the data respectively.
[0126] The deep learning model includes a convolutional neural network, and the neural network includes: an input layer, a graph convolutional layer, a feature aggregation layer, a fully connected layer, and an output layer;
[0127] The input layer receives the preprocessed characteristic data group and initializes it as node features, and the initialization is expressed as: Among them, represents the initial feature of node i, and x i represents the node feature of node i;
[0128] The graph convolutional layer adopts heterogeneous graph convolution operations to aggregate and update node features of different types. The function formula of the graph convolutional layer is:
[0129]
[0130] Among them, represents the feature of node i at the l-th layer, represents the set of neighbor nodes of node i, and c ij is the normalization coefficient, and W (l) and are the weight matrices of the l-th layer, and ReLU is the activation function of the graph convolutional layer;
[0131] The feature aggregation layer aggregates the features of each node to generate a new node feature representation. The function formula of the feature aggregation layer is:
[0132]
[0133] Among them, represents the feature of node i in the last layer, and Aggregate represents the feature aggregation function;
[0134] The fully connected layer inputs the aggregated node features into the fully connected layer for feature transformation and non-linear transformation. The function formula of the fully connected layer is:
[0135]
[0136] where y i represents the predicted output delay value of node i, W fc and b fc are the weights and biases of the fully connected layer, and softmax is the activation function of the fully connected layer; the output layer outputs the predicted value y i of the reclaimed water quality detection data of the node.
[0137] The training process of the deep learning model includes the following steps:
[0138] Step S21: Divide the preprocessed feature data set into a training set and a validation set. The training set is used for model training, and the validation set is used for model validation.
[0139] Step S22: Evaluate the model using the cross-validation method and adjust the model parameters according to the evaluation results. Specifically, it includes: dividing the training set into several subsets for k-fold cross-validation; in each fold, select one subset as the validation set, and the remaining subsets as the training set, train the model and evaluate its performance on the validation set; calculate the average validation error of all folds as the evaluation index of the model.
[0140] Step S23: Evaluate the performance of the model on the validation set, select the best model parameters, and retrain the model using the entire training set to obtain the parameters of the final deep learning model.
[0141] It should be noted that those skilled in the art should understand that various transformations and equivalent substitutions can be made to the present invention without departing from the scope of the present invention. In addition, various modifications can be made to the present invention for specific situations or materials without departing from the scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed, but should include all embodiments falling within the scope of the claims of the present invention.
Claims
1. A method for predicting recycled water quality based on a deep learning algorithm, characterized in that: The method comprises the following steps: Step S1, obtaining the original sewage quality data, water purification treatment data, and reclaimed water quality detection data of the reclaimed water from the factory sewage outlet, the reclaimed water treatment equipment, and the treated reclaimed water samples, respectively, to form a multidimensional data group, and selecting a feature data group that affects water quality prediction from the multidimensional data group; Step S2, preprocessing the data of the feature data group, wherein the preprocessing includes outlier processing and normalization processing, and using the preprocessed feature data group to train a deep learning model to obtain model parameters; Step S3, using the trained model to obtain the water quality prediction result of the recycled water; The original data of sewage water quality include 9 dimensions: chemical oxygen demand of sewage, total nitrogen of sewage, total phosphorus of sewage, suspended solids of sewage, biological oxygen demand of sewage, sewage temperature, pH value of sewage, turbidity of sewage and conductivity of sewage; The water purification treatment data includes four dimensions: settling time, aeration volume, filter rate, and backwash frequency; The reclaimed water quality test data includes 9 dimensions: chemical oxygen demand of reclaimed water, total nitrogen of reclaimed water, total phosphorus of reclaimed water, suspended solids of reclaimed water, biological oxygen demand of reclaimed water, temperature of reclaimed water, pH value of reclaimed water, turbidity of reclaimed water and conductivity of reclaimed water; The multidimensional data set is represented as a matrix : , where each row represents the data at a time point, and each column represents a specific water quality indicator or treatment parameter. Represents the original data of sewage quality, Indicates water purification data. Represents the recycled water quality test data, represents the number of time windows, then the matrix The dimension is equal to 22 ; The feature data set is represented as a matrix : , Characteristic data representing the original data of sewage quality, Characteristic data representing water purification treatment data, Represents the characteristic data of recycled water quality detection data, the total number of characteristic data is k, , then the matrix The dimension is equal to ; By calculating the mutual information between the original sewage quality data or the water quality purification treatment data and the reclaimed water quality detection data, a feature data group that affects the water quality prediction is selected from the multidimensional data group: ; in, Represents the mutual information value between two data, , Indicates the original data of sewage quality or water purification treatment data. , Indicates the recycled water quality test data; is the joint probability distribution, and is the marginal probability distribution; According to the size of the mutual information value, k feature data are selected from high to low to form a feature data group.
2. The method for predicting recycled water quality based on deep learning algorithm according to claim 1, characterized in that: The outlier processing is performed by calculating the standard score for each data point: ; in, is the data point, is the mean of the data, is the standard deviation of the data; When it is greater than the set threshold, the data point is considered an outlier, and the detected outlier is processed and replaced by the mean or median; Use the minimum maximum method to normalize the data: ,in, is the original data, is the normalized data, and are the minimum and maximum values of the data respectively.
3. The method for predicting recycled water quality based on deep learning algorithm according to claim 2, characterized in that: The deep learning model includes a convolutional neural network, which includes: an input layer, a graph convolution layer, a feature aggregation layer, a fully connected layer and an output layer; The input layer receives the preprocessed feature data set and initializes it as node features. The initialization is expressed as: ,in, Representation Node The initial characteristics of Representation Node Node characteristics; The graph convolution layer uses heterogeneous graph convolution operations to aggregate and update different types of node features. The function formula of the graph convolution layer is: ; in, Representation Node In the The characteristics of the layer, Representation Node The set of neighbor nodes of is the normalization coefficient, and It is The weight matrix of the layer, ReLU is the activation function of the graph convolution layer; The feature aggregation layer aggregates the features of each node to generate a new node feature representation. The function formula of the feature aggregation layer is: ; in, Representation Node In the last layer of features, Represents feature aggregation function; The fully connected layer inputs the aggregated node features into the fully connected layer for feature transformation and nonlinear transformation. The function formula of the fully connected layer is: ; in, Representation Node The output delay prediction value of and are the weights and biases of the fully connected layer, is the activation function of the fully connected layer; The output layer outputs the predicted value of the recycled water quality detection data of the node .
4. The method for predicting recycled water quality based on deep learning algorithm according to claim 3 is characterized in that: The training process of the deep learning model includes the following steps: Step S21, dividing the preprocessed feature data set into a training set and a validation set, the training set is used for model training, and the validation set is used for model validation; Step S22, using a cross-validation method to evaluate the model, and adjusting the model parameters according to the evaluation results, specifically including: dividing the training set into several subsets, performing k-fold cross-validation; in each fold, selecting a subset as a validation set and the remaining subsets as training sets, training the model and evaluating its performance on the validation set; calculating the average validation error of all folds as the evaluation index of the model; Step S23: Evaluate the performance of the model on the validation set, select the best model parameters, and retrain the model using the entire training set to obtain the parameters of the final deep learning model.
5. The recycled water quality prediction system based on deep learning algorithm is characterized by: The system includes: a data acquisition module, a model training module and a water quality prediction module; The data acquisition module is used to obtain the original sewage quality data, water purification treatment data, and reclaimed water quality detection data of the reclaimed water from the factory sewage outlet, the reclaimed water treatment equipment, and the processed reclaimed water samples, respectively, to form a multidimensional data group, and select the characteristic data group that affects the water quality prediction from the multidimensional data group; The model training module is used to preprocess the data of the feature data group, wherein the preprocessing includes outlier processing and normalization processing, and the preprocessed feature data group is used to train the deep learning model to obtain model parameters; The water quality prediction module is used to obtain the water quality prediction result of the reclaimed water using the trained model; The original data of sewage water quality include 9 dimensions: chemical oxygen demand of sewage, total nitrogen of sewage, total phosphorus of sewage, suspended solids of sewage, biological oxygen demand of sewage, sewage temperature, pH value of sewage, turbidity of sewage and conductivity of sewage; The water purification treatment data includes four dimensions: settling time, aeration volume, filter rate, and backwash frequency; The reclaimed water quality test data includes 9 dimensions: chemical oxygen demand of reclaimed water, total nitrogen of reclaimed water, total phosphorus of reclaimed water, suspended solids of reclaimed water, biological oxygen demand of reclaimed water, temperature of reclaimed water, pH value of reclaimed water, turbidity of reclaimed water and conductivity of reclaimed water; The multidimensional data set is represented as a matrix : , where each row represents the data at a time point, and each column represents a specific water quality indicator or treatment parameter. Represents the original data of sewage quality, Indicates water purification data. Represents the recycled water quality test data, represents the number of time windows, then the matrix The dimension is equal to 22 ; The feature data set is represented as a matrix : , Characteristic data representing the original data of sewage quality, Characteristic data representing water purification treatment data, Represents the characteristic data of recycled water quality detection data, the total number of characteristic data is k, , then the matrix The dimension is equal to ; By calculating the mutual information between the original sewage quality data or the water quality purification treatment data and the reclaimed water quality detection data, a feature data group that affects the water quality prediction is selected from the multidimensional data group: ; in, Represents the mutual information value between two data, , Indicates the original data of sewage quality or water purification treatment data. , Indicates the recycled water quality test data; is the joint probability distribution, and is the marginal probability distribution; According to the size of the mutual information value, k feature data are selected from high to low to form a feature data group.
6. The reclaimed water quality prediction system based on deep learning algorithm according to claim 5 is characterized in that: The outlier processing is performed by calculating the standard score for each data point: ; in, is the data point, is the mean of the data, is the standard deviation of the data; When it is greater than the set threshold, the data point is considered an outlier, and the detected outlier is processed and replaced by the mean or median; Use the minimum maximum method to normalize the data: ,in, is the original data, is the normalized data, and are the minimum and maximum values of the data respectively.
7. The reclaimed water quality prediction system based on deep learning algorithm according to claim 6 is characterized in that: The deep learning model includes a convolutional neural network, which includes: an input layer, a graph convolution layer, a feature aggregation layer, a fully connected layer and an output layer; The input layer receives the preprocessed feature data set and initializes it as node features. The initialization is expressed as: ,in, Representation Node The initial characteristics of Representation Node Node characteristics; The graph convolution layer uses heterogeneous graph convolution operations to aggregate and update different types of node features. The function formula of the graph convolution layer is: ; in, Representation Node In the The characteristics of the layer, Representation Node The set of neighbor nodes of is the normalization coefficient, and It is The weight matrix of the layer, ReLU is the activation function of the graph convolution layer; The feature aggregation layer aggregates the features of each node to generate a new node feature representation. The function formula of the feature aggregation layer is: ; in, Representation Node In the last layer of features, Represents feature aggregation function; The fully connected layer inputs the aggregated node features into the fully connected layer for feature transformation and nonlinear transformation. The function formula of the fully connected layer is: ; in, Representation Node The output delay prediction value of and are the weights and biases of the fully connected layer, is the activation function of the fully connected layer; The output layer outputs the predicted value of the recycled water quality detection data of the node .
8. The reclaimed water quality prediction system based on deep learning algorithm according to claim 7 is characterized in that: The training process of the deep learning model includes the following steps: Step S21, dividing the preprocessed feature data set into a training set and a validation set, the training set is used for model training, and the validation set is used for model validation; Step S22, using a cross-validation method to evaluate the model, and adjusting the model parameters according to the evaluation results, specifically including: dividing the training set into several subsets, performing k-fold cross-validation; in each fold, selecting a subset as a validation set and the remaining subsets as training sets, training the model and evaluating its performance on the validation set; calculating the average validation error of all folds as the evaluation index of the model; Step S23: Evaluate the performance of the model on the validation set, select the best model parameters, and retrain the model using the entire training set to obtain the parameters of the final deep learning model.
Citation Information
Patent Citations
Treated sewage quality prediction method based on combination of support vector classification and GRU neural network
CN111291937A