Method for intelligently deciding application field and application of reclaimed water
By building a KNN-multi-threaded SVM model, integrating feature space and optimizing the training process, the problem of high-dimensional data processing in the utilization configuration of recycled water is solved, and intelligent decision-making and accurate recommendation of recycled water is achieved.
Patent Information
- Application Number
- CN202510737883.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-08
Smart Images

Figure CN120278342A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water ecological environment, and particularly to a method for intelligent decision-making on the applicable fields and uses of reclaimed water. Background Art
[0002] Reclaimed water refers to sewage (wastewater) that, after specific treatment, meets the water quality standards for specific uses and can thus be reused. According to documents such as the "Regulations on Water Conservation" and the "Regulations on Urban Drainage and Sewage Treatment", it is required to promote reclaimed water in fields such as industry, agriculture, and urban miscellaneous uses, and to give priority to its application in industrial production, urban greening, road cleaning, etc. Currently, the "Water Quality Standard for Reclaimed Water" (SL368) promulgated in 2010 has not been updated and revised, and its applicability is not strong. In recent years, the corresponding standards for the allocation of reclaimed water utilization have been issued in the form of different fields, such as: industrial field (GB / T 1923-2024), municipal field (GB / T 1923-2007), ecological environment field (GB / T 1923-2019), etc. In addition, the number of indicators involved in the standards of each field is inconsistent (differing by dozens at most), and the category differences are large. The urban miscellaneous use field needs to consider sensory indicators (odor). It can be seen that if, according to the effluent conditions of reclaimed water treatment plants (especially when the amount of data is large and the water quality conditions vary greatly), the method of judging its applicable fields and uses by comparing each series of standards one by one has a large workload and high practical operation difficulty.
[0003] Many studies have achieved intelligent recognition of classification problems such as text classification and fault diagnosis by combining the K-Nearest Neighbor Classification Algorithm (KNN) and the Support Vector Machine (SVM). However, most of these methods belong to single-level operation, that is, setting a threshold as the boundary, and using the KNN and SVM models to classify samples respectively. For types with complex characteristics such as reclaimed water utilization configuration, problems such as insufficient high-dimensional data processing ability and superposition of noise sensitivity will occur.
[0004] However, when this double-layer model is applied to the reclaimed water classification problem, it still needs to be innovated and optimized, mainly reflected in: if the conventional method of dealing with missing values is directly used to expand the data set, there is a risk of generating abnormal values, such as values that obviously do not meet the discharge standards; compared with the differences in indicators between different fields, the indicator differences are significantly reduced when classifying the uses of a single field, and often some indicators are used for screening. Running the above hierarchical model will greatly increase the training difficulty of the secondary level, that is, the SVM module; if the reclaimed water conditions simultaneously meet multiple types of standards, running this model is expected to output all target values, which cannot play an effective decision-making role. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the purpose of the embodiments of the present invention is to provide a method for intelligent decision-making on the applicable fields and uses of reclaimed water to solve the problems in the above background art.
[0006] To achieve the above object, the present invention provides the following technical solutions: A method for intelligent decision-making on the applicable fields and uses of reclaimed water, comprising the following steps: Step 1: Construction of a model training data set; Collect the effluent parameter data of reclaimed water treatment plants in different regions, different seasons, and for different uses, compare with the corresponding standards for relevant reclaimed water configuration, respectively judge the applicable fields and uses, and construct a feature data set with reclaimed water conditions as features and applicable fields and uses as targets; Step 2: Construction of a classification-decision model; One is to construct a data processing module, establish a complete feature space, and based on the feature space, perform missing value supplementation, feature standardization and normalization, and feature dimensionality reduction on the collected effluent data of reclaimed water treatment plants to improve the feature data set; The second is to construct a pre-classification module, train and optimize the first-level KNN model based on the feature data set and feature weights, and then combine the results of the KNN model, and use the re-screened feature data set to train the second-level SVM to construct a KNN-multi-threaded SVM model; The third is to construct a decision module, and based on the requirements of priority reclaimed water configuration, perform decision sorting on the output results of the pre-classification module through relevant sorting algorithms; Step 3: Prediction module; Use the classification-decision model to predict the effluent of reclaimed water treatment plants and make intelligent recommendations on the applicable fields and uses of reclaimed water.
[0007] As a further solution of the present invention, the construction of the feature space in the second step includes integrating and statistically analyzing all feature indicators of the existing corresponding standards for reclaimed water configuration, and performing dimension expansion on conditional indicators such as total chlorine and total dissolved solids.
[0008] As a further solution of the present invention, the supplementation of missing values in the data set in the second step includes filling the missing values of the model data set according to the constructed feature space, encoding each feature and output target respectively, and adding boundary values to each feature through the constructed correction module to construct a data set.
[0009] As a further solution of the present invention, the feature standardization and normalization in the second step include processing the data set using Z-score standardization and Min-Max normalization to eliminate the influence of different feature dimensions or magnitude differences on the model.
[0010] As a further solution of the present invention, the feature dimensionality reduction includes using the PCA principal component analysis method to perform dimensionality reduction on the features, removing redundant indicators, and improving the calculation efficiency of subsequent models.
[0011] The method for the applicable fields and uses of intelligent decision-making reclaimed water according to claim 5 is characterized in that in step two, the dataset division includes dividing the constructed characteristic dataset according to different applicable fields and uses, and naming them Data-A and Data-a respectively.
[0012] As a further solution of the present invention, the creation of the KNN-multi-threaded SVM model includes: Construction of the KNN model: Divide the Data-A dataset into a training set and a test set. According to the priority control indicators in the corresponding standards for reclaimed water configuration in different fields, weight the importance of each indicator. Use the reclaimed water conditions as input features and the applicable fields as output features for the training rule. According to the weights of each feature, and using the Manhattan distance as the distance metric method, train the KNN algorithm. Among them, use Bayesian optimization to optimize the model parameters, and use the root mean square error and accuracy as the evaluation indicators for the training set and test set of the model training. Construction of the multi-threaded SVM model: Create SVM models based on different fields. Before training each SVM model, based on the Data-a dataset, divide the required indicators according to the field and use of the SVM model, deeply screen the training features of the SVM model, use the screened indicators as features, different uses as output targets, and SVC as the classification method to train the SVM model. Use grid search to optimize the kernel function, penalty coefficient C, and class-weight hyperparameters of each SVM model, and use accuracy and precision as the model evaluation indicators.
[0013] As a further solution of the present invention, the construction of the decision-making module includes constructing a static weight classification model based on the output results of the KNN-multi-threaded SVM model and based on the priority configuration requirements of reclaimed water in relevant documents to realize the decision-making on the applicable fields and uses of reclaimed water.
[0014] In summary, the embodiments of the present invention have the following beneficial effects compared with the prior art: First, in the data processing module, based on standards such as the "Discharge Standard of Pollutants for Municipal Wastewater Treatment Plants" and the "Environmental Quality Standards for Surface Water", a feature dataset correction module is built, reducing the probability of outliers in the feature dataset and providing a data basis for subsequent model training and prediction. Second, in the construction of the pre-classification module, an SVM model is trained separately for each applicable field of reclaimed water output by the KNN model. At the same time, before training the SVM model, the training features are deeply screened according to the indicators required for the uses divided by the field. Through this strategy, problems such as large training difficulty existing in the SVM model when the data volume is large can be effectively alleviated. Third, a decision-making module is built. The static weight classification method is introduced, and combined with the requirements of relevant documents on the priority allocation of reclaimed water, the output results of the KNN model and the multi-threaded SVM model in the KNN-multi-threaded SVM classification model are innovatively sorted decisionally, effectively improving the decision-making ability of this method.
[0015] To more clearly illustrate the structural features and effects of the present invention, the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. Brief Description of the Drawings
[0016] Figure 1 It is a schematic flow chart of a method for intelligently making decisions on the applicable fields and uses of reclaimed water; Figure 2 It is a schematic diagram of the process of filling in missing values in the feature dataset; Figure 3 It is a schematic diagram of the in-depth screening process of the training features of each SVM model; Figure 4 It is a schematic diagram of the construction of the decision-making module. Detailed Embodiments
[0017] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0018] The following describes the specific implementation of the present invention in detail in conjunction with specific embodiments.
[0019] In one embodiment, a method for intelligently making decisions on the applicable fields and uses of reclaimed water, see Figures 1 to 4 , includes the following steps: Step 1: Construction of the model training dataset Collect the effluent parameter data of reclaimed water treatment plants in different regions, different seasons, and different uses, compare the corresponding standards for relevant reclaimed water allocation, and respectively judge the applicable fields and uses to construct a feature dataset with reclaimed water conditions as features and applicable fields and uses as targets; Step 2: Construction of the classification - decision model First, construct a data processing module to establish a complete feature space. Based on the feature space, perform missing value supplementation, feature standardization, normalization, and feature dimensionality reduction on the collected effluent data from the reclaimed water plant to improve the feature dataset. Second, construct a pre-classification module. Based on the feature dataset and feature weights, train and optimize the first-level KNN model. Then, combined with the results of the KNN model, use the re-screened feature dataset to train the second-level SVM to construct a KNN-multi-threaded SVM model. Third, construct a decision-making module. Based on the requirements of the priority allocation of reclaimed water, perform decision-making sorting on the output results of the pre-classification module through relevant sorting algorithms. Step 3: Prediction module Use the classification-decision model to predict the effluent from the reclaimed water plant and provide intelligent recommendations for the applicable fields and uses of reclaimed water.
[0020] Furthermore, refer to Figures 1 to 4 , in the step of constructing the feature space in step 2, it includes integrating and statistically analyzing all feature indicators corresponding to the existing reclaimed water allocation standards, and performing dimensionality expansion on conditional indicators such as total chlorine and total dissolved solids.
[0021] Furthermore, refer to Figures 1 to 4 , in the step of supplementing missing values in the dataset in step 2, it includes filling in the missing values in the model dataset based on the constructed feature space, encoding each feature and output target respectively, and adding boundary values to each feature through the constructed correction module to construct a complete and highly representative dataset.
[0022] Furthermore, refer to Figures 1 to 4 , in the step of feature standardization and normalization in step 2, it includes using Z-score standardization and Min-Max normalization to process the dataset to eliminate the influence of different feature dimensions or magnitude differences on the model.
[0023] Furthermore, refer to Figures 1 to 4 , the feature dimensionality reduction includes using the PCA principal component analysis method to reduce the dimensionality of the features, removing redundant indicators, and improving the calculation efficiency of the subsequent model.
[0024] Furthermore, refer to Figures 1 to 4 , in the step of dataset division in step 2, it includes dividing the constructed characteristic dataset according to different applicable fields and uses, and naming them Data_A and Data_a respectively.
[0025] Furthermore, refer to Figures 1 to 4 , the creation of the KNN-multi-threaded SVM model includes: Construction of the KNN model (first layer): Divide the Data_A dataset into a training set and a test set. According to the priority control indicators in the corresponding standards for reclaimed water allocation in different fields, weight the importance of each indicator. Use the reclaimed water conditions as input features and the applicable fields as output features to establish training rules. Train the KNN algorithm according to the weights of each feature and use the Manhattan distance as the distance metric. Optimize the model parameters (such as the k value) using Bayesian optimization, and use the root mean square error (RMSE) and accuracy as the evaluation indicators for the training set and test set of the model training. Construction of the multi-threaded SVM model (second layer): Create SVM models based on different fields. Before training each SVM model, based on the Data_a dataset, divide the required indicators according to the field and use of the SVM model, and deeply screen the training features of the SVM model. Use the screened indicators as features and different uses as output targets, and use SVC as the classification method to train the SVM model. Optimize the hyperparameters such as the kernel function, penalty coefficient C, and class_weight of each SVM model using grid search, and use accuracy and precision as the model evaluation indicators.
[0026] Further, refer to Figures 1 to 4 The construction of the decision-making module includes constructing a static weight classification model based on the output results of the KNN-multi-threaded SVM model and the priority allocation requirements of reclaimed water in relevant documents to achieve decision-making on the applicable fields and uses of reclaimed water.
[0027] In this embodiment, step S1 is the construction of the model training dataset. Collect the effluent data of reclaimed water plants under different conditions.
[0028] Step S2 is the construction of the classification-decision model, including: (i) Data processing module. Construct a complete feature space, and process the collected effluent data of reclaimed water plants according to the feature space to improve the feature dataset; (ii) Pre-classification module. Based on the feature dataset and the weights of each feature, train and optimize the first layer (KNN model), and then combine the results of the KNN model to train the second layer (SVM) using the screened feature dataset to construct a KNN-multi-threaded SVM model; (iii) Decision-making module. Based on the priority allocation requirements of reclaimed water in relevant documents, make decisions on the output results of the classification module through relevant sorting algorithms.
[0029] Step S3 is the prediction module, which uses the classification-decision model to predict the effluent of reclaimed water plants to achieve intelligent recommendation of the applicable fields and uses of reclaimed water.
[0030] Further, the construction of the model training dataset in step S1 includes: Collect the effluent data of reclaimed water treatment plants in different regions, seasons, and for different uses, compare with the corresponding standard requirements for relevant reclaimed water allocation, respectively judge the applicable fields and uses, and construct a feature dataset with reclaimed water conditions as features and applicable fields and uses as targets.
[0031] Furthermore, the data processing module in step S2(i) includes parts such as feature space construction, missing value supplementation of the dataset, feature standardization and normalization, and feature dimensionality reduction.
[0032] Furthermore, constructing the feature space includes: integrating and statistically analyzing all feature indicators of the existing corresponding standards for reclaimed water allocation, and expanding the dimensions of conditional indicators.
[0033] The relevant reclaimed water allocation standards are: "Urban Sewage Reclaimed Water Utilization - Quality of Water for Farmland Irrigation" (GB / T 20922 - 2007), "Urban Sewage Reclaimed Water Utilization - Quality of Water for Urban Miscellaneous Use" (GB / T 18920 - 2020), "Urban Sewage Reclaimed Water Utilization - Quality of Water for Landscape and Environmental Use" (GB / T 18921 - 2019), "Urban Sewage Reclaimed Water Utilization - Quality of Water for Groundwater Recharge" (GB / T 18921 - 2005), "Urban Sewage Reclaimed Water Utilization - Quality of Water for Industrial Use" (GB / T 19923 - 2024), "Urban Sewage Reclamation and Utilization - Quality of Water for Green Space Irrigation" (GB / T 25499 - 2010).
[0034] Furthermore, the missing value supplementation of the dataset includes: based on the constructed feature space, using the KNN (K-Nearest Neighbors) imputation method and mode filling to fill the missing values of the feature dataset, and encoding the features and labels. Subsequently, add boundary values to each feature through a correction module (such as the np.clip algorithm) constructed based on the "Reclaimed Water Quality Standard", "Discharge Standard of Pollutants for Municipal Wastewater Treatment Plants", and "Surface Water Environment Quality Standard" to construct a complete and highly representative dataset.
[0035] Mode filling: Use the mode in the dataset to replace all missing values. The steps are as follows:
[0036] First, analyze the target variable (i.e., the column for which missing values need to be filled), find the mode of this variable, and then replace all missing values with the calculated mode (for each missing value, the same mode value is assigned). Among them, for categorical data, count the frequency of each category and select the category with the highest frequency as the mode; for numerical data, count the frequency of each value and select the value with the highest frequency as the mode.
[0037] KNN (K-Nearest Neighbors) imputation method: Find the K nearest neighbors of the missing sample in the feature space through distance measurement, and use the observed values of these neighbors to fill the missing values.
[0038] numpy.clip is a core function in the NumPy library for constraining array elements within a certain range. Its core principle is as follows: By truncating each element, all values in the input array are restricted within the defined range [a_min, a_max]. If the input value is less than the minimum value, it is replaced with the minimum value; if it is greater than the maximum value, it is replaced with the maximum value.
[0039] Furthermore, feature standardization and normalization include: processing the dataset using Z-score standardization and Min-Max normalization to eliminate the impact of different feature dimensions or magnitude differences on the model.
[0040] Furthermore, feature dimensionality reduction includes: using the PCA principal component analysis method to reduce the dimensionality of features, eliminating redundant indicators, and improving the computational efficiency of subsequent models.
[0041] Furthermore, the pre-classification module in step S2 (ii) includes parts such as dataset division and creation of a two-tier model.
[0042] Furthermore, dataset division includes: dividing the characteristic dataset constructed in step S2 (i) according to different application fields and uses, and naming them Data_A and Data_a respectively.
[0043] Furthermore, the KNN model (first tier) includes parts such as dataset division, definition of feature weights, model training, and optimization.
[0044] Furthermore, dataset division includes dividing the Data_A dataset into a training set and a test set.
[0045] Furthermore, defining feature weights includes weighting the importance of each indicator according to the priority control indicators in the standards of each field.
[0046] Furthermore, the model training and optimization steps include: using the reclaimed water conditions as input features and the application field as output features for the training rule. According to each feature weight, and using the Manhattan distance as the distance metric, train the KNN algorithm. Among them, use Bayesian optimization to optimize the model parameters (such as the k value) (perform k-fold cross-validation (using the negative mean squared error (MSE) as the evaluation index)), and use the root mean squared error (RMSE) and accuracy as the evaluation indicators for the model training training set and test set.
[0047] The KNN classification model is a non-parametric supervised learning algorithm based on instances. It realizes classification by calculating the distance between the sample to be classified and all training samples. Among them, the Manhattan distance is a method for measuring the distance between two points in a metric space, especially suitable for spaces with a grid-like structure. Its calculation method is as follows: d(A, B) = |x1 - y1| + |x2 - y2| +... + |x n - y n | where x1, x2,..., x n ; y1, y2,..., y n are the coordinates of two points in n-dimensional space respectively MSE: ; RMSE: ; where yi is the true value, y ^ i is the predicted value, and n is the number of samples. The smaller the RMSE value, the higher the model accuracy, and vice versa.
[0048] Accuracy: ; where TP is predicted as the positive class and is correct; TN is predicted as the negative class and is correct; FP is predicted as the positive class but is wrong; FN is predicted as the negative class but is wrong.
[0049] Furthermore, the construction of the SVM model (the second level) includes: creating SVM models based on different fields. Before training each SVM model, based on the Data_a dataset, divide the required metrics according to the field and usage of the SVM model, and deeply screen the training features of the SVM model. Using the screened metrics as features, different usages as output targets, and Support Vector Classification (SVC) as the classification method, train the SVM model. Use grid search to optimize the hyperparameters such as the kernel function, penalty coefficient C, class_weight, etc. of each SVM model, and use accuracy and precision as the model evaluation metrics.
[0050] SVM is a supervised learning algorithm based on statistical learning theory. Its core idea is to achieve classification by constructing an optimal hyperplane and maximize the margin between classes to improve generalization ability. Its goal is to find a hyperplane (a straight line in two dimensions and a plane in three dimensions) that separates samples of different classes and maximizes the distance from this hyperplane to the nearest sample point (i.e., the support vector), where: Any hyperplane can be described by the following linear equation: ; The distance formula from a two-dimensional space point (x, y) to the straight line Ax + By + C = 0 is: ; After expanding the n-dimensional space, the distance from the point x = (x1, x2... x n ) to the straight line is: ; where is an n-dimensional vector, b is a real number, A and B are coefficients, .
[0051] Precision rate: ; Furthermore, conduct in-depth screening of the SVM model training features. For example, construct an SVM model for identifying different uses in the industrial field, and based on the Data_a dataset, screen out the index data required for use classification belonging to the industrial field (such as turbidity / NTU, total dissolved solids (TDS), etc.).
[0052] Furthermore, the preferred types of kernel functions are: linear, polynomial, Gaussian, and Sigmoid kernel functions.
[0053] Among them, TP is predicted as the positive class and is correct; FP is predicted as the positive class but is wrong.
[0054] Furthermore, in step S2(iii), the decision-making module includes: based on the output results of the KNN-multi-threaded SVM model, construct a static weight classification model (such as sorting weight weighting) based on the requirements for the priority allocation of reclaimed water in documents such as the "Regulations on Water Conservation" and the "Regulations on Urban Drainage and Sewage Treatment", to achieve the decision-making on the applicable fields and uses of reclaimed water.
[0055] Sorting weight weighting: A comprehensive sorting method that combines the sorting order and artificial weight allocation. Its core principle is to transform the original results output by the model through sorting, and then combine the weight values given by business rules or expert experience to generate the final recommended order.
[0056] Furthermore, in step S3, the prediction module includes inputting the outlet parameters of the reclaimed water plant and outputting the recommended applicable fields and uses of reclaimed water through the trained classification-decision model.
[0057] Construction of the model training dataset. First, collect the effluent parameter data of reclaimed water treatment plants in different regions (such as the south and north, water-rich and water-scarce regions, industrial and agricultural developed regions, etc.), different seasons (such as flood season and dry season), and different uses (such as industrial production, urban greening, road cleaning, etc.); Second, compare the "Reclaimed Water Utilization in Urban Sewage - Quality Standards for Agricultural Irrigation Water" (GB / T 20922 - 2007), "Reclaimed Water Utilization in Urban Sewage - Quality Standards for Urban Miscellaneous Water" (GB / T 18920 - 2020), "Reclaimed Water Utilization in Urban Sewage - Quality Standards for Landscape and Environmental Water" (GB / T 18921 - 2019), "Reclaimed Water Utilization in Urban Sewage - Quality Standards for Groundwater Recharge" (GB / T 18921 - 2005), "Reclaimed Water Utilization in Urban Sewage - Quality Standards for Industrial Water" (GB / T 19923 - 2024), and "Reclaimed Water Utilization in Urban Sewage - Quality Standards for Green Irrigation" (GB / T 25499 - 2010), and respectively judge their applicable fields and uses (note: retain the effluent data of reclaimed water that does not meet all standards); Third, construct a feature dataset with reclaimed water conditions (such as BOD5, COD, chromaticity, odor, etc.) as features, and the applicable fields (such as urban miscellaneous use, landscape and environment, industrial water, etc.) and uses (such as toilet flushing, vehicle washing, urban greening, ornamental landscape environmental water (river type), etc.) as targets.
[0058] Step S2: Establish a data processing module. It includes parts such as feature space construction, missing value supplementation of the dataset, feature standardization and normalization, and feature dimensionality reduction.
[0059] Step S201: Feature space construction includes: integrating all feature indicators corresponding to the reclaimed water utilization configuration standards such as (GB / T 19923 - 2024), (GB 20922 - 2007), (GB / T 18920 - 2020), (GB / T 18921 - 2019), (GB / T 19772 - 2005), and (GB / T 25499 - 2010), and performing dimensionality expansion on conditional indicators such as total chlorine and total dissolved solids, such as: classifying total chlorine by factory outlet and network end; classifying total dissolved solids by non-saline-alkali areas and saline-alkali areas.
[0060] Step S202: Based on the constructed feature space, use the KNN nearest neighbor interpolation method and mode filling to fill in the missing values of the feature dataset, where: use the KNN nearest neighbor interpolation method to fill in the numerical features; use mode filling for non-numerical features (such as smell). Feature and label encoding are performed separately for each feature's input and output targets (for different fields and uses), where: one-hot encoding is used for features; LabelEncoder is used for targets. Subsequently, based on the "Reclaimed Water Quality Standard" (SL368), the "Pollutant Discharge Standard for Municipal Wastewater Treatment Plants (GB 18918-2002)" ("Grade A Standard"), and the "Surface Water Environment Quality Standard (GB 3838-2002)" ("Class V") standard requirements, set the range intervals for each indicator and construct the np.clip algorithm to optimize the feature dataset. Among them, numpy.clip restricts all values in the input array to the defined interval [a_min, a_max] through element-wise truncation. If the input value is less than the minimum value, it is replaced with the minimum value; if it is greater than the maximum value, it is replaced with the maximum value. For example: According to the KNN nearest neighbor interpolation method, a pH indicator array [5.9, 6.1, 7.6, 9.1] is generated. After passing through the np.clip correction module, its outliers [5.9, 9.1] are corrected to [6, 9] respectively.
[0061] Step S203: Feature standardization and normalization include: using Z-score standardization and Min-Max normalization to process the dataset to eliminate the impact of different feature dimensions or magnitude differences on the model.
[0062] Step S204: Feature dimensionality reduction includes: using the PCA principal component analysis method to reduce the dimensionality of the features, eliminating redundant indicators, and improving the calculation efficiency of the subsequent model.
[0063] Step S3 Construct a pre-classification model (KNN-multi-threaded SVM model), which includes parts such as dataset division and two-tier model creation.
[0064] Step S301: Dataset division includes dividing the feature dataset according to different applicable fields (such as urban miscellaneous use, landscape environment, industrial water use, etc.) and uses (such as toilet flushing, vehicle washing, urban greening, ornamental landscape environment water use (river type), etc.), and naming them Data_A and Data_a respectively; Step S302: Two-tier model creation includes the construction of a KNN model (the first tier) and an SVM model (the second tier), where: KNN model construction: First, divide the Data_A dataset into a training set and a test set (where the ratio of the number of the training set to the test set is 7:3); Second, perform feature importance weighting on the priority control indicators in different field standards (the weight range is 1.5 - 2.5), such as: for farmland irrigation, the priority control pathogen (such as the number of Ascaris eggs); for urban miscellaneous use, the priority control sensory indicators (such as odor); for landscape environment, the priority control sensory indicators (such as no floating objects, no unpleasant odor and taste); Third, train the KNN model. Taking the reclaimed water parameter conditions as input features and the applicable fields as output features for the training rule, based on the weights of each feature, using the Manhattan distance as the distance metric method, train using the KNN algorithm, and optimize the hyperparameters such as the k value and weight of the model through Bayesian optimization. Select the optimal hyperparameters of the KNN model by performing 5-fold cross-validation and calculating the validation loss (using the negative mean squared error (MSE) as the evaluation index). Subsequently, predict the training set and the test set respectively through the KNN algorithm, and evaluate the model through RMSE and accuracy rate.
[0065] SVM model construction: First, respectively corresponding to the results representing "farmland irrigation", "urban miscellaneous use", "landscape environment", "industrial water use", "groundwater recharge", "greenland irrigation" output by the KNN model, build an SVM classification model (where, since classifying the uses in the fields of farmland irrigation (involving 7 uses) and landscape environment (involving 4 uses) is not a conventional binary classification problem, by adjusting the "decision_function_shape" parameter, call the "One versus One (OVO)" mode to build 21 and 6 SVM classifiers respectively). Second, before training each SVM model, deeply screen the training features according to the indicators required for use division in the field to which each SVM model belongs. For example: the features selected by the SVM (farmland irrigation) model are {"treatment intensity", "BOD5", "COD Cr ", "SS", "TDS", "residual chlorine", "petroleum substances", "anionic surfactant", "As", "total coliform group number"}; the features selected by the SVM (urban miscellaneous use) model are {"chromaticity", "turbidity", "ammonia nitrogen", "Fe", "Mn"}; the features selected by the SVM (landscape environment) model are {"BOD5", "turbidity", "total phosphorus", "total nitrogen", "ammonia nitrogen", "total coliform group number", "residual chlorine"}; the features selected by the SVM (industrial water use) model are {"turbidity", "TDS", "sulfate", "chloride", "Fe", "Mn"}; the features selected by the SVM (groundwater recharge) model are {"BOD5", "COD Cr”, “chromaticity”, “turbidity”, “volatile phenols”, “nitrates”, “ammonia nitrogen”, “animal and vegetable oils”, “petroleum”, “number of Escherichia coli-like groups”}; the SVM (green space irrigation) model selected features as {“turbidity”, “number of Escherichia coli-like groups”, “number of ascarid eggs”}. Third, based on the data sets after deep screening, with screening indicators as features, different uses as output targets, and SVC as the classification method, the SVM model for use classification in various fields was trained, and the kernel functions (linear kernel function, polynomial kernel function, Gaussian kernel function and Sigmoid kernel function) of each SVM model were selected by grid search method. The optimal hyperparameters (including kernel function (linear kernel function, polynomial kernel function, Gaussian kernel function and Sigmoid kernel function), penalty coefficient C) and class_weight (the weight ratio of majority class to minority class is 1:3-1:10)) for each SVM model are optimized by performing 5-fold cross validation and calculating validation loss (using negative mean square error (MSE) as the evaluation indicator). Fourth, each SVM model is used to classify the purpose of the training set and test set, and each SVM model is evaluated using accuracy and precision as model evaluation indicators.
[0066] Step S4, constructing a decision module includes: based on the priority configuration requirements of recycled water in documents such as the "Water Conservation Regulations" and the "Town Drainage and Sewage Treatment Regulations", the "urban greening, road cleaning, vehicle washing, and construction" uses in the "urban miscellaneous uses" field and the "landscape environment" field are recommended as hard rules, and the sorting weight weighting method is adopted to give the output results of the KNN model corresponding to the "urban miscellaneous uses" and "landscape environment" fields higher weights, and at the same time, all output results in SVM (landscape environment) and the uses representing "road cleaning", "vehicle washing", and "construction" in SVM (urban miscellaneous uses) are given high weights (the weight range is set to (0.3, 1]), so as to achieve decision-making sorting of the output results of the KNN-multi-threaded SVM classification model.
[0067] Step S5, the prediction module executes: collecting effluent conditions from different reclaimed water plants (such as organic matter indicators (BOD5, COD), sensory indicators (smell), metal concentrations (Fe, Mn), etc.), constructing a feature data set to be predicted, and then running the classification-decision model constructed in the above steps, and loading the feature data set to be predicted, and making intelligent decisions on the applicable fields and uses of its reclaimed water.
[0068] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for intelligent decision-making on the applicable fields and uses of reclaimed water, characterized in that, It includes the following steps: Step 1: Construction of the model training data set; Collect the effluent parameter data of reclaimed water treatment plants in different regions, different seasons, and for different purposes, compare with the corresponding standards for relevant reclaimed water configuration, respectively judge the applicable fields and purposes, and construct a feature data set with reclaimed water conditions as features and applicable fields and purposes as targets; Step 2: Construction of the classification - decision model; First, construct a data processing module, establish a complete feature space, and according to the feature space, perform missing value supplementation, feature standardization and normalization, and feature dimensionality reduction on the collected effluent data of reclaimed water treatment plants to improve the feature data set; Second, construct a pre-classification module. Based on the feature data set and feature weights, train and optimize the first-level KNN model, and then combine the results of the KNN model. Use the re-screened feature data set to train the second-level SVM to construct a KNN - multi-threaded SVM model; Third, construct a decision module. Based on the requirements of priority reclaimed water configuration, perform decision-making sorting on the output results of the pre-classification module through relevant sorting algorithms; Step 3: Prediction module; Use the classification - decision model to predict the effluent of reclaimed water treatment plants and provide intelligent recommendations for the applicable fields and purposes of reclaimed water.
2. The method for the applicable fields and uses of intelligent decision-making reclaimed water according to claim 1, characterized in that, The construction of the feature space in Step 2 includes integrating all feature indicators of the existing corresponding standards for reclaimed water configuration and expanding the dimensions of conditional indicators such as total chlorine and dissolved total solids.
3. The method for the applicable fields and uses of intelligent decision-making reclaimed water according to claim 2, wherein The supplementation of missing values in the data set in Step 2 includes filling in the missing values of the model data set according to the constructed feature space, encoding each feature and output target respectively, and adding boundary values to each feature through the constructed correction module to construct the data set.
4. The method for the applicable fields and uses of reclaimed water for intelligent decision-making according to claim 3, characterized in that, The feature standardization and normalization in Step 2 include using Z-score standardization and Min-Max normalization to process the data set to eliminate the influence of different feature dimensions or magnitude differences on the model.
5. The method for the applicable fields and uses of intelligent decision-making reclaimed water according to claim 4, wherein The feature dimensionality reduction includes using the PCA principal component analysis method to reduce the dimensionality of the features, removing redundant indicators, and improving the calculation efficiency of the subsequent model.
6. The method for the applicable fields and uses of intelligent decision-making reclaimed water according to claim 5, wherein The division of the data set in Step 2 includes dividing the constructed characteristic data set according to different applicable fields and purposes, and naming them Data-A and Data-a respectively.
7. The method for the applicable fields and uses of intelligent decision-making reclaimed water according to claim 6, characterized in that, The creation of the KNN - multi-threaded SVM model includes: Construction of the KNN model: Divide the Data-A data set into a training set and a test set. According to the priority control indicators in the corresponding standards for reclaimed water configuration in different fields, perform feature importance weighting on each indicator. According to the training rule with reclaimed water conditions as input features and applicable fields as output features, and according to each feature weight, use the Manhattan distance as the distance metric method to train the KNN algorithm, where Bayesian optimization is used to optimize the model parameters, and the root mean square error and accuracy are used as the evaluation indicators for the training set and test set of model training; Multi-threaded SVM model construction: SVM models are created separately based on different fields. Before training each SVM model, based on the Data-a dataset, the required indicators are divided according to the field and usage of the SVM model, and the training features of the SVM model are deeply screened. The screened indicators are used as features, different usages are used as output targets, SVC is used as the classification method to train the SVM model, and the kernel function, penalty coefficient C, and class-weight hyperparameters of each SVM model are optimized using grid search, with accuracy and precision as the model evaluation indicators.
8. The method for the applicable fields and uses of intelligent decision-making reclaimed water according to claim 7, wherein The construction of the decision-making module includes constructing a static weight classification model based on the output results of the KNN-multi-threaded SVM model and the requirements for the priority allocation of reclaimed water in relevant documents, so as to realize the decision-making on the applicable fields and usages of reclaimed water.