Water quality prediction method and apparatus, computer device, storage medium and computer program product
By combining and processing the water quality data of the tailwater to be tested and extracting features, and using multiple pre-trained water quality prediction models for prediction and weighted summation, the problem of low accuracy in traditional water quality detection methods is solved, and higher detection accuracy and reliability are achieved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2026-04-02
AI Technical Summary
Traditional water quality testing methods rely on manual testing, which can lead to low accuracy due to subjective factors.
By acquiring water quality data of the tailwater to be tested, performing combined processing and feature extraction, using multiple pre-trained water quality prediction models for prediction, filtering and weighted summation, the target water quality prediction result is obtained.
It improves the accuracy and reliability of water quality testing, avoids subjective errors in manual testing, and enhances testing accuracy.
Smart Images

Figure CN2024128929_02042026_PF_FP_ABST
Abstract
Description
Water quality prediction method and device, computer device, storage medium and computer program product TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and in particular to a water quality prediction method and device, a computer device, a computer readable storage medium and a computer program product. BACKGROUND
[0002] At present, in order to ensure the smooth progress of water pollution control work, it is essential to accurately detect water quality.
[0003] In the prior art, during the detection of water quality, a manual detection method is generally used. However, this method has subjective factors and is prone to errors, resulting in low accuracy of water quality detection.
[0004] SUMMARY
[0005] Therefore, it is necessary to provide a water quality prediction method and device, a computer device, a computer readable storage medium and a computer program product capable of improving the accuracy of water quality detection.
[0006] In a first aspect, the present application provides a water quality prediction method, comprising:
[0007] obtaining water quality data of a to-be-detected tail water;
[0008] combining the water quality data to obtain combined data corresponding to the water quality data;
[0009] extracting features from the water quality data and the combined data to obtain a feature vector of the water quality data and a feature vector of the combined data;
[0010] splicing the feature vector of the water quality data and the feature vector of the combined data to obtain a target feature vector of the water quality data;
[0011] inputting the target feature vector into a plurality of pre-trained water quality prediction models to obtain a water quality prediction result of the to-be-detected tail water output by each pre-trained water quality prediction model, and a prediction probability corresponding to each water quality prediction result;
[0012] from the each water quality prediction result, filtering out a water quality prediction result corresponding to a prediction probability greater than a preset prediction probability as a filtered water quality prediction result;
[0013] performing weighted summation processing on the filtered water quality prediction result to obtain a target water quality prediction result of the to-be-detected tail water.
[0014] In one of the embodiments, the water quality data of the tail water to be tested is obtained, including:
[0015] The source information of the tail water to be tested is determined;
[0016] According to the source information of the tail water to be tested, the corresponding relationship between the source information and the preset index is queried, and the preset index corresponding to the source information of the tail water to be tested is obtained as the preset index of the tail water to be tested.
[0017] The water quality data of the tail water to be tested under the preset index is obtained as the water quality data of the tail water to be tested.
[0018] In one of the embodiments, the feature extraction processing is performed on the water quality data and the combined data to obtain the feature vector of the water quality data and the feature vector of the combined data, including:
[0019] The water quality data and the combined data are respectively preprocessed to obtain preprocessed water quality data and preprocessed combined data;
[0020] The preprocessed water quality data is taken as main data, and the preprocessed combined data is taken as auxiliary data, and is input into a feature extraction model to obtain the feature vector of the water quality data;
[0021] The preprocessed combined data is taken as main data, and the preprocessed water quality data is taken as auxiliary data, and is input into a feature extraction model to obtain the feature vector of the combined data.
[0022] In one of the embodiments, the weighted sum processing is performed on the screened water quality prediction result to obtain the target water quality prediction result of the tail water to be tested, including:
[0023] The target water quality prediction model corresponding to the screened water quality prediction result is determined;
[0024] The prediction accuracy and the prediction efficiency of the target water quality prediction model corresponding to the screened water quality prediction result are obtained;
[0025] According to the prediction accuracy and the prediction efficiency of the target water quality prediction model corresponding to the screened water quality prediction result, the model weight of the target water quality prediction model corresponding to the screened water quality prediction result is determined;
[0026] According to the model weight of the target water quality prediction model corresponding to the screened water quality prediction result, the weighted sum processing is performed on the screened water quality prediction result to obtain the target water quality prediction result of the tail water to be tested.
[0027] In one of the embodiments, the water quality data is combined to obtain the combined data corresponding to the water quality data, comprising:
[0028] The target combination mode corresponding to the water quality data is determined;
[0029] The water quality data to be combined is determined according to the target combination mode;
[0030] The water quality data and the water quality data to be combined are combined to obtain the combined data corresponding to the water quality data.
[0031] In one of the embodiments, the pre-trained water quality prediction model is trained in the following manner:
[0032] Sample water quality data of sample tail water is obtained;
[0033] The sample water quality data is combined to obtain sample combined data corresponding to the sample water quality data;
[0034] The sample water quality data and the sample combined data are subjected to feature extraction processing to obtain a feature vector of the sample water quality data and a feature vector of the sample combined data;
[0035] The feature vector of the sample water quality data and the feature vector of the sample combined data are spliced to obtain a sample target feature vector of the sample water quality data;
[0036] The sample target feature vector is input into a water quality prediction model to be trained to obtain a water quality prediction result of the sample tail water;
[0037] The actual water quality result of the sample tail water is obtained, and the water quality prediction model to be trained is iteratively trained according to the difference between the water quality prediction result of the sample tail water and the actual water quality result of the sample tail water, to obtain the pre-trained water quality prediction model.
[0038] In one of the embodiments, after the screened water quality prediction results are weighted and summed to obtain the target water quality prediction result of the tail water to be measured, the method further comprises:
[0039] The target water quality prediction result is input into a pre-trained water quality grade prediction model to obtain prediction probabilities of the target water quality prediction result under a plurality of water quality grades;
[0040] The water quality grade with the largest prediction probability is selected from the plurality of water quality grades as the target water quality grade of the tail water to be measured;
[0041] generate tail water treatment instructions corresponding to the to-be-tested tail water according to the target water quality level;
[0042] process the to-be-tested tail water according to the tail water treatment instructions.
[0043] In a second aspect, the present application further provides a water quality prediction device, comprising:
[0044] a data acquisition module configured to acquire water quality data of to-be-tested tail water;
[0045] a data combination module configured to combine the water quality data to obtain combined data corresponding to the water quality data;
[0046] a feature extraction module configured to extract features of the water quality data and the combined data to obtain a feature vector of the water quality data and a feature vector of the combined data;
[0047] a vector splicing module configured to splice the feature vector of the water quality data and the feature vector of the combined data to obtain a target feature vector of the water quality data;
[0048] a model prediction module configured to input the target feature vector into a plurality of pre-trained water quality prediction models to obtain a water quality prediction result of the to-be-tested tail water output by each pre-trained water quality prediction model and a prediction probability corresponding to each water quality prediction result;
[0049] a result screening module configured to screen, from the each water quality prediction result, a water quality prediction result with a prediction probability greater than a preset prediction probability as a screened water quality prediction result;
[0050] a result determination module configured to perform weighted summation processing on the screened water quality prediction result to obtain a target water quality prediction result of the to-be-tested tail water.
[0051] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the following steps when executing the computer program:
[0052] acquire water quality data of to-be-tested tail water;
[0053] combine the water quality data to obtain combined data corresponding to the water quality data;
[0054] extract features of the water quality data and the combined data to obtain a feature vector of the water quality data and a feature vector of the combined data;
[0055] concatenate the feature vector of the water quality data and the feature vector of the combined data to obtain a target feature vector of the water quality data;
[0056] input the target feature vector into a plurality of pre-trained water quality prediction models to obtain a water quality prediction result of the to-be-tested tail water output by each pre-trained water quality prediction model and a prediction probability corresponding to each water quality prediction result;
[0057] from the each water quality prediction result, filter out a water quality prediction result corresponding to a prediction probability greater than a preset prediction probability as a filtered water quality prediction result;
[0058] perform weighted summation processing on the filtered water quality prediction result to obtain a target water quality prediction result of the to-be-tested tail water.
[0059] In a fourth aspect, the present application also provides a computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the following steps:
[0060] obtain water quality data of to-be-tested tail water;
[0061] perform combined processing on the water quality data to obtain combined data corresponding to the water quality data;
[0062] perform feature extraction processing on the water quality data and the combined data to obtain a feature vector of the water quality data and a feature vector of the combined data;
[0063] concatenate the feature vector of the water quality data and the feature vector of the combined data to obtain a target feature vector of the water quality data;
[0064] input the target feature vector into a plurality of pre-trained water quality prediction models to obtain a water quality prediction result of the to-be-tested tail water output by each pre-trained water quality prediction model and a prediction probability corresponding to each water quality prediction result;
[0065] from the each water quality prediction result, filter out a water quality prediction result corresponding to a prediction probability greater than a preset prediction probability as a filtered water quality prediction result;
[0066] perform weighted summation processing on the filtered water quality prediction result to obtain a target water quality prediction result of the to-be-tested tail water.
[0067] In a fifth aspect, the present application also provides a computer program product comprising a computer program, the computer program being executed by a processor to implement the following steps:
[0068] obtain water quality data of to-be-tested tail water;
[0069] combine the water quality data to obtain combined data corresponding to the water quality data;
[0070] extract features from the water quality data and the combined data to obtain a feature vector of the water quality data and a feature vector of the combined data;
[0071] splice the feature vector of the water quality data and the feature vector of the combined data to obtain a target feature vector of the water quality data;
[0072] input the target feature vector into a plurality of pre-trained water quality prediction models to obtain a water quality prediction result of the to-be-tested tail water output by each pre-trained water quality prediction model and a prediction probability corresponding to each water quality prediction result;
[0073] from the each water quality prediction result, filter out the water quality prediction result corresponding to the prediction probability greater than the preset prediction probability as the filtered water quality prediction result;
[0074] perform weighted summation processing on the filtered water quality prediction result to obtain a target water quality prediction result of the to-be-tested tail water.
[0075] The water quality prediction method, device, computer device, storage medium and computer program product obtain water quality data of the to-be-tested tail water, perform combination processing on the water quality data to obtain combined data corresponding to the water quality data, perform feature extraction processing on the water quality data and the combined data to obtain a feature vector of the water quality data and a feature vector of the combined data, perform splicing processing on the feature vector of the water quality data and the feature vector of the combined data to obtain a target feature vector of the water quality data, input the target feature vector into a plurality of pre-trained water quality prediction models to obtain a water quality prediction result of the to-be-tested tail water output by each pre-trained water quality prediction model and a prediction probability corresponding to each water quality prediction result, filter, from the water quality prediction results, a water quality prediction result with a prediction probability greater than a preset prediction probability as a filtered water quality prediction result, and perform weighted summation processing on the filtered water quality prediction result to obtain a target water quality prediction result of the to-be-tested tail water. In this way, in the process of detecting water quality, the target feature vector of the water quality data is obtained by combining the water quality data of the to-be-tested tail water and the combined data corresponding to the water quality data, the target feature vector is predicted by the plurality of pre-trained water quality prediction models, and the plurality of water quality prediction results are filtered and processed to obtain the target water quality prediction result of the to-be-tested tail water, thereby improving the accuracy and reliability of the overall prediction. Moreover, the entire process does not require manual intervention, avoids the subjective factors in the manual detection method, and is less likely to make errors, thereby further improving the detection accuracy of the water quality. BRIEF DESCRIPTION OF DRAWINGS
[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other related drawings can be obtained by those skilled in the art without creative labor.
[0077] FIG. 1 is a flowchart of a water quality prediction method in an embodiment;
[0078] FIG. 2 is a flowchart of a water quality prediction method in another embodiment;
[0079] FIG. 3 is a schematic diagram of a water quality comprehensive soft-sensing method in an embodiment;
[0080] FIG. 4 is a structural schematic diagram of a breeding tail water treatment system in an embodiment;
[0081] FIG. 5 is a schematic diagram of training and prediction results of a LightGBM COD soft measurement model based on Bayesian algorithm optimization on a training set in an embodiment;
[0082] FIG. 6 is a schematic diagram of training and prediction results of a LightGBM TN soft measurement model based on Bayesian algorithm optimization on a training set in an embodiment;
[0083] FIG. 7 is a schematic diagram of training and prediction results of a LightGBM TP soft measurement model based on Bayesian algorithm optimization on a training set in an embodiment;
[0084] FIG. 8 is a structural block diagram of a water quality prediction device in an embodiment;
[0085] FIG. 9 is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION
[0086] In order to make the purposes, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0087] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.
[0088] In an exemplary embodiment, as shown in FIG. 1, a water quality prediction method is provided, and the present embodiment is exemplarily illustrated by applying the method to a server; it can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is realized through the interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, notebook computers, smart phones and tablet computers; the server can be realized by an independent server or a server cluster composed of multiple servers. In the present embodiment, the method includes the following steps:
[0089] In step S101, water quality data of the to-be-tested tail water is acquired.
[0090] The to-be-tested tail water refers to treated wastewater that needs to be tested for water quality.
[0091] The water quality data refers to data associated with the tail water to be measured, such as inlet water DO (Dissolved Oxygen), inlet water SS (Suspended Solids), inlet water pH (potential of hydrogen), inlet water T (temperature), inlet water COD (Chemical Oxygen Demand), inlet water TP (Total Phosphorus), inlet water TN (Total Nitrogen), inlet water NH4 + -N (Ammonium Nitrogen), and the like.
[0092] For example, the server obtains, as the water quality data of the tail water to be measured, water quality data of the tail water to be measured under preset indexes through the water treatment system associated with the tail water to be measured in response to a detection instruction for the tail water to be measured.
[0093] In step S102, the water quality data is combined to obtain combined data corresponding to the water quality data.
[0094] The combined data refers to data obtained by combining the water quality data, such as a ratio of inlet water COD to inlet water TP, a difference between inlet water TN and inlet water NH4 + -N, and a correlation coefficient between inlet water SS and inlet water DO.
[0095] For example, the server determines the water quality data to be combined corresponding to each water quality data, and then combines the water quality data and the water quality data to be combined corresponding to each water quality data to obtain the combined data corresponding to each water quality data.
[0096] In step S103, the water quality data and the combined data are subjected to feature extraction processing to obtain a feature vector of the water quality data and a feature vector of the combined data.
[0097] The feature extraction processing refers to a processing procedure for obtaining a feature vector corresponding to data.
[0098] The feature vector of the water quality data is used to represent a representation vector of the water quality data.
[0099] The feature vector of the combined data is used to represent a representation vector of the combined data.
[0100] Exemplarily, the server determines the data category of the water quality data, and determines, according to a corresponding relationship between the data category and the feature extraction model, the feature extraction model corresponding to the data category of the water quality data as the feature extraction model corresponding to the water quality data; then, the server performs feature extraction processing on the water quality data by using the feature extraction model corresponding to the water quality data, to obtain the feature vector of the water quality data; then, the server determines the data category of the combined data, and determines, according to a corresponding relationship between the data category and the feature extraction model, the feature extraction model corresponding to the data category of the combined data as the feature extraction model corresponding to the combined data; then, the server performs feature extraction processing on the combined data by using the feature extraction model corresponding to the combined data, to obtain the feature vector of the combined data.
[0101] In step S104, the feature vector of the water quality data and the feature vector of the combined data are spliced to obtain a target feature vector of the water quality data.
[0102] The splicing processing refers to a processing process of combining the feature vectors in a specific manner.
[0103] The target feature vector refers to a feature vector obtained by splicing the feature vector of the water quality data and the feature vector of the combined data.
[0104] Exemplarily, the server determines the importance of the water quality data and the importance of the combined data; then, the server determines the splicing order of the feature vector of the water quality data and the feature vector of the combined data according to the importance of the water quality data and the importance of the combined data; then, the server splices the feature vector of the water quality data and the feature vector of the combined data according to the splicing order to obtain the target feature vector of the water quality data.
[0105] In step S105, the target feature vector is input into a plurality of pre-trained water quality prediction models to obtain the water quality prediction result of the to-be-tested tail water output by each pre-trained water quality prediction model and the prediction probability corresponding to each water quality prediction result.
[0106] The water quality prediction model refers to a network model capable of obtaining the water quality prediction result of the to-be-tested tail water and the prediction probability corresponding to the water quality prediction result by using the target feature vector, such as a LightGBM (Light Gradient Boosting Machine) model, an XGBoost (eXtreme Gradient Boosting) model and an RF (Random Forest) model.
[0107] The water quality prediction result is used to represent the water quality prediction of the tail water to be detected, including data such as effluent COD, effluent TP and effluent TN.
[0108] The prediction probability is used to represent the possibility that the water quality prediction model considers the output water quality prediction result to be correct, such as 90%.
[0109] For example, the server obtains a plurality of pre-trained water quality prediction models, and inputs the target feature vector into the plurality of pre-trained water quality prediction models to obtain the water quality prediction result corresponding to the target feature vector output by each pre-trained water quality prediction model as the water quality prediction result of the tail water to be detected. Then, the server takes the prediction probability of the water quality prediction result corresponding to the target feature vector output by each pre-trained water quality prediction model as the prediction probability corresponding to each water quality prediction result.
[0110] In step S106, the water quality prediction result corresponding to the prediction probability greater than the preset prediction probability is selected from each water quality prediction result as the screened water quality prediction result.
[0111] The preset prediction probability refers to a preset prediction probability threshold, such as 80%.
[0112] The screened water quality prediction result refers to the water quality prediction result corresponding to the prediction probability greater than the preset prediction probability.
[0113] For example, the server selects the water quality prediction result corresponding to the prediction probability greater than the preset prediction probability from each water quality prediction result as the screened water quality prediction result. For example, the prediction probability corresponding to the water quality prediction result 1 is 90%, the prediction probability corresponding to the water quality prediction result 2 is 85%, the prediction probability corresponding to the water quality prediction result 3 is 70%, and the prediction probability corresponding to the water quality prediction result 4 is 80%. When the preset prediction probability is 80%, the water quality prediction result 1 and the water quality prediction result 2 are taken as the screened water quality prediction result.
[0114] In step S107, the screened water quality prediction result is weighted and summed to obtain the target water quality prediction result of the tail water to be detected.
[0115] The target water quality prediction result refers to the screened water quality prediction result after the weighted sum processing.
[0116] Exemplarily, the server performs weighted summation processing on the screened water quality prediction results to obtain processed screened water quality prediction results as the target water quality prediction result of the to-be-tested tail water; for example, the effluent COD in the water quality prediction result 1 is 9.30 mg / L, the effluent TP is 0.46 mg / L, and the effluent TN is 1.22 mg / L, the effluent COD in the water quality prediction result 2 is 8.70 mg / L, the effluent TP is 0.44 mg / L, and the effluent TN is 1.08 mg / L, and the effluent COD in the target water quality prediction result of the to-be-tested tail water is 9.00 mg / L, the effluent TP is 0.45 mg / L, and the effluent TN is 1.15 mg / L.
[0117] In the above water quality prediction method, the water quality data of the to-be-tested tail water is first obtained, and the water quality data is combined to obtain combined data corresponding to the water quality data. Then, the water quality data and the combined data are subjected to feature extraction processing to obtain a feature vector of the water quality data and a feature vector of the combined data. The feature vector of the water quality data and the feature vector of the combined data are spliced to obtain a target feature vector of the water quality data. Next, the target feature vector is input into a plurality of pre-trained water quality prediction models to obtain water quality prediction results of the to-be-tested tail water output by each pre-trained water quality prediction model and a prediction probability corresponding to each water quality prediction result. Then, from each water quality prediction result, a water quality prediction result with a prediction probability greater than a preset prediction probability is selected as a screened water quality prediction result. Finally, the screened water quality prediction results are subjected to weighted summation processing to obtain a target water quality prediction result of the to-be-tested tail water. In this way, in the process of detecting the water quality, the water quality data of the to-be-tested tail water and the combined data corresponding to the water quality data are combined to obtain the target feature vector of the water quality data. The target feature vector is predicted by the plurality of pre-trained water quality prediction models, and the plurality of water quality prediction results are screened and processed to obtain the target water quality prediction result of the to-be-tested tail water, thereby improving the accuracy and reliability of the overall prediction. Moreover, the entire process does not require manual intervention, avoiding the subjective factors in the manual detection method, which is prone to errors and leads to low detection accuracy of the water quality, thereby further improving the detection accuracy of the water quality.
[0118] In one exemplary embodiment, the step S101 of obtaining the water quality data of the to-be-tested tail water specifically includes the following contents: determining the source information of the to-be-tested tail water; querying the correspondence between the source information and the preset index according to the source information of the to-be-tested tail water to obtain the preset index corresponding to the source information of the to-be-tested tail water as the preset index of the to-be-tested tail water; and obtaining the water quality data of the to-be-tested tail water under the preset index as the water quality data of the to-be-tested tail water.
[0119] The source information refers to the origin of the to-be-tested tail water, such as a fish pond, a factory, etc.
[0120] The corresponding relationship between the source information and the preset indicators is used to represent the association information between the source information and the preset indicators. For example, when the source information is a fish pond, the corresponding preset indicators are indicators a, b, and c, and when the source information is a factory, the corresponding preset indicators are indicators a and d.
[0121] The preset indicators refer to pre-set indicator information, such as influent DO, influent SS, influent pH, influent T, influent COD, influent TP, influent TN, or influent NH4 + -N indicators.
[0122] Exemplarily, the server determines the source information of the tail water to be tested; then, the server queries the corresponding relationship between the source information and the preset indicators according to the source information of the tail water to be tested, obtains the preset indicators corresponding to the source information of the tail water to be tested as the preset indicators of the tail water to be tested; and then, the server obtains the water quality data of the tail water to be tested under the preset indicators, and takes these water quality data as the water quality data of the tail water to be tested.
[0123] In this embodiment, by determining the preset indicators corresponding to the source information of the tail water to be tested and obtaining the water quality data of the tail water to be tested under the preset indicators, the use of inappropriate indicators for data acquisition is avoided, irrelevant data interference is avoided, the obtained water quality data more accurately reflects the real situation of the tail water to be tested, and the determination accuracy of the water quality data of the tail water to be tested is improved.
[0124] In one exemplary embodiment, the step S103 of performing feature extraction processing on the water quality data and the combined data to obtain the feature vector of the water quality data and the feature vector of the combined data includes the following contents: respectively pre-process the water quality data and the combined data to obtain pre-processed water quality data and pre-processed combined data; input the pre-processed water quality data as main data and the pre-processed combined data as auxiliary data into a feature extraction model to obtain the feature vector of the water quality data; and input the pre-processed combined data as main data and the pre-processed water quality data as auxiliary data into the feature extraction model to obtain the feature vector of the combined data.
[0125] The pre-processed water quality data refers to the water quality data after pre-processing.
[0126] The pre-processed combined data refers to the combined data after pre-processing.
[0127] The main data can refer to data with a larger corresponding weight.
[0128] The auxiliary data can refer to data with a smaller corresponding weight.
[0129] The feature extraction model refers to a network model capable of obtaining a feature vector corresponding to data, such as a transformer model.
[0130] For example, the server respectively pre-processes the water quality data and the combined data to obtain pre-processed water quality data and pre-processed combined data. For example, the server respectively performs outlier identification and elimination and data normalization processing on the water quality data and the combined data to obtain pre-processed water quality data and pre-processed combined data. Then, the server inputs the pre-processed water quality data as main data and the pre-processed combined data as auxiliary data into the feature extraction model for feature extraction processing to obtain a feature vector of the water quality data. For example, the server determines that the first weight corresponding to the pre-processed water quality data is 0.8 and the second weight corresponding to the pre-processed combined data is 0.2, and performs fusion processing on the pre-processed water quality data and the pre-processed combined data according to the first weight and the second weight to obtain first fusion data, and inputs the first fusion data into the feature extraction model for feature extraction processing to obtain a feature vector of the first fusion data as the feature vector of the water quality data. Then, the server inputs the pre-processed combined data as main data and the pre-processed water quality data as auxiliary data into the feature extraction model to obtain a feature vector of the combined data. For example, the server determines that the third weight corresponding to the pre-processed combined data is 0.8 and the fourth weight corresponding to the pre-processed water quality data is 0.2, and performs fusion processing on the pre-processed combined data and the pre-processed water quality data according to the third weight and the fourth weight to obtain second fusion data, and inputs the second fusion data into the feature extraction model for feature extraction processing to obtain a feature vector of the second fusion data as the feature vector of the combined data.
[0131] In this embodiment, the water quality data and the combined data are pre-processed to remove noise, outliers and inconsistencies, which is beneficial to improve the quality of the water quality data and the combined data. Moreover, the feature extraction processing is performed by combining the pre-processed combined data and the pre-processed water quality data, which can more comprehensively utilize the information in the data, and is beneficial to improve the quality and representativeness of the feature vector.
[0132] In an exemplary embodiment, the step S107 of performing weighted summation processing on the screened water quality prediction result to obtain the target water quality prediction result of the tail water to be measured includes the following contents: determining the target water quality prediction model corresponding to the screened water quality prediction result; obtaining the prediction accuracy and prediction efficiency of the target water quality prediction model corresponding to the screened water quality prediction result; determining the model weight of the target water quality prediction model corresponding to the screened water quality prediction result according to the prediction accuracy and prediction efficiency of the target water quality prediction model corresponding to the screened water quality prediction result; and performing weighted summation processing on the screened water quality prediction result according to the model weight of the target water quality prediction model corresponding to the screened water quality prediction result to obtain the target water quality prediction result of the tail water to be measured.
[0133] The target water quality prediction model refers to a water quality prediction model outputting the screened water quality prediction result.
[0134] The prediction accuracy is used to represent the closeness of the predicted value and the actual value of the target water quality prediction model.
[0135] The prediction efficiency is used to represent the time required for the target water quality prediction model to make a prediction.
[0136] The model weight is used to represent the importance of the target water quality prediction model.
[0137] For example, the server determines the water quality prediction model corresponding to the screened water quality prediction result from each water quality prediction model as the target water quality prediction model; then, the server obtains the prediction accuracy and prediction efficiency of the target water quality prediction model; then, the server determines the model weight of the target water quality prediction model according to the prediction accuracy and prediction efficiency of the target water quality prediction model; for example, the prediction accuracy of the target water quality prediction model A is 0.9, the standardized prediction efficiency is 0.7; the prediction accuracy of the target water quality prediction model B is 0.8, the standardized prediction efficiency is 0.9, in the case of the importance coefficient of the prediction accuracy being 0.6 and the importance coefficient of the prediction efficiency being 0.4, the model weight of the target water quality prediction model A is 0.9*0.6+0.7*0.4=0.82, and the model weight of the target water quality prediction model B is 0.8*0.6+0.9*0.4=0.84; then, the server performs weighted summation processing on the screened water quality prediction result according to the model weight of the target water quality prediction model corresponding to the screened water quality prediction result to obtain the processed screened water quality prediction result as the target water quality prediction result of the tail water to be measured.
[0138] In this embodiment, by simultaneously considering the prediction accuracy and the prediction efficiency, the model weight of the target water quality prediction model is determined, so that the water quality prediction result corresponding to the target water quality prediction model with greater prediction accuracy and prediction efficiency occupies a greater proportion in the weighted summation processing, which is beneficial to improve the prediction accuracy of the target water quality prediction result of the to-be-tested tail water.
[0139] In one exemplary embodiment, the step S102 of combining the water quality data to obtain the combined data corresponding to the water quality data includes the following contents: determining a target combination mode corresponding to the water quality data; determining to-be-combined water quality data corresponding to the water quality data according to the target combination mode; and combining the water quality data and the to-be-combined water quality data to obtain the combined data corresponding to the water quality data.
[0140] The target combination mode is used to represent a rule corresponding to the specific combination processing of the water quality data. For example, the influent COD and the influent TP can be combined to obtain the ratio of the influent COD to the influent TP, the influent TN and the influent NH4 + -N can be combined to obtain the difference between the influent TN and the influent NH4 + -N, the influent SS and the influent DO can be combined to obtain the correlation coefficient of the influent SS and the influent DO, and the like.
[0141] The to-be-combined water quality data refers to the water quality data that can be combined. For example, the to-be-combined water quality data corresponding to the influent COD is the influent TP, the to-be-combined water quality data corresponding to the influent TN is the influent NH4 + -N, and the to-be-combined water quality data corresponding to the influent SS is the influent DO.
[0142] Exemplarily, the server determines the target combination mode corresponding to the water quality data; then, the server determines the to-be-combined water quality data corresponding to each water quality data according to the target combination mode; then, the server combines each water quality data and the to-be-combined water quality data corresponding to the water quality data to obtain the combined data corresponding to the water quality data, and further obtains the combined data corresponding to each water quality data.
[0143] In this embodiment, the target combination mode can purposefully integrate and transform the original water quality data, thereby creating more expressive combined data; and through the specific target combination mode, the defect of increasing the complexity of subsequent data analysis caused by a large amount of meaningless or irrelevant combined data is avoided.
[0144] In one exemplary embodiment, the water quality prediction method provided by the present application further comprises a training step of a pre-trained water quality prediction model, specifically comprising the following contents: obtaining sample tail water sample water quality data; combining the sample water quality data to obtain sample combination data corresponding to the sample water quality data; performing feature extraction processing on the sample water quality data and the sample combination data to obtain a feature vector of the sample water quality data and a feature vector of the sample combination data; performing splicing processing on the feature vector of the sample water quality data and the feature vector of the sample combination data to obtain a sample target feature vector of the sample water quality data; inputting the sample target feature vector into the water quality prediction model to be trained to obtain a water quality prediction result of the sample tail water; obtaining an actual water quality result of the sample tail water, and according to the difference between the water quality prediction result of the sample tail water and the actual water quality result of the sample tail water, iteratively training the water quality prediction model to be trained to obtain the pre-trained water quality prediction model.
[0145] Among them, the sample tail water refers to the tail water used for training the water quality prediction model to be trained.
[0146] Among them, the sample water quality data refers to the water quality data of the sample tail water.
[0147] Among them, the sample combination data refers to the combination data obtained by combining the sample water quality data.
[0148] Among them, the sample target feature vector refers to the feature vector obtained by splicing the feature vector of the sample water quality data and the feature vector of the sample combination data.
[0149] Among them, the water quality prediction result of the sample tail water refers to the predicted value of the sample tail water output by the water quality prediction model to be trained.
[0150] Among them, the actual water quality result of the sample tail water refers to the actual value of the sample tail water output by the water quality prediction model to be trained.
[0151] Exemplarily, the server obtains water quality data of the sample tail water from the database as sample water quality data in response to a model training instruction for the water quality prediction model to be trained; then, the server performs combination processing on the sample water quality data to obtain combination data corresponding to the sample water quality data as sample combination data; then, the server inputs the sample water quality data and the sample combination data into the feature extraction model respectively, and performs feature extraction processing on the sample water quality data and the sample combination data by the feature extraction model to obtain a feature vector of the sample water quality data and a feature vector of the sample combination data; then, the server performs splicing processing on the feature vector of the sample water quality data and the feature vector of the sample combination data to obtain a target feature vector of the sample water quality data as a sample target feature vector; then, the server inputs the sample target feature vector into the water quality prediction model to be trained to obtain a water quality prediction result of the sample tail water; then, the server obtains an actual water quality result of the sample tail water, and obtains a loss value according to a difference between the water quality prediction result of the sample tail water and the actual water quality result of the sample tail water; then, the server adjusts model parameters of the water quality prediction model to be trained according to the loss value; then, the server re-trains the water quality prediction model with the adjusted model parameters until a loss value obtained by the trained water quality prediction model is less than a loss value threshold, and then stops the training, and takes the trained water quality prediction model as a pre-trained water quality prediction model.
[0152] In the embodiment, by pre-training the water quality prediction model, it is convenient to predict the target water quality prediction result of the to-be-tested tail water after obtaining the water quality data of the to-be-tested tail water in actual application; moreover, the water quality prediction model receives new data in each iteration, and performs internal improvement and optimization of the model, so as to be more effective in prediction, and to be beneficial to improving the prediction accuracy of the water quality prediction model.
[0153] In one exemplary embodiment, after the step S107, that is, after performing weighted sum processing on the screened water quality prediction results to obtain the target water quality prediction result of the to-be-tested tail water, the step S107 specifically includes the following contents: inputting the target water quality prediction result into the pre-trained water quality grade prediction model to obtain prediction probabilities of the target water quality prediction result under multiple water quality grades; screening a water quality grade with the largest prediction probability from the multiple water quality grades as a target water quality grade of the to-be-tested tail water; generating a tail water processing instruction corresponding to the to-be-tested tail water according to the target water quality grade; and processing the to-be-tested tail water according to the tail water processing instruction.
[0154] The water quality grade prediction model refers to a network model capable of obtaining prediction probabilities of a target water quality prediction result under multiple water quality grades by using the target water quality prediction result.
[0155] The water quality level is used to represent the level information corresponding to the water body quality, such as first-class water quality, second-class water quality, unqualified water quality, etc.
[0156] The target water quality level refers to the water quality level of the tail water to be measured.
[0157] The tail water treatment instruction refers to the instruction information corresponding to the treatment of the tail water to be measured.
[0158] For example, the server inputs the target water quality prediction result into the pre-trained water quality level prediction model, obtains the multiple water quality levels corresponding to the target water quality prediction result and the prediction probability of the target water quality prediction result under the multiple water quality levels through the water quality level prediction model, selects the water quality level with the maximum prediction probability from the multiple water quality levels as the target water quality level of the tail water to be measured, generates the tail water treatment instruction corresponding to the target water quality level as the tail water treatment instruction corresponding to the tail water to be measured according to the target water quality level, and processes the tail water to be measured according to the tail water treatment instruction. For example, if the target water quality level is first-class water quality, the tail water treatment instruction corresponding to the tail water to be measured is “periodically check and maintain the tail water discharge pipeline”, if the target water quality level is second-class water quality, the tail water treatment instruction corresponding to the tail water to be measured is “investigate and control the source that may affect the tail water quality”, and if the target water quality level is unqualified water quality, the tail water treatment instruction corresponding to the tail water to be measured is “immediately stop discharging the tail water”. Finally, the server determines the tail water treatment equipment corresponding to the tail water treatment instruction, and sends the tail water treatment instruction to the tail water treatment equipment, so that the tail water treatment equipment processes the tail water to be measured.
[0159] In this embodiment, the corresponding tail water treatment instruction is generated according to different target water quality levels, which is equivalent to adopting different processing methods and technologies for different levels of water quality problems, so as to improve the pertinence and effectiveness of the tail water treatment to be measured, and is beneficial to improve the processing accuracy of the tail water to be measured.
[0160] In one exemplary embodiment, as shown in FIG. 2, another water quality prediction method is provided, which is applied to the server for illustration, and specifically includes the following steps:
[0161] Step S201, determine the source information of the tail water to be measured; according to the source information of the tail water to be measured, query the corresponding relationship between the source information and the preset index, obtain the preset index corresponding to the source information of the tail water to be measured as the preset index of the tail water to be measured, and obtain the water quality data of the tail water to be measured under the preset index as the water quality data of the tail water to be measured.
[0162] In step S202, a target combination mode corresponding to the water quality data is determined; according to the target combination mode, to-be-combined water quality data corresponding to the water quality data is determined; the water quality data and the to-be-combined water quality data are combined to obtain combined data corresponding to the water quality data.
[0163] In step S203, the water quality data and the combined data are respectively preprocessed to obtain preprocessed water quality data and preprocessed combined data.
[0164] In step S204, the preprocessed water quality data is taken as main data, and the preprocessed combined data is taken as auxiliary data, and is input into a feature extraction model to obtain a feature vector of the water quality data; the preprocessed combined data is taken as main data, and the preprocessed water quality data is taken as auxiliary data, and is input into the feature extraction model to obtain a feature vector of the combined data.
[0165] In step S205, the feature vector of the water quality data and the feature vector of the combined data are spliced to obtain a target feature vector of the water quality data.
[0166] In step S206, the target feature vector is input into a plurality of pre-trained water quality prediction models to obtain a water quality prediction result of the to-be-tested tail water output by each pre-trained water quality prediction model and a prediction probability corresponding to each water quality prediction result.
[0167] In step S207, from each water quality prediction result, a water quality prediction result with a prediction probability greater than a preset prediction probability is selected as a screened water quality prediction result.
[0168] In step S208, a target water quality prediction model corresponding to the screened water quality prediction result is determined; a prediction accuracy and a prediction efficiency of the target water quality prediction model corresponding to the screened water quality prediction result are obtained.
[0169] In step S209, a model weight of the target water quality prediction model corresponding to the screened water quality prediction result is determined according to the prediction accuracy and the prediction efficiency of the target water quality prediction model corresponding to the screened water quality prediction result.
[0170] In step S210, the screened water quality prediction result is weighted and summed according to the model weight of the target water quality prediction model corresponding to the screened water quality prediction result to obtain a target water quality prediction result of the to-be-tested tail water.
[0171] In the water quality prediction method, in the process of detecting the water quality, the water quality data of the to-be-detected tail water and the combination data corresponding to the water quality data are combined to obtain a target feature vector of the water quality data, the target feature vector is predicted through a plurality of pre-trained water quality prediction models, and a plurality of water quality prediction results are screened and processed to obtain a target water quality prediction result of the to-be-detected tail water, so that the accuracy and reliability of the overall prediction are improved. Moreover, the whole process does not need manual intervention, avoids the subjective factors in the manual detection mode, and is prone to errors, thereby improving the detection accuracy of the water quality.
[0172] In one exemplary embodiment, in order to more clearly illustrate the water quality prediction method provided by the embodiments of the present application, the water quality prediction method is specifically described below with one specific embodiment. In one embodiment, the present application also provides a LightGBM water quality soft measurement method for water treatment process based on Bayesian algorithm optimization. In the process of detecting the water quality, a plurality of real-time water quality data in the water treatment process read by the water quality online sensor form an original data set, and an isolation forest algorithm is used to identify and eliminate the outliers in the original data set. Subsequently, a recursive feature elimination algorithm is used for input feature selection, and the selected data set is used as a model sample data set. The sample data set is used to train the LightGBM model. Subsequently, a Bayesian optimization algorithm is used to optimize the model hyperparameters of the LightGBM, to obtain the best hyperparameter combination of the LightGBM model. Finally, the optimal hyperparameter combination is substituted into the LightGBM algorithm to establish chemical oxygen demand, total nitrogen, and total phosphorus soft measurement models, respectively. The trained soft measurement models are stacked to form the water quality soft measurement method for the water treatment process. Specifically, the following contents are included:
[0173] As shown in FIG. 3, the method includes water quality online monitoring data acquisition, outlier identification and elimination, input variable feature elimination and model training, and model hyperparameter optimization. The water treatment system takes a certain aquaculture tail water treatment facility as an example. The structure diagram of the aquaculture tail water treatment system is shown in FIG. 4. The present embodiment specifically includes the following steps:
[0174] (1) Original data acquisition: collect and store the water quality indexes monitored online by the water treatment system every day as the model input data, and the water quality data detected manually as the model output data, to form an original database in a one-to-one correspondence according to the sampling time.
[0175] All data in the present embodiment are collected during the daily operation of the treatment device. The data of DO inf , SS inf are obtained from the water quality online monitoring instruments installed in the fishpond tail water treatment system, and the data of COD inf , TP infand TN inf The target variables such as COD, TN, and TP in the water quality during the treatment process were obtained by water quality analysts based on the national standard monitoring methods, and the unit is mg / L (milligrams per liter).
[0176] (2) Data Preprocessing: An algorithm is used to identify and remove outliers. By randomly selecting data features and random partitioning points within the feature range, a binary tree is gradually constructed. The path length of a sample in the tree is used to measure its degree of anomalousness; a shorter average path length indicates a higher degree of anomalousness. By randomly constructing multiple binary trees and calculating the average path length of each sample in these trees, samples with high anomalousness scores are ultimately considered outliers and removed. The data after outlier removal is normalized, limiting the processed data to the range [0,1]. The mathematical expression is:
[0177] Where x is the original data of the feature variable, x norm These are the normalized data of the feature variables, where max(x) and min(x) are the maximum and minimum values of x.
[0178] In this embodiment, data collection began immediately after the processing system started and stabilized, and continued for 89 days to complete the data collection required for the model, obtaining 152 sets of raw data. After outlier identification and removal, a total of 133 sets of valid raw data were obtained, and 133 sets of data were actually used for modeling, as shown in Table 1 below.
[0179] Table 1 shows some of the original data from the test set used.
[0180] (3) Data Feature Selection: The RFE (Recursive Feature Elimination) algorithm is used to optimize the selection of model variables. After determining the model algorithm, the model is trained using all feature subsets S, and the relative importance score R = {r1, r2, ..., r...} of each feature variable is calculated. k The features are then sorted, where k is the number of features in the feature subset S. The lowest-ranked feature is removed, thus reducing the size of the feature set. The remaining feature variables form a new feature subset S. i For each variable subset S i (i = 1, ..., S), retrain the model using the reduced feature set, recalculate the importance of each feature variable and rank them. Repeat the above steps, retraining the model on the reduced feature set in each iteration, and then setting the model R... 2 A value greater than 0.8 is used as the performance metric for stopping the iteration, and the feature set at the time of iteration termination is selected as the final feature selection result.
[0181] In this embodiment, RFE algorithm is used for model feature selection, normalized data is used to train LightGBM model, the number of times of all features as split point is counted, and the relative importance score R = {r1, r2, ···, r8} is calculated. The current LightGBM model accuracy is evaluated, the feature with the smallest importance in the current set is removed, and the model is re-iterated and trained and evaluated. The model R 2 > 0.8 as the performance indicator of iteration stop, and finally determines the input features as DO inf , SS inf , COD inf , TP inf , TN inf .
[0182] (4) Model learning and training: the model training process includes randomly dividing the model sample data set into training set and test set according to the ratio of 8:2, which are used for model training and testing respectively. The model uses 10-fold cross-validation on the training set, and no separate validation set is set. Bayesian algorithm is used to optimize the hyperparameters of LightGBM model. The hyperparameters to be found include learning rate learning_rate, tree depth max_depth, leaf node number num_leaves, and feature ratio feature_fraction. Gaussian process is used as a probability proxy model in the optimization process, and the acquisition function is PI function. The specific optimization steps are as follows:
[0183] S1, randomly initialize n0 points x init ={x0, x1, ···, x n0-1}.
[0184] S2, get the corresponding function value f(x init ), the initial point set D0 = {x init , f(x init )}.
[0185] S3, let t = 1, D t-1 = D0.
[0186] S4, when t ≤ N, execute step S5.
[0187] S5, according to the current evaluation point set, construct the proxy model g(x).
[0188] S6, based on g(x), maximize the acquisition function a(x|D), get the next evaluation point: x t = arg min a(x|D).
[0189] S7, get the function value f(x t ) of x t), and adding it to the current evaluation point set D t = D t-1 U {x t , f(x t )} t , and determining the loss value corresponding to x t according to the difference between the function value of x t and the actual value.
[0190] S8, t = t + 1, when t ≤ N, return to step S5, when t > N, end.
[0191] S9, output: screen out the x t corresponding to the minimum loss value in S7 and the function value f(x t ) corresponding to the x t , obtain the optimal candidate evaluation point: {x t , f(x t )}, and obtain the optimal hyperparameter set of the LightGBM model based on the optimal candidate evaluation point.
[0192] wherein the input is the number of initialization points n0, the maximum number of iterations N, the proxy model g(x), and the acquisition function a(x|D), and the output is the optimal candidate evaluation point: {x t , f(x t )}.
[0193] In this embodiment, the data set after feature selection is taken as a training set to train the LightGBM model, and the Bayesian algorithm is used to optimize the model hyperparameters. The hyperparameters to be optimized are learning rate (learning_rate), tree depth (max_depth), leaf node number (num_leaves), and feature ratio (feature_fraction). The search ranges of the hyperparameters are set as follows: learning rate is 0.5-1, tree depth is 1-20, leaf node number is 10-100, and feature ratio is 0.1-1. In this embodiment, the ten-fold cross-validation average precision is taken as the objective function of hyperparameter optimization, and the four hyperparameters are optimized. The number of iterations is set to 30 times. The optimal hyperparameter combination obtained by the above method is substituted into the model, and finally the COD, TN, and TP soft measurement models of the trained LightGBM model based on Bayesian optimization are obtained.
[0194] (5) Stack the COD, TN, TP soft measurement models of the LightGBM model based on Bayesian optimization trained in step (4), and combine the models in the order of the cumulative sum of the hyperparameters from small to large. The soft measurement model with the smallest cumulative sum of hyperparameters is placed at the top layer of the comprehensive system, and the one with the largest cumulative sum is placed at the bottom layer, forming a COD, TN, TP comprehensive soft measurement model. For the comprehensive soft measurement model, the input vector is X = [X1, X2, X3, X4, X5], representing DO inf , SS inf , COD inf , TP inf and TN inf , and the output vector is Y = [Y1, Y2, Y3], representing COD, TN and TP, respectively.
[0195] In this embodiment, the trained individual COD, TN, TP soft measurement models are stacked to form a COD, TN, TP comprehensive soft measurement method. The input vector of the stacked comprehensive soft measurement model system remains unchanged, and the output becomes a vector set composed of COD, TN and TP values, i.e. Y = [Y1, Y2, Y3], representing COD, TN and TP, respectively, thereby realizing the multi-input and multi-output process and function of the comprehensive soft measurement method.
[0196] The original data set after removing outliers by algorithm has a total of 133 groups of data as the model sample database. The LightGBM model based on Bayesian optimization is trained and tested for each water quality index. The model training and testing results are shown in FIGS. 5, 6 and 7. As can be seen from the figures, after Bayesian algorithm optimization, the correlation between the COD predicted value and the measured true value of the soft measurement model is 0.902, the correlation between the TN predicted value and the measured true value is 0.874, and the correlation between the TP predicted value and the measured true value is 0.887. The model prediction effect is good.
[0197] In the above embodiment, in the process of detecting water quality, the water quality data of the to-be-detected tail water and the corresponding combined data are combined to obtain the target feature vector of the water quality data, and then the target feature vector is predicted by a plurality of pre-trained water quality prediction models, and the plurality of water quality prediction results are screened and processed to obtain the target water quality prediction result of the to-be-detected tail water, thereby improving the accuracy and reliability of the overall prediction. Moreover, the entire process does not require human intervention, avoiding the subjective factors in manual detection, which can easily lead to errors and result in low water quality detection accuracy, thereby further improving the water quality detection accuracy.
[0198] It should be understood that although the steps in the flowcharts involved in the embodiments described above are shown in sequence according to the arrows, the steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of the steps is not strictly limited in sequence, and the steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the embodiments described above can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of the steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps.
[0199] Based on the same inventive concept, the embodiments of the present application also provide a water quality prediction device for implementing the above-mentioned water quality prediction method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more water quality prediction device embodiments provided below can refer to the limitations of the water quality prediction method described above, which will not be repeated here.
[0200] In an exemplary embodiment, as shown in FIG. 8, a water quality prediction device is provided, comprising: a data acquisition module 801, a data combination module 802, a feature extraction module 803, a vector splicing module 804, a model prediction module 805, a result screening module 806 and a result determination module 807, wherein:
[0201] The data acquisition module 801 is configured to acquire water quality data of the tail water to be measured.
[0202] The data combination module 802 is configured to combine the water quality data to obtain combined data corresponding to the water quality data.
[0203] The feature extraction module 803 is configured to perform feature extraction processing on the water quality data and the combined data to obtain a feature vector of the water quality data and a feature vector of the combined data.
[0204] The vector splicing module 804 is configured to splice the feature vector of the water quality data and the feature vector of the combined data to obtain a target feature vector of the water quality data.
[0205] The model prediction module 805 is configured to input the target feature vector into a plurality of pre-trained water quality prediction models to obtain a water quality prediction result of the tail water to be measured output by each pre-trained water quality prediction model, and a prediction probability corresponding to each water quality prediction result.
[0206] The result screening module 806 is configured to screen, from each water quality prediction result, a water quality prediction result corresponding to a prediction probability greater than a preset prediction probability, as a screened water quality prediction result.
[0207] The result determination module 807 is configured to perform weighted summation processing on the screened water quality prediction results, to obtain a target water quality prediction result of the to-be-tested tail water.
[0208] In an exemplary embodiment, the data acquisition module 801 is further configured to determine source information of the to-be-tested tail water; query a correspondence between the source information and the preset index according to the source information of the to-be-tested tail water, to obtain a preset index corresponding to the source information of the to-be-tested tail water as the preset index of the to-be-tested tail water; and acquire water quality data of the to-be-tested tail water under the preset index as the water quality data of the to-be-tested tail water.
[0209] In an exemplary embodiment, the feature extraction module 803 is further configured to respectively pre-process the water quality data and the combined data, to obtain pre-processed water quality data and pre-processed combined data; input the pre-processed water quality data as main data and the pre-processed combined data as auxiliary data into the feature extraction model, to obtain a feature vector of the water quality data; and input the pre-processed combined data as main data and the pre-processed water quality data as auxiliary data into the feature extraction model, to obtain a feature vector of the combined data.
[0210] In an exemplary embodiment, the result determination module 807 is further configured to determine a target water quality prediction model corresponding to the screened water quality prediction result; acquire a prediction accuracy and a prediction efficiency of the target water quality prediction model corresponding to the screened water quality prediction result; determine a model weight of the target water quality prediction model corresponding to the screened water quality prediction result according to the prediction accuracy and the prediction efficiency of the target water quality prediction model corresponding to the screened water quality prediction result; and perform weighted summation processing on the screened water quality prediction result according to the model weight of the target water quality prediction model corresponding to the screened water quality prediction result, to obtain the target water quality prediction result of the to-be-tested tail water.
[0211] In an exemplary embodiment, the data combination module 802 is further configured to determine a target combination mode corresponding to the water quality data; determine to-be-combined water quality data corresponding to the water quality data according to the target combination mode; and perform combination processing on the water quality data and the to-be-combined water quality data, to obtain combined data corresponding to the water quality data.
[0212] In an example embodiment, the water quality prediction device further comprises a model training module configured to: obtain sample water quality data of sample tail water; perform combination processing on the sample water quality data to obtain sample combination data corresponding to the sample water quality data; perform feature extraction processing on the sample water quality data and the sample combination data to obtain a feature vector of the sample water quality data and a feature vector of the sample combination data; perform splicing processing on the feature vector of the sample water quality data and the feature vector of the sample combination data to obtain a sample target feature vector of the sample water quality data; input the sample target feature vector into a water quality prediction model to be trained to obtain a water quality prediction result of the sample tail water; obtain an actual water quality result of the sample tail water, and perform iterative training on the water quality prediction model to be trained according to a difference between the water quality prediction result of the sample tail water and the actual water quality result of the sample tail water to obtain a pre-trained water quality prediction model.
[0213] In an example embodiment, the water quality prediction device further comprises a tail water processing module configured to: input the target water quality prediction result into the pre-trained water quality grade prediction model to obtain a prediction probability of the target water quality prediction result under a plurality of water quality grades; select a water quality grade with the maximum prediction probability from the plurality of water quality grades as a target water quality grade of the tail water to be measured; generate a tail water processing instruction corresponding to the tail water to be measured according to the target water quality grade; and process the tail water to be measured according to the tail water processing instruction.
[0214] The above-mentioned modules of the water quality prediction device can be realized by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the above-mentioned modules.
[0215] In an example embodiment, a computer device, which can be a server, is provided, and an internal structure diagram of the computer device can be as shown in FIG. 9. The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store water quality data, combination data and the like. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with terminals outside through a network connection. The computer program is executed by the processor to implement a water quality prediction method.
[0216] Those skilled in the art can understand that the structure shown in FIG. 9 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0217] In an example embodiment, a computer device is also provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0218] In an example embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0219] In an example embodiment, a computer program product is provided, which includes a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0220] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0221] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.
[0222] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A water quality prediction method characterized by, The method comprises: acquiring water quality data of the tail water to be tested; combining the water quality data to obtain combined data corresponding to the water quality data; extracting features from the water quality data and the combined data to obtain a feature vector of the water quality data and a feature vector of the combined data; splicing the feature vector of the water quality data and the feature vector of the combined data to obtain a target feature vector of the water quality data; inputting the target feature vector into a plurality of pre-trained water quality prediction models to obtain water quality prediction results of the tail water to be tested output by each pre-trained water quality prediction model, and a prediction probability corresponding to each water quality prediction result; from the each water quality prediction result, filtering out a water quality prediction result corresponding to a prediction probability greater than a preset prediction probability as a filtered water quality prediction result; performing weighted summation processing on the filtered water quality prediction result to obtain a target water quality prediction result of the tail water to be tested.
2. The method of claim 1, wherein, The acquisition of the water quality data of the tail water to be tested comprises: determining source information of the tail water to be tested; querying a correspondence between the source information and a preset index according to the source information of the tail water to be tested to obtain a preset index corresponding to the source information of the tail water to be tested as a preset index of the tail water to be tested; acquiring water quality data of the tail water to be tested under the preset index as the water quality data of the tail water to be tested.
3. The method of claim 1, wherein, The feature extraction processing on the water quality data and the combined data to obtain the feature vector of the water quality data and the feature vector of the combined data comprises: respectively pre-processing the water quality data and the combined data to obtain pre-processed water quality data and pre-processed combined data; inputting the pre-processed water quality data as main data and the pre-processed combined data as auxiliary data into a feature extraction model to obtain the feature vector of the water quality data; inputting the pre-processed combined data as main data and the pre-processed water quality data as auxiliary data into a feature extraction model to obtain the feature vector of the combined data.
4. The method of claim 1, wherein, The weighted summation processing on the filtered water quality prediction result to obtain the target water quality prediction result of the tail water to be tested comprises: determining a target water quality prediction model corresponding to the filtered water quality prediction result; acquiring a prediction accuracy and a prediction efficiency of the target water quality prediction model corresponding to the filtered water quality prediction result; determining a model weight of the target water quality prediction model corresponding to the filtered water quality prediction result according to the prediction accuracy and the prediction efficiency of the target water quality prediction model corresponding to the filtered water quality prediction result; performing weighted summation processing on the filtered water quality prediction result according to the model weight of the target water quality prediction model corresponding to the filtered water quality prediction result to obtain the target water quality prediction result of the tail water to be tested.
5. The method of claim 1, wherein, The combination processing on the water quality data to obtain the combined data corresponding to the water quality data comprises: determining a target combination mode corresponding to the water quality data; determining to-be-combined water quality data corresponding to the water quality data according to the target combination mode; The water quality data and the to-be-combined water quality data are combined to obtain combined data corresponding to the water quality data.
6. The method of claim 1, wherein, The pre-trained water quality prediction model is trained in the following manner: obtain sample water quality data of the sample tail water; combine the sample water quality data to obtain sample combined data corresponding to the sample water quality data; extract features from the sample water quality data and the sample combined data to obtain a feature vector of the sample water quality data and a feature vector of the sample combined data; splice the feature vector of the sample water quality data and the feature vector of the sample combined data to obtain a sample target feature vector of the sample water quality data; input the sample target feature vector into a water quality prediction model to be trained to obtain a water quality prediction result of the sample tail water; obtain an actual water quality result of the sample tail water, and iteratively train the water quality prediction model to be trained according to a difference between the water quality prediction result of the sample tail water and the actual water quality result of the sample tail water to obtain the pre-trained water quality prediction model.
7. The method according to any one of claims 1 to 6, characterized in that, After the weighted sum processing of the filtered water quality prediction results is performed to obtain the target water quality prediction result of the to-be-tested tail water, the method further includes: input the target water quality prediction result into a pre-trained water quality level prediction model to obtain prediction probabilities of the target water quality prediction result under a plurality of water quality levels; from the plurality of water quality levels, filter out a water quality level with the maximum prediction probability as a target water quality level of the to-be-tested tail water; generate a tail water treatment instruction corresponding to the to-be-tested tail water according to the target water quality level; perform treatment on the to-be-tested tail water according to the tail water treatment instruction.
8. A water quality prediction device characterized by comprising: The device includes: a data acquisition module configured to acquire water quality data of a to-be-tested tail water; a data combination module configured to combine the water quality data to obtain combined data corresponding to the water quality data; a feature extraction module configured to extract features from the water quality data and the combined data to obtain a feature vector of the water quality data and a feature vector of the combined data; a vector splicing module configured to splice the feature vector of the water quality data and the feature vector of the combined data to obtain a target feature vector of the water quality data; a model prediction module configured to input the target feature vector into a plurality of pre-trained water quality prediction models to obtain water quality prediction results of the to-be-tested tail water output by each pre-trained water quality prediction model and prediction probabilities corresponding to each water quality prediction result; a result filtering module configured to filter out, from the each water quality prediction result, a water quality prediction result with a prediction probability greater than a preset prediction probability as a filtered water quality prediction result; a result determination module configured to perform weighted sum processing on the filtered water quality prediction results to obtain a target water quality prediction result of the to-be-tested tail water. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.
11. A computer program product comprising a computer program, characterized in that, The computer program, which when executed by the processor, implements the steps of the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Total phosphorus water quality soft measurement and predication method
CN107665363A
Effluent TP interval prediction method in wastewater treatment
CN108898220A
Water quality prediction method and device, electronic equipment and storage medium
CN113159456A
Sewage treatment water quality prediction method and system based on machine learning
CN115470702A
Wastewater treatment multi-output soft measurement method based on XGBoost online monitoring data driving
CN117012307A