Online Intelligent Alum Dosing System Based on Deep Data Cleaning
Through the framework of deep data cleaning and feedforward-feedback control system, a multivariate nonlinear statistical model is established, which solves the intelligence and stability of coagulant injection in tap water treatment, and realizes efficient PAC drug administration control, reduces drug consumption and supports the unmanned operation of digital water plants.
Patent Information
- Application Number
- CN202211196170.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-09-29
AI Technical Summary
The coagulant injection process in the existing tap water treatment relies on manual control, resulting in excessive injection and unstable water quality, which cannot meet the needs of digital smart water plants. The existing data cleaning methods cannot effectively process engineering abnormal data, affecting intelligent and unmanned operations.
Using an online intelligent alum administration system based on deep data cleaning, combined with the feedforward-feedback control system framework, through data mining and modeling, online data cleaning and model update, a multivariate nonlinear statistical model is established, engineering abnormal data is identified and processed, and PAC prediction administration is realized.
On the premise of ensuring the stability of water quality, reasonable intelligent drug administration is achieved, drug consumption is reduced, water plant economic operation level is improved, digital water plant operation is supported, and the stability and accurate drug administration of the system are ensured.
Smart Images

Figure CN116621295B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of optimized tap water treatment, and specifically to an online intelligent alum dosing system based on deep data cleaning. Background Art
[0002] With the improvement of the economic level, people's requirements for the quality of life are also getting higher and higher. As water is an essential part of people's daily life, water quality problems have become a hot issue of social concern. At present, although tap water treatment is a relatively mature technology, there is still a large room for improvement, especially in the coagulant dosing link. Coagulants are common agents in tap water treatment. They can simply treat water, reduce the harm to the human body, and achieve a certain filtration effect. Coagulant dosing, as an important link in the coagulation and sedimentation process of water treatment plants, is the key link affecting the quality of the effluent water. At present, coagulant dosing is mainly controlled manually. Due to the high requirements for water treatment plant operators and frequent problems such as over-dosing and lack of water quality, it cannot meet the operation requirements of digital intelligent water treatment plants in the new era.
[0003] In addition, for the intelligent automatic dosing system, due to the characteristics of "too complicated data and unstable quality" in the engineering data in the manual control stage, it has affected the reliability of the theoretical model in the off-line modeling stage to a certain extent. Therefore, data processing and quality control based on the big data of water treatment plants are undoubtedly important links in theoretical modeling. Existing data cleaning methods such as box plots and moving average outlier processing show stronger performance than traditional hard threshold cleaning and can establish more accurate theoretical models. However, for engineering abnormal data, in addition to obvious statistical abnormal data, it also includes logical abnormal data, such as jump point samples, inverted effluent turbidity samples, etc. Therefore, in the context of today's intelligent development, a more perfect deep data cleaning framework is urgently needed to ensure that the system can obtain higher-quality engineering data in an intelligent and unmanned environment and further assist in establishing a high-precision theoretical model.
[0004] In addition, due to the complexity and "black box" characteristics of artificial intelligence methods, when the quality of historical data is unstable, it is easy to have problems where the prediction results violate engineering logic due to overfitting. For example, Figure 1 as shown, under the condition of increasing raw water turbidity, the predicted dosing amount decreases instead. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide an online intelligent alum dosing system based on deep data cleaning.
[0006] To achieve the above object, the present invention provides the following technical solutions: An online intelligent alum dosing system based on deep data cleaning, based on the framework of a feedforward-feedback control system, includes data mining and modeling of the feedforward system, online data cleaning and model updating, and realizes PAC predictive dosing according to the data mining and modeling of the feedforward system and online data cleaning and model updating.
[0007] Data mining and modeling of the feedforward system:
[0008] (1) Based on the original operation monitoring data of the water plant, data cleaning is carried out.
[0009] (2) After data cleaning, data analysis is carried out on the selected sample data to establish a statistical model.
[0010] (3) According to the requirements of data changes, a multivariate non-linear dosing statistical model is established for working condition simulation evaluation.
[0011] Online data cleaning and model updating:
[0012] (1) Carry out real-time data monitoring and cleaning on the operation data of the water plant.
[0013] (2) According to the real-time online data, carry out real-time data identification and marking.
[0014] (3) According to the real-time data monitoring and cleaning module, the data is divided into identified samples and un-identified samples.
[0015] (4) According to the set model update period, use the un-identified data during the online operation period as the training data set to complete the adaptive update of the theoretical model.
[0016] Preferably, according to step (1) in the data mining and modeling of the feedforward system, the cleaning method for engineering abnormal data is as follows:
[0017] 1) Identify common engineering abnormal data;
[0018] 2) Carry out hard threshold processing to process other samples that obviously do not conform to the operation logic.
[0019] 3. According to the online intelligent alum dosing system based on deep data cleaning described in claim 2, it is characterized in that: the engineering abnormal data includes jump point abnormality and reverse turbidity abnormality of the settled water. Based on the identification of reverse turbidity abnormality data of the settled water, it needs to be carried out after the preprocessing of time offset correction.
[0020] Preferably, according to steps (2)-(3) in the data mining and modeling of the feedforward system, four variables of raw water turbidity, settled water turbidity, flow rate and temperature are set to establish a multivariate non-linear statistical model, and its model expression is:
[0021] PAC = a * I 4 + B * I 3 + c * I 2 + d * I + e
[0022] I = f(TUTs, TUTe, T, pH, t)
[0023] Wherein, I is a comprehensive variable, representing the turbidity removal effect per unit time at different temperatures, TUTs is the raw water turbidity, TUTe is the corresponding turbidity of the settled water after time offset, T is the raw water temperature, pH is the pH value of the raw water, and t is the time offset.
[0024] Preferably, according to steps (1)-(3) in online data cleaning and model updating, three types of data anomalies, namely null values, jump point anomalies, and inverted anomalies of the settled water, are set for data identification response.
[0025] Preferably, for the data identification of null values and inverted anomalies of the settled water, real-time identification and marking are adopted.
[0026] Preferably, for the data identification of jump point anomalies, long-short term mutation identification is adopted, and the steps are as follows:
[0027] 1) If the current sample is a mutant sample, it is marked;
[0028] 2) When the current sample is a mutant sample, trace back to the previous unmarked sample and accumulate the number of marks;
[0029] 3) If the accumulated number is less than the set step size, the current sample is a short-term mutant sample.
[0030] Compared with the prior art, the beneficial effects of the present invention are: introducing deep data cleaning and analysis into the online PAC intelligent dosing system, establishing a highly stable non-linear theoretical model based on physical experience, realizing reasonable and intelligent dosing to achieve the purpose of saving medicine on the premise of ensuring system stability and water quality compliance. At the same time, by accessing the real-time data monitoring and cleaning module, the stability of historical data is further ensured, and a high-quality database is established to assist the self-learning online update of the theoretical model and future data re-analysis. The continuously updated accurate dosing model ensures the long-term safe operation of alum dosing in the few-unmanned environment of the digital water plant and the achievement of water quality targets.
[0031] Establishing an intelligent alum dosing system for waterworks is an effective way to optimize the dosing strategy of water treatment chemicals and reduce water quality risks. By applying scientific deep data cleaning methods and combining the feedforward-feedback statistical model coupling algorithm, it can dynamically provide more accurate dosing instructions according to water quality control objectives and output a self-learning update model. This can not only stabilize the water quality of the effluent and reduce chemical consumption, improving the economic operation level of the waterworks, but also continuously improve the model accuracy and achieve the goal of less attended operation of digital waterworks. Description of the Drawings
[0032] Figure 1 For ANN prediction and raw water turbidity change;
[0033] Figure 2 For the improved feedforward-feedback online PAC intelligent dosing system;
[0034] Figure 3 For the variable scatter plot matrix;
[0035] Figure 4 For the change trend between PAC predicted dosing and three main variables among different models;
[0036] Figure 5 For the jump point abnormal data graph;
[0037] Figure 6 For the predicted PAC dosing monitoring data of the fitting models of different datasets;
[0038] Figure 7 For the environmental data of the test samples;
[0039] Figure 8 For the comparison graph between the intelligent dosing system and manual dosing;
[0040] Figure 9 For the graph of PAC dosing change and settled water turbidity change. Detailed Implementation Modes
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0042] Based on the framework of the feedforward-feedback control system, this invention focuses on the historical big data analysis and in-depth data cleaning and maintenance of modern digital water plants, and proposes an intelligent dosing system for online alum addition in water plants (taking PAC as an example) based on in-depth data cleaning. In the offline modeling stage, through in-depth data cleaning and based on physical experience, a simple and effective theoretical experience model is established; in the online operation stage, by introducing a real-time data monitoring and cleaning module, the intelligent response ability of the system to common abnormal working conditions is improved, and the online closed-loop control of the system is perfected. On this basis, the system can successfully establish a high-quality online database and assist in the autonomous update of the theoretical model. Correspondingly, the theoretical experience model based on physical experience and in-depth data cleaning can better avoid such problems and has high robustness. Therefore, during the transition of the intelligent dosing system from the offline manual stage to the online operation stage, the theoretical experience model can better ensure the safety and reliability of the system operation.
[0043] As part of the intelligent operation of water plants, in order to ensure the safety of project operation and the optimization and autonomous update of the theoretical model, certain requirements are put forward for the intelligent monitoring response to abnormal working conditions and the autonomous update of the theoretical model in the online intelligent dosing system. Abnormal working conditions not only pose a threat to production safety but also pollute the online engineering data set and cannot provide high-quality data support for the autonomous update of the theoretical model. Therefore, it is urgent to establish an early warning module for abnormal working conditions based on real-time engineering data and an online data cleaning module for the online operation system to improve the data quality problem of the system's online operation.
[0044] Based on the above discussion, this invention provides a technical solution: an online intelligent alum dosing system based on in-depth data cleaning, based on the framework of the feedforward-feedback control system, including data mining and modeling of the feedforward system, online data cleaning and model update, and realizing PAC predictive dosing according to the data mining and modeling of the feedforward system and online data cleaning and model update.
[0045] Data mining and modeling of the feedforward system:
[0046] (4) Based on the original operation monitoring data of the water plant, data cleaning is carried out;
[0047] (5) After data cleaning, data analysis is carried out on the selected sample data to establish a statistical model;
[0048] (6) According to the requirements of data change, a multivariate nonlinear dosing statistical model is established for working condition simulation evaluation;
[0049] Online data cleaning and model update:
[0050] (1) Carry out real-time data monitoring and cleaning on the operation data of the water plant;
[0051] (2) Identify and mark data in real time according to real-time online data;
[0052] (3) Divide data into labeled samples and unlabeled samples according to the real-time data monitoring and cleaning module;
[0053] (4) Complete the adaptive update of the theoretical model by using the unlabeled data during the online operation period as the training data set according to the set model update period.
[0054] According to step (1) in the data mining and modeling of the feedforward system, the cleaning method for engineering abnormal data is as follows:
[0055] 1) Identify common engineering abnormal data;
[0056] 2) Perform hard threshold processing to process other samples that obviously do not conform to the operation logic.
[0057] 3. According to the online intelligent alum dosing system based on deep data cleaning described in claim 2, it is characterized in that: the engineering abnormal data includes jump point abnormality and reverse hanging abnormality of the turbidity of the settled water. Based on the identification of the reverse hanging abnormality data of the turbidity of the settled water, it needs to be carried out after the preprocessing of time offset correction.
[0058] According to steps (2)-(3) in the data mining and modeling of the feedforward system, four variables of raw water turbidity, settled water turbidity, flow rate and temperature are set up to establish a multivariate nonlinear statistical model, and its model expression is:
[0059] PAC = a*I 4 +b*I 3 +c*I 2 +d*I+e
[0060] I = f(TUTs, TUTe, T, pH, t)
[0061] Among them, I is a comprehensive variable, representing the turbidity removal effect per unit time at different temperatures, TUTs is the raw water turbidity, TUTe is the corresponding settled water turbidity after time offset, T is the raw water temperature, pH is the raw water pH value, and t is the time offset.
[0062] According to steps (1)-(3) in the online data cleaning and model update, three types of data abnormalities, namely null values, jump point abnormalities and reverse hanging abnormalities of the settled water, are set to perform data identification responses.
[0063] Preferably, for the data identification of null values and reverse hanging of the settled water, real-time identification and marking are adopted.
[0064] Preferably, for the data identification of jump point abnormalities, long-term and short-term mutation identification is adopted, and the steps are as follows:
[0065] 1) If the current sample is a mutant sample, then mark it;
[0066] 2) In the case that the current sample is a mutant sample, trace back to the previous unmarked sample and accumulate the number of marked samples;
[0067] 3) If the accumulated number is less than the set step size, then the current sample is a short-term mutant sample.
[0068] Regarding the above technical solution, during its implementation:
[0069] Based on the original operation monitoring data of the water plant, data cleaning is carried out. The cleaning method for engineering abnormal data is as follows:
[0070] 1) Identify common engineering abnormal data;
[0071] 2) Perform hard threshold processing to handle other samples that obviously do not conform to the operation logic.
[0072] Engineering abnormal data includes jump point anomalies and abnormal inversion of the turbidity of the sedimentation tank effluent. Based on the identification of abnormal inversion data of the turbidity of the sedimentation tank effluent, it needs to be carried out after preprocessing the time offset correction;
[0073] The experimental data of the present invention comes from the operation monitoring data available throughout 2021 of a certain water plant, including raw water pH, raw water turbidity, raw water temperature, inlet flow of the sedimentation tank, turbidity of the sedimentation tank effluent, etc. The data frequency sampling is 10 minutes, and this is used as the original data. Excluding monitoring anomalies and null value phenomena due to monitoring equipment or transmission problems, there are a total of nearly 51,000 complete recorded data.
[0074] There are certain differences between engineering data anomalies and traditional statistical anomalies. General statistical anomaly value processing methods cannot identify such as jump point data, abnormal inversion data of the turbidity of the sedimentation tank effluent, etc.
[0075] Based on the above data cleaning process and precautions, a total of nearly 42,000 effective samples are obtained. A statistical model is established for the screened data, and a multivariate non-linear dosing statistical model is established:
[0076] The statistical relationship between different variables can be simply understood through the scatter plot of variables and the pearson correlation coefficient. Figure 3It can be seen that the manual dosage of PAC has significant statistical correlations with raw water turbidity, settled water turbidity, flow rate, and temperature, and has the largest correlation with raw water turbidity. Therefore, it is necessary to establish a multiple statistical model based on these four variables to simulate the changes in PAC dosage under different working conditions. However, for a multiple linear model, the requirement for each independent variable is that the variables need to be independent of each other, that is, there is no correlation between the independent variables. Obviously, the data of the present invention cannot meet this requirement. For example, the statistical correlation between raw water turbidity and temperature is as high as 0.629. Similarly, from the scatter plot, the linear relationship cannot truly depict the relationship between PAC dosage and each independent variable. Therefore, it is necessary to establish a multiple non-linear statistical model to describe the interaction relationship between each independent variable and the relationship between them and PAC dosage.
[0077] According to physical experience, the present invention first needs to clarify the influence mode of different independent variables on PAC dosage.
[0078] The PAC coagulation effect is the best within a certain temperature range T. To a certain extent, the flow rate represents the reaction time t of the coagulation degree. The longer the reaction time, the appropriate amount of PAC dosage can be reduced. Finally, the physical explanation of the influence of turbidity on PAC dosage is how much the turbidity of the water body can be reduced per unit of PAC dosage. Since the settled water turbidity is the turbidity target value and is relatively stable, the turbidity difference can show the main characteristics of the two.
[0079] Therefore, according to steps (2)-(3), four variables of raw water turbidity, settled water turbidity, flow rate, and temperature are set, and a multiple non-linear statistical model is established. The model expression is:
[0080] PAC = a*I 4 +b*I 3 +c*I 2 +d*I+e
[0081] I = f(TUTs, TUTe, T, pH, t)
[0082] TUTdif = TITs - TUTe
[0083] Among them, I is a comprehensive variable, representing the turbidity removal effect per unit time at different temperatures. TUTs is the raw water turbidity, TUTe is the corresponding settled water turbidity after time offset, T is the raw water temperature, pH is the raw water pH value, and t is the time offset.
[0084] In addition, as a relatively mainstream feedforward theoretical model system at present, the present invention also establishes two
[0085] The ANN model is compared and analyzed with this linear model. The two ANN models have the same structural category. Both are built with 3 hidden layers, with 33 neurons in each layer and the same activation function. The difference lies in the input variables. For MLP1, they are TUTdif (the turbidity difference between inlet and outlet water), T and t, and for MLP2, it is the comprehensive variable I.
[0086] In the present invention, statistical indicators R2 and root mean square error (RMSE) are established, and the model is evaluated in combination with simulation conditions.
[0087] From the perspective of the overall model accuracy, both MLP1 and MLP2 are higher than StaModel (Table 1). Among them, the model accuracies of MPL2 and StaModel are not very different, and the variables used by both are the same. Therefore, to a certain extent, it shows that the non-linear characteristics fitted by ANN are similar to this statistical model. The difference between MLP1 and 2 lies in the difference in input features. MLP1 uses the original variables, while MLP2 uses the comprehensive index I established in the present invention. Only from the model simulation accuracy, the ANN-type models are indeed superior to the statistical models. ANN fits the non-linear characteristics between variables in the form of a "black box", which provides great convenience for us when we cannot accurately describe the variable relationship.
[0088] Table 1 Model fitting degree
[0089]
[0090] On the other hand, based on the safety control of the engineering process, the present invention compares the three models again from the perspective of the engineering logic consistency of engineering variables. Figure 4 It shows the change trends of different models predicting PAC and three variables within 5 hours. Among them, the solid line is for StaModel, the dashed line is for MLP2, and the dotted line is for MLP1. The differences in the predictions of the three models are reflected in the predicted PAC change at 10:00 on June 30th. At this moment, the turbidity difference shows an obvious decrease. Correspondingly, the predicted PAC of StaModel and MLP2 also decreases. However, the PAC predicted by MLP1 increases abnormally, which violates the engineering logic. Therefore, even though the prediction accuracy of MLP1 is higher than that of StaModel and MLP2, we cannot regard this model as the feed-forward theoretical model of the system. On the contrary, when the prediction and fitting accuracies of StaModel and MLP2 are relatively acceptable, the change in the predicted PAC responds reasonably to the changes in all main variables, and they are good alternative models. At the same time, considering that MLP2 also has the "black box" attribute, its performance is also limited by the data quality, and there is no obvious difference in its accuracy from StaModel.
[0091] Therefore, the present invention finally takes StaModel as the offline theoretical model of the system.
[0092] For the online data cleaning and model update module: As an important module of the online PAC intelligent dosing system ( Figure 2 ), this module undertakes the same functions as the data cleaning part in the offline stage. However, due to the particularity of the online attribute, the cleaning logic of this module needs to be different from the traditional thinking and requires the addition of the three processes of "identification, marking, and response".
[0093] (1) Conduct real-time data monitoring and cleaning of the water plant operation data;
[0094] (2) Conduct real-time data identification and marking based on the real-time online data;
[0095] (3) According to the real-time data monitoring and cleaning module, divide the data into identified samples and unidentified samples; according to steps (1)-(3) in the technical content, set three types of data anomalies, namely null values, jump point anomalies, and backflow anomalies of sedimentation water, for data identification response.
[0096] For the data identification of null values and backflow anomalies of sedimentation water, real-time identification and marking are adopted, and the abnormal data are directly identified and marked.
[0097] For the data identification of jump point anomalies, long-term and short-term mutation identification is adopted. Since the identification of jump points must be based on the judgment of its adjacent two points, no one can determine whether the currently observed change point must belong to a jump point. If jump point identification is required, during online monitoring, whether it is manual or an intelligent system, it is necessary to wait for at least one observation period before truly judging a jump point. Therefore, based on this principle, the jump point identification of the online system requires a broader definition, that is, long-term and short-term mutation identification;
[0098] The steps are as follows:
[0099] 1) If the current sample is a mutation sample, mark it;
[0100] 2) When the current sample is a mutation sample, trace back to the previous unmarked sample and accumulate the number of marks;
[0101] 3) If the accumulated number is less than the set step size, the current sample is a short-term mutation sample.
[0102] Through these method steps, the module can judge jump point data within an acceptable time delay and make a response. For long-term mutation points, the module does not need to make a response because long-term mutations represent a change in working conditions rather than abnormal data.
[0103] To ensure the uniformity of data cleaning, long-term and short-term mutation identification is also adopted as the identification method for jump point data during offline data cleaning. Taking the sedimentation water turbidity data as an example ( Figure 5) In the whole year of 2021, there were a total of 2,589 jump point samples, accounting for about 5% of the total sample quantity.
[0104] Finally, according to the set model update period, using the unlabeled data during the online operation period as the training data set, the adaptive update of the theoretical model is completed.
[0105] The online update of the model can not only keep the system synchronized with the current water purification working conditions and environment, but also further improve the accuracy of predicting the chemical dosage. Figure 6 It shows the PAC dosing prediction results of the dosing models fitted by two different data sets for the same period. The dotted line (MJ) is the theoretical model fitted by the 1-minute data set from May to June for the PAC dosing prediction in the next two weeks. The solid line is the predicted value of the theoretical model fitted by the 10-minute data set for the whole year of 2021. The solid dots represent the actual PAC dosing. It is not difficult to find that the model (MJ) obtained from the adjacent data set has a more suitable prediction for the current dosing environment, that is, the prediction accuracy is higher. This also reflects the necessity of the online update of the model.
[0106] According to the real-time data monitoring and cleaning module, the system divides the data into two categories: labeled samples and unlabeled samples. After the system runs online and matures for a period of time, the system can, according to the set model update period, use the unlabeled high-quality data during the online operation period as the training data set to complete the adaptive update of the theoretical model.
[0107] Through the technical solution of the present application, the system has carried out on-site follow-up tests in a certain water plant. At the same time, in order to test the stability and prediction ability of the system in different environments ( Figure 7 ), continuous follow-up experiments were carried out by selecting different time periods and temperature ranges respectively, and manual dosing was used as the control group for comparison. A total of nearly 1,500 groups of data results with a sampling frequency of 1 minute were obtained ( Figure 8 ). Compared with manual dosing, in the test, this system saved an average of 2.36 units of PAC dosage in a single dosing, which is about 15% of the dosage savings. While meeting the turbidity of the settled water, the dosing cost was greatly saved.
[0108] At the same time, in the test, this system also met the expectations in terms of real-time data feedback monitoring and response to dosing changes. Taking Figure 9 as an example, this system found that the original dosing amount was not enough to meet the target turbidity requirement by monitoring the difference between the actual turbidity of the settled water and the target turbidity of the settled water, so a certain amount of dosing was increased. After the necessary sedimentation process, the actual turbidity of the settled water dropped to the target turbidity requirement and the dosing amount was maintained unchanged.
[0109] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An online intelligent alum dosing system based on deep data cleaning, which is based on the framework of a feedforward-feedback control system, is characterized in that: Including data mining and modeling of the feedforward system, online data cleaning and model updating, based on the data mining and modeling of the feedforward system and online data cleaning and model updating, PAC predictive dosing is realized; Data Mining and Modeling of Feedforward Systems: (1) Perform data cleaning based on the original operation monitoring data of the water plant; (2) After data cleaning, the screened sample data is analyzed and a statistical model is established; (3) Establish a multivariate nonlinear dosing statistical model based on data change requirements and conduct operating condition simulation evaluation; Online data cleaning and model updating: (1) Real-time data monitoring and cleaning of water plant operation data; (2) Based on real-time online data, real-time data identification and labeling are performed; (3) According to the real-time data monitoring and cleaning module, the data is divided into identified samples and unidentified samples; (4) According to the set model update cycle, the unlabeled data during the online operation period is used as the training data set to complete the adaptive update of the theoretical model; According to steps (1)-(3) in online data cleaning and model updating, three types of data anomalies, namely null value, jump point anomaly and post-sinking water inversion anomaly, are set for data identification response; For data identification of abnormal jumping points, long-term and short-term mutation identification is adopted. The steps are as follows: 1) If the current sample is a mutation sample, mark it; 2) If the current sample is a mutant sample, trace back to the last unlabeled sample and accumulate the number of labels; 3) If the cumulative number is less than the set step size, the current sample is a short-term mutation sample.
2. The online intelligent alum dosing system based on depth data cleaning according to claim 1, wherein: According to step (1) in data mining and modeling of feedforward system, the method for cleaning engineering abnormal data is as follows: 1) Identify common engineering abnormal data; 2) Perform hard threshold processing to process other samples that obviously do not conform to the operating logic.
3. The online intelligent alum dosing system based on depth data cleaning according to claim 2, characterized in that: The engineering abnormal data include jump point anomaly and post-sinking water turbidity inversion anomaly. Based on the identification of post-sinking water turbidity inversion anomaly data, it is necessary to carry out pre-processing of time offset correction.
4. The online intelligent alum dosing system based on depth data cleaning according to claim 1, characterized in that: According to steps (2)-(3) in the data mining and modeling of the feedforward system, four variables, namely, raw water turbidity, settled water turbidity, flow rate and temperature, are set to establish a multivariate nonlinear statistical model, whose model expression is: PAC = a*I 4 + b*I 3 + c*I 2 + d*I + e I=f(TUTs,TUTe,T,pH,t) Among them, I is a comprehensive variable, representing the deturbidity effect per unit time at different temperatures, TUTs is the raw water turbidity, TUTe is the corresponding turbidity of the settled water after time offset, T is the raw water temperature, pH is the raw water pH value, and t is the time offset.
5. The online intelligent alum dosing system based on depth data cleaning according to claim 1, characterized in that: For the data identification of null values and inverted water after sinking, real-time identification and marking are adopted.
Citation Information
Patent Citations
Device control method based on big data and artificial intelligent water quality prediction
CN110308705A
Accurate dosing method for tap water treatment
CN111859263A