Meteorological-driven turbidity real-time prediction method and device fused with residual correction

By constructing a random forest and residual model based on meteorological data, combining adaptive noise decomposition and wind energy density calculation, the accuracy problem of real-time prediction of turbidity in shallow lakes is solved, real-time and high-precision prediction of turbidity is achieved, and it is suitable for water quality monitoring and embedded monitoring systems.

CN120409270APending Publication Date: 2025-08-01HANGZHOU POLYTECHNIC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510582060.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The prior art cannot achieve real-time prediction of water quality turbidity and poor prediction accuracy, especially in the case of complex wind field driving effects in shallow lakes, the existing models cannot accurately reflect the fluctuations and response patterns of turbidity.

Method used

The sample database is constructed using temperature data, wind speed data, wind direction data and turbidity data. The temperature sequence is reconstructed through fully adaptive noise ensemble empirical modal decomposition and variance contribution rate, combined with the Weibull distribution function to calculate the wind energy density, a random forest model and residual model are constructed, turbidity prediction is performed, and the database is updated in real time.

Benefits of technology

Real-time prediction of water quality turbidity is realized, prediction accuracy and generalization ability are improved, equipment abnormalities can be automatically monitored, suitable for water quality monitoring and embedded monitoring systems, and medium- and short-term forecasts are realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409270A_ABST
    Figure CN120409270A_ABST
Patent Text Reader

Abstract

The invention provides a meteorological-driven turbidity real-time prediction method and device fused with residual correction, and belongs to the technical field of water environment monitoring. Temperature data, wind speed data, wind direction data and turbidity data of a time sequence are collected and preprocessed; completely adaptive noise set empirical mode decomposition is combined with a variance contribution rate to obtain a reconstructed air temperature sequence, and a sample database is constructed; and correcting the sample data by adopting a Weibull distribution function, constructing a random forest model and a residual error model, and carrying out weighted summation on the turbidity predicted by the random forest model and the residual error model at the next moment to obtain a final turbidity prediction result. According to the method, mechanism factors of water quality turbidity change under meteorological driving are fully considered, the random forest model and the residual error model are combined, unbalance caused by nonlinearity and data missing of water quality data is solved, the prediction stability and precision of the model are remarkably improved, and the method is widely applied to water quality monitoring and embedded monitoring systems and has a wide application prospect. And medium-short-term forecasting of water quality is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water environment monitoring, and in particular, to a real-time prediction method and device for meteorological-driven turbidity integrating residual correction. Background Technique

[0002] Due to the topographical characteristics of shallow water bodies, the driving effect of the wind field on the change of turbidity or suspended sediment concentration has become a research hotspot. Existing research mainly falls into two categories: non-mechanistic models and mechanistic models. Non-mechanistic models often use regression equations or machine learning algorithms as tools to establish linear or non-linear relationships between turbidity and its related influencing factors to achieve the purpose of single-site prediction; mechanistic models rely on mathematical models to calculate the resuspension, settlement, migration and other movement processes of suspended sediment particles in the hydrodynamic environment under different driving forces, so as to reproduce the spatio-temporal distribution of suspended sediment in the calculation domain. The non-mechanistic model has well solved the non-linear problem of water environment data and has developed rapidly in recent years.

[0003] However, the non-mechanistic model also has limitations. Different from the turbidity of water environments such as rivers, which is mainly organic turbidity, the hydrodynamic system of shallow lakes is controlled by the wind field, and sediment resuspension caused by the wind field becomes the key factor for the increase in turbidity. Therefore, many studies are based on wind condition data in order to establish an efficient turbidity regression model for shallow water systems.

[0004] Regression models generally take single-site sampling data at a single time as the object, and analyze the influence mechanism of the wind field on turbidity from a statistical perspective to establish linear or non-linear models. However, the action mechanism of the wind field on turbidity is very complex, and differences in water depth, sampling location, etc. will all lead to different response modes of turbidity. The correlation between a single wind speed or wind direction variable and turbidity is poor, and even negative correlation may occur. Further research found that the wind direction changes the response mode of the sampling point turbidity in the form of the fetch length, and there is a critical wind speed value even when the wind directions are the same. Only when this value is exceeded will the turbidity show an obvious positive correlation with the wind force. In addition, the hysteresis of turbidity response has also received more and more attention. Considering the cumulative effect of the previous wind field can effectively improve the model accuracy. Based on the above qualitative research results, various forms of regression equations have been proposed to accurately represent the correlation between turbidity and the wind field.

[0005] For example, Reference 1 (Hurricane impacts on turbidity and sediment in the Rookery Bay National Estuarine Research Reserve, Florida, USA. International Journal of Sediment Research. 2016. 31, 330 - 340) established a regression relationship between the wind speed and turbidity in Rookery Bay based on the wind speed and turbidity data during a hurricane event. After the turbidity sequence was smoothed and the lag response to the wind speed was considered, the linear correlation increased from 0.41 to over 0.95. However, some information was lost in the smoothed sequence, making it difficult to truly reflect the fluctuations of turbidity. Another example is Reference 2 (Response of water turbidity to wind speed, wind direction and their time - cumulative effects - taking Gonghu Bay of Taihu Lake as an example. Journal of Lake Sciences. 2018. 30(6), 1587 - 1598.), which verified the high correlation between the water turbidity and wind speed in Gonghu Bay of Taihu Lake through basic sampling experiments and proposed a trigonometric function formula that couples wind speed, wind direction and time - cumulative effects, increasing the correlation from 0.2 - 0.6 to 0.3 - 0.7. However, this correction formula does not consider the critical wind speed factor. Therefore, there are not many current relevant studies on quantitatively describing the correlation between the wind field and turbidity based on statistical methods. And limited by the ability of the regression equation to solve nonlinear problems, the optimized model still cannot meet the accuracy requirements of prediction.

[0006] With the popularization of intelligent algorithms, artificial neural networks (ANN), genetic algorithms (GA), support vector machine methods (SVM), K - nearest Neighbor (KNN) algorithms, etc. have been increasingly applied to water quality prediction models, providing new ideas for turbidity forecasting. Machine learning models have well solved the complex nonlinear characteristics of the water environment system and have been widely used in the forecasting of water quality indicators such as dissolved oxygen, total phosphorus, ammonia nitrogen, etc. However, they are relatively less applied in turbidity prediction, especially in the application for shallow lakes. Since such models can handle a large amount of multi - dimensional data, in existing research, many factors such as dissolved oxygen, pH, conductivity, chlorophyll a concentration, etc. are included as input variables. Although the accuracy of the model has increased, it does not consider that the monitoring of these water quality indicators and the acquisition of historical data may be more difficult than obtaining turbidity data itself, which limits the practical application. In addition, taking historical turbidity as one of the input variables, although it greatly improves the accuracy of turbidity prediction, it also limits the prediction duration, and can only achieve the prediction for the next moment and still cannot meet the requirements of real - time prediction. Summary of the Invention

[0007] The object of the present invention is to provide a real-time prediction method and device for meteorological-driven turbidity integrating residual correction, which is used to solve the problems of inability to perform real-time prediction of water quality turbidity and poor prediction accuracy.

[0008] A real-time prediction method for meteorological-driven turbidity integrating residual correction provided by an embodiment includes the following steps:

[0009] (1) Collect time series data of air temperature, wind speed, wind direction and turbidity, and perform preprocessing;

[0010] (2) Decompose the time series data of air temperature by using complete adaptive noise ensemble empirical mode decomposition to obtain multiple modal air temperature components and a residual sequence, and then obtain a reconstructed air temperature sequence by using the variance contribution rate;

[0011] (3) Based on the reconstructed air temperature sequence and the preprocessed time series data of wind speed, wind direction and turbidity, construct a sample database, divide it into a training set and a test set, and then divide the training set into several sub-datasets based on the time series data of wind direction. Calculate the average wind energy density of each sub-dataset by using the Weibull distribution function, and sum the weighted average wind energy densities of the sub-datasets to obtain the corrected average wind energy density;

[0012] (4) Determine the decision tree structure with the reconstructed air temperature sequence and the corrected average wind energy density as feature variables to form a random forest model, optimize the parameters of the random forest model, and then use the reconstructed air temperature sequence, the corrected average wind energy density and their linear combination of the training set as inputs, and the preprocessed time series turbidity data as outputs to train the random forest model, and then use the test set as input to predict the turbidity at the next moment;

[0013] (5) Use the residual sequence in step (2) as input and the preprocessed time series turbidity data as output to train a residual model, and then use the test set as input to predict the turbidity at the next moment. Sum the weighted turbidity predictions at the next moment predicted by the random forest model and the residual model to obtain the final turbidity prediction result, and add the final turbidity prediction result to the sample database.

[0014] In one embodiment, the preprocessing is to filter outliers from the air temperature data, wind speed data, wind direction data and turbidity data; the outlier filtering adopts the quartile trimming method based on a moving window:

[0015] L u =Q 0.75 +a·h,

[0016] L d =Q Q.25 -a·h,

[0017] h=Q 0.75 -Q0.25 ,

[0018] wherein, L u , L d are the upper and lower limits for excluding sub-samples of the moving window length, Q 0.75 and Q 0.25 are the upper and lower quartiles of the sub-samples respectively; h is the interquartile range, reflecting the dispersion of the middle 50% of the data; a is a constant determined according to the data distribution and the required accuracy.

[0019] In one embodiment, when performing the quartile exclusion method based on a moving window for temperature data, the value range of a is 1 to 1.5;

[0020] When performing the quartile exclusion method based on a moving window for wind speed data, the value range of a is 1 to 2;

[0021] When performing the quartile exclusion method based on a moving window for wind direction data, the value range of a is 1.5 to 2.5;

[0022] When performing the quartile exclusion method based on a moving window for turbidity data, the value range of a is 1.5 to 2.5.

[0023] In one embodiment, the decomposition of temperature data by using complete adaptive noise ensemble empirical mode decomposition to obtain multiple temperature components and a residual sequence includes:

[0024] Adding white noise to the preprocessed time series of temperature data to generate multiple perturbation signals, performing empirical mode decomposition on each perturbation signal, extracting the ensemble average of the first modal temperature components of the decomposition and calculating the residual;

[0025] Adding adaptive noise to the residual and performing iterative empirical mode decomposition to update the residual sequence and obtain multiple modal temperature components.

[0026] In one embodiment, obtaining the reconstructed temperature sequence by using the variance contribution rate includes: quantifying the influence contribution amount of a single modal temperature component by using the variance contribution rate and selecting a linear combination of modal temperature components to obtain the reconstructed temperature sequence.

[0027] In one embodiment, calculating the average wind energy density of each sub-dataset by using the Weibull distribution function includes:

[0028] Processing the wind speed data by using the Weibull distribution function:

[0029]

[0030] Among them, f(v) is the probability density of the Weibull distribution function based on the wind speed data v of the preprocessed time series, k is the shape parameter of the Weibull distribution function, c is the scale parameter of the Weibull distribution function, and the average wind energy density is calculated through the following formula:

[0031]

[0032] Among them, is the average wind energy density, and ρ is the air density.

[0033] In one embodiment, determining the decision tree structure to form a random forest model with the reconstructed temperature sequence and the corrected average wind energy density as characteristic variables includes:

[0034] Using the reconstructed temperature sequence and the corrected average wind energy density as the first characteristic variables, randomly selecting at least one characteristic variable from all the characteristic variables corresponding to the current node of the decision tree as the second characteristic variable, merging the first characteristic variable and the second characteristic variable, and then selecting a characteristic variable and the corresponding variable value from the merged characteristic variables as the best splitting variable and splitting value, and performing splitting of the left and right child nodes according to the best splitting variable and splitting value to determine the decision tree structure to form a random forest model.

[0035] In one embodiment, determining the decision tree structure to form a random forest model with the reconstructed temperature sequence and the corrected average wind energy density as characteristic variables further includes: measuring the quality of the splitting variable and the splitting value through the splitting gain after the splitting of the left and right child nodes, and optimizing the decision tree structure; the calculation formula of the splitting gain is as follows:

[0036]

[0037] Among them, x i represents the i-th splitting variable, v ij represents the j-th splitting value corresponding to the splitting variable x i , Gain(x i , v ij ) is the splitting gain, N s , n left , n right respectively represent the splitting variables corresponding to the current node and the left and right child nodes after splitting, x left and x right respectively represent the splitting variable sets of the left and right child nodes, MSE(X) represents the mean square error of the current node containing the splitting variable set X, x is a subset of the splitting variable set X, y x represents the actual value based on the sample database, represents the average value of the target splitting variable based on the sample data of the current node.

[0038] In one embodiment, when splitting the left and right child nodes of the current node, if the splitting variable value of the sample data is less than or equal to the splitting value of the current node, access the left child node of the current node; if the splitting variable value of the sample data is greater than or equal to the splitting value of the current node, access the right child node of the current node.

[0039] In one embodiment, the optimization of the random forest model parameters includes: using the reconstructed air temperature sequence and the corrected average wind energy density as characteristic variables, and determining the number of trees, the maximum number of features, and the deepest depth of the decision trees of the random forest model according to the mean square error of the out-of-bag data of the random forest model. Among them, the out-of-bag data of the random forest model uses the bootstrap method to randomly extract data from the sample database to form several sub-sample data sets as the sub-training sets of the decision trees, and the data not extracted each time forms the out-of-bag data.

[0040] To clearly show the meteorological-driven turbidity real-time prediction method with fusion residual correction, a meteorological-driven turbidity real-time prediction device with fusion residual correction is provided, including a memory and a processor. The memory is used to store a computer program, and the processor is used to implement the meteorological-driven turbidity real-time prediction with fusion residual correction when executing the computer program.

[0041] Compared with the prior art, the beneficial effects of the present invention at least include:

[0042] (1) Select meteorological data such as air temperature data, wind speed data, wind direction data, and turbidity data to construct a sample database, overcome the problem of too short prediction duration caused by the turbidity prediction model relying on historical data or non-forecast data, and greatly improve the generalization ability of the turbidity prediction model.

[0043] (2) Preprocess the data in the sample database and perform empirical mode decomposition with added noise, fully consider the mechanistic factors of water quality turbidity change under meteorological driving, and combine the random forest model and the residual model to solve the problem of inaccurate prediction results caused by the nonlinearity of water quality data and the imbalance caused by data missing.

[0044] (3) Update the sample database in real time according to the prediction results, improve the reliability of the prediction results, and at the same time be able to automatically monitor the verification of the data, timely detect the abnormal operation of the monitoring equipment, and be widely applied to water quality monitoring and embedded monitoring systems to realize medium- and short-term forecasting of water quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.

[0046] Figure 1Schematic flow chart of the real-time meteorological-driven turbidity prediction method integrating residual correction provided by the present invention;

[0047] Figure 2 Schematic diagram of the prediction processes of the main model and the residual model provided by the present invention;

[0048] Figure 3 Using the preprocessed sample data as input, predicting the turbidity result at the next moment according to the trained random forest model;

[0049] Figure 4 Using only the reconstructed air temperature sequence, the corrected average wind energy density, and their linear combination as input, predicting the turbidity result at the next moment according to the trained random forest model;

[0050] Figure 5 The final turbidity prediction result obtained by weighted summation of the turbidity predicted at the next moment by the random forest model and the turbidity predicted at the next moment by the residual model. Detailed implementation manners

[0051] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0052] The solution of the embodiment of the present invention is as Figure 1 shown. A real-time meteorological-driven turbidity prediction method integrating residual correction includes the following steps:

[0053] S1. Collect air temperature data, wind speed data, wind direction data, and turbidity data in a time series and perform preprocessing.

[0054] In the embodiment, as Figure 2 shown, collect air temperature data, wind speed data, wind direction data, and turbidity data in a time series to form an initial database. Among them, there is a unique mapping relationship between the air temperature data, wind speed data, and wind direction data at each moment and the turbidity data. When a certain moment's meteorological factors (air temperature data, wind speed data, wind direction data) are determined, the water quality index (turbidity data) at that moment can be correspondingly obtained. The initial sample database contains a total of 34,510 pieces of data, and the recording time interval is 1 hour.

[0055] Then perform outlier filtering to remove invalid data. Specifically, use the quartile trimming method based on a moving window to preprocess the data:

[0056] L u =Q 0.75 +a·h

[0057] L d= Q 0.25 -a·h

[0058] h = Q 0.75 -Q 0.25

[0059] Wherein, L u , L d is the upper and lower limits for excluding sub-samples of the moving window length, Q 0.75 and Q 0.25 are the upper and lower quartiles of the sub-samples respectively; h is the interquartile range, reflecting the dispersion degree of the middle 50% of the data; a is a constant determined according to the data distribution and the required accuracy.

[0060] In the embodiment, for the air temperature data of the time series, the value of the constant a is 1; for the wind speed data of the time series, the value of the constant a is 1.5; for the wind direction data of the time series, the value of the constant a is 2. For the turbidity data of the time series, the value of the constant a is 1. Through outlier filtering, more outliers can be found, and the moving window avoids the influence of the segmentation boundary. After the above outlier filtering, the final available data is 33964 pieces.

[0061] S2. Use complete ensemble empirical mode decomposition with adaptive noise to decompose the air temperature data of the time series, obtain multiple modal air temperature components and a residual sequence, and then use the variance contribution rate to obtain a reconstructed air temperature sequence.

[0062] In the embodiment, after preprocessing the air temperature data, wind speed data, wind direction data and turbidity data of the time series, use complete ensemble empirical mode decomposition with adaptive noise (CEEMDAN) to decompose the preprocessed air temperature data of the time series, and obtain 12 modal (IMF) air temperature components. The specific method is as follows:

[0063] ① Add white noise w i (t) to the air temperature data x(t) of the original time series to generate a perturbed signal x i (t) = x(t) + α·w i (t), where w i (t) represents the value of the i-th white noise sequence at time t, α is the noise factor, and x i (t) represents the perturbed signal set of the i-th white noise sequence;

[0064] ② Perform empirical mode decomposition (EMD) on each perturbed signal, and extract the ensemble average IMF1(t) of the first modal air temperature components of all perturbed signals: The residual is r1(t) = x(t) - IMF1; wherein, The first modal temperature component obtained after the EMD decomposition of the disturbance signal adding the i-th white noise sequence based on the temperature data, N is the total number of disturbance signals, and r1(t) is the residual obtained by separating IMF1 from the temperature data of the original time series;

[0065] ③ Add adaptive noise to the residual r j (t): j is the number of cycles of the EMD decomposition, denotes the h-th disturbance residual after adding adaptive noise to the residual r j (t) in the j-th cycle decomposition;

[0066] ④ Perform EMD decomposition on each disturbance residual and extract the ensemble average IMF j+1 of the modal temperature components in the (j + 1)-th cycle decomposition of all disturbance residuals: Update the residual: r j+1 (t) = r j (t) - IMF j+1 ; where denotes the IMF component in the (j + 1)-th cycle decomposition extracted after performing EMD decomposition on the h-th disturbance residual , H represents the total number of disturbance residuals, and r j+1 (t) is the residual after the cycle decomposition.

[0067] ⑤ Repeat ③ and ④ until the residual r j+1 (t) is a monotonic function or cannot be decomposed into IMFs anymore, and finally a total of 12 modal temperature components are obtained.

[0068] After obtaining multiple temperature components and residual sequences, the variance contribution rate is used to quantify the significance of the influence of a single IMF temperature component on the overall time series. The first two modal temperature components with the largest contribution are selected for linear recombination to obtain the reconstructed temperature sequence. The specific method is as follows:

[0069] ① Variance contribution rate where M i is the variance contribution rate of the i-th IMF temperature component, D i is the corresponding variance, and n is the number of IMF temperature components;

[0070] ② Select the modal temperature components according to the variance contribution. The sum of the selected modal temperature components exceeds 80%, and perform linear combination to obtain the reconstructed temperature sequence Tem = IMF(M max ) + IMF(M sec-max ). In this embodiment, the reconstructed temperature time series is obtained by linearly recombining IMF10, IMF7, and IMF5, and the residual sequence is IMF0.

[0071] S3. Construct a sample database based on the reconstructed air temperature sequence and the wind speed data, wind direction data, and turbidity data of the preprocessed time series, divide the training set and the test set, then divide the training set into several sub-datasets based on the wind direction data of the time series, calculate the average wind energy density of each sub-dataset using the Weibull distribution function, and sum the weighted average wind energy densities of the sub-datasets to obtain the corrected average wind energy density.

[0072] In the embodiment, a sample database is constructed based on the reconstructed air temperature sequence and the wind speed data, wind direction data, and turbidity data of the preprocessed time series, and the training set and the test set are divided. The length of the training set is 360 sample data of consecutive time series, including the reconstructed air temperature values XTEM (TEM tr1 , TEM tr2 , ……, TEM tr359 , TEM tr360 ), the wind direction data XWD of the preprocessed time series (WD tr1 , WD tr2 , ……, WD tr359 , WD tr360 ), the wind speed data XWS of the preprocessed time series (WS tr1 , WS tr2 , ……, WS tr359 , WS tr360 ) and the turbidity YTUR of the preprocessed time series (TUR tr1 , TUR tr2 , ……, TUR tr359 , TUR tr360 ). The length of the test set is 48, that is, 48 consecutive time series samples after the training set time series, including the reconstructed air temperature values X’TEM (TEM te1 , TEM te2 , ……, TEM te47 , TEM te48 ), the wind direction data X’WD of the preprocessed time series (WD te1 , WD te2 , ……, WD te47 , WD te448 ), the wind speed data X’WS of the preprocessed time series (WS te1 , WS te2 , ……, WS te47 , WS te48 ) and the turbidity Y’TUR of the preprocessed time series (TUR te1 , TUR te2 , ……, TUR te47 , TUR te48 ).

[0073] The sample data of 360 consecutive time series in the training set are divided into 16 sub-datasets (S1, S2, ……, S 15 , S 16 ) according to 16 wind directions. Using the Weibull distribution function, calculate the shape parameters k (k1, k2, ……, k 15 , k 16 ), scale parameters c (c1, c2, ……, c 15 , c 16 ) and average wind energy density of each sub-dataset Specifically, the probability density of the Weibull distribution function is calculated as where f(v) is the probability density of the Weibull distribution function based on the wind speed data v of the preprocessed time series, k is the shape parameter of the Weibull distribution function, c is the scale parameter of the Weibull distribution function, and the cumulative distribution function is The parameters c and k are calculated by the gradient descent method: ① After logarithmic transformation, ln(-ln(1 - F(v))) = k×ln(v) - k×ln(c); ② The wind speed data of the preprocessed time series in the training set are divided into average intervals (0~v1, v1~v2, …, v n-1 ~v n ); The probability densities are f(v1), f(v2), …, f(v n ); The cumulative densities are F(v1), F(v2), …, F(v n ); ③ Assume x i = ln(v i ), y i = ln(-ln(1 - F(v i ))), and the above ① can be converted into a linear function Y = a*X + b, where a = k, b = -k×ln(c); ④ The loss function for iterative calculation is where a m and b m are the parameter values at the m-th iteration; ⑤ Iteratively calculate through a m = a m-1 - r*A m-1 and b m = b m-1 - r*B m-1 , where r is the learning rate, A m-1 and B m-1 are the descent gradients of parameters a and b at the m-th iteration, and A m-1 and B m-1 are obtained by taking the partial derivatives of the loss function: and The finally calculated average wind energy density is where is the average wind energy density, and ρ is the air density. The calculation results are shown in Table 1 below:

[0074] Table 1

[0075]

[0076]

[0077] Next, according to the distribution frequencies of the 16 wind direction intervals, calculate the average wind energy density of each wind direction interval, and perform weighted summation to obtain the final corrected wind energy density: where is the average wind energy density of wind direction interval d, and f(D d ) is the distribution frequency of wind direction interval d. The calculation results are shown in Table 2 below:

[0078] Table 2

[0079]

[0080] According to the above calculation method, obtain the training set sequence of the corrected wind energy density through a moving window where According to the wind direction data XWD of the time series after preprocessing, WD tr1 , WD tr2 , ……, WD tr359 , WD tr360 ), and the wind speed data XWS of the time series after preprocessing, WS tr1 , WS tr2 , ……, WS tr359 , WS tr360 ), WD tr-1 and WS tr-1 are the sample data of the first time point of the training set time series, ...... The calculation is the same in the same way; the corrected wind energy density sequence of the test set is the same, which is

[0081] Through linear combination calculation, obtain the training set and test set of the interaction eigenvalues of air temperature and wind field (wind direction data and wind speed data):

[0082] S4. Determine the decision tree structure with the reconstructed temperature sequence and the corrected average wind energy density as characteristic variables to form a random forest model, optimize the parameters of the random forest model, and then use the reconstructed temperature sequence, the corrected average wind energy density of the training set, and their linear combination as inputs, and the turbidity data of the preprocessed time series as the output to train the random forest model, and then use the test set as the input to predict the turbidity at the next moment.

[0083] The random forest algorithm combines multiple weak learners to obtain a strong learner. Since the output of water quality turbidity is a continuous value, the regression algorithm of random forest is used. The air-reconstructed temperature sequence and the corrected average wind energy density are used as characteristic variable inputs, and the model parameters are determined according to the out-of-bag data estimation accuracy of the model (i.e., the mean square error of the out-of-bag data). The main adjustable parameters are the number of trees, the maximum number of features considered when constructing the optimal decision tree model, and the deepest depth of the tree.

[0084] Specifically, first, according to the preset number of decision trees, the bootstrap method is used to randomly extract K different sub-sample data sets from the sample database as the sub-training sets of each decision tree. The sample size of each is the same as that of the original sample database, and the samples can be repeatedly extracted. Each time, the data not sampled forms the out-of-bag data;

[0085] Decision trees are established for the K sub-sample data sets respectively. For each decision tree, the recursive branch growth is continuously looped in process a until the stopping condition is reached, and the decision tree stops growing. The predicted value of the turbidity data of each decision tree is determined according to the final leaf nodes. Among them, process a: use the reconstructed temperature sequence and the corrected average wind energy density as the first characteristic variables, and randomly select at least one characteristic variable from all the characteristic variables corresponding to the current node as the second characteristic variable. Combine the first characteristic variable and the second characteristic variable, and then select 1 characteristic variable and the corresponding variable value from the combined characteristic variables to determine the best split variable and split value. According to the best split variable, split the current node and all the sample variable values that are the same as the best split variable into left and right leaf nodes to determine the decision tree structure combination;

[0086] Combine all the decision trees into a random forest, and take the average value of the predicted values of the turbidity data of the K decision trees as the predicted value of the time series turbidity data to form a random forest model.

[0087] Among them, the split gain after splitting is used to measure the quality of the split variable and split value. The larger the split gain value, the more the total error of the left and right leaf nodes after splitting is reduced compared with that before splitting, and the better the splitting effect:

[0088]

[0089] Among them, x i represents the i-th split variable, and v ijDenote the splitting variable as x i The corresponding j-th splitting value, Gain(x i , v ij ) is the splitting gain, N s , n left , n right respectively represent the splitting variables corresponding to the current node and the left and right child nodes after splitting. x left and x right respectively represent the sets of splitting variables of the left and right child nodes. MSE(X) represents the mean squared error of the current node containing the set of splitting variables X. x is a subset of the set of splitting variables X, and y x represents the actual value based on the sample database. represents the average value of the target splitting variable based on the sample data of the current node.

[0090] When splitting the current node into left and right child nodes, if the value of the splitting variable of the sample data is less than or equal to the splitting value of the current node, access the left child node of the current node; if the value of the splitting variable of the sample data is greater than or equal to the splitting value of the current node, access the right child node of the current node.

[0091] In the embodiment, since there are 3 input variables in the random forest model, namely the reconstructed air temperature sequence, the corrected average wind energy density, and their linear combination, it has little impact on the calculation time. Therefore, the maximum number of features is set to 2. In addition, when the number of trees reaches 150 and the maximum depth reaches 20, further increasing the parameters does not significantly improve the model accuracy but increases the calculation time. So finally, the number of trees is set to 150 and the maximum depth is set to 20.

[0092] In application, first perform outlier processing on the sample data of the training set. Use the reconstructed air temperature sequence, the corrected average wind energy density, and their linear combination (reconstructed air temperature sequence × corrected average wind energy density) as input variables, and the turbidity data of the target time series as output to train the random forest prediction model. Then input the reconstructed air temperature value, the corrected average wind energy density value, and their linear combination (reconstructed air temperature sequence × corrected average wind energy density) at the moment to be predicted in the test set into the trained random forest model to predict the turbidity value P rf . In this embodiment, the turbidity value predicted based on the test set data is MTUR te1 .

[0093] S5: Use the residual sequence in S2 as input and the turbidity data of the preprocessed time series as output to train the residual model. Then use the test set as input to predict the turbidity at the next moment. Weightedly sum the turbidity predicted at the next moment by the random forest model and the residual model to obtain the final turbidity prediction result, and add the final turbidity prediction result to the sample database.

[0094] In the embodiment, in addition to constructing a random forest model as the main model to predict the turbidity data at the next moment, an LSTM residual model is also constructed to improve the prediction accuracy of the turbidity data. The implementation process of the LSTM residual model is as follows:

[0095] Taking the residual sequence r j+1 (t) after the complete adaptive noise ensemble empirical mode decomposition of the air temperature data in S2 as the input, and the turbidity data of the time series as the output, training the long short-term memory network model (LSTM), and inputting the air temperature residual value at the moment to be predicted into the trained model to predict the turbidity value P lstm .

[0096] In this embodiment, training set residual data is established, that is, 360 continuous time series, the input variables are IMF0 TR (IMF0 tr1 , IMF0 tr2 , ……, IMF0 tr359 , IMF0 tr360 ), and the output variables are Y TUR (TUR tr1 , TUR tr2 , ……, TUR tr359 , TUR tr360 ); test set data is established, and the input variables are IMF0 TE (IMF0 te1 , IMF0 te2 , ……, IMF0 te47 , IMF0 te48 ); the turbidity value predicted based on the test set data is CTUR te1 .

[0097] The final predicted value is obtained by weighted summing the turbidity predicted by the random forest model at the next moment and the turbidity predicted by the residual model at the next moment to obtain the final turbidity prediction result: P = P rf + λ·P lstm . In this embodiment, the final turbidity predicted value is TUR te1 = MTUR te1 + 0.3·CTUR te1 .

[0098] After completing one prediction, add the final turbidity prediction result to the sample database, fix the length of the training set, and repeat the above operations to obtain a turbidity prediction sequence of a certain length for updating the database and continue the turbidity prediction at the next moment, that is, cycle the above process to obtain TUR te (TUR te1 , TUR te2 , ……, TUR te47 , TURte48 ) to simulate the real-time update of the database in actual operation once.

[0099] The present invention also provides a real-time prediction device for meteorological-driven turbidity with fusion residual correction, including a memory and a processor. The memory is used to store a computer program, and the processor is used to implement the real-time prediction method for meteorological-driven turbidity with fusion residual correction when executing the computer program.

[0100] To demonstrate the effectiveness of the real-time prediction method for meteorological-driven turbidity with fusion residual correction provided by the present invention, a comparison of turbidity data was carried out, as Figures 3 to 5 shown. From the results in the figure, it can be seen that based on the preprocessed air temperature data, wind speed data, wind direction data, and turbidity data of the time series as inputs, by analyzing the distribution frequency of the turbidity data and the cumulative frequency of the time series for the turbidity prediction results, the cumulative frequency of turbidity prediction within a 10% relative error is only 12.1%. However, in the present invention, the reconstructed air temperature sequence, corrected average wind energy density, and their linear combination are used as inputs, and the turbidity result at the next moment is predicted according to the trained random forest model. Then, the final turbidity prediction result obtained by weighted summing the turbidity predicted at the next moment by the residual model is used. The cumulative frequency of turbidity prediction within a 10% relative error rises to 29.9%, and the relative error is greatly reduced, which is also significantly better than the result predicted by only using the random forest model. Moreover, the turbidity prediction result of the present invention shows that the relative error of more than half of the samples is less than 20%, and the relative error value of turbidity prediction for 91% of the sample data is less than 7%. This is sufficient for the water plant to adjust the dosing amount in advance. This model only uses wind field data and air temperature data as simple inputs, and its effect is better than that of the model based on multi-variable inputs. In the case of a larger data acquisition interval, medium-term prediction can be achieved.

[0101] Based on the analysis results, it can be seen that meteorological data such as air temperature and wind field are selected as input variables to overcome the problems of the prediction model relying on historical data or non-forecast data such as dissolved oxygen and short prediction duration. The air temperature data in the database is decomposed, and the decomposition variables with higher importance are selected through the variance contribution rate. The reconstructed sequence obtained by superposition is used as the input, which maximally retains the periodic information of the air temperature and filters out the noise at the same time. In addition, the average wind energy density in the period before the prediction time point is calculated through the Weibull distribution function and used as an input variable of the model, comprehensively reflecting the influence of wind speed, wind direction, and cumulative effect on turbidity. In addition, considering that single-model prediction is easily affected by the long-term trend or mutation noise in the residual sequence, resulting in error accumulation, the present invention overcomes the time-series dependence relationship of the separately modeled residual sequence by introducing residual correction, and introduces dynamic feature penalty weights and interactive feature generation in the main model to strengthen the sensitivity of the model to key climate features (such as rainfall, wind speed), capture the complex non-linear interaction effects between meteorological factors and turbidity, reduce noise interference, achieve error correction, greatly improve the prediction accuracy of water quality indicators and realize real-time prediction, and can be widely applied to water quality monitoring and embedded monitoring systems.

[0102] The above specific embodiments have elaborated on the technical solutions and beneficial effects of the present invention in detail. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A real-time prediction method for meteorological-driven turbidity that integrates residual correction, characterized in that, Including the following steps: (1) Collect the air temperature data, wind speed data, wind direction data and turbidity data of the time series, and perform preprocessing; (2) Decompose the air temperature data of the time series by using complete ensemble empirical mode decomposition with adaptive noise to obtain multiple modal air temperature components and a residual sequence, and then obtain the reconstructed air temperature sequence by using the variance contribution rate; (3) Based on the reconstructed air temperature sequence and the preprocessed wind speed data, wind direction data and turbidity data of the time series, construct a sample database, divide it into a training set and a test set, and then divide the training set into several sub-datasets based on the wind direction data of the time series. Calculate the average wind energy density of each sub-dataset by using the Weibull distribution function, and sum the weighted average wind energy densities of the sub-datasets to obtain the corrected average wind energy density; (4) Determine the decision tree structure with the reconstructed air temperature sequence and the corrected average wind energy density as feature variables to form a random forest model, optimize the parameters of the random forest model, and then use the reconstructed air temperature sequence, the corrected average wind energy density and their linear combination of the training set as inputs, and the preprocessed turbidity data of the time series as the output to train the random forest model, and then use the test set as the input to predict the turbidity at the next moment; (5) Use the residual sequence in step (2) as the input and the preprocessed turbidity data of the time series as the output to train the residual model, and then use the test set as the input to predict the turbidity at the next moment. Sum the weighted turbidity predicted at the next moment by the random forest model and the residual model to obtain the final turbidity prediction result, and add the final turbidity prediction result to the sample database.

2. The meteorological-driven real-time turbidity prediction method according to claim 1, wherein The preprocessing described is to perform outlier filtering on the air temperature data, wind speed data, wind direction data and turbidity data of the time series; the outlier filtering adopts the quartile removal method based on a moving window: L u = Q 0.75 + a·h, L d = Q 0.25 - a·h, h = Q 0.75 -Q 0.25 , Among them, L u , L d are the upper and lower limits for excluding the moving window length sub-samples, Q 0.75 and Q 0.25 are the upper and lower quartiles of the sub-samples respectively; h is the interquartile range, reflecting the dispersion of the middle 50% of the data; a is a constant determined according to the data distribution and the required accuracy.

3. The meteorological-driven turbidity real-time prediction method according to claim 2, characterized in that When performing the quartile removal method based on a moving window for the air temperature data of the time series, the value range of a is 1 to 1.5; When performing the quartile removal method based on a moving window for the wind speed data of the time series, the value range of a is 1 to 2; When performing the quartile removal method based on a moving window for the wind direction data, the value range of a is 1.5 to 2.5: When performing the quartile removal method based on a moving window for the turbidity data of the time series, the value range of a is 1.5 to 2.

5.

4. The meteorological-driven turbidity real-time prediction method according to claim 1, characterized in that Decompose the preprocessed air temperature data of the time series by using complete ensemble empirical mode decomposition with adaptive noise to obtain multiple modal air temperature components and a residual sequence, including: Add white noise to the preprocessed air temperature data of the time series to generate multiple perturbed signals, perform empirical mode decomposition on each perturbed signal, extract the first modal air temperature component set average of the decomposition and calculate the residual; Add adaptive noise to the residual and perform iterative empirical mode decomposition to update the residual sequence and obtain multiple modal air temperature components.

5. The meteorological-driven turbidity real-time prediction method according to claim 1, wherein, The obtaining of the reconstructed air temperature sequence by using the variance contribution rate described includes: using the variance contribution rate to quantify the influence contribution amount of a single modal air temperature component and select the linear combination of the modal air temperature components to obtain the reconstructed air temperature sequence.

6. The meteorological-driven turbidity real-time prediction method according to claim 1, characterized in that The calculation of the average wind energy density of each sub-dataset by using the Weibull distribution function described includes: The Weibull distribution function is used to process the wind speed data of the preprocessed time series: Among them, f(v) is the probability density of the Weibull distribution function based on the wind speed data v of the preprocessed time series, k is the shape parameter of the Weibull distribution function, c is the scale parameter of the Weibull distribution function, and the average wind energy density is calculated through the following formula: Among them, is the average wind energy density, and ρ is the air density.

7. The meteorological-driven real-time turbidity prediction method according to claim 1, wherein The method of determining the decision tree structure with the reconstructed temperature sequence and the corrected average wind energy density as characteristic variables to form a random forest model includes: Taking the reconstructed temperature sequence and the corrected average wind energy density as the first characteristic variables, randomly selecting at least one characteristic variable from all the characteristic variables corresponding to the current node of the decision tree as the second characteristic variables, combining the first characteristic variables and the second characteristic variables, and then selecting one characteristic variable and the corresponding variable value from the combined characteristic variables as the best splitting variable and splitting value, and performing left and right child node splitting according to the best splitting variable and splitting value to determine the decision tree structure and form a random forest model.

8. The real-time prediction method of meteorological-driven turbidity according to claim 7, characterized in that The method of determining the decision tree structure with the reconstructed temperature sequence and the corrected average wind energy density as characteristic variables to form a random forest model further includes: measuring the quality of the splitting variable and splitting value through the splitting gain after the left and right child node splitting, and optimizing the decision tree structure; the calculation formula of the splitting gain is as follows: Among them, x i represents the i-th splitting variable, and v ij represents the j-th split value corresponding to the splitting variable x i , and Gain(x i , v ij ) is the splitting gain. N s , n left , and n right respectively represent the splitting variables corresponding to the current node and the left and right child nodes after splitting. x left and x right respectively represent the sets of splitting variables of the left and right child nodes. MSE(X) represents the mean squared error of the current node including the set of splitting variables X. x is a subset of the set of splitting variables X, and y x represents the actual value based on the sample database. represents the average value of the target splitting variable based on the sample data of the current node.

9. The meteorological-driven turbidity real-time prediction method according to claim 1, wherein The method of optimizing the random forest model parameters includes: taking the reconstructed temperature sequence and the corrected average wind energy density as characteristic variables, and determining the number of trees, the maximum number of features, and the deepest depth of the decision tree of the random forest model according to the mean square error of the out-of-bag data of the random forest model. Among them, the out-of-bag data of the random forest model is randomly sampled from the sample database by the bootstrap method to form several sub-sample data sets as the sub-training sets of the decision tree, and the data not sampled each time forms the out-of-bag data.

10. A real-time prediction device for meteorological-driven turbidity integrating residual correction, comprising a memory and a processor, wherein the memory is used for storing computer programs, and is characterized in that, When the processor executes the computer program, it implements the real-time prediction method of meteorological-driven turbidity with fusion residual correction according to any one of claims 1 to 9.