Phytoplankton biomass in-situ high-frequency monitoring data processing and predicting method

Through isolated forest algorithm, wavelet threshold noise reduction and XGBoost machine learning to optimize model parameters, outliers and noise problems in high-frequency water ecological monitoring data are solved, and the accuracy and stability of phytoplankton biomass prediction are improved.

CN120277579APending Publication Date: 2025-07-08HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510412227.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing water ecological monitoring data processing methods have problems inaccurate outliers identification, insufficient noise suppression effect and low prediction accuracy in high-frequency data, which affect the accuracy and reliability of data analysis.

Method used

The isolated forest algorithm is used to identify and eliminate outliers, combine wavelet threshold noise reduction technology to remove noise components, and model it through XGBoost machine learning algorithm, and optimize model parameters using Newton-Ravson optimization algorithm to improve prediction accuracy.

Benefits of technology

Effectively identify and remove outliers and noise in high-frequency water ecological monitoring data, improve the accuracy and real-time prediction of phytoplankton biomass, and ensure the stability and applicability of the model under different data sets and environmental conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277579A_ABST
    Figure CN120277579A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of water environment monitoring, in particular to a phytoplankton biomass in-situ high-frequency monitoring data processing and predicting method, and aims to improve the accuracy and reliability of water ecological monitoring and prediction. According to the method, data preprocessing is carried out in combination with an isolated forest algorithm and a wavelet threshold, the isolated forest algorithm effectively detects an abnormal value based on random partition unsupervised learning, the wavelet threshold noise reduction technology analyzes the wavelet coefficient difference between noise of different frequency bands and a signal, the signal quality is optimized, and the signal quality is improved. The common problems of sensor faults, communication problems, external interference and the like in high-frequency monitoring are solved. The phytoplankton biomass prediction model adopts an XGBoost machine learning method, and model parameters are optimized through a Newton-Raphson algorithm, so that the prediction precision is improved. According to the method, through efficient data processing and optimization, a more accurate phytoplankton biomass prediction scheme is provided, and dynamic changes of a water ecological environment can be better reflected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water environment monitoring, and particularly to a method for processing and predicting in-situ high-frequency monitoring data of phytoplankton biomass. Background Art

[0002] With the continuous growth of environmental protection and water ecological monitoring needs, in-situ high-frequency water ecological monitoring technology has been widely used in the ecological environment assessment of lakes, reservoirs and other water bodies. By deploying sensors in the water body to monitor physical, chemical and biological parameters in the water body in real time, such as water temperature, pH value, dissolved oxygen, nitrogen and phosphorus content, and phytoplankton biomass, a large amount of high-frequency data can be obtained. These data help to reveal the dynamic trends of water ecological changes and provide a scientific basis for environmental management, ecological protection and decision-making.

[0003] Phytoplankton biomass is an important part of the water ecosystem, reflecting the nutrient status and productivity of the water body. Its monitoring is crucial for water quality assessment and water ecological management. Through high-frequency monitoring, the changes in phytoplankton biomass can be captured in real time, providing an important reference for taking timely countermeasures.

[0004] However, in-situ high-frequency monitoring data are often affected by factors such as sensor failures, communication delays, and external environmental interferences in practical applications, resulting in outliers and noise in the data, which affect the accuracy and reliability of data analysis. In order to effectively improve the quality and reliability of monitoring data, these outliers and noise must be processed, and predictive analysis of the data must be carried out to provide a more accurate trend of water ecological changes.

[0005] Existing water ecological monitoring data processing methods mainly rely on traditional statistical analysis techniques and signal processing methods. For example, the Isolation Forest algorithm has been widely used in outlier detection, but its application in high-frequency data has not been fully optimized. The wavelet transform denoising technique can effectively extract the effective components of the signal, but how to balance denoising and signal preservation when processing high-frequency data is still a technical difficulty.

[0006] In addition, predictive analysis methods also play an important role in water ecological data processing. Machine learning algorithms such as XGBoost are widely used in predictive modeling of water ecological data. Although these algorithms can provide relatively accurate prediction results, how to improve the performance of the model by optimizing algorithm parameters is still a key problem to be solved.

[0007] The following problems are common in the current water ecological monitoring data processing: inaccurate outlier identification. Traditional outlier detection methods ignore the temporal characteristics of data and are prone to false detection or missed detection. Insufficient noise suppression effect. Existing noise reduction technologies fail to effectively balance short-term fluctuations and long-term trends in data, resulting in unstable noise reduction effects. Low prediction accuracy. Existing prediction models may be affected by outliers and noise, resulting in unstable or biased prediction results. Therefore, how to accurately identify outliers, effectively reduce noise, and improve prediction accuracy has become a core technical problem in the current water ecological monitoring data processing. Summary of the invention

[0008] The purpose of the present invention is to provide a method for processing and predicting in-situ high-frequency monitoring data of phytoplankton biomass to solve the problems raised in the above-mentioned background technology.

[0009] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for processing and predicting in-situ high-frequency monitoring data of phytoplankton biomass, comprising:

[0010] S100, collecting aquatic monitoring data in the water body of the test area in real time through the deployed high-frequency monitoring equipment, and sending the collected aquatic monitoring data to the data processing end through the wireless transmission network;

[0011] S200, the outlier detection module in the data processing end obtains the aquatic monitoring data to be processed, and identifies and removes the abnormal data in the aquatic monitoring data by using a plurality of isolation trees constructed by using an isolation forest algorithm based on a random partitioning method;

[0012] S300, performing wavelet threshold denoising processing on the aquatic monitoring data after removing abnormal data, analyzing the difference in wavelet coefficients between the signal and the noise through wavelet transform, using a threshold method to remove the noise component, retaining the signal part, and obtaining wavelet reconstructed data;

[0013] S400, using XGBoost machine learning algorithm to model the obtained wavelet reconstruction data to obtain an aquatic prediction model; and predicting the phytoplankton biomass based on the aquatic prediction model;

[0014] S500, based on the modeling data of the XGBoost machine learning algorithm and the prediction results of phytoplankton biomass, uses the Newton-Raphson optimization algorithm NRBO to optimize the model parameters of the aquatic prediction model to ensure the optimal performance of the obtained aquatic prediction model under different data sets and environmental conditions, and updates the aquatic prediction model in the current database based on the optimized model parameters.

[0015] The present invention aims to improve the accuracy and reliability of water ecological monitoring and prediction. The invention combines the Isolation Forest algorithm and wavelet threshold for data preprocessing. The Isolation Forest algorithm effectively detects outliers based on unsupervised learning of random partitioning, and the wavelet threshold denoising technology analyzes the differences in wavelet coefficients of noise and signals in different frequency bands to optimize the signal quality, solving common problems such as sensor failures, communication problems, and external interferences in high-frequency monitoring. The phytoplankton biomass prediction model uses the XGBoost machine learning method and optimizes the model parameters through the Newton-Raphson algorithm to improve the prediction accuracy. Through efficient data processing and optimization, this method provides a more accurate phytoplankton biomass prediction scheme, which helps to better reflect the dynamic changes of the water ecological environment.

[0016] Further, the acquisition frequency of the high-frequency monitoring devices deployed in S1 is a preset value. The deployed high-frequency monitoring devices are composed of multiple different types of sensors. The aquatic monitoring data includes water temperature, conductivity, total dissolved solids, permanganate index, dissolved oxygen, pH value, ammonia nitrogen, nitrite, total nitrogen, total phosphorus, and chlorophyll.

[0017] Further, S200 includes:

[0018] S210. Obtain the aquatic monitoring data to be processed;

[0019] S220. Summarize the acquisition results of the same data type in the historical aquatic monitoring data, and use the Isolation Forest model composed of multiple isolation trees constructed by the Isolation Forest algorithm, and train the constructed Isolation Forest model with the summary set of the acquisition results of the same data type in the historical aquatic monitoring data to obtain the corresponding Isolation Forest training model for the corresponding data type;

[0020] The summary set of the acquisition results of the same data type in the historical aquatic monitoring data is used as a data set; the isolation tree is a binary tree, constructed by recursively partitioning the data set, and each isolation tree is trained on a subset of the data set; during the construction process, a split value is randomly selected to divide the data set into two partition data sets; repeat the above operation of recursively partitioning the data set until one of the following conditions is met: (1) Each data point is isolated in a separate leaf node; (2) The predefined maximum tree depth is reached.

[0021] S230. Input each acquisition data in the aquatic monitoring data to be processed into the corresponding Isolation Forest training model for the corresponding data type, aggregate the path lengths of all isolation trees of each acquisition data in the corresponding Isolation Forest training model, and calculate the outlier score of each acquisition data; denote the outlier score of the acquisition data x as s(x).

[0022]

[0023] Among them, h(x) represents the path length of the collected data x in the isolation tree; E[h(x)] represents the average path length of the collected data x in all isolation trees of the isolation forest training model, and c(n) represents the average path length of unsuccessful searches for n pieces of collected data; when E[h(x)] = c(n), s(x) = 0.5;

[0024] S240. Identify and eliminate abnormal data in the aquatic monitoring data, and the abnormal scores corresponding to the abnormal data in the aquatic monitoring data all belong to the interval [0, 0.5];

[0025] The types of the abnormal data include the abnormal data caused by sensor failures and external interferences.

[0026] Further, the S300 includes:

[0027] S310. Obtain the aquatic monitoring data after eliminating the abnormal data, select the wavelet type and the N levels of wavelet decomposition, perform N-layer wavelet decomposition on the aquatic monitoring data with a noisy signal to obtain the wavelet coefficients of each layer, where N is a preset value;

[0028] S320. Quantify the high-frequency wavelet coefficients from the first layer to the Nth layer using an adaptive threshold to remove the relevant noise; the expression for obtaining the adaptive threshold is the root mean square interpolation threshold function:

[0029]

[0030] where ω j,k and are the wavelet coefficients before and after threshold quantization respectively, T represents the threshold, μ represents a preset adjustment parameter and 0 ≤ μ ≤ 1; when μ = 0, the root mean square interpolation threshold function is denoted as the soft threshold function equation; when μ = 1, the root mean square interpolation threshold function is denoted as the hard threshold function equation;

[0031] The soft threshold function equation is:

[0032]

[0033] The hard threshold function equation is:

[0034]

[0035] S330. Reconstruct the processed signal to obtain the wavelet reconstruction data after noise reduction based on the aquatic monitoring data.

[0036] Further, the aquatic prediction model in the S400 is trained by the XGBoost machine learning algorithm for the wavelet reconstruction data corresponding to the historical aquatic monitoring data and the phytoplankton biomass corresponding to the wavelet reconstruction data.

[0037] Furthermore, in the process of optimizing the model parameters of the aquatic prediction model by using the Newton-Raphson optimization algorithm NRBO in S500, two rules, namely the Newton-Raphson search rule and the trap avoidance operator TAO, are used to explore the entire search process;

[0038] The steps for optimizing the model parameters of the aquatic prediction model are as follows:

[0039] S510. NRBO starts to search for the optimal solution by generating an initial random population within the corresponding boundaries of the candidate solutions of the phytoplankton biomass prediction result. Each population consists of fuzzy decision variables. The formula for generating the random population is as follows:

[0040]

[0041] where, represents the j-th dimension position of the n-th population; N p represents the number of populations in the initial random population; Rand represents a random number between (0, 1); dim represents the total number of problem dimensions preset for the n-th population; lb represents the lower bound, which is the minimum value of the decision variable for the corresponding problem dimension of the population; ub represents the upper bound, which is the maximum value of the decision variable for the corresponding problem dimension of the population;

[0042] S520. Control the vector matrix X n through the Newton-Raphson search rule to obtain the position of the feasible region explored; and adjust the global search and local search states during the exploration process through the adaptive coefficient δ;

[0043] S530. Denote the current vector position as the corresponding to an element in the vector matrix X n and use TAO to change the vector position of the previous vector position in the next iteration step

[0044] S540. Complete the update of the model parameters in the aquatic prediction model. The updated model parameters of the aquatic prediction model include the search step size Δx, the search step size of the Newton-Raphson search rule, and the weight of the global or local search ability.

[0045] Furthermore, the calculation formula involved in controlling the vector matrix X n in S520 through the Newton-Raphson search rule is as follows:

[0046]

[0047] yw = r1 × (Mean(Z n+1 + x n )) + r1 × Δx

[0048] y b = r1 × (Mean(Z n+1 + x n )) - r1 × Δx

[0049]

[0050] where randn is a normally distributed random number with a mean of 0 and a variance of 1; Δx represents the search step size; NRSR represents the search step size of the Newton - Raphson search rule; x n represents the position; x w represents the worst position (x n + Δx); Z n+1 represents the new search position; x b represents the best position (x n - Δx); y w and y b represent the positions of two vectors generated using Z n+1 and x n respectively; r1 represents a random number between (0, 1); Rand(1, Dim) is a random number in the interval [1, Dim];

[0051] The calculation formula for the adaptive coefficient δ is as follows:

[0052]

[0053] where IT represents the current iteration, Max_ IT represents the maximum number of iterations.

[0054] Furthermore, in S530, when combining the best position Xb and the current vector position if the value of rand is less than DF, then the solution involves the following formula:

[0055]

[0056] where θ1 and θ2 are uniformly distributed random numbers between (-1, 1) and (-0.5, 0.5) respectively, DF represents the determinant controlling the performance of NRBO, μ1 and μ2 are random numbers; during the exploration iteration process, the obtained vector position value is used as the new current vector position to continue the exploration;

[0057]

[0058] μ1 is a number within the range of (0, 1), and the above equation can be simplified:

[0059] μ1 = β × 3 × rand + (1 - β)

[0060] μ2 = β × rand + (1 - β)

[0061] Where β represents a binary number, and the value of β is 1 or 0; if the value of μ1 is greater than or equal to 0.5, then the value of β is 0; otherwise, the value of β is 1.

[0062] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0063] (1) The present invention can accurately identify abnormal data in high-frequency water ecological data, whether it is caused by sensor failure, data transmission problems, or external environmental interference; at the same time, through the wavelet threshold denoising technology, the noise components in the high-frequency monitoring data are effectively removed, and the important information in the signal is retained; and through the XGBoost framework for phytoplankton biomass prediction, the historical monitoring data and dynamic characteristics are fully utilized to accurately extract the potential laws in the data;

[0064] (2) By combining the isolation forest algorithm and the wavelet threshold denoising technology, the present invention proposes a new processing method, which can not only effectively solve the problems of outliers and noise, but also improve the accuracy and real-time performance of phytoplankton biomass prediction;

[0065] (3) The optimized model shows good stability and real-time prediction ability under the condition of large changes in high-frequency data, and is suitable for long-term monitoring and dynamic assessment of phytoplankton biomass. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the accompanying drawings:

[0067] Figure 1 is the technical roadmap of the in-situ high-frequency monitoring data processing and prediction method for phytoplankton biomass of the present invention;

[0068] Figure 2 is the isolation forest outlier detection diagram in the in-situ high-frequency monitoring data processing and prediction method for phytoplankton biomass of the present invention;

[0069] Figure 3 is the wavelet threshold denoising result diagram in the in-situ high-frequency monitoring data processing and prediction method for phytoplankton biomass of the present invention;

[0070] Figure 4It is a result diagram of the phytoplankton biomass prediction model in the phytoplankton biomass in-situ high-frequency monitoring data processing and prediction method of the present invention. DETAILED DESCRIPTION

[0071] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0072] See also Figures 1-4 The present invention provides a technical solution: a method for processing and predicting in-situ high-frequency monitoring data of phytoplankton biomass, comprising:

[0073] S100, collect aquatic monitoring data in the water body of the test area in real time through the deployed high-frequency monitoring equipment, and send the collected aquatic monitoring data to the data processing end through the wireless transmission network; the collection frequency of the high-frequency monitoring equipment deployed in S100 is a preset value, and the deployed high-frequency monitoring equipment is composed of a plurality of different types of sensors, and the aquatic monitoring data includes water temperature, electrical conductivity, total dissolved solids, permanganate index, dissolved oxygen, pH value, ammonia nitrogen, nitrite, total nitrogen, total phosphorus and chlorophyll.

[0074] S200, the outlier detection module in the data processing end obtains the aquatic monitoring data to be processed, and identifies and removes the abnormal data in the aquatic monitoring data by using a plurality of isolation trees constructed by using an isolation forest algorithm based on a random partitioning method;

[0075] The S200 includes:

[0076] S210, obtaining aquatic monitoring data to be processed;

[0077] S220, summarizing the collection results of the same data type in the historical aquatic monitoring data, and constructing an isolation forest model consisting of multiple isolation trees using an isolation forest algorithm, and training the constructed isolation forest model with the summary set of the collection results of the same data type in the historical aquatic monitoring data, to obtain an isolation forest training model corresponding to the corresponding data type;

[0078] The aggregated collection of the acquisition results of the same data type in the historical aquatic monitoring data is used as a data set; the isolation tree is a binary tree constructed by recursively partitioning the data set, and each isolation tree is trained on a subset of the data set; during the construction process, a split value is randomly selected to divide the data set into two partition data sets; repeat the above operation of recursively partitioning the data set until one of the following conditions is met: (1) each data point is isolated in a separate leaf node; (2) the predefined maximum tree depth is reached.

[0079] S230. Respectively input each acquisition data in the aquatic monitoring data to be processed into the isolation forest training model corresponding to the corresponding data type, aggregate the path lengths of all isolation trees of each acquisition data in the corresponding isolation forest training model, and calculate the anomaly score of each acquisition data; denote the anomaly score of the acquisition data x as s(x).

[0080]

[0081] Among them, h(x) represents the path length of the acquisition data x in the isolation tree; E[h(x)] represents the average path length of the acquisition data x in all isolation trees of the isolation forest training model, and c(n) represents the average path length of unsuccessful searches of n acquisition data; when E[h(x)] = c(n), s(x) = 0.5.

[0082] S240. Identify and remove the abnormal data in the aquatic monitoring data, and the anomaly scores corresponding to the abnormal data in the aquatic monitoring data all belong to the interval [0, 0.5].

[0083] The types of the abnormal data include the abnormal data caused by sensor failures and external interferences.

[0084] S300. Perform wavelet threshold denoising processing on the aquatic monitoring data after removing the abnormal data, analyze the wavelet coefficient differences between the signal and the noise through wavelet transform, adopt the threshold method to remove the noise components, and retain the signal part to obtain the wavelet reconstruction data.

[0085] The S300 includes:

[0086] S310. Obtain the aquatic monitoring data after removing the abnormal data, select the wavelet type and the N levels of wavelet decomposition, perform N-layer wavelet decomposition on the aquatic monitoring data with the noisy signal to obtain the wavelet coefficients of each layer, and N is a preset value.

[0087] S320. Quantify the high-frequency wavelet coefficients from the first layer to the Nth layer with an adaptive threshold to remove the relevant noise; the expression for obtaining the adaptive threshold is the root mean square interpolation threshold function:

[0088]

[0089] Wherein, ω j,k and are the wavelet coefficients before and after threshold quantization respectively, T represents the threshold, μ represents a preset adjustment parameter and 0 ≤ μ ≤ 1; when μ = 0, the root mean square interpolation threshold function is denoted as the soft threshold function equation; when μ = 1, the root mean square interpolation threshold function is denoted as the hard threshold function equation;

[0090] The soft threshold function equation is as follows:

[0091]

[0092] The hard threshold function equation is as follows:

[0093]

[0094] S330. Reconstruct the processed signal to obtain the wavelet reconstruction data after denoising based on the aquatic monitoring data.

[0095] S400. Use the XGBoost machine learning algorithm to model the obtained wavelet reconstruction data to obtain an aquatic prediction model; and predict the phytoplankton biomass based on the aquatic prediction model;

[0096] In the S400, the aquatic prediction model is obtained by training the wavelet reconstruction data corresponding to the historical aquatic monitoring data and the phytoplankton biomass corresponding to the wavelet reconstruction data through the XGBoost machine learning algorithm.

[0097] S500. Based on the modeling data of the XGBoost machine learning algorithm and the prediction results of the phytoplankton biomass, use the Newton - Raphson optimization algorithm NRBO to optimize the model parameters of the aquatic prediction model to ensure the optimal performance of the obtained aquatic prediction model under different data sets and environmental conditions, and update the aquatic prediction model in the current database based on the optimized model parameters;

[0098] In the S500, during the process of using the Newton - Raphson optimization algorithm NRBO to optimize the model parameters of the aquatic prediction model, two rules, namely the Newton - Raphson search rule and the trap avoidance operator TAO, are used to explore the entire search process;

[0099] The steps for optimizing the model parameters of the aquatic prediction model are as follows:

[0100] S510. NRBO starts to search for the optimal solution by generating an initial random population within the corresponding boundaries of the candidate solutions of the phytoplankton biomass prediction results. Each population consists of fuzzy decision variables, and the formula for generating the random population is as follows:

[0101]

[0102] Among them, represents the j-th dimensional position of the n-th population; N p represents the number of populations in the initial random population; Rand represents a random number between (0, 1); dim represents the total number of problem dimensions preset for the n-th population; lb represents the lower bound, which is the minimum value of the decision variable corresponding to the problem dimension of the corresponding population; ub represents the upper bound, which is the maximum value of the decision variable corresponding to the problem dimension of the corresponding population; the model of the aquatic prediction model normalizes the phytoplankton biomass, and the lower bound and upper bound are 0 and 1 respectively;

[0103] S520. Control the vector matrix X through the Newton - Raphson search rule n , and obtain the position of the explored feasible region; and adjust the global search and local search states during the exploration process through the adaptive coefficient δ;

[0104] In the said S520, controlling the vector matrix X through the Newton - Raphson search rule n involves the following calculation formulas:

[0105]

[0106] y w = r1 × (Mean(Z n+1 + x n )) + r1 × △x)

[0107] y b = r1 × (Mean(Z n+1 + x n )) - r1 × △x)

[0108]

[0109] where randn is a normally distributed random number with a mean of 0 and a variance of 1; △x represents the search step size; NRSR represents the search step size of the Newton - Raphson search rule; x n represents the position; x w represents the worst position (x n + △x); Z n+1 represents the new search position; x b represents the best position (x n - △x); y w and y b respectively represent the positions of two vectors generated using Z n+1 and x n ; r1 represents a random number between (0, 1); Rand(1, Dim) is a random number in the interval [1, Dim];

[0110] The calculation formula of the adaptive coefficient δ is as follows:

[0111]

[0112] where IT represents the current iteration, and Max_ IT represents the maximum number of iterations.

[0113] S530. Record the current vector position as the corresponding element in the vector matrix X n and use TAO to change the vector position of the previous vector in the next iteration

[0114] In S530, by combining the best position Xb and the current vector position when the value of rand is less than DF, the solution involves the following formula:

[0115]

[0116] where θ1 and θ2 are uniformly distributed random numbers between (-1, 1) and (-0.5, 0.5) respectively, DF represents the determinant controlling the performance of NRBO, μ1 and μ2 are random numbers; during the exploration iteration process, the obtained vector position value is used as the new current vector position to continue the exploration;

[0117]

[0118] μ1 is a number within the range of (0, 1), and the above equation can be simplified as:

[0119] μ1 = β × 3 × rand + (1 - β)

[0120] μ2 = β × rand + (1 - β)

[0121] where β represents a binary number, and the value of β is 1 or 0; if the value of μ1 is greater than or equal to 0.5, the value of β is 0; otherwise, β is 1.

[0122] S540. Complete the update of the model parameters in the aquatic prediction model. The updated model parameters of the aquatic prediction model include the search step size Δx, the search step size of the Newton - Raphson search rule, and the weights of the global or local search capabilities.

[0123] Taking "Processing of High - Frequency Monitoring Data of Phytoplankton Biomass in Poyang Lake and Model Prediction" as an example, the execution process of the present invention is as follows:

[0124] 1. Acquisition of High-Frequency Monitoring Data at Tangyin Section of Poyang Lake

[0125] Field high-frequency monitoring equipment is deployed in Poyang Lake to automatically collect physical, chemical, and biological parameters related to the biomass of phytoplankton in the water body every hour, obtaining in-situ high-frequency water ecological monitoring results in the field. The monitoring data includes water temperature, conductivity, total dissolved solids, permanganate index, dissolved oxygen, pH value, ammonia nitrogen, nitrite, total nitrogen, total phosphorus, and chlorophyll, etc. These data are sent to the data processing system through wireless transmission.

[0126] 2. Processing of Outliers in High-Frequency Monitoring Data of Tangyin Water Ecology (Taking Chlorophyll as an Example)

[0127] The invention uses the Isolation Forest algorithm to process the original monitoring data and detect abnormal data caused by factors such as sensor failures and external environmental interferences. The core idea of the Isolation Forest algorithm is to separate abnormal points from normal points by constructing multiple isolation trees, thereby achieving anomaly detection (as shown in Figure 1 ). An isolation tree is a special type of binary tree constructed by recursively partitioning the data set. Each isolation tree is trained (sampled, without replacement) on a subset of the data set. During the construction process, the algorithm randomly selects a feature and a split value to divide the data into two parts. This process continues until one of the following conditions is met: (1) Each data point is isolated in a separate leaf node. (2) The predefined maximum tree depth is reached. Each leaf node at the bottom of the tree contains a data point, indicating that the point has been completely isolated. Due to the rarity of anomalies and their significant differences from other points, anomalies are more likely to be quickly isolated during random partitioning and tend to end in leaf nodes closer to the root, resulting in shorter path lengths (the distance from the root to the leaf node). In contrast, densely distributed normal points require more steps to be completely isolated and thus have longer path lengths. The Isolation Forest consists of multiple independently constructed isolation trees. By aggregating the path length results of all isolation trees, the algorithm calculates the anomaly score for each data point. Shorter path lengths correspond to higher anomaly scores, indicating a higher likelihood that the point is an anomaly. By combining the results of multiple isolation trees, the Isolation Forest can effectively detect anomalies and normal data points in high-dimensional complex data distributions.

[0128] The Isolation Forest algorithm calculates the anomaly score s(x) of the observed value x by normalizing the path length h(x):

[0129]

[0130] Among them, h(x) represents the path length of the collected data x in the isolation tree; E[h(x)] represents the average path length of the collected data x in all isolation trees of the isolation forest training model, and c(n) represents the average path length of unsuccessful searches for n pieces of collected data; when E[h(x)] approaches 0, the score approaches 1; therefore, a score value close to 1 indicates an anomaly; when E[h(x)] approaches n - 1, the score approaches 0; in addition, when E[h(x)] approaches c(n), the score approaches 0.5; therefore, a score value less than 0.5 and close to 0 indicates a normal point Figure 2 )

[0131] The optimized isolation forest algorithm in the high-frequency data scenario compared with the traditional algorithm mainly improves the following aspects: (1) Computational efficiency improvement: Improve the path length calculation to enhance the score stability. High-frequency data is susceptible to noise or short-term fluctuations, which may lead to misjudgment of outliers by the isolation forest. The isolation forest method in the present invention can reduce the impact of extreme path lengths on the anomaly score by smoothing the path length calculation E[h(x)], thereby improving the stability of detection. (2) Adaptive forest depth: Reduce the computational latency. Traditional isolation forests use trees with a fixed depth, but in a high-frequency data environment, this may result in unnecessary computational overhead. Setting the maximum depth adaptively based on the data distribution can reduce redundant calculations, improve real-time performance, and maintain the anomaly detection ability. (3) Incremental learning: Improve real-time performance. Due to the large volume of high-frequency data, the model needs to be continuously updated. Using a sliding window to update the forest can avoid reconstructing the entire forest, thus improving efficiency. (4) Enhanced robustness: Reduce false detections and missed detections. High-frequency data is often accompanied by noise, which may lead to false detections or missed detections of outliers. When calculating the score s(x), a density-based anomaly correction term is introduced to reduce the sensitivity to accidental outliers.

[0132] Specific steps for detecting outliers in the chlorophyll data of the Tangyin section of Poyang Lake: (1) Generate data and introduce outliers. In MATLAB, simulate 1000 groups of data and randomly select 5% of the data as outliers. (2) Conduct isolation forest anomaly detection. Use the iforest() function to train the isolation forest and perform anomaly detection. (3) Visualize the outlier detection results. Plot the anomaly score histogram and display the detection threshold. (4) Generate an outlier distribution map. Plot the two-dimensional scatter plot of flow velocity and chlorophyll and mark the outliers. (5) Output the chlorophyll results after processing the outliers.

[0133] 3. Wavelet threshold denoising of high-frequency monitoring data of the aquatic ecosystem in the Tangyin section of Poyang Lake (taking chlorophyll as an example)

[0134] The high-frequency data after the above-mentioned Isolation Forest outlier processing enters the noise reduction module. Wavelet threshold denoising removes the wavelet coefficients corresponding to the noise in each frequency band and retains the wavelet coefficients of the original signal according to the characteristics of the noise coefficient and the different intensity distributions of the wavelet coefficients of the signal in different frequency bands. Then, wavelet reconstruction is performed on the processed coefficients to obtain a pure signal.

[0135] Specific steps for noise reduction processing of Tangyin chlorophyll data:

[0136] Step 1: Obtain the aquatic monitoring data (original signal) after removing abnormal data, select the wavelet type and the N levels of wavelet decomposition, and perform N-layer wavelet decomposition on the aquatic monitoring data with noise to obtain the wavelet coefficients of each layer, where N is a preset value;

[0137] Step 2: Quantify the high-frequency wavelet coefficients from the first layer to the Nth layer using an adaptive threshold (threshold processing) to remove the relevant noise; the expression for obtaining the adaptive threshold is the root mean square interpolation threshold function:

[0138]

[0139] where ω j,k and are the wavelet coefficients before and after threshold quantization respectively, T represents the threshold, μ represents a preset adjustment parameter and 0 ≤ μ ≤ 1. By appropriately adjusting the value of μ, a better noise reduction effect can be achieved; when μ = 0, the root mean square interpolation threshold function is denoted as the soft threshold function equation. The hard threshold directly sets all wavelet coefficients less than the threshold T to 0 without further adjustment, which belongs to a "retain or suppress" operation; when μ = 1, the root mean square interpolation threshold function is denoted as the hard threshold function equation. The soft threshold processing formula subtracts a fixed value T from the coefficients exceeding the threshold, thereby achieving amplitude contraction and achieving a smoothing denoising effect;

[0140] The soft threshold function equation is:

[0141]

[0142] The hard threshold function equation is:

[0143]

[0144] Step 3: Reconstruct the processed signal (wavelet reconstruction) to obtain the wavelet reconstruction data (denoised signal) based on the noise reduction of the aquatic monitoring data. As Figure 3 shown, the first row of the figure represents the actual observed chlorophyll after outlier processing, the second row of the figure is the signal value with a wavelet coefficient of 3, and the third row is the actual chlorophyll used after noise reduction.

[0145] 4. Construction of the prediction model for the biomass of phytoplankton in Poyang Lake

[0146] The outlier and noise-reduced high-frequency monitoring data of Tangyin water ecology are input into the phytoplankton biomass prediction module. All processed data are stored in an excel file, and the data should be time series data, where the first column is the date in the format of yyyy / mm / dd, the second column is the time in the format of hh:mm, the third column is the target variable (i.e., chlorophyll, representing phytoplankton biomass), and the fourth column to the last column are the input features of the model, including water temperature, conductivity, total dissolved solids, permanganate index, dissolved oxygen, pH value, ammonia nitrogen, nitrite, total nitrogen, total phosphorus, and nitrogen-phosphorus ratio.

[0147] This invention is based on the XGBoost machine learning model. The XGBoost algorithm stands out as a powerful learning method, which is rooted in the tree model in the field of machine learning applications. Using gradient boosting, it utilizes the advantages of each weak classifier to create a robust classifier while iteratively minimizing the residuals of the previous model along the gradient direction, thus generating an improved model. XGBoost is more efficient in processing large datasets and complex models and also performs excellently in preventing overfitting and improving generalization ( Figure 1 ).

[0148] To improve the prediction accuracy of the model, this invention uses the Newton-Raphson optimization algorithm to optimize the hyperparameters of the XGBoost model. Although XGBoost reduces the calculation of finding the optimal splitting point through pre-sorting and approximation algorithms, it still needs to traverse the dataset during the node splitting process. The space complexity of the pre-sorting process is high, and it is necessary to store the eigenvalue of the sample and the corresponding gradient statistics, resulting in a large amount of memory consumption. To solve this problem, the software uses the Newton-Raphson optimization algorithm. Two rules are used to explore the entire search process: the Newton-Raphson search rule and the trap avoidance operator (TAO), and several groups of matrices are used to further explore the best results.

[0149] (1) Initialization: NRBO starts searching for the optimal solution by generating an initial random population within the corresponding boundaries of the candidate solutions of the phytoplankton biomass prediction results. Each population consists of fuzzy decision variables, and the formula for generating the random population is as follows:

[0150]

[0151] Among them, represents the j-th dimension position of the n-th population; N prepresents the number of populations in the initial random population; Rand represents a random number between (0, 1); dim represents the total number of problem dimensions preset for the nth population; lb represents the lower bound, which is the minimum value of the decision variable for the corresponding problem dimension of the population; ub represents the upper bound, which is the maximum value of the decision variable for the corresponding problem dimension of the population; the model of the aquatic prediction model normalizes the phytoplankton biomass, and the lower bound and upper bound are 0 and 1 respectively;

[0152] (2) Newton - Raphson search rule: Control the vector matrix X through the Newton - Raphson search rule n , obtain the position of the explored feasible region and get a better position; NRSR is proposed based on the concept of the Newton - Raphson method, aiming to promote the exploration trend and accelerate convergence; and adjust the global search and local search states during the exploration process through the adaptive coefficient δ;

[0153] Control the vector matrix X through the Newton - Raphson search rule n The involved calculation formulas are as follows:

[0154]

[0155] y w = r1×(Mean(Z n+1 + x n ) + r1×△x)

[0156] y b = r1×(Mean(Z n+1 + x n ) - r1×△x)

[0157]

[0158] where randn is a normally distributed random number with a mean of 0 and a variance of 1; △x represents the search step size; NRSR represents the search step size of the Newton - Raphson search rule; x n represents the position; x w represents the worst position (x n + △x); Z n+1 represents the new search position; x b represents the best position (x n - △x); y w and y b respectively represent the positions of two vectors generated using Z n+1 and x n ; r1 represents a random number between (0, 1); Rand(1, Dim) is a random number between the interval [1, Dim], with a weak dimension decision variable;

[0159] Empirically, the proposed algorithm must be able to achieve a balance between diversity and aggregation in order to find the optimal solution in the search space and finally converge to the global solution. The algorithm can be enhanced by applying an adaptive coefficient called δ. The calculation formula of the adaptive coefficient δ is as follows:

[0160]

[0161] where IT represents the current iteration, and Max_ IT represents the maximum number of iterations. To maintain a balance between the exploration phase and the development phase, the parameter δ adjusts itself during the iteration process.

[0162] The adaptive coefficient δ plays a role of dynamic adjustment in the algorithm, aiming to balance global search (exploration) and local search (development) to enhance the search efficiency of the algorithm in different phases (the initial exploration phase and the later convergence phase). This dynamic adjustment helps to avoid the algorithm falling into a local optimum and at the same time improves the convergence speed when approaching the optimal solution.

[0163] The main steps are as follows: A. Exploration phase (early iteration): In the early stage of the algorithm, when IT is small, the value of δ is closer to 1. At this time, the algorithm tends to explore, that is, to conduct a wider search and try different regions of the solution space. In this case, a larger value of δ can prompt the algorithm to conduct a large-scale random search to explore more possible solutions and avoid falling into a local optimum too early. B. Development phase (late iteration): As the number of iterations IT increases, δ will gradually decrease. When IT is close to Max_IT, δ will approach 0 or be close to a negative value. At this time, the algorithm will focus more on the vicinity of the relatively good solutions that have been found and conduct a fine local search, aiming to converge to the global optimal solution. This is the characteristic of the development phase, which means that the algorithm reduces the randomness of exploration and relies more on the solutions that have been obtained. The role of the adjustment method of δ, with a higher value of δ, the algorithm focuses more on exploration, which means higher randomness and is suitable for use in the initial stage of the search space. With a lower value of δ, the algorithm reduces randomness and relies more on the currently known solutions for fine search, which helps to converge to the global optimal solution.

[0164] S530. Record the current vector position as the corresponding to an element in the vector matrix X n Use TAO to change the vector position of the previous vector in the next iteration step

[0165] Combine the best position Xb and the current vector position When the value of rand is less than DF, then the solution involves the following formula:

[0166]

[0167] Among them, θ1 and θ2 are uniformly distributed random numbers between (-1, 1) and (-0.5, 0.5) respectively, DF represents the determinant that controls the performance of NRBO, μ1 and μ2 are random numbers; during the exploration iteration process, the obtained vector position value is used as the new current vector position to continue the exploration;

[0168]

[0169] μ1 is a number within the range of (0, 1), and the above equation can be simplified:

[0170] μ1 = β × 3 × rand + (1 - β)

[0171] μ2 = β × rand + (1 - β)

[0172] where β represents a binary number, and the value of β is 1 or 0; if the value of μ1 is greater than or equal to 0.5, then the value of β is 0; otherwise, β is 1.

[0173] θ1 and θ2 affect the update process by adjusting μ1 and μ2, thereby enhancing the quality of the solution. Specifically, θ1 and θ2 will adjust the solution vector by weighted combination of the values of μ1 and μ2. μ1 and μ2 affect the exploratory and exploitative nature of the search strategy, while θ1 and θ2 regulate the amplitude of this exploration. θ1 and θ2 regulate the amplitude and direction of the update of the solution vector. μ1 and μ2 are determined by the binary β, which controls the exploratory and exploitative nature of the search stage. μ1 is determined by the random number rand and β, and the value of β is dynamically adjusted according to the size of μ1 itself. This process makes the value of μ1 contain both randomness and can adjust the search strategy according to the current value, thereby achieving a certain degree of adaptive optimization. Due to the randomness of the parameter selection of μ1 and μ2, the population becomes more diverse and escapes from the local optimal solution, which helps to improve its diversity.

[0174] S540. Complete the update of the model parameters in the aquatic prediction model. The updated model parameters of the aquatic prediction model include the search step size Δx, the search step size of the Newton - Raphson search rule, and the weights of the global or local search capabilities.

[0175] Finally, after data preprocessing and model hyperparameter optimization, the output result of the phytoplankton biomass prediction model of Poyang Lake is displayed to the user through the output module. Figure 4 Illustrates the comparison between the predicted phytoplankton biomass output by the prediction model and the measured chlorophyll. The small upper - left figure shows the error between the measured value and the model output value.

[0176] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0177] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art may still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. In-situ high-frequency monitoring data processing and prediction method for phytoplankton biomass, characterized in that, include: S100, collecting aquatic monitoring data in the water body of the test area in real time through the deployed high-frequency monitoring equipment, and sending the collected aquatic monitoring data to the data processing end through the wireless transmission network; S200, the outlier detection module in the data processing end obtains the aquatic monitoring data to be processed, and identifies and removes the abnormal data in the aquatic monitoring data by using a plurality of isolation trees constructed by using an isolation forest algorithm based on a random partitioning method; S300, performing wavelet threshold denoising processing on the aquatic monitoring data after removing abnormal data, analyzing the difference in wavelet coefficients between the signal and the noise through wavelet transform, using a threshold method to remove the noise component, retaining the signal part, and obtaining wavelet reconstructed data; S400, using XGBoost machine learning algorithm to model the obtained wavelet reconstruction data to obtain an aquatic prediction model; and predicting the phytoplankton biomass based on the aquatic prediction model; S500, based on the modeling data of the XGBoost machine learning algorithm and the prediction results of phytoplankton biomass, uses the Newton-Raphson optimization algorithm NRBO to optimize the model parameters of the aquatic prediction model to ensure the optimal performance of the obtained aquatic prediction model under different data sets and environmental conditions, and updates the aquatic prediction model in the current database based on the optimized model parameters.

2. The in-situ high-frequency monitoring data processing and prediction method for phytoplankton biomass according to claim 1, characterized in that: The collection frequency of the high-frequency monitoring equipment deployed in S1 is a preset value. The deployed high-frequency monitoring equipment is composed of multiple different types of sensors. The aquatic monitoring data includes water temperature, conductivity, total dissolved solids, permanganate index, dissolved oxygen, pH value, ammonia nitrogen, nitrite, total nitrogen, total phosphorus and chlorophyll.

3. The in-situ high-frequency monitoring data processing and prediction method for phytoplankton biomass according to claim 1, characterized in that: The S200 includes: S210, obtaining aquatic monitoring data to be processed; S220, summarizing the collection results of the same data type in the historical aquatic monitoring data, and constructing an isolation forest model consisting of multiple isolation trees using an isolation forest algorithm, and training the constructed isolation forest model with the summary set of the collection results of the same data type in the historical aquatic monitoring data, to obtain an isolation forest training model corresponding to the corresponding data type; The collection results of the same data type in the historical aquatic monitoring data are summarized as a data set; the isolation tree is a binary tree, which is constructed by recursively partitioning the data set, and each isolation tree is trained on a subset of the data set; during the construction process, a split value is randomly selected to divide the data set into two partitioned data sets; the above recursive partitioning data set operation is repeated until one of the following conditions is met: (1) each data point is isolated in a separate leaf node; (2) a predefined maximum tree depth is reached; S230, input each collected data in the aquatic monitoring data to be processed into the isolation forest training model corresponding to the corresponding data type, aggregate the path lengths of all isolation trees of each collected data in the corresponding isolation forest training model, and calculate the anomaly score of each collected data; record the anomaly score of the collected data x as s(x), Among them, h(x) represents the path length of the collected data x in the isolation tree; E[h(x)] represents the average path length of the collected data x in all isolation trees of the isolation forest training model, and c(n) represents the average path length of unsuccessful searches for n collected data; when E[h(x)] = c(n), s(x) = 0.5; S240. Identify and eliminate abnormal data in the aquatic monitoring data, and the abnormal scores corresponding to the abnormal data in the aquatic monitoring data all belong to the interval [0, 0.5]; The types of the abnormal data include the abnormal data caused by sensor failures and external interferences.

4. The in-situ high-frequency monitoring data processing and prediction method for phytoplankton biomass according to claim 1, characterized in that: The S300 includes: S310. Obtain the aquatic monitoring data after eliminating the abnormal data, select the wavelet type and the N levels of wavelet decomposition, perform N-layer wavelet decomposition on the aquatic monitoring data with a noisy signal to obtain the wavelet coefficients of each layer, where N is a preset value; S320. Quantify the high-frequency wavelet coefficients from the first layer to the Nth layer with an adaptation threshold to remove the relevant noise; the expression for obtaining the adaptation threshold is the root mean square interpolation threshold function: where ω j,k and are the wavelet coefficients before and after threshold quantization respectively, T represents the threshold, μ represents a preset adjustment parameter and 0 ≤ μ ≤ 1; S330. Reconstruct the processed signal to obtain the wavelet reconstruction data after noise reduction based on the aquatic monitoring data.

5. The method for processing and predicting in-situ high-frequency monitoring data of phytoplankton biomass according to claim 1, wherein: The aquatic prediction model in the S400 is trained by the XGBoost machine learning algorithm for the wavelet reconstruction data corresponding to the historical aquatic monitoring data and the phytoplankton biomass corresponding to the wavelet reconstruction data.

6. The in-situ high-frequency monitoring data processing and prediction method for phytoplankton biomass according to claim 1, characterized in that: In the process of optimizing the model parameters of the aquatic prediction model by using the Newton-Raphson optimization algorithm NRBO in the S500, two rules, namely the Newton-Raphson search rule and the trap avoidance operator TAO, are used to explore the entire search process; The steps for optimizing the model parameters of the aquatic prediction model are as follows: S510. NRBO starts searching for the optimal solution by generating an initial random population within the corresponding boundaries of the candidate solutions of the phytoplankton biomass prediction results. Each population consists of fuzzy decision variables, and the formula for generating the random population is as follows: Among them, represents the j-th dimensional position of the n-th population; N p represents the number of populations in the initial random population; Rand represents a random number between (0, 1); dim represents the total number of problem dimensions preset for the n-th population; lb represents the lower bound, which is the minimum value of the decision variable corresponding to the problem dimension of the corresponding population; ub represents the upper bound, which is the maximum value of the decision variable corresponding to the problem dimension of the corresponding population; S520. Control the vector matrix X through the Newton-Raphson search rule n , obtain the position of the explored feasible region; and adjust the global search and local search states during the exploration process through the adaptive coefficient δ; S530. Record the current vector position as The corresponding vector matrix X n an element in, use TAO to change the previous vector position the vector position in the next iteration step S540. Complete the update of the model parameters in the aquatic prediction model. The updated model parameters of the aquatic prediction model include the search step Δx, the search step of the Newton-Raphson search rule, and the weights of the global or local search capabilities.

7. The in-situ high-frequency monitoring data processing and prediction method for phytoplankton biomass according to claim 6, characterized in that: In the S520, the vector matrix X is controlled by the Newton-Raphson search rule n The involved calculation formula is as follows: y w = r1 × (Mean(Z n+1 + x n ) + r1 × Δx) y b = r1 × (Mean(Z n+1 + x n ) - r1 × Δx) where randn is a random number from a normal distribution with mean 0 and variance 1; △x represents the search step size; NRSR represents the search step size of the Newton - Raphson search rule; x n represents the position; x w represents the worst position (x n +△x); Z n+1 represents the new search position; x b represents the best position (x n -△x); y w and y b represent the positions of two vectors generated using Z n+1 and x n respectively; r1 represents a random number between (0, 1); Rand(1, Dim) is a random number in the interval [1, Dim]; The calculation formula for the adaptive coefficient δ is as follows: where IT represents the current iteration, and Max_ IT represents the maximum number of iterations.

8. The in-situ high-frequency monitoring data processing and prediction method for phytoplankton biomass according to claim 7, characterized in that: The optimal combination position Xb and the current vector position in S530 When the value of rand is less than DF, the solution The related formula is as follows: Among them, θ1 and θ2 are uniformly distributed random numbers between (-1, 1) and (-0.5, 0.5) respectively, DF represents the determinant controlling the performance of NRBO, and μ1 and μ2 are random numbers; μ1 = β × 3 × rand + (1 - β) μ2 = β × rand + (1 - β) Among them, β represents a binary number, and the value of β is 1 or 0; if μ1 is greater than or equal to 0.5, the value of β is 0; otherwise, the value of β is 1.