Country wetland water quality monitoring method based on unmanned aerial vehicle multispectral image and machine learning
By using the DUPLEX algorithm to optimize sampling point layout, improved atmospheric radiation correction model and watershed segmentation algorithm in water quality monitoring, a multi-dimensional feature system is built, and a two-stage integrated learning framework combined with spatial interpolation method is built, the problems of time-consuming and labor-intensive and incomplete spatial data are solved, and efficient and accurate water quality monitoring is achieved.
Patent Information
- Application Number
- CN202510143428.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional water quality monitoring methods are time-consuming and labor-intensive, and it is difficult to provide comprehensive spatial data. The existing multi-spectral imaging technology of drone has problems such as insufficient optimization of sampling strategies, imperfect feature extraction, unstable performance of machine learning algorithms and low spatial distribution prediction accuracy.
The DUPLEX algorithm is used to optimize the sampling point layout, and combined with the improved 6S atmospheric radiation correction model and the watershed segmentation algorithm, a multi-dimensional feature system is built, and the optimal feature subset is screened through recursive feature elimination and cross-validation. Then, a two-stage integrated learning framework is constructed and the prediction results are optimized in combination with spatial interpolation methods.
It improves the scientificity and efficiency of sampling and monitoring, improves data processing and feature extraction, improves the accuracy and reliability of water quality parameter estimation, optimizes the accuracy of spatial distribution prediction, and achieves low cost and high efficiency of water quality monitoring.
Smart Images

Figure CN120107828A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of water quality monitoring methods, and specifically relates to a rural wetland water quality monitoring method based on unmanned aerial vehicle multispectral imaging and machine learning, in particular a water quality parameter spatial distribution estimation method using a two-stage integrated learning framework. Background Art
[0002] Rural wetlands are important ecological landscape units that provide important ecosystem services to surrounding residents. Their water quality monitoring is crucial to protecting wetland ecosystems. However, traditional water quality monitoring methods rely on field sampling and laboratory analysis, which is time-consuming and labor-intensive, and it is difficult to provide comprehensive spatial data. This method can usually only cover a relatively small spatial scale. Due to the spatial heterogeneity of water bodies, the accuracy is uncertain when extrapolated to a larger scale. Although satellite remote sensing can provide large-scale water quality information, there is a trade-off between spatial resolution and temporal resolution, and it is difficult to use for small-area water quality monitoring due to weather conditions.
[0003] In recent years, the development of drone technology has provided a new means for water quality monitoring, which can obtain ultra-high-resolution images with strong flexibility and efficiency. Although drone multispectral imaging technology can capture the spatial variability of water bodies and generate spatial distribution maps of water quality parameters, the existing technology still has many shortcomings: first, there is a lack of targeted sampling strategy optimization methods, which affects the representativeness of data acquisition; second, feature extraction and selection methods need to be improved, making it difficult to fully utilize multispectral image information; third, the performance of a single machine learning algorithm is not stable enough, affecting the reliability of the estimation results; finally, the accuracy of spatial distribution prediction needs to be further improved. Summary of the invention
[0004] The purpose of the present invention is to overcome the defects in the prior art and provide a rural wetland water quality monitoring method based on drone multispectral imaging and machine learning. The present invention has made innovations in sampling strategy, data processing, model construction and spatial prediction: the DUPLEX algorithm is used to optimize the sampling point layout to ensure that the samples are evenly distributed in the feature space; the improved 6S atmospheric radiation correction model and watershed segmentation algorithm are introduced to improve data quality; a multidimensional feature system including spectral features, geometric features and texture features is constructed; recursive feature elimination and cross-validation are used for feature optimization; a two-stage integrated learning framework based on XGBoost and LightGBM is constructed, and the prediction results are optimized in combination with spatial interpolation methods.
[0005] The specific technical solutions adopted by the present invention are as follows:
[0006] The present invention provides a rural wetland water quality monitoring method based on drone multispectral imaging and machine learning, which is as follows:
[0007] S1. Adopt a stratified optimization sampling strategy in the target rural wetland research area, collect UAV multispectral remote sensing images and conduct in-situ water quality sampling at each sampling point;
[0008] S2, measuring water quality parameters of the water samples collected in S1;
[0009] S3, image stitching, geometric precision correction and improved 6S atmospheric radiation correction model preprocessing are performed on the multispectral remote sensing images collected in S1 to generate orthophotos, and normalization is performed. Finally, the water body in the study area is extracted by the improved watershed segmentation algorithm;
[0010] S4. Based on the orthophoto obtained in S3, a multidimensional feature system including spectral features, geometric features and texture features is constructed, and the feature information corresponding to the sampling points is extracted using geographic information system software and used as candidate feature variables;
[0011] S5, applying recursive feature elimination and cross-validation algorithm to select the optimal feature subset from the candidate feature variables obtained in S4;
[0012] S6, using stratified random sampling method, divide the water quality parameter data in S2 into training set and validation set;
[0013] S7, based on the optimal feature subset screened by S5 and the training set divided by S6, a two-stage integrated learning framework is constructed to obtain the estimation model of each water quality parameter;
[0014] S8, using the validation set in S6, evaluate the performance of the model constructed in S7 through indicators and select the optimal model;
[0015] S9, based on the optimal model selected in S8, the spatial distribution of water quality parameters is inverted for the entire study area, and a high-resolution thematic map of water quality parameters is generated;
[0016] S10. A strategy combining spatial interpolation and machine learning models was adopted to optimize the inversion results of the spatial distribution of water quality parameters in S9 using Kriging interpolation to improve the prediction accuracy of the spatial distribution of water quality parameters.
[0017] Preferably, in S1, the sampling points are arranged using the DUPLEX algorithm to ensure that the samples are evenly distributed in the feature space; image acquisition is achieved using a DJI Matrice300 RTK drone platform equipped with a Changguang Yuchen AQ600 Pro multispectral camera, which includes five 3.2M pixel multispectral CMOS sensors and one 12.3M pixel panchromatic sensor for collecting reflectance information of ground objects in five narrow bands: blue light, green light, red light, red edge and near infrared.
[0018] Preferably, in S2, the water quality parameters include total nitrogen, total phosphorus, chemical oxygen demand and turbidity; wherein, total nitrogen and total phosphorus are determined by spectrophotometry, chemical oxygen demand is determined by permanganate index method, and turbidity is determined by nephelometer.
[0019] Preferably, in S3, the normalization processing method is as follows:
[0020] Spectral Response=(DN-DNmin) / (DNmax-DNmin)
[0021] Among them, DN represents the digital number value of the pixel; DNmin represents the minimum DN value of the band; DNmax represents the maximum DN value of the band; Spectral Response represents the normalized spectral response value, ranging from [0,1].
[0022] Preferably, in S4, the spectral features include the first-order derivative and the second-order derivative of the spectrum, the geometric features include the shape factor and the complexity index, and the texture features include the gray-level co-occurrence matrix;
[0023] The calculation method of the first-order derivative of the spectrum is as follows:
[0024] FD(λi)=[R(λi+1)-R(λi)] / Δλ
[0025] Wherein, FD(λi) represents the first-order derivative value at wavelength λi; R(λi) represents the reflectivity value at wavelength λi; R(λi+1) represents the reflectivity value at wavelength (λi+1); Δλ represents the interval between adjacent wavelengths; i represents the wavelength number;
[0026] The calculation method of the second-order derivative of the spectrum is as follows:
[0027] SD(λi)=[R(λi+2)-2R(λi+1)+R(λi)] / (Δλ) 2
[0028] Wherein, SD(λi) represents the second-order derivative value at wavelength λi; R(λi+2) represents the reflectivity value at wavelength (λi+2).
[0029] Preferably, in S6, the water quality parameter data is divided into a ratio of 70% training set and 30% validation set.
[0030] Preferably, the S7 is as follows:
[0031] Based on the optimal feature subset screened by S5 and the training set divided by S6, XGBoost and LightGBM were first used to build the base learner. Then, the Stack strategy was adopted to use the prediction results of the base learner as new features, and logistic regression was used as the meta-learner to build the estimation models of various water quality parameters.
[0032] Preferably, in S8, the indicators include a coefficient of determination and a root mean square error;
[0033] The calculation method of the determination coefficient is as follows:
[0034]
[0035] Among them, R 2 represents the coefficient of determination, ytrue,i represents the measured value of the i-th sample, and ypred,i represents the predicted value of the i-th sample. represents the average value of the measured value, i represents the wavelength number;
[0036] The calculation method of the root mean square error is as follows:
[0037]
[0038] Here, RMSE represents the root mean square error, and n represents the total number of samples.
[0039] Preferably, in S10, the calculation method of Kriging interpolation is as follows:
[0040] Z*(s0)=∑λiZ(si)
[0041] Constraint: ∑λi=1
[0042] Among them, Z*(s0) is the estimated value of the point to be predicted, Z(si) is the observed value of the i-th known sampling point, λi is the Kriging weight, si is the spatial position of the sampling point, and s0 is the spatial position of the point to be predicted.
[0043] Compared with the prior art, the present invention has the following beneficial effects:
[0044] (1) Improved the scientificity and efficiency of sampling and monitoring. The present invention adopts the DUPLEX algorithm to arrange sampling points, ensuring that samples are evenly distributed in the feature space and improving the representativeness of sampling. At the same time, it uses drone multispectral remote sensing technology for rapid data collection, which greatly improves the monitoring efficiency compared to traditional field sampling and can obtain continuous spatial distribution data.
[0045] (2) Improved data processing and feature extraction methods. The present invention uses an improved 6S atmospheric radiation correction model for image preprocessing, uses an improved watershed segmentation algorithm to extract water bodies, and constructs a multidimensional feature system including spectral features (first-order derivatives, second-order derivatives), geometric features (shape factor, complexity index) and texture features (gray-level co-occurrence matrix), which significantly improves data quality and information extraction capabilities.
[0046] (3) Improved the accuracy and reliability of water quality parameter estimation. This paper innovatively constructed a two-stage integrated learning framework. First, XGBoost and LightGBM were used to build a base learner. Then, the Stack strategy was used to use the prediction results of the base learner as new features. Logistic regression was used as a meta-learner, which effectively improved the estimation accuracy of water quality parameters such as total nitrogen, total phosphorus, chemical oxygen demand, and turbidity.
[0047] (4) Optimized the accuracy of spatial distribution prediction. The present invention adopts a strategy combining spatial interpolation with machine learning models, and uses Kriging interpolation to optimize the prediction results of the machine learning model, which significantly improves the accuracy of spatial distribution prediction of water quality parameters and provides a reliable basis for understanding the spatiotemporal variation of wetland water quality.
[0048] (5) Low cost and high efficiency of water quality monitoring are achieved. Compared with traditional laboratory analysis methods, the present invention greatly reduces the cost of manpower and material resources; compared with satellite remote sensing, drone remote sensing has higher flexibility and temporal and spatial resolution, and is not restricted by clouds and weather conditions, providing an economically feasible technical solution for the normalized and dynamic monitoring of rural wetland water quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a schematic diagram of the process of the present invention.
[0050] Figure 2 An overview map of the study area.
[0051] Figure 3 This is the water quality measurement data analysis for two sampling periods; 0615 represents June 15, 2023, and 1126 represents November 26, 2023, mainly involving four water quality parameters: Figure 3 a is total nitrogen (TN), Figure 3 b is total phosphorus (TP), Figure 3 c is chemical oxygen demand (COD) and Figure 3 d is turbidity (TUB).
[0052] Figure 4 This is a schematic diagram of water bodies in the study area; Figure 4 a is the water body on June 15, 2023, Figure 4b is the water body on November 26, 2023. DETAILED DESCRIPTION
[0053] The present invention is further described and illustrated below in conjunction with the accompanying drawings and specific embodiments. The technical features of each embodiment of the present invention can be combined accordingly without conflicting with each other.
[0054] like Figure 1 As shown, a rural wetland water quality monitoring method based on drone multispectral imaging and machine learning is provided by the present invention. The method is specifically as follows:
[0055] S1. A stratified optimization sampling strategy was adopted in the target rural wetland study area, and UAV multispectral remote sensing image acquisition and in situ water quality sampling were carried out at each sampling point.
[0056] As a preferred embodiment of the present invention, in this step, image acquisition is achieved by using a DJI Matrice300 RTK drone platform equipped with a Changguang Yuchen AQ600 Pro multispectral camera, wherein:
[0057] 1) DJI Matrice300 RTK drone has the following technical features:
[0058] The RTK (Real Time Kinematic) positioning technology can achieve centimeter-level positioning accuracy. It is equipped with a dual IMU redundant navigation system to ensure flight stability. The flight altitude is controlled within the range of 80-120 meters, ensuring that the ground resolution is better than 0.1 meter.
[0059] 2) Changguang Yuchen AQ600 Pro multispectral camera parameter configuration:
[0060] It is equipped with 5 independent spectral sensors, which collect blue (450±16nm), green (550±16nm), red (660±16nm), red edge (730±16nm) and near-infrared (840±26nm) bands respectively. The spectral resolution is better than 10nm, the radiation resolution is 12bit, and it has a built-in light sensor for radiation correction.
[0061] 3) The sampling points are arranged using the DUPLEX (double optimization) algorithm, which has the following advantages:
[0062] By maximizing the Mahalanobis distance between sample points, the samples are evenly distributed in the feature space. A hierarchical optimization strategy is adopted to balance the spatial representativeness and spectral feature differences of the samples. The Kennard-Stone algorithm is used to divide the sample set and improve the sampling efficiency.
[0063] S2. Conduct laboratory analysis on the water samples collected in S1 and determine the water quality parameters.
[0064] As a preferred embodiment of the present invention, in this step, the water quality parameters include total nitrogen (TN), total phosphorus (TP), chemical oxygen demand (COD) and turbidity (TUB); wherein, total nitrogen and total phosphorus are determined by spectrophotometry, chemical oxygen demand is determined by permanganate index method, and turbidity is determined by nephelometer.
[0065] S3. Perform image stitching, geometric precision correction and improved 6S atmospheric radiation correction model preprocessing on the multispectral remote sensing images collected in S1 to generate orthophotos, and perform normalization. Finally, the water body in the study area is extracted using the improved watershed segmentation algorithm.
[0066] As a preferred embodiment of the present invention, in this step, image preprocessing includes the following specific steps:
[0067] 1) Image stitching and geometric precision correction:
[0068] The SIFT (Scale-Invariant Feature Transform) feature matching algorithm is used to extract the same-name points, the image stitching is performed based on the region growing algorithm, and the ground control points (GCPs) are used for geometric precision correction to ensure that the RMS error is less than 1 pixel.
[0069] 2) Improved 6S (Second Simulation of Satellite Signal in the Solar Spectrum) atmospheric radiation correction model preprocessing: Considering the atmospheric scattering and absorption effects, including Rayleigh scattering, aerosol scattering and ozone absorption, a spatiotemporal adaptive atmospheric parameter estimation method is introduced to improve the correction accuracy.
[0070] The calibration process includes:
[0071] a) Atmospheric parameter acquisition: obtain atmospheric transmittance, scattering reflectivity and other parameters through MODTRAN model simulation;
[0072] b) Surface reflectivity inversion: using iterative optimization algorithm to solve the radiation transfer equation;
[0073] c) Correction coefficient calculation: fast acquisition based on the lookup table (LUT) method;
[0074] d) Introduce BRDF (bidirectional reflectance distribution function) correction to eliminate the influence of solar altitude angle and observation angle 3) Normalize the DN value of each band, and the calculation formula is:
[0075] Spectral Response=(DN-DNmin) / (DNmax-DNmin)
[0076] Among them, DN represents the digital number value of the pixel; DNmin represents the minimum DN value of the band; DNmax represents the maximum DN value of the band; Spectral Response represents the normalized spectral response value, ranging from [0,1].
[0077] The normalization processing has the following technical advantages: it effectively reduces the influence of water body mirror reflection and splash interference, enhances the spectral differences caused by changes in water quality parameters, realizes the standardization of spectral response values, facilitates the comparison of data in different phases, and combines 3×3 window mean filtering to further suppress image noise and more accurately reflect the spectral differences caused by water quality changes.
[0078] 3) Improved watershed segmentation algorithm is used to extract water bodies in the study area: segmentation markers are constructed based on the morphological gradient operator, multi-scale watershed transform is used for initial segmentation, a regional merging strategy is introduced, segmentation results are optimized based on spectral similarity and spatial adjacency, and fragmented areas are eliminated through morphological post-processing to improve the accuracy of water body extraction.
[0079] S4. Based on the orthophotos obtained in S3, a multidimensional feature system including spectral features, geometric features and texture features was constructed, and the feature information corresponding to the sampling points was extracted using geographic information system (GIS) software and used as candidate feature variables.
[0080] As a preferred embodiment of the present invention, in this step, the multidimensional feature system specifically includes:
[0081] 1) Spectral characteristics:
[0082] a) First-order derivative feature: calculate the reflectivity difference between adjacent bands, characterize the changing trend of the spectral curve, and reduce the impact of background noise.
[0083] Finite difference first-order derivative calculation:
[0084] FD(λi)=[R(λi+1)-R(λi)] / [λi+1-λi]
[0085] Where: FD(λi) represents the first-order derivative value at wavelength λi, R(λi) represents the reflectivity at wavelength λi, and [λi+1-λi] represents the interval between adjacent wavelengths.
[0086] Continuous adjacent band combination:
[0087] SD(i)=[R(λi+1)-R(λi-1)] / 2Δλ
[0088] Where: SD(i) represents the first-order derivative value of the i-th band, R(λi+1) represents the reflectivity of the next band, R(λi-1) represents the reflectivity of the previous band, and Δλ represents the band interval.
[0089] Spectral slope characteristics:
[0090] Slope(i)=[R(λi+n)-R(λi)] / (n×Δλ)
[0091] Where n represents the number of interval bands; Δλ represents the basic band interval, which can characterize the spectral variation trend within a larger wavelength range.
[0092] The first-order derivative feature has the following advantages: eliminating the influence of background and baseline drift, enhancing the inflection point characteristics of the spectral curve, and improving the detection sensitivity of spectral changes.
[0093] b) Second-order derivative characteristics: Calculate the rate of change of the first-order derivative, enhance the subtle changes in the spectral curve, and improve the resolution of spectral features.
[0094] Finite difference second-order derivative computation:
[0095] SD(λi)=[R(λi+1)-2R(λi)+R(λi-1)] / (Δλ) 2
[0096] Wherein: SD(λi) represents the second-order derivative value at wavelength λi, R(λi+1), R(λi), R(λi-1) represent the reflectivity at three adjacent wavelengths, and Δλ represents the wavelength interval.
[0097] Second-order difference of continuous bands:
[0098] SSD(i)=[R(λi+2)-2R(λi)+R(λi-2)] / 4(Δλ) 2
[0099] Where: SSD(i) represents the second-order derivative value of the i-th band, R(λi+2), R(λi), and R(λi-2) represent the reflectivity values separated by two bands.
[0100] Spectral curvature characteristics:
[0101] Curvature(i)=|FD(λi+1)-FD(λi)| / Δλ
[0102] Wherein: FD(λi+1), FD(λi) represent the first-order derivative values at adjacent wavelengths, and Δλ represents the wavelength interval.
[0103] The second-order derivative feature has the following advantages: further eliminating linear background interference, enhancing subtle changes in the spectral curve, improving the resolution of overlapping absorption peaks, and improving the recognition accuracy of spectral features.
[0104] 2) Geometric features:
[0105] a) Form Factor:
[0106] Perimeter Area Ratio (PAR):
[0107] PAR=P / A
[0108] Where: P represents the perimeter of the water patch (m), A represents the area of the water patch (m 2 );
[0109] This index reflects the shape complexity of water patches, and the larger the value, the more complex the shape.
[0110] Circularity Index (CI):
[0111] CI=4πA / P 2
[0112] Where: A represents the area of the water patch; P represents the circumference of the water patch; π represents pi;
[0113] This index measures the similarity between the shape of a water patch and a circle, with a value range of (0,1]. The closer the value is to 1, the closer the shape is to a circle.
[0114] Complexity index (FD, Fractal Dimension):
[0115] FD=2ln(P) / ln(A)
[0116] Where: P represents the perimeter of the water patch; A represents the area of the water patch; ln represents the natural logarithm;
[0117] This index describes the complexity of the water patch boundary. The value range is [1,2]. The larger the value, the more complex the boundary.
[0118] b) Spatial characteristics:
[0119] Directional Index (DI):
[0120] DI=L / W
[0121] Where: L represents the length of the minimum enclosing rectangle; W represents the width of the minimum enclosing rectangle;
[0122] This index reflects the spatial extension direction characteristics of water patches.
[0123] Boundary Complexity (BC):
[0124]
[0125] Where: P represents the circumference of the water patch; A represents the area of the water patch; π represents pi;
[0126] This index characterizes the irregularity of water patch boundaries.
[0127] Spatial Aggregation Index (AI):
[0128] AI=(gii / max-gii)×100
[0129] Among them: gii represents the number of adjacent grids of the same type; max-gii represents the maximum possible number of adjacent grids of the same type;
[0130] This index reflects the degree of spatial aggregation of water patches, and its value range is [0,100].
[0131] 3) Texture features:
[0132] The following features are calculated based on the gray-level co-occurrence matrix (GLCM):
[0133] Energy: Energy = ∑∑(P(i,j)) 2 ; Where: P(i,j) represents the normalized GLCM matrix element, i,j represents the grayscale value.
[0134] Entropy: Entropy = -∑∑P(i,j)log(P(i,j)); reflects the randomness of the grayscale distribution of the image, P(i,j) represents the normalized GLCM matrix element, and i,j represents the grayscale value.
[0135] Contrast: Contrast = ∑∑(ij) 2 P(i,j); reflects the clarity of the image and the depth of the texture grooves. P(i,j) represents the normalized GLCM matrix element, and i,j represents the grayscale value.
[0136] Homogeneity: Homogeneity = ∑∑P(i,j) / (1+|ij|); reflects the degree of local change of the image, P(i,j) represents the normalized GLCM matrix element, i,j represents the grayscale value.
[0137] Correlation: Correlation = ∑∑((i-μj)(j-μj)P(i,j)) / (σiσj), where μj, μj represent the edge mean of GLCM; σiσj represents the edge standard deviation of GLCM; the calculation of these texture features is based on the following parameter settings:
[0138] Direction: 0°, 45°, 90°, 135°; Distance: 1-5 pixels; Quantization level: 64 levels; Window size: 7×7 or 9×9 pixels.
[0139] 4) Feature optimization strategy:
[0140] a) Noise suppression: Savitzky-Golay filtering is used for spectral smoothing;
[0141] R′(i)=∑[Cn×R(i+n)] / N
[0142] Where: R′(i) represents the smoothed reflectivity, Cn represents the convolution coefficient, and N represents the filter window size.
[0143] b) Feature enhancement: Continuous wavelet transform (CWT) is used for multi-scale analysis;
[0144] CWT(a,b)=∫[R(λ)ψ((λ-b) / a)]dλ
[0145] Where: a represents the scale parameter, b represents the translation parameter, ψ represents the wavelet mother function, and R(λ) represents the original spectral reflectance.
[0146] c) Feature normalization: using the minimum-maximum normalization method
[0147] R_norm=[R-R_min] / [R_max-R_min]
[0148] Among them: R_norm represents the normalized eigenvalue, R_min and R_max represent the minimum and maximum values of the feature.
[0149] S5. Apply the recursive feature elimination and cross validation (RFECV) algorithm to select the optimal feature subset from the candidate feature variables obtained in S4.
[0150] As a preferred embodiment of the present invention, the steps are as follows:
[0151] 1) Recursive Feature Elimination (RFE) process: Initialize the model using all features, calculate feature importance scores, remove the features with the lowest importance, retrain the model using the remaining features, and repeat the above process until the preset number of features is reached.
[0152] 2) Cross-validation (CV) strategy: k-fold cross-validation (k=5) is used to evaluate the performance of different feature subsets, the average performance index of each feature subset is calculated, and the feature combination with the best performance is selected.
[0153] This method has the following advantages: it can automatically determine the optimal number of features, reduce feature redundancy, improve model interpretability, and reduce the risk of overfitting.
[0154] S6. Using the stratified random sampling method, the water quality parameter data in S2 are divided into a training set and a validation set.
[0155] As a preferred embodiment of the present invention, in this step, the water quality parameter data is divided into a ratio of 70% training set and 30% validation set.
[0156] S7. Based on the optimal feature subset screened by S5 and the training set divided by S6, a two-stage integrated learning framework is constructed to obtain the estimation model of each water quality parameter.
[0157] As a preferred embodiment of the present invention, the steps are as follows:
[0158] 1) Construction of stage-based learner:
[0159] a) XGBoost (Extreme Gradient Boosting) base learner: It uses the second-order Taylor expansion to approximate the objective function, introduces a regularization term to control the model complexity, and uses feature pre-sorting to accelerate the training process.
[0160] Core parameter configuration:
[0161] learning_rate: set to 0.01-0.1;
[0162] max_depth: controls the depth of the tree, set to 3-6;
[0163] min_child_weight: minimum leaf node sample weight sum, set to 1-3;
[0164] subsample: sample sampling ratio, set to 0.8-1.0;
[0165] colsample_bytree: feature sampling ratio, set to 0.8-1.0;
[0166] XGBoost is an optimized distributed gradient boosting library. Its main advantage is that it uses a second-order Taylor expansion to approximate the loss function, which enables the model to better capture nonlinear relationships. It introduces regularization terms to control model complexity and effectively prevent overfitting. In terms of computational efficiency, XGBoost implements parallel computing and distributed training, and uses a "block structure" storage format to compress data, greatly improving training speed. The unique feature binning technology can handle sparse data, and the built-in cross-validation function makes model tuning more convenient. Especially when dealing with complex data such as large-scale water quality parameters, XGBoost shows excellent prediction accuracy and stability.
[0167] b) LightGBM base learner: It uses a histogram-based decision tree algorithm, uses a leaf-first growth strategy, and supports direct input of category features;
[0168] Core parameter configuration:
[0169] num_leaves: the number of leaf nodes, set to 31-127;
[0170] feature_fraction: feature sampling ratio, set to 0.8-1.0;
[0171] bagging_fraction: sample sampling ratio, set to 0.8-1.0;
[0172] bagging_freq: bagging frequency, set to 5-10;
[0173] lambda_l1: L1 regularization parameter, set to 0-0.1;
[0174] lambda_l2: L2 regularization parameter, set to 0-0.1;
[0175] LightGBM uses a histogram-based decision tree algorithm, and significantly improves training efficiency through histogram acceleration technology. Its biggest feature is the introduction of the unique GOSS (Gradient-based One-Side Sampling) algorithm and EFB (Exclusive Feature Bundling) feature bundling technology. These two innovations greatly reduce memory usage and computational complexity. When processing large-scale data, LightGBM adopts a leaf-wise growth strategy instead of the traditional level-wise growth. This strategy can more accurately find the split point and improve the model accuracy. For data with spatiotemporal characteristics such as water quality monitoring, LightGBM can efficiently process high-dimensional features and has significant advantages in memory usage and training speed.
[0176] 2) Second stage Stack strategy implementation:
[0177] a) Processing of prediction results of base learners: Use k-fold cross validation (k=5) to generate predictions for the training set, make overall predictions for the test set, and use the prediction results as the new feature matrix.
[0178] b) Meta-learner construction: Logistic regression is used as the meta-learner, and L1 / L2 regularization is used to control overfitting;
[0179] Optimized configuration:
[0180] Penalty: regularization method, optional 'l1' or 'l2';
[0181] C: the inverse of the regularization strength, ranging from 0.1 to 10;
[0182] solver: optimization algorithm selection, such as 'liblinear' or 'saga';
[0183] max_iter: maximum number of iterations, set to 1000;
[0184] The Stack strategy is an advanced ensemble learning method that uses the prediction results of multiple base learners as new features to train a meta-learner to obtain the final prediction. The advantage of this strategy is that it can fully utilize the advantages of different models and reduce the limitations of a single model. In water quality parameter estimation, XGBoost and LightGBM, as base learners, can capture different features and patterns of the data respectively, while logistic regression, as a meta-learner, can learn how to optimally combine these prediction results. This multi-level learning architecture not only improves the generalization ability of the model, but also can adaptively handle the feature differences of different water quality parameters, ultimately achieving more stable and accurate prediction results. The Stack strategy is particularly suitable for processing water quality monitoring data with complex feature interactions because it can automatically learn the importance weights of prediction results from different models.
[0185] S8. Use the validation set in S6 to evaluate the performance of the model constructed in S7 through indicators and select the optimal model.
[0186] As a preferred embodiment of the present invention, the steps are as follows:
[0187] 1) Coefficient of determination (R 2 )calculate:
[0188] R 2 =1-∑(y_true-y_pred)2 / ∑(y_true-y_mean) 2
[0189] Among them: y_true represents the measured value of water quality parameters; y_pred represents the model predicted value; y_mean represents the mean of the measured value; R 2 It reflects the degree to which the model explains the variability of the data. The value range is [0,1]. The closer it is to 1, the better the model fitting effect.
[0190] 2) Root mean square error (RMSE) calculation:
[0191]
[0192] Where: n represents the number of samples; RMSE reflects the average deviation between the predicted value and the measured value. The unit is the same as the original data. The smaller the value, the higher the prediction accuracy.
[0193] 3) Model performance evaluation process: Calculate the R of each model on an independent validation set 2 and RMSE, comprehensively considering the performance of the two indicators, and selecting the model with the best comprehensive performance for subsequent predictions.
[0194] S9. Based on the optimal model selected in S8, the spatial distribution of water quality parameters is inverted for the entire study area, and a high-resolution thematic map of water quality parameters is generated.
[0195] As a preferred embodiment of the present invention, the steps are as follows:
[0196] 1) Parameter inversion process: The images of the study area are processed according to the preprocessing scheme, the feature information corresponding to the optimal feature subset is extracted, and the trained model is used for prediction to generate high-resolution spatial distribution maps of four water quality parameters (TN, TP, COD, and TUB).
[0197] 2) Spatiotemporal variation analysis: Generate water quality parameter distribution maps for summer and winter respectively, calculate the spatiotemporal variation coefficients of each parameter, analyze the spatial heterogeneity of water quality parameters, and evaluate seasonal variation characteristics.
[0198] S10. A strategy combining spatial interpolation and machine learning models was adopted to optimize the inversion results of the spatial distribution of water quality parameters in S9 using Kriging interpolation to improve the prediction accuracy of the spatial distribution of water quality parameters.
[0199] As a preferred embodiment of the present invention, in this step, the calculation method of Kriging interpolation is as follows:
[0200] Z*(s0)=∑λiZ(si)
[0201] Constraint: ∑λi=1
[0202] Among them, Z*(s0) is the estimated value of the point to be predicted, Z(si) is the observed value of the i-th known sampling point, λi is the Kriging weight, si is the spatial position of the sampling point, and s0 is the spatial position of the point to be predicted.
[0203] As a preferred embodiment of the present invention, in this step, the spatial interpolation optimization strategy is specifically as follows:
[0204] 1) Kriging interpolation implementation:
[0205] a) Variogram analysis: Calculate the experimental variogram, fit the theoretical variogram model (spherical model, exponential model or Gaussian model), and determine the variogram parameters (sill value, partial sill value, range).
[0206] b) Weight determination: Construct the Kriging equations, solve the optimal weight coefficients, and consider spatial autocorrelation.
[0207] 2) Optimization of prediction results: Combining the prediction values of the machine learning model and the kriging interpolation results, a weighted fusion strategy is adopted to generate an optimized spatial distribution map of water quality parameters;
[0208] Evaluate the optimization effect, including: calculating the prediction accuracy before and after optimization, analyzing the continuity of spatial distribution, and evaluating the processing effect of local outliers.
[0209] The method and effects of the method of the present invention will be specifically described below through examples.
[0210] Example
[0211] This embodiment provides a rural wetland water quality monitoring method based on drone multispectral imaging and machine learning, and the specific steps are as follows:
[0212] S1, such as Figure 2 As shown in the figure, 196hm2 was selected in Xiangfudang Wetland Park, Jiashan County, Jiaxing City, Zhejiang Province 2 The typical rural wetland was selected as the study area (the center point is 120°55′20″E, 30°57′17″N, located in the Hangjiahu Plain, with a mild and humid climate, abundant rainfall, a typical subtropical monsoon climate, and the entire study area is flat). A stratified optimization sampling strategy was adopted for drone multispectral remote sensing image acquisition and in-situ water quality sampling. During the specific implementation, on June 15 (wet season) and November 26 (dry season) in 2023, the DJI Matrice300 RTK drone platform was equipped with a Changguang Yuchen AQ600 Pro multispectral camera for image acquisition. The camera includes five 3.2M pixel multispectral CMOS sensors and one 12.3M pixel panchromatic sensor to collect ground reflectance information in five narrow bands: blue light (450±20nm), green light (555±20nm), red light (660±20nm), red edge (720±20nm) and near infrared (840±20nm). The DUPLEX algorithm was used to arrange the sampling points. 31 sampling points were set in the study area in two periods to ensure that the samples were evenly distributed in the feature space.
[0213] S2. Laboratory analysis of the collected water samples: total nitrogen (TN) and total phosphorus (TP) were determined by spectrophotometry, and the concentration was determined by absorbance at a specific wavelength; chemical oxygen demand (COD) was determined by the permanganate index method, using potassium permanganate as an oxidant for digestion under neutral pH conditions; turbidity (TUB) was determined using a TU5200 desktop nephelometer to evaluate the turbidity of suspended and colloidal particles in the water. Figure 3As shown in the figure, the analysis results of water quality measurement data in two periods are shown. Through the distribution characteristics of the four groups of water quality indicators, it can be seen that: the TN (total nitrogen) data range is between 0-1, and TN_0615 and TN_1126 both show obvious bimodal distribution characteristics with similar medians; TP (total phosphorus) is distributed in the range of 0-0.4, and the value of TP_1126 is higher than TP_0615 as a whole, and the distribution is more concentrated; COD values are distributed between -10 and 20, COD_0615 is more dispersed, while COD_1126 shows a relatively concentrated distribution characteristic; TUB values range from 10-30, and the medians of the two periods are close, but the data distribution of TUB_1126 is more concentrated. Figure 3 It clearly shows the distribution characteristics and changing trends of various water quality indicators in different periods.
[0214] S3. Use Yusense Map software to perform image stitching, geometric precision correction and improved 6S atmospheric radiation correction model preprocessing on the collected multispectral images to generate orthophotos. Perform DN value normalization on the images, and the calculation formula is:
[0215] Spectral Response=(DN-DNmin) / (DNmax-DNmin)
[0216] Among them, DN represents the digital number of each pixel, and DNmin and DNmax represent the minimum and maximum DN values of the band respectively. Figure 4 As shown, finally, the water body in the study area was extracted by the improved watershed segmentation algorithm to obtain the complete water body area in the two periods.
[0217] S4. Based on the orthophotos obtained, a multidimensional feature system including spectral features (first-order derivatives, second-order derivatives), geometric features (shape factor, complexity index) and texture features (gray-level co-occurrence matrix) is constructed. The calculation formulas for the first-order derivative and second-order derivative of the spectrum are:
[0218] FD(λi)=[R(λi+1)-R(λi)] / Δλ
[0219] SD(λi)=[R(λi+2)-2R(λi+1)+R(λi)] / (Δλ) 2
[0220] Among them, FD(λi) represents the first-order derivative value at wavelength λi; R(λi) represents the reflectivity value at wavelength λi; R(λi+1) represents the reflectivity value at wavelength (λi+1); Δλ represents the interval between adjacent wavelengths; i represents the wavelength number; SD(λi) represents the second-order derivative value at wavelength λi; R(λi+2) represents the reflectivity value at wavelength (λi+2).
[0221] ARCGIS software was used to extract the characteristic information corresponding to the sampling points as candidate characteristic variables.
[0222] S5. Apply the recursive feature elimination and cross validation (RFECV) algorithm to select the optimal feature subset from the candidate feature variables. Use the random forest model to determine the feature importance and select the top six feature variables.
[0223] S6. Using stratified random sampling method, the water quality parameter data were divided into 70% training set and 30% validation set.
[0224] S7. Based on the screened optimal feature subset and the divided training data set, a two-stage integrated learning framework is constructed: first, XGBoost and LightGBM are used to build a base learner, and then the Stack strategy is used to use the prediction results of the base learner as new features, and logistic regression is used as a meta-learner to build a water quality parameter (TN, TP, COD, TUB) estimation model.
[0225] S8. Use an independent validation set to evaluate model performance by using the coefficient of determination (R 2 ) and root mean square error (RMSE) to select the optimal model. The evaluation index calculation formula is as follows:
[0226]
[0227] Among them, R 2 represents the coefficient of determination, ytrue,i represents the measured value of the i-th sample, and ypred,i represents the predicted value of the i-th sample. represents the average value of the measured value, i represents the wavelength number; RMSE represents the root mean square error, and n represents the total number of samples.
[0228] The results show that the model performs well in estimating various water quality parameters, and the validation set R 2 All reached above 0.72.
[0229] S9. Based on the selected optimal model, the spatial distribution of water quality parameters was inverted for the entire study area, and a high-resolution thematic map of water quality parameters was generated. The results showed that the average TN concentrations in summer (June 15) and winter (November 26) were 1.34 mg / L and 1.16 mg / L, TP concentrations were 0.16±0.06 mg / L and 0.14±0.07 mg / L, COD values were 16.05±9.87 mg / L and 13.02±8.22 mg / L, and TUB values were 18.39 NTU and 20.03 NTU, respectively.
[0230] S10. Adopt a strategy combining spatial interpolation with machine learning models, and use Kriging interpolation to optimize the prediction results of the machine learning model. The basic formula of Kriging interpolation is:
[0231] Z*(s0)=∑λiZ(si)
[0232] Constraint: ∑λi=1
[0233] Among them, Z*(s0) is the estimated value of the point to be predicted, Z(si) is the observed value of the i-th known sampling point; λi is the Kriging weight, which significantly improves the prediction accuracy of the spatial distribution of water quality parameters; si is the spatial position of the sampling point, and s0 is the spatial position of the point to be predicted. Through this method, efficient and accurate monitoring of rural wetland water quality is achieved, providing a scientific basis for wetland management and water environment protection.
[0234] The present invention uses the DUPLEX algorithm to optimize the sampling points and simultaneously collects UAV multispectral images and in-situ water samples. The collected images are geometrically corrected and improved 6S atmospheric radiation correction is performed, and the water body is extracted by the improved watershed segmentation algorithm. A multidimensional feature system is constructed and RFECV is used to screen the optimal features. A two-stage integrated learning framework is constructed with XGBoost and LightGBM as the base learners, and the spatial distribution map of water quality parameters is generated in combination with Kriging interpolation. The method of the present invention has the following advantages: (1) The optimized sampling and improved atmospheric correction model are used to improve data accuracy; (2) A multidimensional feature system is constructed and combined with recursive feature elimination to improve feature selection efficiency; (3) The two-stage integrated learning framework and spatial interpolation optimization are used to improve the accuracy of parameter estimation; (4) Compared with traditional methods, it has the advantages of high spatiotemporal resolution and full domain coverage.
[0235] The above-described embodiment is only a preferred solution of the present invention, but it is not intended to limit the present invention. A person skilled in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.
Claims
1. A rural wetland water quality monitoring method based on drone multispectral imaging and machine learning, characterized in that: The details are as follows: S1. Adopt a stratified optimization sampling strategy in the target rural wetland research area, collect UAV multispectral remote sensing images and conduct in-situ water quality sampling at each sampling point; S2, measuring water quality parameters of the water samples collected in S1; S3, image stitching, geometric precision correction and improved 6S atmospheric radiation correction model preprocessing are performed on the multispectral remote sensing images collected in S1 to generate orthophotos, and normalization is performed. Finally, the water body in the study area is extracted by the improved watershed segmentation algorithm; S4. Based on the orthophoto obtained in S3, a multidimensional feature system including spectral features, geometric features and texture features is constructed, and the feature information corresponding to the sampling points is extracted using geographic information system software and used as candidate feature variables; S5, applying recursive feature elimination and cross-validation algorithm to select the optimal feature subset from the candidate feature variables obtained in S4; S6, using stratified random sampling method, divide the water quality parameter data in S2 into training set and validation set; S7, based on the optimal feature subset screened by S5 and the training set divided by S6, a two-stage integrated learning framework is constructed to obtain the estimation model of each water quality parameter; S8, using the validation set in S6, evaluate the performance of the model constructed in S7 through indicators and select the optimal model; S9, based on the optimal model selected in S8, the spatial distribution of water quality parameters is inverted for the entire study area, and a high-resolution thematic map of water quality parameters is generated; S10. A strategy combining spatial interpolation and machine learning models was adopted to optimize the inversion results of the spatial distribution of water quality parameters in S9 using Kriging interpolation to improve the prediction accuracy of the spatial distribution of water quality parameters.
2. The rural wetland water quality monitoring method based on drone multispectral imaging and machine learning according to claim 1 is characterized in that: In S1, the sampling points are arranged using the DUPLEX algorithm to ensure that the samples are evenly distributed in the feature space; image acquisition is achieved using a DJI Matrice300 RTK drone platform equipped with a Changguang Yuchen AQ600 Pro multispectral camera, which includes five 3.2M pixel multispectral CMOS sensors and one 12.3M pixel panchromatic sensor for collecting ground reflectance information in five narrow bands: blue light, green light, red light, red edge, and near infrared.
3. The rural wetland water quality monitoring method based on drone multispectral imaging and machine learning according to claim 1 is characterized in that: In S2, the water quality parameters include total nitrogen, total phosphorus, chemical oxygen demand and turbidity; wherein, total nitrogen and total phosphorus are determined by spectrophotometry, chemical oxygen demand is determined by permanganate index method, and turbidity is determined by nephelometer.
4. The rural wetland water quality monitoring method based on drone multispectral imaging and machine learning according to claim 1 is characterized in that: In S3, the normalization processing method is as follows: Spectral Response=(DN-DNmin) / (DNmax-DNmin) Among them, DN represents the digital number value of the pixel; DNmin represents the minimum DN value of the band; DNmax represents the maximum DN value of the band; Spectral Response represents the normalized spectral response value, ranging from [0,1].
5. The rural wetland water quality monitoring method based on drone multispectral imaging and machine learning according to claim 1 is characterized in that: In S4, the spectral features include the first-order derivative and the second-order derivative of the spectrum, the geometric features include the shape factor and the complexity index, and the texture features include the gray-level co-occurrence matrix; The calculation method of the first-order derivative of the spectrum is as follows: FD(λi)=[R(λi+1)-R(λi)] / Δλ Wherein, FD(λi) represents the first-order derivative value at wavelength λi; R(λi) represents the reflectivity value at wavelength λi; R(λi+1) represents the reflectivity value at wavelength (λi+1); Δλ represents the interval between adjacent wavelengths; i represents the wavelength number; The calculation method of the second-order derivative of the spectrum is as follows: SD(λi)=[R(λi+2)-2R(λi+1)+R(λi)] / (Δλ) 2 Wherein, SD(λi) represents the second-order derivative value at wavelength λi; R(λi+2) represents the reflectivity value at wavelength (λi+2).
6. The rural wetland water quality monitoring method based on drone multispectral imaging and machine learning according to claim 1 is characterized in that: In S6, the water quality parameter data is divided into a ratio of 70% training set and 30% validation set.
7. The rural wetland water quality monitoring method based on drone multispectral imaging and machine learning according to claim 1 is characterized in that: The S7 is specifically as follows: Based on the optimal feature subset screened by S5 and the training set divided by S6, XGBoost and LightGBM were first used to build the base learner. Then, the Stack strategy was adopted to use the prediction results of the base learner as new features, and logistic regression was used as the meta-learner to build the estimation models of various water quality parameters.
8. The rural wetland water quality monitoring method based on drone multispectral imaging and machine learning according to claim 1 is characterized in that: In said S8, the indicators include the coefficient of determination and the root mean square error; The calculation method of the determination coefficient is as follows: Among them, R 2 represents the coefficient of determination, ytrue,i represents the measured value of the i-th sample, and ypred,i represents the predicted value of the i-th sample. represents the average value of the measured value, i represents the wavelength number; The calculation method of the root mean square error is as follows: RMSE=√[1 / n∑(ytrue,i-ypred,i) 2 ] Here, RMSE represents the root mean square error, and n represents the total number of samples.
9. The rural wetland water quality monitoring method based on drone multispectral imaging and machine learning according to claim 1 is characterized in that: In S10, the calculation method of Kriging interpolation is as follows: Z*(s0)=∑λiZ(si) Constraint: ∑λi=1 Among them, Z*(s0) is the estimated value of the point to be predicted, Z(si) is the observed value of the i-th known sampling point, λi is the Kriging weight, si is the spatial position of the sampling point, and s0 is the spatial position of the point to be predicted.
Citation Information
Cited By
Processing method for water quality detection data of new and old wetlands
CN120766813A