Methods, equipment, and procedures for near-sensory hyperspectral inversion of total phosphorus concentration in eutrophic lakes and reservoirs

CN122677005APending Publication Date: 2026-09-01NANJING INST OF GEOGRAPHY & LIMNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611171686.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-04
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

而低营养湖库总磷浓度普遍较低(通常在0.001-0.08 mg/L),叶绿素a浓度极低(通常<5 μg/L),上述光谱特征极为微弱,易被悬浮物、溶解有机物等干扰信号淹没,反演难度进一步增加

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122677005A_ABST
    Figure CN122677005A_ABST
Patent Text Reader

Abstract

This invention discloses a near-sensory hyperspectral inversion method, equipment, and program for total phosphorus concentration in eutrophic lakes and reservoirs. Based on hyperspectral data of the target lake or reservoir, a feature candidate set is constructed, including single-band, multi-band physical indices, and automatic ratio indices. The multi-band physical indices are spectral indices correlated with water depth, algal uptake, and vegetation activity. The automatic ratio indices are generated by selecting multiple first wavelengths within a preset band range at a preset step size, generating multiple second wavelengths starting from each first wavelength, and then combining the first and second wavelengths pairwise to generate several normalized difference indices. The feature candidate set is repeatedly screened using linear regression models, correlation analysis, and cross-correlation analysis to obtain the final feature subset. A total phosphorus concentration inversion model is trained and constructed based on machine learning methods for use in total phosphorus inversion in eutrophic lakes and reservoirs. This invention eliminates systematic spectral shifts caused by external factors such as sunlight and atmosphere, improving the stability of interannual predictions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of water quality and water environment monitoring technology, specifically relating to a near-sensory hyperspectral inversion method, equipment and program for total phosphorus concentration in low-trophic lakes and reservoirs. Background Technology

[0002] Hypotrophic lakes and reservoirs refer to lakes or reservoirs with low nitrogen and phosphorus concentrations that periodically limit the proliferation of primary producers such as phytoplankton. This includes lakes and reservoirs in oligotrophic and partially mesotrophic states. Hypotrophic lakes and reservoirs generally have excellent water quality and often serve as important drinking water sources, thus their water quality monitoring requirements are far higher than those for eutrophic waters. Total phosphorus, as a key factor limiting lake and reservoir productivity, directly reflects the trophic status and algal proliferation potential of the water body. For hypotrophic lakes and reservoirs, the background value of total phosphorus is low; even a small increase can disrupt the nutrient balance and lead to water quality decline. Therefore, total phosphorus monitoring is of great significance for ensuring drinking water safety and aquatic ecological health. However, traditional total phosphorus monitoring methods mainly employ chemical digestion-spectrophotometry, requiring on-site sampling and laboratory analysis. This process is complex, time-consuming, and costly. Satellite and UAV remote sensing methods are limited by spatial and temporal resolution and weather conditions, making it difficult to meet the high-frequency, high-precision online monitoring needs of hypotrophic lakes and reservoirs.

[0003] Near-sensing hyperspectral inversion technology acquires reflectance data in real time from tens to hundreds of consecutive narrow bands (spectral resolution up to 1 nm) by taking close-up images of the target water surface using hyperspectral sensors. This allows for the capture of subtle spectral differences among different components in the water (such as algae, suspended solids, and dissolved organic matter), providing a new approach for rapid monitoring of water quality parameters. However, existing technologies have significant limitations in applications to hypotrophic lakes and reservoirs. Total phosphorus, as a non-optically active parameter, is difficult to invert due to the lack of direct optical characteristics. Existing total phosphorus inversion techniques are typically developed for eutrophic water bodies, which usually have high total phosphorus concentrations (up to 0.1-1.3 mg / L), large algal biomass, and high chlorophyll a concentrations (up to tens to hundreds of μg / L). These eutrophic water bodies exhibit significant spectral characteristics, typically showing an absorption valley around 675 nm and a fluorescence reflectance peak at 700-710 nm. In eutrophic lakes and reservoirs, total phosphorus concentrations are generally low (typically 0.001-0.08 mg / L), and chlorophyll a concentrations are extremely low (typically <5 μg / L). These spectral characteristics are very weak and easily obscured by interference signals from suspended solids, dissolved organic matter, and other sources, further increasing the difficulty of feature extraction. Regarding feature construction, directly using raw reflectance or relying solely on manual experience to construct input features fails to fully uncover the weak signals associated with low total phosphorus concentrations. Feature selection often employs a single strategy, making it difficult to balance relevance and low redundancy; model validation often uses random partitioning, which cannot accurately reflect cross-year generalization ability and easily overestimates the model's actual application performance. Summary of the Invention

[0004] This application aims to provide a near-sensory hyperspectral inversion method, equipment, and program for total phosphorus concentration in eutrophic lakes and reservoirs. It constructs a hyperspectral data feature library based on the characteristics of low phosphorus concentration in eutrophic lakes and reservoirs, automatically selects feature subsets that are strongly correlated with total phosphorus concentration and have low redundancy from the feature library, and establishes a total phosphorus inversion model with high accuracy, good stability, and strong interpretability verified by an independent test set. It can also realistically evaluate the generalization ability of the model in future predictions.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] A near-sensory hyperspectral inversion method for total phosphorus concentration in eutrophic lakes and reservoirs, the method comprising:

[0007] Acquire near-sensing hyperspectral data and total phosphorus monitoring data of the target lakes and reservoirs, and construct a spatiotemporal matching dataset of hyperspectral and total phosphorus concentrations.

[0008] Based on the hyperspectral data, a feature candidate set is constructed, which includes single-band, multi-band physical indices and automatic ratio indices; wherein, the single band refers to multiple single-band reflectance.

[0009] The multi-band physical index is a spectral index composed of reflectance in different bands, including at least the depth correction index, red edge slope, and red edge ratio index; the depth correction index is the ratio of blue light band reflectance to red light band reflectance.

[0010] The automatic ratio index is as follows: within a preset band range, multiple first wavelengths are selected according to a first preset step size, and with each first wavelength as the starting point, multiple second wavelengths are generated within the preset band range according to a preset difference. The first wavelengths and second wavelengths are combined in pairs to generate several normalized difference indices.

[0011] A linear regression model is used to initially screen each feature in the candidate feature set to obtain an initial feature subset.

[0012] Calculate the correlation between each feature in the initial feature subset and the total phosphorus monitoring data in the spatiotemporal matching dataset, and remove features whose absolute correlation coefficient is lower than the first threshold; perform cross-correlation analysis on the remaining features, and remove features whose absolute correlation coefficient is higher than the second threshold; obtain the final feature subset;

[0013] A matching dataset is constructed using the final feature subset and total phosphorus monitoring data as a training set. A total phosphorus concentration inversion model is trained and constructed based on machine learning methods for the inversion of total phosphorus in low-trophic lakes and reservoirs.

[0014] In some embodiments of the present invention, the single-band feature is selected based on the following method:

[0015] For near-sensory hyperspectral data of the target lake / reservoir, select the spectral range where the reflection peak is located;

[0016] Within the band range where the reflection peak is located, several single-band reflectances are extracted according to a second preset step size.

[0017] In some embodiments of the present invention, the multi-band physical index features also include phycocyanin index, normalized vegetation index, chlorophyll index, reflectance peak height, absorption valley depth, and maximum rate of change of reflectance peak for different combinations of bands.

[0018] In some embodiments of the present invention, the preset band range is a continuous range consisting of green light, yellow light, red light and red edge bands.

[0019] In some embodiments of the present invention, the matching dataset is matched based on the temporal resolution of total phosphorus data, including:

[0020] If the total phosphorus data is sampled at regular intervals, the average value of the hyperspectral data within t minutes before and after the total phosphorus data collection time will be used to match it. Hyperspectral data outside the time window will not be included in the matching. t is a preset window threshold.

[0021] If the total phosphorus data is diurnal, it is matched with the arithmetic mean of the hyperspectral data during a preset diurnal period; the preset diurnal period is determined based on the latitude and longitude of the target lake / reservoir and the sampling season.

[0022] The hyperspectral data is used to remove bands where the system error exceeds a preset error value.

[0023] In some embodiments of the present invention, the hyperspectral-total phosphorus concentration spatiotemporal matching dataset is preprocessed and used for feature calculation;

[0024] The preprocessing includes: calculating the median of hyperspectral reflectance within a sliding window for the hyperspectral time series data in the matching dataset; then subtracting the median from the original hyperspectral reflectance value to obtain the residual sequence; calculating the median absolute deviation (MAD) of the residual sequence; identifying points exceeding a preset deviation range as outliers; and performing linear interpolation replacement using the two nearest non-outliers before and after the outlier.

[0025] In some embodiments of the present invention, the machine learning model is selected as a tree ensemble model.

[0026] In some embodiments of the present invention, each feature in the candidate feature set is initially screened using the LASSO regression method.

[0027] The present invention further provides a computer device including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0028] The present invention further provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the steps of the above-described method.

[0029] The beneficial effects of the present invention include at least the following:

[0030] (1) In view of the characteristics of low total phosphorus concentration and weak spectral signal in low-trophic lakes and reservoirs, a combination of single / multi-band physical indices and automatic ratio indices is adopted as the core inversion features. The candidate features are further screened by combining regression models, correlation and cross-correlation methods to obtain the parameter combination for model training. The introduced automatic ratio index eliminates the systematic spectral shift caused by external factors such as light and atmosphere. Combined with the screening of physical indices, the final model has good prediction accuracy and stability, which can meet the application requirements of online monitoring of total phosphorus in low-trophic lakes and reservoirs.

[0031] (2) The entire process from spectral input to total phosphorus output can be automated without human intervention, making it suitable for embedding in online monitoring systems of low-nutrient lakes and reservoirs. The results show that the prediction accuracy of this invention on independent year test data meets the requirements of online monitoring and has good prospects for promotion. Attached Figure Description

[0032] The accompanying drawings are not intended to be drawn to scale. In the drawings, every identical or nearly identical component shown in each figure can be denoted by the same reference numeral. For clarity, not every component is labeled in each figure. Wherein:

[0033] Figure 1 This is a flowchart illustrating a method for retrieving total phosphorus concentration in eutrophic lakes and reservoirs based on near-sensory hyperspectral feature fusion and intelligent screening, provided in an embodiment of this application.

[0034] Figure 2 This is a hyperspectral image of lake A in an embodiment of the present invention.

[0035] Figure 3 This is a ranking diagram of the feature importance of the RF model in an embodiment of the present invention.

[0036] Figure 4 This is a ranking diagram of the feature importance of the XGBoost model in an embodiment of the present invention.

[0037] Figure 5 This is a comparison chart of the predicted and measured values ​​of total phosphorus concentration for an independent test set using the RF model and XGBoost model in an embodiment of the present invention.

[0038] In the aforementioned Figures 1-5, all the symbols or other representations are well known in the art and will not be described again in this example. Detailed Implementation

[0039] To make the objectives, technical solutions, and beneficial effects of the embodiments of this application clearer, the technical solutions in the embodiments of this application are clearly and completely described below with reference to the accompanying drawings and a practical example of a reservoir. It should be understood that the described embodiments are only some embodiments of this application, and not all embodiments. The following embodiments are only used to illustrate this application, and not to limit the scope of protection of this application.

[0040] Example 1

[0041] This application provides a near-sensory hyperspectral inversion method for total phosphorus concentration in eutrophic lakes and reservoirs. Figure 1 As shown in the flowchart, the method includes at least the following steps:

[0042] S1. Acquire near-sensing hyperspectral data and simultaneous total phosphorus monitoring data of the target lake / reservoir.

[0043] The near-sensing hyperspectral data refers to the real-time high-frequency hyperspectral reflectance data of the water body collected by a near-sensing hyperspectral sensor that has been installed on the surface of the target lake / reservoir; the total phosphorus monitoring data refers to the total phosphorus concentration data collected by traditional monitoring methods or online water quality monitoring instruments.

[0044] The near-sensing hyperspectral data and total phosphorus data in step S1 should be collected at the same sampling point. Before data collection, ensure that the instrument is correctly installed and calibrated according to the equipment installation requirements. Manual sampling and sample determination should be performed in accordance with relevant specifications to ensure data quality.

[0045] In some embodiments of the present invention, the spectral range of the near-sensing hyperspectral data is 400-1000 nm, the spectral resolution is 1 nm, the acquisition time covers the daytime period from 9:00 to 16:00, and the acquisition frequency is not less than 10 minutes / time; the acquisition frequency of the total phosphorus monitoring data is not less than 1 day / time.

[0046] S2. Match the hyperspectral data and total phosphorus data mentioned in S1 by time to form a matching dataset in which the hyperspectral data and total phosphorus data correspond one-to-one, and perform preprocessing such as smoothing and noise reduction and outlier removal on the data.

[0047] In some embodiments of the present invention, during data matching, one of the following two methods is used for matching based on the actual time resolution of the total phosphorus data:

[0048] Method 1: If the total phosphorus data is high-resolution timed sampling data, then the average value of the hyperspectral data within 5 minutes before and after the total phosphorus data acquisition time shall be used to match it. Hyperspectral data outside the time window shall not be included in the matching.

[0049] Method 2: If the total phosphorus data is diurnal, it is matched with the arithmetic mean of the hyperspectral data during the preset daytime period.

[0050] The preset daytime period here is determined based on the latitude and longitude of the lake and reservoir and the sampling season. Since the hyperspectral instrument for collecting near-sensory hyperspectral data has high light requirements, the period with the strongest daytime sunshine and stable lighting conditions is selected as the matching period.

[0051] Furthermore, before data matching, the 400-419nm and 831-1000nm bands with large systematic errors were removed from the collected hyperspectral data, retaining the 420-830nm band. The time span of the matching dataset was no less than one year, and the number of samples was no less than 300, to ensure that the data could reflect seasonal, hydrological, and aquatic environment differences.

[0052] Furthermore, the preprocessing includes: smoothing and denoising the matching dataset using a Savitzky-Golay filter. In this embodiment, the reference window width is set to 11 and the polynomial order is 3.

[0053] A sliding window-based local anomaly detection method is used to remove outliers in order to maintain the continuity and trend of the time series. Specifically, the median of hyperspectral reflectance within the sliding window is calculated for the time series. Then, the median is subtracted from the original value to obtain the residual series. The median absolute deviation (MAD) of the residual series is calculated as median(|residual - median(residual)|). Points exceeding ±2.5 times MAD are identified as outliers. Linear interpolation is performed using the two nearest non-outlier points before and after the outlier point.

[0054] S3. Based on the spectral characteristics of low-nutrient lakes and reservoirs, construct a candidate set of input features from the hyperspectral data in the matching dataset described in S2.

[0055] The input feature candidate set constructed in this application includes single-band (physical index), multi-band physical index, and automatic ratio index, and their construction methods are as follows:

[0056] Single band (physical index): Select the band range where the reflection peak is located for near-sensor hyperspectral data of the target lake / reservoir;

[0057] Within the band range where the reflection peak is located, several single-band reflectances are extracted according to a preset step size.

[0058] Multi-band physical index: a spectral index that is related to water depth, algal absorption, and vegetation activity.

[0059] Automatic ratio index: Within a preset band range, select multiple first wavelengths at a preset step size, and with each first wavelength as the starting point, generate multiple second wavelengths within the preset band range according to a preset difference. Combine the first wavelengths and the second wavelengths in pairs to generate several normalized difference indices.

[0060] The preset band range should cover bands closely related to algal pigment absorption, scattering, and water quality parameter inversion.

[0061] S4. Initial selection of the candidate feature set described in S3 is performed using a linear regression model to obtain an initial feature subset;

[0062] Calculate the correlation coefficient between each feature in the initial feature subset and the total phosphorus concentration, and remove features whose absolute value of the correlation coefficient is lower than the threshold T1;

[0063] Cross-correlation analysis is performed on the remaining features to remove redundant features whose absolute correlation coefficients are higher than the threshold T2, thus obtaining the final feature subset.

[0064] S5. Match the final feature subset described in S4 with the total phosphorus data, select multiple machine learning models for training and multi-dimensional performance evaluation, and give a model recommendation strategy based on the comprehensive evaluation results and different application requirements.

[0065] S6. The optimal model obtained is used for total phosphorus monitoring in the corresponding low-trophic reservoir; after obtaining new spectral samples, features corresponding to the final feature subset are extracted, and the recommended model obtained from S5 evaluation is called to predict total phosphorus concentration.

[0066] Example 2

[0067] This embodiment uses a hypotrophic reservoir A as an example to illustrate the solution and application effects of the present invention.

[0068] The hypotrophic reservoir A is equipped with both a near-sensing hyperspectral sensor and an online water quality monitor at its national monitoring section. Hyperspectral reflectance data and total phosphorus monitoring data (concentration range 0.005 ~ 0.296 mg / L) were collected from June 1, 2024 to December 31, 2025. The near-sensing hyperspectral data had a spectral range of 400-1000 nm and a spectral resolution of 1 nm. Data was collected during the daytime period from 9:00 AM to 4:00 PM at a frequency of 2 minutes per data point. Total phosphorus data was collected once daily.

[0069] Since the total phosphorus data did not record the precise sampling time for each day, the arithmetic mean of the hyperspectral reflectance data in the 420-830 nm band from 12:00 to 14:00 each day was calculated and matched with the total phosphorus data. After removing samples with missing measurements, a total of 489 valid matching samples were obtained.

[0070] The measured spectral data is shown in the figure. Figure 2 As shown, for ease of display, Figure 2 Only spectral data from a specific moment was extracted for illustration; the waveforms of spectra from other moments show largely consistent patterns. The reflection peak is located between 550-580 nm. Other bands lack distinct peaks and troughs, indicating low levels of suspended solids and chlorophyll, suggesting clear water, consistent with the characteristics of eutrophic water bodies. To capture subtle spectral features, a strategy combining single-band data, physical indices, and automatic ratio indices was employed to construct a candidate set of band input features.

[0071] Among them, the single-band spectral reflectance Rλ i Focusing on the known spectral reflectance peak range, with a step size of 10 nm, i.e.:

[0072] λ i = 550 + 10(i-1), i=1,…,4;

[0073] Based on the band range and wavelength, the reflectance of a single band in this embodiment is the reflectance data of four green light bands: R550, R560, R570, and R580.

[0074] Multi-band physical indices mainly reflect the indirect correlation between water depth correction, algal uptake, vegetation activity and total phosphorus, providing physical interpretability and ensuring basic correlation (Table 1).

[0075] Table 1. Detailed Explanation of the Characteristic Bands of Physical Indices

[0076]

[0077] The automatic ratio index is calculated using wavelengths within the 500nm-750nm band. This band covers detailed spectral information across the entire spectrum, including green (500-580nm), yellow (580-620nm), red (620-700nm), and red-edge (700-750nm). These regions are closely related to algal pigment absorption and scattering, as well as water quality parameter inversion, and have a clear physical basis. The automatic ratio index is calculated using the following normalized difference formula:

[0078]

[0079] in Starting at 500nm and using 5nm as the step size; Starting with 501nm and using 5nm as the step size (or you can obtain...) Then, calculate all +1nm added (Dataset). A total of 1275 automatic ratio indices were generated using this method for subsequent filtering.

[0080] The input feature candidate set is standardized to eliminate scale differences between features of different dimensions. The standardized candidate feature set is then initially selected using a LASSO regression model (i.e., L1 regularized regression model). LASSO regression automatically selects features by adding an L1 norm penalty term to the loss function, compressing the regression coefficients of unimportant features to zero.

[0081] The objective function to be optimized is:

[0082] In the formula, the first term is the least squares loss function, which measures the goodness of fit between the model's predicted values ​​and the measured values; the second term is the L1 regularization term, which penalizes the regression coefficients of features, causing the coefficients of unimportant features to be compressed to zero. n is the number of samples, and p is the number of input features. Let be the total phosphorus concentration of the i-th sample. Let i be the spectral feature vector of the i-th sample. For the regression coefficient vector, This is the regularization parameter.

[0083] Regularization parameters The value of is determined using cross-validation: Within the candidate range (10) -4 -10), take 50 values ​​evenly on a logarithmic scale, and calculate each value using 10-fold cross-validation. The model prediction error under the given value is selected to minimize the cross-validation error. The value is used as the optimal regularization parameter. In the optimal... Under the given value, features with non-zero regression coefficients are the candidate features retained in the initial selection. In this embodiment, through LASSO initial selection, a large number of irrelevant or redundant features were automatically compressed from 1289 candidate features, and 25 features with a certain correlation with total phosphorus concentration were initially selected, forming an initial feature subset.

[0084] Calculate the correlation coefficient between each feature in the initial feature subset and the total phosphorus concentration, and remove features with an absolute correlation coefficient lower than T1 = 0.1; then perform cross-correlation analysis on the remaining features, and remove redundant features with an absolute correlation coefficient higher than T2 = 0.9. The logic for determining the correlation coefficient filtering threshold T1 and the cross-correlation redundancy removal threshold T2 is as follows:

[0085] T1 is used to remove features with a very weak linear relationship to total phosphorus concentration. A larger value results in stricter screening, removing more features and potentially losing some features that are weakly correlated with total phosphorus concentration but still have potential value. A smaller value results in looser screening, retaining more features, but may introduce noisy features almost unrelated to total phosphorus concentration, interfering with model training. By adjusting different T1 values ​​and comparing the model accuracy under different T1 values, it was ultimately determined in this embodiment that T1=0.1 can remove most irrelevant features while retaining effective information.

[0086] T2 is used to remove highly correlated redundant features from the candidate feature set. A larger value retains more redundant features, leading to stronger collinearity among features, which may cause model overfitting and decreased interpretability. A smaller value removes more redundant features, resulting in a more concise feature set, but may over-remove complementary feature combinations, losing valuable information. In this embodiment, T2=0.9 achieves feature simplification without losing valuable information.

[0087] After correlation coefficient filtering and cross-correlation redundancy removal, 16 features were finally obtained from the 25 initial features. The final feature subset and the correlation coefficient with total phosphorus are shown in Table 2.

[0088] Table 2 Final Feature Subset Information Table

[0089]

[0090] The final feature subset, along with synchronously matched total phosphorus concentration data, forms the modeling dataset used for training the Random Forest (RF) and XGBoost models. The input features are the 16 final selected features, and the output is the predicted total phosphorus concentration.

[0091] The RF model reduces variance and improves generalization ability by integrating multiple decision trees and averaging their predictions. The XGBoost model, on the other hand, effectively fits complex nonlinear relationships by training weak learners round-by-round using a gradient boosting framework and then weighting and fusing them. In this embodiment, the key parameter settings for both models are shown in Table 3.

[0092] Table 3 Key Parameter Settings for the Model

[0093]

[0094] The accuracy and stability of the model were evaluated using 10-fold cross-validation: the modeling dataset was randomly divided into 10 equal parts, and 9 parts were used for training and 1 part for validation in each iteration, repeated 10 times. The validation determination coefficient R was then calculated. 2 The mean and standard deviation, and the mean root mean square error (RMSE). The specific calculation formula is:

[0095] Coefficient of determination ;

[0096] Coefficient of determination mean ;

[0097] Coefficient of determination and standard deviation

[0098] Root mean square error

[0099] root mean square error

[0100] Where n is the number of training samples; y i The measured total phosphorus concentration for the i-th sample; The predicted total phosphorus concentration for the i-th sample; This is the average of all measured total phosphorus values; Let be the coefficient of determination for the k-th fold verification.

[0101] R 2 A higher mean indicates a better overall model fit; a smaller standard deviation indicates more stable model performance across different data subsets; and a smaller mean RMSE indicates a smaller deviation between predicted and measured values. After training, the R² values ​​of the RF model and the XGBoost model were cross-validated using a 10-fold cross-validation. 2 The mean values ​​were 0.9093 and 0.9134, respectively, the standard deviations were 0.0321 and 0.0467, and the mean RMSE values ​​were 0.0110 and 0.0103, respectively. It is evident that both models achieved excellent fitting results during training, with overall performance at the same level. XGBoost has a slightly better average accuracy, while Random Forest (RF) exhibits stronger stability.

[0102] The interpretability of the feature importance ranking model is analyzed. Figure 3 This is a ranking graph of feature importance for the RF model. For the RF model, feature importance is ranked based on the cumulative contribution of each feature to reducing impurity in all decision tree splits. Figure 4 This is a ranking graph of feature importance for the XGBoost model. For the XGBoost model, feature importance is ranked based on the frequency with which each feature is selected as a split point during the boosting process.

[0103] By comparison, it can be seen that although the feature importance ranking of the two models differs in details, the top 6 features are completely consistent, and their cumulative importance accounts for more than 80%. This indicates that the relationship between these captured features and total phosphorus is robust and independent of a specific model, and can be identified as the main features for total phosphorus inversion in the hypotrophic reservoir in this embodiment. The remaining 10 features serve as auxiliary features, playing a supplementary and corrective role in the model. The physical meaning of the top 6 features and their correlation logic with total phosphorus are shown in Table 4.

[0104] Table 4. List of the top 6 features by feature importance in RF and XGBoost models.

[0105]

[0106] Furthermore, the feature importance distribution in RF is more even (the top 3 features are between 0.16 and 0.17), indicating that each feature contributes to the voting across multiple trees, making the model more robust. XGBoost, on the other hand, is highly dependent on ND_570_666, demonstrating that this feature indeed carries the strongest total phosphorus signal. These two findings corroborate each other, enhancing the credibility of the feature selection results.

[0107] Based on the actual training results of this embodiment, the RF model shows excellent stability in 10-fold cross-validation, and the feature importance ranking is in good agreement with the physical mechanism. It is recommended as the preferred model for business deployment. The XGBoost model has a strong fitting ability within the training set and can be used as an auxiliary comparison model for reference.

[0108] In this embodiment, after the model is deployed, 121 near-sensory hyperspectral reflectance data and synchronously measured total phosphorus concentration data (concentration range 0.011 ~ 0.068 mg / L) of the national control section A of the low-trophic reservoir from January 1 to May 18, 2026 are collected as an independent test set (without overlap with the training data) to test the predictive performance of the model in the low concentration range and time extrapolation scenario.

[0109] The test samples were processed using the same preprocessing procedure as the training set: Savitzky-Golay filters were used for smoothing and noise reduction (window width 11, polynomial order 3), bands with large systematic errors at both ends of the spectral range were removed, the effective band of 420-830nm was retained, and the spectral data were Z-score standardized using the standardized parameters (mean and standard deviation) saved in the training set.

[0110] Sixteen feature values ​​(as shown in Table 2) corresponding to the final screening features were generated from the preprocessed test spectral data to form a test feature matrix. These matrix values ​​were then substituted into the trained RF model and XGBoost model to output the predicted total phosphorus concentration of the test samples, which were then compared with the measured total phosphorus data from the same period. Figure 5 The predicted and measured values ​​of total phosphorus concentration from the two models are compared.

[0111] The same evaluation metrics as the training set were used to evaluate the test results, as shown in Table 5. The test results show that the RF model outperforms the XGBoost model in all accuracy metrics on the independent year test data, with R... 2The concentration was 0.1451 mg / L higher, and the RMSE was 0.0021 mg / L lower. This further verifies that the RF model outperforms the XGBoost model in the overall performance of year-end prediction scenarios.

[0112] Table 5 Comparison of test evaluation results between RF model and XGBoost model

[0113]

[0114] The prediction accuracy of the RF model on the independent test set (R 2 =0.6817) is lower than the accuracy of 10-fold cross-validation on the training set (R² = 0.6817). 2 The mean is 0.9093, and the difference between the two is within a reasonable range. This is mainly because the training and test sets are from different years, resulting in interannual differences that cause a certain shift in the relationship between the spectrum and total phosphorus; the total phosphorus concentration in the test set (0.0110~0.0680 mg / L) is concentrated in the low-value range, making it more difficult for the model to fit the low-concentration range than the full range. However, the model performs well in tests conducted in independent years. 2 The result of 0.68 indicates that the relationship between the spectrum and total phosphorus learned by the model is real and has a certain degree of stability. It also suggests that the accuracy can be improved by accumulating more sample data and continuously updating the model.

[0115] In this embodiment, the percentage of samples with mean absolute percentage error (MAPE) and relative error within 30% is used as the basis for the calculation. 30 The two indicators are used to evaluate the monitoring accuracy of total phosphorus, with the standards being MAPE ≤ 40% and P30 ≥ 60%, respectively. The specific calculation formula is as follows:

[0116]

[0117]

[0118] Where n is the number of samples in the test set; y i The measured total phosphorus concentration for the i-th sample; The predicted total phosphorus concentration for the i-th sample; This is an indicator function that takes the value 1 if the condition is true, and 0 otherwise.

[0119] Calculations show that MAPE = 24.23%, far below the acceptable threshold of 40%. 30 = 67.8%, exceeding the 60% passing mark.

[0120] Both of the above indicators meet the monitoring accuracy requirements, indicating that the prediction error of the RF model on independent year test data is controllable and can meet the application requirement of 60% accuracy for online monitoring of total phosphorus in low-trophic lakes and reservoirs.

[0121] In other embodiments, the present invention tested the impact of different input parameters on the monitoring results.

[0122] Currently, near-sensing hyperspectral imaging is rarely used in total phosphorus monitoring. Research on total phosphorus monitoring primarily focuses on satellite remote sensing, with a small amount of research using UAV remote sensing. Near-sensing hyperspectral imaging is a full-spectrum imaging technique with numerous single-band applications. Existing near-sensing hyperspectral imaging techniques for total phosphorus monitoring select single bands from the full spectrum as inputs using a fixed wavelength step size to achieve inversion. This embodiment uses existing near-sensing hyperspectral imaging input feature selection methods and different feature library settings as comparisons to analyze the impact of different methods on monitoring accuracy. The results are shown in Table 6 below (all using RF models):

[0123] Table 6. Accuracy monitoring results for different feature library settings

[0124]

[0125] It can be seen that some parameter combinations achieve high accuracy on the training set, but completely fail on the independent test set, especially when using single bands as input parameters. Whether using eight bands selected from the 420-830nm single band, or selecting single bands at 10nm intervals within the 420-830nm range as input, good results are achieved on the training set, but they collectively fail on the independent test set. Many current studies divide contemporaneous data into training and test sets, achieving similar results on similar data, but the monitoring performance is poor in practical applications. The method of this invention meets the monitoring accuracy requirements on the independent validation set and conforms to application standards.

[0126] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A near-sensory hyperspectral inversion method for total phosphorus concentration in eutrophic lakes and reservoirs, characterized in that, The method includes: Acquire near-sensing hyperspectral data and total phosphorus monitoring data of the target lakes and reservoirs, and construct a spatiotemporal matching dataset of hyperspectral and total phosphorus concentrations. Based on the hyperspectral data, a feature candidate set is constructed, which includes single-band, multi-band physical indices and automatic ratio indices; wherein, the single band refers to multiple single-band reflectance. The multi-band physical index is a spectral index composed of reflectance in different bands, including at least the depth correction index, red edge slope, and red edge ratio index; the depth correction index is the ratio of blue light band reflectance to red light band reflectance. The automatic ratio index is as follows: within a preset band range, multiple first wavelengths are selected according to a first preset step size, and with each first wavelength as the starting point, multiple second wavelengths are generated within the preset band range according to a preset difference. The first wavelengths and second wavelengths are combined in pairs to generate several normalized difference indices. A linear regression model is used to initially screen each feature in the candidate feature set to obtain an initial feature subset. Calculate the correlation between each feature in the initial feature subset and the total phosphorus monitoring data in the spatiotemporal matching dataset, and remove features whose absolute correlation coefficient is lower than the first threshold; perform cross-correlation analysis on the remaining features, and remove features whose absolute correlation coefficient is higher than the second threshold; obtain the final feature subset; A matching dataset is constructed using the final feature subset and total phosphorus monitoring data as a training set. A total phosphorus concentration inversion model is trained and constructed based on machine learning methods for the inversion of total phosphorus in low-trophic lakes and reservoirs.

2. The method according to claim 1, characterized in that, The single-band feature is selected based on the following method: For near-sensory hyperspectral data of the target lake / reservoir, select the spectral range where the reflection peak is located; Within the band range where the reflection peak is located, several single-band reflectances are extracted according to a second preset step size.

3. The method according to claim 1, characterized in that, The multi-band physical index characteristics also include the phycocyanin index, normalized vegetation index, chlorophyll index, reflectance peak height, absorption valley depth, and maximum rate of change of reflectance peak for different combinations of bands.

4. The method according to claim 1, characterized in that, The preset band range is a continuous range consisting of green light, yellow light, red light, and red-edge bands.

5. The method according to claim 1, characterized in that, The matching dataset is based on temporal resolution matching of total phosphorus data, including: If the total phosphorus data is sampled at regular intervals, the average value of the hyperspectral data within t minutes before and after the total phosphorus data collection time will be used to match it. Hyperspectral data outside the time window will not be included in the matching. t is a preset window threshold. If the total phosphorus data is diurnal, it is matched with the arithmetic mean of the hyperspectral data during a preset diurnal period; the preset diurnal period is determined based on the latitude and longitude of the target lake / reservoir and the sampling season. The hyperspectral data is used to remove bands where the system error exceeds a preset error value.

6. The method according to claim 1, characterized in that, The hyperspectral-total phosphorus concentration spatiotemporal matching dataset was preprocessed and used for feature calculation. The preprocessing includes: calculating the median of hyperspectral reflectance within a sliding window for the hyperspectral time series data in the matching dataset; then subtracting the median from the original hyperspectral reflectance value to obtain the residual sequence; calculating the median absolute deviation (MAD) of the residual sequence; identifying points exceeding a preset deviation range as outliers; and performing linear interpolation replacement using the two nearest non-outliers before and after the outlier.

7. The method according to claim 1, characterized in that, The machine learning model used is a tree ensemble model.

8. The method according to claim 1, characterized in that, The LASSO regression method was used to initially screen each feature in the candidate feature set.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 8.