Coastal wetland grass seed classification method based on sentinel No.2 time sequence analysis and random forest

Through Sentinel 2 timing analysis and random forest method, combined with multi-time phase optical image and radar data, the time-differential vegetation index was constructed, which solved the problems of insufficient utilization of high-cost data and poor tide processing in coastal wetland grass species classification, and achieved high-precision grass species distinction.

CN120355984APending Publication Date: 2025-07-22NANJING FORESTRY UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510415917.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art has problems in the classification of coastal wetland grass species inadequately utilized high-cost data, insufficient utilization of phenological characteristics and limited effect of tidal treatment methods, resulting in low classification accuracy, and it is especially difficult to effectively distinguish reeds and melanchoe with highly similar spectral characteristics.

Method used

Sentinel 2 timing analysis and random forest method were used, combined with multi-time phase optical image and radar data, vegetation characteristics, texture characteristics and polarization characteristics were extracted, time-different vegetation index was constructed, and the SHAP method was used for visual analysis to quantify the importance of characteristics, and classification accuracy was evaluated through ten-fold cross-validation.

Benefits of technology

It significantly improves the accuracy of grass species classification in coastal wetlands, reduces the impact of cloud coverage and tidal interference, enhances the discrimination between reeds and peaches, makes full use of the complementary advantages of multi-source data, and provides more reliable classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355984A_ABST
    Figure CN120355984A_ABST
Patent Text Reader

Abstract

The invention discloses a coastal wetland grass seed classification method based on sentinel No.2 time sequence analysis and a random forest, and the method comprises the steps: extracting spectrum, texture and polarization parameters as classification variables through combining a sentinel No.2 multi-temporal optical image and sentinel No.1 radar data; and the monthly NDVI maximum time sequence of the sentinel No.2 image is utilized to analyze and capture the difference characteristics of the grass seeds in the growth peak period, the aging period and the change rate, a time difference vegetation index is constructed, and the distinguishability of reed and spartina alterniflora with similar spectral characteristics is enhanced. A random forest algorithm is adopted to carry out classification modeling in combination with a suaeda salsa index, an SHAP method is introduced to carry out visual analysis on the model, contribution of characteristic variables in grass seed identification is quantified, and interaction between characteristics and an influence mechanism of the characteristics on a classification result are disclosed. And finally, the classification precision is evaluated through ten-fold cross validation, and the robustness and reliability of the model are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a wetland vegetation classification method based on phenological characteristics, and in particular to a classification method for coastal wetland grass species based on Sentinel-2 time series analysis and random forest, belonging to the technical field of remote sensing images. Background Art

[0002] In recent years, the comprehensive prevention and control of invasive species has become one of the hot issues concerned at home and abroad. In the coastal areas of China, accurately obtaining the distribution information of grass species is of great significance for the effective supervision of invasive Spartina alterniflora and the formulation of coastal wetland ecological protection plans. However, due to the fragmented distribution of grass species in coastal wetlands and the fragile wetland environment, traditional field investigation methods not only have great monitoring difficulty and strong destructiveness, but also are difficult to achieve large-scale and long-term dynamic monitoring. Remote sensing technology, with its wide observation range, multi-spectral information and fixed-period observation ability, provides an efficient and non-destructive technical means for monitoring the distribution of coastal wetland grass species and their long-term dynamic changes.

[0003] Many classification tasks use hyperspectral images combined with vegetation indices for coastal wetland grass species classification. The rich spectral bands of hyperspectral data provide sufficient data support for spectral analysis, but its acquisition cost is high and data processing is complex, which limits its application in large-scale and long-time series studies. In addition, although high-resolution images and high-precision lidar data can provide more detailed ground object information, their high acquisition cost and data processing difficulty also limit their applicability in coastal wetland grass species monitoring. In contrast, Sentinel series medium-resolution satellite data has been widely used in coastal wetland grass species monitoring due to its moderate spatial resolution, good time continuity and easy accessibility.

[0004] Although deep learning algorithms show high accuracy in remote sensing image classification, the interpretability of their abstract features is low, and it is difficult to stably apply them to decision-making support for ecological protection and management. Therefore, how to improve the interpretability of the model while ensuring the classification accuracy of coastal wetland grass species is one of the important challenges in the current coastal wetland grass species classification task. At present, some studies have used the characteristic that the reflectance of the red band is relatively high during the prosperous growth period of Suaeda salsa to construct the Suaeda salsa index, and successfully realized the effective extraction of the distribution of Suaeda salsa in coastal wetlands. However, for reeds and invasive Spartina alterniflora with highly similar spectral characteristics, although time series analysis can capture the subtle differences in spectral information between the two, due to the high overlap of their spectral characteristics and the interference of environmental factors such as tides and soil moisture on spectral signals, the traditional single-phase vegetation index or spectral band combination has low discrimination in distinguishing these two grass species.

[0005] In summary, the existing technologies mainly have the following disadvantages:

[0006] (1) The classification research of multiple grass species relies on high-cost data, and the medium-resolution data is not fully utilized.

[0007] Although there have been attempts at classifying multiple grass species in the current research on coastal wetland grass species, it mainly focuses on hyperspectral data and high-spatial-resolution UAV images. The acquisition cost of such data is high and the coverage is limited, making it difficult to achieve large-scale promotion. In contrast, there are fewer studies on classifying multiple grass species using medium-resolution remote sensing data (such as Sentinel series data), and the advantages of its wide coverage, low cost, and rich time series have not been fully exploited.

[0008] (2) The phenological characteristics are not fully utilized.

[0009] Existing studies mainly rely on single phenological characteristics or single-temporal-phase data, failing to fully exploit the information of phenological curves and effectively analyze the temporal sequence data. This leads to low utilization of the spectral changes of coastal wetland grass species over time, restricting the improvement of classification accuracy.

[0010] (3) The effects of tidal treatment methods are limited.

[0011] The existing tidal treatment methods have poor effects in practical applications. Directly screening low-tide observation data cannot find the most effective observation time, and the times of key phenological period images and low-tide images usually do not fully coincide, resulting in low data utilization. At the same time, due to the cloudy weather in coastal areas, there is insufficient effective observation data, and the phenological curves are prone to underfitting, making the extracted phenological characteristics unstable and affecting the classification accuracy. Summary of the Invention

[0012] Objective of the Invention: Aiming at the above problems existing in the prior art, the objective of the present invention is to provide a classification method for coastal wetland grass species based on Sentinel-2 time series analysis and random forest.

[0013] Technical Solution: The classification method for coastal wetland grass species based on Sentinel-2 time series analysis and random forest described in the present invention includes the following steps:

[0014] (1) Obtain the data of on-site sampling points in coastal wetlands, and combine the multi-temporal optical images of Sentinel-2 and the radar data of Sentinel-1 to extract vegetation characteristics, texture characteristics, and polarization characteristics as alternative variables for grass species classification;

[0015] (2) According to the GEE platform, combine the data of on-site sampling points in coastal wetlands to extract the maximum value of the normalized difference vegetation index for each month, conduct time series analysis, capture the differential characteristics of coastal wetland grass species during the growth peak and senescence periods, construct a time-difference vegetation index, compare and analyze its effectiveness with conventional phenological period variables, and combine them with the alternative variables in step (1) as feature variables;

[0016] (3) Classification modeling is carried out using the data of field sampling points in coastal wetlands and characteristic variables combined with the random forest algorithm, and the SHAP (Shapley Additive Explanations) method is introduced to visually analyze the feature importance, quantify the contributions in different grass species identification, reveal the interaction between features and its impact on the classification results;

[0017] (4) The classification accuracy is evaluated through ten-fold cross-validation, and the classification results are predicted based on the trained random forest model. The Spartina alterniflora, Phragmites australis, and Suaeda glauca categories are extracted to obtain the distribution map of the main grass species in coastal wetlands.

[0018] Furthermore, in step (1), the vegetation features include the normalized difference vegetation index, the Suaeda glauca index, the annual average NDVI, and the enhanced vegetation index; the texture features include the mean, variance, homogeneity, contrast, dissimilarity, entropy, angular second moment, and correlation; the polarization features include the dual-polarization radar vegetation index, the pseudo-scattering entropy, the polarization purity, the equivalent scattering angle, the VV polarization median synthesis, the VH polarization median synthesis, the ratio of VV polarization median synthesis to VH polarization median synthesis, and the ratio of VH polarization median synthesis to VV polarization median synthesis.

[0019] Furthermore, in step (1), obtaining the data of field sampling points in coastal wetlands includes the categories and coordinates of grass species. The grass species include Spartina alterniflora, Phragmites australis, and Suaeda glauca. Artificial annotation of grass species and water body data is added near the sampling points by combining with high-resolution maps.

[0020] Furthermore, in step (1), the specific steps are as follows:

[0021] Using Sentinel-2 data during the growth peak periods (September and October) of coastal wetland grass species, calculate the NDVI pixel by pixel and retrieve the maximum value, while supplemented with quality control conditions (such as cloud cover). Finally, generate a new multispectral image and calculate its vegetation indices, including the normalized difference vegetation index, the Suaeda salsa index, the annual average NDVI, and the enhanced vegetation index, to capture the spectral response characteristics of vegetation during key phenological periods; extract texture features, including mean, variance, uniformity, contrast, dissimilarity, entropy, angular second moment, and correlation, from the weighted grayscale image (0.3×B8 + 0.59×B4 + 0.11×B3) of the new multispectral data to characterize the spatial distribution and structural information of vegetation; use the Sentinel-1 median composite radar data during the growth peak periods (September and October) of the whole year to extract polarization features, including the dual-polarization radar vegetation index, entropy, anisotropy, scattering angle, VV polarization, VH polarization, the ratio of VV polarization to VH polarization, and the ratio of VH polarization to VV polarization, to supplement the microwave scattering characteristics of vegetation; use the above spectral information, texture features, and polarization features as alternative variables for grass species classification, combined with the coastal wetland field sampling point data and the manually annotated data based on the coastal wetland field sampling points, as the input data for the random forest model classification.

[0022] Furthermore, in step (2), according to the distribution trend of the monthly maximum NDVI values of grass species at the coastal wetland sampling points, extract the growth peak period (May) and senescence period (December) of Spartina alterniflora and Phragmites australis to construct the temporal difference vegetation index, and compare it with the month with the largest NDVI difference between Spartina alterniflora and Phragmites australis (May) and the conventional phenological period (October) when vegetation is commonly lush in the study. It is found that the temporal difference vegetation index has higher separability.

[0023] Furthermore, in step (2), the specific steps are as follows:

[0024] Based on the coastal wetland field sampling point data and Sentinel-2 remote sensing images, use the GEE platform to extract the monthly maximum NDVI data of the coastal wetland field sampling points, screen the monthly maximum NDVI data throughout the year, identify and remove outliers through the box plot method, and retain high-quality data for analyzing the temporal spectral differences, construct the temporal difference vegetation index, and enhance the distinguishability between Phragmites australis and Spartina alterniflora;

[0025] Construct box plots of the monthly maximum NDVI data for each grass species according to the data, find the differences between Spartina alterniflora and Phragmites australis during the growth peak and senescence periods. Use the date of the 15th day of the monthly maximum NDVI data as the annual accumulated day, fit a double logistic curve to determine the growth rate and phenological period differences; select the start time and end time of the NDVI change and the time when the maximum NDVI values of different grass species differ significantly. Use this difference to propose the Time Difference Vegetation Index (TDVI), further enhancing the distinguishability between Phragmites australis and Spartina alterniflora. Combine the time difference vegetation index with the alternative variables in the previous step (1) and input them as the characteristic variables for subsequent grass species classification modeling.

[0026] Furthermore, screen the monthly maximum NDVI data within a year, and identify and remove outliers by the box plot method. The specific method is as follows: Calculate the interquartile range IQR of the NDVI values of each pixel in the time series, defined as the difference between the upper quartile Q3 and the lower quartile Q1, that is, IQR = Q3 - Q1; regard the data points with NDVI values less than Q1 - 1.5×IQR or greater than Q3 + 1.5×IQR as outliers and remove them to ensure the stability and reliability of the data.

[0027] Furthermore, the time difference vegetation index is calculated according to the following formula (1):

[0028]

[0029] where NDVI t1 represents the NDVI value in the senescence period, and NDVI t2 represents the NDVI value in the growth peak period.

[0030] Further, in step (3), the SHAP value calculates the contribution of each feature to the model output through the following formula (2):

[0031]

[0032] where φ i represents the SHAP value of feature i, F is the set of all features, S is the subset that does not contain feature i, and f(S) represents the model output based on the feature subset S. By calculating the SHAP value, quantify the contribution of each feature to the classification result and further prove the effectiveness of the feature.

[0033] Further, in step (4), the specific steps are as follows:

[0034] Use the 10 best predictive variables and sample point data screened in step (3) to train a random forest model, and evaluate the model accuracy through ten-fold cross-validation; after the model training is completed, use the trained random forest model to predict the study area, generate a preliminary classification result, extract the target grass species, and finally obtain the distribution map of the main grass species in the coastal wetland.

[0035] Furthermore, ten-fold cross-validation divides the dataset into 10 subsets, successively uses one of the subsets as the validation set, and the remaining 9 subsets as the training set, repeats the training and validation process 10 times, and finally takes the average accuracy as the performance index of the model. The formula (3) for evaluating the classification accuracy by ten-fold cross-validation is as follows:

[0036]

[0037] where the accuracy k represents the classification accuracy of the k-th cross-validation.

[0038] By analyzing the annual changes of the NDVI of Spartina alterniflora and Phragmites australis each month, the present invention constructs a time-difference vegetation index, makes full use of the dynamic differences in spectral information at different times, and enhances the distinguishability between Phragmites australis and Spartina alterniflora. The time-difference vegetation index can effectively capture the spectral change characteristics of grass species in the coastal wetland during the growth cycle, and significantly improve the classification accuracy of grass species in the coastal wetland.

[0039] Advantages: Compared with the prior art, the present invention has the following remarkable advantages:

[0040] (1) Based on the monthly maximum NDVI data for time series analysis, the present invention effectively reduces the probability of cloud cover and tidal interference observations in the coastal area appearing in the data, makes the maximum use of high-quality observation pixels, and solves the problem of unstable fitting of the time series curve caused by insufficient data and tidal masking interference in the traditional time series curve, providing a more reliable data basis for the classification of grass species in the coastal wetland;

[0041] (2) The present invention innovatively extracts the phenological difference vegetation index features, combines the characteristics of the NDVI change rate differences between Spartina alterniflora and Phragmites australis during the growth cycle, and constructs a time-difference vegetation index (TDVI), which significantly enhances the distinguishability between Phragmites australis and Spartina alterniflora;

[0042] (3) By combining Sentinel-2 optical images and Sentinel-1 radar data, the present invention extracts spectral information, texture features and polarization parameters, makes full use of the complementary advantages of multi-source data, and avoids the limitations of a single data source in classification. Description of the Drawings

[0043] Figure 1Classification framework flowchart of coastal wetland grass species based on Sentinel-2 time series analysis and random forest

[0044] Figure 2 Distribution map of the core area of Jiangsu Yancheng National Nature Reserve for Rare Birds and manually annotated data

[0045] Figure 3 Distribution and outlier removal map of the monthly maximum NDVI values at sampling points throughout the year 2023

[0046] Figure 4 Time series curve graph of Spartina alterniflora and Phragmites australis

[0047] Figure 5 Display graph of the distinguishable degree of features of Spartina alterniflora and Phragmites australis. Among them, (a) is the composite of the maximum NDVI value in October; (b) is the composite of the maximum NDVI value in May; (c) is the time difference vegetation index; (d) is the distribution of the maximum NDVI value at sampling point in October; (e) is the distribution of the maximum NDVI value at sampling point in May; (f) is the distribution of the time difference vegetation index at the field sampling points in coastal wetlands

[0048] Figure 6 Contribution degree of different features in the category of Spartina alterniflora

[0049] Figure 7 Distribution map of the main grass species in Jiangsu Yancheng National Nature Reserve for Rare Birds Specific implementation manner

[0050] The present invention will be further described below in conjunction with specific embodiments

[0051] Embodiment 1

[0052] For the classification of coastal wetland grass species, the present invention fully considers the influence of cloud cover and tidal cycle on the observed data, and proposes a classification method of coastal wetland grass species based on Sentinel-2 time series analysis and random forest. Through the monthly maximum NDVI composite image, the problem of insufficient observed data caused by cloudy weather and tidal masking in coastal wetlands during time series analysis, and thus inaccurate phenological analysis, is significantly reduced; a time difference vegetation index is constructed, which significantly enhances the distinguishable degree between Phragmites australis and Spartina alterniflora, and improves the accuracy and efficiency of Sentinel-2 data in the classification of coastal wetland grass species

[0053] 1. Study area and data preprocessing

[0054] Jiangsu Yancheng Rare Birds National Nature Reserve is located on the east coast of China, in the eastern part of Yancheng City, Jiangsu Province. It is one of the important coastal wetland ecosystems along the Yellow Sea. The region belongs to the warm temperate monsoon climate, with distinct seasons, abundant rainfall, diverse vegetation types, and a complex and rich ecosystem. The core area of the reserve is located in the Yellow Sea sedimentary plain, with sediments mainly consisting of silt and sand, forming a large area of tidal flat wetlands and salt marsh vegetation, providing an ideal habitat for numerous rare birds and plants, and having extremely high ecological protection value and significance for biodiversity. The study area is located in the core area of the reserve, which is a key area where rivers meet the sea, and is distributed with typical salt marsh vegetation such as Spartina alterniflora, Suaeda glauca, and Phragmites australis. It is one of the areas most severely invaded by Spartina alterniflora in recent years. The location of the study area is as Figure 2 shown.

[0055] The data sources of this embodiment include field sampling data and satellite remote sensing data. (1) Field sampling point data of coastal wetlands: including grass species (Spartina alterniflora, Phragmites australis, Suaeda glauca), and the distribution coordinates of each grass species. To expand the training data samples, this embodiment also uses high-resolution Google Maps to add artificial annotation data near the sampling points, records the grass species (Spartina alterniflora, Suaeda glauca, Phragmites australis) and coordinate information, and at the same time adds water body sampling points, such as Figure 2 shown.

[0056] (2) Satellite remote sensing data includes Sentinel-1 and Sentinel-2 remote sensing images. The Sentinel-1 data converts the DN value of the SAR signal amplitude into the backscattering coefficient. The Sentinel-2 images are obtained through the Google Earth Engine platform, and are geometrically corrected, radiometrically calibrated, and atmospherically corrected to obtain surface reflectance data. All preprocessed remote sensing images are cropped with reference to the study area boundary.

[0057] 2. Calculation of alternative feature variables

[0058] In this study, Sentinel-2 data during the growth peak period (September 1 - October 31) of coastal wetland grass species was used to calculate NDVI for each pixel and retrieve the maximum value, while supplemented with quality control conditions (Cloud Score + Sentinel-2 dataset). Finally, a new multi-spectral image was generated, and its original band information was extracted and its vegetation index and texture features were calculated (Table 1). Polarization features (Table 1) were extracted from Sentinel-1 data synthesized by the median from September 1 to October 31. All alternative variables for grass species classification in this embodiment are shown in Table 1.

[0059] Table 1 Variables Extracted from Sentinel Remote Sensing Images

[0060]

[0061]

[0062]

[0063] 3. Construction of Temporal Difference Vegetation Index

[0064] Based on the data of field sampling points in coastal wetlands and Sentinel-2 remote sensing images, using the GEE platform, the maximum NDVI data of each month for the field sampling points in coastal wetlands are extracted. First, the maximum NDVI data of each month are screened, and outliers are identified and removed by the box plot method. The specific method is as follows: Calculate the interquartile range (IQR) of the NDVI values of each pixel in the time series, which is defined as the difference between the upper quartile (Q3) and the lower quartile (Q1), that is, IQR = Q3 - Q1. Data points with NDVI values less than Q1 - 1.5×IQR or greater than Q3 + 1.5×IQR are regarded as outliers and removed to ensure the stability and reliability of the data. Through the above screening method, high-quality data are retained for time series analysis, and the outlier removal is as follows Figure 3 .

[0065] This embodiment is based on the synthetic data of the maximum NDVI values of 12 months within 2023 of Sentinel-2 (see Figure 3 ). After removing outliers, the mean value of the maximum NDVI of each month for all field sampling points in coastal wetlands is calculated, and the 15th day of each month is used as the representative value of the day of year (DOY). The double logistic function is used to fit the time series curve to determine the growth rate and phenological differences. The double logistic function can effectively capture the key phenological stages in the vegetation growth cycle, including the initial growth stage, the peak growth stage, and the senescence stage (see Figure 4 ). Through the analysis of the time series curve, the NDVI change trends and phenological characteristics of different grass species are reflected, providing more accurate time series change information and laying a foundation for subsequent grass species classification and dynamic monitoring.

[0066] By analyzing the time series curves of the mean values of the maximum NDVI of each month for Spartina alterniflora and Phragmites australis (such as Figure 4 ) and the box plots of the maximum NDVI of each month for different grass species (see Figure 3), the difference in the change rate of NDVI between Spartina alterniflora and Phragmites australis was found. Phragmites australis showed a faster rising rate of NDVI in the early growth stage, while it showed a faster declining rate of NDVI in the late growth stage. This difference provides an important basis for distinguishing different grass species in terms of temporal characteristics. Based on this finding, in this embodiment, December at the end of the NDVI growth period was selected as the initial value, and May at the maximum growth period was selected as the end value. The difference between the two periods was used to amplify the impact of the change rate, and normalization was used to construct the Time Difference Vegetation Index (TDVI). The calculation formula is as follows in Formula (1), and a TDVI distribution map ( Figure 5 in (c), (f) of

[0067]

[0068] where NDVI t1 and NDVI t2 respectively represent the NDVI values at the key time phases with relatively large differences selected. In this embodiment, NDVI t1 represents the NDVI value in the senescence period, and NDVI t2 represents the NDVI value in the peak growth period. In this embodiment, t1 is December and t2 is May.

[0069] The time difference vegetation index constructed in this embodiment was compared with the NDVI in October during the lush growth period (conventional phenological period) commonly used in the study ( Figure 5 in (a), (d) of Figure 5 in (b), (e) of

[0070] 4. Variable Screening and Grass Species Classification

[0071] In this embodiment of the present invention, the random forest algorithm in the scikit-learn package is used in the Python environment to perform two steps: variable screening and classification modeling.

[0072] (1) Variable screening. Stack the obtained time-difference vegetation index features with the above-mentioned alternative variables, extract the values of each feature variable according to the coordinate information of the manually marked points at the field sampling points, and use them as the input for the screening of grass species classification variables in coastal wetlands. Select the top 10 feature variables in terms of importance, and introduce the SHAP (Shapley Additive Explanations) method to further evaluate the contribution of feature variables to the classification results. By calculating the Shapley values of the ten features, quantify the importance of different features in the grass species classification of coastal wetlands. Among them, the SHAP value calculates the contribution of each feature to the model output through the following formula (2):

[0073]

[0074] Among them, φi represents the SHAP value of feature i, F is the set of all features, S is the subset that does not include feature i, and f(S) represents the model output based on the feature subset S.

[0075] By calculating the SHAP values, quantify the contribution of each feature to the classification results, and further prove the effectiveness of the features. The results show that the time-difference vegetation index has the highest contribution rate to the extraction of Spartina alterniflora (see Figure 6 ), indicating the effectiveness of the time-difference vegetation index. This method is implemented through the shap package in Python.

[0076] (2) Grass species classification. Use the above ten important feature values as the input of the random forest model. The model parameters are set as n_estimators = 100 and random_state = 42. Among them, the value of n_estimators is based on the balance of literature query and preliminary experimental results, aiming to ensure the prediction performance and computational efficiency of the model; the fixing of random_state ensures the repeatability of the experimental process, so that the results can be reproduced and verified under the same conditions. Finally, a distribution map of the main grass species in coastal wetlands is obtained ( Figure 7 ). To more clearly display the spatial distribution of the main grass species, Figure 7 the water body is masked.

[0077] 5. Accuracy verification and spatial mapping

[0078] During the model training process, the present invention uses the ten-fold cross-validation method to evaluate the performance of the random forest model to ensure the stability and scientificity of the classification results. The ten-fold cross-validation divides the data set into 10 subsets, sequentially uses one of the subsets as the validation set, and the remaining 9 subsets as the training set, repeats the training and validation process 10 times, and finally takes the average accuracy as the performance index of the model. The ten-fold cross-validation evaluates the classification accuracy using the following formula (3):

[0079]

[0080] Among them, the accuracy k represents the classification accuracy of the k-th cross-validation.

[0081] The final average classification accuracy is used as the accuracy of the classification result (see Table 2).

[0082] Table 2 Average Classification Accuracy of Ten-Fold Cross-Validation

[0083]

Claims

1. A classification method for coastal wetland grass species based on Sentinel-2 time series analysis and random forest, characterized in that, The method includes the following steps: (1) Obtain the field sampling point data of coastal wetlands, and combine the Sentinel-2 multi-temporal optical images and Sentinel-1 radar data to extract vegetation characteristics, texture characteristics, and polarization characteristics as alternative variables for grass species classification; (2) According to the GEE platform, combine the field sampling point data of coastal wetlands to extract the maximum value of the normalized difference vegetation index (NDVI) for each month, conduct time series analysis, capture the differential characteristics of grass species in coastal wetlands during the growth peak and senescence periods, construct the time-differential vegetation index, compare and analyze its effectiveness with conventional phenological variables, and merge it with the alternative variables in step (1) as feature variables; (3) Use the field sampling point data of coastal wetlands and feature variables to combine with the random forest algorithm for classification modeling, and introduce the SHAP method to visualize the feature importance, quantify the contributions in different grass species identifications, and reveal the interaction between features and its impact on the classification results; (4) Evaluate the classification accuracy through ten-fold cross-validation, and predict the classification results based on the trained random forest model, and extract the Spartina alterniflora, Phragmites australis, and Suaeda salsa categories to obtain the distribution map of the main grass species in coastal wetlands.

2. The coastal wetland grass species classification method based on Sentinel-2 time series analysis and random forest according to claim 1, characterized in that In step (1), the vegetation characteristics include the normalized difference vegetation index, Suaeda salsa index, annual average NDVI, and enhanced vegetation index; the texture characteristics include mean, variance, homogeneity, contrast, dissimilarity, entropy, angular second moment, and correlation; the polarization characteristics include the dual-polarization radar vegetation index, pseudo-scattering entropy, polarization purity, equivalent scattering angle, median synthesis of VV polarization, median synthesis of VH polarization, and the ratio of median synthesis of VV polarization to median synthesis of VH polarization, and the ratio of median synthesis of VH polarization to median synthesis of VV polarization.

3. The coastal wetland grass species classification method based on Sentinel-2 time series analysis and random forest according to claim 1, wherein In step (1), obtaining the field sampling point data of coastal wetlands includes the categories and coordinates of grass species. The grass species include Spartina alterniflora, Phragmites australis, and Suaeda salsa. Manually combine with a high-spatial-resolution map to add manually labeled grass species and water body data near the sampling points.

4. The coastal wetland grass species classification method based on Sentinel-2 time series analysis and random forest according to claim 1, characterized in that In step (1), the specific steps are as follows: Using Sentinel-2 data during the growth peak period of coastal wetland grass seeds, calculate the NDVI pixel by pixel and retrieve the maximum value. At the same time, supplemented by quality control conditions (such as cloud cover), finally generate a new multispectral image, extract its original spectral information, and calculate its vegetation indices, including the normalized difference vegetation index, the Suaeda salsa index, the annual average NDVI, and the enhanced vegetation index, to capture the spectral response characteristics of vegetation during the key phenological periods; generate a weighted grayscale image based on the new multispectral image and extract texture features, including mean, variance, uniformity, contrast, dissimilarity, entropy, angular second moment, and correlation, to characterize the spatial distribution and structural information of vegetation; use the Sentinel-1 median composite radar data during the growth peak period of the whole year to statistically screen the number of available pixels and the total NDVI value, and calculate the mean NDVI; combine the polarization features extracted from the Sentinel-1 median composite radar data during the growth peak period, including the dual-polarization radar vegetation index, the pseudo-scattering entropy, the polarization purity, the equivalent scattering angle, the VV polarization median composite, the VH polarization median composite, and the ratio of VV polarization to VH polarization, the ratio of VH polarization to VV polarization, to supplement the microwave scattering characteristics of vegetation; use the above spectral information, texture features, and polarization features as alternative variables for grass seed classification, combined with the data of field sampling points in coastal wetlands and the manually annotated data based on the field sampling points in coastal wetlands, as the input data for the random forest model classification.

5. The coastal wetland grass species classification method based on Sentinel-2 time series analysis and random forest according to claim 1, characterized in that In step (2), according to the distribution trend of the monthly NDVI maximum value of grass seeds at the coastal wetland sampling points, extract the growth peak period and senescence period of Spartina alterniflora and Phragmites australis to construct the temporal difference vegetation index, and compare it with the month with the largest NDVI difference between Spartina alterniflora and Phragmites australis and the conventional phenological periods with lush vegetation growth commonly used in the study, to prove that the temporal difference vegetation index has higher separability.

6. The method for classifying coastal wetland grass species based on Sentinel-2 time series analysis and random forest according to claim 1, wherein In step (2), the specific steps are as follows: According to the manually annotated data based on the field sampling point data in coastal wetlands and the Sentinel-2 remote sensing images, use the GEE platform to extract the monthly NDVI maximum value data of the field sampling points in coastal wetlands, screen the monthly NDVI maximum value data throughout the year, identify and remove outliers through the box plot method, and retain high-quality data for analyzing the temporal spectral differences, construct the temporal difference vegetation index, and enhance the distinguishability between grass seeds with similar spectral characteristics; Construct box plots of the monthly NDVI maximum value data distributed by different grass species categories, find the differences between the growth peak period and senescence period of coastal wetland grass seeds, use the date of the 15th day of the monthly NDVI maximum value data as the annual accumulated day, fit the double logistic curve to determine the growth rate and phenological period differences, select the starting time and ending time of the NDVI change and the time with a large difference in the NDVI maximum value of different grass species, use this difference to propose the temporal difference vegetation index TDVI, reflect the dynamic change characteristics of grass seeds during the key phenological periods through the temporal difference vegetation index, further enhance the distinguishability between Phragmites australis and Spartina alterniflora, and merge with the alternative variables in the previous step (1) as the input of the feature variables for the subsequent grass seed classification modeling.

7. The method for classifying coastal wetland grass seeds based on Sentinel-2 time series analysis and random forest according to claim 6, wherein The maximum NDVI data of each month of the year were screened, and outliers were identified and removed through the box plot method. The specific method was as follows: the interquartile range (IQR) of the NDVI value of each pixel in the time series was calculated, which was defined as the difference between the upper quartile Q3 and the lower quartile Q1, that is, IQR = Q3-Q1; data points with NDVI values less than Q1-1.5×IQR or greater than Q3+1.5×IQR were regarded as outliers and removed to ensure the stability and reliability of the data.

8. The method for classifying coastal wetland grass seeds based on Sentinel-2 time series analysis and random forest according to claim 6, wherein The time series difference vegetation index is calculated according to the following formula (1): Among them, NDVI t1 represents the NDVI value in the senescence period, and NDVI t2 represents the NDVI value in the peak growth period.

9. The method for classifying coastal wetland grass seeds based on Sentinel-2 time series analysis and random forest according to claim 1, wherein In step (3), the SHAP value is calculated by the following formula (2) to calculate the contribution of each feature to the model output: Among them, φ i represents the SHAP value of feature i, F is the set of all features, S is the subset excluding feature i, and f(S) represents the model output based on the feature subset S.

10. The method for classifying coastal wetland grass seeds based on Sentinel-2 time series analysis and random forest according to claim 1, characterized in that, In step (4), the ten-fold cross validation divides the data set into 10 subsets, one of which is used as the validation set and the other nine subsets as the training set. The training and validation process is repeated 10 times, and the average accuracy is finally taken as the performance indicator of the model. The ten-fold cross validation evaluates the classification accuracy using the following formula (3): Among them, the accuracy k represents the classification accuracy of the k-th cross-validation.

Citation Information

Cited By

  • Fast-growing forest felling detection method and device considering sentinel No.1 time sequence autocorrelation

    CN121596281A

  • Method and device for detecting fast-growing forest felling taking into account sentinel-1 temporal autocorrelation

    CN121596281B

  • Rice information extraction method and equipment based on multi-source remote sensing and machine learning

    CN121661528A

  • Spartina alterniflora identification method and system based on multi-source remote sensing image

    CN121767859A