Rubber forest extraction method based on GEE platform in combination with multi-source data and double-layer integration model

By using multi-source data and a two-layer integrated model on the GEE platform, combined with Stacking integrated learning model, the problem of existing rubber forest remote sensing extraction methods relying on image during the fallen leaf period is solved, and high-precision and full-season rubber forest extraction is achieved.

CN120107778APending Publication Date: 2025-06-06YUNNAN NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510060472.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing rubber forest remote sensing extraction method relies on imagery during the fallen leaf period and is susceptible to cloud occlusion and missing problems, resulting in a decrease in spatial continuity and accuracy of classification results.

Method used

The rubber forest extraction method based on the GEE platform is adopted to build a multi-source remote sensing data set and a phenological four-time integrated model (PFT-EM), combined with the Stacking integrated learning model, and the output results of different machine learning models are fused to improve the extraction accuracy and generalization capabilities.

Benefits of technology

It effectively compensates for the shortage of a single data source, improves the accuracy and reliability of rubber forest extraction, can observe rubber forests throughout the season, reduces dependence on images during the fallen leaves, and improves the spatial continuity and accuracy of classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107778A_ABST
    Figure CN120107778A_ABST
Patent Text Reader

Abstract

The invention discloses a rubber forest extraction method based on a GEE platform in combination with multi-source data and a double-layer integration model, and the method comprises the following steps: obtaining and preprocessing the multi-source data through the GEE platform, and constructing a rubber forest sample data set in combination with field survey data; extracting multi-dimensional features of the rubber collar sample data set; recursive training is carried out on the multi-dimensional data set of each stage by adopting a recursive feature elimination algorithm RFE, and feature selection is carried out; dividing phenology of the rubber tree into four stages by utilizing inflection points and morphological characteristics of the vegetation index curve; constructing a phenological four-period model PFT-EM of multi-category classification decision, inputting the images of the four stages into the model, and training machine learning models according to different phenological stages of rubber; and the Stacking ensemble learning model SEL fuses an output result of the basic learning model through a training element learning device. PFT-EM information based on different machine learning is integrated again through Stacking ensemble learning, so that the accuracy and generalization ability of the whole model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of forestry remote sensing technology, and in particular to a rubber forest extraction method based on a GEE platform combined with multi-source data and a double-layer integrated model. Background Art

[0002] Rubber trees have extremely high economic value. Since the 1990s, stimulated by market demand, the global rubber planting area has expanded rapidly, which has brought huge economic income to the local area. However, the expansion of rubber forests has also led to changes in regional microclimate, increased soil erosion, weakened water conservation functions, excessive deforestation, and reduced biodiversity. Therefore, it is urgent to use effective monitoring methods to grasp the spatiotemporal pattern of regional rubber forests, so as to provide important support for the development of the rubber industry and environmental protection.

[0003] In the current remote sensing extraction research of rubber plantations, the advantages of multi-source and different resolutions of remote sensing data have gradually become the key factors to improve the extraction accuracy and reliability. First of all, the main resolutions of remote sensing data include spatial, spectral, temporal and radiometric resolution. The higher the spatial resolution, the richer the data details, but in the study of large-scale rubber plantations, high resolution is not always necessary because the processing cost of large-scale remote sensing data is high. In addition, the Landsat series of satellites (Landsat 4, 5, 7, 8) have become the mainstream data source in remote sensing extraction research due to their free data and moderate spatial resolution, especially when monitoring vegetation and land changes, multispectral data has been widely used. In terms of temporal resolution, the temporal revisit periods of MODIS and Landsat are 1 day and 16 days respectively, which are suitable for multi-temporal analysis, especially under the influence of clouds and precipitation in tropical areas. The Sentinel series of satellites, especially Sentinel-1 and Sentinel-2, have become an important data source for remote sensing research on rubber plantations due to their high temporal and spatial resolution and free data acquisition advantages. The synthetic aperture radar (SAR) data of Sentinel-1 can penetrate clouds and is suitable for remote sensing monitoring in tropical areas, while the high spatial resolution and short revisit period provided by Sentinel-2 have greatly improved the ability to monitor vegetation and analyze changes. In the remote sensing extraction of rubber plantations, the multi-source nature of data has become an important factor in improving classification accuracy and research reliability. Combining optical, radar and multi-temporal data, especially taking advantage of the Sentinel series of satellites, can effectively make up for the shortcomings of a single data source, especially under the influence of clouds and precipitation in tropical areas. The free and open data of Sentinel-1 and Sentinel-2 provide remote sensing images with high temporal and spatial resolution, which significantly improves the extraction capabilities of large-scale rubber plantations in the study area. Therefore, the integration of multi-source data provides new research directions and technical approaches for remote sensing extraction of rubber plantations.

[0004] It would be more economical and efficient to extract rubber forests using remote sensing technology. At present, the remote sensing identification method of rubber forests is mainly based on phenology. Although the phenology-based method can better distinguish rubber forests from natural forests and farmland, it still relies heavily on the leaf-falling period of rubber forests. Once the remote sensing images of the leaf-falling period are blocked by clouds or missing, the spatial continuity and accuracy of the classification results will be greatly affected. Therefore, how to overcome the limitation of the scarcity of remote sensing images during the leaf-falling period and expand the effective observation time to the entire growth stage of the rubber forest still faces huge challenges. Summary of the invention

[0005] The purpose of the present invention is to provide a rubber forest extraction method based on GEE platform combined with multi-source data and double-layer integrated model to solve the above problems. A double-layer integrated model framework is constructed on Google Earth Engine platform based on multi-source data for rubber forest extraction in large areas.

[0006] The technical solution of the present invention is as follows:

[0007] The rubber forest extraction method based on GEE platform combined with multi-source data and double-layer integrated model includes the following steps:

[0008] Data collection and feature extraction: Multi-source data was acquired and preprocessed through the GEE platform. Google Earth high-definition images and Sentinel-2 satellite images were used in combination with field survey data to construct a rubber forest sample dataset. Multi-dimensional features of the rubber forest sample dataset were extracted.

[0009] Feature selection: The recursive feature elimination algorithm RFE is used to recursively train the multidimensional data set at each stage for feature selection;

[0010] The rubber tree phenology is divided into four stages using the inflection points and morphological characteristics of the vegetation index curve, and the images of the four stages are input into the model;

[0011] A four-period phenological model PFT-EM for multi-category classification decision-making was constructed, and machine learning models were trained according to different phenological stages of rubber. The PFT-EM model is:

[0012]

[0013] Among them, P i,c is the probability that pixel i belongs to category c in the jth image, T i is the number of images used to classify pixel i; c represents different categories, ranging from 1 to C, C is the total number of categories; calculate the average probability P that pixel i belongs to each category c i,c, use function D(i) to make classification decisions;

[0014] Construct Stacking ensemble model: The Stacking ensemble learning model SEL integrates the output results of the basic learning model by training the meta-learner. SEL includes the basic learning model and the meta-learner. The basic learning model is PFT-EM trained by different machine learning methods. The configuration scheme with the highest accuracy is selected for the extraction of rubber forest through the SEL configuration method.

[0015] Furthermore, the SEL configuration method includes:

[0016] Sort the training results of the basic learning models by accuracy, sort the configurations of the basic learning models from high to low according to accuracy, and select them in sequence. The number of basic learning models is 1-5;

[0017] The classification accuracy of different configuration schemes was compared, and the configuration scheme with the highest accuracy was selected for rubber forest extraction.

[0018] Furthermore, the SEL evaluates the accuracy of the training results of the basic learning model through the confusion matrix and derived parameters, and the derived parameters include: overall accuracy OA, Kappa coefficient, user accuracy UA, producer accuracy PA and F1 score; the confusion matrix is ​​used to calculate the user accuracy and mapping accuracy of the rubber forest, and the F1 score of the rubber forest is calculated:

[0019]

[0020] Among them, β represents the weight relationship between UA and PA. When β is set to 1, it means that the weights of UA and PA are equal.

[0021] Furthermore, the confusion matrix is ​​constructed by the model classification results and the number of validation sample points. The sample points are randomly divided into five equal-sized subsets using the five-fold cross-validation method. One of them is selected as the validation set in turn, and the rest are used as training sets. This cycle is repeated five times so that each subset can be used as a validation set once. After each round of training, the model performance is evaluated on the validation set, and finally the results of the five validations are averaged as the comprehensive performance indicator of the model.

[0022] Furthermore, the RFE algorithm comprises the following steps:

[0023] Sort all features involved in training according to feature weights or importance scores and remove the lowest ranked features;

[0024] Keep the remaining features and repeat the training and sorting process, continue to eliminate the features with the lowest weights, until all features have been iteratively screened;

[0025] The feature set with the highest classification accuracy in the iterative process is selected as the best feature subset.

[0026] Furthermore, the RFE algorithm performs feature selection specifically including the following steps:

[0027] The machine learning models in the GEE platform were selected for training, including random forest RF, maximum entropy model MaxEnt, gradient boosting tree GTB and classification regression tree CART;

[0028] The feature weights or importance scores obtained through training are subjected to RFE algorithms, namely RF-RFE, MaxEnt-RFE, GTB-RFE and CART-RFE algorithms, and the algorithms are compared to select the algorithm with the highest accuracy improvement for feature selection.

[0029] Furthermore, the multi-source data includes: Sentinel-1 synthetic aperture radar SAR, Sentinel-2 multispectral, bioclimatic and DEM data;

[0030] Sentinel-1 synthetic aperture radar SAR data comes from ground range detection products. The data spatial resolution is 10 meters. The data is a first-level image dataset obtained through Doppler centroid estimation, single-view composite focusing and post-processing. The dataset images are pre-processed by the Sentinel-1 toolkit, including speckle filtering, thermal noise removal, terrain correction and radiation correction.

[0031] Sentinel-2 multispectral data uses 2A-level surface reflectance data products obtained after atmospheric and orthorectification correction. The blue, green, red and near-infrared bands have a spatial resolution of 10m, the red edge / shortwave-infrared band has a spatial resolution of 20m, and the atmospheric correction band has a spatial resolution of 60m. The images are pre-processed in time and space by screening, splicing, cloud removal, and cropping.

[0032] The bioclimatic data uses the "WorldClim version 1" dataset with a spatial resolution of 927.67 meters. This dataset contains information representing annual trends: annual average temperature, annual precipitation, seasonality: annual temperature difference and precipitation, and extreme or limiting environmental factors;

[0033] The DEM data used is the Shuttle Radar Topography Mission V3 product with a spatial resolution of 30 meters, provided by NASA. All raster data were resampled to 10 m using bilinear interpolation, and the geographic coordinate system was unified to WGS_1984.

[0034] Furthermore, the multi-dimensional features include backscattering features BS, spectral features SP, texture features TX, bioclimatic features BC and terrain features T0.

[0035] Furthermore, the four stages of rubber forest phenology include:

[0036] The pre-defoliation period t1 means that the crown leaves change from dark green to light yellow;

[0037] The leaf-falling stage t2 means that the leaves change from light yellow to dark yellow until they fall off;

[0038] The initial stage of leaf expansion, t3, means the beginning of germination until new green leaves grow;

[0039] The intensive leaf expansion period t4 indicates that the new leaves have fully grown, the forest canopy is lush, and the leaves are dark green.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] 1. A rubber forest extraction method based on the GEE platform combined with multi-source data and a two-layer integrated model was constructed to construct a multi-source remote sensing dataset to make up for the deficiency of a single data source in remote sensing extraction of rubber forests under the influence of clouds and precipitation in tropical regions;

[0042] 2. Based on the GEE platform, a rubber forest extraction method combining multi-source data and a two-layer integrated model was developed to construct a phenological four-period integrated model (PFT-EM). The classification probability model of each phenological stage was integrated with machine learning, and the preliminary classification results containing all-season information were output to make up for the defect that the phenological method is overly dependent on images in a single time period.

[0043] 3. A rubber forest extraction method based on the GEE platform combined with multi-source data and a two-layer integrated model, which integrates PFT-EM information based on different machine learning methods through Stacking integrated learning to improve the accuracy and generalization ability of the overall model. This method can not only be used for remote sensing extraction research and application of rubber forests, but also has wide applicability for the extraction of other tree species with typical phenological change characteristics, providing decision support for forest resource inventory and forestry ecological protection management. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is a flow chart of the rubber forest extraction method based on the GEE platform combining multi-source data and a two-layer integrated model.

[0045] Figure 2 This is a conceptual diagram of the Stacking integration algorithm of this application.

[0046] Figure 3 Comparison of the classification results of the basic learning model and SEL in areas a, b, and c for this application.

[0047] Figure 4 This is a visual comparison of the classification results of the four sub-areas of this application. DETAILED DESCRIPTION

[0048] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0049] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.

[0050] See also Figure 1-4 , a rubber forest extraction method based on GEE platform combined with multi-source data and two-layer integrated model, such as Figure 1 As shown, the following steps are included:

[0051] Data collection and feature extraction: Multi-source data were acquired and preprocessed through the GEE platform. Based on field surveys, the land use and land cover (LULC) types in the study area were divided into five categories, namely rubber forest, natural forest, cultivated land, water surface and impervious surface. Google Earth high-definition images and Sentinel-2 satellite images were used in combination with field survey data to construct a rubber forest sample dataset; and the multidimensional features of the rubber forest sample dataset were extracted;

[0052] Multi-source data include: Sentinel-1 synthetic aperture radar SAR, Sentinel-2 multispectral, bioclimatic and DEM data;

[0053] Sentinel-1 synthetic aperture radar SAR data comes from ground range detection products. The data spatial resolution is 10 meters. The data is a first-level image dataset obtained through Doppler centroid estimation, single-view composite focusing and post-processing. The dataset images are pre-processed by the Sentinel-1 toolkit, including speckle filtering, thermal noise removal, terrain correction and radiation correction.

[0054] Sentinel-2 multispectral data uses 2A-level surface reflectance data products obtained after atmospheric and orthorectification correction. The blue, green, red and near-infrared bands have a spatial resolution of 10m, the red edge / shortwave-infrared band has a spatial resolution of 20m, and the atmospheric correction band has a spatial resolution of 60m. The images are pre-processed in time and space by screening, splicing, cloud removal, and cropping.

[0055] The bioclimatic data uses the "WorldClim version 1" dataset with a spatial resolution of 927.67 meters. This dataset contains information representing annual trends: annual average temperature, annual precipitation, seasonality: annual temperature difference and precipitation, and extreme or limiting environmental factors;

[0056] The DEM data used is the Shuttle Radar Topography Mission V3 product with a spatial resolution of 30 meters, provided by NASA. All raster data were resampled to 10 m using bilinear interpolation, and the geographic coordinate system was unified to WGS_1984.

[0057] The multidimensional features include backscattering features BS, spectral features SP, texture features TX, bioclimatic features BC and topographic features T0. A feature set containing five types of features was constructed, with a total of 61 feature factors, as shown in Table 1. Among them, backscattering features (BS), spectral features (SP) and texture features (TX) describe their physical properties, and bioclimatic features (BC) and topographic features (TO) can reflect their geographical distribution.

[0058] Table 1 Description and calculation method of characteristic factors

[0059]

[0060]

[0061] Feature selection: The recursive feature elimination algorithm RFE is used to recursively train the multidimensional data set at each stage for feature selection;

[0062] The RFE algorithm consists of the following steps:

[0063] Sort all features involved in training according to feature weights or importance scores and remove the lowest ranked features;

[0064] Keep the remaining features and repeat the training and sorting process, continue to eliminate the features with the lowest weights, until all features have been iteratively screened;

[0065] The feature set with the highest classification accuracy in the iterative process is selected as the best feature subset.

[0066] The RFE algorithm performs feature selection in the following steps:

[0067] The machine learning models in the GEE platform were selected for training, including random forest RF, maximum entropy model MaxEnt, gradient boosting tree GTB and classification regression tree CART;

[0068] The feature weights or importance scores obtained through training are subjected to RFE algorithms, namely RF-RFE, MaxEnt-RFE, GTB-RFE and CART-RFE algorithms, and the algorithms are compared to select the algorithm with the highest accuracy improvement for feature selection.

[0069] The rubber tree phenology is divided into four stages using the inflection point and morphological characteristics of the vegetation index curve. The physiological structure and morphological characteristics of rubber trees vary greatly during different phenological periods. The use of information from multiple phenological stages is conducive to the identification of rubber forests. At the same time, in order to ensure the full-season input of remote sensing images, the images of the four stages are input into the model; the vegetation index is the normalized difference vegetation index (NDVI) that can effectively reflect the density and intensity of the vegetation growth process; the land surface water index (LSWI) that reflects the changes in plant canopy, soil moisture and soil surface water content. The specific formulas and explanations of each index are shown in Table 1.

[0070] The phenology of rubber trees is the growth rhythm of leaves that changes with the environment and seasons, mainly including budding, leaf expansion (copper brown), color change (light green), yellowing leaves, leaf fall and dormancy.

[0071] The four phases of rubber forest phenology include:

[0072] The pre-defoliation period t1 refers to the period during which the crown leaves change from dark green to light yellow;

[0073] The leaf-falling period t2 refers to the period from when the leaves change from light yellow to dark yellow until they fall off;

[0074] The initial stage of leaf expansion, t3, refers to the period from the beginning of germination to the growth of new green leaves, when the new leaves are light green;

[0075] The intensive leaf expansion period t4 indicates that the new leaves have fully grown, the forest canopy is lush, and the leaves are dark green.

[0076] A four-period phenological model PFT-EM for multi-category classification decision-making was constructed, and machine learning models were trained according to different phenological stages of rubber. The PFT-EM model is:

[0077]

[0078] Among them, P i,c is the probability that pixel i belongs to category c in the jth image, T iis the number of images used to classify pixel i; c represents different categories, ranging from 1 to C, C is the total number of categories; calculate the average probability P that pixel i belongs to each category c i,c , function D(i) is used for classification decision; PFT-EM uses five machine learning classifiers, namely RF, MaxEnt, GTB, SVM and CART, and their hyperparameter settings are shown in Table 2.

[0079] Table 2 Hyperparameter settings of machine learning classifiers in GEE platform

[0080]

[0081] Construct Stacking ensemble model: The Stacking ensemble learning model SEL integrates the output results of the basic learning model by training the meta-learner. SEL includes the first-layer basic learning model and the second meta-learner. The basic learning model is PFT-EM trained by different machine learning. The configuration scheme with the highest accuracy is selected for rubber forest extraction through the SEL configuration method.

[0082] like Figure 2 As shown, the SEL configuration method includes:

[0083] The training results of the basic learning models are sorted by accuracy, and the first "n" (the number of basic learning models) configurations of the basic learning models ranked from high to low in terms of accuracy are selected in turn. The value range of "n" is 1 to 5. There are 5 configuration schemes at the basic learning model level, corresponding to 5 different basic learning models, for a total of 25 configuration schemes.

[0084] The classification accuracy of different configuration schemes was compared, and the configuration scheme with the highest accuracy was selected for rubber forest extraction.

[0085] SEL evaluates the accuracy of the basic learning model training results through the confusion matrix and derived parameters, including: overall accuracy OA, Kappa coefficient, user accuracy UA, producer accuracy PA and F1 score; the confusion matrix is ​​used to calculate the user accuracy and mapping accuracy of the rubber forest, and the F1 score of the rubber forest is calculated:

[0086]

[0087] Among them, β represents the weight relationship between UA and PA. When β is set to 1, it means that the weights of UA and PA are equal.

[0088] The confusion matrix is ​​constructed through the model classification results and the number of validation sample points. In order to effectively evaluate the generalization ability of the model and reduce the accidental errors caused by single data division, the five-fold cross-validation method is used to randomly divide the sample points into five equal-sized subsets, and one of them is selected as the validation set in turn, and the rest are used as training sets. This cycle is repeated five times so that each subset can be used as a validation set once. After each round of training, the model performance is evaluated on the validation set, and finally the average of the five validation results is taken as the comprehensive performance indicator of the model.

[0089] Experimental verification:

[0090] Ranking results of different PFT-EM accuracy:

[0091] The statistical results show that PFT-EM based on GTB (PFT-GTB) has the highest accuracy, with an F1 score of 0.945 and an OA of 92.72%. The highest F1 score of PFT-MaxEnt is 0.931 and the highest OA is 91.80%; the highest F1 score of PFT-RF is 0.915 and the highest OA is 92.16%; the highest F1 score of PFT-SVM is 0.893 and the highest OA is 87.05%; the highest F1 score of PFT-CART is 0.877 and the highest OA is 86.63%. We can get the ranking of PFT-EM accuracy from high to low: PFT-GTB>PFT-MaxEnt>PFT-RF>PFT-SVM>PFT-CART.

[0092] SEL model configuration: The accuracy ranking of PFT-EM is obtained. The basic learning model in the SEL model is configured by changing the parameter "n" (the number of basic learning models). The metamodels are RF, MaxEnt, GTB, SVM and CART, and each metamodel corresponds to the configuration of five basic learning models. The average accuracy comparison of 5 years is obtained by SEL training based on this configuration mode. Table 3 shows the impact of "n" on F1 score and OA. From the perspective of parameter "n", that is, the configuration of the basic learning model, when "n" is 4, it represents the configuration of the top 4 PFT-EM combinations in the accuracy ranking. SEL obtains the highest average accuracy, with an F1 score of 0.920 and an OA of 92.57%. From the perspective of metamodel configuration, SEL with RF as the metamodel obtains the highest average accuracy, with an F1 score of 0.917 and an OA of 92.10%. When RF is used as the meta-model and the basic learning model is a combination of PFT-GTB, PFT-MaxEnt, PFT-RF and PFT-SVM (n=4), SEL can achieve the highest accuracy, with an F1score of 0.932 and an OA of 93.70%. Therefore, the configuration with RF as the meta-model and n=4 (denoted as RF-SEL (n=4)) will be selected for the classification and extraction of rubber forests.

[0093] Table 3 Effect of meta-model and “n” (number of base learning models) on SEL classification accuracy

[0094]

[0095] Comparison of SEL model and PFT-EM results: The classification accuracy of RF-SEL (n = 4) was compared with that of the basic learning model. Compared with PFT-GTB, which has the highest accuracy in PFT-EM, the F1 score was improved by 0.027 (0.905 VS 0.932), and the OA was 5.70% (88.00% VS 93.70%). Comparison of classification results between RF-SEL (n = 4) and the configured basic learning model ( Figure 3), it can be seen that the classification results of SEL are more spatially consistent with satellite images. Specifically, the PFT-RF basic learning model easily misclassifies natural forests as rubber forests, especially in the forest areas of low independent plots, as shown in partition (a). The PFT-MaxEnt classifier tends to misclassify rubber forests with mountain shadows and narrow valleys as forest land. The performance of the PFT-GTB model is relatively moderate. In areas with fragmented plots, there is still a mixed distribution between cultivated land, natural forests and rubber forests, especially in areas where land types intersect. This phenomenon is more obvious. The PFT-SVM model easily identifies rubber forests in areas with gentle slopes as cultivated land, and also easily identifies cultivated land as forest land. However, in areas with less fragmentation, SVM can accurately distinguish between entire natural forests and rubber forests, as shown in partition (b), and the recognition of water surfaces is also clearer.

[0096] Comparison with the phenology-based method: The accuracy of the results obtained by the RESI-based phenology algorithm was verified using sample points, and the F1 score was slightly lower than that obtained by the proposed two-layer ensemble model (0.928 VS 0.932). The classification results of four sub-regions were selected for visual comparison ( Figure 4 ), Figure 4 c, d show that the results of RESI still have mixed classification of natural forests and rubber forests, and are prone to produce sporadic and fragmented patches. Figures a and b show that when the images in the rubber forest defoliation period (t2) are blocked by clouds or missing, the classification results of the RESI algorithm are incomplete. This is because the phenology-based algorithm is constructed based on images of a specific phenological period and will be seriously affected by the image quality of that phenological period. The image input period of the dual integration model is all-season, and it can make full use of image information from other periods to compensate for and alleviate the image quality problem of a single phenological period.

[0097] Compared with phenology-based methods, machine learning methods have no specific restrictions on the input time period of remote sensing images. Many machine learning algorithms have been applied to the identification of remote sensing rubber forests, such as random forests, decision trees, K-nearest neighbor algorithms, and support vector machines. However, the spectral characteristics of rubber trees vary in different regions. The generalization ability of a single machine learning model is limited for the differences in classification scenarios and vegetation types in different regions. Ensemble learning is widely used in the study of automated classification of remote sensing images. Ensemble learning methods can utilize the complementary information of different machine models and have great potential in improving classification accuracy and generalization. The stacked ensemble learning algorithm uses a high-level meta-classification model to train the output features of the low-level basic model, enhance generalization ability, improve classification accuracy, and reduce the risk of overfitting.

[0098] The first layer of PFT-EM of the proposed model is constructed based on the image data of the rubber forest growth period, which can effectively alleviate the cloud occlusion and missing problems of remote sensing data in specific periods. The study has obtained good results, which is better than the traditional phenology-based method. In the classification and extraction of rubber forests, Stacking ensemble learning has strong performance, with a classification accuracy of F1 score of 0.932 and OA of 93.70%. It shows that the combined multi-source data features and the two-layer ensemble model are helpful to distinguish rubber forests from other land types and realize large-scale extraction of rubber forests. In addition, the model proposed by the invention has better generalization ability and stronger robustness, and the model can also be well applied to multi-classification tasks in other scenarios. Using the model proposed by the invention to classify and extract rubber forests can provide important support for the development of the rubber industry and environmental protection.

[0099] The above-mentioned embodiments only express the specific implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the protection scope of the present application. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the technical solution concept of the present application, and these all belong to the protection scope of the present application.

Claims

1. A rubber forest extraction method based on GEE platform combined with multi-source data and double-layer integrated model, characterized in that: The following steps are involved: Data collection and feature extraction: Multi-source data was acquired and preprocessed through the GEE platform. Google Earth high-definition images and Sentinel-2 satellite images were used in combination with field survey data to construct a rubber forest sample dataset. Multi-dimensional features of the rubber forest sample dataset were extracted. Feature selection: The recursive feature elimination algorithm RFE is used to recursively train the multidimensional data set at each stage for feature selection; The rubber tree phenology is divided into four stages using the inflection points and morphological characteristics of the vegetation index curve, and the images of the four stages are input into the model; A four-period phenological model PFT-EM for multi-category classification decision-making was constructed, and machine learning models were trained according to different phenological stages of rubber. The PFT-EM model is: Among them, P i,c is the probability that pixel i belongs to category c in the jth image, T i is the number of images used to classify pixel i; c represents different categories, ranging from 1 to C, where C is the total number of categories; calculate the average probability P that pixel i belongs to each category c i,c , use function D(i) to make classification decisions; Construct Stacking ensemble model: The Stacking ensemble learning model SEL integrates the output results of the basic learning model by training the meta-learner. SEL includes the basic learning model and the meta-learner. The basic learning model is PFT-EM trained by different machine learning methods. The configuration scheme with the highest accuracy is selected for the extraction of rubber forest through the SEL configuration method.

2. The rubber forest extraction method based on GEE platform combined with multi-source data and double-layer integrated model according to claim 1 is characterized in that, The SEL configuration method includes: Sort the training results of the basic learning models by accuracy, sort the configurations of the basic learning models from high to low according to accuracy, and select them in sequence. The number of basic learning models is 1-5; The classification accuracy of different configuration schemes was compared, and the configuration scheme with the highest accuracy was selected for rubber forest extraction.

3. The rubber forest extraction method based on GEE platform combined with multi-source data and double-layer integrated model according to claim 2 is characterized in that, The SEL evaluates the accuracy of the basic learning model training results through the confusion matrix and derived parameters, and the derived parameters include: overall accuracy OA, Kappa coefficient, user accuracy UA, producer accuracy PA and F1 score; the confusion matrix is ​​used to calculate the user accuracy and mapping accuracy of the rubber forest, and the F1 score of the rubber forest is calculated: Among them, β represents the weight relationship between UA and PA. When β is set to 1, it means that the weights of UA and PA are equal.

4. The rubber forest extraction method based on GEE platform combined with multi-source data and double-layer integrated model according to claim 3 is characterized in that, The confusion matrix is ​​constructed by the model classification results and the number of validation sample points. The sample points are randomly divided into five equal-sized subsets using the five-fold cross-validation method. One of them is selected as the validation set in turn, and the rest are used as training sets. This cycle is repeated five times so that each subset can be used as a validation set once. After each round of training, the model performance is evaluated on the validation set, and finally the average of the five validation results is taken as the comprehensive performance indicator of the model.

5. The rubber forest extraction method based on GEE platform combined with multi-source data and double-layer integrated model according to claim 1 is characterized in that, The RFE algorithm The following steps are involved: Sort all features involved in training according to feature weights or importance scores and remove the lowest ranked features; Keep the remaining features and repeat the training and sorting process, continue to eliminate the features with the lowest weights, until all features have been iteratively screened; The feature set with the highest classification accuracy in the iterative process is selected as the best feature subset.

6. The rubber forest extraction method based on GEE platform combined with multi-source data and double-layer integrated model according to claim 5 is characterized in that, The RFE algorithm performs feature selection specifically including the following steps: The machine learning models in the GEE platform were selected for training, including random forest RF, maximum entropy model MaxEnt, gradient boosting tree GTB and classification regression tree CART; The feature weights or importance scores obtained through training are subjected to RFE algorithms, namely RF-RFE, MaxEnt-RFE, GTB-RFE and CART-RFE algorithms, and the algorithms are compared to select the algorithm with the highest accuracy improvement for feature selection.

7. The rubber forest extraction method based on GEE platform combined with multi-source data and double-layer integrated model according to claim 1 is characterized in that, The multi-source data include: Sentinel-1 synthetic aperture radar SAR, Sentinel-2 multispectral, bioclimatic and DEM data; Sentinel-1 synthetic aperture radar SAR data comes from ground range detection products. The data spatial resolution is 10 meters. The data is a first-level image dataset obtained through Doppler centroid estimation, single-view composite focusing and post-processing. The dataset images are pre-processed by the Sentinel-1 toolkit, including speckle filtering, thermal noise removal, terrain correction and radiation correction. Sentinel-2 multispectral data uses 2A-level surface reflectance data products obtained after atmospheric and orthorectification correction. The blue, green, red and near-infrared bands have a spatial resolution of 10m, the red edge / shortwave-infrared band has a spatial resolution of 20m, and the atmospheric correction band has a spatial resolution of 60m. The images are pre-processed in time and space by screening, splicing, cloud removal, and cropping. The bioclimatic data uses the "WorldClim version 1" dataset with a spatial resolution of 927.67 meters. This dataset contains information representing annual trends: annual average temperature, annual precipitation, seasonality: annual temperature difference and precipitation, and extreme or limiting environmental factors; The DEM data used is the Shuttle Radar Topography Mission V3 product with a spatial resolution of 30 meters, provided by NASA. All raster data were resampled to 10 m using bilinear interpolation, and the geographic coordinate system was unified to WGS_1984.

8. The rubber forest extraction method based on GEE platform combined with multi-source data and double-layer integrated model according to claim 1 or 7, characterized in that, The multi-dimensional features include backscattering features BS, spectral features SP, texture features TX, bioclimatic features BC and terrain features T0.

9. The rubber forest extraction method based on GEE platform combined with multi-source data and double-layer integrated model according to claim 1 is characterized in that, The four stages of rubber forest phenology include: The pre-defoliation period t1 means that the crown leaves change from dark green to light yellow; The leaf-falling stage t2 means that the leaves change from light yellow to dark yellow until they fall off; The initial stage of leaf expansion, t3, means the beginning of germination until new green leaves grow; The intensive leaf expansion period t4 indicates that the new leaves have fully grown, the forest canopy is lush, and the leaves are dark green.