Pinus kesiya forest age remote sensing estimation method for long-time-sequence repeated survey data
By selecting important factors using the Boruta and INLA-SPDE algorithms, a forest age estimation model considering spatiotemporal correlation was constructed, which solved the accuracy problem of long-term repeated survey data, achieved high-precision forest age estimation, and supported scientific forest management.
Patent Information
- Application Number
- CN202511686648.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-17
AI Technical Summary
Existing forest age estimation models suffer from weak generalization and accuracy due to the complexity of forest structure and the influence of external conditions on remote sensing spectra. The spatiotemporal correlation of long-term repeated survey data is difficult to be effectively processed by traditional statistical regression algorithms, resulting in low estimation accuracy.
The Boruta algorithm was used to select important factors for forest age estimation, and the INLA-SPDE algorithm was used to consider the spatiotemporal correlation of the data. A forest age estimation model was constructed by combining random forest and INLA-SPDE algorithm, and the best method was selected by ten-fold cross-validation.
It significantly improves the accuracy of forest age estimation, especially when using multi-year sample data, enabling stable and efficient large-scale forest age estimation and providing scientific data support for forest management.
Smart Images

Figure CN121542622A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of forest age estimation technology, and in particular to a remote sensing estimation method for the age of Pinus simaoensis forests based on long-term repeated survey data. Background Technology
[0002] Forest age is a key factor affecting the carbon storage of an ecosystem. Accurately estimating forest age information is crucial for a proper understanding of the ecosystem's carbon cycle and carbon balance, and for developing scientific forest management plans.
[0003] Constructing regional-scale forest age estimation models using remote sensing imagery and measured plots is a primary technical means for achieving large-scale non-destructive forest monitoring. However, the complexity of forest structure and the non-stationarity of remote sensing spectra during ground object detection, influenced by external conditions such as the atmosphere and solar radiation, lead to weak generalization and accuracy of forest age estimation models. Long-term Landsat imagery and repeatedly surveyed plot data can provide more stable ground object information than single-period data; however, these data exhibit significant spatiotemporal correlations, and traditional statistical regression algorithms struggle to handle the non-independence issues arising from these correlations, resulting in lower accuracy in the constructed models.
[0004] Therefore, there is an urgent need for a forest age estimation method that can take into account the spatiotemporal correlation of long-term repeated survey data in order to achieve high-precision estimation of regional forest age. Summary of the Invention
[0005] This invention provides a remote sensing method for estimating the age of *Pinus simonii* forests based on long-term repeated survey data. For age datasets composed of different survey time spans, the Boruta algorithm is used to select age estimation variables for long-term Landsat imagery and climate, topography, and soil datasets. The INLA-SPDE algorithm is then used to further screen for factors significant and important for age estimation, taking into account the spatiotemporal correlation of the data. Finally, a forest age estimation model is constructed using random forest and the INLA-SPDE algorithm. The accuracy evaluation results of ten-fold cross-validation are used to select the optimal method for estimating the age of *Pinus simonii* forests.
[0006] To achieve the above objectives, this invention provides a remote sensing estimation method for the age of *Pinus simaoensis* forests based on long-term repeated survey data. The method includes the following steps:
[0007] S1: Acquire multi-period forest inventory data, Landsat multispectral remote sensing data, and environmental data consisting of climate, topography, and soil, and complete the corresponding matching of the three types of data according to the year and location of the sample plot survey;
[0008] S2: Perform log10 transformation and standard deviation standardization on the matched forest age data, and generate long-term repeated survey data scenarios with different spans by combining the survey years;
[0009] S3: Based on various data combination scenarios, the Boruta algorithm is used to initially screen important factors for forest age estimation, and then the INLA-SPDE algorithm is used to further screen variables considering spatiotemporal correlation.
[0010] S4: Based on the carefully selected variables, construct random forest and INLA-SPDE age estimation models, determine the optimal model through accuracy evaluation, and draw the age distribution map of the study area.
[0011] Optionally, step S1: Acquire multi-period forest inventory data, Landsat multispectral remote sensing data, and environmental data consisting of climate, topography, and soil. Complete the matching of these three types of data according to the year and location of the sample plot survey. Specifically, this includes:
[0012] S11: Collect data from multiple consecutive forest inventory surveys and extract the forest age, survey year, sample plot latitude and longitude coordinates, and unique plot number from each record;
[0013] S12: Based on the GEE platform, retrieve Landsat multispectral remote sensing data corresponding to each survey year, and perform cloud removal and cloud shadow masking preprocessing on the images.
[0014] S13: Obtain climate data, topographic data, and layered soil data from open-source databases respectively, and complete the standardization of environmental data format;
[0015] S14: Using ArcMap, with the sample plot coordinates as the data collection reference, the preprocessed remote sensing data, standardized environmental data, and survey data are accurately matched according to the year-location dimension, and the matching dataset is output.
[0016] Optionally, in step S11, the collected multi-period forest continuous survey data is configured to meet the following conditions: a first condition of including survey records from at least two different years; a second condition of data fields including forest age, survey year, sample plot latitude and longitude coordinates, unique plot number, and basic site conditions of the sample plot; and a third condition of outlier removal processing.
[0017] Optionally, step S12: Based on the GEE platform, retrieve Landsat multispectral remote sensing data corresponding to each survey year, and perform cloud removal and cloud shadow masking preprocessing on the images, specifically including:
[0018] S121: Select Landsat images with cloud cover below 10% in the GEE platform; the selected Landsat images must include the original bands Blue, Green, Red, NIR, SWIR1, and SWIR2.
[0019] S122: Using the QA quality band of Landsat imagery combined with the CFmask algorithm, interference information such as clouds, cloud shadows, and snow in the imagery is masked and removed.
[0020] S123: Call GEE to perform median synthesis on the masked annual image set to generate annual remote sensing images corresponding to the survey year, calculate derived data including vegetation index and texture features, and supplement the remote sensing dataset with the derived data.
[0021] Optionally, step S13: Obtain climate data, topographic data, and layered soil data from open-source databases respectively, and complete the environmental data format standardization, specifically including:
[0022] S131: Obtain climate data from the WorldClim website; wherein, the climate data includes 19 bioclimatic factors under the current climate scenario and habitat suitability data of Pinus simaoensis calculated based on these factors;
[0023] S132: Obtain resolution terrain data from geospatial data cloud and process it using SAGA software to generate 11 commonly used terrain representation data;
[0024] S133: Obtain soil data from HWSD and extract core indicators of soil texture and organic matter content at the 0-30cm and 30-100cm locations;
[0025] S134: Convert climate data, topographic data, and layered soil data into a unified spatial resolution and projection coordinate system consistent with Landsat imagery.
[0026] Optionally, step S2: Perform log10 transformation and standard deviation standardization on the matched forest age data, and generate long-term repeated survey data scenarios with different spans according to the survey years, specifically including:
[0027] S21: Convert the output matching dataset into CSV format, and use Excel software to perform the log10 function on the forest age field to obtain logarithmic forest age data;
[0028] S22: The logarithmic forest age data was processed using the standard deviation standardization formula, which is as follows:
[0029] ;
[0030] Where, x i Here are the measured forest ages at the sample points, and n is the number of sample plots. S represents the mean age of the sample plots in a specific year, and S represents the standard deviation of the age of the sample plots in a specific year.
[0031] S23: Arrange all standardized data in ascending order by survey year and combine them into a basic CSV file containing all survey periods;
[0032] S24: Delete the latest data for the survey year from the base CSV file one by one to generate scenario data with different survey spans from multiple periods to two periods. Each scenario is stored in an independent CSV file.
[0033] Optionally, step S3: Based on each data combination scenario, the Boruta algorithm is used to initially screen important factors for forest age estimation, and then the INLA-SPDE algorithm is used to refine the variable screening considering spatiotemporal correlation. Specifically, this includes:
[0034] S31: For each data combination scenario, use the Boruta algorithm to traverse remote sensing factors and environmental factors, calculate the importance score of each factor to forest age estimation, and screen out the set of important factors with scores higher than the threshold.
[0035] S32: Using standardized forest age data as the dependent variable and important factor set as the independent variable, construct the INLA-SPDE mixed effect model, use the iid sub-model to handle the spatial location correlation of the sample plots, use the ar sub-model to handle the time series correlation of the survey year, and conduct a significance test on the important factor set;
[0036] S33: Retain the factors that pass the significance test (P<0.05) to form the final variable set for forest age estimation.
[0037] Optionally, step S31: For each data combination scenario, the Boruta algorithm is used to iterate through remote sensing factors and environmental factors, calculate the importance score of each factor for forest age estimation, and screen out the set of important factors with scores higher than the threshold, specifically including:
[0038] S311: Build a screening environment based on the Boruta package in R language, and set the algorithm iteration count to 1000 and the confidence level to 95%;
[0039] S312: For each dataset of data combination scenarios, remote sensing factors and environmental factors are used as input variables, and forest age-standardized data are used as response variables;
[0040] S313: Randomly generate shadow variables and compare them with the original variables, calculate the importance Z-score of each original variable, filter out factors with Z-scores higher than the maximum value of the shadow variables, and form an important factor set corresponding to each scenario;
[0041] S314: Summarize the important factors of all scenarios, eliminate the accidental factors that only appear in a single scenario, and retain the core important factors that appear stably across scenarios for subsequent fine screening.
[0042] Optionally, step S32: Using standardized forest age data as the dependent variable and the important factor set as the independent variable, construct an INLA-SPDE mixed-effects model. The iid sub-model handles the spatial correlation of sample plots, and the ar sub-model handles the time-series correlation of the survey years. A significance test is performed on the important factor set, specifically including:
[0043] S321: Based on the INLA package and SPDE extension module of R language, construct a hybrid model framework that includes spatial effects, time effects and fixed effects;
[0044] S322: Standardized forest age data is used as the dependent variable in the model, the core important factor set is used as the fixed effect independent variable, the latitude and longitude coordinates of the sample plots are mapped to spatial random effects, and the survey year sequence is mapped to temporal random effects;
[0045] S323: The iid sub-model assumes that the spatial correlation of sample plots follows an independent and identically distributed distribution, and the ar sub-model assumes that the annual temporal correlation follows a first-order autoregressive distribution. The model parameters are then solved iteratively.
[0046] S324: Based on the factor significance P-values output by the model, remove insignificant factors with P≥0.05 to form the final variable set, so as to stabilize the linear or nonlinear association between variables and forest age.
[0047] Optionally, step S4: Based on the finely screened variables, construct random forest and INLA-SPDE age estimation models, determine the optimal model through accuracy evaluation, and draw an age distribution map of the study area, specifically including:
[0048] S41: Using the final set of variables as independent variables and standardized forest age data as dependent variables, construct random forest models and INLA-SPDE models respectively;
[0049] S42: Using the ten-fold cross-validation method, the matching dataset is randomly divided into training and validation sets. The model training and validation are completed in 10 iterations. The fitting coefficient R² and root mean square error RMSE are calculated for each validation.
[0050] S43: Determine the optimal forest age estimation model from the two models based on the criteria of maximizing the mean R² and minimizing the mean RMSE after 10 validations.
[0051] S44: Apply the optimal forest age estimation model to remote sensing data and environmental data of the entire study area, calculate the forest age estimate pixel by pixel, and draw a spatial distribution map of the forest age of Pinus simaoensis in the study area using geographic information system tools.
[0052] The beneficial effects of this invention are as follows:
[0053] To improve the accuracy of regional forest age estimation, this invention constructs a forest age estimation method based on long-term repeated survey data that considers the spatiotemporal correlation of such data. The accuracy of *Pinus simonii* forest age estimation using this method is significantly improved compared to traditional regression statistical methods. Furthermore, it was found that the accuracy of forest age estimation is closely related not only to the model estimation method but also to the time span of the sample data used in the estimation; using more years of survey data yields better estimation accuracy. In addition, this invention uses open-source datasets widely used in related studies, ensuring high data availability and low cost. In conclusion, the forest age estimation method proposed in this invention, which considers the spatiotemporal correlation of long-term repeated survey data, can stably, efficiently, and cost-effectively conduct large-scale *Pinus simonii* forest age estimation, providing accurate data support for the scientific formulation of regional forest management plans. Attached Figure Description
[0054] Figure 1 This is a flowchart illustrating a remote sensing method for estimating the age of Pinus simaoensis forests based on long-term repeated survey data.
[0055] Figure 2 A schematic diagram illustrating the forest age estimation error under different sample combination scenarios;
[0056] Figure 3 Scatter plots of forest age estimates under different estimation models;
[0057] Figure 4 This is a schematic diagram showing the estimated age of Simao pine forests in Pu'er City in 2017. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0059] This invention provides a remote sensing estimation method for the age of *Pinus simaoensis* forests based on long-term repeated survey data, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the remote sensing estimation method for the age of Pinus simaoensis forests based on long-term repeated survey data, according to an embodiment of the present invention.
[0060] In this embodiment, a remote sensing estimation method for the age of *Pinus simaoensis* forests based on long-term repeated survey data is provided, the method comprising the following steps:
[0061] S1: Acquire multi-period forest inventory data, Landsat multispectral remote sensing data, and environmental data consisting of climate, topography, and soil, and complete the corresponding matching of the three types of data according to the year and location of the sample plot survey;
[0062] S2: Perform log10 transformation and standard deviation standardization on the matched forest age data, and generate long-term repeated survey data scenarios with different spans by combining the survey years;
[0063] S3: Based on various data combination scenarios, the Boruta algorithm is used to initially screen important factors for forest age estimation, and then the INLA-SPDE algorithm is used to further screen variables considering spatiotemporal correlation.
[0064] S4: Based on the carefully selected variables, construct random forest and INLA-SPDE age estimation models, determine the optimal model through accuracy evaluation, and draw the age distribution map of the study area.
[0065] In a preferred embodiment, step S1: acquire multi-period forest continuous survey data, Landsat multispectral remote sensing data, and environmental data consisting of climate, topography, and soil; and complete the corresponding matching of the three types of data according to the year and location of the sample plot survey, specifically including:
[0066] S11: Collect data from multiple consecutive forest inventory surveys and extract the forest age, survey year, sample plot latitude and longitude coordinates, and unique plot number from each record;
[0067] S12: Based on the GEE platform, retrieve Landsat multispectral remote sensing data corresponding to each survey year, and perform cloud removal and cloud shadow masking preprocessing on the images.
[0068] S13: Obtain climate data, topographic data, and layered soil data from open-source databases respectively, and complete the standardization of environmental data format;
[0069] S14: Using ArcMap, with the sample plot coordinates as the data collection reference, the preprocessed remote sensing data, standardized environmental data, and survey data are accurately matched according to the year-location dimension, and the matching dataset is output.
[0070] In a preferred embodiment, in step S11, the collected multi-period forest continuous survey data is configured to meet the following conditions: a first condition of including survey records from at least two different years; a second condition of data fields including forest age, survey year, sample plot latitude and longitude coordinates, unique plot number, and basic site conditions of the sample plot; and a third condition of outlier removal processing.
[0071] In a preferred embodiment, step S12: Based on the GEE platform, retrieve Landsat multispectral remote sensing data corresponding to each survey year, and perform cloud removal and cloud shadow masking preprocessing on the images, specifically including:
[0072] S121: Select Landsat images with cloud cover below 10% in the GEE platform; the selected Landsat images must include the original bands Blue, Green, Red, NIR, SWIR1, and SWIR2.
[0073] S122: Using the QA quality band of Landsat imagery combined with the CFmask algorithm, interference information such as clouds, cloud shadows, and snow in the imagery is masked and removed.
[0074] S123: Call GEE to perform median synthesis on the masked annual image set to generate annual remote sensing images corresponding to the survey year, calculate derived data including vegetation index and texture features, and supplement the remote sensing dataset with the derived data.
[0075] In a preferred embodiment, step S13: Obtain climate data, topographic data, and layered soil data from open-source databases respectively, and complete the standardization of environmental data format, specifically including:
[0076] S131: Obtain climate data from the WorldClim website; wherein, the climate data includes 19 bioclimatic factors under the current climate scenario and habitat suitability data of Pinus simaoensis calculated based on these factors;
[0077] S132: Obtain resolution terrain data from geospatial data cloud and process it using SAGA software to generate 11 commonly used terrain representation data;
[0078] S133: Obtain soil data from HWSD and extract core indicators of soil texture and organic matter content at the 0-30cm and 30-100cm locations;
[0079] S134: Convert climate data, topographic data, and layered soil data into a unified spatial resolution and projection coordinate system consistent with Landsat imagery.
[0080] In a preferred embodiment, step S2: The matched forest age data undergoes log10 transformation and standard deviation standardization to generate long-term repeated survey data scenarios with different spans, based on the survey years. Specifically, this includes:
[0081] S21: Convert the output matching dataset into CSV format, and use Excel software to perform the log10 function on the forest age field to obtain logarithmic forest age data;
[0082] S22: The logarithmic forest age data was processed using the standard deviation standardization formula, which is as follows:
[0083] (1)
[0084] Where, x i Here are the measured forest ages at the sample points, and n is the number of sample plots. S represents the mean age of the sample plots in a specific year, and S represents the standard deviation of the age of the sample plots in a specific year.
[0085] S23: Arrange all standardized data in ascending order by survey year and combine them into a basic CSV file containing all survey periods;
[0086] S24: Delete the latest data for the survey year from the base CSV file one by one to generate scenario data with different survey spans from multiple periods to two periods. Each scenario is stored in an independent CSV file.
[0087] In a preferred embodiment, step S3: Based on each data combination scenario, the Boruta algorithm is used to initially screen important factors for forest age estimation, and then the INLA-SPDE algorithm is used to refine the variable screening considering spatiotemporal correlation. Specifically, this includes:
[0088] S31: For each data combination scenario, use the Boruta algorithm to traverse remote sensing factors and environmental factors, calculate the importance score of each factor to forest age estimation, and screen out the set of important factors with scores higher than the threshold.
[0089] S32: Using standardized forest age data as the dependent variable and important factor set as the independent variable, construct the INLA-SPDE mixed effect model, use the iid sub-model to handle the spatial location correlation of the sample plots, use the ar sub-model to handle the time series correlation of the survey year, and conduct a significance test on the important factor set;
[0090] S33: Retain the factors that pass the significance test (P<0.05) to form the final variable set for forest age estimation.
[0091] In a preferred embodiment, step S31: For each data combination scenario, the Boruta algorithm is used to traverse remote sensing factors and environmental factors, calculate the importance score of each factor for forest age estimation, and screen out the set of important factors with scores higher than a threshold, specifically including:
[0092] S311: Build a screening environment based on the Boruta package in R language, and set the algorithm iteration count to 1000 and the confidence level to 95%;
[0093] S312: For each dataset of data combination scenarios, remote sensing factors and environmental factors are used as input variables, and forest age-standardized data are used as response variables;
[0094] S313: Randomly generate shadow variables and compare them with the original variables, calculate the importance Z-score of each original variable, filter out factors with Z-scores higher than the maximum value of the shadow variables, and form an important factor set corresponding to each scenario;
[0095] S314: Summarize the important factors of all scenarios, eliminate the accidental factors that only appear in a single scenario, and retain the core important factors that appear stably across scenarios for subsequent fine screening.
[0096] In a preferred embodiment, step S32: using standardized forest age data as the dependent variable and the important factor set as the independent variable, an INLA-SPDE mixed-effects model is constructed. The spatial location correlation of the sample plots is handled through the iid sub-model, and the time series correlation of the survey year is handled through the ar sub-model. A significance test is performed on the important factor set, specifically including:
[0097] S321: Based on the INLA package and SPDE extension module of R language, construct a hybrid model framework that includes spatial effects, time effects and fixed effects;
[0098] S322: Standardized forest age data is used as the dependent variable in the model, the core important factor set is used as the fixed effect independent variable, the latitude and longitude coordinates of the sample plots are mapped to spatial random effects, and the survey year sequence is mapped to temporal random effects;
[0099] S323: The iid sub-model assumes that the spatial correlation of sample plots follows an independent and identically distributed distribution, and the ar sub-model assumes that the annual temporal correlation follows a first-order autoregressive distribution. The model parameters are then solved iteratively.
[0100] S324: Based on the factor significance P-values output by the model, remove insignificant factors with P≥0.05 to form the final variable set, so as to stabilize the linear or nonlinear association between variables and forest age.
[0101] In a preferred embodiment, step S4: Based on the finely screened variables, construct random forest and INLA-SPDE age estimation models, determine the optimal model through accuracy evaluation, and draw an age distribution map of the study area, specifically including:
[0102] S41: Using the final set of variables as independent variables and standardized forest age data as dependent variables, construct random forest models and INLA-SPDE models respectively;
[0103] S42: Using the ten-fold cross-validation method, the matching dataset is randomly divided into training and validation sets. The model training and validation are completed in 10 iterations. The fitting coefficient R² and root mean square error RMSE are calculated for each validation.
[0104] S43: Determine the optimal forest age estimation model from the two models based on the criteria of maximizing the mean R² and minimizing the mean RMSE after 10 validations.
[0105] S44: Apply the optimal forest age estimation model to remote sensing data and environmental data of the entire study area, calculate the forest age estimate pixel by pixel, and draw a spatial distribution map of the forest age of Pinus simaoensis in the study area using geographic information system tools.
[0106] It should be noted that constructing regional-scale forest age estimation models using remote sensing imagery and measured sample plots is the main technical means to achieve large-scale non-destructive forest monitoring. However, due to the complexity of forest structure and the non-stationarity of remote sensing spectra during ground object detection, which is easily affected by external conditions such as the atmosphere and solar radiation, the generalization and accuracy of forest age estimation models are relatively weak. Long-term Landsat imagery and repeated survey sample plot data can provide more stable ground object information than single-period data, but these types of data have obvious spatiotemporal correlations. Traditional statistical regression algorithms are difficult to handle the non-independence problems caused by data correlations, thus leading to lower accuracy in the constructed models.
[0107] To address the aforementioned issues, this embodiment proposes a remote sensing method for estimating the age of *Pinus simonii* trees by considering the spatiotemporal correlation of long-term repeated survey data. This method uses the Boruta algorithm to select variables important for age estimation, and then employs the INLA-SPDE algorithm to further refine these important variables by considering the spatiotemporal correlation of long-term repeated survey data, obtaining variables that are both important and significant for age estimation. Furthermore, it selects *Pinus simonii*, an important pioneer tree species in the mountainous region of southwestern China, to construct an age estimation model based on random forest (RF) and the INLA-SPDE algorithm, and selects the fitting coefficient R0. 2 The root mean square error (RMSE) was used as the evaluation index to verify the estimation accuracy. Compared with traditional remote sensing methods for forest age estimation, this invention selects long-term time-series data for variable selection. The variables obtained are more stable than those selected using single-period data. Furthermore, the spatiotemporal correlation between long-term repeated survey data is considered in both variable selection and model construction, ensuring the prediction accuracy and transferability of the method. This provides an efficient and feasible method for large-scale lossless estimation of forest age.
[0108] To explain this application more clearly, a specific application example of a remote sensing estimation method for the age of Pinus simaoensis forests based on long-term repeated survey data is provided below.
[0109] 1. Overview of the study area:
[0110] Pu'er City is located in the southern Hengduan Mountains, with over 90% of its area being mountainous. It has a subtropical plateau monsoon climate, with most areas experiencing no frost throughout the year. Annual precipitation ranges from 1100mm to 2780mm, making it the second largest forest area in Yunnan Province. The dominant tree species in the region is *Pinus simonii*, which covers more than half of the area. This species is mainly found in the low hills surrounding wide valleys and riverbanks, with an altitude ranging from 600 meters to 2000 meters, exhibiting a clear vertical zonation. It is a major afforestation species in southern, central, and western Yunnan Province below 1800 meters in altitude.
[0111] 2. Data collection and processing:
[0112] The measured sample plot data in this embodiment comes from the continuous forest survey conducted in Pu'er City from 1992 to 2017. This survey carried out continuous and periodic investigations of forest resources at 5-year intervals by setting up fixed sample plots of 28.5m*28.5m. The survey recorded information such as the location, plot number, and forest age of the sample plots. Specific sample plot statistics are shown in Table 1.
[0113] Table 1. Statistical information on forest age of fixed sample plots in Pu'er City from 1992 to 2017
[0114]
[0115] The imagery data used for stand age estimation consisted of Landsat Level 2 Tier 1 data with cloud cover below 10% for the corresponding data collection year, acquired using the GEE platform. The images were then filtered for clouds, shadows, and snow using the CFmask algorithm, and median-filled to synthesize annual remote sensing images corresponding to the field survey year, along with corresponding image transformation data. Data representing the regional environment were derived from 19 bioclimatic datasets under the current WorldClim climate scenario and their generated Simao pine habitat suitability data, 30m ASTER GDEM data, and soil data from the Harmonized World Soil Database (HWSD) version 2.0. The specific estimation variables used for stand age estimation and their sources are shown in Table 2.
[0116] Table 2. Estimated variables for the age of Pinus simaoense forests
[0117]
[0118] 3. Selection of estimation variables:
[0119] Six periods of measured data obtained from forest fixed plot surveys conducted every five years from 1992 to 2017 were used to create five estimation sample combinations in descending order of year: i. a combination of six survey periods from 1992 to 2017; ii. a combination of five survey periods from 1992 to 2012; iii. a combination of four survey periods from 1992 to 2007; iv. a combination of three survey periods from 1992 to 2002; and v. a combination of two survey periods from 1992 to 1997. The Boruta package in R was used to select important variables for age estimation for each combination. Subsequently, the INLA-SPDE package, along with the iid and ar sub-models, was used to consider the spatiotemporal correlation of data representing plot location and survey year as covariates in the selection of age estimation variables. This resulted in variables that were significant and important for age estimation over long time series, thus completing the selection of age estimation variables for different survey time spans.
[0120] 4. Estimation of forest age in sample plots:
[0121] Using the selected age estimation variables, age estimation models based on the random forest algorithm and the INLA-SPDE algorithm were constructed for measured sample plot data from 1992 to 2017. The INLA-SPDE algorithm model construction can be divided into two cases: considering the spatiotemporal correlation of the data and not considering it. The model estimation accuracy was validated using a 10-fold cross-validation method, with the validation index including the coefficient of determination (R²). 2 The formulas for calculating the root mean square error (RMSE) are shown in equations (2) and (3).
[0122] (2)
[0123] (3)
[0124] In the formula and y i These are the predicted and measured ages of the sample points. denoted as the average forest age of the measured sample plots within the study area, and n represents the number of sample plots.
[0125] 5. Results and Analysis:
[0126] (1) Selection of forest age variable based on long-term repeated survey data:
[0127] Table 3 shows the selection results of age estimation variables obtained from sample scenarios composed of five different survey years. It can be seen that the Landsat factors important for estimating the age of *Pinus simonii* are mostly image-transformed variables, primarily based on texture information. Climate factors are predominantly the driest month's precipitation and factors reflecting climate change. Topographic information is concentrated on elevation and humidity index (WI). Soil factors include four factors from the topsoil: available water content (AWC), carbon-nitrogen ratio, subsoil coarse debris content, and aluminum saturation. Furthermore, in terms of the number of variables obtained, scenario i, using the most survey years, yielded the most age estimation variables; the other variable sample combinations all showed a reduction in the number of variables compared to scenario i. Meanwhile, the INLA-SPDE algorithm, as a further selection method for important variables obtained from the Boruta algorithm, can reduce the number of variables by at least 10%, even reducing it by 9 variables in scenario i, which has the most sample years.
[0128] Table 3. Results of selecting the age variable for Pinus simaoensis forests
[0129]
[0130] (2) Comparison results of forest age estimation models:
[0131] Table 4 shows the age estimation results obtained by constructing a random forest model using only the Boruta algorithm for age estimation variables and refined variables obtained by further considering the spatiotemporal correlation of data using the INLA-SPDE algorithm. The consideration of spatiotemporal correlation is further divided into two categories: considering only long-term Landsat imagery and considering the influence of all estimation data. Table 4 shows that the model fitting accuracy of the three estimation variable selection methods generally increases with the number of survey years used for the variables. The model with variables considering the spatiotemporal correlation of remote sensing factors has the highest fitting accuracy, followed by the method considering only variable importance, while the model with variables considering the spatiotemporal correlation of all estimation data has the worst fitting result. Therefore, the optimal method for selecting Simao pine age estimation variables is the variable selection result obtained by using the Boruta algorithm on six periods of measured data from 1992 to 2017 and considering the spatiotemporal correlation of long-term Landsat imagery.
[0132] Table 4. Age estimation results of Pinus simonii forests under different variable selection methods
[0133]
[0134] Using the variables obtained from the optimal selection method for Pinus sylvestris age estimation, age estimation models were constructed and their accuracy validated using random forest and INLA-SPDE algorithms on measured data from 1992 to 2017. The results are as follows: Figure 2 As shown, the INLA-SPDE algorithm, which does not consider the spatiotemporal correlation of the data, yields the lowest model estimation accuracy, with age estimation errors exceeding 15 years. In contrast, the INLA-SPDE algorithm, which considers the spatiotemporal correlation of the data, significantly reduces the age estimation error. Specifically, the age estimation using sample combination scenario i reduces the error by nearly half. Furthermore, similar to random forest, this model construction method achieves higher model fitting accuracy with more years of sample combination used. The INLA-SPDE algorithm, which sets the survey time and plot location as random effects in the model, effectively addresses the nonlinearity between age and estimation factors. Moreover, the age estimation model based on the random forest algorithm can achieve high-precision age estimation without considering the spatiotemporal correlation of the data when using repeated survey data from periods 2, 5, and 6 (i.e., sample combination scenarios i, ii, and v). Scatter plots of age estimation using different estimation methods are shown below. Figure 3 It can also be seen that the random forest algorithm is the optimal age estimation scheme using six periods of repeated survey data. This invention uses the optimal variable selection results and estimation model to map the age estimation of *Pinus simonii* forests in Pu'er City in 2017, and the results are as follows: Figure 4 As shown.
[0135] Overall, in the estimation of Simao pine forest age based on long-term repeated survey data, the Boruta method combined with INLA-SPDE, considering the spatiotemporal correlation of Landsat imagery, yielded the best variable selection results for six repeated survey periods. The selected estimation variables covered remote sensing factors, climate, topography, and soil information. Furthermore, the random forest algorithm was the most effective in model building for the longest survey period. The method proposed in this invention meets the requirement for non-destructive and reliable age estimation at the regional scale.
[0136] It is understood that in the description of this specification, references to terms such as "one embodiment," "another embodiment," "other embodiments," or "first embodiment to Nth embodiment," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0137] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0138] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A remote sensing estimation method for the age of *Pinus simaoensis* forests based on long-term repeated survey data, characterized in that, The method includes the following steps: S1: Acquire multi-period forest inventory data, Landsat multispectral remote sensing data, and environmental data consisting of climate, topography, and soil, and complete the corresponding matching of the three types of data according to the year and location of the sample plot survey; S2: Perform log10 transformation and standard deviation standardization on the matched forest age data, and generate long-term repeated survey data scenarios with different spans by combining the survey years; S3: Based on various data combination scenarios, the Boruta algorithm is used to initially screen important factors for forest age estimation, and then the INLA-SPDE algorithm is used to further screen variables considering spatiotemporal correlation. S4: Based on the carefully selected variables, construct random forest and INLA-SPDE age estimation models, determine the optimal model through accuracy evaluation, and draw the age distribution map of the study area.
2. The remote sensing estimation method for the age of *Pinus simaoensis* forests based on long-term repeated survey data as described in claim 1, characterized in that... Step S1: Acquire multi-period forest inventory data, Landsat multispectral remote sensing data, and environmental data consisting of climate, topography, and soil. Match these three types of data according to the year and location of the sample plot survey. Specifically, this includes: S11: Collect data from multiple consecutive forest inventory surveys and extract the forest age, survey year, sample plot latitude and longitude coordinates, and unique plot number from each record; S12: Based on the GEE platform, retrieve Landsat multispectral remote sensing data corresponding to each survey year, and perform cloud removal and cloud shadow masking preprocessing on the images. S13: Obtain climate data, topographic data, and layered soil data from open-source databases respectively, and complete the standardization of environmental data format; S14: Using ArcMap, with the sample plot coordinates as the data collection reference, the preprocessed remote sensing data, standardized environmental data, and survey data are accurately matched according to the year-location dimension, and the matching dataset is output.
3. The remote sensing estimation method for the age of *Pinus simaoensis* forests based on long-term repeated survey data as described in claim 2, characterized in that... In step S11, the collected multi-period forest continuous survey data is configured to meet the following conditions: first condition: it contains survey records from at least two different years; second condition: data fields include forest age value, survey year, sample plot latitude and longitude coordinates, unique plot number, and basic site condition information of the sample plot. The third condition for outlier removal.
4. The remote sensing estimation method for the age of *Pinus simaoensis* forests based on long-term repeated survey data as described in claim 2, characterized in that... Step S12: Based on the GEE platform, retrieve Landsat multispectral remote sensing data corresponding to each survey year, and perform cloud removal and cloud shadow masking preprocessing on the images, specifically including: S121: Select Landsat images with cloud cover below 10% in the GEE platform; the selected Landsat images must include the original bands Blue, Green, Red, NIR, SWIR1, and SWIR2. S122: Using the QA quality band of Landsat imagery combined with the CFmask algorithm, interference information such as clouds, cloud shadows, and snow in the imagery is masked and removed. S123: Call GEE to perform median synthesis on the masked annual image set to generate annual remote sensing images corresponding to the survey year, calculate derived data including vegetation index and texture features, and supplement the remote sensing dataset with the derived data.
5. The remote sensing estimation method for the age of *Pinus simaoensis* forests based on long-term repeated survey data as described in claim 2, characterized in that... Step S13: Obtain climate data, topographic data, and stratified soil data from open-source databases, and standardize the environmental data format, specifically including: S131: Obtain climate data from the WorldClim website; wherein, the climate data includes 19 bioclimatic factors under the current climate scenario and habitat suitability data of Pinus simaoensis calculated based on these factors; S132: Obtain resolution terrain data from geospatial data cloud and process it using SAGA software to generate 11 commonly used terrain representation data; S133: Obtain soil data from HWSD and extract core indicators of soil texture and organic matter content at the 0-30cm and 30-100cm locations; S134: Convert climate data, topographic data, and layered soil data into a unified spatial resolution and projection coordinate system consistent with Landsat imagery.
6. The remote sensing estimation method for the age of *Pinus simaoensis* forests based on long-term repeated survey data as described in claim 1, characterized in that... Step S2: Perform log10 transformation and standard deviation standardization on the matched forest age data, and generate long-term repeated survey data scenarios with different spans by combining survey years, specifically including: S21: Convert the output matching dataset into CSV format, and use Excel software to perform the log10 function on the forest age field to obtain logarithmic forest age data; S22: The logarithmic forest age data was processed using the standard deviation standardization formula, which is as follows: ; Where, x i Here are the measured forest ages at the sample points, and n is the number of sample plots. S represents the mean age of the sample plots in a specific year, and S represents the standard deviation of the age of the sample plots in a specific year. S23: Arrange all standardized data in ascending order by survey year and combine them into a basic CSV file containing all survey periods; S24: Delete the latest data for the survey year from the base CSV file one by one to generate scenario data with different survey spans from multiple periods to two periods. Each scenario is stored in an independent CSV file.
7. The remote sensing estimation method for the age of *Pinus simaoensis* forests based on long-term repeated survey data as described in claim 1, characterized in that... Step S3: Based on various data combinations, the Boruta algorithm is used to initially screen important factors for forest age estimation, and then the INLA-SPDE algorithm is used to refine the variable selection by considering spatiotemporal correlation. Specifically, this includes: S31: For each data combination scenario, use the Boruta algorithm to traverse remote sensing factors and environmental factors, calculate the importance score of each factor to forest age estimation, and screen out the set of important factors with scores higher than the threshold. S32: Using standardized forest age data as the dependent variable and important factor set as the independent variable, construct the INLA-SPDE mixed effect model, use the iid sub-model to handle the spatial location correlation of the sample plots, use the ar sub-model to handle the time series correlation of the survey year, and conduct a significance test on the important factor set; S33: Retain the factors that pass the significance test (P<0.05) to form the final variable set for forest age estimation.
8. The remote sensing estimation method for the age of *Pinus simaoensis* forests based on long-term repeated survey data as described in claim 7, characterized in that... Step S31: For each data combination scenario, use the Boruta algorithm to iterate through remote sensing factors and environmental factors, calculate the importance score of each factor for forest age estimation, and filter out the important factor set with scores higher than the threshold, specifically including: S311: Build a screening environment based on the Boruta package in R language, and set the algorithm iteration count to 1000 and the confidence level to 95%; S312: For each dataset of data combination scenarios, remote sensing factors and environmental factors are used as input variables, and forest age-standardized data are used as response variables; S313: Randomly generate shadow variables and compare them with the original variables, calculate the importance Z-score of each original variable, filter out factors with Z-scores higher than the maximum value of the shadow variables, and form an important factor set corresponding to each scenario; S314: Summarize the important factors of all scenarios, eliminate the accidental factors that only appear in a single scenario, and retain the core important factors that appear stably across scenarios for subsequent fine screening.
9. The remote sensing estimation method for the age of *Pinus simaoensis* forests based on long-term repeated survey data as described in claim 7, characterized in that... Step S32: Using standardized forest age data as the dependent variable and the important factor set as the independent variable, construct an INLA-SPDE mixed-effects model. The iid sub-model handles the spatial correlation of sample plots, and the ar sub-model handles the time-series correlation of the survey years. A significance test is performed on the important factor set, specifically including: S321: Based on the INLA package and SPDE extension module of R language, construct a hybrid model framework that includes spatial effects, time effects and fixed effects; S322: Standardized forest age data is used as the dependent variable in the model, the core important factor set is used as the fixed effect independent variable, the latitude and longitude coordinates of the sample plots are mapped to spatial random effects, and the survey year sequence is mapped to temporal random effects; S323: The iid sub-model assumes that the spatial correlation of sample plots follows an independent and identically distributed distribution, and the ar sub-model assumes that the annual temporal correlation follows a first-order autoregressive distribution. The model parameters are then solved iteratively. S324: Based on the factor significance P-values output by the model, remove insignificant factors with P≥0.05 to form the final variable set, so as to stabilize the linear or nonlinear association between variables and forest age.
10. The remote sensing estimation method for the age of *Pinus simaoensis* forests based on long-term repeated survey data as described in claim 9, characterized in that... Step S4: Based on the refined variables, construct random forest and INLA-SPDE age estimation models, determine the optimal model through accuracy evaluation, and draw an age distribution map of the study area. This includes: S41: Using the final set of variables as independent variables and standardized forest age data as dependent variables, construct random forest models and INLA-SPDE models respectively; S42: Using the ten-fold cross-validation method, the matching dataset is randomly divided into training and validation sets. The model training and validation are completed in 10 iterations. The fitting coefficient R² and root mean square error RMSE are calculated for each validation. S43: Determine the optimal forest age estimation model from the two models based on the criteria of maximizing the mean R² and minimizing the mean RMSE after 10 validations. S44: Apply the optimal forest age estimation model to remote sensing data and environmental data of the entire study area, calculate the forest age estimate pixel by pixel, and draw a spatial distribution map of the forest age of Pinus simaoensis in the study area using geographic information system tools.