A regional tree number monitoring method based on a cross-scale nested model

By combining a cross-scale nested model with UAV deep learning and satellite remote sensing machine learning, the challenges of high precision, low cost, and full coverage in tree number monitoring in existing technologies have been solved, enabling dynamic monitoring at the township and county levels and improving monitoring accuracy and model interpretability.

CN122176523APending Publication Date: 2026-06-09SHENYANG INST OF APPL ECOLOGY CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENYANG INST OF APPL ECOLOGY CHINESE ACAD OF SCI
Filing Date
2026-03-23
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing tree count and inversion technologies are insufficient to meet the requirements of high accuracy, low cost, and full coverage at the regional scale. Traditional methods are time-consuming and labor-intensive, while single remote sensing technologies are costly and have low accuracy. Existing models lack interpretability and adaptability, and cannot meet the dynamic monitoring needs at the township and county levels.

Method used

A cross-scale nested model is adopted, combining UAV deep learning and satellite remote sensing machine learning. The UAV acquires sample plot labels, integrates spectral, texture and spatial pattern features, and constructs a random forest regression model to achieve seamless scale transformation from sample plots to regions. The interpretability of the model is enhanced through the SHAP interpretation mechanism.

Benefits of technology

It significantly improves the inversion accuracy and model generalization ability of sparse forest areas, provides a reliable and interpretable regional tree resource dynamic monitoring scheme, reduces costs and improves monitoring efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122176523A_ABST
    Figure CN122176523A_ABST
Patent Text Reader

Abstract

This invention discloses a method for monitoring the number of trees in a region based on a cross-scale nested model. The method includes: regional partitioning and quadrat layout based on satellite remote sensing data; acquiring UAV imagery of the quadrats, automatically identifying and counting the number of trees within each quadrat using a deep learning object detection model, and generating quadrat labels; extracting the corresponding satellite multidimensional features of the quadrats and aggregating them to the quadrat scale; constructing and training a machine learning regression model using the quadrat feature vector as the independent variable and the quadrat label as the dependent variable; and applying the trained model to global satellite feature data to achieve spatial inversion and mapping of the number of trees in the region. This invention, through a two-layer nested architecture of UAV deep learning and satellite machine learning, achieves a scale conversion from high-precision quadrat calibration to high-efficiency regional inversion, solving the problems of high cost and low efficiency of traditional survey methods, as well as insufficient accuracy of single remote sensing methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of forestry resource remote sensing monitoring technology, specifically to a method for monitoring the number of trees in a region based on a cross-scale nested model. Background Technology

[0002] Tree number and spatial density are core benchmark indicators for forest ecosystem management, ecological resource assessment, and evaluation of the effectiveness of ecological engineering construction. Accurately grasping the number and distribution characteristics of individual trees at the regional scale is a crucial prerequisite for scientifically assessing the windbreak and sand-fixing ecological effectiveness of windbreak forest systems, accurately calculating forest carbon sequestration, and formulating precise tending and ecological restoration plans. Given the current needs for digital transformation of forestry and grassland resources in my country, precise management of key ecological projects such as the Three-North Shelterbelt Project, and precise measurement of terrestrial ecosystem carbon sequestration under the "dual carbon" target, achieving high-precision, high-efficiency, and low-cost dynamic inventory of forest resources at the township and county levels has become a critical technical issue urgently needing to be addressed in the field of forestry ecological management.

[0003] However, existing tree count and inversion technologies generally suffer from a core problem of imbalance between efficiency, accuracy, and cost when applied at the regional scale. They struggle to meet the dual needs of large-scale, comprehensive coverage and refined, precise monitoring. The core shortcomings are concentrated in the following aspects:

[0004] First, traditional ground-based survey methods are ill-suited to regional-scale monitoring needs. Traditional methods of manual quadrat measurement and visual interpretation rely heavily on manual fieldwork, which is not only time-consuming, labor-intensive, and costly, but also limited by the complexity of arid terrain and poor accessibility in forest areas, making it difficult to achieve full coverage of the study area. At the same time, the sample size obtained manually is limited, which cannot support the massive and representative samples required for large-scale inversion models, and there are subjective errors in manual visual interpretation, making it difficult to meet the needs of routine and high-frequency dynamic monitoring.

[0005] Secondly, there is an irreconcilable scale contradiction in single remote sensing monitoring technologies. Although UAV remote sensing can acquire sub-decimeter-level high-resolution images and achieve accurate identification of individual trees, the coverage area of ​​a single operation is limited. In large-scale scenarios at the township and county levels, the time and economic costs of full-area aerial surveys are extremely high, making large-scale promotion impossible. On the other hand, while medium-resolution satellite remote sensing data such as Sentinel-2 have the advantages of full-area coverage, free access, and high-frequency updates, their 10m spatial resolution pixels often contain mixed spectral signals from multiple trees. In arid and sparsely forested areas, the spectral signals are easily interfered with by non-vegetation elements such as bare ground and weeds. The generalization ability of single-tree identification or tree count inversion models based directly on satellite imagery is weak and the estimation accuracy is low, which cannot meet the accuracy requirements of grassroots forestry management.

[0006] Furthermore, existing tree count inversion models have technical shortcomings and are difficult to adapt to the monitoring needs of heterogeneous forest stands. Firstly, existing inversion models mostly focus on the statistical characteristics of spectral bands and conventional vegetation indices, generally neglecting key parameters reflecting the growth environment and spatial distribution patterns of trees, such as stand texture index, forest edge spatial pattern, and topographic site conditions. However, the tree distribution in arid and sparsely forested areas exhibits strong spatial heterogeneity and significant edge effects. Models driven by single spectral features show severely uneven estimation accuracy under different stand density gradients, easily leading to underestimation in sparse areas and underestimation in dense areas. Secondly, the training of existing machine learning inversion models is highly dependent on human intervention. The "ground truth" label obtained through field measurements suffers from low sample acquisition efficiency, high cost, and limited sample size, resulting in insufficient robustness and spatial generalization ability of the model, and poor adaptability to different study areas. Thirdly, although existing technologies have attempted to integrate UAV and satellite remote sensing data, a mature and standardized two-layer nested coupling technology system has not been formed. The scale matching degree between the high-precision calibration results of UAVs and the global features of satellites is insufficient, and the problem of error accumulation is prominent during cross-scale conversion, making it impossible to achieve a robust leap from "precise calibration of quadrat points" to "global inversion of the entire region".

[0007] Furthermore, existing plant number inversion models generally suffer from a "black box effect," lacking a robust mechanism for interpreting the models. Most studies focus only on the model's fitting accuracy and prediction results, failing to quantitatively analyze the contribution and influence mechanism of each input feature on plant number prediction. This makes it impossible to verify whether the model captures the correct ecological laws, resulting in difficulties in ensuring the scientific validity and reliability of the prediction results and failing to provide rigorous scientific support for forestry management decisions. Summary of the Invention

[0008] To address the aforementioned technical problems, this invention provides a method for monitoring the number of trees in a region based on a cross-scale nested model.

[0009] The specific details of the invention are as follows:

[0010] A method for monitoring the number of trees in a region based on a cross-scale nested model includes the following steps:

[0011] S1. Regional Zoning and Plot Layout: Based on satellite remote sensing data of the study area, spatial heterogeneity zoning is carried out, and multiple standard-area plots are laid out in each zoning.

[0012] S2, First-level nesting - Automatic labeling of tree count in quadrats: High-resolution UAV images of the corresponding areas of each quadrat are acquired and input into a trained deep learning target detection model to automatically identify and count the number of trees in each quadrat as quadrat labels;

[0013] S3. Satellite Multidimensional Feature Extraction and Aggregation: Based on satellite remote sensing data, multidimensional features of the corresponding regions of each sample plot are extracted. The multidimensional features include at least spectral features, vegetation index, texture features, and spatial pattern features. The extracted pixel-level features are aggregated to the sample plot scale to form a sample plot feature vector.

[0014] S4. Second-layer nested-region inversion model construction and training: Using the sample feature vector as the independent variable and the corresponding sample label as the dependent variable, construct and train a machine learning regression model to establish a nonlinear mapping relationship between sample scale satellite features and tree number.

[0015] S5. Regional Inversion and Mapping: The trained machine learning regression model is applied to the full-domain satellite multidimensional feature data of the study area to predict the number of trees on a unit-by-unit basis and generate a spatial distribution map of the number of trees in the region.

[0016] In step S1, the spatial heterogeneity partitioning specifically involves: calculating vegetation cover based on satellite remote sensing data, and dividing the study area into multiple stand density level zones according to the value range of vegetation cover; the quadrats are deployed in each stand density level zone using a stratified random sampling method.

[0017] In step S2, the deep learning target detection model is the DeepForest model; the automatic identification and statistics specifically include: the DeepForest model outputs the detection bounding box of a single tree in the high-resolution image of the UAV, and the quadrat label is obtained by counting the total number of detection bounding boxes within the geographical range of each quadrat.

[0018] In step S3, the spatial pattern features include edge distance features, which are obtained by calculating the geometric distance from the center point of each cell to the nearest forest edge.

[0019] In step S3, the step of aggregating the extracted pixel-level features to the quadrat scale specifically involves: for all pixels within the geographic range of each quadrat, calculating the arithmetic mean of the features in each dimension, and using the arithmetic mean as the feature value of the quadrat in that dimension.

[0020] In step S4, the machine learning regression model is a random forest regression model; during the training process, the K-fold cross-validation method is used to evaluate the generalization performance of the model.

[0021] Furthermore, step S4 also includes a model interpretation step: after the random forest regression model is trained, the SHAP interpretation mechanism is used to analyze the contribution of the satellite multidimensional features of each dimension to the tree number prediction result.

[0022] Furthermore, in step S4, before constructing the machine learning regression model, the quadrat labels are subjected to square root transformation, and the values ​​after square root transformation are used as the dependent variable for training; before the prediction results are output in step S5, the model prediction values ​​are squared to restore them to the original plant number scale.

[0023] Step S5 is followed by:

[0024] S6. Total Statistical Analysis: Sum the predicted values ​​of all spatial units in the spatial distribution map of the number of trees in the region to obtain an estimated total number of trees in the study area, and mark the estimated value on the spatial distribution map.

[0025] The satellite remote sensing data is Sentinel-2 multispectral image data; the standard area of ​​the sample plot is 1 hectare.

[0026] The beneficial effects of this invention lie in its two-layer nested architecture of UAV deep learning and satellite remote sensing machine learning, which effectively reconciles the contradictions between high accuracy, high efficiency, and full coverage in regional tree surveys. It utilizes UAVs to automate and reduce the cost of obtaining true values ​​from sample plots, replacing traditional manual surveys. Furthermore, by fusing multi-dimensional satellite features such as spectral density, texture, and spatial pattern into a machine learning model, it achieves seamless scale conversion from point calibration to area inversion. This method not only significantly improves the inversion accuracy and model generalization ability in complex scenarios such as sparse forest areas, but its embedded feature contribution interpretation mechanism also enhances the model's interpretability and ecological rationality. Ultimately, it provides grassroots forestry departments with a reliable, cost-effective, and interpretable solution for dynamic monitoring of regional tree resources. Attached Figure Description

[0027] Figure 1 Monitoring method flowchart;

[0028] Figure 2 : Schematic diagram of multi-source data and detection results, including (a) satellite image of the study area and heat map of quadrat distribution and tree number, (b) schematic diagram of the overall distribution of satellite image of the study area, and (c) overlay map of satellite image of the study area and quadrat boundary;

[0029] Figure 3 Map showing the spatial distribution and total number estimation of tree density in the study area;

[0030] Figure 4 : Ranking of variable importance for model input features;

[0031] Figure 5 Summary chart of SHAP values;

[0032] Figure 6 Close-up image of single-tree detection effect from high-resolution drone imagery;

[0033] Figure 7 Figures showing model fitting and cross-validation results, where (a) is the model training set fitting result and (b) is the 5-fold cross-validation result. Detailed Implementation

[0034] The following embodiments illustrate the present invention in detail. In the description of these embodiments, specific details such as particular system structures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail.

[0035] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0036] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0037] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0038] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0039] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0040] Example 1

[0041] Reference Appendix Figure 1 The core of the method lies in constructing two model layers that are closely coupled in terms of spatial scale and function: a single tree detection and calibration layer at the high-resolution scale of UAVs and a feature regression and inversion layer at the medium-resolution scale of satellites.

[0042] The first nested layer, the labeling layer, automatically generates high-precision tree number labels for quadrats from UAV imagery. This process can be formally represented as:

[0043] Let the first The drone imagery corresponding to each sample plot is Using a pre-trained deep learning object detection model The model processes the input image and infers from it, outputting a set of bounding boxes for detected tree targets. :

[0044]

[0045] in, This represents the set of parameters for which the model has been optimized. Then the... Tree count labels for each sample plot This can be obtained by counting the number of bounding boxes within its geographical range:

[0046]

[0047] This formula realizes the mapping from image pixel space to tree count, replacing inefficient manual counting and forming the high-precision data foundation of the entire method.

[0048] The second nested layer, the inversion layer, establishes a generalizable mapping relationship between satellite features and quadrat labels. First, it extracts the quadrat labels from the satellite data. corresponding 3D feature vector This includes spectral, textural, topographical, and spatial pattern features. To stabilize the variance, a square root transformation is performed on the labels to obtain the transformed labels. :

[0049]

[0050] Subsequently, a machine learning regression model was used. Learning from eigenvectors After transformation, the label The model exhibits complex nonlinear relationships. The training process aims to minimize the prediction error.

[0051]

[0052] in, For loss function, This represents the optimal model obtained through training. For any unsampled location within the study area, its corresponding satellite feature vector... The number of trees can be predicted using the following formula. :

[0053]

[0054] This formula expands the scale from limited "point" quadrat labels to full-area "area" plant count estimation. The two-layer model achieves seamless nesting and information transfer by sharing the "quadrat" as a geostatistical unit. Satellite features are aggregated within the quadrat, making its representation scale consistent with the scale of UAV labels, thus forming a complete "point-area" cross-scale monitoring chain.

[0055] Example 2

[0056] Reference Appendix Figure 2 Appendix Figure 7 and appendix Figure 3 The implementation process and intermediate results of the method are shown in detail.

[0057] First, Sentinel-2 satellite imagery of the study area was acquired. After preprocessing, vegetation cover was calculated and stratified accordingly. A stratified random sampling method was used to establish 88 square quadrats (100m x 100m) throughout the town. (See attached reference.) Figure 2 (b) and appendix Figure 2 (c) The overall scope of the study area and the spatial location of the quadrats are clearly visible. Aerial surveys of each quadrat were conducted using drones to acquire orthophotos with a ground resolution better than 5 cm. These images were then input into a deep learning model trained on the DeepForest architecture to automatically identify individual tree canopies. (Appendix) Figure 7The results demonstrate the effect of automatically generated tree detection bounding boxes from drone imagery and models, intuitively verifying the feasibility of automatic calibration of the first-layer nested model.

[0058] Secondly, satellite data is processed in parallel. Multidimensional features are extracted for each quadrat, including 12 original spectral bands, vegetation indices such as NDVI, EVI, NDVI705, and IRECI, texture features calculated from the gray-level co-occurrence matrix, topographic factors such as elevation, slope, and aspect, and a specially calculated distance feature to the nearest forest edge. All 10-meter resolution features are arithmetically averaged and aggregated within the quadrat area to generate an 88×p-dimensional feature matrix X at the quadrat scale. Simultaneously, based on the detection results of the first-layer nested model, a tree count label vector Y is generated for each quadrat. This constitutes the complete dataset used to train the second-layer nested model, as shown in Tables 1 and 2 below.

[0059] Table 1 Sample Plot Data Table 1 - Environmental Background and Spectral Characteristics

[0060]

[0061] Table 2 Sample Plot Data Table 2 - Derivative Index and Sample Plot Response Table

[0062]

[0063] Subsequently, a random forest regression model was constructed in the R language environment as... The label vector Y is transformed by the square root and then input into the model along with the feature matrix X for training. A five-fold cross-validation strategy is used to evaluate the model's generalization performance. (See attached...) Figure 3 As shown in (b), the mean R² of the five-fold cross-validation is 0.766, indicating that the model maintains stable predictive ability across different data subsets. The final model was trained using all 88 samples, and its fit on the training set is shown in the attached figure. Figure 3 As shown in (a), the goodness of fit R² reaches 0.959, proving that the model successfully captures the core association between satellite features and the number of trees.

[0064] Example 3

[0065] This embodiment proposes an adaptive feature enhancement mechanism based on local contribution awareness, building upon the core framework of the aforementioned double-layer nested model, as an innovative optimization of the second-layer nested regression model. This mechanism aims to address the problem of insufficient global model response to local heterogeneity in complex landscape regions; its principle can be found in the appendix. Figure 5 With appendix Figure 6 The differences in the importance of the revealed features and their contribution patterns.

[0066] The core idea of ​​this mechanism is to dynamically adjust the weights of the input feature vectors during the model inference phase, based on the local feature environment of the prediction unit, so that the model's attention is more focused on the features most sensitive to changes in plant number within that local environment. The specific implementation involves two steps:

[0067] The first step is local contribution pattern extraction. After the model training is completed, the SHAP interpretation framework is used to calculate the contribution pattern for each training sample. SHAP value vector ,in Indicates the first Each feature for the sample The contribution of the predicted value. For any new unit feature vector to be predicted. Find its K nearest neighbors in the training set to form a neighborhood set. Calculate the absolute mean of the SHAP values ​​of each feature within the neighborhood, and use it as the feature sensitivity benchmark vector for that local region. :

[0068]

[0069]

[0070] The second step is adaptive feature weight generation and inference. This involves designing a feature weighting algorithm based on the baseline vector. Weight function For example, the Softmax function can be used for normalization to obtain the weight vector. :

[0071]

[0072] in, This is a scaling factor used to control the concentration of the weight distribution. In the final prediction, it is not directly applied... Input Model Instead, it performs a Hadamard product (element-wise multiplication) with the weight vector to obtain the enhanced feature vector. :

[0073]

[0074] The final predicted value is calculated using the following formula:

[0075]

[0076] This mechanism is equivalent to adding an adaptive feature selection and enhancement layer based on the local context before model inference. It allows a globally fixed random forest model to adaptively focus on the most important features in the current local context when applied to different regions. For example, it might focus more on canopy closure-related features within the forest floor, and emphasize edge distance features in the forest edge. This dynamic adjustment capability significantly improves the model's detail capture accuracy and ecological interpretation rationality in areas with strong spatial heterogeneity, and can be continuously optimized through incremental learning.

[0077] Example 4

[0078] Reference Appendix Figure 4 Appendix Figure 5 Appendix Figure 6 and appendix Figure 2 (a) After completing model training in Example 2, an in-depth interpretive analysis of the second-layer nested model is performed. First, an evaluation is conducted based on the variable importance metric built into the Random Forest algorithm. (See attached...) Figure 5 As shown, features such as forest cover, distance from forest edge, and NDVI contribute most significantly to reducing the mean square error of the model, which statistically confirms the key role of integrating spatial pattern information and spectral information in improving the accuracy of tree number inversion.

[0079] To further understand how each feature specifically affects the predicted value, the SHAP framework is used for case-by-case explanation. (Appendix) Figure 6 The SHAP beehive plot visualizes the distribution of SHAP values ​​for each feature, with each point representing a sample. It can be observed that high SHAP values ​​for the forest cover feature are mainly distributed on the right side, indicating that increasing its value has a stable positive driving effect on increasing the predicted number of trees. Conversely, high SHAP values ​​for the distance to the edge feature are mostly concentrated on the left side, indicating that increasing distance usually leads to a decrease in the predicted number of trees. This is highly consistent with the edge effect theory in ecology, greatly enhancing the credibility and interpretability of the model's prediction results. (Appendix) Figure 2 (a) By overlaying the plant number labels of quadrats onto satellite imagery in the form of a heat map, the spatial distribution heterogeneity of the quadrat data itself is visually displayed, providing a real "ground" reference for model training.

[0080] Finally, the trained and validated model was applied to satellite feature raster data covering the entire Zhanggutai Town area, performing pixel-by-pixel prediction to generate a spatial distribution map of tree density at a 100-meter grid scale. (See attached image) Figure 4As shown, the resulting map clearly reveals the spatial aggregation and dispersion patterns of tree resources, with color gradients intuitively reflecting density levels. By summing the predicted values ​​from all effective grid cells across the entire area, a refined estimate of the total number of trees in the study area was obtained: 3,618,033 trees. This total is prominently marked on the distribution map. This achievement represents a leap from limited, automatically labeled quadrat data to clear quantitative mapping of the entire spatial area, providing valuable technical support for the dynamic inventory, conservation planning, and benefit assessment of regional forestry resources.

[0081] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for monitoring the number of trees in a region based on a cross-scale nested model, characterized in that, Includes the following steps: S1. Regional Zoning and Plot Layout: Based on satellite remote sensing data of the study area, spatial heterogeneity zoning is carried out, and multiple standard-area plots are laid out in each zoning. S2, First layer of nesting: Automatic labeling of tree count in quadrats: High-resolution UAV images of the corresponding areas of each quadrat are acquired and input into a trained deep learning target detection model to automatically identify and count the number of trees in each quadrat as quadrat labels; S3. Satellite Multidimensional Feature Extraction and Aggregation: Based on satellite remote sensing data, extract multidimensional features of the corresponding regions of each sample plot. The multidimensional features include at least spectral features, vegetation index, texture features and spatial pattern features. The extracted pixel-level features are aggregated to the quadrat scale to form a quadrat feature vector; S4. Second layer of nesting: Construction and training of regional inversion model: Using the sample plot feature vector as the independent variable and the corresponding sample plot label as the dependent variable, construct and train a machine learning regression model to establish a nonlinear mapping relationship between sample plot scale satellite features and tree number. S5. Regional Inversion and Mapping: The trained machine learning regression model is applied to the full-domain satellite multidimensional feature data of the study area to predict the number of trees on a unit-by-unit basis and generate a spatial distribution map of the number of trees in the region.

2. The method for monitoring the number of trees in a region based on a cross-scale nested model according to claim 1, characterized in that, In step S1, the spatial heterogeneity partitioning specifically involves: calculating vegetation cover based on satellite remote sensing data, and dividing the study area into multiple stand density level zones according to the value range of vegetation cover; the quadrats are deployed in each stand density level zone using a stratified random sampling method.

3. The method for monitoring the number of trees in a region based on a cross-scale nested model according to claim 1, characterized in that, In step S2, the deep learning target detection model is the DeepForest model; the automatic identification and statistics specifically include: the DeepForest model outputs the detection bounding box of a single tree in the high-resolution image of the UAV, and the quadrat label is obtained by counting the total number of detection bounding boxes within the geographical range of each quadrat.

4. The method for monitoring the number of trees in a region based on a cross-scale nested model according to claim 1, characterized in that, In step S3, the spatial pattern features include edge distance features, which are obtained by calculating the geometric distance from the center point of each cell to the nearest forest edge.

5. The method for monitoring the number of trees in a region based on a cross-scale nested model according to claim 1, characterized in that, In step S3, the step of aggregating the extracted pixel-level features to the quadrat scale specifically involves: for all pixels within the geographic range of each quadrat, calculating the arithmetic mean of the features in each dimension, and using the arithmetic mean as the feature value of the quadrat in that dimension.

6. The method for monitoring the number of trees in a region based on a cross-scale nested model according to claim 1, characterized in that, In step S4, the machine learning regression model is a random forest regression model; during the training process, the K-fold cross-validation method is used to evaluate the generalization performance of the model.

7. A method for monitoring the number of trees in a region based on a cross-scale nested model according to claim 6, characterized in that, Step S4 also includes a model interpretation step: after the random forest regression model is trained, the SHAP interpretation mechanism is used to analyze the contribution of the satellite multidimensional features of each dimension to the tree number prediction results.

8. The method for monitoring the number of trees in a region based on a cross-scale nested model according to claim 1, characterized in that, In step S4, before constructing the machine learning regression model, the quadrat labels are subjected to square root transformation, and the values ​​after square root transformation are used as the dependent variable for training; before the prediction results are output in step S5, the model prediction values ​​are squared to restore them to the original plant number scale.

9. A method for monitoring the number of trees in a region based on a cross-scale nested model according to claim 1, characterized in that, Step S5 is followed by: S6. Total Statistical Analysis: Sum the predicted values ​​of all spatial units in the spatial distribution map of the number of trees in the region to obtain an estimated total number of trees in the study area, and mark the estimated value on the spatial distribution map.

10. A method for monitoring the number of trees in a region based on a cross-scale nested model according to any one of claims 1-9, characterized in that, The satellite remote sensing data is Sentinel-2 multispectral image data; the standard area of ​​the sample plot is 1 hectare.