Forest tree species identification method and system in combination with unmanned aerial vehicle RGB image and semi-supervised learning
By combining UAV RGB imagery with a pseudo-label teacher-student framework based on semi-supervised learning, this method solves the problem of difficult dataset construction in traditional tree species identification methods, thereby improving the efficiency and accuracy of forest tree species identification. It is applicable to forest ecosystem research and sustainable forestry management.
Patent Information
- Application Number
- CN202510995355.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-31
AI Technical Summary
Traditional tree species identification methods rely on manual field surveys, which make it difficult to complete macro-scale monitoring in a short period of time. Furthermore, high-quality ground annotation is difficult to achieve, which makes it difficult to construct datasets and affects the accuracy and generalization ability of the models.
By combining UAV RGB imagery and semi-supervised learning, and utilizing a pseudo-labeled teacher-student framework, the Efficient-Tree model is jointly trained in both supervised and semi-supervised phases, making full use of unlabeled data to improve the model's robustness and generalization ability.
It improves the accuracy and generalization of forest tree species identification, enabling efficient identification of tree species in heterogeneous forest environments, and is applicable to forest ecosystem research and sustainable forestry management.
Smart Images

Figure CN120877112A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of forest remote sensing monitoring technology, and in particular to a method and system for forest tree species identification that combines UAV RGB imagery and semi-supervised learning. Background Technology
[0002] Trees, as the main body of forest vegetation, are key factors influencing the stability, productivity, and carbon storage of forest ecosystems, and their diversity characterizes the level of forest biodiversity at the stand scale. Identifying tree species at the individual tree level can further refine stand composition and is of great significance for forestry (such as timber supply for specific species), biogeographical assessment (such as climate-induced species distribution changes), or biodiversity monitoring, and is also the primary step in forest monitoring. However, traditional tree species identification mainly relies on manual field surveys. This method is usually limited by the complex terrain of forest areas, making it difficult to complete macro-scale monitoring in the short term, and even more difficult to reflect the spatiotemporal dynamics of species in a timely manner.
[0003] Forest structure can exhibit significant diversity across different locations and time periods. This heterogeneity poses a significant challenge to the development of supervised models, as models may not easily generalize to areas outside the training data distribution range. Therefore, a large dataset containing data from different locations, years, seasons, lighting conditions, and forest stands is needed to train efficient and robust models. While advancements in sensor technology have made acquiring high-resolution aerial imagery more efficient and economical, high-quality ground annotation relies on time-consuming field surveys and visual interpretation. The complex geographical environments and community structures of some forests may limit the acquisition of in-situ data. Furthermore, the consistency and accuracy of data annotation are difficult to guarantee due to complex forest stand structures (such as vertical stratification), intraspecific variation (such as morphological changes in canopy size), and dynamic weather conditions. Therefore, these factors collectively hinder the construction of large-scale datasets, thus limiting the accuracy and generalization ability of supervised models.
[0004] Many studies employ transfer learning strategies to address the limitations of the aforementioned labeled datasets. This strategy uses a "pre-training-fine-tuning" paradigm, learning general feature representations on large-scale datasets (such as crowdsourced datasets like iNaturalist or various forest datasets) and then adapting them to small-scale labeled data to achieve cross-domain knowledge transfer, thereby achieving competitive model performance with far less data than training from scratch. However, the effectiveness of transfer learning largely depends on the representativeness of the source domain dataset used for pre-training. When there is a significant domain gap between the source and target domains, simple fine-tuning is often insufficient to bridge this gap, resulting in limited improvements in model performance. Summary of the Invention
[0005] The purpose of this application is to provide a method and system for forest tree species identification that combines UAV RGB imagery and semi-supervised learning, aiming to make full use of a large amount of unlabeled data to improve the robustness of the model, thereby improving the accuracy and generalization ability of the model.
[0006] To achieve the above objectives, this application provides the following solution:
[0007] In a first aspect, this application provides a forest tree species identification method that combines UAV RGB imagery and semi-supervised learning, wherein the forest tree species identification method that combines UAV RGB imagery and semi-supervised learning includes:
[0008] Acquire drone aerial images of multiple sample areas; each sample area's drone aerial imagery includes multiple forest canopy RGB images.
[0009] Multiple forest canopy RGB images of each sample area were stitched together to obtain digital orthophotos of each sample area.
[0010] The digital orthophotos of each sample region are resampled and rotated to obtain optimized orthophotos of each sample region.
[0011] All optimized orthophotos are divided into a first image set and a second image set according to the proportions, and the geographical location and species of each tree in the corresponding sample area of the first image set are recorded to obtain tree species location record information.
[0012] Based on the first image set and the tree species location record information, the dominant tree species in the sample area corresponding to the first image set are outlined to obtain a sample image annotation set.
[0013] The optimized orthophotos in the second image set are cropped in sections to obtain an unlabeled sample image set;
[0014] An Efficient-Tree model is constructed based on a pseudo-labeled teacher-student framework, and trained using the labeled sample image set and the unlabeled sample image set. The pseudo-labeled teacher-student framework uses the YOLO-Tree model as a baseline detector. The Efficient-Tree model uses the labeled sample image set in the supervised learning phase and introduces the unlabeled sample image set for joint training in the semi-supervised learning phase.
[0015] The trained Efficient-Tree model and SAHI database are used to identify digital orthophotos of the target forest area, and the detection results of each tree species in the target forest area are obtained.
[0016] Secondly, this application also provides a computer system, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the forest tree species identification method combining UAV RGB imagery and semi-supervised learning as described in the first aspect.
[0017] According to the specific embodiments provided in this application, the following technical effects are disclosed:
[0018] This application employs a pseudo-label-based teacher-student framework that fully utilizes a large amount of unlabeled data. Specifically, the Efficient-Tree model does not rely entirely on the labeled image set during training; instead, it uses the labeled image set during supervised training and introduces the unlabeled image set for joint training during semi-supervised learning. The introduction of the unlabeled image set further enhances the robustness of the Efficient-Tree model. Simultaneously, this application uses the YOLO-Tree model as a baseline detector within the pseudo-label-based teacher-student framework, which also strengthens the basic detection capabilities of the Efficient-Tree model, ensuring the quality of pseudo-labels and training stability. Based on these features, the accuracy and generalization ability of the Efficient-Tree model in identifying tree species within the target forest area are significantly improved. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart illustrating a forest tree species identification method combining UAV RGB imagery and semi-supervised learning, provided in an embodiment of this application.
[0021] Figure 2 This is an overall framework diagram of the Efficient-Tree model provided in the embodiments of this application;
[0022] Figure 3 This is an overall architecture diagram of the YOLO-Tree model provided in the embodiments of this application;
[0023] Figure 4 This is an internal structure diagram of a computer system provided in an embodiment of this application. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0025] In recent years, with the rapid development of remote sensing technology, various remote sensing platforms, including drones, aircraft, and satellites, have been used for forest monitoring at different scales. In particular, drones, due to their flexibility, are widely used to acquire multi-temporal, high-resolution remote sensing images. Furthermore, with the continuous updates to deep learning model architectures in the field of computer vision, especially the emergence of some state-of-the-art (SOTA) models in object detection or semantic segmentation, the effectiveness of tree species mapping tasks based on drone remote sensing images has been confirmed.
[0026] Currently, research on forest tree species identification using UAV remote sensing technology is still dominated by passive remote sensing methods such as visible light, multispectral, and hyperspectral imaging, especially RGB images, which are widely used due to their low cost and high resolution. With the continuous advancement of sensor technology, the acquisition of high-resolution aerial imagery has become more efficient and economical; however, high-quality ground annotation still relies on manual field surveys, which are time-consuming, labor-intensive, and costly to obtain in-situ data. Semi-supervised learning (SSL), as an effective method to address the bottleneck of data annotation, has received widespread attention in the field of remote sensing. Through strategies such as pseudo-label optimization, SSL can fully utilize large amounts of unlabeled data to help the model learn more robust feature representations, thereby improving the model's accuracy and transferability. In recent years, SSL has been widely used in land cover classification and small target detection, but research on its application to forest remote sensing monitoring tasks is very limited.
[0027] Therefore, combining UAV RGB imagery and semi-supervised learning to detect and identify multiple tree species in heterogeneous forests to obtain refined information on the spatial distribution of forest tree species is of great significance for forest ecosystem research, biodiversity monitoring, and sustainable forestry management.
[0028] The purpose of this application is to provide a method and system for forest tree species identification that combines UAV RGB imagery and semi-supervised learning, aiming to make full use of a large amount of unlabeled data to improve the robustness of the model, thereby improving the accuracy and generalization ability of the model.
[0029] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0030] In one exemplary embodiment, such as Figure 1 As shown, this embodiment provides a forest tree species identification method that combines UAV RGB imagery and semi-supervised learning. This method includes:
[0031] Step S1: Acquire drone aerial images of multiple sample areas.
[0032] In this embodiment, the UAV aerial images of each sample area include multiple forest canopy RGB images. Step S1 specifically includes:
[0033] (1) Set the flight parameters of the UAV. Flight parameters include: flight altitude, flight speed, flight mode, and lateral and directional overlap.
[0034] In a preferred embodiment, the flight altitude of the UAV is set to 60m, 100m and 150m respectively, and the corresponding flight speeds are set to 3m / s, 6m / s and 8m / s respectively; the flight mode of the UAV is set to ground-following flight mode; and the lateral and directional overlap rates of the UAV are set to 80%.
[0035] (2) The forest canopy of multiple sample areas was photographed by drones month by month. Each time the drone was photographed, it was necessary to photograph the forest canopy of each sample area from multiple flight altitudes to obtain multiple forest canopy RGB images of each sample area.
[0036] Step S2: Stitch together multiple forest canopy RGB images of each sample area to obtain a digital orthophoto of each sample area.
[0037] In this embodiment, DJI Terra software is used to stitch together multiple forest canopy RGB images of each sample area.
[0038] Step S3: Resample and rotate the digital orthophotos of each sample region to obtain the optimized orthophotos of each sample region.
[0039] In this embodiment, ArcGIS Pro software is used to resample and rotate the digital orthophotos of each sample region.
[0040] Step S4: Divide all optimized orthophotos into a first image set and a second image set according to the proportion, and record the geographical location and species of each tree in the corresponding sample area of the first image set to obtain tree species location record information.
[0041] In this embodiment, a manual ground survey is required for the sample area corresponding to the first image set. First, all trees visible in the forest canopy of the sample area are marked in ArcGIS Pro software, and the latitude and longitude of the center of the canopy outline of each tree are recorded. Then, a manual ground survey is conducted, and a handheld GPS measuring device is used to locate the geographical location of each tree and match it with the latitude and longitude. Tree species are identified by checking the characteristics of the bark, branches and leaves.
[0042] Step S5: Based on the first image set and tree species location record information, perform contour annotation on the dominant tree species in the sample area corresponding to the first image set to obtain the sample image annotation set.
[0043] In this embodiment, the annotation of the first image set is mainly accomplished using Labelme software. In actual operation, the first image set and tree species location recording information need to be input into Labelme software, and then the following steps are performed:
[0044] (1) Based on the tree species location record information and the optimized orthophoto JSON file of the first image set, generate the corresponding single-channel image.
[0045] (2) The single-channel image is cropped into multiple small blocks by framing.
[0046] (3) Extract the position coordinates and label information of all contours based on the image labels of each small patch, and write them into the corresponding label file.
[0047] (4) Convert the JSON file of each small image patch to the corresponding target detection format to obtain the sample image annotation set.
[0048] Step S6: Perform frame cropping on the optimized orthophotos in the second image set to obtain an unlabeled sample image set.
[0049] In this embodiment, ArcGIS Pro software is used to perform frame cropping on the optimized orthophotos in the second image set. Since the optimized orthophotos after stitching, resampling, and rotation are usually large in size, they need to be cropped to a tile size of 640×640 for subsequent model training. A 50% overlap cropping ratio is set to expand the size of the dataset.
[0050] Step S7: Construct an Efficient-Tree model based on a pseudo-labeled teacher-student framework, and train the Efficient-Tree model using a set of labeled sample images and a set of unlabeled sample images.
[0051] In this embodiment, the Efficient-Tree model adopts the following... Figure 2The pseudo-label-based teacher-student framework shown aims to fully utilize a large amount of unlabeled data to improve the model's detection accuracy and generalization ability in heterogeneous forest environments. In this framework, the entire training process is divided into supervised (burn-in) and semi-supervised phases. The student model is the actual participant in the training, while the teacher model is updated based on the student model's parameters using an exponential moving average technique. The supervised phase is an initialization process, primarily aimed at enabling both the student and teacher models to acquire preliminary detection capabilities through supervised learning. In the subsequent semi-supervised phase, unlabeled data enhanced with mosaic effects is input into the teacher model to generate pseudo-labels. The pseudo-label assigner (PLA) uses a dual threshold to classify non-maximum suppressed pseudo-labels into two categories: reliable and uncertain. Based on the confidence score, uncertain pseudo-labels are further subdivided into high classification score and high confidence score pseudo-labels. PLA also introduces a soft loss to increase the consistency constraint of uncertain pseudo-labels; specifically, uncertain pseudo-labels with high classification scores are only used to calculate the confidence loss. The uncertain pseudo-label of the high confidence score is then used to calculate its regression loss. This pseudo-label allocation mechanism and soft loss calculation method help alleviate the pseudo-label inconsistency problem in semi-supervised object detection training. Subsequently, the strongly enhanced unlabeled data (with pseudo-labels) and the mosaic-enhanced labeled data serve together as supervision signals to guide the student model's updates and optimizations. Furthermore, an EpochAdaptor (EA) is designed to stabilize and accelerate the training process. In the burn-in phase, it uses a domain adaptation method to reduce the distributional differences between labeled and unlabeled datasets, preventing the student model from overfitting to labeled data and thus improving the model's generalization ability. In the semi-supervised phase, it uses a redistribution-based distribution adaptation method to automatically calculate the pseudo-label threshold for each epoch.
[0052] Furthermore, the pseudo-label-based teacher-student framework uses the YOLO-Tree model as the baseline detector; its model architecture can be found here. Figure 3The YOLO-Tree model uses a ResNet-34 network as its backbone, while the neck component is largely inherited from the YOLOV8 network. The neck structure embeds a Channel Attention Module (CAM) and a Multi-scale Attention Module (MSAM) to mitigate semantic mismatch caused by upsampling / downsampling during the layer-by-layer fusion of high- and low-level features. CAM uses global average pooling to extract global channel information, and then generates channel weights through two 1×1 convolutions and a sigmoid activation function, thereby dynamically adjusting the feature responses between feature channels. CAM has low computational cost and flexible input size, making it more suitable for processing features from different levels. MSAM uses parallel multi-branch depthwise separable convolutions to capture features at different scales. Small 3×3 convolutional kernels are used to extract local details, while large 5×5 and 7×7 convolutional kernels are used to expand the receptive field and aggregate broader contextual information. 1×1 convolutions are then used to compress the channel dimensions to generate a spatial attention mask, and soft feature selection is achieved by weighted fusion of the outputs of different branches. CAM and MSAM alternately process high- and low-level features in a top-down and bottom-up process. During upsampling, MSAM dynamically adjusts the spatial distribution of high-level features to correct for positional shifts caused by interpolation, while CAM is used to suppress noise in low-level features. During downsampling, MSAM expands the receptive field of low-level features and reduces detail loss by integrating global information, while CAM is used to focus on key regions of high-level semantic features. Since the number of feature channels in adjacent layers is not consistent, a simple concatenation method is used to aggregate the processed features to avoid introducing additional parameters.
[0053] Step S8: Use the trained Efficient-Tree model and the Slicing Aided Hyper Inference (SAHI) library to identify the digital orthophoto of the target forest area and obtain the detection results of each tree species in the target forest area.
[0054] In this embodiment, step S8 specifically includes:
[0055] (1) Using the SAHI library, set the image size and overlap ratio to crop the digital orthophoto of the target forest area into multiple tiles.
[0056] (2) Use the trained Efficient-Tree model to perform reasoning on each tile to obtain the corresponding detection results.
[0057] (3) Use the SAHI library to post-process all detection results and merge overlapping detection boxes to generate detection results for each tree species in the entire target forest area.
[0058] Based on the above analysis, the Efficient-Tree model proposed in this embodiment integrates the improved baseline detector (i.e., the YOLO-Tree model) into a pseudo-label-based teacher-student semi-supervised learning framework. By enhancing the basic detection capabilities, it ensures the quality of pseudo-labels and training stability, so as to give full play to the performance improvement potential brought by semi-supervised learning, which is beneficial to the practice of remote sensing monitoring of forest tree species diversity.
[0059] In another exemplary embodiment, a practical application scenario is provided for a forest tree species identification method that combines UAV RGB imagery and semi-supervised learning, specifically including:
[0060] The first step is the acquisition and processing of aerial imagery from drones. In this embodiment, a DJI Mavic 3E drone was used to collect monthly RGB images of the forest canopy in three sample areas within the Tianmu Mountain Nature Reserve in Zhejiang Province. First, DJI Terra software was used to stitch together multiple RGB images of the forest canopy in each sample area to obtain a digital orthophoto of that area. Then, ArcGIS Pro software was used to resample and rotate the digital orthophotos of each sample area to obtain an optimized orthophoto. Two sample areas were selected for standardization of their optimized orthophotos. The optimized orthophoto of the remaining sample area was left unlabeled and directly cropped to obtain an unlabeled sample image set.
[0061] The second step is field survey. For the two selected sample areas, all trees visible in the canopy are first marked in ArcGIS Pro software, and the latitude and longitude of the center of the canopy outline of each tree are recorded. Then, an actual ground survey is conducted, and the geographical location of each tree is located using a handheld GPS measuring device to match its latitude and longitude. Tree species are identified by examining the characteristics of the bark, branches and leaves, and the tree species location records for the two sample areas are obtained.
[0062] The third step is data annotation. The optimized orthophotos and tree species location records of the two sample areas are input into Labelme software. Labelme software is used to annotate the outlines of the dominant tree species in each sample area. To improve annotation efficiency, for orthophotos of the same location at different times, the orthophoto of a specific month selected during resampling is used as a baseline to annotate the visible canopy outlines of the surveyed trees. This annotation result is then applied to orthophotos of all other months. Each outline is then manually modified and adjusted to avoid subtle differences in tree outlines caused by different seasons, ultimately resulting in a sample image annotation set.
[0063] The fourth step is semi-supervised object detection. The Efficient-Tree model is trained using a set of labeled sample images. In the supervised learning phase, the Efficient-Tree model uses the full-site dataset (i.e., the labeled sample image set). In the semi-supervised learning phase, data from other sites (i.e., unlabeled sample images) are additionally introduced as unlabeled data for joint training. Fine-tuning is performed based on a pre-trained YOLOv8 model to accelerate convergence and improve performance. Finally, the trained Efficient-Tree model is used for slice inference of the entire orthophoto image to improve the detection capability of small targets in ultra-high resolution images.
[0064] In another exemplary embodiment, a computer system is provided, which may be a server or a terminal, and its internal structure diagram may be as follows: Figure 4 As shown, the computer system includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores unmanned aerial imagery of multiple sample areas. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the aforementioned forest tree species identification method combining UAV RGB imagery and semi-supervised learning.
[0065] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer system to which the present application is applied. A specific computer system may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0066] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0067] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0068] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0069] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0070] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A forest tree species identification method combining UAV RGB imagery and semi-supervised learning, characterized in that, The forest tree species identification method combining UAV RGB imagery and semi-supervised learning includes: Acquire drone aerial images of multiple sample areas; each sample area's drone aerial imagery includes multiple forest canopy RGB images. Multiple forest canopy RGB images of each sample area were stitched together to obtain digital orthophotos of each sample area. The digital orthophotos of each sample region are resampled and rotated to obtain optimized orthophotos of each sample region. All optimized orthophotos are divided into a first image set and a second image set according to the proportions, and the geographical location and species of each tree in the corresponding sample area of the first image set are recorded to obtain tree species location record information. Based on the first image set and the tree species location record information, the dominant tree species in the sample area corresponding to the first image set are outlined to obtain a sample image annotation set. The optimized orthophotos in the second image set are cropped in sections to obtain an unlabeled sample image set; An Efficient-Tree model is constructed based on a pseudo-labeled teacher-student framework, and trained using the labeled sample image set and the unlabeled sample image set. The pseudo-labeled teacher-student framework uses the YOLO-Tree model as a baseline detector. The Efficient-Tree model uses the labeled sample image set in the supervised learning phase and introduces the unlabeled sample image set for joint training in the semi-supervised learning phase. The trained Efficient-Tree model and SAHI database are used to identify digital orthophotos of the target forest area, and the detection results of each tree species in the target forest area are obtained.
2. The forest tree species identification method combining UAV RGB imagery and semi-supervised learning according to claim 1, characterized in that, The acquisition of unmanned aerial images of multiple sample areas specifically includes: Set the flight parameters of the drone; the flight parameters include: flight altitude, flight speed, flight mode, lateral and directional overlap ratio; The forest canopy of multiple sample areas was photographed monthly using drones. Each time the drones were used, they were required to photograph the forest canopy of each sample area from various flight altitudes, resulting in multiple RGB images of the forest canopy of each sample area.
3. The forest tree species identification method combining UAV RGB imagery and semi-supervised learning according to claim 2, characterized in that, The drone's flight altitudes were set to 60m, 100m, and 150m, respectively, and the corresponding flight speeds were set to 3m / s, 6m / s, and 8m / s, respectively. The drone's flight mode was set to terrain-following flight mode, and the drone's lateral and directional overlap rates were set to 80%.
4. The forest tree species identification method combining UAV RGB imagery and semi-supervised learning according to claim 1, characterized in that, Multiple forest canopy RGB images of each sample area were stitched together using DJI Terra software.
5. The forest tree species identification method combining UAV RGB imagery and semi-supervised learning according to claim 1, characterized in that, The digital orthophotos of each sample region were resampled and rotated using ArcGIS Pro software; the optimized orthophotos in the second image set were then cropped using ArcGIS Pro software.
6. The forest tree species identification method combining UAV RGB imagery and semi-supervised learning according to claim 1, characterized in that, The YOLO-Tree model uses a ResNet-34 network as the backbone network and a YOLOV8 network structure as the neck structure, embedding CAM and MSAM in the neck structure.
7. The forest tree species identification method combining UAV RGB imagery and semi-supervised learning according to claim 1, characterized in that, In the pseudo-label-based teacher-student framework, the entire training process is divided into two stages: supervised and semi-supervised. The student model actually participates in the training, while the teacher model updates its parameters based on the student model using the exponential moving average technique.
8. The forest tree species identification method combining UAV RGB imagery and semi-supervised learning according to claim 1, characterized in that, Based on the first image set and the tree species location record information, the dominant tree species in the sample area corresponding to the first image set are outlined to obtain a sample image annotation set. Based on the tree species location record information and the JSON file of the optimized orthophoto from the first image set, a corresponding single-channel image is generated; The single-channel image is cropped into multiple small image blocks by framing; Based on the image labels of each small image patch, the position coordinates and label information of all contours are extracted and written to the corresponding label file; The JSON files of each small image patch are converted to the corresponding target detection format to obtain the sample image annotation set.
9. The forest tree species identification method combining UAV RGB imagery and semi-supervised learning according to claim 1, characterized in that, The trained Efficient-Tree model and the SAHI database are used to identify digital orthophotos of the target forest area, obtaining the detection results of each tree species within the target forest area, specifically including: Using the SAHI library, the image size and overlap ratio are set to crop the digital orthophoto of the target forest area into multiple tiles; The trained Efficient-Tree model is used to perform inference on each tile to obtain the corresponding detection results; The SAHI library is used to post-process all detection results, and overlapping detection boxes are merged to generate detection results for each tree species in the entire target forest area.
10. A computer system, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the forest tree species identification method combining UAV RGB imagery and semi-supervised learning as described in any one of claims 1-9.