Larix singe tree segmentation and phenotype extraction method based on unmanned aerial vehicle multi-source remote sensing
By using UAV multi-source remote sensing technology and dynamic sparse gated multimodal feature fusion network, the accuracy and automation issues of larch single-tree segmentation and phenotypic extraction were solved, achieving efficient and accurate single-tree identification and parameter extraction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INSTITUTE OF ECOLOGICAL PROTECTION & RESTORATION CHINESE ACADEMY OF FORESTRY SCIENCE
- Filing Date
- 2025-12-02
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies, when dealing with tree species like larch with their complex canopy structure and dense foliage, suffer from limitations in single-modal data and insufficient utilization of features. This results in low accuracy of individual tree segmentation and phenotypic extraction, insufficient automation, and difficulty in efficient implementation in large-scale forest areas.
Using UAV multi-source remote sensing technology, a dynamic sparse gated multimodal feature fusion network is constructed to integrate optical images, multispectral images, and lidar point cloud data to achieve semantic segmentation and instance segmentation of single tree canopies. Phenotypic parameters such as tree height, crown width, and diameter at breast height are extracted by combining lidar point clouds.
It significantly improves the accuracy of single-tree segmentation, realizes fully automatic extraction from images to parameters, avoids errors caused by manual intervention and step-by-step processing, has strong robustness, adapts to different forest land conditions, and meets the needs of forestry scientific research and production.
Smart Images

Figure CN121962882A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of forestry remote sensing, and in particular to a method for segmenting and extracting the phenotype of larch trees based on multi-source remote sensing from unmanned aerial vehicles. Background Technology
[0002] Larch is an important timber species in my country, and precise segmentation and phenotypic parameter extraction at the individual tree scale are of great significance for forest resource surveys, genetic breeding, and carbon sequestration. Traditional manual field measurement methods are inefficient, costly, and difficult to implement in large-scale forest areas.
[0003] With the development of UAV remote sensing technology, high-resolution images and 3D point cloud data of forest stands can be quickly acquired by equipping them with sensors such as optical cameras, multispectral cameras, and lidar. However, existing methods still face challenges when dealing with tree species like larch, which have complex canopy structures and dense foliage. Limitations of single-modal data: Optical images alone are susceptible to the effects of light and shadow, and it is difficult to accurately distinguish the boundaries of several adjacent trees; LiDAR point clouds alone have limited ability to identify lower branches and trunks under dense canopies.
[0004] Insufficient feature utilization: Existing methods are mostly based on manually designed features or simple feature stitching, failing to fully explore the deep complementary information between multi-source remote sensing data (such as spectrum, texture, and 3D structure).
[0005] Insufficient feature utilization: Existing methods are mostly based on manually designed features or simple feature stitching, failing to fully explore the deep complementary information between multi-source remote sensing data (such as spectrum, texture, and 3D structure).
[0006] The automation level of phenotypic extraction is low: segmentation and parameter extraction are often carried out in separate steps, the process is complex, and the extraction accuracy of key phenotypic parameters such as tree height, crown width, and diameter at breast height needs to be improved.
[0007] Therefore, this invention proposes a method for larch single-tree segmentation and phenotypic extraction based on UAV multi-source remote sensing. Summary of the Invention
[0008] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method for larch single-tree segmentation and phenotypic extraction based on UAV multi-source remote sensing.
[0009] To achieve the above objectives, the present invention adopts the following technical solution: A method for larch tree segmentation and phenotypic extraction based on UAV multi-source remote sensing includes the following steps: S1: Use an unmanned aerial vehicle (UAV) platform to collect multi-source remote sensing data of larch forest areas, and perform data preprocessing and registration; S2: Construct a dynamic sparse gated multimodal feature fusion network. The dynamic sparse gated multimodal feature fusion network calculates sparse attention weights to fuse features from different remote sensing modalities. S3: Using the dynamic sparse gated multimodal feature fusion network constructed in step S2, semantic segmentation and instance segmentation of the single-canopy layer are performed on the input multi-source data to obtain the outline of each larch tree; S4: Based on the individual tree segmentation results obtained in step S3, and combined with lidar point cloud and image information, extract the tree height, crown width, crown area and diameter at breast height phenotypic parameters of the individual trees.
[0010] Preferably, in step S1, the drones use a swarm for synchronous data collection, and the swarm performs mutual positioning and calibration through communication, specifically as follows: A1: Input the drone flight mission and path planning information into the control system of each drone; A2: The GPS positioning module acquires the drone's flight position in real time during flight and corrects the drone according to the planned path if deviation occurs; A3: The signal strength detection unit within the GPS positioning module detects the positioning signal strength in real time; A4: Set a lower limit threshold for signal strength based on the positioning deviation caused by signal strength and the required positioning deviation accuracy; A5: When the signal strength of the positioning signal is lower than the threshold, a positioning algorithm based on mutual voting is used to utilize the network of drones for positioning.
[0011] Preferably, in step A5, the method of using a network of drones for positioning based on a mutual voting algorithm includes the following steps: A51: Based on the path planning, the theoretical positions between each UAV are obtained; A52: Determine the actual location of each drone based on the visual acquisition module's acquisition of visual markers; A53: By using the deviation between theoretical and actual azimuth, the positional accuracy of the target UAV is voted on, and the benchmark UAV is determined based on the voting results; A54: Combined with the acquisition of visual markers and orientation calculations on the reference UAV by the visual acquisition module, the orientation of the UAV is adjusted.
[0012] Preferably, step A53 includes the following steps: A531: Set the allowable deviation coefficient k; A532: Decompose the orientation along three-dimensional space to obtain the theoretical distances along the X, Y, and Z axes respectively. , , and actual distance , , ; A533: Judgment ①: 、②: , ③ ; A535: Average the votes of all drones for the target drone to obtain the target drone's vote score.
[0013] Preferably, in step S1, the multi-source remote sensing data includes visible light images, multispectral images, and lidar point clouds.
[0014] Preferably, in step S1, data preprocessing includes: Point cloud denoising and filtering: The raw point cloud acquired by lidar is processed to remove flying points (noise points floating in the air) and noise points (abnormal points caused by dust, insects, etc.), and smoothing is performed using statistical filtering or radius filtering algorithms. Spectral radiometric calibration involves processing image data acquired by a multispectral camera to convert the raw DN values recorded by the camera into reflectance values. Image orthorectification processes visible light and multispectral images, using positioning and attitude system (POS) data and digital surface model (DSM) acquired during flight to eliminate distortions caused by terrain undulations and camera perspective.
[0015] Preferably, in step S2, the dynamic sparse gating calculation process in the dynamic sparse gating multimodal feature fusion network is represented as follows: ,in: , , The feature map is a three-dimensional tensor derived from different modalities, and includes at least a feature extractor derived from optical images, features derived from multispectral images, and features derived from lidar point clouds. M is the total number of modes, M≥3; To be Perform vector concatenation; Here are the weight matrix and bias vector of the first layer of the gated network G; for Activation function ; Here are the weight matrix and bias vector of the output layer of the gated network G; Sigmoid is a sigmoid activation function. ; g is a dynamically generated sparse attention graph.
[0016] Preferably, in step S2, the specific process of feature fusion is as follows: The attention map g is fused based on... Percentage sparsity processing yields a binary mask, which is then used to calculate the fused features: ,in: To create a binary sparse mask, the attention map g is processed... Choose to obtain; This is the default background feature vector; This is the final feature map after fusion; and They represent two different modes.
[0017] Preferably, in step S5, the overall loss function for multi-task joint training is: ,in: For semantic segmentation loss, The weights are for the semantic segmentation loss; For instance segmentation loss, Instance segmentation loss weights; Loss estimation for diameter at breast height , The loss weight for estimating the diameter at breast height.
[0018] A forestry survey system, comprising: The data acquisition module is used to control the UAV to collect remote sensing data from multiple sources. The processing module integrates a dynamic sparse gated multimodal feature fusion network trained using a larch single-tree segmentation and phenotypic extraction method based on UAV multi-source remote sensing. The phenotypic extraction module is used to automatically calculate individual tree phenotypic parameters based on the segmentation results of the processing module. The results output module is used to generate individual tree distribution maps, phenotypic parameter reports, and statistical analysis results.
[0019] The beneficial effects of this invention are as follows: This invention effectively utilizes the advantages of multi-source data through a dynamic sparse gating fusion mechanism, significantly improving the segmentation accuracy of individual trees in complex scenarios. It also achieves fully automated extraction from images to parameters, avoiding the accumulation of errors caused by cumbersome manual intervention and step-by-step processing. Furthermore, it can adapt to different forest conditions and has strong robustness in the identification and parameter extraction of individual larch trees. Finally, it can simultaneously acquire various phenotypic parameters such as tree height, crown width, crown area, and even estimated diameter at breast height, meeting the diverse needs of forestry research and production. Attached Figure Description
[0020] Figure 1 This is a flowchart of a method for segmenting and extracting the phenotype of larch trees based on multi-source remote sensing from unmanned aerial vehicles (UAVs) proposed in this invention. Detailed Implementation
[0021] The technical solution of the present invention will be further described in detail below with reference to specific embodiments.
[0022] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "setting" should be interpreted broadly. For example, they can refer to a fixed connection or setting, a detachable connection or setting, or an integral connection or setting. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0023] Example 1: A method for larch tree segmentation and phenotypic extraction based on UAV multi-source remote sensing includes the following steps: S1: Use an unmanned aerial vehicle (UAV) platform to collect multi-source remote sensing data of larch forest areas, and perform data preprocessing and registration; S2: Construct a dynamic sparse gated multimodal feature fusion network. The dynamic sparse gated multimodal feature fusion network calculates sparse attention weights to fuse features from different remote sensing modalities. S3: Using the dynamic sparse gated multimodal feature fusion network constructed in step S2, semantic segmentation and instance segmentation of the single-canopy layer are performed on the input multi-source data to obtain the outline of each larch tree; S4: Based on the individual tree segmentation results obtained in step S3, and combined with lidar point cloud and image information, extract the tree height, crown width, crown area and diameter at breast height phenotypic parameters of the individual trees.
[0024] In step S1, the drones use a swarm for synchronous data collection. The swarm performs mutual positioning and calibration through communication. Specifically: A1: Input the drone flight mission and path planning information into the control system of each drone; A2: The GPS positioning module acquires the drone's flight position in real time during flight and corrects the drone according to the planned path if deviation occurs; A3: The signal strength detection unit within the GPS positioning module detects the positioning signal strength in real time; A4: Set a lower limit threshold for signal strength based on the positioning deviation caused by signal strength and the required positioning deviation accuracy; A5: When the signal strength of the positioning signal is lower than the threshold, a positioning algorithm based on mutual voting is used to utilize the network of drones for positioning.
[0025] In step A5, the method of using a network of drones for positioning based on a mutual voting algorithm includes the following steps: A51: Based on the path planning, the theoretical positions between each UAV are obtained; A52: Determine the actual location of each drone based on the visual acquisition module's acquisition of visual markers; A53: By using the deviation between theoretical and actual azimuth, the positional accuracy of the target UAV is voted on, and the benchmark UAV is determined based on the voting results; A54: Combined with the acquisition of visual markers and orientation calculations on the reference UAV by the visual acquisition module, the orientation of the UAV is adjusted.
[0026] Step A53 includes the following steps: A531: Set the allowable deviation coefficient k; A532: Decompose the orientation along three-dimensional space to obtain the theoretical distances along the X, Y, and Z axes respectively. , , and actual distance , , ; A533: Judgment ①: 、②: ③ ; A534: Determine the voting score based on the conditions met in ①②③; A535: Average the votes of all drones for the target drone to obtain the target drone's vote score.
[0027] Example 2: A method for larch tree segmentation and phenotypic extraction based on UAV multi-source remote sensing includes the following steps: S1: Use an unmanned aerial vehicle (UAV) platform to collect multi-source remote sensing data of larch forest areas, and perform data preprocessing and registration; S2: Construct a dynamic sparse gated multimodal feature fusion network. The dynamic sparse gated multimodal feature fusion network calculates sparse attention weights to fuse features from different remote sensing modalities. S3: Using the dynamic sparse gated multimodal feature fusion network constructed in step S2, semantic segmentation and instance segmentation of the single-canopy layer are performed on the input multi-source data to obtain the outline of each larch tree; S4: Based on the individual tree segmentation results obtained in step S3, and combined with lidar point cloud and image information, extract the tree height, crown width, crown area and diameter at breast height phenotypic parameters of the individual trees.
[0028] In step S1, the multi-source remote sensing data includes visible light images, multispectral images, and lidar point clouds.
[0029] In step S1, data preprocessing includes: Point cloud denoising and filtering: The raw point cloud acquired by lidar is processed to remove flying points (noise points floating in the air) and noise points (abnormal points caused by dust, insects, etc.), and smoothing is performed using statistical filtering or radius filtering algorithms. Spectral radiometric calibration involves processing image data acquired by a multispectral camera to convert the raw DN values recorded by the camera into reflectance values. Image orthorectification processes visible light and multispectral images, using positioning and attitude system (POS) data and digital surface model (DSM) acquired during flight to eliminate distortions caused by terrain undulations and camera perspective.
[0030] In step S2, the dynamic sparse gating calculation process in the dynamic sparse gating multimodal feature fusion network is represented as follows: ,in: , , The feature map is a three-dimensional tensor derived from different modalities, and includes at least a feature extractor derived from optical images, features derived from multispectral images, and features derived from lidar point clouds. M is the total number of modes, M≥3; To be , , Perform vector concatenation; , Here are the weight matrix and bias vector of the first layer of the gated network G; ReLU(·) is the ReLU activation function, ReLU(x) = max(0, x); , Here are the weight matrix and bias vector of the output layer of the gated network G; Sigmoid is a sigmoid activation function. ; g is a dynamically generated sparse attention graph.
[0031] In step S2, the specific process of feature fusion is as follows: the attention map g is sparsified based on the Top-k percentage to obtain a binary mask, and then the fused features are calculated: ,in: The binary sparse mask is obtained by performing a Top-k% selection on the attention map g. This is the default background feature vector; This is the final feature map after fusion; and They represent two different modes.
[0032] In step S5, the overall loss function for multi-task joint training is: ,in: For semantic segmentation loss, The weights are for the semantic segmentation loss; For instance segmentation loss, Instance segmentation loss weights; To estimate the loss in diameter at breast height, The loss weight for estimating the diameter at breast height.
[0033] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for larch single-tree segmentation and phenotypic extraction based on UAV multi-source remote sensing, characterized in that, Includes the following steps: S1: Use an unmanned aerial vehicle (UAV) platform to collect multi-source remote sensing data of larch forest areas, and perform data preprocessing and registration; S2: Construct a dynamic sparse gated multimodal feature fusion network. The dynamic sparse gated multimodal feature fusion network calculates sparse attention weights to fuse features from different remote sensing modalities. S3: Using the dynamic sparse gated multimodal feature fusion network constructed in step S2, semantic segmentation and instance segmentation of the single-canopy layer are performed on the input multi-source data to obtain the outline of each larch tree; S4: Based on the individual tree segmentation results obtained in step S3, and combined with lidar point cloud and image information, extract the tree height, crown width, crown area and diameter at breast height phenotypic parameters of the individual trees.
2. The method for larch single-tree segmentation and phenotypic extraction based on UAV multi-source remote sensing according to claim 1, characterized in that, In step S1, the drones use a swarm for synchronous data collection. The swarm performs mutual positioning and calibration through communication. Specifically: A1: Input the drone flight mission and path planning information into the control system of each drone; A2: The GPS positioning module acquires the drone's flight position in real time during flight and corrects the drone according to the planned path if deviation occurs; A3: The signal strength detection unit within the GPS positioning module detects the positioning signal strength in real time; A4: Set a lower limit threshold for signal strength based on the positioning deviation caused by signal strength and the required positioning deviation accuracy; A5: When the signal strength of the positioning signal is lower than the threshold, a positioning algorithm based on mutual voting is used to utilize the network of drones for positioning.
3. The method for larch single-tree segmentation and phenotypic extraction based on UAV multi-source remote sensing according to claim 2, characterized in that, In step A5, the method of using a network of drones for positioning based on a mutual voting algorithm includes the following steps: A51: Based on the path planning, the theoretical positions between each UAV are obtained; A52: Determine the actual location of each drone based on the visual acquisition module's acquisition of visual markers; A53: By using the deviation between theoretical and actual azimuth, the positional accuracy of the target UAV is voted on, and the benchmark UAV is determined based on the voting results; A54: Combined with the acquisition of visual markers and orientation calculations on the reference UAV by the visual acquisition module, the orientation of the UAV is adjusted.
4. The method for larch single-tree segmentation and phenotypic extraction based on UAV multi-source remote sensing according to claim 3, characterized in that, Step A53 includes the following steps: A531: Set the allowable deviation coefficient k; A532: Decompose the orientation along three-dimensional space to obtain the theoretical distances along the X, Y, and Z axes respectively. , , and actual distance ; A533: Judgment ①: ②: ③ ; A534: Determine the voting score based on the conditions met in ①②③; A535: Average the votes of all drones for the target drone to obtain the target drone's vote score.
5. The method for larch single-tree segmentation and phenotypic extraction based on UAV multi-source remote sensing according to claim 1, characterized in that, In step S1, the multi-source remote sensing data includes visible light images, multispectral images, and lidar point clouds.
6. The method for larch single-tree segmentation and phenotypic extraction based on UAV multi-source remote sensing according to claim 1, characterized in that, In step S1, data preprocessing includes: Point cloud denoising and filtering: The raw point cloud acquired by lidar is processed to remove flying points and noise points, and smoothed using statistical filtering or radius filtering algorithms. Spectral radiometric calibration involves processing image data acquired by a multispectral camera to convert the raw DN values recorded by the camera into reflectance values. Image orthorectification processes visible light and multispectral images, using positioning and attitude system (POS) data and digital surface model (DSM) acquired during flight to eliminate distortions caused by terrain undulations and camera perspective.
7. The method for larch single-tree segmentation and phenotypic extraction based on UAV multi-source remote sensing according to claim 1, characterized in that, In step S2, the dynamic sparse gating calculation process in the dynamic sparse gating multimodal feature fusion network is represented as follows: in: , , The feature map is a three-dimensional tensor derived from different modalities, and includes at least a feature extractor derived from optical images, features derived from multispectral images, and features derived from lidar point clouds. M is the total number of modes, M≥3; To be Perform vector concatenation; , Here are the weight matrix and bias vector of the first layer of the gated network G; It is the ReLU activation function. ; , Here are the weight matrix and bias vector of the output layer of the gated network G; Sigmoid is a sigmoid activation function. ; g is a dynamically generated sparse attention graph.
8. The method for larch single-tree segmentation and phenotypic extraction based on UAV multi-source remote sensing according to claim 1, characterized in that, In step S2, the specific process of feature fusion is as follows: the attention map g is sparsified based on the Top-k percentage to obtain a binary mask, and then the fused features are calculated: ,in: The binary sparse mask is obtained by performing a Top-k% selection on the attention map g. This is the default background feature vector; This is the final feature map after fusion; and They represent two different modes.
9. The method for larch single-tree segmentation and phenotypic extraction based on UAV multi-source remote sensing according to claim 1, characterized in that, In step S5, the overall loss function for multi-task joint training is: ,in: For semantic segmentation loss, The weights are for the semantic segmentation loss; For instance segmentation loss, Instance segmentation loss weights; To estimate the loss in diameter at breast height, The loss weight for estimating the diameter at breast height.
10. A forestry survey system, characterized in that, include: The data acquisition module is used to control the UAV to collect remote sensing data from multiple sources. The processing module integrates a dynamic sparse gated multimodal feature fusion network trained by the larch single-tree segmentation and phenotypic extraction method based on UAV multi-source remote sensing as described in any one of claims 1 to 9. The phenotypic extraction module is used to automatically calculate individual tree phenotypic parameters based on the segmentation results of the processing module. The results output module is used to generate individual tree distribution maps, phenotypic parameter reports, and statistical analysis results.