A hyperspectral image dense region tree individual plant segmentation method, system and terminal
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-27
- Publication Date
- 2026-08-11
AI Technical Summary
[0006]本发明的主要目的在于提供一种高光谱图像密集区域树木单株分割方法、系统及终端,旨在解决现有技术中树种分割方法在面对密集林地或高郁闭度森林时分割边缘不准确的问题
[0017]有益效果:本发明提供一种高光谱图像密集区域树木单株分割方法、系统及终端,该方法通过立体注意力模块提取不同单株树木的光谱差异特征及空间轮廓细节特征,以适应不同光照场景和树种变化的情况,并且周密自注意力模块能够在同一网络结构中同时处理全局和局部信息,从而有效提高了针对密集林地或高郁闭度森林进行单株树木分割的准确性,增强了分割效果。
Smart Images

Figure CN118506000B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of forest tree species identification technology, and in particular to a method, system, terminal, and computer-readable storage medium for segmenting individual trees in dense areas of hyperspectral images. Background Technology
[0002] Forests, as indispensable economic and environmental assets, play a crucial role in absorbing and fixing atmospheric carbon. In this process, individual tree segmentation studies have become a key technological tool, enabling the acquisition of high-precision information on the distribution and attributes of individual trees. This research provides crucial support for scientific research and management across various fields, offering a robust data foundation for sustainable forest management and resource planning. Through individual tree segmentation studies, we can gain a deeper understanding of the location, morphology, and characteristics of each tree in the forest. This not only helps scientists better understand the structure and function of forest ecosystems but also has wide applications in monitoring forest health, controlling forest pests and diseases, and conducting environmental impact assessments. This data has a profound impact not only on scientific research but also provides substantial support for governments and businesses in formulating forest protection and management policies. In summary, individual tree segmentation studies are a vital tool in forest ecosystem management, playing an irreplaceable role in promoting carbon sequestration and protecting the ecological environment. Collaboration between governments and research institutions will further advance this field, providing strong support for better protection and utilization of forest resources.
[0003] Tree species segmentation refers to dividing an image of a forest or woodland area into individual tree specimens, identifying and classifying each tree to determine its species, thereby enabling automated analysis and monitoring of the distribution of different tree species in the forest. Traditional tree species identification methods typically rely on manual field surveys. However, this approach faces challenges such as enormous workload, high costs, and time-consuming processes, proving inefficient, especially for mapping tasks covering large areas. In manual field surveys, professional surveyors need to spend considerable time and effort recording the species and distribution of each tree. This not only requires significant human resources but may also face challenges such as complex terrain and inconvenient transportation, limiting the depth and breadth of the survey. Furthermore, field surveys are subject to subjectivity and limitations, as the surveyor's experience and subjective judgment can affect the accuracy and consistency of the results. The limitation of this approach lies in its inability to achieve timely monitoring and comprehensive analysis of large-scale areas, particularly for scenarios requiring frequent updates and dynamic monitoring.
[0004] With the rapid advancement of remote sensing technology, a wealth of technical means have been provided for forest tree species identification. Remote sensing technology, with its macroscopic, real-time, and periodic characteristics, creates favorable conditions for the rapid, accurate, and efficient acquisition of large-scale forest resource information, and is therefore widely used in research on forest tree species segmentation, identification, and information extraction. Remote sensing data encompasses various sensor types, including visible light, multispectral, hyperspectral, and lidar data. Many scholars both domestically and internationally have successfully applied these remote sensing technologies to forest tree species identification and have achieved phased research results in this field. The application of these technologies enables the automatic acquisition of important information about forest tree species distribution through remote sensing images, not only improving work efficiency and reducing the burden of manual surveys, but also providing strong support for more comprehensive forest management and protection. However, despite these technologies providing significant progress for forest resource management, some shortcomings and challenges remain. First, they are affected by light and cloud cover. Under strong light or cloud cover, visible light images may become unclear, making it difficult to accurately capture tree features. Furthermore, vegetation shading is also a problem, especially in densely vegetated areas, which may cause trees to be partially or completely obscured, thus affecting the comprehensive identification of tree species. This results in visible light data performing worse than expected in tree species segmentation under complex weather and vegetation conditions. Secondly, LiDAR (Light Detection and Ranging) data also presents challenges in individual tree segmentation. Topographical complexity is a major issue, as undulating terrain in mountainous or canyon areas can introduce errors into LiDAR data, affecting the accuracy of 3D tree modeling and segmentation. Vegetation occlusion is another problem for LiDAR, especially in high-density vegetation areas, which may prevent LiDAR from fully acquiring structural information of all trees.
[0005] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0006] The main objective of this invention is to provide a method, system, and terminal for segmenting individual trees in dense areas of hyperspectral images, aiming to solve the problem of inaccurate segmentation edges in existing tree species segmentation methods when dealing with dense woodlands or high-density forests.
[0007] The first aspect of this application provides a method for segmenting individual trees in a dense area of a hyperspectral image. The method includes the following steps: acquiring an original image of a forest; extracting features from the original image to obtain pixel-level features and three-dimensional features; inputting the three-dimensional features and the pixel-level features into a stereo attention module to generate three-dimensional attention features; inputting the three-dimensional attention features into a comprehensive attention module to generate a first output feature; wherein the three-dimensional attention features are used to reflect the spectral differences and spatial contour details of different individual trees; extracting features from the first output feature to obtain two-dimensional features; inputting the two-dimensional features into a two-dimensional attention module to generate two-dimensional attention features; inputting the two-dimensional attention features into the comprehensive attention module to generate a second output feature; and inputting the second output feature into a segmentation module to generate segmented and recognized images of different individual trees in the forest.
[0008] Optionally, in one embodiment of this application, the original image is a hyperspectral image sample; the feature extraction of the original image to obtain pixel-level features and three-dimensional features specifically includes: inputting the hyperspectral image sample into a linear layer to obtain pixel-level features; and inputting the hyperspectral image sample into a three-dimensional convolutional layer to obtain three-dimensional features.
[0009] Optionally, in one embodiment of this application, the three-dimensional feature is a spatial spectral feature, and the three-dimensional attention feature includes a target spatial spectral feature and a target pixel-level embedding; the step of inputting the three-dimensional feature and the original image into a stereo attention module to generate three-dimensional attention features specifically includes: the stereo attention module processing the spatial spectral feature in the depth dimension, channel dimension, and spatial dimension respectively to obtain three-dimensional depth attention, three-dimensional channel attention, and three-dimensional spatial attention, and generating the target spatial spectral feature of the tree based on the spatial spectral feature, the three-dimensional depth attention, the three-dimensional channel attention, and the three-dimensional spatial attention; the stereo attention module acquiring the attention image corresponding to the three-dimensional feature, and performing weighted processing on the pixel-level feature and the attention image to obtain the target pixel-level embedding.
[0010] Optionally, in one embodiment of this application, the stereo attention module processes the spatial spectral features in the depth dimension, channel dimension, and spatial dimension respectively to obtain three-dimensional depth attention, three-dimensional channel attention, and three-dimensional spatial attention, and generates target spatial spectral features of trees based on the spatial spectral features, the three-dimensional depth attention, the three-dimensional channel attention, and the three-dimensional spatial attention, specifically including: The stereo attention module performs mean pooling on the spatial spectral features in the depth dimension to obtain three-dimensional depth attention; the formula for the three-dimensional depth attention is as follows: ; The 3D attention module performs mean pooling and max pooling on the spatial spectral features in the channel dimension to obtain 3D channel attention; the formula for the 3D channel attention is as follows: ; The stereo attention module performs mean pooling and max pooling on the spatial spectral features in the spatial dimension to obtain three-dimensional spatial attention; the formula for the three-dimensional spatial attention is as follows: ; The calculation formula for the target spatial spectral features is as follows: ; in, Indicates spatial spectral characteristics, Represents 3D depth attention. for function, This represents a shared multilayer perceptron module. and These represent mean pooling and max pooling operations, respectively. Represents 3D channel attention. Represents attention in three-dimensional space. Indicates the filter size is Convolution operation, Indicates the spatial spectral characteristics of the target. This represents element-wise multiplication; The calculation formula for the target pixel-level embedding is as follows: ; in, Indicates target pixel-level embedding, Represents pixel-level features. This represents an attention image.
[0011] Optionally, in one embodiment of this application, the meticulous attention module includes a dilated convolution module and a self-attention mechanism; the step of inputting the three-dimensional attention features into the meticulous attention module to generate the first output feature specifically includes: recoding the target spatial spectral features and the target pixel-level embedding to generate recoded features; the self-attention mechanism flattens the recoded features in the spatial dimension to generate flattened features; the dilated convolution module generates detail features based on the recoded features; and the flattened features and the detail features are concatenated and fused to obtain the first output feature.
[0012] Optionally, in one embodiment of this application, the calculation formula for the recoded feature is as follows: ; The calculation formula for the flattening feature is as follows: ; The calculation formula for the detailed features is as follows: ; The formula for calculating the first output feature is as follows: ; in, This indicates recoded features. Indicates the filter size is Convolution operation, This indicates a channel dimension splicing operation. Indicates a flat or extended feature. Represents the standard after removing position coding Encoder module, Indicate detailed features, Indicates the filter size is dilated convolution function, and These are operations for flattening and restoring the vector in space, respectively. This represents the first output feature.
[0013] Optionally, in one embodiment of this application, the step of inputting the two-dimensional features into a two-dimensional attention module to generate two-dimensional attention features specifically includes: The two-dimensional attention module performs mean pooling and max pooling on the two-dimensional features along the channel dimension to obtain two-dimensional channel attention; the formula for the two-dimensional channel attention is as follows: ; The two-dimensional attention module performs mean pooling and max pooling on the two-dimensional features in the spatial dimension to obtain two-dimensional spatial attention; the formula for the three-dimensional spatial attention is as follows: ; The formula for calculating the two-dimensional attention features is as follows: ; in, Representing two-dimensional attention features, Representing two-dimensional features, Represents two-dimensional channel attention. Represents attention in two-dimensional space. The kernel size is The convolutional layer.
[0014] A second aspect of this application also provides a tree segmentation system for densely populated regions of hyperspectral images, wherein the tree segmentation system for densely populated regions of hyperspectral images includes: A multi-level attention connection module is used to acquire the original image of the forest, extract features from the original image to obtain pixel-level features and three-dimensional features, and input the three-dimensional features and the pixel-level features into a stereo attention module to generate three-dimensional attention features. The three-dimensional attention features are then input into a comprehensive attention module to generate the first output feature. The spatial attention enhancement module is used to extract features from the first output feature to obtain two-dimensional features, and input the two-dimensional features into the two-dimensional attention module to generate two-dimensional attention features. The two-dimensional attention features are then input into the comprehensive attention module to generate the second output feature. The segmentation module is used to input the second output feature into the segmentation module to generate segmentation and recognition images of different individual trees in the forest.
[0015] A third aspect of this application also provides a terminal, wherein the terminal includes: a memory, a processor, and a hyperspectral image dense area tree segmentation program stored in the memory and executable on the processor, wherein when the hyperspectral image dense area tree segmentation program is executed by the processor, it implements the steps of the hyperspectral image dense area tree segmentation method as described above.
[0016] A fourth aspect of this application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a tree segmentation program for dense areas of hyperspectral images, and when the tree segmentation program for dense areas of hyperspectral images is executed by a processor, it implements the steps of the tree segmentation method for dense areas of hyperspectral images as described above.
[0017] Beneficial effects: This invention provides a method, system, and terminal for segmenting individual trees in dense areas of hyperspectral images. The method extracts spectral difference features and spatial contour detail features of different individual trees through a stereo attention module to adapt to different lighting scenarios and tree species variations. Furthermore, the comprehensive self-attention module can process global and local information simultaneously in the same network structure, thereby effectively improving the accuracy of segmenting individual trees in dense woodlands or high-density forests and enhancing the segmentation effect. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart of a preferred embodiment of the method for segmenting individual trees in dense areas of hyperspectral images according to this application; Figure 2 This is a schematic diagram of the expression form and design method based on the stereo attention module in a preferred embodiment of the tree segmentation method for dense areas of hyperspectral images in this application; Figure 3 This is a preferred embodiment of the tree segmentation method for dense areas of hyperspectral images in this application, which is a hybrid structure enhancement model framework based on a stereo attention module; Figure 4 This is a schematic diagram of a preferred embodiment of the tree segmentation system for dense areas of hyperspectral images in this application; Figure 5 This is a schematic diagram of a preferred embodiment of the terminal of this application.
[0020] Explanation of reference numerals in the attached figures: 10. Tree segmentation system for dense areas of hyperspectral images; 100. Multi-level attention connection module; 200. Spatial attention enhancement module; 300. Segmentation module.
[0021] The accompanying drawings have illustrated specific embodiments of the invention, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the invention in any way, but rather to illustrate the concept of the invention to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0022] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. The described embodiments are merely possible technical implementations of this invention and not all possible implementations. Based on the embodiments of this invention, those skilled in the art can obtain other embodiments without creative effort, and these embodiments are also within the protection scope of this invention.
[0023] Among related technologies, hyperspectral data has significant advantages in tree species segmentation. One of its main advantages is the richness of spectral information. By capturing reflectance spectra in multiple bands, including visible and near-infrared light, hyperspectral data provides more detailed and comprehensive spectral characteristics of trees. This allows hyperspectral data to distinguish subtle spectral differences between different tree species, providing a powerful tool for accurate tree species identification. Hyperspectral sensors have high spectral resolution, enabling observations within a finer spectral range. This high resolution allows hyperspectral data to more accurately distinguish the spectral characteristics of trees, improving the accuracy and precision of tree species identification. Another advantage is the mitigation of vegetation shading. Compared to visible light data, hyperspectral data is less sensitive to vegetation shading. Because it provides more diverse information across different bands, hyperspectral data can better penetrate vegetation, acquiring the spectral characteristics of trees beneath it, effectively reducing the impact of vegetation shading on tree species segmentation. Furthermore, hyperspectral data can be used to calculate various vegetation indices, such as NDVI, reflecting the physiological state of vegetation and providing additional ecological information for tree species classification. This capability enables hyperspectral data to play a crucial role in gaining a deeper understanding of forest ecosystems, monitoring tree growth, and promoting sustainable forest management. Therefore, hyperspectral data is widely regarded as a powerful tool for individual tree segmentation, providing more comprehensive and accurate data support for scientific research and resource management.
[0024] While existing tree segmentation methods have achieved satisfactory results in low canopy closure conditions, further in-depth research is needed on tree canopy segmentation under high canopy closure. Currently, research on high-precision identification of individual tree species in high-density forest stands using remote sensing data is relatively limited, necessitating the development of tree species segmentation methods suitable for high-density stands and capable of handling complex canopy structures. These methods need to consider the data processing challenges under high canopy closure, such as accurate canopy boundary detection, handling of occlusion areas, and efficient image segmentation algorithms. High-precision remote sensing data and advanced image processing technologies will provide crucial support for solving this problem, thereby achieving the goal of accurate identification of individual tree species in high-density forest stands.
[0025] In existing design schemes, firstly, a fully convolutional neural network model is used to obtain a fractional map from Amazon palm tree data in UAV-RGB images, and then morphological operations are used to refine the boundaries to obtain individual tree detection and tree species classification results. Secondly, spatial geometric features are first used for coarse segmentation of the canopy, and then multispectral information is used to further segment the under-segmented canopy. The segmentation results of this method are better than those using only spatial features. However, due to the lack of labeled data, the results of existing methods for single tree canopy segmentation are mixed. Thirdly, deep learning models can use existing unsupervised delineation based on light detection and ranging (LIDAR) to generate trees for training the initial RGB canopy detection model. Fourthly, when using UAV ranging (LIDAR) data to perform single tree detection and delineation of mangroves, oversegmentation may occur due to the high clump density and limited height difference between neighboring mangroves. Fifth, a set of tree clusters is obtained by processing urban mobile laser scanning point cloud data using semantic segmentation networks and Euclidean distance methods. Then, a point orientation embedding deep network is used to predict the direction vector of each tree cluster pointing to the tree center to enhance the boundary of instance-level trees. The clusters are divided into single-tree clusters and multi-tree clusters based on the number of tree centers. Single-tree clusters are individual trees, while multi-tree clusters are further segmented. Sixth, a masked region convolutional neural network (mask R-CNN) is used to automatically detect the individual canopy and height of plantation Chinese fir trees.
[0026] However, this method struggles to achieve precise segmentation when canopies overlap significantly. Although the aforementioned studies primarily focused on segmenting individual trees in complex and dense forests, some cases of oversegmentation or undersegmentation still exist.
[0027] Specifically, addressing the shortcomings mentioned in points one through four above, current tree segmentation methods face a common challenge: inaccurate segmentation edges, making it impossible to extract the precise shape of the tree. This problem is particularly pronounced when dealing with dense woodlands or high-canopy-density forests. The main reason is that existing methods are insufficient for handling tree edge segmentation in complex environments. Traditional segmentation algorithms may not effectively handle complex situations such as tree occlusion, lighting variations, and canopy overlap, leading to inaccurate edge segmentation. Furthermore, the complexity of environmental conditions also increases the difficulty of segmentation. In high-canopy-density forests, the intersection and occlusion between trees blur tree edges, thus affecting the accurate extraction of tree shapes.
[0028] To address the shortcomings mentioned in points five and six above, existing tree segmentation methods suffer from insufficient feature extraction when dealing with dense woodlands or high-canopy-density forests, leading to inaccurate segmentation and recognition. Dense woodlands are characterized by closely spaced trees, overlapping canopies, and occlusion, making it difficult for traditional feature extraction methods to fully capture subtle tree features. Furthermore, the complex lighting conditions in high-canopy-density forests, with shadows easily appearing under the tree canopy, further increase the difficulty of segmentation and recognition. Because existing methods are insufficient in feature extraction under these complex environments, the accuracy of segmentation and recognition is limited.
[0029] To address the shortcomings of the existing technologies mentioned above, the present invention aims to solve the problem of inaccurate edge segmentation caused by complex situations such as occlusion between trees, changes in lighting, and overlapping canopies, as described in points one through four above. This invention proposes a multi-layered attention connection module. This module captures the multi-layered spatial structure of trees, such as the distribution of trunks, branches, and leaves, through 3D convolution, facilitating a more comprehensive utilization of the spatial-spectral information in hyperspectral images and a more complete understanding of the tree's morphological structure. Simultaneously, the use of a stereo attention mechanism and a meticulous self-attention module assists the network in focusing more intently on regions more critical to the image segmentation task, enhancing the model's perception of different regions in hyperspectral images, strengthening the tree's contour features, and contributing to more accurate capture of tree contour information, thereby improving the accuracy of image segmentation.
[0030] To address the shortcomings of existing tree segmentation methods mentioned in points 5 and 6 above in feature extraction, a hybrid structure-enhanced neural network architecture is proposed. This architecture can capture multi-scale features, better adapt to targets or structures at different scales, and learn fine-grained tree structure and shape. This enables the model to more comprehensively understand tree structure and variations, and helps it consider the structure and layout of different trees throughout a dense area. This improvement enhances the model's ability to extract tree species features and segmentation accuracy in complex, dense forests.
[0031] The following describes a method, system, and terminal for segmenting individual trees in dense areas of hyperspectral images, based on embodiments of this application, with reference to the accompanying drawings. Addressing the problem of inaccurate segmentation edges in tree species segmentation methods in the aforementioned related technologies when dealing with dense forests or high-canopy-density forests, this application provides a method for segmenting individual trees in dense areas of hyperspectral images. In this method, a stereo attention module extracts the spectral difference features and spatial contour detail features of different individual trees to adapt to different lighting scenarios and tree species variations. Furthermore, a comprehensive self-attention module can simultaneously process global and local information within the same network structure, thereby effectively improving the accuracy of individual tree segmentation in dense forests or high-canopy-density forests and enhancing the segmentation effect. Thus, this solves the technical problem of inaccurate segmentation edges in tree species segmentation methods in the related technologies when dealing with dense forests or high-canopy-density forests.
[0032] The tree segmentation method for dense regions in hyperspectral images presented in this application is based on a Hybrid Structure Enhanced Network (HSEN). The backbone feature extractor of the HSEN consists of two branches: a Multilevel Attention Aggregation Module (MAAM) and a Spatial Attention-Enhanced Module (SAEM). MAAM includes a stereo attention module and a comprehensive attention module, while SAEM includes a two-dimensional attention module and a comprehensive attention module. This multi-scale hybrid convolutional module enables the model to have better perceptual capabilities in both spatial and depth dimensions through attention mechanisms. By using 3D and 2D convolutions, multi-scale features can be captured. Connecting these two modules allows for more effective learning of fine-grained tree structure and shape, better adaptation to targets or structures at different scales, and a more comprehensive understanding of tree structure and changes. This addresses problems such as inaccurate canopy segmentation at different scales and inaccurate canopy edge segmentation. The MAAM module first captures the multi-layered spatial structure of trees, such as the distribution of trunks, branches, and leaves, through 3D convolution, helping to more comprehensively utilize the spatial-spectral information in hyperspectral images and gain a more complete understanding of tree morphology. Then, it uses a Cubic Attention Module (CAM) to extract spectral differences and spatial contour details (canopy contour and morphological features) from different individual trees, enhancing the model's perception of the spectrum, local structure, and shape of different trees, improving its robustness, and making it more adaptable to different lighting, scenes, and tree species. Finally, a meticulous self-attention module combines self-attention mechanisms and dilated convolution modules to process global and local information simultaneously within the same network structure, further and more effectively solving the problem of tree segmentation at different scales of canopies. The SAEM module, on the other hand, first performs convolution and pooling operations on the image through 2D convolutional layers, effectively extracting various features such as texture, shape, and color from tree images. Then, it uses channel attention and spatial attention mechanisms to further enhance the extraction of spatial features of trees, more accurately delineating the boundary between trees and the background. Finally, the features from the 2D attention module are transferred to the meticulous attention module. The attention module further processes these features to extract more discriminative features. Finally, the extracted features are input into the segmentation module for segmentation and recognition, resulting in images of individual trees with high segmentation accuracy.
[0033] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0034] The preferred embodiment of the present invention describes a method for segmenting individual trees in densely populated regions of hyperspectral images, such as... Figure 1 As shown, the method for segmenting individual trees in dense areas of a hyperspectral image includes the following steps: In step S101, the original image of the forest is acquired, and feature extraction is performed on the original image to obtain pixel-level features and three-dimensional features. The three-dimensional features and the pixel-level features are then input into the stereo attention module to generate three-dimensional attention features. The three-dimensional attention features are then input into the comprehensive attention module to generate the first output feature. The three-dimensional attention features are used to reflect the spectral differences and spatial contour details of different individual trees.
[0035] In one implementation, after acquiring hyperspectral image samples, the hyperspectral image samples are input into a linear layer to obtain pixel-level features; the hyperspectral image samples are then input into a three-dimensional convolutional layer to obtain three-dimensional features.
[0036] Specifically, let Represents a hyperspectral image sample, where, , , These represent the length, width, and number of bands of the input sample, respectively. The samples are input into a linear layer and a 3D convolutional layer to obtain pixel-level embeddings. Spatial spectral features .
[0037] ; ; in, Refers to a two-dimensional convolution function with a kernel size of 1×1. The term refers to a three-dimensional convolutional function with a kernel size of 3×3×3. The three-dimensional convolutional module network mainly consists of two layers of convolutional neural networks, each of which has four three-dimensional convolutional kernels with a size of 3×3×3. This represents the batch normalization function.
[0038] Understandably, 3D convolutional layers can also be replaced with volumetric convolutional networks (V-Net), etc.
[0039] In other words, the 3D convolution in the multi-layered attention connection module can capture the multi-layered spatial structure of trees, such as the distribution of trunks, branches, and leaves, which helps to more comprehensively utilize the spatial-spectral information in hyperspectral images and gain a more comprehensive understanding of the morphological structure of trees. Simultaneously, the addition of an attention mechanism helps the network focus more intently on regions more critical to the image segmentation task, enhancing the extraction of detailed features such as canopy contours, refining canopy edges, and improving the accuracy of image segmentation. This module also helps the network understand complex scenes and improves the model's expressive and generalization abilities.
[0040] In one implementation, the three-dimensional features are spatial-spectral features, and the three-dimensional attention features include target spatial-spectral features and target pixel-level embeddings.
[0041] Specifically, such as Figure 2 and Figure 3 As shown, the stereo attention module (CAM) further refines the spatial-spectral features in three dimensions—space, channel, and depth—to obtain the target spatial-spectral features. ; ; ; ; ; in, and These represent mean pooling and max pooling operations, respectively. This represents a shared multilayer perceptron module. for function, Indicates the filter size is Convolution operation, Indicates the spatial spectral characteristics of the target. This indicates element-wise multiplication.
[0042] Specifically, embedding pixels Attention map in CAM Perform dot product calculations to achieve embedding weighting operations and obtain the target pixel-level embedding. This will fully integrate the spectral and spatial information of hyperspectral images, enhancing the model's ability to perceive different regions within the hyperspectral images, strengthening the contour features of trees, and helping to capture tree contour information more accurately.
[0043] ; in, Indicates target pixel-level embedding, Represents pixel-level features. This represents an attention image.
[0044] In other words, this application embodiment designs a stereo attention module. This module fully considers the importance of depth, spatial, and channel information, enhancing the model's perception of the spectrum, local structure, and shape of different trees, and comprehensively extracting the spectral differences and spatial contour details (crown outline and morphological features) of different individual trees. The stereo attention module helps improve the robustness of the model, making it more adaptable to different lighting, scene, and tree species variations.
[0045] In one implementation, the meticulous attention module combines a self-attention mechanism and a dilated convolution module to process global and local information simultaneously within the same network structure, thereby more effectively solving the tree segmentation problem of canopies at different scales. Specifically, this module uses a 1×1 two-dimensional convolutional layer to recode the embedded features. Next, the features are flattened in the spatial dimension to create the input sequence for the Transformer encoder. The encoder generates features that maintain relative positional invariance by modeling global relationships within this sequence. Simultaneously, by using dilated convolutional layers with varying void ratios, multi-scale receptive fields can be obtained, thereby generating features. This helps capture detailed features of trees at different scales, such as leaf texture and small-scale structure. The two outputs are concatenated through dense connections, and the module's input information is introduced through skip connections to obtain the module's output. .in, This indicates recoded features. This indicates a channel dimension splicing operation. Indicates a flat or extended feature. Represents the standard after removing position coding Encoder module, Indicate detailed features, Represents the dilated convolution function. and These are operations for flattening and restoring the vector in space, respectively. This represents the first output feature.
[0046] Understandably, the self-attention mechanism in the meticulous attention module can be replaced with something like transformer or ViT.
[0047] In other words, this application's embodiments employ a meticulously designed self-attention module. This module combines the self-attention mechanism with a dilated convolution module, enabling the model to process both global and local information simultaneously within the same network structure. This allows for a more effective solution to tree segmentation problems across canopies at different scales. When processing hyperspectral images of tree species, the self-attention mechanism helps the model consider the structure and layout of different trees throughout a dense region, while the dilated convolution module, through varying dilation rates, can obtain receptive fields at multiple scales. This facilitates the capture of detailed features of trees at different scales, such as leaf texture and small-scale structures.
[0048] In step S102, feature extraction is performed on the first output feature to obtain a two-dimensional feature, and the two-dimensional feature is input into the two-dimensional attention module to generate a two-dimensional attention feature. The two-dimensional attention feature is then input into the meticulous attention module to generate a second output feature.
[0049] In one implementation, such as Figure 2 As shown, the input features are passed to a two-dimensional convolutional layer, and after processing, two-dimensional features are obtained. , Next, channel attention and spatial attention operations are performed using this two-dimensional feature to further extract feature information.
[0050] ; ; ; in, Representing two-dimensional attention features, Representing two-dimensional features, Represents two-dimensional channel attention. This module represents two-dimensional spatial attention. Through spatial attention, the model can focus on features in different regions of an image, thereby better understanding the shape and structure of trees. Channel attention helps the model learn the relationships between different channels, better capturing detailed features such as texture and color of trees. By combining spatial and channel attention, the model can more accurately segment different types of trees, improving the accuracy and efficiency of tree species segmentation.
[0051] Subsequently, similar to the second step, the features acquired from the 2D attention module are passed to the comprehensive attention module. The comprehensive attention module further processes these features to extract more discriminative second output features. .
[0052] Specifically, a two-dimensional convolutional layer with a kernel size of 1×1 is used to re-encode the embedded features. Next, the features are flattened in the spatial dimension to create the input sequence for the Transformer encoder. The encoder generates features that maintain relative positional invariance by modeling global relationships within this sequence. Simultaneously, by using dilated convolutional layers with varying void ratios, multi-scale receptive fields can be obtained, thereby generating features. This helps capture detailed features of trees at different scales, such as leaf texture and small-scale structure. The two outputs are concatenated through dense connections, and the module's input information is introduced through skip connections to obtain the module's output. .in, Represents two-dimensional recoding features. This indicates a channel dimension splicing operation. Represents two-dimensional flattened features, Represents the standard after removing position coding Encoder module, Representing two-dimensional detail features, Represents the dilated convolution function. and These are operations for flattening and restoring the vector in space, respectively. This represents the second output feature.
[0053] Understandably, the spatial feature extraction of the 2D attention module can be replaced by, for example, the Squeeze-and-Excitation (SE) block. This enhances the model's focus on important features, thereby improving performance.
[0054] In step S103, the second output feature is input to the segmentation module to generate segmentation and recognition images of different individual trees in the forest.
[0055] In one implementation, low-resolution features extracted from the image are input into a pixel decoder, which is then progressively upsampled to generate a feature pyramid composed of high-resolution features at resolutions of 1 / 32, 1 / 16, and 1 / 8 of the original image. These high-resolution features provide more detailed information, making it easier for the model to distinguish and recognize small-scale objects (such as small-scale tree canopies). Subsequently, the different scale features are fed into different Transformer decoder layers. A sinusoidal position embedding is added for each resolution. and a learnable scale-level embedding Therefore, the first three layers achieve the desired resolution. = / 32, = / 16, = / 8 and = / 32, = / 16, = / 8 feature map, where and This is the original image resolution. This 3-layer Transformer decoder is repeated L times, resulting in a final Transformer decoder with 3L layers.
[0056] The Transformer decoder utilizes image features to process object queries. The final binary mask prediction is decoded from the pixel-by-pixel embeddings via object queries. A key component of the Transformer decoder includes a masked attention operator, which extracts local features by confining cross-attention to the foreground region of the predicted mask for each query, rather than focusing on the entire feature map. Masked attention modulates the attention matrix by adding mask values; specifically, ; ; in, For the number of floors, For the first Layer indivual dimensional features, , and These represent the information to be queried, the information being queried, and the value obtained from the query, respectively. It was before The Transformer decoder adjusts the binarized output of the mask prediction, which is adjusted to be similar to... Same resolution. From The binary mask prediction obtained is before the query features are input into the Transformer decoder.
[0057] It should be noted that the segmentation and recognition step in step S103 is the same as the process for processing the same image in the prior art, and will not be described again here. The second output feature obtained by the multi-level attention connection module and the spatial attention enhancement module in steps S101 and S102 of the present invention enables the tree segmentation image generated by the recognition module to have a high segmentation accuracy.
[0058] Understandably, segmentation and recognition operations can be replaced, such as Maskformer, K-Net, etc.
[0059] Next, referring to the accompanying drawings, a tree segmentation system for dense areas of hyperspectral images according to an embodiment of this application is described.
[0060] Figure 4 This is a block diagram of a tree segmentation system for densely populated areas in a hyperspectral image, according to an embodiment of this application.
[0061] like Figure 4 As shown, the tree segmentation system 10 for dense areas of hyperspectral images includes: a multi-level attention connection module 100, a spatial attention enhancement module 200, and a segmentation module 300.
[0062] Specifically, the multi-level attention connection module 100 is used to acquire the original image of the forest, extract features from the original image to obtain pixel-level features and three-dimensional features, input the three-dimensional features and the pixel-level features into the stereo attention module to generate three-dimensional attention features, and input the three-dimensional attention features into the comprehensive attention module to generate the first output feature. The spatial attention enhancement module 200 is used to extract features from the first output feature to obtain two-dimensional features, and input the two-dimensional features into the two-dimensional attention module to generate two-dimensional attention features, and input the two-dimensional attention features into the comprehensive attention module to generate a second output feature; The segmentation module 300 is used to input the second output feature into the segmentation module to generate segmentation and recognition images of different individual trees in the forest.
[0063] This solves the technical problem of inaccurate segmentation edges in tree species segmentation methods when dealing with dense forests or high-density forests.
[0064] Figure 5 A schematic diagram of the structure of a terminal provided in an embodiment of this application. The terminal may include: The memory 501, the processor 502, and the computer program stored on the memory 501 and capable of running on the processor 502.
[0065] When the processor 502 executes the program, it implements the method for segmenting individual trees in dense areas of hyperspectral images provided in the above embodiments.
[0066] Furthermore, the terminal also includes: Communication interface 503 is used for communication between memory 501 and processor 502.
[0067] The memory 501 is used to store computer programs that can run on the processor 502.
[0068] Memory 501 may include high-speed RAM memory, and may also include non-volatile memory. volatile memory), for example, at least one disk storage.
[0069] If the memory 501, processor 502, and communication interface 503 are implemented independently, then the communication interface 503, memory 501, and processor 502 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EIS) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 5 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0070] Optionally, in a specific implementation, if the memory 501, processor 502, and communication interface 503 are integrated on a single chip, then the memory 501, processor 502, and communication interface 503 can communicate with each other through an internal interface.
[0071] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of this application.
[0072] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for segmenting individual trees in dense regions of a hyperspectral image.
[0073] One embodiment of this application provides a computer program product, including a computer program that, when executed by a processor, implements the features described in this application. Figure 1 The corresponding embodiments provide a method for segmenting individual trees in dense areas of hyperspectral images.
[0074] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0075] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0076] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0077] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable storage medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable storage medium could be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0078] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0079] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0080] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0081] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
[0082] It should be understood that the application of this application is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for segmenting individual trees in dense regions of hyperspectral images, characterized in that, The method for segmenting individual trees in dense regions of hyperspectral images includes: The original image of the forest is acquired, and feature extraction is performed on the original image to obtain pixel-level features and three-dimensional features. The three-dimensional features and the pixel-level features are then input into a stereo attention module to generate three-dimensional attention features. The three-dimensional attention features are then input into a comprehensive attention module to generate the first output feature. The three-dimensional attention features are used to reflect the spectral differences and spatial contour details of different individual trees. Feature extraction is performed on the first output feature to obtain a two-dimensional feature, and the two-dimensional feature is input into the two-dimensional attention module to generate a two-dimensional attention feature. The two-dimensional attention feature is then input into the meticulous attention module to generate a second output feature. The second output feature is input into the segmentation module to generate segmentation and recognition images of different individual trees in the forest; The three-dimensional feature is a spatial-spectral feature, and the three-dimensional attention feature includes target spatial-spectral features and target pixel-level embedding; The step of inputting the three-dimensional features and the original image into the stereo attention module to generate three-dimensional attention features specifically includes: The stereo attention module processes the spatial spectral features in the depth dimension, channel dimension, and spatial dimension respectively to obtain three-dimensional depth attention, three-dimensional channel attention, and three-dimensional spatial attention, and generates the target spatial spectral features of the tree based on the spatial spectral features, the three-dimensional depth attention, the three-dimensional channel attention, and the three-dimensional spatial attention. The stereo attention module acquires the attention image corresponding to the three-dimensional feature, and performs weighted processing on the pixel-level feature and the attention image to obtain the target pixel-level embedding; The meticulous attention module includes a dilated convolution module and a self-attention mechanism; The step of inputting the three-dimensional attention features into the meticulous attention module to generate the first output feature specifically includes: The target spatial spectral features and the target pixel-level embedding are recoded to generate recoded features; The self-attention mechanism flattens the recoded features in the spatial dimension to generate flattened features, and the dilated convolution module generates detailed features based on the recoded features. The flattened feature and the detailed feature are spliced and fused together to obtain the first output feature; The formula for calculating the recoded features is as follows: ; The calculation formula for the flattening feature is as follows: ; The calculation formula for the detailed features is as follows: ; The formula for calculating the first output feature is as follows: ; in, This indicates recoded features. Indicates the filter size is Convolution operation, This indicates a channel dimension splicing operation. Indicates a flat or extended feature. Represents the standard after removing position coding Encoder module, Indicate detailed features, Indicates the filter size is dilated convolution function, and These are operations for flattening and restoring the vector in space, respectively. This represents the first output feature.
2. The method for segmenting individual trees in dense areas of hyperspectral images according to claim 1, characterized in that, The original image is a hyperspectral image sample; The step of extracting features from the original image to obtain pixel-level features and three-dimensional features specifically includes: The hyperspectral image samples are input into a linear layer to obtain pixel-level features; The hyperspectral image samples are input into a three-dimensional convolutional layer to obtain three-dimensional features.
3. The method for segmenting individual trees in dense regions of hyperspectral images according to claim 1, characterized in that, The stereo attention module processes the spatial spectral features in the depth, channel, and spatial dimensions respectively to obtain three-dimensional depth attention, three-dimensional channel attention, and three-dimensional spatial attention. Based on the spatial spectral features, the three-dimensional depth attention, the three-dimensional channel attention, and the three-dimensional spatial attention, the target spatial spectral features of the tree are generated, specifically including: The stereo attention module performs mean pooling on the spatial spectral features in the depth dimension to obtain three-dimensional depth attention; the formula for the three-dimensional depth attention is as follows: ; The 3D attention module performs mean pooling and max pooling on the spatial spectral features in the channel dimension to obtain 3D channel attention; the formula for the 3D channel attention is as follows: ; The stereo attention module performs mean pooling and max pooling on the spatial spectral features in the spatial dimension to obtain three-dimensional spatial attention; the formula for the three-dimensional spatial attention is as follows: ; The calculation formula for the target spatial spectral features is as follows: ; in, Indicates spatial spectral characteristics, Represents 3D depth attention. for function, This represents a shared multilayer perceptron module. and These represent mean pooling and max pooling operations, respectively. Represents 3D channel attention. Represents attention in three-dimensional space. Indicates the filter size is Convolution operation, Indicates the spatial spectral characteristics of the target. This represents element-wise multiplication; The calculation formula for the target pixel-level embedding is as follows: ; in, Indicates target pixel-level embedding, Represents pixel-level features. This represents an attention image.
4. The method for segmenting individual trees in dense areas of hyperspectral images according to claim 1, characterized in that, The step of inputting the two-dimensional features into the two-dimensional attention module to generate two-dimensional attention features specifically includes: The two-dimensional attention module performs mean pooling and max pooling on the two-dimensional features along the channel dimension to obtain two-dimensional channel attention; the formula for the two-dimensional channel attention is as follows: ; The two-dimensional attention module performs mean pooling and max pooling on the two-dimensional features in the spatial dimension to obtain two-dimensional spatial attention; the formula for the three-dimensional spatial attention is as follows: ; The formula for calculating the two-dimensional attention features is as follows: ; in, Representing two-dimensional attention features, Representing two-dimensional features, Represents two-dimensional channel attention. Represents attention in two-dimensional space. The kernel size is The convolutional layer.
5. A tree segmentation system for densely populated areas in hyperspectral images, characterized in that, The hyperspectral image dense region tree segmentation system is used to implement the hyperspectral image dense region tree segmentation method according to any one of claims 1-4, and the hyperspectral image dense region tree segmentation system includes: A multi-level attention connection module is used to acquire the original image of the forest, extract features from the original image to obtain pixel-level features and three-dimensional features, and input the three-dimensional features and the pixel-level features into a stereo attention module to generate three-dimensional attention features. The three-dimensional attention features are then input into a comprehensive attention module to generate the first output feature. The spatial attention enhancement module is used to extract features from the first output feature to obtain two-dimensional features, and input the two-dimensional features into the two-dimensional attention module to generate two-dimensional attention features. The two-dimensional attention features are then input into the comprehensive attention module to generate the second output feature. The segmentation module is used to input the second output feature into the segmentation module to generate segmentation and recognition images of different individual trees in the forest.
6. A terminal, characterized in that, The terminal includes: a memory, a processor, and a hyperspectral image dense area tree segmentation program stored in the memory and executable on the processor. When the hyperspectral image dense area tree segmentation program is executed by the processor, it implements the steps of the hyperspectral image dense area tree segmentation method as described in any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a tree segmentation program for dense areas of hyperspectral images, which, when executed by a processor, implements the steps of the tree segmentation method for dense areas of hyperspectral images as described in any one of claims 1-4.