Multi-spectral tree species classification method based on separable cavity ASPP

Through the multi-spectral tree species classification method based on separable cavity ASPP, the images are acquired by drones and data preprocessing and model optimization are solved, and the problems of traditional tree species recognition are achieved with low efficiency and poor accuracy, and efficient tree species classification is achieved.

CN120279338AActive Publication Date: 2025-07-08NORTH CHINA INST OF AEROSPACE ENG
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510573915.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-08
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

In the prior art, traditional tree species recognition methods have low efficiency and poor accuracy, small sample data sets of deep learning models lead to insufficient generalization capabilities, DeepLabv3+ network computing cost is high and the recognition accuracy is insufficient, and feature fusion effect is poor.

Method used

Using a multi-spectral tree species classification method based on separable hollow ASPP, images are acquired through drones and data preprocessing are performed. Combined with principal component analysis and adaptive noise cancellation technology, the DeepLabV3+ model is optimized using the EnhancedMobileNetV2 module, SEBlock module, SeparableASPP module and FPNFusion feature fusion module to optimize the DeepLabV3+ model for data enhancement, feature extraction and fusion.

Benefits of technology

It effectively eliminates noise interference, improves computing efficiency and image quality, enhances data set diversity and model robustness, and improves the accuracy and efficiency of tree species classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279338A_ABST
    Figure CN120279338A_ABST
Patent Text Reader

Abstract

The invention discloses a multispectral tree species classification method based on separable cavity ASPP, and relates to the technical field of semantic segmentation. According to the method, a PCA process and ANR are combined, on the basis of principal components of data, high-frequency noise is removed through an adaptive technology, utilization of near-infrared band information is enhanced, the quality of the data is improved, and then the feature expression of tree species to be classified is enhanced; according to the method, a basic DeepLabV3 + model is improved and used for semantic segmentation, a plurality of modules including an EnhanceMobi leNetV2 module, an SEBlock module, a SeparableASPP module and an FPNFus ion feature fusion module are introduced, and accurate classification of different tree species is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic segmentation, specifically a multi-spectral tree species classification method based on separable atrous ASPP. Background Art

[0002] In the field of forest resource monitoring, accurate tree species classification is of great significance for ecological research, rational planning of forest resources, and sustainable development. Different tree species play different functions in the ecosystem. Accurately identifying and classifying tree species can provide key scientific basis for quantifying the ecological contributions of tree species and optimizing the forest community structure, thus promoting the scientific management and protection of forest ecosystems.

[0003] Traditional forest surveys rely on manual fieldwork, which has problems such as long cycle, difficulty in accurate statistics in densely forested areas, and inability to comprehensively monitor large areas of forest resources in a timely manner. The application of deep learning algorithms brings new opportunities. Among them, the DeepLabv3+ network is widely used, and unmanned aerial vehicle (UAV) multi-spectral images also provide rich data for tree species classification. However, the existing technologies still face some challenges: traditional tree species recognition methods are inefficient and have poor accuracy; small deep learning sample data sets lead to insufficient model generalization ability; the DeepLabv3+ network model has high computational costs, insufficient recognition accuracy, and poor feature fusion effects. Therefore, there is an urgent need for a new method to solve these problems and improve the accuracy and efficiency of tree species classification. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-spectral tree species classification method based on separable atrous ASPP to solve the problems raised in the existing technologies.

[0005] To achieve the above purpose, the present invention provides the following technical solution, a multi-spectral tree species classification method based on separable atrous ASPP, the classification method includes the following steps:

[0006] Step S1: Obtain tree species images through an unmanned aerial vehicle; perform data preprocessing on the tree species images to obtain label images;

[0007] Step S2: Perform data augmentation on the label images to obtain a sample data set;

[0008] Step S3: Introduce an EnhancedMobileNetV2 module, an SEBlock module, a SeparableASPP module, and an FPNFusion feature fusion module into the basic DeepLabV3+ model to obtain an optimized DeepLabV3+ model; input the sample data set into the optimized DeepLabV3+ model to obtain a semantic segmentation result.

[0009] Further, step S1 includes:

[0010] Step S1-1: Collect tree species images through a drone, splice the tree species images collected by the drone, and obtain a complete tree species image in tif format;

[0011] Step S1-2: Use the combination of principal component analysis technology and adaptive noise cancellation technology to preprocess the complete tree species image;

[0012] Step S1-3: Crop the preprocessed complete tree species image according to a preset first size;

[0013] Step S1-4: Perform data annotation on the cropped complete tree species image, and convert the complete tree species image after data annotation from tif format to json format to obtain a labeled image;

[0014] Since the multispectral image consists of multiple spectral bands, each spectral band corresponds to a specific spectral range, including visible light (red, green, blue), near-infrared, and other bands; in the multispectral image, the near-infrared band contains a large amount of useful information, but the near-infrared band is also easily affected by noise interference. In order to effectively utilize the information in the near-infrared band, the present invention adopts the method of principal component analysis and adaptive noise removal (PCA-ANR) to process the band information of the multispectral image.

[0015] Further, the process of preprocessing the complete tree species image includes:

[0016] Standardize the data of the complete tree species image; calculate the covariance between every two bands to obtain a covariance matrix, and perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues;

[0017] Select the principal components greater than the eigenvalues for data reconstruction to remove the noise part, so that the complete tree species image is displayed in the near-infrared manner;

[0018] The above-mentioned standardization of the data of the complete tree species image aims to make the mean of each band 0 and the variance 1; select the principal components greater than the eigenvalues and discard the principal components with smaller eigenvalues, because the principal components with smaller eigenvalues correspond to noise, and the principal components with larger eigenvalues correspond to meaningful information; after such processing, the weight of the near-infrared band in the image is more, and then the image is displayed in the near-infrared manner.

[0019] Further, step S2 includes:

[0020] Step S2-1: Further crop the complete tree species image and label image that have been cropped according to the first size by the bilinear interpolation method according to a preset second size, and convert the complete tree species image and label image that have been cropped according to the second size into both jpg format and png format at the same time to obtain the scaled image data and label data;

[0021] Step S2-2: Linearly weighted synthesize the image data and the label data through the Mixup technology to obtain a synthesized image; perform geometric transformation operations on the synthesized image to obtain a sample data set;

[0022] The above steps are mainly for the label image preprocessed by the PCA-ANR method to expand the sample size of the data set; the Mixup technology randomly generates a proportionality coefficient (the value range of the proportionality coefficient is between 0 and 1) by selecting two images, linearly interpolates and mixes the two images according to the ratio, and then linearly weighted synthesizes the two images, and the target labels are also weighted; the above steps are used to enhance the diversity of the data set and avoid overfitting.

[0023] Further, step S3 includes:

[0024] The sample data set extracts low-level features and high-level features through the EnhancedMobileNetV2 module; the SEBlock module processes the extracted features by adjusting the feature responses of each channel; the features processed by the SEBlock module are passed to the SeparableASPP module, and the SeparableASPP module uses preset depthwise separable atrous convolutions to extract spatial context information; the FPNFusion feature fusion module generates a feature map by upsampling layer by layer and fusing shallow features and deep features; the feature map passes through the DeepLabV3+ decoder to obtain the semantic segmentation result.

[0025] Further, the EnhancedMobileNetV2 module includes

[0026] Extract low-level features and high-level features through an inverted residual structure; among them, the inverted residual structure expands the input feature map through an expansion layer using a 1×1 convolution kernel, and then performs a depthwise separable convolution operation on the expanded feature map using a 3×3 convolution kernel through a depthwise separable convolution layer; the dimensionality reduction layer performs a linear projection operation on the feature map after the depthwise separable convolution operation using a 1×1 convolution kernel; among them, the 1×1 convolution kernel of the dimensionality reduction layer does not contain an activation function; among them, the depthwise separable convolution layer performs an independent convolution operation on each input channel using depthwise convolution and maps the output of the depthwise convolution to a new channel space using a 1×1 pointwise convolution;

[0027] The above-mentioned EnhancedMobileNetV2 module is an enhanced version based on the MobileNetV2 module. It reduces the network computing overhead through the inverted residual structure and depthwise separable convolution, and adds SEBlock to both shallow and deep features to enhance the channel weights and improve the model's attention to important features.

[0028] Furthermore, the SEBlock module is integrated into the low-level and high-level features of the backbone network of the EnhancedMobileNetV2 module; among them, the processing of the extracted features by the SEBlock module includes the following steps:

[0029] Step S7-1: Compress the input feature map in the spatial dimension through global average pooling to generate global statistics for each channel;

[0030] Step S7-2: Learn the relationship between channels through two convolutional layers; among them, the first convolutional layer uses the reduction parameter to control the compression ratio to reduce the number of channels of the global statistics, and uses the activation function ReLU to strengthen the channels; the second convolutional layer restores the number of channels of the global statistics to the initial size and restricts the channel weights to between [0,1] through the activation function to obtain the channel weights for each channel;

[0031] Step S7-3: Weight the input feature map using the obtained channel weights;

[0032] The above-mentioned SEBlock module is integrated into the low-level and high-level features of the backbone network of the EnhancedMobileNetV2 module, which can enhance the expression ability of the feature map.

[0033] Furthermore, the SeparableASPP module includes:

[0034] Step S8-1: Extract multi-scale features using depthwise separable convolution; among them, depthwise separable convolution consists of depth convolution and pointwise convolution; among them, depth convolution performs convolution operations on each input channel; pointwise convolution mixes information between channels through a 1×1 convolution kernel;

[0035] Step S8-2: Obtain global context information through global average pooling, and perform convolution processing and upsampling processing; among them, by calculating the average value of all pixels in each channel, the spatial dimension information is compressed into a single value to extract global context information; after the input feature map is globally averaged pooled, each channel of the output only retains one average value, and each of the average values represents the global information of the corresponding channel; after global average pooling, the features are processed through a 1×1 convolution kernel and upsampled to the initial size;

[0036] Step S8-3: Concatenate the multi-scale features extracted by depthwise separable convolution and the global context information obtained by global average pooling along the channel dimension to obtain the concatenated features; use a 1×1 convolutional kernel to fuse the concatenated features;

[0037] The above steps use depthwise separable convolution to extract multi-scale features, mainly using depthwise separable convolution with different dilation rates (such as 6, 12, 18) to extract multi-scale features, so as to capture features from spatial regions of different scales.

[0038] Furthermore, the FPNFusion feature fusion module is used to fuse the features output by the SeparableASPP module and the EnhancedMobileNetV2 module, including the following steps:

[0039] Step S9-1: Process the low-level features through a 1×1 convolutional kernel, convert the number of channels of the low-level features from 24 to 48 to match the number of channels of the high-level features;

[0040] Step S9-2: Use bilinear interpolation for upsampling to upsample the high-level features to the same spatial resolution as the low-level features; use a concatenation function to concatenate the high-level features and the low-level features to obtain a fused feature map;

[0041] Step S9-3: Fuse the fused feature map through a 3×3 convolutional kernel to obtain a feature map with spatial details and semantic information;

[0042] In the above steps, since the semantic features of the high level have a smaller spatial resolution, while the spatial detail features of the low level have a larger spatial resolution, bilinear interpolation is used for upsampling.

[0043] Compared with the prior art, the beneficial effects of the present invention are:

[0044] 1. The image preprocessing part adopts principal component analysis and adaptive noise removal (PCA-ANR). Compared with traditional methods such as median filtering and wavelet transform, it effectively removes the interference caused by sensor noise and outliers, and dynamically adjusts parameters adaptively to match the spatial heterogeneity of remote sensing images; this method can more effectively remove the noise interference in the near-infrared band, make the information in the near-infrared band more prominent, thereby better retaining the band information and improving the calculation efficiency.

[0045] 2. The present invention expands the dataset by using the bilinear interpolation method, the Mixup method, and geometric transformation. The bilinear interpolation method not only reduces the image size but also smooths the pixel values of the image, avoiding rough distortion during image scaling. This ensures that even after reducing the image size, the image can still maintain high quality, facilitating subsequent processing. The Mixup method is used to mix two different images to generate multiple new samples with different combinations of image features and labels. By generating synthetic samples in this way, the diversity of training data is increased, enhancing the robustness of the model. The present invention gives full play to the advantages of the bilinear interpolation method and the Mixup method, providing stronger generalization ability and robustness during the data augmentation process.

[0046] 3. The present invention introduces the EnhancedMobileNetV2 module, the SEBlock module, the SeparableASPP module, and the FPNFusion feature fusion module into the basic DeepLabV3+ model. The EnhancedMobileNetV2 module is used as the backbone neural network. Compared with the traditional MobileNetV2 module, the enhanced module can help the network focus on the channels of important information and ignore the less useful channels, thus improving the performance of the model. The SEBlock module is introduced to enhance the expressive ability of the model by adaptively readjusting the channel feature responses, enabling the model to pay more attention to useful features and suppress less useful features. The SeparableASPP module is used to replace the traditional ASPP module. Compared with the traditional ASPP module, depthwise separable convolution not only reduces the computational amount and improves the classification performance but also can effectively extract features with different scales. The FPNFusion feature fusion module is used to output a feature map containing rich information to strengthen the model's segmentation performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a structural diagram of the improved DeepLabV3+ model for the multi-spectral tree species classification method based on separable atrous ASPP of the present invention;

[0048] Figure 2 It is an image before data processing of the data collected by the unmanned aerial vehicle for the multi-spectral tree species classification method based on separable atrous ASPP of the present invention;

[0049] Figure 3 It is an image after data processing of the multi-spectral tree species classification method based on separable atrous ASPP of the present invention;

[0050] Figure 4 It is the original label map of the multi-spectral tree species classification method based on separable atrous ASPP of the present invention;

[0051] Figure 5Prediction result graph of the improved DeepLabV3+ model for the multi-spectral tree species classification method based on separable atrous ASPP of the present invention;

[0052] Figure 6 Prediction result graph of the DeepLabV3+ model which is the basis of the multi-spectral tree species classification method based on separable atrous ASPP of the present invention;

[0053] Figure 7 Model accuracy verification parameter graph after data augmentation of the multi-spectral tree species classification method based on separable atrous ASPP of the present invention;

[0054] Figure 8 Model accuracy verification parameter graph before data augmentation of the multi-spectral tree species classification method based on separable atrous ASPP of the present invention.

[0055] Figure 1 : Image represents the input image, SEBlock represents squeeze-and-excitation, SeparableASPP represents separable atrous pyramid pooling, FPNFusion concat represents feature pyramid network fusion and concatenation operation, Encoder represents the encoder, Decoder represents the decoder, 1×1Conv represents a 1×1 convolutional kernel, Upsample by4 represents upsampling by 4, 3×3Conv represents a 3×3 convolutional kernel, Prediction represents the predicted image;

[0056] Figures 2 - 3 : MPA represents mean pixel accuracy, MioU represents mean intersection over union, TotalLoss represents training loss, ValLoss represents validation loss. Detailed implementation manners

[0057] All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0058] Embodiment: As Figures 1 - 8 shown, the present invention provides a technical solution, a multi-spectral tree species classification method based on separable atrous ASPP, and the classification method includes the following steps:

[0059] Step S1: Obtain tree species images by an unmanned aerial vehicle; perform data preprocessing on the tree species images to obtain label images;

[0060] Among them, step S1 includes:

[0061] Step S1-1: Collect tree species images by an unmanned aerial vehicle, splice the tree species images collected by the unmanned aerial vehicle to obtain a complete tree species image in tif format;

[0062] Step S1-2: Use the combination of principal component analysis technology and adaptive noise cancellation technology to preprocess the complete tree species image;

[0063] Step S1-3: Crop the preprocessed complete tree species image according to a preset first size;

[0064] Step S1-4: Perform data annotation on the cropped complete tree species image, and convert the data-annotated complete tree species image from the tif format to the json format to obtain a labeled image;

[0065] Among them, the process of preprocessing the complete tree species image includes: First, standardize the data of the complete tree species image; calculate the covariance between every two bands to obtain a covariance matrix, and perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues; then, select the principal components greater than the eigenvalues for data reconstruction to remove the noise part, so that the complete tree species image is displayed in the near-infrared manner;

[0066] In the embodiment of the present invention, broad-leaved forests are selected as the research object, a multi-spectral image dataset is collected by an unmanned aerial vehicle, the PCA process is first performed on the collected multi-spectral image dataset, and then, based on the principal components of the data (i.e., the components that retain the most variance), by dynamically adjusting the denoising strategy, optimize for the local features and noise distribution of the data, and achieve adaptive denoising in the unmanned aerial vehicle multi-spectral data; dynamically adjust the neighborhood and weights for different land cover types (such as trees, water bodies, buildings, etc.) and noise sources (sensors, shadows), eliminate the noise in the near-infrared band of the multi-spectral image, and make the image retain as much near-infrared band information as possible. Combine the characteristics of the unmanned aerial vehicle data (such as the number of bands, resolution) to select appropriate parameters, and verify the effect through ground truth to ensure that the data after noise removal meets the requirements for subsequent classification of broad-leaved tree species; the specific effects are as Figure 2 、 Figure 3 shown;

[0067] Step S2: Perform data augmentation on the labeled image to obtain a sample dataset;

[0068] Among them, Step S2 includes:

[0069] Step S2-1: Further crop the complete tree species image and the labeled image cropped according to the first size by the bilinear interpolation method according to a preset second size, and convert both the complete tree species image and the labeled image cropped according to the second size into the jpg format and the png format to obtain scaled image data and labeled data;

[0070] Step S2-2: Linearly weighted synthesis is performed on the image data and the label data through the Mixup technique to obtain a synthesized image; geometric transformation operations are performed on the synthesized image to obtain a sample data set;

[0071] In an embodiment of the present invention, data augmentation operations are performed on the preprocessed tree species data set. By performing transformations such as scaling, horizontal flipping, and rotation on the original images, the data set originally with only 150 images is expanded to 4160 images.

[0072] Step S3: The EnhancedMobileNetV2 module, SEBlock module, SeparableASPP module, and FPNFusion feature fusion module are introduced into the basic DeepLabV3+ model to obtain an optimized DeepLabV3+ model; the sample data set is input into the optimized DeepLabV3+ model to obtain a semantic segmentation result;

[0073] Among them, Step S3 includes:

[0074] The sample data set extracts low-level features and high-level features through the EnhancedMobileNetV2 module; the SEBlock module processes the extracted features by adjusting the feature responses of each channel; the features processed by the SEBlock module are passed to the SeparableASPP module, and the SeparableASPP module uses a preset depthwise separable dilated convolution to extract spatial context information; the FPNFusion feature fusion module generates a feature map by upsampling layer by layer and fusing shallow features and deep features; the feature map passes through the DeepLabV3+ decoder to obtain a semantic segmentation result;

[0075] In an embodiment of the present invention, in the improved DeepLabV3+ model under different computational amounts, by introducing modules such as SEBlock, SeparableASPP, and FPNFusion, the model's ability in feature extraction and fusion is strengthened, and the segmentation accuracy is improved while the performance is enhanced. The specific segmentation results and accuracy verification parameters are as Figure 4 、 Figure 5 、 Figure 6 、 Figure 7 、 Figure 8 shown; among them, by Figure 7 and Figure 8It can be seen that the overall performance of the improved DeepLabV3+ model has been significantly improved in the tree species dataset after data augmentation; among them, the overall accuracy of the improved DeepLabV3+ model in training the dataset after data augmentation reaches 93.73%, which is about 2.84% higher than the overall accuracy of the basic DeepLabV3+ model (DeepLabV3+(mobileV2)) in training the tree species dataset before data augmentation, and about 0.51% higher than the overall accuracy of the basic DeepLabV3+ model in training the tree species dataset after data augmentation; among them, Figure 7 and Figure 8 DeepLabV3+(mobileV2), DeepLabV3+(Xception), Hrnet, and PSPnet in

[0076] are all existing models;

[0077] Among them, the EnhancedMobileNetV2 module includes:

[0078] Extract low-level features and high-level features through an inverted residual structure; among them, after the inverted residual structure passes through an expansion layer, a 1×1 convolutional kernel is used to expand the input feature map, and then a depthwise separable convolutional layer uses a 3×3 convolutional kernel to perform depthwise separable convolution operations on the expanded feature map; the dimensionality reduction layer uses a 1×1 convolutional kernel to perform a linear projection operation on the feature map after the depthwise separable convolution operation; among them, the 1×1 convolutional kernel of the dimensionality reduction layer does not contain an activation function; among them, the depthwise separable convolutional layer uses depth convolution to perform independent convolution operations on each input channel, and uses a 1×1 pointwise convolution to map the output of the depth convolution to a new channel space;

[0079] Step S7-1: Compress the input feature map in the spatial dimension through global average pooling to generate global statistics for each channel;

[0080] Step S7-2: Learn the relationship between channels through two convolutional layers; among them, the first convolutional layer uses the reduction parameter to control the compression ratio to reduce the number of channels of the global statistics, and uses the activation function ReLU to strengthen the channels; the second convolutional layer restores the number of channels of the global statistics to the initial size, and restricts the channel weights to [0,1] through the activation function to obtain the channel weights for each channel;

[0081] Step S7-3: Weight the input feature map using the obtained channel weights;

[0082] Among them, the SeparableASPP module includes:

[0083] Step S8-1: Use depthwise separable convolution to extract multi-scale features; among them, depthwise separable convolution consists of depthwise convolution and pointwise convolution; among them, depthwise convolution performs convolution operations on each input channel; pointwise convolution mixes information between channels through a 1×1 convolution kernel;

[0084] Step S8-2: Obtain global context information through global average pooling, and perform convolution processing and upsampling processing; among them, by calculating the average value of all pixels in each channel, the spatial dimension information is compressed into a single value, thereby extracting global context information; after the input feature map undergoes global average pooling, each channel of the output only retains one average value, and each of the average values represents the global information of the corresponding channel; after global average pooling, the features are processed through a 1×1 convolution kernel and upsampled to the initial size;

[0085] Step S8-3: Concatenate the multi-scale features extracted by depthwise separable convolution and the global context information obtained through global average pooling along the channel dimension to obtain the concatenated features; use a 1×1 convolution kernel to fuse the concatenated features;

[0086] Among them, the FPNFusion feature fusion module is used to fuse the features output by the SeparableASPP module and the EnhancedMobileNetV2 module, including the following steps:

[0087] Step S9-1: Process the low-level features through a 1×1 convolution kernel, convert the number of channels of the low-level features from 24 to 48 to match the number of channels of the high-level features;

[0088] Step S9-2: Use bilinear interpolation for upsampling to upsample the high-level features to the same spatial resolution as the low-level features; use a concatenation function to concatenate the high-level features and the low-level features to obtain the fused feature map;

[0089] Step S9-3: Fuse the fused feature map through a 3×3 convolution kernel to obtain a feature map with spatial details and semantic information.

[0090] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multispectral tree species classification method based on separable atrous ASPP, characterized in that The classification method includes the following steps: Step S1: Obtain a tree species image; perform data preprocessing on the tree species image to obtain a label image; Step S2: Perform data augmentation on the label image to obtain a sample data set; Step S3: Introduce the EnhancedMobileNetV2 module, SEBlock module, SeparableASPP module, and FPNFusion feature fusion module into the basic DeepLabV3+ model to obtain an optimized DeepLabV3+ model; input the sample data set into the optimized DeepLabV3+ model to obtain a semantic segmentation result.

2. The multispectral tree species classification method based on separable hole ASPP according to claim 1, wherein Step S1 includes: Step S1-1: Collect a tree species image by drone, splice the tree species images collected by the drone to obtain a complete tree species image in tif format; Step S1-2: Use the combination of principal component analysis technology and adaptive noise cancellation technology to preprocess the complete tree species image; Step S1-3: Crop the preprocessed complete tree species image according to a preset first size; Step S1-4: Perform data annotation on the cropped complete tree species image, and convert the data-annotated complete tree species image from tif format to json format to obtain a label image.

3. The multispectral tree species classification method based on separable hole ASPP according to claim 2, characterized in that, The process of preprocessing the complete tree species image includes: Normalize the data of the complete tree species image; calculate the covariance between every two bands to obtain a covariance matrix, and perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues; Select the principal components greater than the eigenvalues for data reconstruction to remove the noise part, and display the complete tree species image in the near-infrared mode.

4. The multi-spectral tree species classification method based on separable hole ASPP according to claim 1, characterized in that Step S2 includes: Step S2-1: Further crop the complete tree species image and the label image cropped according to the first size by bilinear interpolation method according to a preset second size, and convert the complete tree species image and the label image cropped according to the second size into jpg format and png format at the same time to obtain scaled image data and label data; Step S2-2: Linearly weighted synthesize the image data and the label data by Mixup technology to obtain a synthesized image; perform geometric transformation operations on the synthesized image to obtain a sample data set.

5. The multispectral tree species classification method based on separable hole ASPP according to claim 1, wherein, Step S3 includes: The sample data set extracts low-level features and high-level features through the EnhancedMobileNetV2 module; the SEBlock module processes the extracted features by adjusting the feature responses of each channel; the features processed by the SEBlock module are passed to the SeparableASPP module, and the SeparableASPP module uses preset depthwise separable atrous convolutions to extract spatial context information; the FPNFusion feature fusion module generates a feature map by upsampling layer by layer and fusing shallow features and deep features; the feature map passes through the DeepLabV3+ decoder to obtain a semantic segmentation result.

6. The multispectral tree species classification method based on separable hole ASPP according to claim 1, wherein, The EnhancedMobileNetV2 module includes: Extract low-level features and high-level features through an inverted residual structure. Among them, after the inverted residual structure passes through an expansion layer and uses a 1×1 convolutional kernel to expand the input feature map, a depthwise separable convolutional layer performs a depthwise separable convolution operation on the expanded feature map using a 3×3 convolutional kernel. The dimensionality reduction layer performs a linear projection operation on the feature map after the depthwise separable convolution operation using a 1×1 convolutional kernel. Among them, the 1×1 convolutional kernel of the dimensionality reduction layer does not include an activation function. Among them, the depthwise separable convolutional layer uses depthwise convolution to perform an independent convolution operation on each input channel and uses a 1×1 pointwise convolution to map the output of the depthwise convolution to a new channel space.

7. The multispectral tree species classification method based on separable hole ASPP according to claim 1, characterized in that The SEBlock module is integrated into the low-level features and high-level features of the backbone network of the EnhancedMobileNetV2 module. Among them, the steps for the SEBlock module to process the extracted features include the following: Step S7-1: Compress the input feature map in the spatial dimension through global average pooling to generate global statistics for each channel. Step S7-2: Learn the relationship between channels through two convolutional layers. Among them, the first convolutional layer uses a reduction parameter to control the compression ratio to reduce the number of channels of the global statistics and uses the activation function ReLU to strengthen the channels. The second convolutional layer restores the number of channels of the global statistics to the initial size and restricts the channel weights to the range [0,1] through the activation function to obtain the channel weights for each channel. Step S7-3: Perform weighted processing on the input feature map using the obtained channel weights.

8. The multispectral tree species classification method based on separable hole ASPP according to claim 1, characterized in that The SeparableASPP module includes: Step S8-1: Use depthwise separable convolution to extract multi-scale features. Among them, depthwise separable convolution consists of depthwise convolution and pointwise convolution. Among them, depthwise convolution performs a convolution operation on each input channel. Pointwise convolution mixes information between channels through a 1×1 convolutional kernel. Step S8-2: Obtain global context information through global average pooling and perform convolution processing and upsampling processing. Among them, by calculating the average value of all pixels in each channel, the spatial dimension information is compressed into a single value to extract global context information. After the input feature map passes through global average pooling, each channel of the output only retains one average value, and each of the average values represents the global information of the corresponding channel. After global average pooling, the feature is processed through a 1×1 convolutional kernel and upsampled to the initial size. Step S8-3: Concatenate the multi-scale features extracted through depthwise separable convolution and the global context information obtained through global average pooling along the channel dimension to obtain the concatenated features. Use a 1×1 convolutional kernel to fuse the concatenated features.

9. The multispectral tree species classification method based on separable atrous ASPP according to claim 1, wherein, The FPNFusion feature fusion module is used to fuse the features output by the SeparableASPP module and the EnhancedMobileNetV2 module, including the following steps: Step S9-1: Process the low-level features through a 1×1 convolutional kernel, convert the number of channels of the low-level features from 24 to 48, and match the number of channels with the high-level features; Step S9-2: Use bilinear interpolation for upsampling, upsample the high-level features to the same spatial resolution as the low-level features; use a concatenation function to concatenate the high-level features and the low-level features to obtain a fused feature map; Step S9-3: Fuse the fused feature map through a 3×3 convolutional kernel to obtain a feature map with spatial details and semantic information.

Citation Information

Patent Citations

  • Multispectral unmanned aerial vehicle remote sensing image crop segmentation method

    CN116503590A

  • Improved lightweight efficient semantic segmentation algorithm based on Deeplabv3 +

    CN118429643A

  • DeeplabV3 +-based unmanned aerial vehicle image multi-crop plot identification and productivity prediction method

    CN119478736A

  • Convolutional neural network-based multi-spectral image coniferous forest tree species classification method

    CN119904701A

  • Image data enhancement method and apparatus, computer device, and storage medium

    US20240265502A1