Tree Species Classification Method and System Based on Spectral Depth Extraction Convolutional Neural Network SDA-CNN
Through the SDA-CNN model combined with two-dimensional and depth-by-depth convolution and adjacent pixel perception module, the problem of low spectral information utilization efficiency in hyperspectral images is solved, and more efficient tree classification performance and accuracy are achieved.
Patent Information
- Application Number
- CN202411493850.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-10-24
AI Technical Summary
The prior art has low spectral information utilization efficiency in hyperspectral images, and it is impossible to effectively extract the effective features of each spectral channel, resulting in limited tree species classification performance.
A tree species classification method based on spectral depth extraction convolution neural network SDA-CNN is adopted, combining the hybrid neural network module of two-dimensional convolution and depth-by-deep convolution, a domain pixel perception module is introduced to capture global spatial information through two-dimensional convolution, and the correlation between spectral channels is captured through depth-by-deep convolution, and local detailed features are extracted using adjacent pixel perception modules.
It improves the performance and accuracy of tree species classification, can better capture the correlation between channels and local details, and improves the automation and accuracy of classification.
Smart Images

Figure CN119273994B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical fields of computer vision and deep learning hyperspectral imaging technology, and in particular to a tree species classification method and system based on a spectral depth extraction convolutional neural network SDA-CNN. Background Art
[0002] As a hub of material cycle and energy exchange, forests have key functions of regulating climate, conserving water sources, and maintaining biodiversity. Conducting forest resource surveys and monitoring is of great significance for the protection and sustainable development of forest resources. To improve the automation and intelligence level of forest tree surveys, this patent aims to conduct refined image classification of forest trees based on hyperspectral drone image data to achieve automatic identification of forest trees, reduce waste of human resources, and achieve a win-win situation for ecological protection and cost savings.
[0003] Traditional tree species identification methods mainly rely on manual field surveys to determine tree species by observing the external morphology of trees. Although this method performs well in terms of accuracy, it has significant limitations in practical applications, especially in areas with inconvenient transportation and complex terrain, where the survey difficulty and risk increase significantly.
[0004] In early forest remote sensing vegetation classification and mapping applications, spectral, vegetation indices, and temporal differences were used as features through the analysis of single-temporal or multi-temporal multispectral images, and forest vegetation classification was achieved through unsupervised clustering (such as K-means), syntactic classification (such as decision trees), or statistical classifiers (such as maximum likelihood classification). However, due to the relatively small spectral differences between forest tree species, traditional methods are difficult to distinguish the subtle spectral differences between different tree species types, and heavily rely on the tuning of artificial algorithm parameters, resulting in problems such as low classification accuracy and fragmented map patches.
[0005] Methods based on deep learning have received extensive attention in remote sensing data classification research. Compared with traditional machine learning methods, deep learning models are more suitable for processing high-dimensional, non-linear, and complex data. When traditional two-dimensional convolution processes hyperspectral data, it focuses on the extraction of spatial features, and the utilization efficiency of spectral information in hyperspectral images is relatively low. Three-dimensional convolution can make up for the deficiencies of two-dimensional convolution when processing hyperspectral data, and improve the classification performance by simultaneously capturing spatial and spectral features. However, its design is not specifically optimized for information redundancy between bands, which results in the model being unable to effectively extract the effective features of each spectral channel, thus limiting the performance of the tree species classification task. Summary of the Invention
[0006] To overcome the above technical defects of low utilization efficiency of spectral information in hyperspectral images and inability to effectively extract effective features of each spectral channel, resulting in limited tree species classification performance, the embodiments of the present application provide a tree species classification method based on spectral depth extraction convolutional neural network SDA-CNN, and its specific technical solutions are as follows:
[0007] S1. Construction of deep learning sample set:
[0008] First, use the PCA method to reduce the dimension of the hyperspectral forest image data. Secondly, label the dimension-reduced image data based on the results of the image-forest information mapping diagram. Finally, divide the labeled image data into a training set, a validation set, and a test set according to a ratio of 1:1:8 to complete the construction of the deep learning sample set;
[0009] S2. Training of SDA-CNN model:
[0010] Use the deep learning sample set constructed by the hyperspectral forest image data to train the spectral depth extraction convolutional neural network SDA-CNN model. Among them, the SDA-CNN model adopts a hybrid neural network module including two-dimensional convolution and depthwise convolution, and a domain pixel perception module is newly set;
[0011] Among them, two-dimensional convolution is used to extract the global spatial information of the image to generate global features; depthwise convolution is used to independently perform convolution operations on each spectral channel without sharing weights between spectral channels to extract the correlation information between spectral channels; the domain pixel perception module is used to extract the environmental information around the target pixel contained in adjacent pixels, and then capture the detailed information of local pixels to improve the model's ability to extract local detailed features.
[0012] S3. Tree species classification based on the SDA-CNN model:
[0013] Input the collected forest image data into the trained SDA-CNN model to classify the tree species in the forest image data, and obtain the tree species classification result.
[0014] Optionally, in a possible implementation manner of the above tree species classification method,
[0015] The hybrid neural network module includes a two-dimensional convolutional layer and a depthwise convolutional layer, and the two convolutional layers are alternately connected to each other.
[0016] Optionally, in a possible implementation manner of the above tree species classification method,
[0017] The domain pixel perception module includes a pointwise convolution layer. In the pointwise convolution layer, pointwise convolution is used to perform convolution operations on all input channels through a 1×1 convolution kernel to fuse the features of different channels.
[0018] Optionally, in a possible implementation of the above tree species classification method,
[0019] The domain pixel perception module further includes a neighboring pixel perception activation layer. After the features fused by the pointwise convolution operation are input into the neighboring pixel perception activation layer, a neighborhood pixel perception activation function is used to process the input features to enhance the extraction of local spatial information in the hyperspectral forest image data.
[0020] Optionally, in a possible implementation of the above tree species classification method,
[0021] The domain pixel perception module further includes an SE layer for performing Squeeze and Excitation operations. The Squeeze operation generates a channel descriptor by aggregating feature maps across spatial dimensions, and the Excitation operation weights the channel descriptor obtained by Squeeze to emphasize important channels.
[0022] Optionally, in a possible implementation of the above tree species classification method,
[0023] The SDA-CNN model further includes a fully connected module composed of two linear layers for mapping the high-dimensional feature space to the probability space of classification labels for classification and outputting the classification result.
[0024] Optionally, in a possible implementation of the above tree species classification method,
[0025] The hyperspectral forest image data is reduced to 30 bands, and one band corresponds to one data dimension.
[0026] Optionally, in a possible implementation of the above tree species classification method,
[0027] The specific steps of data dimensionality reduction include:
[0028] S1.1. Data standardization: Perform mean centering and standard deviation normalization on the original image data to make the scales between different bands the same;
[0029] S1.2. Covariance calculation: Calculate the covariance matrix of the standardized image data to measure the correlation between different bands;
[0030] S1.3. Eigenvalue decomposition: Perform eigenvalue decomposition on the covariance matrix to obtain eigenvectors and eigenvalues;
[0031] S1.4, Principal Component Selection: Select the top k eigenvectors with the largest eigenvalues as the principal components, where k is the expected dimension after dimensionality reduction;
[0032] S1.5, Feature Space Transformation: Multiply the original data by the selected principal components to obtain the representation in the new feature space.
[0033] An embodiment of the present application also provides a tree species classification system, which is applicable to the Spectral Depth Extraction Convolutional Neural Network SDA-CNN and includes:
[0034] A processor and a memory, with a communication connection established between the two;
[0035] The memory is used to store computer instructions, and at the same time, it also stores hyperspectral forest image data and its corresponding deep learning sample set;
[0036] When the computer instructions are called, the processor is caused to execute the technical solution described in any one of the above-mentioned tree species classification methods and their implementation manners.
[0037] The embodiment of the present application adopting the above technical solution can achieve the following technical effects: The Spectral Depth Extraction Convolutional Neural Network SDA-CNN model in the present application is a hybrid convolutional neural network architecture capable of capturing the correlation between channels and local detail features. A hybrid neural network module capable of capturing the correlation between channels and local detail features is constructed using two-dimensional convolution and depthwise convolution. For the rich spectral and spatial information contained in airborne hyperspectral images, a strategy of alternating two-dimensional convolution and depthwise convolution is adopted: two-dimensional convolution is used to capture global features, while depthwise convolution performs independent convolution operations on each channel to capture the correlation information between channels. In addition, a neighboring pixel perception module is introduced in the above SDA-CNN model. By considering the information of neighboring pixels, the understanding and perception ability of the spatial information and local detail features of the input data are enhanced. Finally, for the defects that the utilization efficiency of the spectral information in hyperspectral images is low and the effective features of each spectral channel cannot be effectively extracted, resulting in the limitation of tree species classification performance, the tree species classification performance and classification accuracy are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The drawings exemplarily show the embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0039] Figure 1 It is a schematic flowchart of a tree species classification method based on the Spectral Depth Extraction Convolutional Neural Network SDA-CNN in the present application;
[0040] Figure 2 It is a schematic diagram of a network structure of the SDA-CNN model in this application;
[0041] Figure 3 It is a schematic diagram of a structure of the hybrid neural network module in the SDA-CNN model of this application;
[0042] Figure 4 It is a schematic diagram of a structure of the domain pixel perception module in the SDA-CNN model of this application;
[0043] Figure 5 It is the Gaofeng Forest Farm - A forest image data;
[0044] Figure 6 It is the hyperspectral forest image data of the southern Sierra Nevada in California, USA;
[0045] Figure 7 It is the classification result map of the Gaofeng Forest Farm - A hyperspectral forest dataset;
[0046] Figure 8 It is the classification result map of the hyperspectral forest dataset of the southern Sierra Nevada in California, USA;
[0047] Figure 9 It is a schematic diagram of a structure of the tree species classification system in this application. Detailed implementation manners
[0048] In order to make the purpose, technical solutions and advantages of this application clearer, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of this application.
[0049] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of this application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by this application.
[0050] In the description of this application, it should be understood that the numerical labels before the steps do not indicate the order of execution of the steps. They are only used to facilitate the description of this application and to distinguish each step, and therefore should not be construed as a limitation of this application.
[0051] First, the following are the term explanations related to this application:
[0052] Principal Component Analysis (PCA): Principal Component Analysis (PCA) is a commonly used dimensionality reduction technique that is widely applied in data preprocessing and pattern recognition, especially suitable for the analysis of high-dimensional data such as remote sensing images and genomic data. In the fields of remote sensing and image processing, the main purpose of PCA is to reduce the dimensionality of the data while retaining as much valid information as possible from the original data. The core idea of PCA is to project the original high-dimensional data onto a new set of mutually orthogonal coordinate axes through a linear transformation. These new coordinate axes are called principal components. The principal components are the eigenvectors of the data covariance matrix, sorted by the amount of variance they explain. The first principal component captures the largest variance in the data, the second principal component captures the largest of the remaining variance, and so on.
[0053] Two-dimensional Convolution (2D Convolution): It is one of the core operations in Convolutional Neural Networks (CNNs) and is widely used in the fields of image processing and pattern recognition. Two-dimensional convolution slides a convolutional kernel (filter) over a two-dimensional image to extract local features such as edges and textures. A small matrix, such as a 3x3 or 5x5 weight matrix, is defined, which represents the features we want to extract from the image. The size of the convolutional kernel is usually much smaller than the image. The convolutional kernel starts from the upper left corner of the image and slides step by step according to the defined stride. Each time, the dot product of the area covered by the convolutional kernel is calculated, and the result is stored in a new matrix, which is called the feature map or convolutional output. Let the input image be I and the convolutional kernel be K. The operation of two-dimensional convolution can be expressed as:
[0054] where I(i,j) represents the pixel value of the input image at position (i,j), and K(m,n) represents the weight value of the convolutional kernel.
[0055] Depthwise Convolution: It is a type of convolution method in deep learning and a special convolution operation. The difference between depthwise convolution and standard convolution lies in that the convolution kernel of 2D convolution is in single-channel mode, and it is necessary to perform convolution on each channel of the input, so that the output feature map with the same number of channels as the input feature map can be obtained. Depthwise convolution independently performs convolution operations on each spectral channel without sharing weights between channels. Specifically, depthwise convolution uses multiple independent convolution kernels, and each kernel only acts on one spectral channel. This method ensures the information independence between channels, thereby better retaining and extracting the unique information of each channel. This characteristic enables depthwise convolution to extract high-level features in each spectral channel. Since the information is not mixed between channels, depthwise convolution can refine the feature processing of each channel, helping the model to better capture the deep correlations between spectral channels.
[0056] Neighborhood Pixel Perception Activation Module (NPPM): The neighborhood pixel perception module first uses a pointwise convolution module to efficiently adjust the number of channels of the input feature map while keeping the spatial dimension unchanged through a serial connection method. Subsequently, it uses a neighborhood pixel perception activation function to enhance the extraction of local spatial information of hyperspectral images. Finally, the SELayer adaptively emphasizes important features and suppresses unimportant features, so that the model pays more attention to the features most useful for tree species classification and improves the accuracy of the tree species classification task.
[0057] Secondly, to facilitate those skilled in the art to understand the technical solutions provided in the embodiments of the present application, the related technologies will be described below. The technical solutions in the present application will be described in detail below in conjunction with the accompanying drawings and their specific embodiments, as follows:
[0058] Embodiment 1
[0059] As Figure 1 shown, the tree species classification method based on the Spectral Depth Extraction Convolutional Neural Network (SDA-CNN) in the present application includes the following steps:
[0060] S1. Construction of the deep learning sample set:
[0061] First, use the PCA method to reduce the dimensionality of the hyperspectral forest image data. Secondly, label the dimensionality-reduced image data based on the results of the image-forest information mapping diagram. Finally, divide the labeled image data into a training set, a validation set, and a test set according to a ratio of 1:1:8 to complete the construction of the deep learning sample set.
[0062] Optionally, in an implementation manner of the embodiments of the present application, data dimensionality reduction reduces the hyperspectral forest image data to 30 bands, and one band corresponds to one data dimension. PCA is a statistical technique that projects high-dimensional data into a low-dimensional space while maximizing the preservation of the variance of the original data. Specifically, in this solution, the first 30 principal components are selected, and these principal components together contain 98% of the spectral information in the original data.
[0063] Further optionally, the specific steps for data dimensionality reduction of hyperspectral forest image data include: data standardization, covariance calculation, eigenvalue decomposition, principal component selection, and feature space transformation. The specific technical solution is described in detail in Embodiment 2 below.
[0064] S2. Training of the SDA-CNN model:
[0065] Use the deep learning sample set constructed from the hyperspectral forest image data to train the spectral depth extraction convolutional neural network SDA-CNN model. In the SDA-CNN model, a hybrid neural network module including two-dimensional convolution and depthwise convolution is adopted, and a domain pixel perception module is newly set.
[0066] Specifically, two-dimensional convolution is used to extract the global spatial information of the image to generate global features; depthwise convolution is used to independently perform convolution operations on each spectral channel without sharing weights among the spectral channels to extract the correlation information among the spectral channels; the domain pixel perception module is used to extract the environmental information around the target pixel contained in the neighboring pixels, and then capture the detailed information of the local pixels to improve the model's ability to extract local detailed features.
[0067] Further optionally, the domain pixel perception module includes a pointwise convolution layer, a neighboring pixel perception activation layer, and an SE layer.
[0068] Among them, in the pointwise convolution layer, pointwise convolution is used to perform convolution operations on all input channels through a 1×1 convolution kernel to fuse the features of different channels.
[0069] After the features fused by the pointwise convolution operation are input into the neighboring pixel perception activation layer, the neighborhood pixel perception activation function is used to process the input features to enhance the extraction of local spatial information in the hyperspectral forest image data, so as to realize the function of extracting the environmental information around the target pixel contained in the neighboring pixels.
[0070] The SE layer is used to perform Squeeze and Excitation operations. The Squeeze operation generates a channel descriptor by aggregating feature maps across the spatial dimensions (H×W). The Excitation operation weights the channel descriptor obtained by Squeeze to emphasize important channels. The SE layer can adaptively emphasize important features and suppress unimportant features, thus enabling the model to pay more attention to the features most useful for tree species classification and improving the accuracy of the tree species classification task.
[0071] Further optionally, the SDA-CNN model further includes a fully connected module composed of two linear layers, which is used to map the high-dimensional feature space to the probability space of classification labels for classification and output the classification result.
[0072] S3. Tree species classification based on the SDA-CNN model:
[0073] Input the collected forest image data into the trained SDA-CNN model to classify the tree species in the forest image data and obtain the tree species classification result.
[0074] In the technical solution of the tree species classification method based on the spectral depth extraction convolutional neural network SDA-CNN of the present application, a hybrid neural network module capable of capturing the correlation between channels and local detail features is constructed using two-dimensional convolution and depthwise convolution. For the rich spectral and spatial information contained in airborne hyperspectral images, a strategy of alternating two-dimensional convolution and depthwise convolution is adopted: two-dimensional convolution is used to capture global features, while depthwise convolution performs independent convolution operations on each channel to capture the correlation information between channels. In addition, a neighboring pixel perception module is introduced into the above SDA-CNN model, which enhances the understanding and perception ability of the spatial information and local detail features of the input data by considering the information of neighboring pixels.
[0075] Next, the important technical solutions involved in the first embodiment of the present application will be further described as follows:
[0076] Embodiment 2
[0077] Hyperspectral image data has great potential in many applications due to its rich spectral information, but it is also accompanied by high-dimensional and data redundancy problems. This not only increases the computational complexity but also poses high requirements for storage. Without a doubt, the hyperspectral forest image data in this application also belongs to a type of hyperspectral image data. To solve the above high-dimensional and data redundancy problems and ensure the extraction of the main spectral information, the PCA method is used in this application to reduce the dimension of the hyperspectral forest image to 30 bands. One band corresponds to one data dimension, that is, the first 30 principal components in the hyperspectral forest image data are selected. These principal components together contain a large amount of spectral information in the original data, and their proportion can reach 98%.
[0078] Specifically, the specific steps for dimensionality reduction of hyperspectral forest image data include:
[0079] S1.1. Data standardization: Perform mean centering and standard deviation normalization on the original image data to make the scales between different bands the same.
[0080] Specifically, the following calculation formula is used for data standardization: In the formula, X i is the image of the i-th band, μ i is the mean of the i-th band, σ i is the standard deviation of the i-th band, and the value range of i is [1, 30].[[]END]]
[0081] S1.2. Covariance calculation: Calculate the covariance matrix of the standardized image data to measure the correlation between different bands.
[0082] Specifically, the calculation formula for covariance is: In the formula, C is the covariance matrix, N is the number of bands, X i is the standardized image of the i-th band, and μ is the mean vector.
[0083] S1.3. Eigenvalue decomposition: Perform eigenvalue decomposition on the covariance matrix to obtain eigenvectors and eigenvalues.
[0084] The eigenvectors represent the main directions of the original data, and the eigenvalues represent the importance of the data in these directions. The expression for its decomposition is: C = UΛU T , where U is the eigenvector matrix and Λ is the diagonal eigenvalue matrix.
[0085] S1.4. Principal component selection: Select the first k eigenvectors with the largest eigenvalues as the principal components, where k is the expected dimension after dimensionality reduction.
[0086] Specifically, the principal components are arranged in descending order according to the eigenvalues. Among them, W is a matrix composed of the eigenvectors corresponding to the top k largest eigenvalues, and its expression is: W = U(:, 1:k).
[0087] S1.5. Feature space transformation: Multiply the original data by the selected principal components to obtain the representation in the new feature space.
[0088] Specifically, the representation in the new feature space is: F = XW, where F is the image in the new feature space and X is the original standardized image data.
[0089] By using the above scheme to perform data dimensionality reduction on the hyperspectral forest image data, the problems of high dimensionality and data redundancy existing in the hyperspectral image data can be effectively solved. At the same time, the first 30 principal components in the hyperspectral forest image data are ensured to be extracted to retain a large amount of spectral information in the original data, and even the proportion can reach 98%.
[0090] Example Three
[0091] Deep learning technology can effectively process the complex spectral and spatial information in hyperspectral images, automatically extract features, and achieve high-precision performance in tree species classification. This method greatly improves the accuracy and efficiency of tree species classification on the basis of traditional methods.
[0092] The performance of deep learning in hyperspectral tree species classification is better than that of traditional classification methods, mainly due to its strong feature learning ability. Typical deep learning network architectures include convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), etc. Specifically, a suitable framework is selected according to the task requirements. Research on the hyperspectral image tree species classification of deep learning network models is continuously carried out, and the application scope of the models is continuously expanded.
[0093] This application improves the deep learning network architecture to overcome the disadvantage that traditional convolutional neural networks are mainly designed for natural images, focus on the extraction of spatial features, and have low utilization efficiency of spectral information when processing hyperspectral images.
[0094] Next, in combination with the attached Figure 2 The network structure of the improved SDA-CNN model in this application will be described as follows:
[0095] As Figure 2As shown in the figure, the SDA-CNN model in the embodiments of this application includes: a hybrid neural network module, a domain pixel perception module, and a fully connected module. As described in Embodiment 1 above, the hybrid neural network module includes a two-dimensional convolutional layer and a depthwise convolutional layer, and the two convolutional layers are alternately connected to each other; the domain pixel perception module includes a pointwise convolutional layer, a neighboring pixel perception activation layer, and an SE layer; the fully connected module consists of two linear layers.
[0096] The network structure of the hybrid neural network module is as Figure 3 shown:
[0097] Two-dimensional convolution is an operation commonly used in deep learning for image processing and can be used to extract spatial features in hyperspectral images. Hyperspectral images contain not only a spectral dimension (spectral information of each pixel) but also a spatial dimension (spatial relationship between pixels). Two-dimensional convolution is typically used to process spatial information. The following is the specific process of how two-dimensional convolution extracts features from hyperspectral data:
[0098] A hyperspectral image can be regarded as a three-dimensional data cube, where two dimensions are spatial dimensions (horizontal and vertical coordinates), and the third dimension is the spectral dimension (i.e., spectral information of each pixel). For a two-dimensional convolution operation, the convolutional kernel (filter) K ∈ R K×K×D is applied to the spatial dimensions of the entire image. The convolutional kernel slides over the image to extract local spatial features.
[0099] The operation of two-dimensional convolution can be described as:
[0100]
[0101] where X is the input image, K is the convolutional kernel, and Y is the output feature map
[0102] Although this convolution operation helps capture the global spatial information of the image, it has limitations in extracting independent information between spectral channels. To address this problem, depthwise convolution is introduced.
[0103] Different from two-dimensional convolution, depthwise convolution performs convolution operations independently on each spectral channel without sharing weights between channels.
[0104] Specifically, depthwise convolution uses multiple independent convolutional kernels, and each kernel only acts on one spectral channel. For the input tensor X ∈ R K×K×D , depthwise convolution will apply an independent convolutional kernel to each band to generate the spatial feature map of each band. Its convolution formula is:
[0105]
[0106] In the formula, X dis the input data of the d-th band, K d is the convolution kernel of the d-th band, Y d is the output feature map of the d-th band
[0107] Depthwise convolution specifically processes the spectral dimension of hyperspectral images. The key idea of depthwise convolution is to perform independent convolution operations on each band using a single convolution kernel without mixing information across bands. This approach ensures information independence between channels, thus better preserving and extracting the unique information of each channel. This property enables depthwise convolution to extract high-level features in each spectral channel. Since information is not mixed across channels, depthwise convolution can refine the feature processing for each channel, helping the model better capture the deep correlations between spectral channels.
[0108] As Figure 4 shown, the domain pixel perception module includes a pointwise convolution layer, a neighboring pixel perception activation layer, and an SE layer.
[0109] Neighboring pixels contain the environmental information around the target pixel, and this contextual information helps the network extract local detailed features. First, the pointwise convolution module is used to efficiently adjust the number of channels of the input feature map while keeping the spatial dimensions unchanged. Pointwise convolution performs convolution operations on all input channels through a 1x1 convolution kernel, thereby fusing the features of different channels. For hyperspectral images, each channel usually represents different spectral band information, and pointwise convolution can integrate the feature information of these bands together.
[0110] The specific operation of pointwise convolution can be expressed as:
[0111]
[0112] X d (i,j) represents the value of the d-th channel of the input feature map at position (i, j). K d is a 1×1×D convolution kernel used to generate the value Y(i, j) of the output channel by weighted summation of the values of all input channels.
[0113] The neighborhood pixel perception activation function adopted in this embodiment uses a square window centered on x′ to define the neighborhood range of x′, and then all the pixels n in the window are sequentially labeled as x1, x2, xn, where n is the number of pixels. Denote the NPAF of the input x′ as A(x′), and its calculation process can be summarized as follows:
[0114] y i = f(x i + b c ); In the above two formulas, w irepresents the weight of each activated pixel, f is the basic ReLU activation function, and b c serves as the channel-shared bias.
[0115] The neighborhood pixel-aware activation function is based on the interaction with adjacent pixels. If several adjacent pixels have positive values and larger weights, the value of the pixel can be "restored" from 0, which produces a clustering correction effect. By introducing a neighborhood pixel-aware activation layer after the convolution operation, the module can better capture the detailed information of local pixels, thereby improving the extraction effect of the model on local detailed features.
[0116] Finally, the SE layer (i.e., SELayer) adaptively emphasizes important features and suppresses unimportant features, so that the model pays more attention to the features most useful for tree species classification, improving the accuracy of the tree species classification task.
[0117] The two main operations of the SE module are the Squeeze and Excitation operations. The corresponding formulas are as follows: s = F ex (z, W) = σ(g(z, W)) = σ(W2δ(W1z)), where z represents the input feature vector or feature map, W1 and W2 represent the weight matrices of the fully connected layers, δ represents the ReLU activation function, and σ represents the sigmoid activation function.
[0118] First is the Squeeze operation, which generates a channel descriptor by aggregating the feature maps across the spatial dimensions (H×W).
[0119] Suppose there is an input feature map with dimensions H×W×C. For each channel, a global average pooling operation is performed. Specifically, for the i-th channel, the average value of all spatial positions on this channel is calculated. This can be achieved by taking the average of all elements in each channel, where each element represents the average value on the corresponding channel. After performing the above operation on all channels in the input feature map, C average values will be obtained, forming a vector of length C. Thus, the Squeeze operation compresses the input H×W×C feature map into a C-dimensional vector, where each element represents the average value on the corresponding channel. This vector can be regarded as a "descriptor" of the entire feature map in the channel dimension.
[0120] Next is the Excitation operation, which is the second step in the SE module. Its main purpose is to weight the channel descriptor obtained from the previous Squeeze operation to emphasize important channels.
[0121] Further, the C-dimensional vector obtained by Squeeze is input into a fully connected layer (FC layer) with the aim of learning channel weights. This fully connected layer includes a non-linear activation function, such as ReLU, to introduce non-linear transformation. Through learning, the channel weights obtained by the fully connected layer pass through a Sigmoid activation function to limit their range between 0 and 1. The learned weights can be regarded as the activation degree of each channel. Finally, the original C-dimensional vector is multiplied by the learned channel weights to obtain a weighted vector. This weighted vector reflects the importance and contribution degree of each channel in the task. This process enables the network to dynamically adjust the attention of channels, thereby more effectively using information to complete the task.
[0122] The domain pixel perception module first uses the point convolution module in a serial connection manner to efficiently adjust the number of channels of the input feature map while keeping the spatial dimension unchanged, and then uses the neighborhood pixel perception activation function to enhance the extraction of local spatial information of the hyperspectral image. Finally, the SELayer adaptively emphasizes important features and suppresses unimportant features, so that the model pays more attention to the features most useful for tree species classification and improves the accuracy of the tree species classification task.
[0123] Finally, for the convenience of understanding, the technical solutions in the present application are described below in conjunction with specific examples as follows:
[0124] Exemplarily, in some specific embodiments of deep learning sample construction, the used Gaofeng Forest Farm-A hyperspectral forest dataset, 9 forest vegetation categories, and 3 non-forest vegetation categories such as construction land, roads, and logging sites. The size of this hyperspectral open set is 906 * 572 pixels and has 125 spectral bands. To ensure data independence between the test set and the training set and accurately evaluate the model performance, this study randomly divides the samples into a training set, a validation set, and a test set according to a ratio of 1:1:8, and the sample numbers of each set are shown in Table 1.
[0125] Table 1
[0126]
[0127]
[0128] Further, in a specific SDA-CNN model, 4 convolutional layers, 3 depthwise convolutional layers, 1 domain pixel perception module, 1 flattening layer, and 2 linear connection layers are set. It is easy to know that the flattening layer is the above-mentioned SE layer, and the 2 linear connection layers constitute the fully connected module described above.
[0129] After referring to well-known CNN models such as ImageNet, the number of convolutional kernels is determined by following the rule that the number of the next layer is twice that of the previous layer. The algorithm process and parameters are shown in Table 2.
[0130] Table 2
[0131]
[0132]
[0133] By combining depthwise convolution and two-dimensional convolution, this model enhances the ability to extract relevant features between channels. When traditional two-dimensional convolution is used to process hyperspectral data, it usually mixes the information of all channels, which may lead to the loss or confusion of spectral features. To solve this problem, this study introduces the depthwise convolution module in MobileNet. This module generates feature maps by using independent convolutional kernels on each input channel, thus retaining the independent features of the channels. By alternately using depthwise convolution and general convolution, the model can better capture the deep correlations between channels at different levels.
[0134] Subsequently, a two-dimensional convolution with 256 convolutional kernels is used to combine the extracted spectral features and spatial features for input into the neighborhood pixel perception module, thereby enhancing the extraction of spatial information from hyperspectral images. In further processing, a flattening layer is used to flatten the downsampled feature maps into one-dimensional vectors.
[0135] The ReLU activation function is applied to all convolutional layers and dense layers to introduce non-linear transformations and help extract complex and rich features, where:
[0136] f(x) = max(0, x);
[0137] As can be seen from the above formula, if the input is greater than 0, the output is the original value; if the input is 0 or negative, the output is 0. For values greater than 0, the function is linear. Therefore, when training a neural network using backpropagation, it has the characteristics of a linear activation function. For values less than or equal to 0, the function is non-linear and is also called a piecewise linear function. In terms of the function model, the ReLU function is simpler to calculate, has lower computational costs, can accelerate learning and simplify the model, and avoid problems such as gradient disappearance.
[0138] Finally, multi-class classification is performed through two linear layers. A linear layer, also known as a fully connected layer or a fully connected module, plays a crucial role in classification tasks, especially in the final stage of a neural network. Its working principle is to map the high-dimensional feature space to the probability space of classification labels. In a classification task, the features extracted by the neighborhood pixel perception module are flattened into a feature vector of 30976. This feature vector contains important features of the input data (such as an image or spectrum). It is passed to the linear layer. The operation performed by the linear layer is a simple linear transformation. Given an input vector, the output of the linear layer is calculated by the following formula:
[0139] y = Wx + b; where x is the input feature vector, W is the weight matrix, with size m×n, where m is the number of output classes and n is the number of dimensions of the input features. b is the bias term, which is used to enhance the fitting ability of the model. y is the output of the linear layer and is also known as logits.
[0140] The output y of the linear layer does not directly give the classification result but is an original score (logits). To convert these scores into a probability distribution, an activation function is usually used.
[0141] For multi-class classification problems, the commonly used activation function is the Softmax function:
[0142]
[0143] The Softmax function converts the output of the linear layer into a probability distribution, representing the probability values of each class. These probability values sum to 1.
[0144] Once the probabilities of each class are obtained through the Softmax function, the model can classify based on these probabilities. Usually, the class corresponding to the maximum probability is the prediction result, that is:
[0145] To train the model, we usually use the cross-entropy loss function. It measures the difference between the probability distribution of the model output and the probability distribution of the true labels.
[0146] The formula for cross-entropy loss is: where y i is the one-hot encoding of the true class is the probability value predicted by the model
[0147] Through backpropagation, the error propagates from the loss function back to the linear layer, and the weight matrix W and bias b in the linear layer are adjusted by gradient descent methods (such as SGD, Adam, etc.) to improve the classification accuracy.
[0148] The features are finally classified into 12 categories through the linear layer.
[0149] After determining the network structure, the model training parameters are configured and optimized. According to the best parameters obtained from training, the batch size is 64, the number of epochs is set to 300, the window size is 11, the learning rate is 0.0001, and SGD is used as the optimizer.
[0150] The tree species classification experiment of hyperspectral data for the deep learning network is programmed using Python (3.8), the running hardware is Tesla T4, the Cuda version is 11.8, and all deep learning models are based on the open-source Pytorch (2.0) deep learning framework.
[0151] To comprehensively evaluate the classification accuracy, the overall accuracy (OA), average accuracy (AA), and Kappa coefficient are used as model evaluation indicators.
[0152] Overall accuracy (OA): OA represents the ratio of correctly classified pixels to the total number of pixels in all classes.
[0153] Average accuracy (AA): AA represents the average accuracy of each class.
[0154] Kappa coefficient, as a quantitative evaluation indicator of the model:
[0155] The Kappa coefficient measures the consistency between the predicted classification and the ground truth classification, taking into account the consistency that may occur by chance.
[0156] In addition, precision, recall, and F1 score are used as evaluation indicators to evaluate each tree species:
[0157]
[0158] The higher the accuracy, recall, OA, and F1 score, the closer the predicted value is to the true value.
[0159] The macro-average is used to evaluate the average performance of the model across all classes, and the weighted average reflects the overall performance of the model. The classification accuracy of each tree species using this method is shown in Table 3
[0160] Table 3
[0161] Tree species Precision Recall F1 score Quantity Chinese fir 0.99 0.99 0.99 64594 Masson pine 1.00 0.99 1.00 4272 Slash pine 0.99 0.99 0.99 3340 Eucalyptus grandis × urophylla 1.00 1.00 1.00 11945 Eucalyptus urophylla 1.00 1.00 1.00 32220 Castanopsis hystrix 0.97 0.98 0.98 20369 Mytilaria laosensis 1.00 1.00 1.00 2400 Camellia oleifera 1.00 1.00 1.00 6659 Other broad-leaved trees 1.00 0.98 0.99 8593 Road 0.99 1.00 1.00 8378 Logging site 1.00 1.00 1.00 8136 Construction land 0.96 0.98 0.97 54 Precision 0.99 170960 Macro-average 0.99 0.99 0.99 170960 Weighted average 0.99 0.99 0.99 170960
[0162] Based on the above, this patent realizes the automatic classification of tree species in UAV hyperspectral images based on an improved SDA-CNN network model architecture, achieving high-precision and highly automated vegetation classification.
[0163] To more clearly display the classification results, the following actual research cases illustrate the classification effect of the technical solution of this application. Figure 5 and Figure 6 are hyperspectral images of the study area, where Figure 5 is the hyperspectral forest image data of Gaofeng Forest Farm - A; Figure 6 is the hyperspectral forest image data of the southern Sierra Nevada in California, USA. Correspondingly, the technical solution of this application is used for the above Figure 5 and Figure 6 The classification result maps obtained by classifying the forest data in are shown in Figure 7 The classification result map of the Gaofeng Forest Farm - A hyperspectral forest data set is shown in Figure 8 The classification result map of the hyperspectral forest data set of the southern Sierra Nevada in California, USA is shown in.
[0164] As Figure 7 shown, on the Gaofeng Forest Farm - A hyperspectral forest data set, it can be seen that the classification boundary is clearly extracted and the fragmented patches are reduced, proving the effectiveness of the proposed method in the open set classification of hyperspectral images.
[0165] SDA-CNN model adaptability:
[0166] To further test the stability of the SDA-CNN network structure and determine its applicability, the hyperspectral forest data set of the southern Sierra Nevada in California, USA as shown in Figure 6 is used, and the data is substituted into the SDA-CNN model for tree species classification.
[0167] In this data set, the hyperspectral data includes a wavelength range of 280 - 2510 nm, 426 bands, a sampling interval of 5 nm, and a spatial resolution of 1 m. The hyperparameters are set as follows: the batch size is 64, the learning rate is 0.0001, the number of iterations is 300, and the window size is 13. In the case where the ratio of the training set to the validation set is 6:4, the overall accuracy is 99.71% and the average accuracy is 98.19%. The prediction results of the model are as shown in Figure 8 This study shows the universality and potential of the SDA-CNN model, indicating that it can be applied to multiple fields and tasks rather than being limited to specific applications.
[0168] Example 4
[0169] Figure 9Schematically shown is a hardware architecture diagram of a computer device 10000 suitable for implementing a tree species classification method according to an embodiment of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc. As Figure 9 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can be communicatively linked to each other through a system bus. Among them:
[0170] The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disc, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the tree species classification method. In addition, the memory 10010 may also be used to temporarily store various data that have been output or will be output.
[0171] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.
[0172] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, the Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.
[0173] It should be noted that Figure 9 only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0174] In this embodiment, the tree species classification method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.
[0175] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device, so that they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0176] It should be noted that the above is only the preferred embodiment of the present application, and does not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A tree species classification method based on the spectral depth extraction convolutional neural network SDA-CNN, characterized in that Including: S1. Construction of deep learning sample set: First, use the PCA method to reduce the dimensionality of the hyperspectral forest image data. Secondly, label the dimensionality-reduced image data based on the results of the image-tree information mapping diagram. Finally, divide the labeled image data into a training set, a validation set, and a test set according to the ratio of 1:1:8 to complete the construction of the deep learning sample set; S2. Training of SDA-CNN model: Use the deep learning sample set constructed from the hyperspectral forest image data to train the spectral depth extraction convolutional neural network SDA-CNN model. In the SDA-CNN model, a hybrid neural network module including two-dimensional convolution and depthwise convolution is adopted, and a domain pixel perception module is newly set; Among them, the two-dimensional convolution is used to extract the global spatial information of the image to generate global features; the depthwise convolution is used to independently perform convolution operations on each spectral channel without sharing weights between spectral channels to extract the correlation information between spectral channels; the domain pixel perception module is used to extract the environmental information around the target pixel contained in adjacent pixels, and then capture the detailed information of local pixels to improve the model's ability to extract local detailed features; S3. Tree species classification based on the SDA-CNN model: Input the collected forest image data into the trained SDA-CNN model to classify the tree species in the forest image data and obtain the tree species classification result.
2. The tree species classification method according to claim 1, wherein The hybrid neural network module includes a two-dimensional convolutional layer and a depthwise convolutional layer, and the two convolutional layers are alternately connected to each other.
3. The tree species classification method according to claim 1 or 2, wherein The domain pixel perception module includes a pointwise convolutional layer. In the pointwise convolutional layer, pointwise convolution is used to perform convolution operations on all input channels through a 1×1 convolutional kernel to fuse the features of different channels.
4. The tree species classification method according to claim 3, wherein The domain pixel perception module further includes a neighboring pixel perception activation layer. When the features fused by the pointwise convolution operation are input into the neighboring pixel perception activation layer, a neighborhood pixel perception activation function is used to process the input features to enhance the extraction of local spatial information in the hyperspectral forest image data.
5. The tree species classification method according to claim 4, wherein The domain pixel perception module further includes an SE layer for performing Squeeze and Excitation operations. The Squeeze operation is to generate a channel descriptor by aggregating the feature maps across the spatial dimensions, and the Excitation operation is to weight the channel descriptor obtained by Squeeze to emphasize important channels.
6. The tree species classification method according to claim 1, wherein The SDA-CNN model further includes a fully connected module composed of two linear layers, which is used to map the high-dimensional feature space to the probability space of classification labels for classification and output the classification result.
7. The tree species classification method according to claim 1, wherein The hyperspectral forest image data is reduced to 30 bands, and each band corresponds to one data dimension.
8. The tree species classification method according to claim 7, characterized in that The specific steps of data dimensionality reduction include: S1.
1. Data standardization: Perform mean centering and standard deviation normalization on the original image data to make the scales between different bands the same; S1.
2. Covariance calculation: Calculate the covariance matrix of the standardized image data to measure the correlation between different bands; S1.
3. Eigenvalue decomposition: Perform eigenvalue decomposition on the covariance matrix to obtain eigenvectors and eigenvalues; S1.
4. Principal component selection: Select the top k eigenvectors with the largest eigenvalues as the principal components, where k is the expected dimension after dimensionality reduction; S1.
5. Feature space transformation: Multiply the original data by the selected principal components to obtain the representation in the new feature space.
9. A tree species classification system, characterized in that, The system is applicable to the Spectral Depth-based Convolutional Neural Network SDA-CNN. The system includes: A processor and a memory, with a communication connection established between the two; The memory is used to store computer instructions and also stores the hyperspectral forest image data and its corresponding deep learning sample set; When the computer instructions are called, the processor executes the tree species classification method described in any one of claims 1-8.
Citation Information
Patent Citations
Hyperspectral remote sensing data depth spectral feature extraction method based on one-dimensional group convolutional neural network
CN111062403A
Hyperspectral remote sensing image classification method based on hybrid convolutional neural network
CN115909052A