Hyperspectral tree identification method and system
Through the dual-path convolutional neural network fusion space and spectral features, the high-dimensional data processing and information redundancy problems in hyperspectral tree recognition are solved, and efficient and accurate tree recognition and health assessment are achieved, which is suitable for real-time monitoring of transmission line corridors.
Patent Information
- Application Number
- CN202511055617.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-08-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, high-spectral tree identification has challenges in high-dimensional data processing and information redundancy, insufficient joint modeling of spatial and spectral features, high consumption of deep learning models, insufficient data annotation and training samples, noise and data inconsistency, and multimodal data fusion.
Dual-path convolutional neural network (DPN) is used to combine residual learning and dense connections to extract the morphological characteristics of trees and spectral paths through spatial paths, and analyze their spectral reflection characteristics, and perform feature fusion, and use hyperspectral data for tree recognition.
It improves the accuracy and efficiency of tree identification, can accurately distinguish different tree species in complex environments, provide tree health assessment and disaster prediction, and is suitable for real-time tree monitoring in dynamic monitoring scenarios such as transmission line corridors.
Smart Images

Figure CN120564055A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a hyperspectral tree recognition method and system. Background Art
[0002] Tree management in transmission line corridors is a critical issue for the safe operation of power systems. Trees growing near transmission lines can cause branches to come into contact with power lines, leading to safety hazards such as short circuits, faults, and fires. Therefore, timely and accurate identification and management of these trees is crucial to ensuring the stable operation of power lines. Traditional manual inspection methods are not only inefficient but also susceptible to factors such as weather and terrain, making efficient and comprehensive monitoring difficult. With the development of drone, satellite, and ground-based hyperspectral remote sensing technology, using remote sensing data to identify trees and analyze their distance from transmission lines has become an effective solution. Hyperspectral remote sensing technology has demonstrated significant advantages in vegetation monitoring and tree identification. Unlike traditional RGB images, hyperspectral imagery captures richer spectral information, enabling more precise differentiation of characteristics such as tree species and tree health. However, due to the high dimensionality and complexity of hyperspectral data, effectively extracting key features from this massive amount of spectral information remains a core challenge in hyperspectral tree identification.
[0003] In this context, the application of deep learning technology, especially convolutional neural networks (CNNs) and their derivative models, such as dual-path convolutional neural networks (DPNs), in hyperspectral image processing provides a solution. DPNs combine the advantages of residual learning and dense connections, can effectively process the spatial and spectral features in hyperspectral data, improve the model's feature extraction capabilities, and better cope with noise and redundant information in hyperspectral images. Moreover, DPNs have strong multi-path learning capabilities and can extract features at different levels through multiple paths, thereby improving the model's recognition accuracy for trees in complex environments. This invention, based on a dual-path convolutional neural network, proposes a hyperspectral recognition method for trees in transmission line corridors. By combining a deep learning model with the spectral information and spatial structure characteristics of hyperspectral data, efficient tree recognition can be achieved, thereby providing precise technical support for tree management in transmission line corridors. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by the present invention is: to solve the problems of high-dimensional data processing and information redundancy in the existing technology, insufficient joint modeling of spatial and spectral features, high computing resource consumption of deep learning models, insufficient data annotation and training samples, noise and data inconsistency, poor model interpretability and explainability, and challenges of multimodal data fusion.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: a hyperspectral tree recognition method, comprising: Acquiring hyperspectral image data, and preprocessing the hyperspectral image data to obtain first hyperspectral image data; Performing feature extraction on the first hyperspectral image data, wherein the feature extraction includes spatial feature extraction and spectral feature extraction, and fusing the spatial features and the spectral features to obtain a fused feature map; Based on the fused feature map, the dual-path convolutional neural network model is trained to obtain a first dual-path convolutional neural network model; Acquire real-time hyperspectral image data, perform tree recognition on the real-time hyperspectral image data using the first dual-path convolutional neural network model to obtain a probability that each pixel belongs to a specific tree species, and obtain a first category of the tree species for each pixel based on the probability; The spatial positions of trees are extracted by image segmentation method, and combined with the first category, the tree species category of each pixel is obtained.
[0007] As a preferred solution of the hyperspectral tree recognition method of the present invention, preprocessing the hyperspectral image data includes: The hyperspectral image data is corrected by radiation and atmosphere correction to remove errors, and then denoised by adaptive filters and data enhancement is performed. The data of each band is standardized to obtain standardized hyperspectral image data. The dimensionality reduction method is used to select the bands of the standardized hyperspectral image data to obtain the selected hyperspectral image data; The convolutional autoencoder is used to reduce the dimension of the selected hyperspectral image data to obtain the first hyperspectral image data.
[0008] As a preferred solution of the hyperspectral tree recognition method of the present invention, the feature extraction of the first hyperspectral image data includes: The first hyperspectral image data is divided into image blocks using a superpixel segmentation method, and band features are extracted using principal component analysis; Each image block is processed through a convolutional layer to extract the spatial features of the local area and add an activation function; The spatial features of the local area are pooled to reduce the size of the spatial features of the local area; The spatial features of all local areas are fused to obtain the spatial features.
[0009] As a preferred embodiment of the hyperspectral tree recognition method of the present invention, the method further includes: The reflectance spectrum of each pixel of the first hyperspectral image data is used as the input of the spectral path, the spectral reflectance data of each pixel point is converted using discrete cosine transform, and the band characteristics are extracted using independent component analysis; The spectral data of each pixel is filtered through multiple convolution kernels to extract the characteristics of spectral reflectance, and an activation function is added to obtain the spectral features.
[0010] As a preferred embodiment of the hyperspectral tree recognition method of the present invention, fusing spatial features and spectral features includes: The spatial features and spectral features are added through residual connections and spliced using dense connections to obtain the fused feature map.
[0011] As a preferred embodiment of the hyperspectral tree recognition method of the present invention, the method of using the first dual-path convolutional neural network model to perform tree recognition on the real-time hyperspectral image data includes: Input the fused feature map into the fully connected layer for classification; The original output of the first dual-path convolutional neural network model is converted into a probability distribution through an activation function; The first category of the tree species for each pixel is determined based on the maximum probability output by the activation function.
[0012] As a preferred embodiment of the hyperspectral tree recognition method of the present invention, the tree species category of each pixel includes: Using the semantic segmentation method of deep learning, the encoder part of the network extracts the features of the image layer by layer, and the decoder maps the features to an image with higher spatial resolution to obtain the category label of each pixel; According to the probability of each pixel belonging to a specific tree species, the category probability map is combined with the first category and a threshold is set to obtain the tree species category of each pixel.
[0013] The present invention provides a hyperspectral tree recognition system.
[0014] To solve the above technical problems, the present invention provides the following technical solutions: a hyperspectral tree recognition system, comprising: a preprocessing module, configured to obtain hyperspectral image data, and preprocess the hyperspectral image data to obtain first hyperspectral image data; a feature fusion module, configured to perform feature extraction on the first hyperspectral image data, wherein the feature extraction includes spatial feature extraction and spectral feature extraction, and fuse the spatial features and the spectral features to obtain a fused feature map; A model training module, configured to train a dual-path convolutional neural network model based on the fused feature map to obtain a first dual-path convolutional neural network model; a tree identification module, configured to acquire real-time hyperspectral image data, perform tree identification on the real-time hyperspectral image data using the first dual-path convolutional neural network model, obtain a probability that each pixel belongs to a specific tree species, and obtain a first category of the tree species for each pixel based on the probability; The category determination module is used to extract the spatial position of trees through an image segmentation method, and combine the first category to obtain the tree species category of each pixel.
[0015] The present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and is characterized in that the processor implements the steps of the hyperspectral tree recognition method when executing the computer program.
[0016] The present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the hyperspectral tree recognition method when executed by a processor.
[0017] Beneficial effects of the present invention: The present invention adopts a dual-path convolutional neural network to process spatial information and spectral information respectively, effectively improving the accuracy of tree identification. The spatial path extracts the morphological characteristics of trees through convolution, and the spectral path analyzes its spectral reflectance characteristics. The two complement each other to ensure the comprehensive identification of trees by the model. The deep learning model combined with image segmentation technology can not only identify the overall area of the tree, but also finely distinguish the boundaries between trees, avoiding the ambiguity in traditional methods; through hyperspectral images, the trees are comprehensively monitored and evaluated, not only limited to the identification of a single tree species, but also can provide tree health assessment and disaster prediction and other information. This method is suitable for application in dynamic monitoring scenarios, such as real-time tree monitoring in transmission line corridors. It can regularly update the health status of trees, quickly detect potential risk trees, and take timely measures to ensure the safety of transmission lines. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 This is a logic diagram of the overall process of a hyperspectral tree recognition method according to an embodiment of the present invention; Figure 2 A cross entropy loss function curve diagram of a hyperspectral tree recognition method according to an embodiment of the present invention; Figure 3 This is a flowchart of spatial path and spectral path feature extraction for a hyperspectral tree recognition method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0020] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0021] Example 1, with reference to Figures 1 to 3 Table 1 is an embodiment of the present invention, which provides a hyperspectral tree recognition method, including: S100: Acquire hyperspectral image data, and preprocess the hyperspectral image data to obtain first hyperspectral image data; S200: performing feature extraction on the first hyperspectral image data, the feature extraction including spatial feature extraction and spectral feature extraction, and fusing the spatial features and the spectral features to obtain a fused feature map; S300: Based on the fused feature map, the dual-path convolutional neural network model is trained to obtain a first dual-path convolutional neural network model; Specifically, the first dual-path convolutional neural network model is a trained dual-path convolutional neural network model; The labeled hyperspectral image data is divided into an 80% training set and a 20% validation set to ensure that the data distribution of the training set and the validation set is consistent. The validation set is used to evaluate the generalization ability of the model. The dual-path convolutional neural network model is initialized using He, and its rule is expressed as: , in, For the The weights of the layers, is the number of neurons in the previous layer, is a normal distribution; During the initialization phase, initial weights are randomly assigned to each convolution kernel. This is to break the symmetry in the network and prevent neurons from learning the same features in the early stages of training, thereby improving the convergence speed and expressiveness of the model. Forward propagation is the process of obtaining output from the network through input data during training. The input is standardized hyperspectral image data. The feature maps obtained by processing the spatial path and spectral path of the network are fused. In the final fully connected layer, the network outputs the category to which each pixel belongs based on the fused features. The Softmax function is used to convert the output into a probability representation as follows: , in, For the score of the fir category, is the number of categories, The probability of each category is is the index of the score for the fir category, is the category index, from 1 to K, is the sum of the index scores for all categories; like Figure 2 As shown, for the prediction result of each pixel, the cross entropy loss function calculated with the true label is expressed as: , in, is the true label, is the probability of the network output, is the sample size, Take the natural logarithm of the probability output by the network, is the cross entropy loss function, To sum all samples; Backpropagation is a key step in the training process for updating network weights. Through backpropagation, the network updates the weights of each layer according to the gradient of the loss function to minimize the loss. Calculate the partial derivative of the loss function with respect to each weight, expressed as: , in, is the weight of the current layer, is the cross entropy loss function, is the partial derivative of the cross entropy loss function with respect to the weight; Backpropagate the gradient of each layer through the chain rule and update the weights based on the gradient; The Adam optimizer is used to update the weights. The Adam optimizer combines the advantages of momentum and RMSprop to calculate the adaptive learning rate for each weight. Its update rule is as follows: , in, and are first-order and second-order momentum estimates, initialized to zero, 、 are the decay coefficients of the first-order and second-order momentum estimates, is the learning rate, is a small constant introduced for numerical stability, is the loss function Parameters The gradient information is used to adjust the weight of the model to minimize the loss function. and Respectively for and The bias-corrected estimate is For the The first-order momentum estimate at the iteration, For the The second-order momentum estimate at the iteration, is the square root of the bias-corrected second-order momentum estimate, and are the bias correction factors for the first-order momentum estimation and the second-order momentum estimation, For the The parameters at the iteration, For the Parameters at the iteration; During training, if the gradient is too large and the weights are updated too quickly, it will cause training instability. Use gradient clipping to limit the size of the gradient and avoid gradient explosion. To avoid overfitting, a portion of the neurons in the network are randomly dropped during each training iteration. The dropout rate is usually set to 0.5, that is, 50% of the neurons are randomly dropped during each training to prevent the network from relying on certain specific neurons and improve the generalization ability of the model. Normalizing the input of each layer alleviates the gradient vanishing and gradient exploding problems in model training, increases training speed, and helps prevent overfitting; By adding the L2 regularization term to the loss function, the weight of the model will not be too large, thus suppressing overfitting; The training process typically involves multiple training epochs, each of which iterates through the training dataset. After each epoch, the model's performance on the validation set is evaluated to monitor its generalization capabilities. If overfitting occurs on the validation set, with the training set accuracy significantly higher than the validation set, hyperparameters such as the learning rate and dropout rate may need to be adjusted for optimization.
[0022] It should be noted that data augmentation and regularization techniques effectively prevent overfitting of the model and improve its generalization ability on unknown data. This means that the technology performs well in different environments and on different types of hyperspectral imagery. The use of the Adam optimizer and backpropagation algorithm, combined with batch normalization and gradient clipping, effectively improves the stability of the training process, ensures that the network converges to an optimal solution, and further enhances the robustness of the model.
[0023] S400: Acquire real-time hyperspectral image data, perform tree recognition on the real-time hyperspectral image data using a first dual-path convolutional neural network model, obtain a probability that each pixel belongs to a specific tree species, and obtain a first category of the tree species for each pixel based on the probability; S500: Extracting the spatial position of trees by an image segmentation method, and combining it with the first category to obtain the tree species category of each pixel.
[0024] It should be noted that the hyperspectral tree recognition technology based on dual-path convolutional neural networks performs particularly well in complex environments. Hyperspectral images provide rich spectral information that can accurately distinguish different tree species. Even similar tree species can be accurately distinguished through differences in their spectral characteristics. When dealing with complex scenes, existing technologies can often only extract spatial features or spectral features separately, resulting in low precision or inaccurate recognition. Through the dual-path design of DPN, spatial features and spectral features are effectively integrated, avoiding the information loss caused by a single path and improving the model's expressive power and accuracy. Traditional methods often perform poorly in fine-grained tree species classification. The subtle differences provided by hyperspectral images can more accurately distinguish tree species, especially when the spectral differences between tree species are small, and effective identification can still be achieved.
[0025] In the embodiment of the present invention, the above step S100 includes the following sub-steps A1-A3; In A1: Radiation correction and atmospheric correction are used to remove errors from the hyperspectral image data, adaptive filters are used to perform denoising, and data enhancement is performed. The data of each band is standardized to obtain standardized hyperspectral image data. In A2: using the dimension reduction method to perform band selection on the standardized hyperspectral image data to obtain the selected hyperspectral image data; In A3: using a convolutional autoencoder to reduce the dimension of the selected hyperspectral image data to obtain first hyperspectral image data.
[0026] Specifically, using fir trees as an example, hyperspectral images of transmission corridors are acquired using drones, satellite remote sensing equipment, or ground-based sensors. Each image includes multiple bands, typically between 100 and 200, from visible to infrared light. The spatial resolution of each image is typically set between 5 and 30 meters, enabling coverage of a large area while ensuring individual tree identification.
[0027] The flight route is planned based on the target area to ensure data coverage and uniformity. Here, an area with a distribution of fir trees is selected as the target area. During flight, the remote sensing equipment collects hyperspectral images according to a predetermined plan. Each capture is performed simultaneously in multiple bands to ensure coverage of multi-dimensional spectral information. Each hyperspectral image contains spectral information from 150 bands, and the image size is 224×224×150, indicating a spatial resolution of 224×224 and 150 bands. During the image acquisition process, the sensor needs to be calibrated regularly to ensure data accuracy and consistency. This is accomplished using reflectors or ground calibration points.
[0028] After acquisition, hyperspectral image data often contains various noises, radiation errors, and other environmental interference factors, requiring preprocessing to ensure image data quality and adapt to subsequent deep learning models. The first hyperspectral image data is pre-processed hyperspectral image data; For the original image, radiation correction and atmospheric correction are used to remove errors caused by environmental and sensor differences.
[0029] Radiation correction includes, radiation correction algorithm; Through internal and external correction models, such as reflectors or zenith radiation data based on ground observations, the radiation values in the image are adjusted to restore the true spectral information, which is expressed as: , in, is the corrected radiation value, is the radiation reflectance of the original remote sensing image (uncorrected), is the background noise, is the correction factor of the sensor; Atmospheric correction models include FLAASH and ACORN models, which remove atmospheric effects based on atmospheric correction models; By constructing an atmospheric radiation transfer model, the atmospheric influence is analyzed and removed from the original data, which is expressed as: , in, is the atmospheric influence function, is the temperature, For humidity, The parameters are adjusted according to the atmospheric environment changes. is the radiation reflectance of the original remote sensing image (uncorrected), It is the true reflectivity value of the ground object after removing atmospheric interference, indicating the true reflectivity characteristics of the target after eliminating atmospheric influence factors (such as gas absorption, scattering, etc.); It should be noted that hyperspectral images will be affected by sensor errors, ground reflection differences and the atmosphere during the acquisition process, resulting in inaccurate radiation values of the images. The radiation correction algorithm aims to correct these errors and restore the true reflectivity of the image; water vapor, dust and other components in the atmosphere will interfere with the spectral signals of the hyperspectral images and affect the final tree recognition effect, so atmospheric correction is needed to remove the influence of the atmosphere.
[0030] In an embodiment of the present invention, an adaptive filter is used, including a median filter or wavelet denoising, to ensure the accuracy of the data; Wavelet denoising includes,wavelet transform can process signals simultaneously in time domain and frequency domain,,thus effectively removing high frequency noise; Median filtering involves performing local window processing on the image and replacing the value of each pixel in the image with the median of its neighboring pixels. This can effectively remove salt and pepper noise, which can be expressed as: , in, is the denoised image, is the original image, The median operation means that for each pixel value in the image, the middle value (i.e., median) of the neighboring pixels in a local window (such as 3×3 or 5×5) is taken to replace the current pixel value. Perform operations such as rotation, flipping, cropping, and scaling on training images to increase the robustness of the model; The data of each band is standardized and its range is unified to the interval [0,1], which is expressed as: , in, is the mean pixel value, is the standard deviation of pixel values, is the pixel value, which is the value of each pixel in the original image in a certain band (grayscale value or spectral reflectance, etc.), the original data that has not been standardized. Normalized pixel values represent pixel values after normalization. Their values are usually normalized to a certain range (such as [0,1] or a normal distribution with 0 as the mean and 1 as the standard deviation). It should be noted that hyperspectral images usually contain noise, such as sensor noise, atmospheric noise, and environmental interference. De-noising is a key step to ensure data quality and subsequent model accuracy.
[0031] In the embodiment of the present invention, band selection is required to extract the most representative bands from the hyperspectral data through dimensionality reduction techniques such as principal component analysis (PCA) and independent component analysis (ICA). The principal component analysis method includes,centering the data and subtracting the mean of each band so that the mean of each band is zero; Calculate the covariance matrix. For a standardized data set, the covariance matrix describes the correlation between the bands. Each element of the covariance matrix Indicates the Band and The covariance between the bands is expressed as: , in, For the The band in The value of the sample, For the The mean of the bands, For the Band and The covariance between the bands, is the sample size, For the The band in The value of the sample, For the The mean of the bands; Solve the eigenvalues and eigenvectors. The eigenvalues represent the variance of each principal component, and the eigenvectors represent the direction of each principal component. The larger the eigenvalue, the greater the data variance contained in the principal component and the greater the amount of information. The eigenvalues and eigenvectors of the covariance matrix are obtained by solving the following characteristic equation: , in, is the covariance matrix, is the eigenvector, is the characteristic value; Select the principal components for dimensionality reduction. The principal components to be selected are determined by the size of each eigenvalue. The principal components with larger eigenvalues correspond to larger data changes and represent the most important information in the data. The principal components with the largest eigenvalues can be selected for data dimensionality reduction. Usually, the principal components with cumulative variance reaching a certain threshold (such as 95%) are selected. For example, if the cumulative variance of the first two principal components accounts for more than 95% of the variance, then these two principal components can be selected to represent the data, and the remaining bands can be discarded; It should be noted that principal component analysis (PCA) is a very effective dimensionality reduction technique. Through the PCA method, high-dimensional spectral data can be reduced to a lower dimension, retaining the most important features in the data and reducing redundant information. The PCA method reduces the dimensionality of the original hyperspectral data. The reduced dimensionality not only retains most of the variation information in the data, but also reduces the number of bands and removes redundant information. The bands corresponding to these selected principal components are the most representative bands that need to be retained.
[0032] The independent component analysis (ICA) method involves standardizing the hyperspectral data to ensure that the mean of each band is 0 and the standard deviation is 1; using the ICA algorithm to decompose the standardized hyperspectral data into independent components, decomposing the input data into multiple statistically independent components to reveal potential independent features in the data; selecting the independent components most useful for tree identification and removing irrelevant components; and reconstructing the spectral data using the selected independent components, retaining the most informative components and removing redundancy and noise. It should be noted that ICA differs from PCA in that it emphasizes finding independent components in the data, rather than simply retaining the component with the largest variance. ICA can further optimize band selection, especially when processing tree reflectance spectra. By decomposing hyperspectral data into independent components, deeper nonlinear relationships between different bands can be discovered. ICA helps the model discover more independent and informative components in the data, further improving tree identification accuracy, especially when dealing with tree species classification and health assessment in complex environments.
[0033] In the embodiment of the present invention, the convolutional autoencoder is an unsupervised learning algorithm based on deep learning, which consists of two parts: an encoder and a decoder; The dimensionality reduction process includes: the data input to the convolutional autoencoder is a pre-processed hyperspectral image, each image consists of multiple bands, and each band corresponds to a feature map of the hyperspectral image; The encoder extracts the spatial features of the hyperspectral data through a series of convolutional layers and pooling layers, while performing dimensionality reduction. Each convolutional layer convolves the input data with a filter, gradually extracting more abstract features, including, using a standard convolution operation, expressed as: , in, is the output of the convolutional layer, is the convolution kernel, For input data, is the bias term, is the activation function, is the convolution operation; The pooling layer compresses the data by selecting the maximum value or average value, thereby achieving a gradual reduction in dimension. In the dimensionality reduction process, the feature map is compressed using maximum pooling or average pooling as follows: , in, is the area within the pooling window, is the result after pooling, To obtain the maximum value, For the to OK, for List List; Through multiple layers of convolution and pooling, the spatial information of the image is gradually compressed into a smaller latent space representation, forming a dimensionality-reduced feature. The final output of the encoder is a low-dimensional latent space representation, that is, the output of the encoding layer, which is usually a small feature quantity, for example, ,in It is the dimension of the latent space, which contains the main features of the input image. At this stage, the redundant information of the original data is removed and only the most discriminative features are retained; The decoder part maps the low-dimensional features of the latent space back to the high-dimensional space through a series of deconvolution layers to reconstruct the original data. The deconvolution layer includes the transposed convolution represented as: , in, is the transposed convolution kernel, is the latent space feature output by the encoder, The image reconstructed by the decoder, is the bias term, is the activation function, is the convolution operation; The decoder gradually restores the latent space representation to the spatial dimensions of the original image through deconvolution layers; The trained convolutional autoencoder can map the original hyperspectral image into a low-dimensional latent space through the encoder part; During the training process, the network calculates the difference between the reconstructed image and the original image, and uses the mean square error as the loss function to measure it: , in, is the reconstruction error, is the original input image, is the reconstructed image output by the decoder, is the total number of pixels or samples, To take the average value; The network is trained through the back-propagation algorithm to minimize the reconstruction error, so that the encoder can learn the optimal dimensionality reduction representation; The trained convolutional autoencoder can map the original hyperspectral image into a low-dimensional latent space through the encoder part, obtaining a concise feature representation. This low-dimensional feature representation retains the main information of the original data, removes redundancy and noise, and provides optimized input features for the subsequent tree recognition task. Assume that the latent space output by the encoder is represented as , whose dimensions can be , which means that the dimension of the hyperspectral image has been compressed from the original B bands to N key features; After preprocessing, hyperspectral images are typically stored in standard hyperspectral data formats, such as ENVI (.hdr and .dat files), and labeled according to training requirements. During the labeling process, target regions (such as trees and non-tree areas) are annotated in each image, providing realistic sample data for subsequent model training.
[0034] It should be noted that after selecting the most representative bands, the convolutional autoencoder dimensionality reduction algorithm is used to reduce the dimension of the hyperspectral data. Hyperspectral images usually contain hundreds of bands, which leads to excessively high data dimensions and huge computational complexity. In addition, the redundant information of some bands has limited contribution to the tree recognition task. Through dimensionality reduction processing, the main features of the image can be retained, redundant information can be removed, and the efficiency and accuracy of subsequent model training can be improved.
[0035] like Figure 3 As shown, in the embodiment of the present invention, the above step S200 includes the following sub-steps B1-B4; In B1: the first hyperspectral image data is divided into image blocks using a superpixel segmentation method, and band features are extracted using principal component analysis; In B2: Each image block is processed through a convolutional layer to extract the spatial features of the local area and add an activation function; In B3: Pooling is performed on the spatial features of the local area to reduce the size of the spatial features of the local area; In B4: the spatial features of all local regions are fused to obtain the spatial features.
[0036] Specifically, the first hyperspectral image data is preprocessed hyperspectral image data. The spatial path uses traditional 2D convolution operations to extract spatial information from the image. The preprocessed hyperspectral image data is input into the model. The image is divided into multiple smaller regions or superpixels through the superpixel segmentation method. The most representative band features are extracted from the hyperspectral data through the principal component analysis method. Each image block is processed through a convolution layer to extract the spatial features of the local area. The convolution layer uses multiple convolution kernels to capture spatial information such as edges, textures, and morphology in the image. Each convolution operation filters the pixels of the image and outputs a new feature map that describes the local spatial structure in the image block. The tree crowns, branches, leaves, and other morphological features will be extracted through the convolution operation. The convolution formula is expressed as: , in, is the input image block, is the convolution kernel, is the convolution operation, is the output spatial feature map; After the convolutional layer, the ReLU activation function is added to enhance the nonlinear expression ability of spatial features. The pooling layer further reduces the size of the feature map through maximum pooling or average pooling operations while retaining the most important spatial features. The spatial features extracted by multiple convolutional layers and pooling layers will be fused together to form the final spatial feature representation; It should be noted that each superpixel area contains a group of similar pixels, which helps to reduce the complexity of processing and focus on extracting spatial features in more representative areas, extracting the most representative band features to reduce data redundancy and improve the efficiency of tree recognition; plus the activation function increases the expressive power of the model, enabling it to better learn complex tree morphology. Through pooling, the model not only reduces the amount of calculation, but also reduces unnecessary spatial redundancy, thereby improving the computational efficiency of the model. By fusing features at multiple levels, the model can accurately extract the spatial distribution and morphology of trees.
[0037] In an embodiment of the present invention, after completing steps B1-B4 in the above step S200, the following steps B5-B6 are further included; In B5: the reflectance spectrum of each pixel of the first hyperspectral image data is used as the input of the spectral path, the spectral reflectance data of each pixel point is converted using discrete cosine transform, and the band characteristics are extracted using independent component analysis; In B6: The spectral data of each pixel is filtered through multiple convolution kernels to extract the characteristics of spectral reflectance, and an activation function is added to obtain the spectral features.
[0038] Specifically, the spectral data of each pixel is extracted through a 1D convolution (3×1 convolution kernel) operation, and the reflectance spectrum (spectral reflectance values of multiple bands) of each pixel of the hyperspectral image is used as the input of the spectral path; The discrete cosine transform maps these data from the original band space to the frequency domain to obtain components of different frequencies; the spectral signal is represented as different frequency components, where the low-frequency components contain the main change information in the image, while the high-frequency components represent details or noise. The discrete cosine transform is expressed as: , in, Input spectrum data elements, is the first frequency components, is the length of the input signal, The index of the input signal, ranging from 0 to , is the frequency index, ranging from 0 to , is the cosine function, is the normalized frequency factor, is the offset of the input signal; Optimize band selection through independent component analysis method to decompose hyperspectral data into the most useful independent components; Using 1D convolution operation, the spectral data of each pixel is processed by the convolution kernel. Through 1D convolution operation, the spectral data of each pixel will be filtered by multiple convolution kernels to extract the characteristics of spectral reflectance. The spectral path convolution is expressed as: , in, For the convolution kernels, is the input spectral data, is the number of convolution kernels, is the convolution output feature of the spectral path, The index of the convolution kernel, ranging from 1 to ; Activation functions such as ReLU are used to enhance the nonlinear expression ability of spectral features. Through the action of convolutional layers and activation functions, spectral features are effectively enhanced, helping the network to distinguish different spectral characteristics of trees.
[0039] It should be noted that discrete cosine transform is used to process hyperspectral data in order to extract more representative features from it. The spectral feature map after discrete cosine transform conversion will contain the frequency components of each pixel point, which can be used to further enhance the spectral feature performance in the tree identification task; when processing the reflectance spectrum of trees, by decomposing the hyperspectral data into the most useful independent components, it is possible to discover deeper nonlinear relationships between different bands, remove irrelevant components, and reduce redundant information; 1D convolution is specifically used to extract spectral information and capture the differences in reflectance between different bands.
[0040] In an embodiment of the present invention, after completing steps B5-B6 in the above step S200, the following step B7 is further included; In B7: spatial features and spectral features are added through residual connections and spliced using dense connections to obtain a fused feature map.
[0041] Specifically, the residual connection includes that the spatial features and spectral features are directly added in one or more layers, and the fused feature map is represented as: , in, is the output feature map of the spatial path, is the output feature map of the spectral path, It is the feature map after residual connection fusion; Dense connection involves concatenating the feature maps extracted from the spatial path and the spectral path into a longer feature map represented as: , in, is the feature map after dense connection fusion, is the output feature map of the spatial path, is the output characteristic graph of the spectral path; After residual connection and dense connection, the spatial features and spectral features are fused into a new feature map, which will be passed to the subsequent network layers for further processing.
[0042] It should be noted that residual connections, through skip connections, allow input features to skip certain network layers and be passed to subsequent layers. This helps alleviate the vanishing gradient problem in deep networks, improves training efficiency, and ensures that important information is not lost. Through residual connections, the model can retain more low-level features, reduce information loss, and improve training stability. This enables the model to effectively integrate spatial and spectral features while preventing the vanishing gradient problem in deep networks. Dense connections enhance the fluidity between features by connecting the output of each previous layer with the input of the subsequent layer. This connection method helps to efficiently transmit information, allowing the network to learn richer and more diverse features; this splicing operation not only maintains the independence of each feature, but also allows spatial and spectral features to be shared and combined in subsequent network layers. In this way, the model can learn more feature relationships and further improve recognition accuracy. In the embodiment of the present invention, the above step S400 includes the following sub-steps C1-C3; In C1: the fused feature map is input into the fully connected layer for classification; In C2: the original output of the first dual-path convolutional neural network model is converted into a probability distribution through an activation function; In C3: Determine the first category of the tree species for each pixel based on the maximum probability output by the activation function.
[0043] Specifically, the first dual-path convolutional neural network model is a trained dual-path convolutional neural network model. In this trained network, feature maps from the spatial and spectral paths are fused and fed into a fully connected layer for final classification. Activated by the Softmax function, the model outputs the probability of each pixel belonging to a different tree species. Based on the maximum probability output by the Softmax function, the first category of each pixel, i.e., the inferred category, is obtained. Taking fir trees as an example, if the probability of fir trees is the highest, the pixel is classified as a fir tree. If it is another tree species or background, it is classified into the corresponding category. Post-processing steps to accurately extract tree locations, classify tree species, and assess health status; It should be noted that changes in the spectral characteristics of tree species can reflect the health status of trees, and promptly detect diseased trees or trees affected by the external environment. It has a strong early warning function and can not only identify the health status of trees, but also quantitatively assess the health status through spectral characteristics, further improving the accuracy of monitoring and management.
[0044] In the embodiment of the present invention, the above step S500 includes the following sub-steps D1-D2; In D1: Using the semantic segmentation method of deep learning, the encoder part of the network extracts image features layer by layer, and the decoder maps the features to an image with higher spatial resolution to obtain the category label of each pixel; In D2: Based on the probability of each pixel belonging to a specific tree species, the category probability map is combined with the first category and a threshold is set to obtain the tree species category of each pixel.
[0045] Specifically, semantic segmentation methods may also include threshold method, region growing method, etc. Using the class probability map output by model inference, we can use thresholding to select the class with the highest probability for each pixel and map it to the image. In a certain part of the image, the model classifies a group of consecutive pixels as "fir tree," and thresholding creates a continuous canopy region. Using image segmentation algorithms, we can further extract all pixels in this region, obtaining the precise spatial location of each fir tree. The tree species classification uses the probability value of the inference output to make the final tree species confirmation. Based on the inference output probability, the threshold is set to 0.5 to determine the category of each pixel; For example, if the probability of a pixel being classified as a fir tree is 0.75, which is greater than 0.5, the pixel is classified as a fir tree; Ultimately, each pixel is classified as a specific tree species, forming a region for each tree species in the image. In the resulting image, pixels representing fir trees will be labeled with a specific color, while pixels representing other trees will be labeled with other colors.
[0046] It should be noted that by combining a deep learning model with image segmentation technology, the spatial location of each tree can be accurately extracted. This segmentation method not only identifies the entire area of the tree, but also carefully distinguishes the boundaries between trees, avoiding the ambiguity found in traditional methods.
[0047] Specifically, taking fir trees as an example, this method obtains the recognition results of each tree and outputs the tree species category of each pixel. For each pixel, the tree species category to which it belongs is output. The tree recognition results are shown in Table 1. Table 1 Tree recognition results , It should be noted that through a hyperspectral tree recognition method based on a dual-path convolutional neural network, the model successfully identified and classified trees in the image, was able to accurately classify each pixel into tree species, and output the tree species category and its corresponding prediction probability.
[0048] Experimental results demonstrate the high accuracy and robustness of this technology in practical applications, particularly in tree identification and location extraction. This process is of great significance for transmission corridor management, helping to promptly identify and address tree problems that may pose a threat to power transmission security.
[0049] The above is a schematic diagram of a hyperspectral tree identification method according to this embodiment. It should be noted that the technical solution of this hyperspectral tree identification system and the technical solution of the aforementioned hyperspectral tree identification method are based on the same concept. For details not described in detail in the technical solution of the hyperspectral tree identification system according to this embodiment, please refer to the description of the technical solution of the aforementioned hyperspectral tree identification method.
[0050] In this embodiment, a hyperspectral tree recognition system includes: a preprocessing module, configured to obtain hyperspectral image data, and preprocess the hyperspectral image data to obtain first hyperspectral image data; a feature fusion module, configured to perform feature extraction on the first hyperspectral image data, wherein the feature extraction includes spatial feature extraction and spectral feature extraction, and fuse the spatial features and the spectral features to obtain a fused feature map; A model training module, configured to train a dual-path convolutional neural network model based on the fused feature map to obtain a first dual-path convolutional neural network model; a tree identification module, configured to acquire real-time hyperspectral image data, perform tree identification on the real-time hyperspectral image data using the first dual-path convolutional neural network model, obtain a probability that each pixel belongs to a specific tree species, and obtain a first category of the tree species for each pixel based on the probability; The category determination module is used to extract the spatial position of trees through an image segmentation method, and combine the first category to obtain the tree species category of each pixel.
[0051] This embodiment further provides a computer device applicable to a hyperspectral tree identification method, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement a hyperspectral tree identification method as proposed in the above embodiment.
[0052] This embodiment further provides a storage medium storing a computer program. When the program is executed by a processor, the hyperspectral tree recognition method proposed in the above embodiment is implemented.
[0053] The storage medium proposed in this embodiment and the method for implementing a hyperspectral tree recognition proposed in the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0054] From the above description of the embodiments, those skilled in the art will clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, can also be implemented using hardware. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This software product can be stored on a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk, or optical disk, and includes instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0055] Example 2, referring to Tables 2 to 3, this example is different from the first example in that it provides a verification test of a hyperspectral tree recognition method to verify and illustrate the technical effects of the method.
[0056] The results identified by this method are compared with those of traditional methods: This method uses a dual-path convolutional neural network to extract spatial and spectral features, and performs residual connections and dense connections; Traditional method 1, using traditional convolutional neural network (CNN) to rely only on spatial features for tree recognition; Traditional method 2 uses support vector machine (SVM) combined with principal component analysis (PCA) to perform classification using only low-dimensional principal components in hyperspectral data; The comparison results of tree recognition accuracy under different methods are shown in Table 2; Table 2 Comparison of tree recognition accuracy using different methods , As shown in Table 1, our dual-path convolutional neural network model surpasses the traditional CNN and SVM+PCA methods in recognition accuracy, improving accuracy by 7.8% and 12.2%. This demonstrates that the dual-path convolutional neural network model is able to better integrate spatial and spectral features, providing more comprehensive tree recognition information. The F1-Score, which reflects the balance between the model's precision and recall, is higher than that of traditional methods, demonstrating its superiority in both precision and recall. The training time of the dual-path convolutional neural network model is slightly longer than that of the traditional CNN, but it is still relatively efficient compared to other complex network models; the inference time of the dual-path convolutional neural network model is slightly lower than that of the traditional CNN, indicating that it is more efficient in the inference stage and suitable for real-time applications.
[0057] Table 3 Robustness comparison results of different methods in different environments , As shown in Table 3, the dual-path convolutional neural network model demonstrates significantly better robustness across various test environments than traditional methods. It maintains high recognition accuracy in real-world conditions such as noise interference, low light conditions, and occlusion. Compared to traditional CNN and SVM+PCA methods, the dual-path convolutional neural network model demonstrates enhanced robustness and adaptability to diverse environmental challenges.
[0058] The comparative experiments above demonstrate that the dual-path convolutional neural network model outperforms traditional single convolutional neural network (CNN) and support vector machine (SVM) + principal component analysis (PCA) methods in terms of tree recognition accuracy, computational efficiency, and robustness. This method not only accurately extracts both spatial and spectral features from hyperspectral imagery but also effectively fuses these two types of features, demonstrating higher accuracy and robustness in tree recognition tasks.
[0059] Due to the efficiency and stability of the dual-path convolutional neural network model, it is very suitable for application in large-scale, dynamically changing environments, especially in practical scenarios such as tree monitoring on transmission lines. Combined with the continuous development of deep learning technology and the improvement of hardware acceleration, the dual-path convolutional neural network model is expected to be widely used and further optimized in the future.
[0060] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A hyperspectral tree recognition method, characterized in that: include: Acquiring hyperspectral image data, and preprocessing the hyperspectral image data to obtain first hyperspectral image data; Performing feature extraction on the first hyperspectral image data, wherein the feature extraction includes spatial feature extraction and spectral feature extraction, and fusing the spatial features and the spectral features to obtain a fused feature map; Based on the fused feature map, the dual-path convolutional neural network model is trained to obtain a first dual-path convolutional neural network model; Acquire real-time hyperspectral image data, perform tree recognition on the real-time hyperspectral image data using the first dual-path convolutional neural network model to obtain a probability that each pixel belongs to a specific tree species, and obtain a first category of the tree species for each pixel based on the probability; The spatial positions of trees are extracted by image segmentation method, and combined with the first category, the tree species category of each pixel is obtained.
2. A hyperspectral tree recognition method according to claim 1, characterized in that: Preprocessing the hyperspectral image data includes: The hyperspectral image data is corrected by radiation and atmosphere correction to remove errors, and then denoised by adaptive filters and data enhancement is performed. The data of each band is standardized to obtain standardized hyperspectral image data. The dimensionality reduction method is used to select the bands of the standardized hyperspectral image data to obtain the selected hyperspectral image data; The convolutional autoencoder is used to reduce the dimension of the selected hyperspectral image data to obtain the first hyperspectral image data.
3. A hyperspectral tree recognition method according to claim 2, characterized in that: Extracting features from the first hyperspectral image data includes: The first hyperspectral image data is divided into image blocks using a superpixel segmentation method, and band features are extracted using principal component analysis; Each image block is processed through a convolutional layer to extract the spatial features of the local area and add an activation function; The spatial features of the local area are pooled to reduce the size of the spatial features of the local area; The spatial features of all local areas are fused to obtain the spatial features.
4. A hyperspectral tree recognition method according to claim 3, characterized in that: Also includes: The reflectance spectrum of each pixel of the first hyperspectral image data is used as the input of the spectral path, the spectral reflectance data of each pixel point is converted using discrete cosine transform, and the band characteristics are extracted using independent component analysis; The spectral data of each pixel is filtered through multiple convolution kernels to extract the characteristics of spectral reflectance, and an activation function is added to obtain the spectral features.
5. The hyperspectral tree recognition method according to claim 4, wherein: The fusion of spatial features and spectral features includes: The spatial features and spectral features are added through residual connections and spliced using dense connections to obtain the fused feature map.
6. The hyperspectral tree recognition method according to claim 5, wherein: Performing tree recognition on the real-time hyperspectral image data using the first dual-path convolutional neural network model includes: Input the fused feature map into the fully connected layer for classification; The original output of the first dual-path convolutional neural network model is converted into a probability distribution through an activation function; The first category of the tree species for each pixel is determined based on the maximum probability output by the activation function.
7. The hyperspectral tree recognition method according to claim 6, wherein: The tree species categories for each pixel include: Using the semantic segmentation method of deep learning, the encoder part of the network extracts the features of the image layer by layer, and the decoder maps the features to an image with higher spatial resolution to obtain the category label of each pixel; According to the probability of each pixel belonging to a specific tree species, the category probability map is combined with the first category and a threshold is set to obtain the tree species category of each pixel.
8. A hyperspectral tree recognition system, using a hyperspectral tree recognition method according to any one of claims 1 to 7, characterized in that: include: a preprocessing module, configured to obtain hyperspectral image data, and preprocess the hyperspectral image data to obtain first hyperspectral image data; a feature fusion module, configured to perform feature extraction on the first hyperspectral image data, wherein the feature extraction includes spatial feature extraction and spectral feature extraction, and fuse the spatial features and the spectral features to obtain a fused feature map; A model training module, configured to train a dual-path convolutional neural network model based on the fused feature map to obtain a first dual-path convolutional neural network model; a tree identification module, configured to acquire real-time hyperspectral image data, perform tree identification on the real-time hyperspectral image data using the first dual-path convolutional neural network model, obtain a probability that each pixel belongs to a specific tree species, and obtain a first category of the tree species for each pixel based on the probability; The category determination module is used to extract the spatial position of trees through an image segmentation method, and combine the first category to obtain the tree species category of each pixel.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the hyperspectral tree recognition method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the hyperspectral tree recognition method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Hyperspectral image feature extraction method, classification model construction method and classification method
CN110110596A
Hyperspectral image classification method based on double-path convolution and double attention and storage medium
CN115272776A
Tree species identification, classification and counting method based on PLS feature fusion convolutional neural network
CN116958630A
Hyperspectral tree species high-precision feature extraction method and system based on improved BP neural network
CN118657952A
Tree species classification method and system based on spectral depth extraction convolutional neural network SDA-CNN
CN119273994A