Method for detecting and identifying weak and small target in hyperspectral image
Through the adaptive global dense nested inference network combined with biochemical parameters and spatial knowledge, the problem of feature loss and poor detection effect in weak object detection of hyperspectral images is solved, and efficient object recognition and detection is achieved.
Patent Information
- Application Number
- CN202510233391.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-08
AI Technical Summary
The existing weak object detection technology for hyperspectral images has insufficient in small object detection, deep feature maintenance, mining of relationships between targets and surrounding elements, and utilization of prior knowledge, resulting in poor detection results.
Using an adaptive global dense nested inference network method, human prior knowledge and spectral features are used to combine biochemical parameters and spatial knowledge, a global semantic pool and knowledge graph are constructed to enhance feature extraction and target recognition.
It effectively protects the high-level characteristics of weak targets, improves the correlation between spectral information and targets, improves the accuracy and robustness of detection, reduces network noise, and achieves better detection and recognition of small targets.
Smart Images

Figure CN120279243A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of detection and recognition of small and weak targets in hyperspectral images, and particularly relates to a method for detecting and recognizing small and weak targets in hyperspectral images. Background Technique
[0002] With the rapid development of remote sensing technology, hyperspectral imaging technology has become an important means to obtain spectral information of ground objects. Hyperspectral images (HSIs) provide rich spectral information for the identification and analysis of surface materials by capturing hundreds of continuous spectral bands. This enables HSIs to show broad application prospects in many fields such as reconnaissance, environmental monitoring, agricultural management, and resource exploration. Different from optical images, due to the need of hyperspectral remote sensing sensors to balance the energy in the spatial and spectral dimensions, although they have relatively rich spectral information, their spatial resolution is generally low. This makes small targets such as cars and airplanes occupy only a few pixels in space. This poses a challenge to the accurate detection task of hyperspectral small targets.
[0003] In recent years, convolutional neural networks have made many breakthroughs in image processing. Many scholars have tried to apply convolutional neural networks to the detection of small targets in hyperspectral images. For example, methods such as CNNTD, HTD-IRN-Net, and TSCNTD have been proposed. However, these CNN-based methods are often limited by the insufficient samples of hyperspectral small targets. To solve the problem of insufficient samples, self-supervised learning techniques have been introduced. For example, self-supervised learning methods based on spectral mixture features, HTD-IRN methods based on self-supervised spectral-level contrast learning, etc. However, these methods usually need to be trained and tested on the same dataset, and the problems of applicability and computational consumption are significant. However, the existing hyperspectral image target detection technologies still have deficiencies in small target detection, deep feature preservation, mining the relationship between targets and surrounding elements, and utilization of prior knowledge, and further improvement and optimization are urgently needed. Summary of the Invention
[0004] The present invention aims to solve the problem of poor detection effect of small and weak targets in hyperspectral images. By using human prior knowledge and the spatial and spectral characteristics of spectra, a method for detecting and recognizing small and weak targets in hyperspectral images based on an adaptive global dense nested inference network (AGDNR) is proposed.
[0005] To achieve the above object, the present invention provides the following technical solution: A method for detecting and recognizing small and weak targets in hyperspectral images, the method comprising the following steps:
[0006] Step 1: Obtain the hyperspectral image HSI to be detected. Calculate five surface parameters based on the infrared and near-infrared bands of the hyperspectral image. The Normalized Difference Vegetation Index (NDVI) is used to classify the surface. The remaining four surface parameters, namely Leaf Area Index (LAI), Ratio Vegetation Index (RVI), Difference Vegetation Index (DVI), and Temperature-Vegetation Drought Index (TDVI), are spliced into the original image as additional spectra. Then, use a feature extractor to extract the semantic features of different surface types from the spliced hyperspectral image.
[0007] Step 2: Input the hyperspectral image into the densely nested model for feature extraction to obtain the original features of the image.
[0008] Step 3: Extract the semantic information of the original features based on the weights of the classifier as the semantic features of the target class, and jointly construct a global semantic pool with the semantic features of the surface types.
[0009] Step 4: Hard-link the surface categories and pixels. For each pixel, it can be corresponded to a unique surface type through its NDVI value to obtain the probability matrix P2. Soft-link the target categories and pixels, that is, obtain the classification probability matrix P1 on all C target categories. Finally, obtain a total probability matrix P = [P1, P2].
[0010] Step 5: Generation of enhanced features: Combine the adjacency matrix constructed by the knowledge graph, the global semantic pool, the total probability matrix, and the weight matrix of the convolutional neural network to obtain an enhanced feature vector. Use the attention mechanism to reduce network noise to obtain the final enhanced features.
[0011] Step 6: Splice the enhanced features obtained in Step 5 with the original features of the image obtained in Step 2, and then input the spliced features into the convolutional neural network to obtain the final detection result.
[0012] Preferably, the densely nested model is a stacked U-shaped fully convolutional neural network, and its specific structure is I downsampling layers. The first downsampling layer has J horizontal dense blocks, and then each layer decreases by one horizontal dense block. Each horizontal dense block is specifically a convolutional attention module.
[0013] Specifically, the hyperspectral image is input into the first horizontal dense block of the first layer. The input of the first horizontal dense block of each layer comes from the output of the first horizontal dense block of the previous layer after pooling. The input of the j-th horizontal dense block of the first layer comes from the outputs of all previous horizontal dense blocks in the same layer and the upsampling result of the output of the j-1-th node in the second layer. The input of the j-th horizontal dense block of the remaining layers comes from the outputs of all previous horizontal dense blocks in the same layer, the pooling result of the j-th horizontal dense block in the previous layer, and the upsampling result of the j-1-th horizontal dense block in the next layer.
[0014] The structure of the convolutional attention module specifically includes a convolutional layer, a channel attention module, and a spatial attention module. Let a horizontal dense block be L i,j The output after the input of L passes through the convolutional layer is L, and the calculation formula for the output of this horizontal dense block is:
[0015]
[0016] where represents element-wise multiplication, M c is the channel attention module, and M s is the spatial attention module.
[0017] Preferably, step 1 is specifically as follows:
[0018] First, calculate five surface parameters, and then splice the surface parameters to the original image. The process is expressed as:
[0019] input' = [input, LAI, RVI, DVI, TDVI]
[0020] where input is hyperspectral image data, and LAI, RVI, DVI, and TDVI are surface parameters calculated from the input image respectively. [,] represents the splicing operation; the input' obtained after splicing is input into the feature extraction network for feature extraction;
[0021] Use the NDVI parameter value to divide the surface type. The feature extraction process is as follows:
[0022] First, segment the image using the divided surface type:
[0023]
[0024] where is the hyperspectral value after image segmentation, input′ i,j is the hyperspectral value at the (i, j) position of the spliced image, and condition k is the range of the NDVI value corresponding to the k-th surface type in the divided surface type;
[0025] After that, input the segmented image into the convolutional neural network, which is expressed as:
[0026] M k = MLP(Conv(F k ))
[0027] where Conv represents a multi-layer deep convolutional neural network for feature extraction, and MLP represents a fully connected layer with one hidden layer to reduce the features to one dimension to obtain one-dimensional features Denote the real number field as \(\mathbb{R}\), and \(L\) is the dimension of the semantic features of different surface types.
[0028] Preferably, step 3 is specifically as follows:
[0029] First, for the original features, two convolutional layers with a convolutional kernel size of \(1\times1\) are used for pixel-level classification, and the weights of the convolutional kernels corresponding to different category targets in the last convolutional layer are extracted and expressed as \(C\) semantic information of the target categories, constructing a global semantic pool \(M_1\in\mathbb{R}\) C×L , where \(C\) is the number of categories of the targets to be detected, and \(L\) is the length of the weights of the convolutional kernels;
[0030] Then, the results after NDVI surface classification and feature extraction in step 1 are used as the semantic features \(M_2\in\mathbb{R}\) of various surface types 5×L , where \(L\) is the length of the semantic features of the surface types;
[0031] Finally, the global semantic pool \(M = [M_1, M_2]\) of the targets and surface types is obtained.
[0032] Preferably, step 4 is specifically as follows:
[0033] First, the process of establishing the enhanced feature vector is represented by matrix multiplication as follows:
[0034] \(F' = PEMW\) G
[0035] where \(E\) is the adjacency matrix constructed by the knowledge graph, \(M\) is the global semantic pool of the target categories and surface types, \(P\) represents the weight matrix of the convolutional neural network, \(D\) represents the number of channels of the enhanced features, then the dimension of the enhanced feature \(F'\) is
[0036] Then, the attention mechanism is used to capture the target categories and surface categories that the input image mainly focuses on, obtaining the attention weights \(\alpha=\text{softmax}(z\) α \(P\) α \(M\) T ), where \(z\) is the feature vector obtained after convolution of the original image, \(P\) is the fully connected weight, and the attention weights of each category are obtained as
[0037] The final enhanced feature is obtained:
[0038]
[0039] where, represents the element-wise product.
[0040] Preferably, the calculation formula of the channel attention module is as follows:
[0041] M c (L) = σ(MLP(avgPool(L))) + σ(MLP(maxPool(L)))
[0042] Wherein, MLP represents a fully connected layer with one hidden layer, avgPool and maxPool represent global average pooling and max pooling for the channels of each dimension L respectively, and σ represents an activation function;
[0043] The calculation formula of the spatial attention module is as follows:
[0044] M s (L) = σ(Conv[CavgPool(L), CmaxPool(L)])
[0045] Wherein, Conv represents a convolutional layer, and CavgPool and CmaxPool represent channel average pooling and channel max pooling respectively.
[0046] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0047] 1. A high-dimensional multi-layer nested U-net network is proposed: The traditional U-net network will cause the loss of deep features of small targets. The present invention adopts a dense nesting method to transfer features to deeper or shallower nodes. At the same scale, target features can also be transferred at the same layer. Each layer of the network contains feature information at other scales, protecting the high-level features of weak targets from being lost to the greatest extent.
[0048] 2. Biochemical parameters are used to enhance the spectral characteristics of targets: Due to the differences in hyperspectral sensor parameters, the spectral curves of the same ground object are different. In order to improve the correlation between spectral information and targets and the robustness of the algorithm, the present invention proposes to utilize a variety of biochemical parameters to mine the correlation between bands and combine spectral information with biochemical parameters to enhance the spectral characteristics of ground objects and targets.
[0049] 3. A target inference algorithm integrating spatial knowledge information is proposed: In order to associate the ground objects around the target, the environment in which the target is located with the target, the present invention constructs a geographical map of target and environment information. And combines the global semantic pool with the geographical map, propagates the target semantic representation according to the geographical map, and then conducts global inference; uses the attention mechanism to adaptively emphasize the category of the inference target. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a flowchart of hyperspectral small target detection for the present invention;
[0051] Figure 2 U-shaped fully convolutional neural network diagram of the present invention;
[0052] Figure 3 Flowchart of the densely nested network of the present invention;
[0053] Figure 4 Schematic diagram of the spatial and channel attention mechanism of the present invention;
[0054] Figure 5 ROC curves of the AGDNR model of the present invention and other models on three datasets;
[0055] Figure 6 Classification detection diagram of the AGDNR model of the present invention on the San Diego dataset;
[0056] Figure 7 Classification ROC curve diagrams of the AGDNR model of the present invention on three datasets;
[0057] Figure 8 Schematic diagrams of the structures of the AGDNR model, the AGDNR w / o SC model, and the AGDNR w / o SC&DS model of the present invention;
[0058] Figure 9 RGB diagram and ground truth diagram of the San Diego dataset of the present invention;
[0059] Figure 10 Detection diagram of the ARGNR model of the present invention on the San Diego dataset;
[0060] Figure 11 Detection diagram of the AGDNR w / o LSF model of the present invention on the San Diego dataset. Detailed implementation manners
[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0062] A method for detecting and recognizing small and weak targets in hyperspectral images, the method is as Figure 1 shown, and includes the following steps:
[0063] Step 1: Obtain the hyperspectral image (HSI) to be detected. Calculate five surface parameters based on the infrared and near-infrared bands of the hyperspectral image. The Normalized Difference Vegetation Index (NDVI) is used to classify the surface, and the remaining four surface parameters are spliced into the original image as additional spectra. Meanwhile, a feature extractor is used to extract the semantic features of different surface types from the spliced hyperspectral image.
[0064] The value range of the surface parameter NDVI is [-1, 1]. According to the calculated different NDVI values, the surface can be divided into five different types. The numerator of the calculation formula of NDVI is the reflection value of the infrared band minus the reflection value of the red light band. Since water, clouds, oceans, etc. have a high reflection of the visible light band, when NDVI < 0.2, it can be considered that the area is ocean, water or cloud. Vegetation has a high reflection value for infrared light but absorbs red light. Therefore, when NDVI > 0, it can be considered that there is vegetation coverage, and the vegetation coverage degree increases with the increase of the NDVI value. When NDVI < 0 and NDVI > -0.2, the surface can be regarded as an urban area. The specific classification method is shown in Table 1.
[0065] Table 1 Land type classification based on NDVI
[0066]
[0067]
[0068] First, calculate five surface parameters: NDVI, Leaf Area Index (LAI), Ratio Vegetation Index (RVI), Difference Vegetation Index (DVI), and Temperature-Vegetation Drought Index (TDVI). Among these five surface parameters, NDVI is used for the first step of surface division, while parameters such as LAI, RVI, DVI, and TDVI are used as enhanced features and spliced into the original hyperspectral image to increase the richness of information. The process of splicing the surface parameters into the original image can be expressed as:
[0069] input' = [input, LAI, RVI, DVI, TDVI] (1)
[0070] where input is the hyperspectral image (HSI) data, and LAI, RVI, DVI, TDVI are the surface parameters calculated from the input image respectively, and [,] represents the splicing operation. The spliced input' is input into the feature extraction network for feature extraction.
[0071] The process of using NDVI for image segmentation is as follows:
[0072]
[0073] Among them, is the hyperspectral value after image segmentation, input′ i,j is the hyperspectral value of the input image at the (i, j) position, condition k is the range of NDVI values corresponding to the k-th surface type among the above 5 surface types. Then, the image segmented according to NDVI is input into the convolutional neural network, which can be expressed as:
[0074] M k = MLP(Conv(F k ))(3)
[0075] Among them, Conv represents a multi-layer deep convolutional neural network for feature extraction. MLP represents a fully connected layer with one hidden layer that reduces the features to one dimension. At the same time, in the real number domain the one-dimensional feature L is the dimension of the semantic features of different surface types, and the one-dimensional feature after dimensionality reduction can be used for pixel-level knowledge reasoning.
[0076] Step 2: Input the hyperspectral image into the densely nested model for feature extraction to obtain the original features, enabling the model to maintain the features of small targets in the deep convolution and enabling the model to obtain more accurate recognition ability.
[0077] As Figure 2 shown, the U-shaped fully convolutional neural network gradually reduces the image size and increases the number of feature channels of the original image by using convolution and pooling operations to extract the high-level features of the image, and then increases the image size and reduces the number of feature channels through methods such as deconvolution or interpolation to perform upsampling. Finally, the low-level and high-level features are fused to extract the image features.
[0078] Specifically, the U-shaped fully convolutional neural network has M layers. The input image is input from the first convolutional layer of the first layer and output from the second convolutional layer of the first layer; there are two convolutional layers in each of the first layer to the M-1 layer, and one convolutional layer in the M layer; the input of the first convolutional layer of the M-i layer comes from the features after pooling of the output of the first convolutional layer of the previous layer; the input of the second convolutional layer of the M-i layer is obtained by splicing and combining the features after pooling of the output of the first convolutional layer of this layer and the features after upsampling of the output of the second convolutional layer of the next layer.
[0079] As Figure 3As shown, multiple U-shaped fully convolutional neural networks are stacked to obtain a dense nested convolutional neural network structure. In the detection and recognition of small targets in hyperspectral images, since the sizes of different targets are different, the receptive fields required are also different. Therefore, at different levels of the convolutional neural network, the features of different-sized targets may be included. So, in the network structure adopted in the present invention, multiple nodes are added at each level. These nodes can fuse the output features of the nodes in the same layer and the adjacent layers, thereby achieving multi-layer feature fusion. This structure is beneficial to maintaining the features of small targets in the deep layer, thus obtaining better small target detection performance. Taking the five-layer feature extraction module as an example, Figure 3 The specific structure of this module is shown.
[0080] As Figure 3 shown, this dense convolutional neural network has I = 5 downsampling layers. The first downsampling layer has J horizontal dense blocks, and then each layer decreases by one horizontal dense block. Then each node can be represented as L i,j (i = 0, 1, …, I; j = i = 0, 1, …, J), where i represents the i-th downsampling layer and j represents the j-th horizontal dense block. Then, let L i,j be the output of the node L i,j at the (i + 1)-th row and (j + 1)-th column. For the output of L i,j , the nodes can be discussed in four cases.
[0081] The first type of node, i.e., the first node: i = 0, j = 0
[0082] The output of the first node can be expressed in the following form:
[0083] L i,j = F(input) (4)
[0084] where input is the input image and F(.) represents the convolutional attention module of this node.
[0085] The second type of node: i ≠ 0, j = 0
[0086] L i,0 represents the output result of the downsampling layer along the encoding part of the U-shaped network, that is, the first horizontal dense block of each layer. The input of the L i,0 node only comes from the output of the L i-1,0 node. Therefore, the output of these nodes can be expressed in the following form:
[0087] L i,0 = maxPool(F(L i-1,0 )) (4)
[0088] where maxPool(.) represents the max pooling layer with a stride of 2.
[0089] The 3rd type of node: i = 0, j ≠ 0
[0090] L 0,j represents the dense block of the first layer. The inputs of these dense blocks are the outputs of all the previous modules in the same layer and the upsampling of the output result of the (j - 1)-th node in the second layer. Therefore, the outputs of these nodes can be expressed as:
[0091] L 0,j = F([L 0,0 , L 0,1 ,..., L 0,j-1 , L 0,0 , U(L 1,j-1 )) (5)
[0092] where F(.) represents the convolutional attention module of this layer, [.,.] represents the concatenation module, and U(.) represents the upsampling layer with a scaling factor of 2.
[0093] The 4th type of node: i ≠ 0, j ≠ 0
[0094] L i,j 's input comes from the outputs of all the previous modules in the same layer, the pooling result of the j-th module in the previous layer, and the upsampling result of the (j - 1)-th module in the next layer. Therefore, the outputs of these nodes can be expressed as:
[0095] L i,j = F([L i,0 , L i,1 ,..., L i,j-1 , D(L i-1,j ), U(L i+1,j-1 )) (6)
[0096] where F(.) represents the convolutional attention module of this layer, [.,.] represents the concatenation module, U(.) represents the upsampling with a scaling factor of 2, and D(.) represents the downsampling with a scaling factor of 2.
[0097] In each node of the densely nested convolutional neural network, the input needs to be subjected to multiple layers of convolution, and then sequentially input into the channel and spatial attention modules to enhance the feature extraction effect, as Figure 4 shown. The channel attention module and the spatial attention module can extract the channels and pixels that are more interesting for the hyperspectral image detection task through the attention calculation of the channel and the space, thereby improving the detection accuracy. Let the channel attention map and the spatial attention map be M c and M s . If the dimension of the input feature Let \(C\) denote the number of bands of the hyperspectral image, \(H\) denote the width of the spatial dimension of the hyperspectral image, and \(W\) denote the height of the spatial dimension of the hyperspectral image. Channel attention map \(M\) c The calculation formula is as follows:
[0098] \(M\) c (L)=\(\sigma\)(MLP(avgPool(L)))+\(\sigma\)(MLP(maxPool(L))) (7)
[0099] Among them, MLP represents a fully connected layer with one hidden layer, avgPool and maxPool represent global average pooling and max pooling for each channel of \(L\), that is, taking the average or maximum value of all pixel values in the same channel, and \(\sigma\) represents the ReLU activation function.
[0100] Spatial attention map \(M\) s The calculation formula is as follows:
[0101] \(M\) s (L)=\(\sigma\)(Conv 7×7 [CavgPool(L),CmaxPool(L)]) (8)
[0102] Among them, Conv 7×7 represents a convolutional layer with 7 kernels. This convolutional layer converts the input of two channels into an output on one channel. CavgPool and CmaxPool represent channel average pooling and channel max pooling respectively, that is, taking the average or maximum value of the pixels in each feature map on the channel.
[0103] As Figure 4 shown, assume that the output of the input of a node \(L\) i,j after passing through the convolutional layer is \(L'\), then the calculation formula for the output of this node is:
[0104]
[0105] Among them, represents element-wise multiplication. Before multiplying \(M\) s and \(M\) c with \(L\) and \(L'\), the sizes of \(M\) c and \(M\) s have already been replicated to
[0106] Step 3: Construction of the global semantic pool: First, semantic features of the target and surface types are constructed in two ways respectively. Inspired by works such as Reasoning-RCNN, which use the weights of classifiers to identify some unknown targets, the weights of the classifier can, to a certain extent, represent the semantic features of various targets. In the present invention, two convolutional layers with a kernel size of 1×1 are used for pixel-level classification of the original features after Step 2, and the weights of the convolutional kernels corresponding to different category targets in the last convolutional layer are extracted and represented as C semantic information of the target categories, constructing the global semantic pool M1 ∈ R C×L . Where C is the number of categories of the targets to be detected, and L is the weight length of the convolutional kernel. At the same time, for different surface types, the results after NDVI surface classification and feature extraction in Step 1 are used as the semantic features M2 ∈ R 5 ×L , and L is the semantic feature length of the surface categories. Finally, the global semantic pool M = [M1, M2] of the target and surface types can be obtained. The global semantic pool constructed by this method will be continuously updated during the network training process, and as the number of iterations increases, the semantic pool will become more accurate.
[0107] Step 4: Generation of enhanced features: After creating semantics for all target categories and surface types, these semantics can be propagated between related categories through the edges of the prior knowledge graph. However, a method for linking pixels and categories in HSI still needs to be found.
[0108] In this embodiment, a hard link is selected between the surface category and the pixel, and a soft link is selected between the target category and the pixel. Among them, the soft link between the target category and the pixel is the classification probability over all C categories This classification probability can flexibly represent the category to which each pixel point belongs and its category proportion. Where N is the number of pixels in the image, and C is the number of categories of the targets to be detected. The link between the pixel and the surface category is a hard link. For each pixel, through its NDVI value, it can be corresponded to a unique surface type, thus establishing a hard link between the surface type and the pixel. The hard link between the pixel and the surface type can be regarded as the probability that the pixel belongs to this surface type is 1, and the probability of belonging to other surface types is 0. Therefore, a probability matrix P2 ∈ R N×5 can be obtained, where N is the number of pixels in the image. Then, a total probability matrix can be obtained.
[0109]
[0110] Finally, the process of establishing the enhanced feature vector can be represented by matrix multiplication: F′ = PEMW GAmong them, E is the adjacency matrix constructed by the knowledge graph, and M is the global semantic pool of the target category and the surface type. represents the weight matrix of the convolutional neural network, D represents the number of channels of the enhanced feature, then the dimension of the enhanced feature F′ is That is, each pixel has an enhanced feature.
[0111] Step 5: Since all categories are used to generate the enhanced feature, there will be noise in the network. Therefore, an attention module needs to be introduced to extract the attention patterns of the network for different categories to reduce the noise introduced by targets that do not exist in different hyperspectral images.
[0112] The present invention uses an attention mechanism to capture the target categories and surface categories that each input image mainly focuses on. Global pooling is performed on the entire graph through a 3×3 convolutional kernel, and then a fully connected layer is connected. Finally, softmax is applied to the output result of the fully connected layer to obtain the attention weight α of each category = softmax(z α W α M T ), where is the feature vector obtained after convolution of the original image, is the fully connected weight, and finally the attention weight of each category is The generation of the entire enhanced feature is Among them, represents element-wise multiplication, and the other parts are matrix multiplications. E is the adjacency matrix constructed by the knowledge graph, and M is the global semantic pool of the target category and the surface type.
[0113] Step 6: Feature concatenation: Concatenate the enhanced feature F′ obtained in Step 5 with the original feature F of the image obtained in Step 2 to obtain an enhanced feature [F, F′]. Finally, input this concatenated feature into a new 1×1 convolutional neural network with a hidden layer to obtain the final detection result. Since the enhanced feature F′ is the pixel feature obtained through prior knowledge, it can be used as a supplement to the features of some targets interfered by the background, thereby increasing the detection accuracy.
[0114] To fully verify the effectiveness and superiority of the method we proposed, the present invention uses two hyperspectral datasets, the San Diego dataset and the Mosaic Avon dataset, and a synthetic hyperspectral dataset for training and testing of the target detection task.
[0115] An experiment was conducted by selecting a sub-region containing 256×256 pixels from the Mosaic Avon dataset. This region consists of a large area of grassland and small areas of land and roads. There are a total of 228 target pixels to be detected in this dataset.
[0116] For the synthetic simulated dataset background image, the hyperspectral dataset TG1HRSSC collected by Tiangong-1 and released by the Space Application Engineering and Technology Center of the Chinese Academy of Sciences in 2021 for scene classification was selected. This dataset contains images in three spectral ranges, namely the panchromatic band (PAN), visible near-infrared band (VNIR), and short-wave infrared band (SWIR). The dataset captured nine geographical categories such as towns, farmlands, ports, airports, etc. Three hyperspectral images of ports were selected from this dataset to construct the training set and the test set. The size of this hyperspectral image is 256×256 pixels and it contains 54 band information from 400 to 1000 nanometers. The port image consists of sea areas, vegetation, and urban surface types. The targets of this dataset were extracted from the ABU dataset. The Airport-Beach-Urban (ABU) dataset was manually cropped after being downloaded from the AVIRIS website. The size of each image varies and it covers about 200 spectral information. These datasets contain targets such as airplanes, ships, and cars that can be used for detection. In this invention, targets such as airplanes, ships, and cars were randomly selected from these targets, and after adjusting their sizes and spectra, they were cropped into the water areas and urban areas of the background dataset. Thus, three hyperspectral images were synthesized. Each of these three images contains more than a dozen targets to be detected, and there are no duplicates among these targets.
[0117] To verify the effectiveness of the algorithm proposed in this invention, the AGDNR model of this invention and five other comparison models were used to conduct target detection experiments on three datasets.
[0118] Figure 5 The ROC curves of the six models on the Sandigo dataset (a), Masic Avon dataset (b), and synthetic dataset (c) are shown. It can be seen from the curves that on the San Diego dataset, Avon dataset, and synthetic dataset, AGDNR has the best ROC curve, with a higher detection probability and a lower false alarm rate. This shows that the AGDNR model has stable detection accuracy and optimal detection effect.
[0119] Table 2 shows the AUC (D,F) and AUC (τ,F) values of the six models, and the optimal values among all models are bolded. Generally speaking, this method performs best in the vast majority of AUC values. Among them, on the Avon dataset and the synthetic dataset, all the AAUC of AGDNR (τ,F)UAUC (τ,D) The C values are all optimal, showing good model detection accuracy, effectiveness, and background suppression ability. On the San Diego dataset, the value of AGDNR is slightly higher, indicating that the model has poor background suppression ability on the San Diego dataset.
[0120] Table 2 Comparison of AUC values of different algorithms in three kinds of data
[0121]
[0122] The model constructed by the method of the present invention can not only detect the positions of small targets, but also has the ability to infer the categories of small targets. In order to analyze the classification and detection ability of AGDNR, the detection result graphs and ROC curves of three datasets are calculated and the three AUC values of different types of targets are given to verify the target category detection performance of the model constructed by the present invention.
[0123] Figure 6 Shows the classification detection graph of the AGDNR model on the San Diego dataset, where (a) is the original RGB image; (b) is the ground truth image, and different colors represent different target categories; (c) is the target detection image; (d)-(f) are the detection images of different categories.
[0124] As Figure 6 shown, for the targets of category three in the San Diego dataset, there is no missed detection, but there are a small number of misdetections distributed on the background. It is speculated that this may be because the targets of category three are too small and their spectra are greatly affected by the background. For misdetected pixels, the probability value of being classified as an incorrect target is low and within an acceptable range, which does not affect the final classification result.
[0125] Figure 7 Shows the classification ROC curves of the AGDNR model on the San Diego dataset, the Masic Avon dataset, and the synthetic dataset. Table 3 shows the A UC (τ,D) 、AUC (τ,D) and AUC (τ,F) values.
[0126] Table 3 AUC values of AGDNR in the detection results of different category targets in three datasets
[0127]
[0128] Figure 7As can be seen from the ROC curve, the model can accurately classify target pixels on all three datasets. As shown in Table 3, on the three datasets, the AUC (D,F) values for each category of the model are all close to 1, indicating that the model can accurately detect and classify targets. The AUC (τ,D) values of the model on the three datasets are all relatively high, indicating that the pixels detected by the model are highly effective. On Category 3 of the San Diego dataset, Category 4 of the Masic Avon dataset, and Category 3 of the simulation dataset, the effectiveness of the model detection is slightly weaker. The AUC (τ,F) values on the three datasets are all relatively close to the AUC (D,F) close to 0, indicating that the model has strong background suppression ability. Generally speaking, the model has obtained good results on the three datasets and can accurately and effectively detect and classify targets.
[0129] The deep nested convolutional neural network adopted by the present invention can realize the direct interaction between shallow features and deep features, can effectively input the features of small targets directly into the deep network, and can prevent the disappearance of the features of small targets due to deep convolution and pooling. The present invention removes the skip connections and downsampling connections in the deep connection convolutional neural network to form two variants of the network, with other network structures and output scales remaining unchanged, to study the role of skip connections and downsampling connections in maintaining the features of small targets.
[0130] AGDNR w / o SC: As Figure 8 shown, all skip connections in the model are removed to obtain a variant. Here it is called AGDNR w / o SC, and its structure is shown in Figure 8(a).
[0131] AGDNR w / o SC&DS: The sampling connection layer is used to maintain the features of small targets in the deep network. All downsampling layers except the first column are removed to form a variant. Here it is called AGDNR w / o SC&DS, and its structure is as Figure 8 (b) shown.
[0132] The knowledge reasoning part of the present invention aims to improve the detection accuracy by using surface categories. During the reasoning process, the learned surface features will be propagated on the features of target pixels according to the weights of the knowledge graph adjacency matrix, thereby forming enhanced features and increasing the detection accuracy. To study the impact of surface reasoning on the model performance, the present invention removes both the surface semantic feature generation part and the surface classification part during the reasoning process, and keeps other model parts unchanged.
[0133] Figure 9 are the RGB image and the ground truth map of San Diego. Figure 10Shows the object classification and detection results of AGDNR on the San Diego dataset, while Figure 11 Shows the object classification and detection results of AGDNR on the San Diego dataset. Among them Figure 11 (a) is the object detection map, and (b)-(d) are the object detection maps of different categories.
[0134] For both models, most of the target pixels can be detected. For AGDNR, there are a small number of misdetected targets. For AGDNR w / o LSF, in addition to the occurrence of target misdetection, the targets of category one are wrongly detected as targets of category two. At the same time, for each correctly detected target, the target brightness of the model without surface participation in the reasoning is lower than that of the model with surface category participation in the reasoning, indicating that the model with surface category participation in the reasoning can significantly increase the accuracy and effectiveness of target recognition and classification.
[0135] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0136] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for detecting and recognizing small and weak hyperspectral targets, characterized in that, It includes the following steps: Step 1: Obtain the hyperspectral image to be detected, calculate five surface parameters based on the infrared and near-infrared bands of the hyperspectral image. The Normalized Difference Vegetation Index (NDVI) is used to classify the surface. The remaining four surface parameters, namely Leaf Area Index (LAI), Ratio Vegetation Index (RVI), Difference Vegetation Index (DVI), and Temperature-Vegetation Dryness Index (TDVI), are spliced into the original image as additional spectra. A feature extractor is used to extract the semantic features of different surface types from the spliced hyperspectral image; Step 2: Input the hyperspectral image into a densely nested model for feature extraction to obtain the original features of the image; Step 3: Extract the semantic information of the original features based on the weights of the classifier as the semantic features of the target category, and jointly construct a global semantic pool with the semantic features of the surface types; Step 4: Hard-link the surface categories and pixels. For each pixel, it can be corresponding to a unique surface type through its NDVI value to obtain the probability matrix P2; Soft-link the target category and pixels, that is, obtain the classification probability matrix P1 on all C target categories; Finally, obtain a total probability matrix P = [P1, P2]; Step 5: Generation of enhanced features: Combine the adjacency matrix constructed by the knowledge graph, the global semantic pool, the total probability matrix, and the weight matrix of the convolutional neural network to obtain an enhanced feature vector, and adopt an attention mechanism to reduce network noise to obtain the final enhanced features; Step 6: Splice the enhanced features obtained in Step 5 and the original features obtained in Step 2, and then input the spliced features into the convolutional neural network to obtain the final detection result.
2. The hyperspectral dim and small target detection and recognition method according to claim 1, wherein The densely nested model is a stacked U-shaped fully convolutional neural network. The specific structure is I downsampling layers. The first downsampling layer has J horizontal dense blocks, and then each layer decreases by one horizontal dense block. Each horizontal dense block is specifically a convolutional attention module; Specifically, the hyperspectral image is input into the first horizontal dense block of the first layer; The input of the first horizontal dense block of each layer comes from the output of the first horizontal dense block of the previous layer after pooling; The input of the j-th horizontal dense block of the first layer comes from the outputs of all previous horizontal dense blocks in the same layer and the upsampling result of the output of the j-1-th node in the second layer; The input of the j-th horizontal dense block of the remaining layers comes from the outputs of all previous horizontal dense blocks in the same layer, the pooling result of the j-th horizontal dense block in the previous layer, and the upsampling result of the j-1-th horizontal dense block in the next layer.
3. A hyperspectral dim and small target detection and recognition method according to claim 2, characterized in that, The structure of the convolutional attention module specifically includes a convolutional layer, a channel attention module, and a spatial attention module; assume a horizontal dense block L i,j Let the output after passing through the convolutional layer of the input of i,j be L, then the calculation formula for the output of this horizontal dense block is: Among them, denotes element-wise multiplication, and M c is the channel attention module, and M s is the spatial attention module.
4. A method for detecting and recognizing small and weak hyperspectral targets according to claim 3, characterized in that, The specific content of Step 1 is as follows: First, calculate the five surface parameters, and then splice the surface parameters into the original image. The process is expressed as: input' = [input, LAI, RVI, DVI, TDVI] where input is the hyperspectral image data, and LAI, RVI, DVI, and TDVI are the surface parameters calculated from the input image respectively. [,] represents the splicing operation; The input' obtained after splicing is input into the feature extraction network for feature extraction; Use the NDVI parameter value to divide the surface types. The feature extraction process is as follows: First, the image is segmented by classifying the surface types: Among them, is the spectral value after image segmentation, input' i,j is the spectral value of the spliced image at the (i, j) position, condition k is the range of NDVI values corresponding to the k-th surface type in the divided surface types; After that, the segmented image is input into a convolutional neural network, which is expressed as: M k = MLP(Conv(F k )) Among them, Conv represents a multi-layer deep convolutional neural network for feature extraction, and MLP represents a fully connected layer with one hidden layer, which can reduce the dimension of features to one dimension, thereby obtaining one-dimensional features. represents the real number field, and L is the dimension of the semantic features of different land surface types.
5. A hyperspectral dim small target detection and recognition method according to claim 4, characterized in that, The specific steps of step 3 are as follows: First, for the original features, two convolutional layers with a convolutional kernel size of 1×1 are used for pixel-level classification. The weights of the convolutional kernels corresponding to different category targets in the last convolutional layer are extracted and represented as C semantic information of the target category, and the global semantic pool M1∈R of the target category is constructed. C×L , where C is the number of categories of the targets to be detected, and L is the weight length of the convolutional kernel; Then, the results after NDVI land surface classification and feature extraction in step 1 are used as the semantic features M2 ∈ R of various land surface types 5×L , where L is the length of the semantic features of the land surface types; Finally, the global semantic pool M = [M1, M2] of the target and the surface type is obtained.
6. The hyperspectral dim small target detection and recognition method according to claim 5, characterized in that The specific steps of step 4 are as follows: First, the process of establishing the enhanced feature vector is represented by matrix multiplication as follows: F' = PEMW G Among them, E is the adjacency matrix constructed by the knowledge graph, and M is the global semantic pool of the target category and the surface type. represents the weight matrix of the convolutional neural network, D represents the number of channels of the enhanced features, and the dimension of the enhanced feature F' is Then, the attention mechanism is used to capture the target categories and surface categories that the input image mainly focuses on, and the attention weight α of each category is obtained as α = softmax(z α W α M T ), where is the feature vector obtained after convolution of the original image, is the fully connected weight, and the attention weight of each category is The final enhanced feature is obtained: Among them, represents an element product.
7. A hyperspectral small target detection and recognition method according to claim 6, characterized in that The calculation formula of the channel attention module is shown as follows: M c (L) = σ(MLP(avgPool(L))) + σ(MLP(maxPool(L))) Among them, MLP represents a fully connected layer with one hidden layer, avgPool and maxPool represent global average pooling and max pooling for the channels of each dimension L, and σ represents the activation function; The calculation formula of the spatial attention module is shown as follows: M s (L) = σ(Conv[CavgPool(L), CmaxPool(L)]) (8) Among them, Conv represents the convolutional layer, CavgPool and CmaxPool represent channel average pooling and channel max pooling respectively.
Citation Information
Cited By
Farmland multi-modal data fusion and decision attribution method, electronic equipment and computer readable storage medium
CN122287922A
A method for multimodal data fusion and decision attribution in farmland, electronic equipment and computer-readable storage medium
CN122287922B