A method and device for target detection of air-spectrum multi-scale hyperspectral remote sensing images

By adopting a multi-scale space-spectral object detection network in hyperspectral object detection, combined with twin networks and attention mechanisms, the problem of insufficient application of spatial information in the existing technology is solved, and more efficient hyperspectral object detection is achieved.

CN116246171BActive Publication Date: 2025-06-13HARBIN ENG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310217028.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-08
Publication Date
2025-06-13
Estimated Expiration
2043-03-08

AI Technical Summary

Technical Problem

The existing hyperspectral object detection methods are not sufficient in the application of spatial information, resulting in a decrease in detection accuracy and making it difficult to be applicable to hyperspectral object detection of small-volume targets.

Method used

A multi-scale spatial-spectral object detection network is adopted to extract and fuse the spatial-spectral joint features through the similarity measurement module of the twin network structure, the spatial-spectral feature extraction module based on the attention mechanism, and the multi-scale spatial-spectral difference feature mixing module, combined with a one-dimensional convolutional neural network and visual attention mechanism.

Benefits of technology

It improves the accuracy and ability of hyperspectral object detection, can more effectively extract and distinguish targets from backgrounds, and improves the reliability of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246171B_ABST
    Figure CN116246171B_ABST
Patent Text Reader

Abstract

An empty-spectrum multi-scale hyperspectral remote sensing image target detection method and device, belonging to the technical field of hyperspectral remote sensing image target detection. In order to solve the problem that the existing hyperspectral target detection does not make full use of spatial information and the resulting weak hyperspectral target detection ability. Based on the hyperspectral data of the target area to be detected, the present invention constructs a sample from three-dimensional spatial pixel blocks, which is called a patch; constructs sample pairs from the patches of all pixels in the image to be detected and the patches of the target prior; inputs the sample pairs into a multi-scale empty-spectrum target detection network for target detection. First, two independent branches with shared structures and parameters are used to process the two patches in the input sample pair respectively. On each independent branch, spectral features and spatial features are extracted, and then they are merged after the last downsampling unit in the two independent branches. The output features are given the similarity of the two inputs after multi-scale differential feature mixing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hyperspectral remote sensing image target detection, and particularly relates to a hyperspectral remote sensing image target detection method and device. Background Art

[0002] A hyperspectral image is a three-dimensional data cube containing dozens or even hundreds of bands, and simultaneously contains the spatial information and spectral information of ground objects. Benefiting from the rich spectral information of hyperspectral images, hyperspectral images have important applications in fields such as geological exploration, precision agriculture, and urban monitoring. Hyperspectral target detection is a technology that identifies and separates targets from the background based on a small amount of prior information of the targets, and it can achieve fine discrimination of specific targets. Its task characteristics are limited prior information, small target proportion, and small size.

[0003] Currently, a large number of deep learning-based algorithms have been applied to hyperspectral target detection. However, these deep learning-based target detection methods still have the problem of insufficient application of spatial information. On the one hand, the method of introducing spatial information by using post-processing means such as guided filtering increases the operation complexity of the detection method and is difficult to extract non-linearly correlated features; on the other hand, the spatial information extraction methods based on two-dimensional convolution and three-dimensional convolution are applicable to larger spatial sizes, while in hyperspectral target detection, the target size is small and the number is small. The introduction of larger spatial size information may instead cause a decrease in detection accuracy. Therefore, the two-dimensional convolution and three-dimensional convolution widely used in hyperspectral classification and hyperspectral change detection are not applicable to hyperspectral target detection of small-volume targets, and their detection results are difficult to be ideal. Summary of the Invention

[0004] The present invention aims to solve the problem of insufficient application of spatial information in existing hyperspectral target detection and the resulting weak hyperspectral target detection ability.

[0005] A spatial-spectral multi-scale hyperspectral remote sensing image target detection method, based on the hyperspectral data of the target area to be detected, forms a sample from three-dimensional spatial pixel blocks, and this pixel block is called a patch; constructs sample pairs from the patches corresponding to all pixels in the image to be detected and the patches corresponding to the target prior; inputs the sample pairs into a multi-scale spatial-spectral target detection network for target detection. The multi-scale spatial-spectral target detection network is composed of a similarity measurement module with a siamese network structure, a spatial-spectral feature extraction module based on an attention mechanism, and a multi-scale spatial-spectral difference feature mixing module. The specific structure is as follows:

[0006] Similarity Measurement Module Based on Siamese Network Structure: The similarity measurement module based on the siamese network structure is implemented based on the siamese network; the siamese network is a kind of conjoined network structure. It first processes two patch pixel blocks in the input sample pair through two independent branches with shared structure and parameters respectively, and then merges them after the last downsampling unit in the two independent branches. The output features are given in the form of similarity scores after passing through the upsampling unit to represent the similarity between the two inputs;

[0007] On two independent branches with shared structure and parameters, a spatial-spectral feature extraction module based on the attention mechanism is set; each spatial-spectral feature extraction module contains two branches - a spectral feature extraction branch and a spatial feature extraction branch. First, the spectral and spatial information of the pixel is extracted through spectral feature extraction and spatial feature extraction respectively, and then the two are fused to obtain the spatial-spectral fusion feature;

[0008] The spectral feature extraction branch takes the central pixel p ∈ R of the patch as the input, and band is the number of bands; the spectral feature extraction branch includes four downsampling units; each downsampling unit consists of 2 or 3 convolutional layers and a max pooling layer. After each convolutional layer, a batch normalization layer and a dropout layer are added to prevent overfitting. The batch normalization layer is the BN layer; after spectral feature extraction, the spectral feature of the pixel is obtained; 1×band The spatial feature extraction branch takes the entire pixel block patch ∈ R as the input, where l is the spatial size of the pixel block; the spatial feature extraction includes a two-dimensional convolutional layer and a visual attention unit, and the visual attention unit is the ViT unit;

[0009] After spectral feature extraction and spatial feature extraction, the spectral feature obtained from the convolutional output of the fourth downsampling unit and the finally output spatial feature are processed by contact, and then sent to the max pooling layer in the fourth downsampling unit for processing to obtain the spatial-spectral joint feature of the pixel; l×l×band The multi-scale spatial-spectral difference feature mixing module includes four upsampling units. Each upsampling unit contains 2 or 3 transposed convolutions. The first transposed convolution performs upsampling, and the remaining transposed convolutions realize feature aggregation and extraction; each transposed convolution is followed by a BN layer and a dropout layer;

[0010]

[0011]

[0012] ​​The output of each upsampling unit is connected to the output of the downsampling stage at the corresponding scale in the spatial-spectral feature extraction module through a long skip connection, serving as the input for the next upsampling unit; after passing through four upsampling units, four different-scale spatial-spectral difference mixed features are obtained; finally, the features at three different scales are summed with weights, and the resulting tensor passes through a fully connected layer and an activation function layer to obtain a similarity score; object detection is achieved based on the similarity score.

[0013] Further, the number of convolution kernels in the convolution layers of the four downsampling units in the spectral feature extraction branch increases successively as 8, 16, 32, and 64.

[0014] Further, the number of convolution kernels in the four upsampling units in the multi-scale spatial-spectral difference feature mixing module decreases in the order of 64, 32, 16, and 8.

[0015] Further, the features at the last three different scales among the four upsampling units are summed with weights of 1 / 8, 1 / 4, and 1.

[0016] Further, the ViT unit includes a multi-head attention module and a feed-forward network.

[0017] Further, the training process of the multi-scale spatial-spectral object detection network includes the following steps:

[0018] First, construct a training dataset and a test dataset: use object prior information and background prior information to construct two types of sample pairs, including sample pairs from the same-class prior and sample pairs from different-class priors; sample pairs from the same class are labeled with a pseudo-label of 0, representing similarity; sample pairs from different classes are labeled with a pseudo-label of 1, representing difference;

[0019] Then, use the constructed training dataset to train the built multi-scale spatial-spectral object detection network; in the training stage, send the training sample set and its pseudo-labels into the multi-scale spatial-spectral object detection network to train the network; the trained network can distinguish whether two input pixels are similar and perform a similarity score.

[0020] Further, binary cross-entropy is used as the loss function during the training of the multi-scale spatial-spectral object detection network.

[0021] A computer storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the described method for object detection of spatial-spectral multi-scale hyperspectral remote sensing images.

[0022] An air-spectrum multi-scale hyperspectral remote sensing image target detection device, the device includes a processor and a memory, and at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the air-spectrum multi-scale hyperspectral remote sensing image target detection method described above.

[0023] Beneficial effects:

[0024] The hyperspectral target detection method proposed by the present invention jointly utilizes a one-dimensional convolutional neural network and visual attention to extract air-spectrum joint features. Compared with the networks that only consider spectral information and use two-dimensional / three-dimensional convolutional neural networks (convolutional attention) to extract spatial information, this method can not only specifically extract spectral features through a one-dimensional convolutional neural network, but also adaptively extract spatial features through spatial attention within a small-size spatial range. In order to further distinguish the target from the background, this method adopts a siamese network to expand the intra-class similarity and inter-class difference, so that the target is easier to distinguish. In addition, the present invention adopts multi-scale fusion features instead of single-scale features to obtain more discriminative air-spectrum difference features. The entire detector finally gives the detection result in the form of a similarity score. The higher the similarity score, the more likely it is to be a target. Through experimental analysis, the method proposed by the present invention can obtain AUC values of 0.9992, 0.9937, and 0.9983 on three datasets respectively. Description of the drawings

[0025] Figure 1 It is a flowchart of the multi-scale air-spectrum target detection method.

[0026] Figure 2 It is a structural diagram of the multi-scale air-spectrum hyperspectral target detection network based on the attention mechanism.

[0027] Figure 3(a), Figure 3(b), and Figure 3(c) are respectively the false color image, ground truth image, and detection result of Dataset I.

[0028] Figure 4(a), Figure 4(b), and Figure 4(c) are respectively the false color image, ground truth image, and detection result of Dataset II.

[0029] Figure 5(a), Figure 5(b), and Figure 5(c) are respectively the false color image, ground truth image, and detection result of Dataset III. Detailed implementation manners

[0030] The present invention proposes a spatial-spectral multi-scale hyperspectral target detection method based on an attention mechanism, which uses a one-dimensional convolutional neural network and an attention mechanism to extract spectral information and spatial information respectively and fuse them to obtain spatial-spectral hybrid features. Compared with two-dimensional convolution and three-dimensional convolution, the Vision Transformer attention used in the present invention can adaptively extract spatial information within a smaller spatial size range, thus more completely representing pixel features and improving the detection results of the method.

[0031] The present invention uses a siamese network as the framework of this method. Different from the structure that directly extracts pixel features, the siamese network extracts the differential features of pixels. Therefore, it can expand the inter-class differences and reduce the intra-class differences, thereby promoting the comparison between the target and the background and better separating the target.

[0032] In order to better extract spatial-spectral differential features, the present invention also designs a multi-scale feature extraction module. The multi-scale feature extraction module overcomes the problem that most deep learning methods only use the single-scale features output by the last layer of the network and ignore the important features of other layers, thereby obtaining more discriminative differential features and further promoting the discrimination and detection of the target.

[0033] The following will describe the present invention in detail in conjunction with specific embodiments.

[0034] Specific Embodiment 1: In conjunction with Figure 1 to illustrate this embodiment, Figure 1 Band selection in [reference] represents band selection, Train stage represents the training stage, Testing stage represents the testing stage, Background prior patches represents the pixel blocks of background prior information, Target prior patches represents the pixel blocks of target prior information, testpatches represents the test pixel blocks, Attention-based multiscale spectral-spatial detector (AMSSD) represents the attention-based multi-scale spectral-spatial detector, that is, the multi-scale spatial-spectral target detection network of the present invention, and Similarity score represents the similarity score.

[0035] This embodiment is a spatial-spectral multi-scale hyperspectral remote sensing image target detection method, including the following steps:

[0036] Step 1: Construct a training data set and a test data set:

[0037] Deep learning-based networks require a large amount of training data, while the prior target information for hyperspectral target detection is often limited. Based on this, the present invention expands the number of training samples by pixel block pairing to meet the network training requirements:

[0038] In the present invention, a three-dimensional spatial pixel block of a certain size constitutes a sample, and this pixel block is called a patch.

[0039] To construct a training sample set, two types of sample pairs are constructed using target prior information and background prior information - sample pairs from the same category prior (target and target, background and background) and sample pairs from different category priors (target and background). Sample pairs from the same category are labeled with a pseudo-label of 0, representing similarity; sample pairs from different categories are labeled with a pseudo-label of 1, representing difference.

[0040] To construct a test sample set, sample pairs are constructed by pairing all the patches corresponding to the pixels in the image to be tested with the patches corresponding to the target prior.

[0041] In the training stage, the training sample set and its pseudo-labels are fed into the designed network for training. The trained network can distinguish whether two input pixels are similar or different. In the test stage, the test sample set is fed into the trained network, and the network will give the similarity score between the test pixel and the target prior. A high similarity score is judged as a target, otherwise it is judged as background.

[0042] Step 2: Build a multi-scale spatio-spectral target detection network, including the following steps:

[0043] Step 2.1: Build a similarity measurement module based on the Siamese network structure:

[0044] The similarity measurement module based on the Siamese network structure designed in the present invention is implemented based on the Siamese network. The Siamese network is a kind of conjoined network structure. It first processes the two patch pixels in the input sample pair through two independent branches with shared structure and parameters, and then merges them after the last downsampling unit (see Step 2.2) in the two independent branches. The output features are given in the form of a similarity score after passing through the upsampling (see Step 2.3) module.

[0045] For binary classification problems, the Siamese network often uses binary cross entropy as the loss function, and its expression is

[0046]

[0047] l n =-ω[y n lopxn +(1 - y n ) log(1 - x n )]

[0048] Where x n and y n are the output and label value of the Siamese network respectively, is a hyperparameter, for single - label binary classification, it can be ignored. Different from the network that directly extracts input features, the Siamese network extracts the similarity features of two inputs. Therefore, it can expand the differences between different classes and the similarities between the same classes, thereby promoting the comparison between the target and the background, and enabling the target to be better extracted from the background. At the same time, the characteristic of parameter sharing in the Siamese network structure can reduce the network computation amount and improve the algorithm efficiency.

[0049] Step 2.2: Build an attention - based spatial - spectral feature extraction module on two independent branches with shared structure and parameters respectively:

[0050] Most of the existing deep - learning - based object detection methods only consider spectral feature extraction, and for the methods that utilize both spatial and spectral information, there are still problems of insufficient utilization of spatial information. For example, the methods of using two - dimensional or three - dimensional convolution to extract spatial information can extract spatial information, but they are often applicable to large - size spatial pixel blocks, and it is difficult to extract useful spatial information or even cause interference when the spatial size is small. Therefore, the present invention designs a new spatial - spectral joint feature extraction method. As Figure 2 shown, each spatial - spectral feature extraction module contains two branches - a spectral feature extraction branch and a spatial feature extraction branch. First, the spectral and spatial information of the pixel are extracted through spectral feature extraction and spatial feature extraction respectively, and then the two are fused to obtain the spatial - spectral fusion feature.

[0051] The spectral feature extraction branch takes the central pixel p ∈ R 1×band of the patch as the input, and band is the number of bands. The spectral feature extraction branch includes four down - sampling units. Each down - sampling unit consists of 2 or 3 convolutional layers and a max - pooling layer. A batch normalization layer and a dropout layer are added after each convolutional layer to prevent overfitting. The batch normalization layer is the BN (batch normalization) layer. The number of convolutional kernels in the convolutional layers of the down - sampling unit increases successively according to the rule of 8, 16, 32, 64, and the spectral feature of the pixel is obtained after spectral feature extraction.

[0052] The spatial feature extraction branch takes the entire pixel block patch ∈ R l×l×bandis the input, where l is the spatial dimension of the pixel block. Spatial feature extraction includes a two-dimensional convolutional layer and a visual attention unit; the visual attention unit is the ViT (vision Transformer) unit.

[0053] The two-dimensional convolutional layer aggregates the spatial information of the pixel blocks while adjusting the dimension of the patch to better achieve subsequent spatial feature extraction. ViT is used to adaptively capture spatial features in a small-size space. It utilizes the encoder of the Transformer and uses an additional token to represent the global feature. ViT contains a multi-head attention module and a feed-forward network. The tensor obtained after two-dimensional convolution is position-encoded after passing through a linear mapping layer. The encoded embedding vectors embeddings are then fed into the Transformer encoder, and finally, the global spatial feature is given by the additional token.

[0054] The multi-head attention is composed of self-attention. In self-attention, each embedding vector is first multiplied by three different learnable matrices to obtain three different vectors, namely the query vector q, the key vector k, and the value vector v. All q, k, and v form tensors Q, K, and V. The self-attention formula is

[0055]

[0056] where d is the length of the query vector and the key vector. To better extract relevant information and prevent self-attention from over-focusing on itself, multi-head attention is proposed. The multi-head attention expression is

[0057] MultiHead(Q,K,V)=Concat(head 1 ,head 2 ,…,head h )×W o

[0058] head i =Attention(Q,K,V)

[0059] where h is the number of heads in the multi-head attention, and W o is the learnable parameter matrix.

[0060] After spectral feature extraction and spatial feature extraction, the global spatial feature and the spectral feature (obtained from the convolution output of the fourth downsampling unit) are subjected to contact processing, and then sent to the max-pooling layer in the fourth downsampling unit for processing to obtain the spatial-spectral joint feature of the pixel.

[0061] Step 2.3: Build a multi-scale spatial-spectral difference feature mixing module:

[0062] Most deep learning-based algorithms only utilize the features output by the last layer of the network, while ignoring the important features of different scales in other layers. Therefore, the present invention designs a multi-scale spatial-spectral difference feature mixing module to extract more distinguishable difference features, so as to better distinguish the target from the background.

[0063] The multi-scale spatial-spectral difference feature mixing module includes four upsampling units. Each upsampling unit contains 2 or 3 transposed convolutions. The first transposed convolution performs upsampling, and the remaining transposed convolutions realize feature aggregation and extraction. A BN layer and a dropout layer follow each transposed convolution. The number of convolution kernels of the upsampling units decreases in the order of 64, 32, 16, 8.

[0064] The difference between the output of each upsampling unit and the output of the corresponding scale's downsampling stage in the spatial-spectral feature extraction module is connected through a long skip connection and used as the input of the next upsampling unit. The long skip connection can fuse the features of the deep layer and the shallow layer, so as to better represent the features. After four upsampling units, four different scales of spatial-spectral difference mixed features are obtained. The last three different scales of features are added with weights of 1, and the resulting tensor passes through a fully connected layer and a sigmoid layer to obtain a similarity score.

[0065] Step 3: Use the constructed training dataset to train the built multi-scale spatial-spectral object detection network; the trained network can distinguish whether two input pixels are similar and give a similarity score.

[0066] Step 4: Use the trained multi-scale spatial-spectral object detection network to process the test dataset for object detection:

[0067] In the object detection stage, all pixels in the image to be detected are respectively paired with the target prior and input into the network. The network will give a similarity score between the pixel to be detected and the target. The higher the similarity score (the closer the pseudo-label is to 0), the more likely the pixel to be detected is the target, otherwise it is judged as the background.

[0068] In the hyperspectral remote sensing image object detector, each pixel to be detected is paired with multiple prior target pixels, and the final similarity score is the average of the similarities between the pixel to be detected and each prior target pixel.

[0069] The present invention adaptively extracts spatial information by introducing an attention mechanism, and integrates the extracted spatial-spectral difference information through a siamese network structure to achieve better object detection. Aiming at the problem that most deep learning-based networks only utilize the extracted single-scale features, a multi-scale feature fusion module is designed to more completely extract multi-scale features and reasonably utilize the important features of different scales in the network, thereby improving the detection accuracy. Specific Embodiment 2:

[0071] This embodiment is a computer storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the method for object detection of spatial-spectral multi-scale hyperspectral remote sensing images.

[0072] It should be understood that the instructions include computer program products, software or computerized methods corresponding to any method described in the present invention; the instructions can be used to program a computer system or other electronic devices. The computer storage medium may include a readable medium on which instructions are stored, which may include but are not limited to magnetic storage media and optical storage media; magneto-optical storage media include read-only memory ROM, random access memory RAM, erasable programmable memory (e.g., EPROM and EEPROM), and flash memory layers, or other types of media suitable for storing electronic instructions. Specific Embodiment 3:

[0074] This embodiment is a device for object detection of spatial-spectral multi-scale hyperspectral remote sensing images. The device includes a processor and a memory. It should be understood that it includes any device including a processor and a memory described in the present invention. The device may also include other units and modules for display, interaction, processing, control, etc. and other functions through signals or instructions;

[0075] At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the method for object detection of spatial-spectral multi-scale hyperspectral remote sensing images.

[0076] Embodiment

[0077] In this embodiment, three hyperspectral data are used to illustrate the effect of the multi-scale spatial-spectral detection method based on the attention mechanism proposed by the present invention. The detailed information of the three data used is listed in Table 1. The experimental results use the area under the receiver operating characteristic curve (AUC) as the evaluation index. The higher the value of AUC, the better the detection effect.

[0078] Table 1 Detailed information of the hyperspectral images used

[0079]

[0080] For different data, the optimal parameter settings of the method of the present invention are shown in Table 2. Learning rate is the learning rate, Batch size is the minimum batch size, and Epoch is the training cycle. Patch size is the size of the spatial block. The language environment of this invention is python, and the experimental hardware platform is an Intel(R) Core(TM) i5-7200U CPU with 8GB of memory.

[0081] Optimal parameters and AUC values on three groups of experimental data in Table 2

[0082]

[0083] Through experimental analysis, the method proposed by the present invention can obtain AUC values of 0.9992, 0.9937, and 0.9983 on three datasets respectively. The false color images, true value images, and detection results of Dataset I, Dataset II, and Dataset III are shown in Figures Figures 3(a) to 3(c) , Figures 4(a) to 4(c) , Figures 5(a) to 5(c) as follows.

[0084] The above examples of the present invention are only to illustrate in detail the calculation model and calculation process of the present invention, rather than to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is impossible to list all the implementation manners here. Any obvious changes or modifications derived from the technical solutions of the present invention still fall within the protection scope of the present invention.

Claims

1. An object detection method for air-spectrum multi-scale hyperspectral remote sensing images, characterized in that, based on the hyperspectral data of the target area to be detected, a three-dimensional spatial pixel block forms a sample, and this pixel block is called a patch; sample pairs are constructed from the patches corresponding to all pixels in the image to be detected and the patches corresponding to the target prior; the sample pairs are input into a multi-scale air-spectrum object detection network for object detection. The multi-scale air-spectrum object detection network consists of a similarity measurement module with a siamese network structure, an air-spectrum feature extraction module based on the attention mechanism, and a multi-scale air-spectrum difference feature mixing module. The specific structure is as follows: Similarity measurement module based on the siamese network structure: The similarity measurement module based on the siamese network structure is implemented based on the siamese network; the siamese network is a kind of conjoined network structure. It first processes the two patch pixel blocks in the input sample pair through two independent branches with shared structure and parameters, and then merges after the last downsampling unit in the two independent branches. The output feature is given as the similarity of the two inputs in the form of a similarity score after passing through the upsampling unit; Air-spectrum feature extraction modules based on the attention mechanism are respectively set on the two independent branches with shared structure and parameters; each air-spectrum feature extraction module contains two branches - a spectral feature extraction branch and a spatial feature extraction branch. First, the spectral and spatial information of the pixel is extracted through spectral feature extraction and spatial feature extraction respectively, and then the two are fused to obtain the air-spectrum fusion feature; The spectral feature extraction branch takes the central pixel p ∈ R of the patch 1×band as the input, and band is the number of bands; the spectral feature extraction branch includes four downsampling units; each downsampling unit consists of 2 or 3 convolutional layers and a max pooling layer, and a batch normalization layer and a dropout layer are added after each convolutional layer to prevent overfitting, and the batch normalization layer is the BN layer; after spectral feature extraction, the spectral features of the pixels are obtained; The spatial feature extraction branch takes the entire pixel patch patch∈R l×l×band as input, where l is the spatial dimension of the pixel patch; the spatial feature extraction includes a two-dimensional convolutional layer and a visual attention unit, that is, a ViT unit; After spectral feature extraction and spatial feature extraction, the spectral feature obtained from the convolution output of the fourth downsampling unit and the finally output spatial feature are processed by contact, and then sent to the max pooling layer in the fourth downsampling unit for processing to obtain the air-spectrum joint feature of the pixel; The multi-scale air-spectrum difference feature mixing module includes four upsampling units. Each upsampling unit contains 2 or 3 transposed convolutions. The first transposed convolution performs upsampling, and the remaining transposed convolutions realize feature aggregation and extraction; a BN layer and a dropout layer follow each transposed convolution; The output of each upsampling unit is connected to the difference between the output of the corresponding scale in the downsampling stage of the air-spectrum feature extraction module through a long skip connection as the input of the next upsampling unit; after passing through the four upsampling units, four different scales of air-spectrum difference mixed features are obtained; finally, the features of the three different scales are added by weight, and the resulting tensor passes through a fully connected layer and an activation function layer to obtain a similarity score; object detection is realized based on the similarity score.

2. An object detection method for air-spectrum multi-scale hyperspectral remote sensing images according to claim 1, characterized in that, the number of convolution kernels in the convolution layer of the downsampling unit in the four downsampling units of the spectral feature extraction branch increases successively as 8, 16, 32, 64.

3. An object detection method for air-spectrum multi-scale hyperspectral remote sensing images according to claim 2, characterized in that, Among the four upsampling units of the multi-scale spatial-spectral difference feature mixing module, the number of convolution kernels of the upsampling units decreases in the order of 64, 32, 16, and 8.

4. An object detection method for spatial-spectral multi-scale hyperspectral remote sensing images according to claim 1, 2, or 3, characterized in that the features of the last three different scales in the four upsampling units are added with weights of 1 / 8, 1 / 4, and 1.

5. An object detection method for spatial-spectral multi-scale hyperspectral remote sensing images according to claim 4, characterized in that the ViT unit includes a multi-head attention module and a feed-forward network.

6. An object detection method for spatial-spectral multi-scale hyperspectral remote sensing images according to claim 5, characterized in that the training process of the multi-scale spatial-spectral object detection network includes the following steps: First, construct a training data set and a test data set: use target prior information and background prior information to construct two types of sample pairs, the sample pairs from the same-class prior and the sample pairs from different-class priors; the sample pairs from the same class are marked with a pseudo-label of 0, representing similarity; the sample pairs from different classes are marked with a pseudo-label of 1, representing difference; Then, use the constructed training data set to train the constructed multi-scale spatial-spectral object detection network; in the training stage, send the training sample set and its pseudo-labels into the multi-scale spatial-spectral object detection network to train the network; the trained network can distinguish whether two input pixels are similar and perform similarity scoring.

7. An object detection method for spatial-spectral multi-scale hyperspectral remote sensing images according to claim 6, characterized in that binary cross-entropy is used as the loss function during the training of the multi-scale spatial-spectral object detection network.

8. A computer storage medium, characterized in that at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement an object detection method for spatial-spectral multi-scale hyperspectral remote sensing images according to any one of claims 1 to 7.

9. An object detection device for spatial-spectral multi-scale hyperspectral remote sensing images, characterized in that the device includes a processor and a memory, and at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement an object detection method for spatial-spectral multi-scale hyperspectral remote sensing images according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Hyperspectral target detection method based on multi-example twin network

    CN113723482A

  • Hyperspectral image saliency map generation method based on semi-supervised neural network

    CN114359675A