Hyperspectral image target detection method based on spatial-spectral joint dual networks
Through the dual-network method of null spectrum combined with the generative adversarial network and convolutional neural network, the problems of insufficient samples and incomplete information utilization in hyperspectral object detection are solved, and high-precision object detection is achieved, which improves the detection effect.
Patent Information
- Application Number
- CN202411884564.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-07-08
AI Technical Summary
There are problems in the existing hyperspectral object detection methods such as insufficient samples, high computational complexity and ignoring spatial information, resulting in poor detection results.
The dual-network method based on null spectrum is adopted to obtain background samples through PCA and superpixel segmentation, and the generative adversarial network and convolutional neural network are trained. The adversarial loss function and focus loss function of the generator and discriminator are used to optimize the network, and the results are fused with the guide filter to realize joint detection of null spectrum.
It improves the accuracy and robustness of hyperspectral image object detection, makes full use of spectral and spatial information, alleviates the overfitting problem, and enhances the stability and feature extraction ability of the detection results.
Smart Images

Figure CN120279403A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of microscopic hyperspectral image processing, and particularly relates to a dual-network hyperspectral image target detection method based on the combination of spatial and spectral information. Background Art
[0002] Hyperspectral images (HSIs) are images that contain hundreds of narrow bands and can be used to measure the reflection characteristics of imaging objects. Different materials have unique electromagnetic reflection characteristics, so hyperspectral images can be used to detect and identify objects. At the same time, these images cover a wide range of the electromagnetic spectrum, and each pixel provides continuous spectral information, which helps to distinguish different targets. Therefore, hyperspectral images have a wide range of applications in fields such as mineral identification, urban planning, change detection, and land cover classification. In these applications, the detection of targets of interest has received increasing attention, and target detection is one of the most important applications.
[0003] Hyperspectral target detection is a method of detecting and locating targets in an image under the condition of a given target prior, and its ultimate goal is to classify pixels into two categories: targets and backgrounds. In recent years, many methods have flourished in hyperspectral target detection. The sparsity-based method assumes that the spectra of targets or backgrounds can be sparsely represented by target or background dictionaries. The sparse representation model has received extensive attention because it can obtain better results without assuming the statistical distribution characteristics of data compared with traditional methods. Only a small number of elements in the over-complete dictionary composed of training samples can be used to approximately represent the test samples.
[0004] Although progress has been made and many problems in the detection task have been solved, there are still challenges in the field of hyperspectral target detection. First, due to the influence of sensor noise and other factors, the target spectra of the same material may vary greatly, which makes it difficult to represent the target spectra using only a single spectrum or subspace. In addition, since the spatial resolution of HSI is usually lower than that of multispectral and color images, there are still mixed pixels covering multiple materials, which makes the detection task more complex. In addition, most target detection problems are non-linear, but the methods mentioned above are more likely to make decisions under linear assumptions, which does not conform to the complex reality of HSI processing. In addition, it is worth noting that when the non-linearity increases, many of these methods do not work well.
[0005] Nowadays, due to the powerful feature extraction ability of neural networks, deep learning-based methods have been widely applied to the field of hyperspectral image target detection. The above situation indicates the importance of directly using neural networks as detectors. However, deep learning-based target detection algorithms have problems such as insufficient training samples and high computational complexity, and most detectors only utilize the spectral information of individual pixel points in hyperspectral images while ignoring the spatial information. Summary of the Invention
[0006] Object of the Invention: In order to overcome the deficiencies of the prior art and solve problems such as insufficient samples, high computational complexity in object detection methods based on deep learning, and most detectors only utilizing the spectral information of hyperspectral images while ignoring the spatial information, a hyperspectral image object detection method based on a joint spatial-spectral dual network is proposed.
[0007] Technical Solution: To achieve the above object, the present invention is implemented by the following technical methods. A hyperspectral image object detection method based on a joint spatial-spectral dual network includes the following steps:
[0008] S1: First, perform PCA by eigenvalue decomposition of the covariance matrix on the hyperspectral image. After obtaining the dimensionality-reduced image, use the simple linear iterative clustering algorithm for superpixel segmentation to obtain a superpixel image, and then obtain background samples through sparse representation and sample selection strategies.
[0009] S2: Based on the background samples obtained above for training, send them into a generative adversarial network to learn the background distribution. At the same time, send the rough detection results obtained by minimizing the constrained energy of the original hyperspectral image through a classical detector into a convolutional neural network and perform spatial filtering, thereby realizing joint spatial-spectral, and finally adaptively fuse the results of the above two to obtain a trained model.
[0010] S3: Send the original hyperspectral image data into the trained network model, and through adaptive fusion. At the same time, in order to suppress the background, perform nonlinear suppression and guided filtering to obtain the final detection result.
[0011] Furthermore, the hyperspectral image object detection method based on a joint spatial-spectral dual network is characterized in that the following sub-steps are included in the step S1:
[0012] S11: In order to construct an object detector based on a neural network, it is first necessary to construct a representative training set. Since classical detectors can effectively suppress the background, a classical detector is used for pre-detection, and then the results of the pre-detection are used to construct the training set.
[0013] Suppose \(X\in R\) L×N represents the spectra of all \(N\) pixels with \(L\) bands. After pre-detection, we will obtain \(N\) outputs, and the pixels with larger values are more likely to be targets. Then, we sort these pixels in descending order, that is, \(X\) s \(=[x_1,\) x2 ,\(...\), \(x\) n , which satisfies where, represents the pixel \(x\)i Output. By setting two different thresholds T a and T b (T a >T b ), these pixels can be divided into an ordered target set X t , an ordered mixed set, and an ordered background set X b . For simplicity, T a and T b are respectively set to (1 / 4) and (1 / 5)
[0014] S12: After constructing the training set, we attempt to build a dual parallel network for HTD. The first network is a GAN-based network, and the other is a CNN network. Regarding the specific network structure, it should be noted that the GAN network is divided into a generator network and a discriminator network. In addition, we followed different criteria when constructing these two object detection networks.
[0015] Furthermore, the above-mentioned hyperspectral image object detection method based on an empty-spectrum joint dual network is characterized in that the step S2 includes the following sub-steps:
[0016] S21: The generator adopts a residual structure, which has strong expressiveness and can well simulate the complex distributions of target and background spectra. At the same time, in order to stabilize the training process, the spectral normalization layer proposed in is added between every two fully connected layers. For the discriminator, it provides two probabilities, namely the probability that the sample is a real sample and the probability that the sample is a target.
[0017] For the loss function of the whole process, there is no difference from the traditional GAN, and the overall formula is as follows:
[0018]
[0019] Optimizing the adversarial loss V(D, G) can achieve two purposes at the same time. The first purpose is to enable the generator G to generate real samples, and the second purpose is to enable the discriminator D to better distinguish between real samples and generated samples.
[0020] More specifically, the above-mentioned adversarial loss function can be divided into a generator G loss function and a discriminator D loss function, as follows:
[0021] Generator G loss function: The loss function for training the generator is actually the term related to the noise z in the adversarial loss V(D, G), and its loss function is:
[0022]
[0023] The goal of the generator is to make the samples generated by the generator as realistic as possible, which is specifically reflected in the mathematical formula:
[0024]
[0025] where BCE() represents the binary cross-entropy loss function. From the above, it can be seen that the generator in GAN minimizes the loss function L G which can be written in the form of minimizing the binary cross-entropy function.
[0026] Discriminator D loss function: The loss function for training the discriminator is actually the term related to the sample x in the adversarial loss V(D, G), and its loss function is:
[0027]
[0028] The goal of the discriminator is to make the discriminator better distinguish between generated samples and real samples, which is specifically reflected in the mathematical formula as (5):
[0029]
[0030]
[0031] where x represents the real data sample and x' represents the sample generated by the generator. From the above, it can be found that the discriminator in GAN maximizes the loss function L D which can be written in the form of minimizing two binary cross-entropy functions.
[0032] For and To alleviate the overfitting problem caused by finite and unbalanced samples, focal loss is adopted, which is expressed as follows:
[0033]
[0034] where y c represents the class output of x provided by the discriminator, indicating the probability that x belongs to the target class, and y ∈ {0, 1} represents the class labels for real and generated samples. γ represents an adjustable focusing parameter and is set to 2 in our experiment.
[0035]
[0036] As shown above, L p is the absolute cosine between two generated spectra, which allows the generator to generate samples with a high level of divergence in order to alleviate the problem of mode collapse and produce better results.
[0037] By minimizing these modified objective functions, as in traditional GANs, the filter and the generator are alternately optimized. When the network converges, the neural network can provide a class label for each sample, representing the probability that the sample belongs to the target.
[0038] S22: Different from typical CNNs, the proposed CNN is entirely constructed from CNNs without spatial pooling layers or fully connected layers, which ensures that the spatial resolution is constant across all layers. A hierarchical structure is adopted in the network, and the feature structures of different layers are cascaded for target detection. Meanwhile, a channel pooling layer is adopted in the internal layers. To address the overfitting problem, the dropout method is adopted during network training.
[0039] The loss of the objective function of the CNN is the same as the above formula, and the network is trained through backpropagation. As the network converges, the internal layers can extract useful features, and the last layer can detect the target based on the spatial distribution of the training samples and their spectral information.
[0040] Furthermore, in the described hyperspectral image target detection method based on an empty-spectrum joint dual network, it is characterized in that the step S3 includes the following sub-steps:
[0041] S31: Fusion of results: The dual networks discussed above are quite different, so there are small differences in the results. The GAN-based network focuses on spectral information and is more sensitive to single pixels, while the CNN processes both spatial and spectral information and shows smoother results. To fuse the results, a pixel-level fusion strategy is adopted, which is expressed as follows:
[0042] D f = μD G +(1 - μ)D c
[0043] where D G and D C represent the detection results of the GAN-based network and the CNN respectively, and μ represents the average weight and is set to 0.5. After fusing the detection results, more distinctive and robust detection results can be obtained.
[0044] S32: Guided Filter: In addition, since spatial correlation is useful in detecting certain objects in HSI, a set of guided filters can be utilized to smooth the initial detection results. The initial detection maps obtained at different times are first filtered separately by the filter bank. In this case, all filters will use the initial input image as the guidance image, and they only differ from each other in terms of the radius r and the regularization coefficient ε. Then, the initial detection maps are smoothed to different degrees and scales. After that, max pooling is performed on the filtered images to collect the most valuable information. Finally, the overall average value of the images obtained at different times is calculated to obtain a relatively stable detection result.
[0045] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0046] (1) In order to fully obtain the spectral information and spatial information of the hyperspectral image, modules of superpixel segmentation, position encoding, and multi-head attention mechanism are migrated.
[0047] (2) A new dual-branch generative adversarial network is proposed, which can more fully learn different features of the samples and perform adaptive fusion of the detection results, making the detection results more accurate.
[0048] (3) The guided filter is utilized, which is a commonly used image enhancement and edge-preserving filtering method. It combines the guidance image and the target image, and filters the target image through the structural information of the guidance image to improve the performance of target detection. Then, by using the non-linear suppression function, the separability between the target and the background is further improved. Description of the Drawings
[0049] Figure 1 This is the flowchart of the present invention.
[0050] Figure 2 This is the schematic diagram of the entire model framework in the inventive method.
[0051] Figure 3 This is the detailed information of the layers in the GAN network.
[0052] Figure 4 This is the detailed information of the layers in the CNN network Detailed Embodiments
[0053] The following further clarifies the present invention in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification made by those skilled in the art fall within the scope defined by the appended claims of this application.
[0054] As Figure 1As shown in the figure, a hyperspectral image target detection method based on the joint of spatial and spectral features and dual networks disclosed in an embodiment of the present invention specifically includes the following steps:
[0055] S1: Construct the schematic diagram of the entire model framework as shown in Figure 2 the figure. The entire process is mainly divided into two parts: one is the training and detection stage of the generative adversarial network, and the other is the training and detection stage of the convolutional neural network.
[0056] Specifically, step S1 includes the following sub-steps:
[0057] S11: In order to construct a target detector based on a neural network, it is first necessary to construct a representative training set. Since classical detectors can effectively suppress the background, classical detectors are used for pre-detection, and then the results of the pre-detection are used to construct the training set.
[0058] Suppose X ∈ R L×N represents the spectra of all N pixels with L bands. After pre-detection, we will obtain N outputs, and the pixels with larger values are more likely to be targets. Then, we sort these pixels in descending order, that is, X s = [x1, x2,..., x n , which satisfies Among them, represents the output of pixel x i . By setting two different thresholds T a and T b (T a > T b ), these pixels can be divided into an ordered target set X t , an ordered mixed set, and an ordered background set X b . For simplicity, T a and T b are respectively set to (1 / 4) and (1 / 5)
[0059] S12: After constructing the training set, we try to build a dual parallel network for HTD. The first network is a GAN-based network, and the other is a CNN network. Regarding the specific network structure, it should be noted that the GAN network is divided into a generator network and a discriminator network. In addition, different criteria are followed when constructing these two target detection networks.
[0060] S2: During the training process of the generative adversarial network, the target spectrum to be detected and the unlabeled spectra are used as the target and the background respectively and fed into the adversarial network for training. Although it is a bit risky to consider all unlabeled spectra as the background, simulation samples prove that the generative adversarial network can very effectively offset misassignments. When the training process converges, the discriminator of the GAN is selected as the detector in the detection stage. In this process, we obtain the initial detection result map through the discriminator based on spectral information.
[0061] Specifically, step S2 specifically includes the following steps:
[0062] S21: As Figure 2 shown, the generator adopts a residual structure, has strong expressiveness, and can well simulate the complex distributions of the target and background spectra. At the same time, to stabilize the training process, the spectral normalization layer proposed in is added between every two fully connected layers. For the discriminator, it provides two probabilities, namely the probability that the sample is a real sample and the probability that the sample is the target. Our network parameters are as Figure 3 shown.
[0063] Adopt a generative adversarial network, including a generator and a discriminator. The generator mainly includes an input layer, a convolutional layer, a spectral layer, an activation function, and an output layer. First, receive a random vector as the input, extract the features of the input image through convolutional operations, gradually increase the depth and complexity of the image, introduce a residual connection structure, which helps to alleviate the problem of gradient disappearance during the training process, improve the training speed and stability of the model, then introduce non-linearity using the activation function to help the network learn complex image features, and finally generate a high-resolution hyperspectral image, hoping that this image is as similar as possible to the real hyperspectral image.
[0064] The discriminator mainly includes an input layer, a convolutional layer, a spectral layer, an activation function, and an output layer. First, receive two types of image inputs, one is a real hyperspectral image and the other is the hyperspectral image of the generator. Extract the features of the image through convolutional operations, use pooling to reduce the size and complexity of the feature map, flatten the feature map into a one-dimensional vector for discriminating the authenticity of the input image, and then use the activation function to output a probability value between 0 and 1, representing the probability that the input image is a real hyperspectral image. Through continuous iterative training, the generative adversarial network can generate more realistic and high-quality hyperspectral images, which helps to improve the accuracy and performance of target detection.
[0065] For the loss function of the whole process, there is no difference from the traditional GAN, and the overall formula is as follows:
[0066]
[0067] Optimizing the adversarial loss V(D, G) can achieve two purposes simultaneously. The first purpose is to enable the generator G to generate real samples, and the second purpose is to enable the discriminator D to better distinguish between real samples and generated samples.
[0068] More specifically, the above adversarial loss function can be divided into the generator G loss function and the discriminator D loss function, as follows:
[0069] Generator G loss function: The loss function for training the generator is actually the term related to the noise z in the adversarial loss V(D, G), and its loss function is:
[0070]
[0071] The goal of the generator is to make the samples generated by the generator as realistic as possible, which is specifically reflected in the mathematical formula:
[0072]
[0073] where BCE() represents the binary cross-entropy loss function. From the above, it can be seen that the generator in GAN minimizes the loss function L G can be written in the form of minimizing the binary cross-entropy function.
[0074] Discriminator D loss function: The loss function for training the discriminator is actually the term related to the sample x in the adversarial loss V(D, G), and its loss function is:
[0075]
[0076] The goal of the discriminator is to enable the discriminator to better distinguish between generated samples and real samples, which is specifically reflected in the mathematical formula as (5):
[0077]
[0078]
[0079] where x represents the real data sample and x' represents the sample generated by the generator. From the above, it can be found that the discriminator in GAN maximizes the loss function L D can be written in the form of minimizing two binary cross-entropy functions.
[0080] For and To alleviate the overfitting problem caused by finite and unbalanced samples, focal loss is adopted, which is expressed as follows:
[0081]
[0082] where y cdenotes the class output of x provided by the discriminator, indicating the probability that x belongs to the target class, and y ∈ {0, 1} represents the class labels for real and generated samples. γ represents the adjustable focusing parameter and is set to 2 in our experiments.
[0083]
[0084] As shown above, L p is the absolute cosine between two generated spectra, which allows the generator to generate samples with a high level of divergence in order to alleviate the problem of mode collapse and produce better results.
[0085] By minimizing these modified objective functions, as in traditional GANs, the filter and the generator are alternately optimized. When the network converges, the neural network can provide a class label for each sample, indicating the probability that the sample belongs to the target.
[0086] S22: Different from typical CNNs, the proposed CNN is entirely constructed from CNNs without spatial pooling layers or fully connected layers, which ensures that the spatial resolution is constant in all layers. A hierarchical structure is adopted in the network, and the feature structures of different layers are cascaded for target detection. At the same time, channel pooling layers are adopted in the internal layers. To address the overfitting problem, the dropout method is adopted when training the network.
[0087] The loss of the objective function of the CNN is the same as the above formula, and the network is trained by backpropagation. As the network converges, the internal layers can extract useful features, and the last layer can detect the target based on the spatial distribution of the training samples and their spectral information.
[0088] S3: Finally, since both networks can detect the target independently, a pixel-level fusion strategy is finally adopted to fuse the results of both. In addition, to further suppress the background and small outliers, a guided filter is used to improve the smoothness and robustness of the detection results, thereby enhancing the target.
[0089] Specifically, step 4 specifically includes the following steps:
[0090] S31: Fusion of results: The dual networks discussed above are very different, so there are small differences in the results. The GAN-based network focuses on spectral information and is more sensitive to single pixels, while the CNN processes both spatial and spectral information and shows smoother results. To fuse the results, a pixel-level fusion strategy is adopted, which is expressed as follows:
[0091] D f = μD G +(1 - μ)D C
[0092] Among them, D G and D C respectively represent the detection results of the GAN-based network and the CNN. μ represents the average weight and is set to 0.5. After fusing the detection results, more distinctive and robust detection results can be obtained.
[0093] S32: Guided filter: In addition, since spatial correlation is useful for detecting certain objects in the HSI, a set of guided filters can be used to smooth the initial detection results. The initial detection maps obtained at different times are first filtered separately by the filter bank. In this case, all filters will use the initial input image as the guidance image, and they only differ from each other in the radius r and the regularization coefficient ε. Then, the initial detection maps are smoothed to different degrees and scales. After that, max pooling is performed on the filtered images to collect the most valuable information. Finally, the overall average value of the images obtained at different times is calculated to obtain relatively stable detection results.
[0094] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can still be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A hyperspectral image target detection method based on an empty-spectrum joint dual network, characterized in that Including the following steps: S1: First, perform PCA by eigenvalue decomposition of the covariance matrix of the hyperspectral image. After obtaining the dimensionality-reduced image, use the simple linear iterative clustering algorithm for superpixel segmentation to obtain the superpixel image, and then obtain the background samples through sparse representation and sample selection strategies. S2: Based on the background samples obtained above for training, feed them into the generative adversarial network to learn the background distribution. At the same time, feed the rough detection results obtained by minimizing the constrained energy of the original hyperspectral image through the classical detector into the convolutional neural network and perform spatial filtering, thereby realizing the joint spatial-spectral. Finally, adaptively fuse the results of the above two to obtain a trained model. S3: Feed the original hyperspectral image data into the trained network model, and through adaptive fusion. At the same time, in order to suppress the background, perform nonlinear suppression and guided filter to obtain the final detection result.
2. The hyperspectral image target detection method based on an empty-spectrum joint dual network according to claim 1, wherein The step S1 includes the following sub-steps: S11: To construct an object detector based on a neural network, first, a representative training set must be constructed. Since the classical detector can effectively suppress the background, the classical detector is used for pre-detection, and then the results of the pre-detection are used to construct the training set. Assume X ∈ R L×N represents the spectra of all N pixels with the L band. After pre-detection, we will obtain N outputs, where pixels with larger values are more likely to be targets. Then, we sort these pixels in descending order, i.e., X s = [x1, x2,..., x n , which satisfies where represents the output of pixel x i . By setting two different thresholds T a and T b (T a > T b ), these pixels can be divided into an ordered target set X t , an ordered mixed set, and an ordered background set X b . For simplicity, T a and T b are set to and S12: After constructing the training set, we attempt to build a dual-parallel network for HTD. The first network is a GAN-based network, and the other is a CNN network. Regarding the specific network structure, it should be noted that the GAN network is divided into a generator network and a discriminator network. In addition, we follow different criteria when constructing these two object detection networks.
3. A hyperspectral image target detection method based on an empty-spectrum joint dual network according to claim 1, characterized in that, The step S2 includes the following sub-steps: S21: The generator adopts a residual structure, which has strong expressiveness and can well simulate the complex distributions of the target and background spectra. At the same time, to stabilize the training process, the spectral normalization layer proposed in is added between every two fully connected layers. For the discriminator, it provides two probabilities, namely the probability that the sample is a real sample and the probability that the sample is a target. For the loss function of the whole process, it is no different from the traditional GAN, and the overall formula is as follows: Optimizing the adversarial loss V(D, G) can achieve two purposes at the same time. The first purpose is to make the generator G able to generate real samples, and the second purpose is to make the discriminator D better distinguish between real samples and generated samples. More specifically, the above adversarial loss function can be divided into the generator G loss function and the discriminator D loss function, as follows: Generator G loss function: The loss function for training the generator is actually the term about the noise z in the adversarial loss V(D, G), and its loss function is: The goal of the generator is to hope that the samples generated by the generator are as real as possible, which is specifically reflected in the mathematical formula: Among them, BCEO represents the binary cross-entropy loss function. It can be seen from the above that the generator in GAN minimizes the loss function L G can be written in the form of minimizing the binary cross-entropy function. Discriminator D loss function: The loss function for training the discriminator is actually the term about the sample x in the adversarial loss V(D, G), and its loss function is: The goal of the discriminator is to hope that the discriminator can better distinguish between generated samples and real samples, which is specifically reflected in the mathematical formula as (5): where \(x\) represents the real data sample and \(x'\) represents the sample generated by the generator. As can be seen from the above, the discriminator in GAN maximizes the loss function \(L\). D It can be written in the form of minimizing two binary cross - entropy functions. For and To alleviate the overfitting problem caused by limited and imbalanced samples, focal loss is adopted and expressed as follows: where y c represents the class output of x provided by the discriminator, indicating the probability that x belongs to the target class, and y ∈ {0, 1} represents the class label for real and generated samples. γ represents an adjustable focusing parameter and is set to 2 in our experiments. As shown above, L p is the absolute cosine between two generated spectra, which allows the generator to generate samples with a high level of divergence in order to alleviate the problem of mode collapse and produce better results. By minimizing these modified objective functions, as in traditional GANs, the filter and the generator are alternately optimized. When the network converges, the neural network is able to provide a class label for each sample, representing the probability that the sample belongs to the target. S22: Different from typical CNNs, the proposed CNN is completely constructed by CNNs without spatial pooling layers or fully connected layers, which ensures that the spatial resolution is constant in all layers. A hierarchical structure is adopted in the network, and the feature structures of different layers are cascaded for object detection. At the same time, channel pooling layers are adopted in the internal layers. To address the overfitting problem, the dropout method is adopted when training the network. The loss of the objective function of the CNN is the same as the above formula, and the network is trained by backpropagation. As the network converges, the internal layers can extract useful features, and the last layer can detect the target according to the spatial distribution of the training samples and their spectral information.
4. A hyperspectral image target detection method based on an empty-spectrum joint dual network according to claim 1, characterized in that, The step S3 includes the following sub-steps: S31: Fusion of results: The two networks discussed above are quite different, so there are small differences in the results. The GAN-based network focuses on spectral information and is more sensitive to single pixels, while the CNN processes both spatial and spectral information, showing smoother results. To fuse the results, a pixel-level fusion strategy is adopted, which is expressed as follows: D f = μD G + (1 - μ)D C Among them, D G and D C respectively represent the detection results of the GAN-based network and the CNN. μ represents the average weight and is set to 0.
5. After fusing the detection results, more distinctive and more robust detection results can be obtained. S32: Guided filter: In addition, since spatial correlation is useful for detecting certain objects in HSI, a group of guided filters can be used to smooth the initial detection results. The initial detection maps obtained at different times are first filtered separately by the filter group. In this case, all filters will use the initial input image as the guidance image, and they only differ from each other in the radius r and the regularization coefficient ε. Then, the initial detection maps are smoothed to different degrees and scales. After that, max pooling is performed on the filtered images to collect the most valuable information. Finally, the overall average value of the images obtained at different times is calculated to obtain a relatively stable detection result.