Hyperspectral Target Detection Method Based on Hyperspectral-Spatial Joint Sparse Guided Deep Neural Network
By combining adaptive spatial spectroscopy with sparse model and sparse representation theory with neural network, the problems of uninterpretation and parameter dependence of deep learning models in hyperspectral object detection are solved, and high-precision and high generalization object detection are achieved.
Patent Information
- Application Number
- CN202211180618.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-09-27
AI Technical Summary
The existing deep learning models have problems of uninterpretation and dependence on parameter settings in hyperspectral object detection, resulting in poor generalization performance, and model-based methods require manual adjustment of parameters, which makes the detection effect unstable.
Adaptive spatial spectroscopy combined with sparse representation theory and neural network, and adaptively weighted adjacent pixels and spatial enhancement images, use CEM detector to select samples, introduce SMOTE algorithm to expand the target samples, and combine residual graphs through background and target dictionary training networks to finally generate object detection images.
The accuracy of hyperspectral object detection and the generalization performance of network models are improved, the uninterpretation of deep neural networks and the dependence of parameter settings based on model methods are overcome, and the accuracy of object detection and background suppression capabilities are enhanced.
Smart Images

Figure CN115841606B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing technology, and in particular, to a hyperspectral target detection method based on an air-spectral joint sparse-guided deep neural network. Background Art
[0002] Hyperspectral images with rich spectral information have been widely used in fields such as biomedicine, military operations, agricultural monitoring, and geological surveys. In hyperspectral data, spectral reflectance is an inherent property of each substance, and spectral reflectance contains information such as the physical structure and chemical composition of ground objects. The rich spectral information provides unique advantages for hyperspectral target detection. For example, in military reconnaissance, camouflaged targets can be easily identified in hyperspectral images, which is difficult to achieve in RGB or multispectral images. Due to the unique advantages of hyperspectral images, hyperspectral target detection has received increasing attention and has become one of the most important applications of hyperspectral images.
[0003] Hyperspectral target detection is a technique that separates targets from the background using prior information of the targets. Spectral features are the inherent properties of each substance. By matching the pixel points in the image with the spectral curves of target substances in the spectral library, targets in the scene can be effectively identified. There are a large number of algorithms in the field of remote sensing for detecting targets of interest in hyperspectral images. The existing methods are mainly divided into two categories: model-based methods and deep learning-based methods.
[0004] Model-based methods are easy to understand and have excellent detection performance. Model-based methods include: methods based on probability statistical models, which design the target detection problem as a binary hypothesis testing problem and construct a detector through the generalized likelihood ratio test (GLRT); methods based on subspace models, which suppress background information to a specific subspace by projecting image information, thereby achieving target detection; methods based on signal processing models, which design a linear filter to constrain the energy of the background signal and highlight the target; methods based on sparse representation, which represent pixels as a linear combination of several dictionary atoms and detect targets while reducing complexity.
[0005] With the proposal of artificial neural networks, deep learning methods have been effectively applied to hyperspectral image processing, such as panchromatic sharpening, image fusion, and change detection. The exploration and utilization of deep learning have also brought many new ideas to the research of hyperspectral target detection. Convolutional neural network-based detection (CNND) uses the differences between pixel pairs to train a deep convolutional neural network to achieve supervised target detection. HTD-Net expands the training set by generating pixel pairs to overcome the difficulty that the existing spectral features are insufficient to meet the requirements of deep neural network training.
[0006] Model-based methods have strict logical reasoning. However, many parameters need to be set according to experience or through additional manual calculations. Once the parameters are not set properly, the detection effect will be very poor. Deep neural networks have strong non-linear modeling capabilities and adaptive prior fitting capabilities, and can be effectively applied in hyperspectral target detection. However, due to the lack of interpretability of existing deep learning models for target detection problems, the internal structure of the network cannot be intuitively observed, and its implementation mechanism cannot be intuitively analyzed and understood, resulting in poor generalization performance.
[0007] Aiming at the non-interpretability of deep neural networks and the need to set parameters of model-based methods according to experience or through additional calculations, the present invention combines the rigorous sparse representation theory with neural networks, and proposes an adaptive spatial-spectral joint sparse model, which improves the generalization performance of the network model and the accuracy of target detection. Summary of the Invention
[0008] The purpose of the present invention is to solve the problems existing in the prior art, and proposes a hyperspectral target detection method based on an empty-spectrum joint sparse-guided deep neural network.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] A hyperspectral target detection method based on an empty-spectrum joint sparse-guided deep neural network, the specific steps are as follows:
[0011] S101: Input the hyperspectral image to be subjected to target detection, and perform adaptive spatial-spectral joint on it to obtain an enhanced hyperspectral image
[0012] S102: Use classical methods to generate background samples and an initial background dictionary Target samples and a target dictionary
[0013] S103: Use the background samples and the initial background dictionary to train the background reconstruction network, and continuously update the background dictionary during the training process;
[0014] S104: Use the target samples and the target dictionary to train the target reconstruction network;
[0015] S105: Use the trained background reconstruction network and the final background dictionary to reconstruct the hyperspectral image to be detected, and obtain a background reconstruction residual map;
[0016] S106: Use the trained target reconstruction network and the target dictionary to reconstruct the hyperspectral image to be detected, and obtain a target reconstruction residual map;
[0017] S107: Combine the two residual maps to obtain the final target detection map.
[0018] As a further technical solution of the present invention, in the step S101, an enhanced hyperspectral image is obtained by adaptively combining the input hyperspectral image in space and spectrum Specifically:
[0019] An adaptive spatial-spectral joint sparse model is established, and weights are calculated according to the similarity measure between adjacent pixels and the central pixel to adaptively weight the adjacent pixels; the spatially enhanced pixels Are obtained by the following formula:
[0020]
[0021] Where h i Represents the i-th pixel adjacent to h, n is the number of adjacent pixels, ω i Is the weight corresponding to h i And the weight is adaptively adjusted for each adjacent pixel;
[0022]
[0023] Where σ is the bandwidth of the Gaussian kernel; the more similar h i And h are, the larger ω i Is; the spatially enhanced hyperspectral image is obtained by adaptively weighting the neighborhood spatial information of each pixel in H
[0024] As a further technical solution of the present invention, in the step S102, background samples, an initial background dictionary Target samples and a target dictionary Specifically:
[0025] In the training of the neural network, training samples are used for the neural network model fitting, and a sample selection module based on CEM is introduced for sample selection; the CEM detector is used for image pre-detection, and its output is as follows:
[0026]
[0027] Where Represents the correlation matrix, Represents the spatially enhanced hyperspectral image, ω represents the vector of the FIR linear filter, d is the given target spectrum; N pixels with small output values are selected b As the background training samples; the selected background samples are used for the training of the background reconstruction network; the initial background dictionary Is obtained by clustering the background samples;
[0028] After the image is pre-detected by the CEM detector, several pixels with larger output values are selected as target samples; an algorithm for synthesizing new samples from minority class samples is introduced to generate features for target detection by synthesizing new samples; the nearest neighbor sample of a target sample t is found, and then a new target sample t is synthesized using the following formula new :
[0029] t new = t + rand(0, 1) × (t - t n )
[0030] where t n is randomly selected from the samples with the closest Euclidean distance to t; by performing such an operation on each target sample, the number of target samples can be expanded to the required number; the expanded target samples are used for the training of the target reconstruction network; the target dictionary is obtained by clustering the target samples.
[0031] As a further technical solution of the present invention, the background reconstruction network in step S103 can reconstruct the background spectral vector pixel by pixel according to the dictionary, and the structure of the background reconstruction network is as follows:
[0032] The reconstruction process of the background is driven by the model, and the solution process is iteratively optimized by the following three equations; z is generated by (k+1) , and then (k+1) and are used to generate Then is used to generate Iterate until convergence;
[0033]
[0034]
[0035]
[0036] The background reconstruction network consists of k stages, where the input of each stage is all background samples and the output of the previous stage; each stage consists of a deep convolutional denoising network, a pixel vector reconstruction module, and a dictionary update module; the structure of each module corresponds to the solution process of the equation;
[0037] The formula for generating z by (k+1) can be equivalent to The denoising process; using a deep convolutional denoising network to simulate the formula; selecting U-Net as the backbone network of DCDN, which is a symmetric structure composed of an encoder and a decoder, and contains a total of seven convolutional layers and two skip connections; after the encoder, the decoder gradually repairs the details of the sparse vector, and two skip connections are adopted; the information is directly transmitted from the encoder to the corresponding decoder for feature fusion;
[0038] In the target sparse representation, the process of generating by z (k+1) and is completed through the pixel vector reconstruction module; μ is set as a learnable parameter that can be fitted by the network; One-dimensional convolutional kernels are used for learning and fitting; the input of PVRM includes the output of DCDN and the background dictionary In hyperspectral target detection, the background dictionary should contain background information of all categories; when directly selecting background samples for dictionary filling, construct
[0039] A dictionary update strategy is adopted to obtain the background dictionary; dictionary update is to update according to the formula of generating which ensures the optimal background dictionary; the structure of the dictionary update module is transformed into the form of a neural network according to the formula of generating for operation and solution. generating As a further technical solution of the present invention, the target reconstruction network can reconstruct the target spectral vector pixel by pixel according to the dictionary, and the target reconstruction network is as follows;
[0040] In the target reconstruction network, the objective function is solved through the iterative optimization of two sub-models (the following formula), without updating the target dictionary
[0041]
[0042]
[0043]
[0044] The target reconstruction network is designed according to the iterative solution process of the equation; the target reconstruction network performs k iterations through k epochs, and the network architectures of DCDN and PVRM of the target reconstruction network and the background reconstruction network are the same.
[0045] As a further technical solution of the present invention, the training method of the background reconstruction network and the target reconstruction network;
[0046] The background samples, initial background dictionary, target samples, and target dictionary obtained from the hyperspectral image are respectively fed into the target reconstruction network and the background reconstruction network for network training, and finally, the trained target reconstruction network, background reconstruction network, and the updated dictionary are obtained.
[0047] As a further technical solution of the present invention, the target reconstruction network and the background reconstruction network respectively reconstruct the hyperspectral image to obtain the final target detection map.
[0048] The hyperspectral image is input pixel by pixel into the background reconstruction network and the target reconstruction network, and the reconstruction of the target network and the background network is carried out simultaneously to obtain the target reconstruction residual and the background reconstruction residual respectively. The final target detection map can be obtained according to these two residuals.
[0049]
[0050] Where is the background residual, is the target residual, is the final target detection map.
[0051] Advantages of the present invention:
[0052] 1. The present invention uses the sample selection module of CEM for the selection of target samples and background samples. The image is pre-detected using the CEM detector to select samples, and high-confidence training samples can be obtained, which is beneficial to the fitting of the model.
[0053] 2. The present invention adopts the SMOTE algorithm to expand the target samples. This method can generate more features beneficial to target detection by synthesizing new samples, enhancing the accuracy of target detection.
[0054] 3. The present invention uses the background dictionary update strategy, which can overcome the disadvantages that in hyperspectral target detection, the background objects are diverse and scattered, and the initial background dictionary cannot contain all categories of background information. Using this strategy can construct an accurate background dictionary and enhance the accuracy of target detection.
[0055] 4. The present invention combines the rigorous sparse representation theory with the neural network, which can solve the problem of the non-interpretability of the deep neural network and the problem that the parameters of the model-based method need to be manually set according to experience or additional calculations, improving the generalization performance of the network model and the accuracy of target detection. Description of the Drawings
[0056] Figure 1 It is a flowchart of the hyperspectral image target detection method provided by the embodiment of the present invention.
[0057] Figure 2It is a schematic diagram of the background reconstruction network structure provided by an embodiment of the present invention.
[0058] Figure 3 It is a schematic diagram of the target reconstruction network structure provided by an embodiment of the present invention.
[0059] Figure 4 It is a schematic diagram of the deep convolutional denoising network structure provided by an embodiment of the present invention.
[0060] Figure 5 It is a schematic diagram of the pixel vector reconstruction module structure provided by an embodiment of the present invention.
[0061] Figure 6 It is a schematic diagram of the dictionary update module structure provided by an embodiment of the present invention.
[0062] Figure 7 It is a comparison of the target detection result diagrams provided by an embodiment of the present invention.
[0063] Figure 7 Among them: (a)-(j) are the image to be detected, the reference detection result, the CEM detection result, the ACE detection result, the MF detection result, the hCEM detection result, the CSCR detection result, the CNND detection result, the HTD-Net detection result, and the ASJ-DNN detection result in sequence. Detailed implementation manners
[0064] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features and their effects of the present invention as follows.
[0065] Refer to Figures 1 - 7 As Figure 1 shown, the hyperspectral target detection method based on the spatial-spectral joint sparse-guided deep neural network provided by the present invention includes the following steps
[0066] S101: Input the hyperspectral image to be subjected to target detection.
[0067] S102: Perform adaptive spatial-spectral joint on the input hyperspectral image to obtain an enhanced hyperspectral image
[0068] S103: Use classical methods to generate background samples and an initial background dictionary target samples and a target dictionary
[0069] S104: Train the background reconstruction network using the background samples and the initial background dictionary, and continuously update the background dictionary during the training process.
[0070] S105: Train a target reconstruction network using the target samples and the target dictionary.
[0071] S106: Use the trained background reconstruction network and the final background dictionary to reconstruct the hyperspectral image to be detected, and obtain a background reconstruction residual map.
[0072] S107: Use the trained target reconstruction network and the target dictionary to reconstruct the hyperspectral image to be detected, and obtain a target reconstruction residual map.
[0073] S108: Combine the two residual maps to obtain a final target detection map.
[0074] As Figure 1 shown, the hyperspectral target detection method based on the spatial-spectral joint sparse-guided deep neural network provided by the present invention is implemented as follows:
[0075] (1) Perform adaptive spatial-spectral joint on the input hyperspectral image to obtain an enhanced hyperspectral image
[0076] In order to be able to adaptively adjust the weights of each adjacent pixel, the present invention proposes an Adaptive Spatial-Spectral Joint Sparsity Model (AS 2 JSM), which adaptively weights adjacent pixels by calculating weights based on the similarity measure (Euclidean distance) between adjacent pixels and the central pixel. The spatially enhanced pixel is obtained by the following formula:
[0077]
[0078] where h i represents the i-th pixel adjacent to h, n is the number of adjacent pixels, and ω i is the weight corresponding to h i , and this weight is adaptively adjusted for each adjacent pixel.
[0079]
[0080] where σ is the bandwidth of the Gaussian kernel. The more similar h i and h are, the more likely they are to belong to the same object, and the larger ω i is. The spatially enhanced hyperspectral image is obtained by adaptively weighting the neighborhood spatial information of each pixel in H In this way, each pixel vector in not only contains spectral features, but also adaptively introduces the spatial information of the neighborhood, so as to make full and reasonable use of spatial-spectral information.
[0081] (2) Obtain background samples and an initial background dictionary from spatially enhanced hyperspectral images Target samples and a target dictionary
[0082] (2a) During the training of a neural network, training samples are used for model fitting. Therefore, it is very important to select samples with high confidence. Thus, a sample selection module based on CEM is introduced for sample selection. The CEM detector is used for image pre-detection, and its output is as follows:
[0083]
[0084] Where represents the correlation matrix, represents the spatially enhanced hyperspectral image, ω represents the FIR linear filter algorithm vector, and d is the given target spectrum. The larger the output value corresponding to each pixel point after CEM detection, the greater the possibility that the pixel belongs to the target. Therefore, N b pixels with small output values are selected as background training samples with high confidence. The selected background samples are used for the training of the background reconstruction network. The initial background dictionary is obtained by clustering the background samples.
[0085] (2b) After the image is pre-detected by the CEM detector, several pixel points with larger output values are selected as target samples. For hyperspectral target detection, the number of target samples is usually much less than that of background samples. The present invention introduces the Synthetic Minority Oversampling Technique (SMOTE), which can generate more features beneficial to target detection by synthesizing new samples. Specifically, find the nearest neighbor sample of a target sample t, and then synthesize a new target sample t using the following formula new :
[0086] t new = t + rand(0,1)×(t - t n )
[0087] where t n is randomly selected from the samples with the closest Euclidean distance to t. By performing such an operation on each target sample, the number of target samples can be expanded to the required number. The expanded target samples are used for the training of the target reconstruction network. The target dictionary is obtained by clustering the target samples.
[0088] (3) The background reconstruction network can reconstruct the background spectral vector pixel by pixel according to the dictionary. The structure of the background reconstruction network is as Figure 2 shown.
[0089] (3a) The background reconstruction process is driven by a model, and the solution process is iteratively optimized by the following three equations. From Generate z (k+1) , and then from z (k+1) and Generate Then from Generate Iterate until convergence.
[0090]
[0091]
[0092]
[0093] The background reconstruction network consists of k stages, where the input of each stage is all background samples and the output of the previous stage. Each stage consists of a Deep Convolution Denoising Network (DCDN), a Pixel Vector Reconstruction Module (PVRM), and a Dictionary Update Module (DUM). As Figure 4 , Figure 5 , Figure 6 shown. The structure of each module corresponds to the solution process of the equation.
[0094] (3b) The formula for generating z from (k+1) can be equivalent to the denoising process of . Therefore, a deep convolution denoising network is used to simulate the formula. In the hyperspectral image denoising task, classical methods have high generation ability and low sample dependence. However, performing high-quality denoising requires complex numerical iterations and accurate model construction. Compared with classical methods, deep learning-based methods have better generalization and higher efficiency. Therefore, the present invention selects U-Net as the backbone network of DCDN because it has excellent performance in pixel-level restoration. Similar to U-Net, DCDN is a symmetric structure composed of an encoder and a decoder, and contains a total of seven convolutional layers and two skip connections. After the encoder, the decoder gradually repairs the details of the sparse vector. Since the convolution operation will cause to lose some details, two skip connections are adopted. They directly transfer information from the encoder to the corresponding decoder for feature fusion and provide more detailed features for information restoration.
[0095] (3c) In the target sparse representation, from z(k+1) and generate The formula for is the reconstruction problem, which can be solved by the designed pixel vector reconstruction module. The penalty parameter μ that controls the similarity between the auxiliary variable and the original variable should theoretically be large enough. However, this may lead to a very slow convergence rate of the algorithm. Manual adjustment is time-consuming and laborious, and it may not necessarily find the optimal value. Therefore, in the present invention, μ is set as a learnable parameter that can be fitted by the network. Additionally, has a very high computational complexity because it involves the inversion process of high-dimensional data. Therefore, the present invention uses a one-dimensional convolution kernel to learn and fit its value. In addition, the input of the PVRM includes the output of the DCDN and the background dictionary
[0096] In hyperspectral target detection, background objects are diverse and scattered. The background dictionary should contain background information of all categories. When directly selecting background samples for dictionary filling, some background features of certain categories may be omitted. In order to be able to construct an accurate Therefore, the present invention adopts a dictionary update strategy to obtain a more accurate background dictionary. Dictionary update is based on generate formula to update to ensure the optimal background dictionary. The structure of the dictionary update module is transformed into the form of a neural network according to generate formula for operation and solution.
[0097] (4) The target reconstruction network can reconstruct the target spectral vector pixel by pixel according to the dictionary. The structure of the target reconstruction network is as Figure 3 shown.
[0098] (4a) In the target reconstruction network, the objective function is solved by the iterative optimization of two sub-models (the following formula), without updating the target dictionary
[0099]
[0100]
[0101] The target reconstruction network is designed according to the solution iterative process of the equation. The target reconstruction network performs k iterations through k epochs. The following takes the (k + 1)-th iteration as an example for detailed description. generate υ (k+1) is denoising process, which is simulated by using DCDN in the same way as the background reconstruction network. From υ (k+1) and generate It is reconstruction, and the reconstruction can be carried out through PVRM. The network architectures of DCDN and PVRM in the target reconstruction network and the background reconstruction network are the same. In addition, the most obvious difference between the target reconstruction network and the background reconstruction network is that DUM is not required in the target reconstruction network.
[0102] (5) Training methods for the background reconstruction network and the target reconstruction network.
[0103] The background samples, initial background dictionary, target samples, and target dictionary obtained from the hyperspectral image are respectively fed into the target reconstruction network and the background reconstruction network for network training, and finally the trained target reconstruction network, background reconstruction network, and the updated dictionary are obtained.
[0104] (6) The target and background reconstruction networks respectively reconstruct the hyperspectral image to obtain the final target detection map.
[0105] The hyperspectral image is input pixel by pixel into the background reconstruction network and the target reconstruction network, and the reconstruction of the target network and the background network is carried out simultaneously to obtain the target reconstruction residual and the background reconstruction residual respectively.
[0106] The final target detection map can be obtained based on these two residuals.
[0107]
[0108] Among them is the background residual, is the target residual, is the final target detection map.
[0109] The technical effects of the present invention will be described in detail below in combination with simulation experiments:
[0110] 1. Datasets and simulation experiment conditions:
[0111] This experiment uses the San Diego dataset, which is obtained by the Airborne Visible / Infrared Imaging Spectrometer (AVIRIS) sensor over San Diego, California, USA. The original image contains 224 spectral bands, covering a wavelength range of 0.4 - 2.5 μm. After removing the noise bands, 189 available bands are retained. The entire image contains 400×400 pixels, and sub-images of size 100×100 are cropped from it in this experiment. Three airports in the image are regarded as the detection targets of interest.
[0112] Table 1 Experimental basic environment and configuration
[0113]
[0114] 2. Evaluation criteria
[0115] In addition to qualitative observation of subjective results, quantitative evaluation should also be carried out to measure the detection performance of target detection. In this paper, the area under the curve (AUC) of the 2D receiver operating characteristic curve (ROC) generated in three two-dimensional spaces (i.e., (PD, PF), (PD, τ), and (PF, τ)) (denoted as AUC DF , AUC Dτ , AUC Fτ ) is used to quantify the detection results. These three AUC values can evaluate the effectiveness of the target detection method, the probability of target detection, and the background suppression ability, respectively. The better the performance of the target detection method, the higher the AUC DF and AUC Dτ , while the lower the AUC Fτ . Therefore, an overall detection AUC value, denoted as AUC OD , can be designed to comprehensively evaluate the detection performance. The higher the AUC OD , the better the detection performance.
[0116] AUC OD = AUC DF + AUC Dτ - AUC Fτ
[0117] 3. Test Results
[0118] For the San Diego dataset, the detection graphs of different target detection methods are as shown in Figure 7 . (a)-(j) are the images to be detected, reference detection results, CEM detection results, ACE detection results, MF detection results, hCEM detection results, CSCR detection results, CNND detection results, HTD-Net detection results, and ASJ-DNN detection results in sequence. It can be seen from the figure that CSCR and this method have the best effects, and they can detect all target pixels with a high probability; while CNND performs the worst, and it is difficult to distinguish between targets and backgrounds. Compared with other detection methods, hCEM performs outstandingly in background suppression. CEM, ACE, and MF all have varying degrees of omission of target pixels. It is difficult to intuitively and accurately compare the detection accuracies of detection methods (such as CSCR and AJM-DNN). Under this test, this target detection method obtained the largest AUC DF and AUC OD , and achieved satisfactory detection performance.
[0119] Table 2 Detection Results
[0120]
[0121] In summary, the present invention realizes a deep learning network driven by a sparse representation model for hyperspectral target detection. Since each node corresponds to an implementation operator in the sparse representation model, the network is interpretable while completing the end-to-end detection task.
[0122] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments with equivalent changes within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A hyperspectral target detection method based on an empty-spectrum joint sparse guided deep neural network, characterized in that, The specific steps are as follows: S101: Input the hyperspectral image to be subjected to target detection, and perform adaptive spatial-spectral joint processing on it to obtain an enhanced hyperspectral image S102: Generate a background sample and an initial background dictionary respectively using a classical method a target sample and a target dictionary Specifically: Introduce a sample selection module based on CEM for sample selection; the CEM detector is used for image pre-detection, and its output is as follows: Among them represents the correlation matrix represents the spatially enhanced hyperspectral image, ω represents the vector of the FIR linear filter, and d is the given target spectrum; select N b pixels with small output values as background training samples The selected background samples are used for the training of the background reconstruction network; the initial background dictionary is obtained by clustering the background samples; After the image is pre-detected by the CEM detector, several pixels with larger output values are selected as target samples; the minority class sample synthesis algorithm is introduced to generate features for target detection by synthesizing new samples; find the nearest neighbor sample of a target sample t, and then synthesize a new target sample t using the following formula new :[[]]END]] t new = t + rand(0,1)×(t - t n ) where t n is randomly selected from the samples with the closest Euclidean distance to t; by performing such an operation on each target sample, the number of target samples can be expanded to the required number; the expanded target samples are used for the training of the target reconstruction network; the target dictionary is obtained by clustering the target samples; S103: Train the background reconstruction network using background samples and the initial background dictionary, and continuously update the background dictionary during the training process; S104: Train the target reconstruction network using target samples and the target dictionary; S105: Use the trained background reconstruction network and the final background dictionary to reconstruct the hyperspectral image to be detected, and obtain the background reconstruction residual map; S106: Use the trained target reconstruction network and the target dictionary to reconstruct the hyperspectral image to be detected, and obtain the target reconstruction residual map; S107: Combine the two residual maps to obtain the final target detection map.
2. The hyperspectral target detection method based on the joint sparse-guided deep neural network of spatial and spectral domains according to claim 1, wherein In the step S101, the input hyperspectral image is adaptively processed in a spatial-spectral joint manner to obtain an enhanced hyperspectral image Specifically: An adaptive spatio-spectral joint sparse model is established, and weights are calculated based on the similarity measure between adjacent pixels and the central pixel to adaptively weight the adjacent pixels; the spatially enhanced pixels are obtained by the following formula: where h i represents the i-th pixel adjacent to h, n is the number of adjacent pixels, and ω i is the weight corresponding to h i and the weight is adaptively adjusted for each adjacent pixel; where σ is the bandwidth of the Gaussian kernel; h i The more similar h is, the larger ω i is; the spatially enhanced hyperspectral image is obtained by adaptively weighting the neighborhood spatial information of each pixel in H 3. The hyperspectral target detection method based on the joint sparse guidance deep neural network of spatial and spectral domains according to claim 1, wherein, In the step S103, the background reconstruction network can reconstruct the background spectral vector pixel by pixel according to the dictionary. The structure of the background reconstruction network is as follows: The background reconstruction process is driven by a model, and the solution process is iteratively optimized by the following three equations; from Generate z (k+1) , and then from z (k+1) and Generate Then from Generate Iterate until convergence; The background reconstruction network consists of k stages. The input of each stage is all background samples and the output of the previous stage; each stage consists of a deep convolutional denoising network, a pixel vector reconstruction module, and a dictionary update module; the structure of each module corresponds to the solution process of the equation; Generated by to generate z (k+1) The formula can be equivalent to The denoising process; using a deep convolutional denoising network to simulate the formula; selecting U-Net as the backbone network of the DCDN, the DCDN is a symmetric structure composed of an encoder and a decoder, and contains a total of seven convolutional layers and two skip connections; after the encoder, the decoder gradually repairs the details of the sparse vector, and two skip connections are adopted; the information is directly transmitted from the encoder to the corresponding decoder for feature fusion; In the target sparse representation, by z (k+1) and generate The process is completed by the pixel vector reconstruction module; μ is set as a learnable parameter that can be fitted through the network; Use a one-dimensional convolutional kernel to learn and fit; The input of PVRM includes the output of DCDN and the background dictionary In hyperspectral target detection, the background dictionary should contain background information of all categories; when directly selecting background samples for dictionary filling, construct Adopt a dictionary update strategy to obtain the background dictionary; dictionary update is based on Generate Update according to the formula of to ensure the optimal background dictionary; the structure of the dictionary update module is transformed into the form of a neural network according to Generate The formula of to perform operation and solution.
4. The hyperspectral target detection method based on the joint sparse-guided deep neural network of spatial and spectral domains as claimed in claim 1, wherein The target reconstruction network can reconstruct the target spectral vector pixel by pixel according to the dictionary. The target reconstruction network is as follows: In the target reconstruction network, the objective function is solved by iterative optimization of two sub-models (as shown in the following formula), without updating the target dictionary The target reconstruction network is designed according to the iterative process of solving the equation; the target reconstruction network performs k iterations through k epochs. The network architectures of the DCDN and PVRM of the target reconstruction network and the background reconstruction network are the same.
5. The hyperspectral target detection method based on the joint sparse guidance deep neural network of spatial and spectral domains according to claim 1; characterized in that, The training methods of the background reconstruction network and the target reconstruction network; The background samples, initial background dictionary, target samples, and target dictionary obtained from hyperspectral images are respectively fed into the target reconstruction network and the background reconstruction network for network training, and finally, the trained target reconstruction network, background reconstruction network, and the updated dictionary are obtained.
6. The hyperspectral target detection method based on the joint sparse guidance deep neural network of spatial and spectral domains as claimed in claim 1, wherein The target reconstruction network and the background reconstruction network respectively reconstruct the hyperspectral image to obtain the final target detection map; Input the hyperspectral image pixel by pixel into the background reconstruction network and the target reconstruction network, and simultaneously perform the reconstruction of the target network and the background network to obtain the target reconstruction residual and the background reconstruction residual respectively; the final target detection map can be obtained based on these two residuals; wherein is the background residual, is the target residual, is the final object detection map.
Citation Information
Patent Citations
Hyperspectral abnormal object detection method based on structure sparse representation and internal cluster filtering
CN106919952A
Method and device for building neural network
WO2018227801A1