Hyperspectral image specific target detection method and system based on pixel weighted background reconstruction network
The network is constructed through cell-weighted background, and the dual-window blind spot structure and surrounding pixel weighting strategy are used to enhance the background suppression ability, solving the problem of insufficient background suppression in hyperspectral image detection and improving the accuracy and stability of the detection.
Patent Information
- Application Number
- CN202510333436.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-08
AI Technical Summary
The existing hyperspectral image specific object detection methods have insufficient background suppression capabilities, resulting in the impact of detection results.
The background reconstruction network is adopted based on cell weighting, and the background reconstruction network is built through the convolutional network. The dual-window blind spot structure and surrounding pixel weighting strategy are used to generate supervised labels, which are fused with pre-detection results to enhance the background suppression ability.
The problem of sparse target prior information was effectively solved, the background suppression ability was improved, and the accuracy and stability of detection were improved. The experimental results reached 0.9988, 0.9997, 0.9985 and 0.9999 on multiple data sets.
Smart Images

Figure CN120279409A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection based on hyperspectral images, and particularly relates to a method and system for detecting specific targets in hyperspectral images. Background Technique
[0002] Hyperspectral image is an imaging technology that uses multiple continuous narrow bands to obtain the spectral information of an object, which can provide rich spectral data. It usually contains dozens to hundreds of spectral bands, covering the range from visible light to near-infrared. By analyzing the spectral characteristics of each pixel, this technology is widely used in fields such as agricultural monitoring, environmental assessment, mineral exploration, and medical imaging. Hyperspectral target detection is an important hyperspectral analysis task, which focuses on detecting whether a pixel is a target of interest based on known prior information.
[0003] The rich spectral information in hyperspectral images can effectively distinguish materials with similar appearances but different spectral characteristics. At present, specific target detection methods for hyperspectral images have been developed. However, these methods still have deficiencies in using the background information of hyperspectral images, resulting in insufficient background suppression and affected detection results. Summary of the Invention
[0004] The present invention aims to solve the problems of scarce target prior information and poor background suppression ability of conventional methods in the field of hyperspectral target detection.
[0005] A method for detecting specific targets in hyperspectral images based on a background reconstruction network with pixel weighting includes the following steps:
[0006] For the area to be detected, obtain the original hyperspectral image X, and use the background reconstruction network based on a convolutional network to reconstruct the entire image to obtain a reconstructed image:
[0007]
[0008] Wherein, is the reconstructed image obtained through the network, are the network parameters;
[0009] The reconstructed image is obtained by the following formula of the pixel and the pixel x of the original hyperspectral image X in the corresponding area v The reconstruction difference
[0010]
[0011] Wherein, v = 1, 2,..., N represents each pixel of the image X, and N is the total number of pixels in the hyperspectral image;
[0012] Simultaneously obtain the pre-detection result of the original hyperspectral image X through pre-detection, and the pre-detection result y v Use the following background suppression function:
[0013]
[0014] Then perform the following fusion:
[0015]
[0016] where q v is the finally obtained detection result vector;
[0017] Combine the N result vectors q v to obtain the final detection image.
[0018] Furthermore, the process of obtaining the pre-detection result from the original hyperspectral image X through pre-detection includes:
[0019] The original hyperspectral image is represented, where H, W, and L represent the number of rows, columns, and channels of the image respectively; the hyperspectral image is two-dimensionalized and normalized to obtain where N = H × W is the total number of pixels in the hyperspectral image;
[0020] According to the data r, obtain the autocorrelation matrix R of the image:
[0021]
[0022] where r i is the pixel of the data r, and i = 1, 2,..., N;
[0023] According to the known prior spectral information of the target pixel perform pre-detection on the original hyperspectral image:
[0024]
[0025] In the formula, is the filter, and y is the detection result map of the pre-detection.
[0026] Further, the background reconstruction network built based on the convolutional network includes a convolutional unit and a residual connection module network; the convolutional unit is arranged before the residual connection module network, the convolutional unit includes two convolutional layers, a rectified linear unit is connected after the first convolutional layer, and a batch normalization layer is connected after the second convolutional layer; the residual connection module network includes a plurality of residual connection modules, a convolutional layer arranged after the residual connection module, a batch normalization layer, and a convolutional layer arranged after the batch normalization layer; after the input and output of each residual connection module are element-wise summed, the sum is used as the input of the next residual connection module; the output of the normalization layer after a plurality of residual connection modules is element-wise summed with the input of the first residual connection module, and finally fed into the last separate convolutional layer as the output of the residual connection module network.
[0027] Preferably, the residual connection module network is provided with three residual connection modules.
[0028] Further, the background reconstruction network built based on the convolutional network is pre-trained, and the specific training process includes:
[0029] S1. The pre-detection strategy generation network learns samples:
[0030] First, use the pre-detection algorithm to pre-detect the original hyperspectral image to obtain the pre-detection result map y, and then process it using the following strategy:
[0031]
[0032] Among them, y(i,j) is the result of the i-th row and j-th column in the pre-detection result map, α is the threshold; p tg represents the target pixel of the pre-detection result in the image, and p bg represents the background pixel of the pre-detection result;
[0033] S2. Double-window blind block processing:
[0034] Divide the original hyperspectral image X into small blocks with each pixel as the central pixel, the width and height are both Wo, denoted as First, use a double-window structure on the image block to generate a blind spot area, where the outer window size is the same as the image block size, which is W out ×W out ; the inner window size is W in ×W in , denoted by , and the blind spot area generation process is as follows:
[0035] Based on the target sample p tg and background sample p bg constructed by pre-detection, randomly select W in ×W in the area between the inner window and the outer window.in Replace each pixel within the inner window with a background pixel; Denote the image block that has undergone double-window blind blocking as Denote the inner window of the image block that has undergone double-window blind blocking as
[0036] S3. Adopt the surrounding pixel replacement strategy to generate a supervised label for reconstructing the central pixel:
[0037] Take out the inner window P of the image block that has undergone double-window blind blocking ib and define the pixels except the central pixel as surrounding pixels. Denote the set of these surrounding pixels as P s For each surrounding pixel p s in P and the prior spectral information d of the target pixel, calculate the spectral angle difference L si' through the following formula: i :
[0038]
[0039] where i' = 1, 2,..., n represents the i'-th pixel within the inner window. The inner window has n pixels, and n = W in ×W in -1 is the number of surrounding pixels;
[0040] After obtaining the spectral differences between each surrounding pixel and the prior target pixel, construct the weighted pixel p w as the label of the central pixel p c of the image block through the following strategy:
[0041]
[0042] where j′ = 1, 2,..., n, j' represents the pixel within the inner window, n = W in ×W in -1 is the number of surrounding pixels, and β is the weight;
[0043] S4. The background reconstruction network performs background reconstruction on the data:
[0044] For each input image block P ob , denote the proposed background reconstruction network model as N br , and the network during the training process can be represented by the following formula:
[0045]
[0046] where p br represents the central pixel of P reconstructed by the network; ob and the pixel p
[0047] obtained by weighting the surrounding pixels within the inner window of the image blockw As the supervision information of p br , the training objective is to make p br converge to p w ; the specific loss objective function is as follows:
[0048]
[0049] Among them, represents the loss between the reconstructed p br and p w at the center pixel of the image patch. The l1 norm based on the mean absolute error is used in the loss to measure the error to a lesser extent;
[0050] The trained background reconstruction network is obtained based on the network gradient descent method.
[0051] A hyperspectral image specific target detection system based on a pixel-weighted background reconstruction network, including:
[0052] Hyperspectral image acquisition unit: For the area to be detected, the original hyperspectral image X is acquired;
[0053] Image reconstruction and reconstruction difference calculation unit: The original hyperspectral image X is reconstructed using the background reconstruction network built based on the convolutional network to obtain the reconstructed image From the reconstructed image pixels and the pixels x of the original hyperspectral image X in the corresponding area v calculate the reconstruction difference
[0054] Reconstructed image Among them, is the reconstructed image obtained through the network, is the network parameter;
[0055] Reconstruction difference Among them, v = 1, 2,..., N represents each pixel of the image X, and N is the total number of pixels in the hyperspectral image;
[0056] Background suppression unit: Perform background suppression on the pre-detection result y obtained by pre-detecting the original hyperspectral image X v . When performing background suppression on the pre-detection result y v , the following background suppression function is used:
[0057]
[0058] Fusion detection unit: Based on B(y v ) and perform fusion to obtain the detection result vector Combine N result vectors q v to obtain the final detected image.
[0059] Furthermore, the process of obtaining the pre-detection result from the original hyperspectral image X through pre-detection includes:
[0060] The original hyperspectral image is denoted as, where H, W, and L represent the number of rows, columns, and channels of the image respectively; the hyperspectral image is two-dimensionalized and normalized to obtain where N = H×W is the total number of pixels in the hyperspectral image;
[0061] According to the data r, obtain the autocorrelation matrix R of the image:
[0062]
[0063] where, r i is the pixel of the data r, and i = 1, 2,..., N;
[0064] According to the known prior spectral information of the target pixel perform pre-detection on the original hyperspectral image:
[0065]
[0066] In the formula, is the filter, and y is the detection result map of the pre-detection.
[0067] Furthermore, the background reconstruction network built based on the convolutional network includes a convolutional unit and a residual connection module network; the convolutional unit is set before the residual connection module network, and the convolutional unit includes two convolutional layers. The first convolutional layer is followed by a rectified linear unit, and the second convolutional layer is followed by a batch normalization layer; the residual connection module network includes multiple residual connection modules, a convolutional layer set after the residual connection module, a batch normalization layer, and a convolutional layer set after the batch normalization layer; after the element-wise summation of the input and output of each residual connection module, it is used as the input of the next residual connection module; the output of the normalization layer after multiple residual connection modules is element-wise summed with the input of the first residual connection module and finally sent into the last separate convolutional layer as the output of the residual connection module network.
[0068] Preferably, the residual connection module network is provided with three residual connection modules.
[0069] Furthermore, the background reconstruction network built based on the convolutional network is pre-trained, and the specific training process includes:
[0070] S1. The pre-detection strategy generation network learns samples:
[0071] First, use a pre-detection algorithm to pre-detect the original hyperspectral image to obtain the pre-detection result map y, and then process it using the following strategy:
[0072]
[0073] where y(i,j) is the result of the i-th row and j-th column in the pre-detection result map, and α is the threshold; p tg represents the target pixel of the pre-detection result in the image, and p bg represents the background pixel of the pre-detection result;
[0074] S2. Double-window blind block processing:
[0075] Divide the original hyperspectral image X into small blocks with each pixel as the central pixel, where the width and height are both Wo, denoted as First, use a double-window structure on the image block to generate a blind spot area. The size of the outer window is the same as that of the image block, which is W out ×W out ; the size of the inner window is W in ×W in , denoted by . The process of generating the blind spot area is as follows:
[0076] Based on the target sample p tg constructed by pre-detection and the background sample p bg , randomly select W in ×W in background pixels in the area between the inner window and the outer window to replace each pixel in the inner window; Denote the image block that has undergone double-window blind block as Denote the inner window of the image block that has undergone double-window blind block as
[0077] S3. Adopt the surrounding pixel replacement strategy to generate a supervised label for reconstructing the central pixel:
[0078] Take out the inner window P ib of the image block that has undergone double-window blind block, and define the pixels except the central pixel as surrounding pixels. The set of these surrounding pixels is denoted by P s ; For each surrounding pixel p s in P si' and the prior spectral information d of the target pixel, calculate the spectral angle difference L i through the following formula:
[0079]
[0080] where i' = 1, 2,..., n represents the i'-th pixel in the inner window, the inner window has n pixels, and n = W in ×W in-1 is the number of surrounding pixels;
[0081] After obtaining the spectral differences between each surrounding pixel and the prior target pixel, a weighted pixel p is constructed through the following strategy w As the central pixel p of the image patch c label:
[0082]
[0083] where j′ = 1, 2, ..., n, j' represents the pixels within the inner window, and n = W in ×W in -1 is the number of surrounding pixels, and β is the weight;
[0084] S4. The background reconstruction network reconstructs the background for the data:
[0085] For each input image patch P ob , denote the proposed background reconstruction network model as N br , and the network during the training process can be represented by the following formula:
[0086]
[0087] where p br represents the central pixel of P reconstructed by the network ob ;
[0088] The pixel p obtained by weighting the surrounding pixels within the inner window of the image patch w is used as the supervision information for p br , and the training objective is to make p br converge to p w ; The specific loss objective function is as follows:
[0089]
[0090] where represents the loss between the reconstructed p br and p w at the center of the image patch, and the l1 norm based on the mean absolute error is used in the loss to measure the smaller degree of error;
[0091] The trained background reconstruction network is obtained based on the network gradient descent method.
[0092] Beneficial effects:
[0093] The present invention designs a background reconstruction network framework guided by pre-detection, and designs a double-window blind spot structure in the training stage to guide the network to receive pure background information, enhancing the expression of background features. Moreover, the present invention designs a surrounding pixel weighting module to generate weighted pixels that fuse all spectral information around the central pixel as the label for network learning, weakening the fitting ability for target pixels. In the detection stage, a reconstruction-detection result fusion module is proposed for result fusion, which further enhances background suppression on the premise of ensuring the ability to detect specific targets. Therefore, the present invention can effectively solve the problems of scarce prior information of targets and poor background suppression ability of conventional methods in the field of target detection. Through experimental analysis, the method proposed by the present invention can obtain AUC values of 0.9988, 0.9997, 0.9985, and 0.9999 on four datasets respectively. Description of the Drawings
[0094] Figure 1 It is a flowchart of a method for detecting specific targets in hyperspectral images based on a double-window blind spot background reconstruction network with pixel weighting.
[0095] Figure 2 It is a false color image, ground truth image, and detection result image of Dataset I, where (a) is the false color image, (b) is the ground truth image, and (c) is the detection image.
[0096] Figure 3 It is a false color image, ground truth image, and detection result image of Dataset II, where (a) is the false color image, (b) is the ground truth image, and (c) is the detection image.
[0097] Figure 4 It is a false color image, ground truth image, and detection result image of Dataset III, where (a) is the false color image, (b) is the ground truth image, and (c) is the detection image.
[0098] Figure 5 It is a false color image, ground truth image, and detection result image of Dataset IV, where (a) is the false color image, (b) is the ground truth image, and (c) is the detection image. Detailed Embodiments
[0099] The present invention proposes a method for detecting specific targets in hyperspectral images based on a double-window blind spot background reconstruction network with pixel weighting. First, the present invention uses a pre-detection strategy and a double-window blind spot structure to generate network learning samples, and uses a surrounding pixel replacement strategy to generate a supervised label for reconstructing the central pixel. Then, a background reconstruction network is used to reconstruct the background of the data, and a detection and reconstruction result fusion module is used for processing, which improves the effectiveness of target detection while ensuring sufficient background suppression ability. The present invention can make better use of the background information of hyperspectral images to better detect specific targets. Next, detailed descriptions will be given in combination with specific embodiments. Detailed Embodiment 1:
[0101] This embodiment is a hyperspectral image specific target detection method based on a pixel weighted background reconstruction network, which is essentially a hyperspectral image specific target detection method based on a dual-window blind spot background reconstruction network with pixel weighting, and includes the following steps:
[0102] S1. The pre-detection strategy generation network learns samples:
[0103] First, use a pre-detection algorithm to process the input hyperspectral image:
[0104] Represent the original hyperspectral image with , where H, W, and L represent the number of rows, columns, and channels of the image respectively. Two-dimensionalize and normalize the hyperspectral image to obtain where N = H×W is the total number of pixels in the hyperspectral image.
[0105] According to the data r, the autocorrelation matrix R of the image can be obtained:
[0106]
[0107] where r i is the pixel of the data r, and i = 1, 2,..., N.
[0108] According to the known prior spectral information of the target pixels Use the fast CEM algorithm given by the following formula to pre-detect the original hyperspectral image:
[0109]
[0110] In the formula, is the optimal filter, and y is the detection result map of the pre-detection.
[0111] Theoretically, the response value corresponding to the background pixels in y should be close to 0, while the response value corresponding to the target pixels will be close to 1. Use the following strategy:
[0112]
[0113] where y(i, j) is the result of the i-th row and j-th column in the detection result map of the pre-detection, α is the threshold (empirically set); p tg represents the target pixels in the pre-detection result of the image, and p bg represents the background pixels in the pre-detection result. After this pre-detection step, only using one prior spectral information d of the target pixel, the target samples composed of the target pixels p tg and a large number of background samples that can be used for training and composed of the background pixels p bg are quickly and roughly obtained.
[0114] S2. Dual-window blind block processing:
[0115] After pre-detection, the original hyperspectral image is divided into small blocks with each pixel as the central pixel, and the width and height are both Wo, denoted as Subsequently, the subsequent background reconstruction network receives N = H × W image blocks, and uses the spatial and spectral information of each image block to complete the reconstruction of N central pixels, thereby obtaining the final detection image. Before feeding into the background reconstruction network, a dual-window structure is first used on the image block to generate a blind spot area, where the outer window size is the same as the image block size, which is W out ×W out . The inner window size is W in ×W in , denoted by . The blind spot area generation process is as follows:
[0116] The background reconstruction network usually reconstructs the central pixel by learning the information of the entire image block, but this may lead to inaccurate background reconstruction because it learns the target information. According to the characteristics of hyperspectral remote sensing, the pixels in the hyperspectral image and the surrounding pixels usually reflect the same or similar ground objects. Therefore, if there is an interference situation where the central pixel is a target pixel, the target with a certain spatial scale often falls within the inner window. Thus, the present invention constructs the target sample p tg and the background sample p bg according to the pre-detection, and randomly selects W out -P in background pixels in the area between the inner window and the outer window (P in ×W in ) to replace each pixel within the inner window. This makes the network unable to perceive the information within the inner window, that is, when the network reconstructs the central pixel of the image block, it only uses the surrounding background pixels between the inner and outer windows, thereby reducing the participation of the target pixel information in the network, which is equivalent to creating a blind spot area during network training. At the same time, since the inner window size is variable, the network is applicable to detecting targets of different sizes.
[0117] The present invention denotes the image block that has undergone dual-window blind block as The inner window of the image block that has undergone dual-window blind block is denoted as
[0118] S3. Adopt the surrounding pixel replacement strategy to generate a supervised label for reconstructing the central pixel:
[0119] Take out the inner window P ib of the image block that has undergone dual-window blind block, and define the pixels except the central pixel as the surrounding pixels. The set of these surrounding pixels is denoted by P sRepresentation. In the process of using the surrounding pixels to represent the central pixel label, different weights are assigned to them by evaluating the spectral differences between each surrounding pixel and the target prior information. In the present invention, spectral angle detection is used to measure the spectral differences. Using this method has two obvious advantages: on the one hand, it takes into account the absolute intensity differences and relative shape similarities between spectral features, effectively captures the wavelength correlation, can effectively distinguish substances with similar spectra but different properties, and enhances the accuracy of detection. On the other hand, it is insensitive to changes in illumination conditions. This enables high detection accuracy and stability to be maintained under different environmental conditions.
[0120] For each surrounding pixel p s in P si' and the prior spectral information d of the target pixel, the spectral angle difference L is calculated by the following formula i :
[0121]
[0122] where i' = 1, 2,..., n represents the i'-th pixel in the inner window, the inner window has n pixels, and n = W in ×W in -1 is the number of surrounding pixels.
[0123] After obtaining the spectral differences between each surrounding pixel and the prior target pixel, the weighted pixel p w is constructed as the label of the central pixel p c of the image patch by the following strategy:
[0124]
[0125] where j′ = 1, 2,..., n, j' represents the pixel in the inner window, n = W in ×W in -1 is the number of surrounding pixels, and β is the weight.
[0126] This strategy uses the n pixels adjacent to the central pixel in the inner window, assigns different weights according to the spectral similarity with the prior information, and then mixes them with the central pixel. Each pixel satisfies the background distribution of the outer window, thus weakening the network's fitting ability to the target.
[0127] S4. The background reconstruction network reconstructs the background of the data:
[0128] The background reconstruction network of the present invention is constructed based on the residual connection module. The receptive field of the network is consistent with the size of the image patch, and the input spectral information can be fully utilized. Due to the function of the blind spot area, for each input image patch P ob , only the pixels in its inner window participate in the learning of the network.
[0129] The structure of the network is as Figure 1 shown, mainly including a convolutional unit and a residual connection module network;
[0130] The convolutional unit is set before the residual connection module network. The convolutional unit includes two convolutional layers. After the first convolutional layer, a rectified linear unit (Relu) is connected, and after the second convolutional layer, a batch normalization layer (Batchnorm) is connected. The rectified linear unit and the batch normalization layer reduce the computational complexity, alleviate the vanishing gradient problem, and help the network accelerate training and improve the model stability.
[0131] The residual connection module network includes multiple residual connection modules, a convolutional layer set after the residual connection modules, a batch normalization layer, and a convolutional layer set after the batch normalization layer;
[0132] The specific structure of the residual connection module is as Figure 1 shown. After the element-wise summation of the input and output of each residual connection module, it is used as the input of the next residual connection module to continue learning. The output of the normalization layer after multiple residual connection modules is also element-wise summed with the input of the first residual connection module and finally fed into the last separate convolutional layer as the output of the residual connection module network;
[0133] In this embodiment, the residual connection module network is set with three residual connection modules. In other embodiments, other numbers of residual connection modules can be adopted.
[0134] In a neural network, low-level features capture some simple and local patterns, while high-level features represent more complex and abstract concepts. Through the residual module, the low-level feature maps can be directly connected to the high-level feature maps, enabling the network to refer to the low-level features when learning high-level features. At the same time, the skip connection provides a direct information flow path for the network, so that the low-level features will not be forgotten or lost after passing through multiple layers of transformations. In this way, even if the high-level features undergo multiple non-linear transformations, the original information at the bottom layer can still be retained and utilized. This direct connection method enables the network to better maintain the integrity of features.
[0135] Denote the proposed background reconstruction network model as N br , and the network parameters are Then the network during the training process can be expressed by the following formula:
[0136]
[0137] where p br represents the central pixel of P ob reconstructed by the network.
[0138] The network uses formulas (4) and (5) to obtain pixel p by weighting the surrounding pixels within the inner window of the image patch. w As p br For the supervision information, the training objective is to make p br Converge to p w . The specific loss objective function is as follows:
[0139]
[0140] Among them, Represents the loss between the reconstructed p br And p w . In the loss, the l1 norm based on the mean absolute error is used to measure the error to a lesser extent.
[0141] The process of network gradient descent uses the adaptive moment estimation stochastic optimization algorithm of Pytorch to obtain the trained background reconstruction network.
[0142] S5. Fusion of Detection and Reconstruction Results:
[0143] Hyperspectral images containing a large amount of data have the characteristics of data redundancy, scarce labels, and high dimensionality of the feature space. These problems make it quite difficult for the network to perform background reconstruction. Usually, the ideal state of the network is to have a small reconstruction difference for the background and a large reconstruction difference for the target, so as to detect the target by comparing the reconstructed image and the original image. Sometimes, the reconstruction accuracy of some regions is low, and some target samples are reconstructed by the network together. This leads to a reduction in the reconstruction error of the target, and the final detection result will have the problem that the target pixels are not prominent.
[0144] To solve this problem, a solution is proposed to merge and output the pre-detection result and the network reconstruction result after fusion.
[0145] In the actual detection process, the trained background reconstruction network is used to reconstruct the entire image to obtain the reconstructed image. The background reconstruction network model is denoted as N br , and the network parameters are Then the network in the detection process can be expressed by the following formula:
[0146]
[0147] Among them, Is the reconstructed image obtained through the network, Represents the original hyperspectral image input into the network during the detection process.
[0148] The pixels of the reconstructed image Can be obtained from the following formula The reconstruction difference between the pixel x of the original hyperspectral image X in the corresponding region v and
[0149]
[0150] where v = 1, 2,..., N represents each pixel of the image and N = H × W.
[0151] In addition, for the pre-detection result y obtained by formula (2) v the following background suppression function is used:
[0152]
[0153] where v = 1, 2,..., N and N = H × W.
[0154] The function B(x) can suppress the pixels with low target response values and retain as much as possible the pixels with high target response values. Applying B(x) to the pre-detection result achieves the purpose of filtering out misdetected targets and suppressing background pixels.
[0155] It should be noted that: the detection result y of the pre-detection here v can be obtained in various ways. For the original hyperspectral image X in the corresponding region, the corresponding pre-detection result y can be directly stored after being processed by formula (2) v and then directly used here. In fact, the original hyperspectral image X in the corresponding region can also be directly stored, and when used here, the pre-detection result y is obtained after processing X by formula (2) v and then the background suppression function is applied for suppression.
[0156] Subsequently, the reconstruction result is fused with it, and the fusion strategy is as follows:
[0157]
[0158] where v = 1, 2,..., N, N = H × W, B(y v ) is the pre-detection result after background suppression, is the reconstruction result generated by the network, and q v is the final detection result vector obtained by fusing the pre-detection result and the reconstruction result.
[0159] Combining the N result vectors q v gives the final detection image. Using this method of fusing the detection result and the reconstruction result not only effectively improves the effectiveness of target detection but also ensures sufficient background suppression ability.
[0160] Embodiment
[0161] Four hyperspectral data were used to illustrate the effect of the hyperspectral image specific target detection method based on the pixel-weighted dual-window blind spot background reconstruction network proposed by the present invention. The detailed information of the four data used is listed in Table 1. The area under the receiver operating characteristic curve (AUC) was used as the evaluation index for the experimental results. The higher the value of AUC, the better the detection effect.
[0162] Table 1 Details of the hyperspectral images used
[0163]
[0164] For different data, the optimal parameter settings of the method of the present invention are shown in Table 2. The parameter β is the weight coefficient, and W in and W out are the sizes of the inner and outer windows. The language environment is python, and the experimental hardware platform: CPU: i5-7200U, graphics card: GTX2080Ti, memory: 8G.
[0165] Table 2 Optimal parameters and AUC values on three groups of experimental data
[0166] Specific implementation method 2:
[0168] This implementation method is a hyperspectral image specific target detection system based on a pixel-weighted background reconstruction network, including:
[0169] Hyperspectral image acquisition unit: For the area to be detected, the original hyperspectral image X is acquired;
[0170] Image reconstruction and reconstruction difference calculation unit: Using the background reconstruction network built based on the convolutional network to reconstruct the original hyperspectral image X to obtain the reconstructed image From the obtained reconstructed image of the pixel and the pixel x of the original hyperspectral image X in the corresponding area v Calculate the reconstruction difference
[0171] Reconstructed image Among them, is the reconstructed image obtained through the network, is the network parameter;
[0172] Reconstruction difference Among them, v = 1, 2,..., N represents each pixel of the image X, and N is the total number of pixels of the hyperspectral image;
[0173] The background reconstruction network built based on the convolutional network includes a convolutional unit and a residual connection module network; the convolutional unit is arranged before the residual connection module network, the convolutional unit includes two convolutional layers, the first convolutional layer is followed by a rectified linear unit, and the second convolutional layer is followed by a batch normalization layer; the residual connection module network includes a plurality of residual connection modules, a convolutional layer arranged after the residual connection modules, a batch normalization layer, and a convolutional layer arranged after the batch normalization layer; after the input and output of each residual connection module are element-wise summed, it is used as the input of the next residual connection module; the output of the normalization layer after a plurality of residual connection modules is element-wise summed with the input of the first residual connection module and finally sent into the last separate convolutional layer as the output of the residual connection module network.
[0174] In this embodiment, three residual connection modules are provided in the residual connection module network.
[0175] The background reconstruction network built based on the convolutional network is pre-trained, and the specific training process includes:
[0176] S1. The pre-detection strategy generation network learns samples:
[0177] First, use the pre-detection algorithm to pre-detect the original hyperspectral image to obtain the pre-detection result map y, and then process it using the following strategy:
[0178]
[0179] Among them, y(i,j) is the result of the i-th row and j-th column in the pre-detection result map, and α is the threshold; p tg represents the target pixel of the pre-detection result in the image, and p bg represents the background pixel of the pre-detection result;
[0180] S2. Double-window blind block processing:
[0181] Divide the original hyperspectral image X into small blocks with each pixel as the central pixel, and the width and height are both Wo, denoted as First, use a double-window structure on the image block to generate a blind spot area, where the outer window size is the same as the image block size, which is W out ×W out ; the inner window size is W in ×W in , denoted by . The blind spot area generation process is as follows:
[0182] Based on the target sample p tg and background sample p bg constructed by pre-detection, randomly select W in ×W inReplace each pixel inside the inner window with a background pixel; Denote the image block that has undergone double-window blind blocking as Denote the inner window of the image block that has undergone double-window blind blocking as
[0183] S3. Adopt the surrounding pixel replacement strategy to generate a supervised label for reconstructing the central pixel:
[0184] Take out the inner window P of the image block that has undergone double-window blind blocking ib and define the pixels except the central pixel as surrounding pixels. The set of these surrounding pixels is denoted by P s For each surrounding pixel p s in P si' and the prior spectral information d of the target pixel, calculate the spectral angle difference L through the following formula i :
[0185]
[0186] where i' = 1, 2,..., n represents the i'-th pixel inside the inner window, the inner window has n pixels, and n = W in ×W in -1 is the number of surrounding pixels;
[0187] After obtaining the spectral differences between each surrounding pixel and the prior target pixel, construct the weighted pixel p w as the label of the central pixel p c of the image block through the following strategy:
[0188]
[0189] where j′ = 1, 2,..., n, j' represents the pixel inside the inner window, n = W in ×W in -1 is the number of surrounding pixels, and β is the weight;
[0190] S4. The background reconstruction network performs background reconstruction on the data:
[0191] For each input image block P ob , denote the proposed background reconstruction network model as N br , and the network during the training process can be represented by the following formula:
[0192]
[0193] where p br represents the central pixel of P ob reconstructed by the network;
[0194] The pixel p obtained by weighting the surrounding pixels inside the inner window of the image block wAs the supervision information of p br , the training objective is to make p br converge to p w ; specifically, the loss objective function is as follows:
[0195]
[0196] where, represents the loss between the reconstructed p br at the center pixel of the image patch and p w . The l1 norm based on the mean absolute error is used in the loss to measure the error to a lesser extent;
[0197] The trained background reconstruction network is obtained based on the network gradient descent method.
[0198] Background suppression unit: Perform background suppression on the pre-detection result y v obtained by pre-detecting the original hyperspectral image X. When performing background suppression on the pre-detection result y v , the following background suppression function is used:
[0199]
[0200] Fusion detection unit: Based on B(y v ) and perform fusion to obtain the detection result vector Combine N result vectors q v to obtain the final detection image.
[0201] The process of obtaining the pre-detection result by pre-detecting the original hyperspectral image X includes:
[0202] Represent the original hyperspectral image as , where H, W, and L represent the number of rows, columns, and channels of the image respectively; Two-dimensionalize and normalize the hyperspectral image to obtain where N = H×W is the total number of pixels in the hyperspectral image;
[0203] According to the data r, obtain the autocorrelation matrix R of the image:
[0204]
[0205] where r i is the pixel of the data r, and i = 1, 2,..., N;
[0206] According to the known prior spectral information of the target pixel perform pre-detection on the original hyperspectral image:
[0207]
[0208] In the formula, is a filter, and y is the detection result graph to be pre-detected.
[0209] The above calculation examples of the present invention are only for explaining in detail the calculation model and calculation process of the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made on the basis of the above description. It is impossible to list all the implementation manners here. Any obvious changes or variations derived from the technical solutions of the present invention still fall within the protection scope of the present invention.
Claims
1. A method for detecting specific targets in hyperspectral images based on a background reconstruction network with pixel weighting, characterized in that, It includes the following steps: For the area to be detected, obtain the original hyperspectral image X, and use the background reconstruction network built based on the convolutional network to reconstruct the entire image to obtain the reconstructed image: Among them, is the reconstructed image obtained through the network, is the network parameter; The reconstructed image is obtained from the following formula of the pixel and the pixel x of the original hyperspectral image X in the corresponding region v The reconstruction difference between Where v = 1, 2,..., N represents each pixel of the image X, and N is the total number of pixels in the hyperspectral image; Simultaneously obtain the pre-detection result obtained by pre-detecting the original hyperspectral image X, and use the pre-detection result y v Use the following background suppression function: Then the following fusion is performed: where q v is the finally obtained detection result vector; Combining N result vectors q v results in the final detected image.
2. A method for detecting specific targets in hyperspectral images of a background reconstruction network based on pixel weighting according to claim 1, characterized in that, The process of obtaining the pre-detection result after the original hyperspectral image X undergoes pre-detection includes: The original hyperspectral image is denoted as, where H, W, and L represent the number of rows, columns, and channels of the image respectively; the hyperspectral image is two-dimensionalized and normalized to obtain where N = H × W is the total number of pixels in the hyperspectral image; According to the data r, obtain the autocorrelation matrix R of the image: where r i is the pixel of data r, and i = 1, 2, ..., N; According to the prior spectral information of known target pixels Perform pre-detection on the original hyperspectral image: In the formula, is a filter, and y is the detection result graph to be pre-detected.
3. A hyperspectral image specific target detection method based on a background reconstruction network with pixel weighting as claimed in claim 2, characterized in that, The background reconstruction network built based on the convolutional network includes a convolutional unit and a residual connection module network; the convolutional unit is set before the residual connection module network, the convolutional unit includes two convolutional layers, the first convolutional layer is followed by a rectified linear unit, and the second convolutional layer is followed by a batch normalization layer; the residual connection module network includes multiple residual connection modules, a convolutional layer set after the residual connection module, a batch normalization layer, and a convolutional layer set after the batch normalization layer; after the input and output of each residual connection module are element-wise summed, it serves as the input of the next residual connection module; The output of the normalization layer after multiple residual connection modules is element-wise summed with the input of the first residual connection module and finally fed into the last separate convolutional layer as the output of the residual connection module network.
4. A hyperspectral image specific target detection method based on a background reconstruction network with pixel weighting according to claim 3, characterized in that, The residual connection module network is set with three residual connection modules.
5. A method for detecting a specific target in a hyperspectral image of a background reconstruction network based on pixel weighting according to any one of claims 2 to 4, characterized in that, The background reconstruction network built based on the convolutional network is pre-trained, and the specific training process includes: S1. The pre-detection strategy generation network learns samples: First, use the pre-detection algorithm to pre-detect the original hyperspectral image to obtain the pre-detection result map y, and then use the following strategy to process: Among them, y(i,j) is the result of the i-th row and j-th column in the pre-detected detection result map, α is the threshold; p tg represents the target pixel of the pre-detection result in the image, p bg represents the background pixel of the pre-detection result; S2. Double-window blind block processing: Divide the original hyperspectral image X into small patches centered on each pixel, with both the width and height being Wo, denoted as First, use a double-window structure on the image patches to generate blind spot regions, where the outer window size is the same as the image patch size, which is W out ×W out ; the inner window size is W in ×W in , denoted by . The blind spot region generation process is as follows: Target sample p constructed based on pre-detection tg and background sample p bg , randomly select W in ×W in background pixels in the area between the inner window and the outer window to replace each pixel in the inner window; Denote the image block that has undergone double-window blind block as The inner window of the image block that has undergone double-window blind block is denoted as S3. Adopt the surrounding pixel replacement strategy to generate a supervised label for reconstructing the central pixel: Take out the inner window P of the image block passing through the double-window blind block ib and define the pixels except the central pixel as surrounding pixels, and the set of these surrounding pixels is represented by P s ; For each surrounding pixel p s in P si' and the prior spectral information d of the target pixel, calculate the spectral angle difference L through the following formula i : where \(i' = 1, 2, \cdots, n\) represents the \(i'\)-th pixel inside the inner window, the inner window has \(n\) pixels, and \(n = W in \times W in - 1 is the number of surrounding pixels; After obtaining the spectral differences between each surrounding pixel and the prior target pixel, the weighted pixel p is constructed through the following strategy w As the central pixel p of the image patch c Label of where \(j' = 1, 2, \cdots, n\), \(j'\) represents the pixel within the inner window, and \(n = W\) in \(\times W\) in \(- 1\) is the number of surrounding pixels, and \(\beta\) is the weight; S4. The background reconstruction network performs background reconstruction on the data: For each input image patch P ob , denote the proposed background reconstruction network model as N br , and the network during the training process can be expressed by the following formula: where p br represents the central pixel of P ob reconstructed by the network; Pixel p obtained by weighting surrounding pixels within the inner window of the image block w As p br The supervision information, and the training objective is to make p br Converge to p w ; The specific loss objective function is as follows: Among them, represents the loss of the reconstructed pixel p at the center of the image block br and p w The loss uses the l1 norm based on the mean absolute error to measure the error to a lesser extent; Obtain the trained background reconstruction network based on the network gradient descent method.
6. A hyperspectral image specific target detection system based on a pixel-weighted background reconstruction network, characterized in that, It includes: Hyperspectral image acquisition unit: For the area to be detected, obtain the original hyperspectral image X; Image reconstruction and reconstruction difference calculation unit: The original hyperspectral image X is reconstructed using a background reconstruction network based on a convolutional network to obtain a reconstructed image From the obtained reconstructed image Pixels And the pixels x of the original hyperspectral image X in the corresponding region v Calculate the reconstruction difference Reconstructed image Among them, is the reconstructed image obtained through the network, is the network parameter; Reconstruction difference where \(v = 1,2,\cdots,N\) represents each pixel of the image \(X\), and \(N\) is the total number of pixels in the hyperspectral image; Background suppression unit: the preliminary detection result y obtained by preliminary detection of the original hyperspectral image X v Perform background suppression on the preliminary detection result y v The following background suppression function is used for background suppression: Fusion detection unit: Based on B(y v ) and perform fusion to obtain the detection result vector Combine the N result vectors q v to obtain the final detected image.
7. A hyperspectral image specific target detection system based on a pixel weighted background reconstruction network according to claim 6, characterized in that, The process of obtaining the pre-detection result after the original hyperspectral image X undergoes pre-detection includes: The original hyperspectral image is represented, where H, W, and L represent the number of rows, columns, and channels of the image respectively; the hyperspectral image is two-dimensionalized and normalized to obtain where N = H × W is the total number of pixels in the hyperspectral image; According to the data r, obtain the autocorrelation matrix R of the image: where r i is the pixel of data r, and i = 1, 2,..., N; According to the known prior spectral information of the target pixel Perform pre-detection on the original hyperspectral image: In the formula, is a filter, and y is the detection result graph to be pre-detected.
8. A hyperspectral image specific target detection system based on a background reconstruction network with pixel weighting as claimed in claim 7, characterized in that, The background reconstruction network built based on the convolutional network includes a convolutional unit and a residual connection module network; the convolutional unit is set before the residual connection module network, the convolutional unit includes two convolutional layers, the first convolutional layer is followed by a rectified linear unit, and the second convolutional layer is followed by a batch normalization layer; the residual connection module network includes multiple residual connection modules, a convolutional layer set after the residual connection module, a batch normalization layer, and a convolutional layer set after the batch normalization layer; after the input and output of each residual connection module are element-wise summed, it serves as the input of the next residual connection module; The output of the normalization layer after multiple residual connection modules is element-wise summed with the input of the first residual connection module and finally fed into the last separate convolutional layer as the output of the residual connection module network.
9. A hyperspectral image specific target detection system based on a pixel-weighted background reconstruction network according to claim 8, characterized in that, The residual connection module network is set with three residual connection modules.
10. A hyperspectral image specific target detection system based on a pixel-weighted background reconstruction network according to any one of claims 7 to 9, characterized in that, The background reconstruction network built based on the convolutional network is pre-trained, and the specific training process includes: S1. Generate network learning samples for pre-detection strategy: First, use the pre-detection algorithm to perform pre-detection on the original hyperspectral image to obtain the pre-detection result map y, and then process it using the following strategy: Among them, y(i,j) is the result of the i-th row and j-th column in the pre-detected detection result map, α is the threshold; p tg represents the target pixel of the pre-detected result in the image, p bg represents the background pixel of the pre-detected result; S2. Dual-window blind block processing: Divide the original hyperspectral image X into small patches with each pixel as the central pixel, where the width and height are both Wo, denoted as First, use a double-window structure on the image patches to generate blind spot regions. The size of the outer window is the same as that of the image patch, which is W out ×W out ; the size of the inner window is W in ×W in , denoted by . The blind spot region generation process is as follows: Target sample p constructed based on pre-detection tg and background sample p bg , randomly select W in ×W in background pixels in the area between the inner window and the outer window to replace each pixel inside the inner window; Denote the image block that has undergone double-window blind block as Denote the inner window of the image block that has undergone double-window blind block as S3. Adopt the surrounding pixel replacement strategy to generate supervised labels for reconstructing the central pixel: Take out the inner window P of the image block passing through the double-window blind block, and define the pixels except the central pixel as surrounding pixels. The set of these surrounding pixels is represented by P ib ; For each surrounding pixel p in P s and the prior spectral information d of the target pixel, calculate the spectral angle difference L through the following formula s : si' i where i' = 1, 2,..., n represents the i'-th pixel within the inner window, the inner window has n pixels, and n = W in ×W in -1 is the number of surrounding pixels; After obtaining the spectral differences between each surrounding pixel and the prior target pixel, the weighted pixel p is constructed through the following strategy w As the central pixel p of the image patch c Label of where j′ = 1, 2,..., n, j' represents the pixel within the inner window, and n = W in × W in −1 is the number of surrounding pixels, and β is the weight; S4. The background reconstruction network performs background reconstruction on the data: For each input image patch P ob , denote the proposed background reconstruction network model as N br , and the network during the training process can be expressed by the following formula: where p br represents the central pixel of P ob reconstructed by the network; Pixel p obtained by weighting surrounding pixels within the inner window of the image block w As p br The supervision information, and the training objective is to make p br Converge to p w ; The specific loss objective function is as follows: Among them, represents the loss of the central pixel of the image block after reconstruction, where p br and p w The loss uses the l1 norm based on the mean absolute error to measure the error to a lesser degree; Obtain the trained background reconstruction network based on the network gradient descent method.