A Single-Frame Image Infrared Dim and Small Target Detection Method Based on Deep U-Net
Through the combination of deep U-shaped network and dense feature coding module, the problem of insufficient sample size and insufficient feature resolution in infrared weak target detection is solved, and efficient and accurate infrared weak target detection is achieved.
Patent Information
- Application Number
- CN202210947869.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-09
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-08-09
AI Technical Summary
The existing infrared weak target detection methods have poor detection performance under complex backgrounds, and due to insufficient sample size and small target size, it is difficult to train deep learning models, resulting in low detection accuracy and insufficient feature resolution, making it difficult to effectively extract multi-scale and high-resolution features.
The infrared weak object detection method based on deep U-shaped network is adopted. By building a deep supervised U-shaped network, combining dense feature coding modules, multi-level and multi-scale feature extraction and reduction are carried out, and residual U-shaped blocks and attention-oriented learning is used to improve feature resolution and target representation capabilities.
It improves the accuracy and efficiency of infrared weak target detection, solves the problems of insufficient sample size and small target size, enhances the global and local context information representation of the target, and reduces the difficulty of detection in complex contexts.
Smart Images

Figure CN115311508B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of single-frame image infrared small and weak target detection, and particularly to the technical field of infrared small and weak target detection methods based on deep learning. Background Art
[0002] The infrared band is an electromagnetic wave with a frequency between microwave and visible light, and its frequency range in the electromagnetic spectrum is 0.3 THz to 400 THz. Due to the advantage that the data in this band is not affected by the environment, light, occlusion and other conditions, it is widely used in military applications, such as infrared guidance, early warning, etc. However, small and weak targets in infrared images generally do not exceed 30 pixels, are usually covered in complex backgrounds, and lack color and texture information, making detection difficult.
[0003] Currently, the methods for detecting small and weak infrared targets in single-frame images mainly include model-driven methods and data-driven methods. When the small and weak targets in the infrared image are brighter than the background, some methods based on the human visual system, such as block-based methods, local contrast strategy-based methods, and multi-scale block contrast-based methods, etc., can accurately screen out the regions where the bright targets are located through the visual attention mechanism, so as to achieve the final detection. Usually, the feature representation ability of such methods is better than that of methods based on filtering and low-rank sparse decomposition. However, model-driven methods are easily affected by clutter and noise, which greatly limits the detection performance of such methods for small and weak infrared targets in single-frame images or video sequence images in complex scenarios.
[0004] The rapid development of deep learning has made it widely used in infrared small and weak target detection. However, due to the limited publicly available infrared data and the small size of the targets, it cannot be directly used to train large-scale detection networks. Therefore, most of the existing infrared small and weak target detection methods based on deep learning usually achieve the final detection by directly migrating or fine-tuning pre-trained models generated based on natural scene images. However, due to the different distributions of infrared small and weak target data and natural scene data, the false negative rate of the detection model is very high. In fact, modeling infrared small and weak target detection as a semantic segmentation problem, rather than a typical target detection problem, helps to better solve the problem of limited detection performance of the model caused by the small size of the targets. However, almost all of the existing networks rely on network architectures suitable for image classification with classical downsampling schemes. As the network deepens, the resolution of the target features is greatly reduced or even lost, which is particularly disadvantageous for the detection of infrared small and weak targets.
[0005] Summarizing the above existing problems, it can be seen that for the problem of infrared small and weak target detection, it is urgent to construct a learning network with deep multi-scale, high-resolution features, and capable of enhancing the representation of the global and local context information of the target. This network also needs to adapt to situations such as insufficient sample size and small target size. Therefore, the present invention fully considers the above problems existing in the infrared small and weak target detection task, and proposes an infrared small and weak target detection method based on a deep U-shaped network. Summary of the Invention
[0006] The purpose of the present invention is to overcome the problems in existing infrared small and weak target detection methods, such as low local contrast of small and weak targets in images and the contradiction between the depth of the network and the feature resolution. A deep U-shaped network infrared small and weak target detection method for single-frame infrared video images is proposed. This method can also solve the problem that the insufficient sample size of infrared small and weak targets is difficult to support the training of complex deep network models, and at the same time has the ability to extract multi-scale and highly distinguishable features, and finally generates a detection model with high accuracy and low complexity.
[0007] The technical solution of the present invention is as follows:
[0008] A single-frame image infrared small and weak target detection method based on a deep U-shaped network, which includes:
[0009] S1: Construct a single-frame image infrared small and weak target detection model based on a deep U-shaped network;
[0010] S2: Train the detection model through the labeled single-frame infrared image sample set or its enhanced sample set after enhancement processing;
[0011] S3: Implement the detection of infrared small and weak targets in the video sequence image set through the trained detection model;
[0012] Among them, the deep U-shaped network includes a deep supervised U-shaped network and a dense feature encoding module integrated into the deep supervised U-shaped network;
[0013] Among them, the deep supervised U-shaped network includes a compression path network for obtaining multi-level and multi-scale image extraction features, and an expansion path network for restoring the image accuracy at multi-level and multi-scale. The dense feature encoding module is located between the compression path network and the expansion path network;
[0014] The dense feature encoding module includes a first encoding module that performs channel attention cross-guided learning on the low-level detail features of multi-level and multi-scale extracted features obtained by the compression path network to obtain multi-level low-level features; a second encoding module that performs spatial attention interaction-guided learning on the low-level features to obtain multi-level high-level features; and a third encoding module that cascades and fuses the multi-level low-level features and the multi-level high-level features to obtain dense encoded features.
[0015] According to some preferred embodiments of the present invention, the deep U-shaped network sequentially includes an input layer, the compression path network, the dense feature encoding module, the expansion path network, and an output layer.
[0016] According to some preferred embodiments of the present invention, the deep supervised U-shaped network is a nested U-shaped network composed of the compression path network and the expansion path network. Among them, the compression path network includes multi-level and multi-scale extraction and compression modules connected in sequence. Except for the last extraction and compression module that only contains 1 residual U-shaped block, each of the remaining extraction and compression modules includes at least 1 residual U-shaped block and at least 1 downsampling layer connected thereto; the expansion path network includes multi-level and multi-scale expansion and reduction modules connected in sequence. Each expansion and reduction module includes at least 1 upsampling layer and at least 1 residual U-shaped block connected thereto; among them, the residual U-shaped block of the first extraction and compression module is connected to the input layer, the downsampling layer of the last extraction and compression module is connected to the first encoding module in the dense feature encoding module, the residual U-shaped block of each extraction and compression module in between is connected to the downsampling layer of its previous extraction and compression module, the upsampling layer of the first expansion and reduction module is connected to the third encoding module of the dense feature encoding module, the residual U-shaped block of the last expansion and reduction module is connected to the output layer, the upsampling layer of each expansion and reduction module in between is connected to the residual U-shaped block of its previous expansion and reduction module, and the downsampling layer of each extraction and compression module is also connected to the upsampling layer of the expansion and reduction module corresponding to its size, and the residual U-shaped block of each expansion and reduction module is also connected to the first encoding module in the dense feature encoding module.
[0017] According to some preferred embodiments of the present invention, the downsampling layer uses max-pooling downsampling.
[0018] According to some preferred embodiments of the present invention, the compression path network includes a 6-layer network structure, each layer constitutes one of the extraction and compression modules, and the dilation rates of the dilated convolutions of the residual U-shaped blocks included in each layer are different.
[0019] According to some preferred embodiments of the present invention, the extended path network includes a 5-layer network structure, each layer constitutes one of the extended reduction modules, and the dilation rates of the dilated convolutions of the residual U-shaped blocks included in each layer are different.
[0020] According to some preferred embodiments of the present invention, in the 6-layer network structure of the compression path network, each of the first to fifth layers of the network only includes a plurality of convolutional layers with a dilation rate of 1, that is, a traditional convolutional network, and a dilated convolutional layer with a dilation rate of 2, and the sixth layer of the network includes 4 dilated convolutional layers with dilation rates of 1, 2, 4, and 8 respectively.
[0021] According to some preferred embodiments of the present invention, in the 5-layer network structure of the extended path network, each of the first to fifth layers of the network only includes a plurality of convolutional layers with a dilation rate of 1, that is, a traditional convolutional network, and a dilated convolutional layer with a dilation rate of 2.
[0022] According to some preferred embodiments of the present invention, in the 6-layer network structure of the compression path network, the depths of the residual U-shaped blocks of the first to sixth layers of the network are 7, 6, 5, 4, 3, and 3 in sequence.
[0023] According to some preferred embodiments of the present invention, in the 6-layer network structure of the compression path network, the number of input channels of the residual U-shaped block included in the first layer of the network is 3, the number of input channels of the residual U-shaped block included in the second layer of the network is 64, the number of input channels of the residual U-shaped block included in the third layer of the network is 128, the number of input channels of the residual U-shaped block included in the fourth layer of the network is 256, the number of input channels of the residual U-shaped block included in the fifth layer of the network is 512, and the number of input channels of the residual U-shaped block included in the sixth layer of the network is 512.
[0024] According to some preferred embodiments of the present invention, the method for target detection according to the target detection model is as follows:
[0025] S11 Obtain multi-level and multi-scale depth features of small and weak targets in single-frame infrared image data according to the compression path network;
[0026] S12 Perform channel attention cross-guided learning, spatial attention interaction-guided learning, and cascade fusion on the optimized multi-level and multi-scale depth features according to the dense feature encoding module to obtain dense encoded features;
[0027] S13 Perform multi-level and multi-scale depth supervised decoding on the dense encoded features according to the extended path network to obtain the detection result.
[0028] According to some preferred embodiments of the present invention, the detection method further includes: performing a first convolution process and a second convolution process by the compression path network, wherein the first convolution process includes a residual U-shaped block process and a rectified linear unit activation process, and the second convolution process includes a residual U-shaped block process and a max-pooling downsampling process; performing a third convolution process by the expansion path network, which includes a residual U-shaped block process and an upsampling process.
[0029] According to some preferred embodiments of the present invention, S13 further includes activating the non-linear feature expression of the dense coding feature to obtain a final dense coding feature, and obtaining a detection result according to the multi-level and multi-scale depth supervision decoding of the final dense coding feature.
[0030] According to some preferred embodiments of the present invention, the depth feature is obtained through the following calculation model:
[0031]
[0032] wherein, F k represents the feature obtained after learning by the residual U-shaped block of the k-th layer of the compression path network, and σ(·) represents the sigmoid activation function.
[0033] According to some preferred embodiments of the present invention, the depth feature is obtained through the following calculation model:
[0034]
[0035] wherein, U k (·)(k = 1, 2, …, K) represents the unfolded representation of the feature obtained after learning by the residual U-shaped block of the k-th layer in the compression path network, and K represents the maximum number of layers of the residual U-shaped block.
[0036] According to some preferred embodiments of the present invention, the obtaining of the dense coding feature includes:
[0037] Inputting the feature obtained from the k-th layer of the compression path network into the dense feature encoding module, performing adaptive average pooling to obtain the feature F k ′;
[0038] Inputting the feature F k ′ after the adaptive average pooling into a two-layer network with different weights and numbers of neurons, and feeding it into the ReLu activation function to obtain a first transformed feature;
[0039] Activating the first transformed feature through the Sigmoid function to obtain a weight coefficient A1, and further obtaining the feature of channel attention cross-guided learning as follows:
[0040]
[0041] The features of the cross-guided learning of channel attention Perform a global maximum pooling and average pooling in the spatial domain to obtain two feature maps with 1 channel number;
[0042] Concatenate the two feature maps, and pass the obtained concatenated feature map through a neural network with a channel transformation and a ReLu activation function to obtain a second transformed feature;
[0043] Pass the second transformed feature through a 7×7 convolutional layer and activate it with a Sigmoid function to obtain a weight coefficient A2, and further obtain the feature after the spatial attention interaction-guided encoding As follows:
[0044]
[0045] According to some preferred embodiments of the present invention, the feature F k ' after adaptive average pooling is obtained through the following calculation model:
[0046]
[0047] Wherein, F k represents the feature extracted by the k-th layer compression path network, W and H represent the width and height of the image corresponding to the feature, and i, j represent the position coordinates of the target, and their values are all positions traversed in the image.
[0048] According to some preferred embodiments of the present invention, the weight coefficient A1 is obtained through the following calculation model:
[0049] A1 = σ(Β(W2δ(Β(W1F k '))) (4)
[0050] Wherein, δ(·) and Β(·) respectively represent the rectified linear unit ReLU and batch normalization, and respectively represent the weight coefficients when the number of channels after passing through the first layer of neurons is reduced to 1 / 4 of the initial value and the number of channels is restored to the initial channel number after passing through the second layer of neurons, wherein C represents the number of channels, represents the channel compression or stretching factor.
[0051] According to some preferred embodiments of the present invention, the weight coefficient A2 is obtained through the following calculation model:
[0052]
[0053] Among them, is the feature for channel attention cross-direction learning The transformed feature after channel transformation, 1×1 and 3×3 represent convolution operations, P avg and P max represent the average pooling and max pooling operators respectively.
[0054] According to some preferred embodiments of the present invention, the final dense coding feature is obtained through the following calculation model:
[0055]
[0056] Among them, F DFE represents the dense coding feature.
[0057] According to some preferred embodiments of the present invention, in the training, the loss function of the detection model is set as:
[0058]
[0059] Among them, represents the loss function of the m-th layer in the deep supervised U-shaped network, corresponds to the weight of the loss function of the m-th layer, Loss fuse represents the loss function of the final fusion output, ω fuse corresponds to the weight of the loss function of the fusion output, and M represents the total number of layers of the deep supervised U-shaped network.
[0060] According to some preferred embodiments of the present invention, each loss function in the loss function of the detection model is calculated using the standard binary cross-entropy, as follows:
[0061]
[0062] Among them, i and j represent the coordinates of pixels in the image, H and W are the sizes of the image, p G(i,j) and p S(i,j) represent the probability values of the output maps of the reference pixel value and the predicted pixel value respectively.
[0063] According to some preferred embodiments of the present invention, in the training, after a certain number of training rounds, the training results are verified once, and the best IoU and nIoU values in each verification are recorded. When the IoU and nIoU values of the current round are greater than the existing best values, the current learning rate, the parameters of the network, as well as the IoU value and the nIoU value are saved.
[0064] According to some preferred embodiments of the present invention, the IoU value is obtained through the following calculation model:
[0065]
[0066] Among them, M represents the number of targets in a single-frame image, N represents the total number of samples, A int er and A all respectively represent the intersection and union of the real object and the predicted object, T, P, and TP respectively represent the real label, the predicted result, and the number of correctly detected pixels, and i and m respectively represent the i-th sample and the m-th object.
[0067] According to some preferred embodiments of the present invention, the nIoU value is obtained through the following calculation model:
[0068]
[0069] Among them, represents the IoU value of each sample, and i in T[i], P[i], and TP[i] represents the i-th sample.
[0070] The present invention has the following beneficial effects:
[0071] The present invention fully considers the problems of insufficient sample size, weak target size, and complex and changeable target background in the detection of infrared small and weak targets, and proposes an infrared small and weak target detection method based on a deep U-shaped network, which has high detection efficiency, high accuracy, and strong generalization ability.
[0072] Compared with the classical segmentation network and detection network suitable for image classification, the deep supervision U-shaped network structure based on residual U-shaped blocks proposed by the present invention can not only solve the contradiction between network depth and feature resolution, indirectly improve the global context representation of the target, but also learn multi-scale features of the target between the same layer and different layers of this network, effectively alleviating the low distinguishability of the deep features of small and weak targets and solving the problem that targets are easily lost in deep networks. Moreover, the depth of the residual block can be dynamically adjusted according to the size of the feature map and the number of downsamplings, so that the network structure and the learning ability of the target features reach the optimal state.
[0073] The structural design of channel attention cross-guided learning for low-level detailed features and spatial attention interaction-guided learning for high-level semantic features of the present invention effectively solves the problems of poor distinguishability and weak representation ability of deep features of infrared small and weak targets, and effectively constructs long-range dependence relationships between pixel-based targets, improving the local context representation of targets. Among them, channel attention cross-guided learning effectively mines the feature representations of those potential, implicit, and diagnostic targets through a top-down encoding method, transmits high-level semantic information into low-level features, and optimizes low-level features. Spatial attention interaction-guided learning integrates the features after channel attention cross-learning and the context feature representation of spatial attention, and further enriches the detailed representation of high-level semantic features through a bottom-up pixel-based encoding method, transmits low-level detailed information into high-level features, and optimizes the local context representation of targets.
[0074] In a further specific implementation, the loss function of the present invention is different from a binary cross-entropy loss function at the end of a standard network. The loss function of the deep U-shaped network is essentially the sum of multiple loss functions, including the output results of each layer of the deep supervision network plus the result after feature fusion. This not only avoids the low discriminability or absence of small and weak targets due to the single output of deep network features, but also solves problems such as the vanishing gradient and slow convergence speed in the training of deep neural networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 It is a flowchart of an infrared small and weak target detection method based on a deep U-shaped network in a specific implementation.
[0076] Figure 2 It is a detailed structure diagram of the deep U-shaped network constructed in a specific implementation.
[0077] Figure 3 It is a structure diagram of a deep supervision U-shaped network based on a residual U-shaped block in a specific implementation.
[0078] Figure 4 It is a schematic diagram of a dense feature encoding module in a specific implementation.
[0079] Figure 5 It is a flowchart of detecting small and weak targets in an infrared video sequence image by the deep U-shaped network detection method in an embodiment.
[0080] Figure 6 It is a visualization result of detecting small and weak targets in an infrared video sequence by the deep U-shaped network detection method in an embodiment. SPECIFIC IMPLEMENTATION
[0081] The present invention will be described in detail below in conjunction with embodiments and the accompanying drawings. However, it should be understood that the embodiments and the drawings are only used for exemplary description of the present invention, and do not constitute any limitation to the protection scope of the present invention. All reasonable transformations and combinations within the scope of the inventive concept of the present invention fall within the protection scope of the present invention.
[0082] According to the technical solution of the present invention, a specific implementation manner includes a detection and recognition process as shown in the accompanying Figure 1 drawings, which specifically includes the following steps:
[0083] S1: Construct a single-frame image infrared small and weak target detection model based on a deep U-shaped network.
[0084] More specifically, referring to the accompanying Figure 2 drawings, the detection model is constructed based on a deep supervised U-shaped network and a dense feature encoding module added to the deep supervised U-shaped network.
[0085] This structure grafted with a dense feature encoding module can effectively improve the interactive perception ability between the low-level detail features and high-level semantic features of the target, further improve the representation of the target's context information, and solve the problems of weak deep semantic features of infrared small targets and easy confusion between target and background features.
[0086] Furthermore, in the Figure 2 specific embodiment shown in the drawings, the deep supervised U-shaped network is a classic fully convolutional network, including an input layer, a compression path network, a dense feature encoding module, an expansion path network, and an output layer. The dense encoding module is located between the compression path and the expansion path.
[0087] Furthermore, in the Figure 2 specific embodiment shown in the drawings, the compression path network is composed of 6 network modules. Except for the last module that does not require max-pooling downsampling, each module uses a residual U-shaped block and 1 max-pooling-based downsampling. After each downsampling, the size of the feature map is reduced by 1 / 2. Among them, the residual U-shaped block in the first module is connected to the input layer, and the input of the remaining other modules is the output of the previous module. The expansion path network is composed of 5 network modules. Before each module starts, the size of the feature map is multiplied by 2 through deconvolution (i.e., upsampling), and then cascaded with the feature map of the symmetric compression path on the left and sent to the dense encoding module. The transformed feature is the output of this module.
[0088] Furthermore, the above structure based on the multi-scale residual U-shaped block can effectively solve the contradiction between the depth of the network and the feature resolution, while enhancing the representation of the global context information of small and weak targets. In this structure, the feature maps of each module in the compression path can be used as the input of the residual U-shaped block to obtain the depth multi-scale features of each module in the compression path, and a multi-level, multi-scale high-resolution depth feature map of small and weak targets can be obtained.
[0089] Among them, furthermore, the multi-level, multi-scale compression path network may include 1 to 4 dilated convolutional layers with different dilation rates, and the number of dilated convolutional layers can be selected according to the size of the target feature map.
[0090] Furthermore, in the Figure 3 specific embodiment shown, the compression path network includes 6 layers of networks, each layer of network is a nested network structure, and the internal structure of each layer of network includes 1 residual U-shaped block, including operations such as convolution and pooling with different dilation rates. Correspondingly, the expansion path network includes 5 layers of networks. The dilation rates of the dilated convolutions of the residual U-shaped blocks in the compression path network and the expansion path network can be set differently. For example, the residual U-shaped blocks in the 1st to 5th layers of the compression path network and the expansion path network only include multiple convolutional layers with a dilation rate of 1, that is, a traditional convolutional network, and a dilated convolutional layer with a dilation rate of 2. The residual U-shaped block in the 6th layer of the compression path includes 4 dilated convolutional layers with dilation rates of 1, 2, 4, and 8 respectively. Furthermore, the depths of the residual U-shaped blocks inside the 1st to 6th layers can be set to 7, 6, 5, 4, 3, and 3 in sequence.
[0091] In a further specific implementation manner, the residual U-shaped block may include an encoding layer and a decoding layer. The structure of the encoding part mainly includes a first convolution operation (conv-1) and a second convolution operation (conv-2). The first convolution operation indicates that this encoding layer uses a residual U-shaped block and 1 rectified linear unit (ReLu) operation. The second convolution operation indicates that this encoding layer uses a residual U-shaped block and 1 downsampling operation based on max pooling. After each downsampling, the size of the feature map decreases by 1 / 2. The structure of the decoding part mainly includes a third convolution operation (conv-3). The third convolution operation indicates that this decoding layer in the expansion path uses a residual U-shaped block and 1 upsampling operation. After each upsampling, the size of the feature map increases by 2 times.
[0092] In the above structure, dilated convolutions with different dilation rates are introduced, which can maintain the feature resolution while increasing the depth of the network and reduce the memory consumption of the network because the internal structure of each layer of the network is downsampled and then upsampled back to the resolution of the input features of this layer of the network.
[0093] Further, in some specific embodiments, for the 6-layer compression path network, the number of input channels of the residual U-shaped block in the first-layer network structure is 3, the number of input channels of the residual U-shaped block in the second layer can be set to 64, the number of input channels of the residual U-shaped block in the third layer can be set to 128, the number of input channels of the residual U-shaped block in the fourth layer can be set to 256, the number of input channels of the residual U-shaped block in the fifth layer can be set to 512, and the number of input channels of the residual U-shaped block in the sixth layer can be set to 512.
[0094] Further, in some specific embodiments, for the 6-layer compression path network, in the shallow networks such as the 1st to 5th layer networks, a max pooling operation with a stride of 2, i.e., a downsampling operation, is introduced after the residual U-shaped blocks contained therein to reduce the size of the image, thereby reducing the computational cost of the network.
[0095] Further, the dense feature encoding module includes a first encoding module that performs channel attention cross-guided learning on multi-level and multi-scale low-level detail features to obtain the first encoding of multi-level low-level features; a second encoding module that performs spatial attention interaction-guided learning on the low-level features to obtain multi-level high-level features; and a third encoding module that cascades the features of the first encoding module and the features of the second encoding module to obtain dense encoding features.
[0096] Among them,
[0097] Further, the input of the first encoding module of the dense feature encoding module is the deep semantic feature, and the output is the feature after channel attention cross-guided learning. In this structure, first, adaptive average pooling is applied to the high-level semantic feature to achieve feature compression and simplify the network complexity, and it is sent into a two-layer neural network. After the first layer of neurons, the number of channels is reduced to 1 / 4 of the initial value, and the activation function is the rectified linear unit ReLu. After the second layer of neurons, the number of channels is restored to the initial channel number, and the weight coefficient of channel attention cross-guided learning is calculated. Finally, the weighted feature is the output of the first encoding module.
[0098] Further, the input of the second encoding module of the dense feature encoding module is the feature after channel attention cross-guided encoding, and the output is the feature after spatial attention interaction-guided learning. In this structure, first, global max pooling and average pooling in the spatial domain are used to obtain two features with 1 channel number, and the two features are concatenated together. After passing through a layer of neural network, the number of channels is reduced to 1 / 4 of the initial value, and the activation function is the rectified linear unit ReLu. The weight coefficient of spatial attention interaction-guided learning is calculated. Finally, the weighted feature is the output of the second encoding module.
[0099] Among them, further, the input of the third encoding module of the dense feature encoding module is the first encoding module and the second encoding module, and the output is the final dense feature encoding feature. In this structure, the features after cascading the second encoding module and the second encoding module are sent into the activation function sigmoid to obtain the final dense feature encoding feature.
[0100] For example, in some specific embodiments, the dense feature encoding module includes performing channel attention cross-guided learning on the low-level detail features of the multi-level and multi-scale features obtained from the compression path to obtain the first encoding module of the dense feature encoding module Performing spatial attention interaction-guided learning on the low-level features to obtain the second encoding module of the dense feature encoding module And a third encoding module that cascades and fuses the multi-level low-level features and multi-level high-level features to obtain dense encoding features That is, the output of the final dense feature encoding module.
[0101] Among them, the input of the first encoding module of the dense feature encoding module is the deep semantic features obtained from the extended path network (decoding layer) The output is the feature after channel attention cross-guided learning In this structure, first, adaptive average pooling is applied to the high-level semantic features to achieve feature compression and simplify the network complexity, and it is sent into a two-layer neural network. After the first layer of neurons, the number of channels is reduced to 1 / 4 of the initial value, and the activation function is the rectified linear unit ReLu. After the second layer of neurons, the channels return to the initial channel number, and the weight coefficients of channel attention cross-guided learning are calculated. The finally weighted feature is the output of the first encoding module.
[0102] Further, the feature after channel attention cross-guided encoding Is used as the input of the second encoding module, and the obtained output is the feature after spatial attention interaction-guided learning In this structure, first, global maximum pooling and average pooling in the spatial dimension are used to obtain two features with 1 channel number, and the two features are concatenated together. After passing through a layer of neural network, the number of channels is reduced to 1 / 4 of the initial value, and the activation function is the rectified linear unit ReLu. The weight coefficients of spatial attention interaction-guided learning are calculated, and the finally weighted feature is the output of the second encoding module.
[0103] Further, the output features of the first encoding module and the second encoding module Are used as the input of the third encoding module, and the features after cascading the first encoding module and the second encoding module And send it into the sigmoid activation function to obtain the final dense feature encoding features.
[0104] Through the above structure, the dense feature encoding module can enhance the local context representation of small and weak targets through channel attention cross-guidance learning of low-level detailed features and spatial attention interaction guidance learning of high-level semantic features.
[0105] Under the above structure, the process of detection according to the detection model of the present invention is as follows:
[0106] S11 Obtain multi-level and multi-scale high-resolution depth features of small and weak targets in the single-frame infrared image data according to the compression path network, wherein the compression path network is a 6-layer U-shaped network, and each layer of the network contains 1 residual U-shaped block to optimize the depth multi-scale features of each layer of the network, and obtain an optimized multi-level and multi-scale feature map.
[0107] S12 Set the residual U-shaped blocks in each layer of the network to 7, 6, 5, 4, 3, and 3-layer structures respectively according to the scale of the input image of the network to enrich the feature capabilities of small infrared targets.
[0108] S13 Obtain dense coding features according to the dense feature encoding module by cross-fusing and interacting low-level features and high-level features of the optimized multi-level and multi-scale feature maps, as well as in a cascaded manner.
[0109] S14 Perform multi-level and multi-scale depth supervised decoding on the densely encoded features according to the expansion path network to obtain the detection result.
[0110] In some more specific embodiments, in S11, when the input single-frame infrared image data is X W×H×C where W, H, and C respectively represent the width, height, and number of channels of the input image, W×H×C represents the input dimension of the input layer, and the general representation of the output multi-level and multi-scale high-resolution depth feature map O is as follows:
[0111]
[0112] In the formula, F k represents the feature extracted from the kth layer in the compression path network, and σ(·) represents the sigmoid activation function.
[0113] In some more specific embodiments, in S11, expand the internal network structure of the residual U-shaped block, and the representation of the output multi-level and multi-scale high-resolution depth feature map O is as follows:
[0114]
[0115] Where U k (·)(k = 1, 2, …, K) represents the expanded form of the features learned by the k-th residual U-shaped block of the compressed path network, and K represents the maximum number of layers of the residual U-shaped blocks.
[0116] In some more specific embodiments, such as Figure 4 As shown, the multi-scale supervised encoding performed by the dense feature encoding module in S13 may include:
[0117] (1) Perform channel attention cross-guided learning on the low-level detail features to enrich the cross relationships between the channels of the low-level detail features.
[0118] For example, according to some specific embodiments, a high-level semantic feature F of H×W×C with a pixel size of H, W, and a channel number of C obtained by inputting a compressed path network may be k sent to the first encoding module of the dense encoding module.
[0119] Subsequently, the high-level semantic feature F k is subjected to an adaptive average pooling operation to obtain the feature F k ′;
[0120] Subsequently, the feature F k ′ is respectively sent into a two-layer neural network. After passing through the first layer of neurons, the number of channels of the feature is reduced to 1 / 4 of the initial value, and the activation function is the rectified linear unit ReLu. After passing through the second layer of neurons, the number of channels is restored to the initial value to achieve interactive learning between channel features;
[0121] Subsequently, the above feature is activated by a Sigmoid function to obtain a weight coefficient A1, and the feature after channel attention cross-guided learning is the result of multiplying the weight coefficient by the corresponding feature F k ′ That is
[0122] (2) Perform spatial attention interactive guidance learning on the features after channel attention cross-guided learning to enhance the perception of the high-level semantic features for target details, thereby further enriching the local context information of the target in the global context features of deep infrared small and weak targets.
[0123] For example, according to some specific embodiments, it further includes:
[0124] The feature Perform global maximum pooling and average pooling in a space to obtain two features with 1 channel each; and concatenate the two parts of the features, then after passing through a layer of network, the number of feature channels is reduced to 1 / 4 of the initial value, and the activation function is the rectified linear unit ReLu neural network to obtain the transformed features;
[0125] Pass the transformed features through a 7×7 convolutional layer and activate them with a Sigmoid function to obtain the weight coefficient A2, and the features after spatial attention interaction-guided learning That is, the result of multiplying the weight coefficient by the features Namely
[0126] The features after channel attention cross-guided learning And the features after spatial attention interaction-guided learning Are cascaded to obtain the dense feature encoded features, and they are fed into the sigmoid activation function to obtain the output of the final network.
[0127] In some more specific embodiments, the feature F k ′ is obtained by the following formula:
[0128]
[0129] Wherein, F k Represents the high-level semantic feature of the k-th layer, W and H are the width and height of the corresponding image, and i and j represent the position coordinates of the target, and its value traverses all positions in the image.
[0130] In some more specific embodiments, the weight coefficient A1 obtained by the channel attention cross mechanism of the feature F k ′ is obtained by the following formula:
[0131] A1 = σ(Β(W2δ(Β(W1F k ′)))) (4)
[0132] In the formula, δ(·) and Β(·) respectively represent the rectified linear function (ReLu) and batch normalization (BatchNormalization). And Respectively represent the excitation operators of C→C / r and C / r→C, where Represents the channel compression or stretching factor, and r is preferably 4.
[0133] In some more specific embodiments, the weight coefficient A2 of the spatial attention mechanism of the feature Is obtained by the following formula:
[0134]
[0135] In the formula, is the feature after channel transformation (C→C / r), 1×1 and 3×3 represent convolution operations. P avg and P max represent the average pooling and max pooling operators respectively.
[0136] (3) Through the feature after cross - oriented learning of channel attention and the feature
[0137] after interactive - oriented learning of spatial attention
[0138]
[0139] In the formula, F DFE represents the feature after dense coding.
[0140] S2 trains the detection model with single - frame infrared video image data.
[0141] Furthermore,
[0142] the single - frame infrared image data can be collected by an unmanned aerial vehicle optoelectronic sensor, etc.
[0143] The image data can include samples that have undergone data augmentation pre - processing, and the data augmentation methods can include mirror flipping, random cropping, random contrast change, etc.
[0144] For example, in some specific embodiments, the training set of the detection model can select an infrared image small - target detection dataset with a sample size of less than 600. This dataset is collected and labeled for specific targets, and the labeling format can adopt the general labelme labeling format.
[0145] Furthermore, in some specific embodiments, a validation set can be set during training to test the training results. For example, in the infrared image small - target detection dataset with a sample size of less than 600, the number of training set samples is set to 400, and the number of validation set samples is 100. The size of both is 320×320×3.
[0146] The training can include saving the parameters of the learned deep U - shaped network into the trained detection model, inputting the infrared small - target image to be predicted, and performing target detection.
[0147] According to some more specific embodiments, during training, the loss function of the deep U-shaped network can be set as the sum of 7 loss functions, which is different from the standard method of setting a single binary cross-entropy loss function at the end of the network. The 7 loss functions include the loss function for the output result of the 6-layer compression path network and the loss function for the feature fusion result of the last layer of the expansion path network. This not only avoids the low discriminability of weak targets or target loss due to the single output of features in the deep network, but also solves problems such as the vanishing gradient and slow convergence speed in the training of deep neural networks.
[0148] Furthermore, the loss function of the detection model can be set as:
[0149]
[0150] Wherein, is the loss function of the m-th hidden layer in the middle of the deep supervision network, corresponds to the weight of the m-th hidden layer loss function, Loss fuse is the loss function of the final fusion output, ω fuse corresponds to the weight of the fusion output loss function, and M represents the total number of outputs of the deep U-shaped supervision network.
[0151] Each loss function term in Equation (9) can be further calculated using the standard binary cross-entropy as follows:
[0152]
[0153] Wherein, i, j represent the coordinates of pixels, H, W are the dimensions of the image, p G(i,j) and p S(i,j) respectively represent the output maps of the reference pixel value and the predicted pixel value.
[0154] In some specific embodiments, the binary cross-entropy loss function adopts the method of mini-batch training, continuously looping until the loss function converges. For example, the maximum number of training epochs is set to 500, and validation is performed every 10 epochs. During validation, all parameters in the network are switched to the evaluation mode, and the IoU and nIoU of the evaluation metrics are used as reference values. Record the best IoU and nIoU values for each validation. If the IoU and nIoU of the current epoch are greater than the existing best values, then save the current learning rate and the parameters of the network, the corresponding training epoch, as well as the IoU and nIoU values.
[0155] Among them, the IoU and nIoU values can be further obtained as follows:
[0156]
[0157] In the formula, M represents the number of targets in each image in the test set, and N represents the total number of samples in the test set. Aint er and A all represent the intersection and union of the real object and the predicted object respectively. T, P, and TP represent the true label, the predicted result, and the number of correctly detected pixels respectively. and in which i and m represent the m-th object in the i-th sample.
[0158]
[0159] In the formula, represents the IoU value of each sample, and the i in T[i], P[i], and TP[i] represents the i-th sample.
[0160] S3 loads the trained detection model to realize the detection of infrared small and weak targets in the video sequence image set.
[0161] In some specific embodiments, it may further include loading the detection model generated by training on a computer platform of NVIDIA GeForce GTX1080 (8GB memory). The video sequence images can be collected by the optoelectronic sensor of the unmanned aerial vehicle.
[0162] Embodiment 1
[0163] Detect the infrared small and weak targets in the video sequence data set captured by the unmanned aerial vehicle on the computer platform of NVIDIA GeForce GTX1080 (8GB memory). The implementation process is as Figure 5 shown, and the specific process includes:
[0164] Step 1: Collect the infrared video sequence images captured by the unmanned aerial vehicle, and construct a training and test sample set after data preprocessing.
[0165] Among them, the data annotation is carried out by using the 0, 1 annotation method for distinguishing the foreground and the background. The data preprocessing includes obtaining enhanced samples by means of mirror flipping, random cropping, and random contrast change. The data after preprocessing is the real input data of the network.
[0166] Step 2: Train the deep supervised U-shaped network based on U-shaped blocks according to the specific implementation manner through the training sample set.
[0167] Specifically, it includes:
[0168] Through the compression path network of the U-shaped network, extract the multi-scale deep high-resolution target features of the input image, and at the same time enhance the global context representation of the target.
[0169] For example, in this network structure, a sample with an input size of 320×320×3 is input. After passing through a 6-layer deep U-shaped network encoding and decoding structure, feature maps of 160×160×64, 80×80×128, 40×40×256, 20×20×512, and 10×10×512 are obtained respectively, and are output after weighted fusion to obtain the final detection result.
[0170] Through the dense feature encoding module of the U-shaped network, each layer of features obtained by the compression path network is encoded to enhance the local context representation of the target.
[0171] For example, for the third layer of the network, the feature sizes of both the encoding structure and the output stage are 80×80×128, denoted as F l 3 and First, channel attention cross-encoding is performed on the output of F l 3 On this basis, spatial attention interaction encoding is performed on the encoded features, and finally the features of both parts are jointly used as the final dense encoding features.
[0172] The U-shaped network is trained and adjusted based on the set loss function.
[0173] Among them, the set loss function includes the sum of the losses of 7 parts, specifically including the output results of 6 deep supervision networks plus the result after feature fusion.
[0174] During the training process, continuous cyclic iteration is performed until the loss function converges or reaches the default cumulative number of times set by the model. For example, considering the limited training data of infrared small and weak targets, to avoid overfitting of the model, the maximum number of training times of the model is set to 500.
[0175] In the training of this embodiment, the batch size is set to 3, the initial learning rate is 1×10 -3 , and the final learning rate is 1×10 -8 , and the cosine annealing Adam iteration algorithm is used to adjust the learning rate. In addition, due to the small size of infrared small and weak targets and the use of a segmentation-based network for the detection method, the threshold for the network to output the segmentation result is set to 0, and the overlap rate with the true value is set to 0.9.
[0176] After using multi-loss accumulation, due to the small size of infrared small and weak targets, integrating the multi-level loss function after deep supervision encoding can not only avoid the low discriminability or lack of small and weak targets due to the single output of deep network features, but also solve problems such as the disappearance of training gradients and slow convergence speed of deep neural networks.
[0177] Step 3: Load the trained deep supervised U-shaped network model and use it to detect small and weak targets in the UAV infrared sequence images.
[0178] In this embodiment, the visualization result of the detection result is as Figure 6 shown, which are representative infrared sequence images, including single-target scenarios, multi-target scenarios, simple-background scenarios, and complex-background scenarios. The targets existing in the images are mainly UAVs flying in the air. The targets surrounded by the circles in the figure are the targets detected in this embodiment.
[0179] As can be seen from the figure, compared with the existing model-driven method and data-driven method, the detection model provided by the present invention can more accurately locate the measured targets in the sequence video, has high detection accuracy, good model generalization performance, and is suitable for the detection tasks of small and weak infrared targets under various complex backgrounds.
[0180] The above embodiments are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, improvements and refinements should also be regarded as within the protection scope of the present invention.
Claims
1. A single-frame image infrared small and weak target detection method based on a deep U-shaped network, characterized in that, It includes: S1: Construct a single-frame image infrared small and weak target detection model based on a deep U-shaped network; S2: Train the detection model with the labeled single-frame infrared image sample set or its enhanced sample set after enhancement processing; S3: Detect the infrared small and weak targets in the video sequence image set through the trained detection model; Among them, the detection model is constructed based on a deep supervised U-shaped network and a dense feature encoding module added to the deep supervised U-shaped network; Among them, the deep supervised U-shaped network includes a compression path network for obtaining multi-level and multi-scale image extraction features, and an expansion path network for restoring the image accuracy at multiple levels and scales. The dense feature encoding module is located between the compression path network and the expansion path network; The dense feature encoding module includes a first encoding module that performs channel attention cross-guidance learning on the low-level detail features of the multi-level and multi-scale extraction features obtained by the compression path network to obtain multi-level low-level features; a second encoding module that performs spatial attention interaction guidance learning on the low-level features to obtain multi-level high-level features; and a third encoding module that cascades the low-level features and the high-level features to obtain dense encoded features; Among them, the acquisition of the dense encoded features includes: Adaptive average pooling is performed on the features obtained from the k-th layer of the compressed path network to obtain the features F k ′; Input the feature F after the adaptive average pooling k ' into a two-layer network with weight sharing and the same number of neurons, and feed it into the Relu activation function to obtain the first transformed feature; The first transformed feature is activated by the Sigmoid function to obtain a weight coefficient A1, and further obtain the feature of channel attention cross-guided learning as follows: The features learned by cross-guiding channel attention Perform global maximum pooling and average pooling in the spatial dimension to obtain two feature maps with 1 channel number; Stitch the two feature maps, and pass the obtained stitched feature map through a neural network that performs channel transformation and contains a ReLu activation function to obtain a second transformed feature; The second transformed feature is passed through a 7×7 convolutional layer and activated by a Sigmoid function to obtain a weight coefficient A2, and further obtain the feature after spatial attention interaction-guided encoding as follows:
2. The detection method according to claim 1, characterized in that, The deep U-shaped network includes an input layer, the compression path network, the dense feature encoding module, the expansion path network, and an output layer. Among them, the compression path network includes sequentially connected multi-level and multi-scale extraction and compression modules. Except for the last extraction and compression module that only contains 1 residual U-shaped block, each of the other extraction and compression modules includes at least 1 residual U-shaped block and at least 1 downsampling layer connected thereto; the expansion path network includes sequentially connected multi-level and multi-scale expansion and restoration modules, and each expansion and restoration module includes at least 1 upsampling layer and at least 1 residual U-shaped block connected thereto; among them, the residual U-shaped block of the first extraction and compression module is connected to the input layer, the downsampling layer of the last extraction and compression module is connected to the first encoding module in the dense feature encoding module, the residual U-shaped block of each extraction and compression module in between is connected to the downsampling layer of its previous extraction and compression module, the upsampling layer of the first expansion and restoration module is connected to the third encoding module of the dense feature encoding module, the residual U-shaped block of the last expansion and restoration module is connected to the output layer, the upsampling layer of each expansion and restoration module in between is connected to the residual U-shaped block of its previous expansion and restoration module, and the downsampling layer of each extraction and compression module is also connected to the upsampling layer of the expansion and restoration module with a corresponding size, and the residual U-shaped block of each expansion and restoration module is also connected to the first encoding module in the dense feature encoding module.
3. The detection method according to claim 2, wherein The downsampling layer uses max-pooling downsampling.
4. The detection method according to claim 2, wherein The compression path network includes a 6-layer network structure, and each layer of the network constitutes one of the extraction and compression modules. Among them, each of the first to fifth layers of the network only includes multiple convolutional layers with a dilation rate of 1 and one dilated convolutional layer with a dilation rate of 2, and the sixth layer of the network includes 4 dilated convolutional layers with dilation rates of 1, 2, 4, and 8 respectively.
5. The detection method according to claim 4, wherein Among them, The depths of the residual U-shaped blocks of the first to sixth layers of the network are 7, 6, 5, 4, 3, and 3 in sequence.
6. The detection method according to claim 4, wherein Among them, The number of input channels of the residual U-shaped block contained in the first layer of the network is 3, the number of input channels of the residual U-shaped block contained in the second layer of the network is 64, the number of input channels of the residual U-shaped block contained in the third layer of the network is 128, the number of input channels of the residual U-shaped block contained in the fourth layer of the network is 256, the number of input channels of the residual U-shaped block contained in the fifth layer of the network is 512, and the number of input channels of the residual U-shaped block contained in the sixth layer of the network is 512.
7. The detection method according to claim 2, wherein The expansion path network includes a 5-layer network structure. Among them, each of the first to fifth layers of the network only includes multiple convolutional layers with a dilation rate of 1 and one dilated convolutional layer with a dilation rate of 2.
8. The detection method according to any one of claims 2-7, characterized in that, The process of target detection according to the target detection model is as follows: S11 Obtain multi-level and multi-scale depth features of small and weak targets in single-frame infrared image data according to the compression path network; S12 Perform channel attention cross-guided learning, spatial attention interaction-guided learning, and cascade fusion on the multi-level and multi-scale depth features according to the dense feature encoding module to obtain dense encoded features; S13 Perform multi-level and multi-scale depth supervised decoding on the dense encoded features according to the expansion path network to obtain detection results.
9. The detection method according to claim 8, characterized in that, S13 also includes activating the non-linear feature expression of the dense encoded features to obtain the final dense encoded features, and obtaining detection results according to the multi-scale depth supervised decoding of the final dense encoded features.
10. The detection method according to claim 8, wherein Among them, The multi-level and multi-scale depth features are obtained through the following calculation model: Among them, F k represents the feature obtained after learning by the residual U-shaped block of the k-th layer of the compression path network, and σ(·) represents the sigmoid activation function; And / or, the multi-level and multi-scale depth features are obtained through the following calculation model: Among them, U k (·)(k = 1, 2, …, K) represents the unfolded representation of the feature learned by the k-th layer residual U-shaped block in the compression path network, and K represents the total number of layers of the residual U-shaped block.
11. The detection method according to claim 10, wherein Among them, The feature F after adaptive average pooling k ' is obtained through the following calculation model: Among them, F k represents the feature extracted by the k-th layer compression path network, W and H represent the width and height of the image corresponding to this feature, and i, j represent the position coordinates of the target, and their values are all positions traversed in the image; And / or, the weight coefficient A1 is obtained through the following calculation model: Among them, δ(·) and Β(·) represent the rectified linear unit (ReLU) and batch normalization respectively, and represent the weight coefficients for reducing the number of channels to 1 / 4 of the initial value after the first layer of neurons and restoring the number of channels to the initial value after the second layer of neurons respectively. Here, C represents the number of channels, represents the channel compression or expansion factor; And / or, the weight coefficient A2 is obtained through the following calculation model: Among them, is the feature of the channel attention cross-guided learning The transformed feature after channel transformation, 1×1 and 3×3 represent convolutional operations, and P avg and P max represent the average pooling and maximum pooling operators respectively; And / or, the final dense encoded features are obtained through the following calculation model: Among them, F DFE represents the dense coding feature.
12. The detection method according to claim 8, characterized in that, During the training, set the loss function of the detection model as: Among them, represents the loss function of the m-th layer in the depth-supervised U-shaped network, corresponds to the weight of the loss function of the m-th layer, Loss fuse represents the loss function of the final fusion output, ω fuse corresponds to the weight of the fusion output loss function, and M represents the total number of layers of the depth-supervised U-shaped network.
13. The detection method according to claim 12, wherein Each loss function is calculated using the standard binary cross-entropy as follows: Among them, i and j represent the coordinates of pixels in the image, H and W are the dimensions of the image, and p G(i,j) and p S(i,j) represent the output maps of the reference pixel value and the predicted pixel value, respectively.
14. The detection method according to claim 12, characterized in that, During the training, after a certain number of training rounds, verify the training results once, and record the best IoU and nIoU values in each verification. When the IoU and nIoU values of the current round are greater than the existing best values, save the current learning rate and the parameters of the network. Among them, the IoU value is obtained through the following calculation model: Among them, M represents the number of targets in a single-frame image, N represents the total number of samples, A inter and A all respectively represent the intersection and union of the real object and the predicted object. T, P, and TP respectively represent the true label, the predicted result, and the number of correctly detected pixels. i and m respectively represent the i-th sample and the m-th object; The nIoU value is obtained through the following calculation model: Among them, represents the IoU value of each sample, and the i in T[i], P[i], and TP[i] represents the i-th sample.