An infrared dim small target segmentation method based on improved U-Net
Patent Information
- Application Number
- CN202410546876.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-06
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-05-06
AI Technical Summary
[0002]随着红外成像技术的发展,红外探测器凭借着隐蔽性好、全天候工作、成像清楚、抗电磁干扰能力强、结构简单等优点,被广泛的应用于军事、医疗、安防以等领域,由于在红外图像中目标尺寸较小、信号微弱,同时背景环境复杂,信噪比低,这导致了检测难度较高
[0027]本发明的方法在U-Net编码器中采用基于Haar小波变换的下采样方式,能够在特征提取过程中保留更多细节信息,减少信息丢失。
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence for infrared weak target detection technology, specifically to an infrared weak target segmentation method based on an improved U-Net. Background Technology
[0002] With the development of infrared imaging technology, infrared detectors, due to their advantages such as good concealment, all-weather operation, clear imaging, strong resistance to electromagnetic interference, and simple structure, have been widely used in military, medical, and security fields. However, because targets in infrared images are small in size, have weak signals, and the background environment is complex with a low signal-to-noise ratio, detection is quite difficult. Most existing methods fail to achieve ideal results and have many problems. In practical applications, they often generate a large number of false alarms, and the small size of the target also makes it easy to miss detections. In recent years, with the increasing complexity of weak target detection application scenarios, the requirements for weak target detection algorithms have become increasingly stringent. Therefore, researching robust and robust infrared weak target detection algorithms has become an urgent problem to solve. With the development of computer hardware and related theories of artificial intelligence, the use of deep learning methods to implement infrared weak target detection has become a development trend. Researching weak target detection algorithms that meet the needs of practical engineering has significant practical implications. Summary of the Invention
[0003] The present invention aims to provide an infrared weak target segmentation method based on an improved U-Net. This method can better focus on multi-dimensional useful information, enhance the expressive power of features, and improve the ability to segment weak targets.
[0004] The technical solution of the present invention is as follows:
[0005] The infrared weak eye segmentation method based on the improved U-Net includes the following steps:
[0006] A. Construct an improved U-Net network, which includes an encoding network and a decoding network;
[0007] The encoding network includes five sets of multi-scale residual modules (MSRB) connected in sequence, with a downsampling module (HWD) between two adjacent sets of MSRB; the decoding network includes four upsampling operations, after which the signal is processed sequentially by a triple attention module (TRA) and a multi-scale residual module (MSRB).
[0008] B. Train the improved U-Net network model to obtain the trained improved U-Net network model;
[0009] C. Based on the trained deep neural network, the infrared image input encoding network sequentially passes through each group of multi-scale residual modules (MSRB) and downsampling modules (HWD). The processing procedure in the multi-scale residual module (MSRB) is as follows:
[0010] The input result is first processed by a 3×3 convolution, and the resulting feature map is divided into four parts according to channels. The first part of the feature map is not processed and is used as a feature reuse layer. The second, third and fourth parts of the feature map are processed by depth convolution with dilation rates of 1, 3 and 5 respectively. The three dilated convolution results are concatenated with the first part of the feature map. Then, a 1×1 convolution unit is used to fuse the concatenated feature map between channels. At the same time, a residual connection is made with the input feature map to obtain the output result.
[0011] D. The processing results of each multi-scale residual module (MSRB) are input into the decoding network, which performs feature recovery to obtain the final prediction result.
[0012] The processing procedure in the HWD downsampling module is as follows:
[0013] The input results are divided into two paths. The first path is processed by a low-pass filter H0 and then downsampled by a factor of two to obtain the first downsampled result. The second path is processed by a high-pass filter H1 and then downsampled by a factor of two to obtain the second downsampled result.
[0014] The first downsampling result is divided into two branches. The first branch is processed by a low-pass filter H0 to obtain the low-frequency component A, and the second branch is processed by a high-pass filter H1 to obtain the horizontal component H.
[0015] The second downsampling result is divided into two branches. The first branch is processed by a low-pass filter H0 to obtain the vertical component V, and the second branch is processed by a high-pass filter H1 to obtain the diagonal component D.
[0016] The low-frequency component A, horizontal component H, vertical component V, and diagonal component D are concatenated. The concatenated result is then processed by a 1×1 convolutional layer, batch normalization, and ReLU activation function to adjust the number of channels, resulting in the output.
[0017] The processing procedure in the decoding network is as follows:
[0018] The output of the fifth multi-scale residual module MSRB in the coding network is upsampled by two times and then concatenated with the output of the fourth multi-scale residual module MSRB in the coding network using the cat function. After being processed by the triple attention module TRA and the multi-scale residual module MSRB in sequence, the first TRA+MSRB processing result is obtained.
[0019] The first TRA+MSRB processing result is upsampled by two times and then spliced and fused with the output of the third multi-scale residual module MSRB in the coding network. After being processed by the triple attention module TRA and the multi-scale residual module MSRB in sequence, the second TRA+MSRB processing result is obtained.
[0020] The second TRA+MSRB processing result is upsampled by two times and then spliced and fused with the output of the second multi-scale residual module MSRB in the coding network. After being processed by the triple attention module TRA and the multi-scale residual module MSRB in sequence, the third TRA+MSRB processing result is obtained.
[0021] The third TRA+MSRB processing result is upsampled by two times and then spliced and fused with the output of the first multi-scale residual module MSRB in the coding network. After being processed by the triple attention module TRA and the multi-scale residual module MSRB in sequence, the fourth TRA+MSRB processing result is obtained.
[0022] The fourth TRA+MSRB processing result is fused between channels using 1×1 convolutional units to obtain the final prediction result.
[0023] The processing procedure in the Triple Attention Module (TRA) is as follows: The input feature tensor is processed in three branches. The first branch rotates the tensor of size C×H×W counterclockwise by 90° around the H axis, resulting in a tensor of size W×H×C. The rotation result is then processed sequentially using Z-pooling, 7×7 convolution, batch normalization, and Sigmoid activation function to generate attention weights. These weights are then multiplied by the input result. Finally, the tensor is rotated 90° clockwise around the H axis to restore its initial shape, resulting in the first branch result.
[0024] After the second branch is rotated 90° counterclockwise around the W axis, the resulting tensor has a size of H×C×W. The rotation result is then subjected to Z-pooling, 7×7 convolution, batch normalization, and Sigmoid activation function calculation in sequence to generate attention weights. These weights are then multiplied by the input result. Finally, the tensor is rotated 90° clockwise around the W axis to restore it to its initial shape, thus obtaining the result of the second branch.
[0025] The third branch sequentially performs Z-pool calculation, 7×7 convolution, batch normalization, and Sigmoid activation function calculation to generate attention weights. These weights are then multiplied by the input results to obtain the output results, which is the result of the third branch.
[0026] Finally, the results of the first branch, the second branch, and the third branch are added together and the average is taken to obtain the output result.
[0027] The method of this invention employs a downsampling approach based on Haar wavelet transform in the U-Net encoder, which can retain more detailed information and reduce information loss during feature extraction.
[0028] The present invention designs a multi-scale residual unit as the basic building block of U-Net to more efficiently acquire multi-scale contextual information and improve feature extraction efficiency; finally, a triple attention mechanism is added to the decoder structure to enable the network to focus on useful information in multiple dimensions and enhance the expressive power of features.
[0029] This invention uses the NUDT-SIRST infrared weak target dataset for model training, and then evaluates and predicts the trained model on the test set. Experiments show that the algorithm invented in this paper can effectively improve the ability to segment weak targets. Attached Figure Description
[0030] Figure 1 Here is a flowchart of the improved U-Net algorithm in Example 1;
[0031] Figure 2 This is a schematic diagram of the structure of the multi-scale residual module MSRB in Example 1;
[0032] Figure 3 This is a schematic diagram of the structure of the Haar wavelet transform-based downsampling module HWD in Example 1;
[0033] Figure 4 This is a schematic diagram of the triple attention module TRA in Example 1;
[0034] Figure 5 This example compares the segmentation results of the method in Example 1 with other existing U-Net segmentation methods on the NUDT-SIRST dataset. Detailed Implementation
[0035] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0036] Example 1
[0037] The infrared weak eye segmentation method based on the improved U-Net in this embodiment includes the following steps:
[0038] A. As Figure 1 As shown, an improved U-Net network is constructed, which includes an encoding network and a decoding network.
[0039] The encoding network includes five sets of multi-scale residual modules (MSRBs) connected in sequence, with a downsampling module (HWD) between adjacent sets of MSRBs; see the diagram of the MSRB module structure. Figure 2The HWD structure diagram of the downsampling module is shown below. Figure 3 The decoding network consists of four upsampling operations. After each upsampling, the data is processed sequentially by a Triple Attention Module (TRA) and a Multi-Scale Residual Module (MSRB). The structure diagram of the Triple Attention Module (TRA) is shown below. Figure 4 ;
[0040] B. Train the improved U-Net network model, randomly select 60% of the samples for training the model and 40% of the samples for testing, with a sample resolution of 256. Use the trained network model to predict the test samples to achieve automatic segmentation of infrared weak targets.
[0041] C. Based on the trained deep neural network, the 256×256 infrared image is input into the encoding network and sequentially processed by each group of multi-scale residual modules (MSRB) and downsampling modules (HWD). The processing procedure in the multi-scale residual module (MSRB) is as follows:
[0042] The input result is first processed by a 3×3 convolution, and the resulting feature map is divided into four parts according to channels. The first part of the feature map is not processed and is used as a feature reuse layer. The second, third and fourth parts of the feature map are processed by depth convolution with dilation rates of 1, 3 and 5 respectively. The three dilated convolution results are concatenated with the first part of the feature map. Then, a 1×1 convolution unit is used to fuse the concatenated feature map between channels. At the same time, a residual connection is made with the input feature map to obtain the output result.
[0043] The processing procedure in the HWD downsampling module is as follows:
[0044] The input results are divided into two paths. The first path is processed by a low-pass filter H0 and then downsampled by a factor of two to obtain the first downsampled result. The second path is processed by a high-pass filter H1 and then downsampled by a factor of two to obtain the second downsampled result.
[0045] The first downsampling result is divided into two branches. The first branch is processed by a low-pass filter H0 to obtain the low-frequency component A. The second branch is processed by a high-pass filter H1 to obtain the horizontal component H.
[0046] The second downsampling result is divided into two branches. The first branch is processed by a low-pass filter H0 to obtain the vertical component V, and the second branch is processed by a high-pass filter H1 to obtain the diagonal component D.
[0047] The low-frequency component A, horizontal component H, vertical component V, and diagonal component D are all of equal magnitude. These are concatenated, and the concatenated result is then processed through a 1×1 convolutional layer, batch normalization, and ReLU activation function to adjust the number of channels, resulting in the output.
[0048] D. The processing results of each multi-scale residual module (MSRB) are input into the decoding network, and the final prediction result is obtained after decoding by the decoding network.
[0049] The processing procedure in the decoding network is as follows:
[0050] The output of the fifth multi-scale residual module MSRB in the coding network is upsampled by two times and then concatenated and fused with the output of the fourth multi-scale residual module MSRB in the coding network. After being processed by the triple attention module TRA and the multi-scale residual module MSRB in sequence, the first TRA+MSRB processing result is obtained.
[0051] The first TRA+MSRB processing result is upsampled by two times and then spliced and fused with the output of the third multi-scale residual module MSRB in the coding network. After being processed by the triple attention module TRA and the multi-scale residual module MSRB in sequence, the second TRA+MSRB processing result is obtained.
[0052] The second TRA+MSRB processing result is upsampled by two times and then spliced and fused with the output of the second multi-scale residual module MSRB in the coding network. After being processed by the triple attention module TRA and the multi-scale residual module MSRB in sequence, the third TRA+MSRB processing result is obtained.
[0053] The third TRA+MSRB processing result is upsampled by two times and then spliced and fused with the output of the first multi-scale residual module MSRB in the coding network. After being processed by the triple attention module TRA and the multi-scale residual module MSRB in sequence, the fourth TRA+MSRB processing result is obtained.
[0054] The fourth TRA+MSRB processing result is fused between channels using 1×1 convolutional units to obtain the final prediction result.
[0055] like Figure 4As shown, the processing procedure in the Triple Attention Module (TRA) is as follows: The input feature map needs to go through three branch structures to output the result. For the input tensor X of size C×H×W, the interaction between the channel dimension C and the height dimension H is established in the first branch. First, the input tensor X is rotated 90° counterclockwise around the H axis to obtain a tensor of size W×H×C. After Z-pooling, a tensor of size 2×H×C is obtained. Then, convolution, batch normalization, and Sigmoid activation function are performed to generate attention weights, which are multiplied with the input. Finally, the tensor is rotated 90° clockwise around the H axis to restore the initial shape. The second branch establishes an interaction between the channel dimension C and the width dimension W. The input tensor is rotated 90° counterclockwise around the W axis to obtain a tensor with a shape and size of H×C×W. After Z-pooling, a tensor of 2×C×W is obtained. Then, convolution, batch normalization, and sigmoid activation function calculation are performed to generate attention weights, which are then multiplied by the input. Finally, the tensor is rotated 90° clockwise around the W axis to restore its initial shape. The third branch performs Z-pooling on the input tensor X to obtain a tensor with a shape and size of 2×H×W. Then, convolution, batch normalization, and sigmoid activation function calculation are performed to generate attention weights, which are then multiplied by the input to obtain the output result. Finally, the output results of the three branches are added together and the average is taken to obtain the output feature.
[0056] Z-pooling calculates both max pooling and average pooling on the 0th dimension of the tensor, then concatenates the results to reduce the dimensionality to two. This operation allows the tensor to retain rich feature information while reducing computational cost, and can be expressed as: where 0d represents the 0th dimension of the tensor.
[0057] Z-pool(x) = [Maxpool] 0d (x),Avgpool 0d (x)] (1)
[0058] Example 2
[0059] Using the method in Example 1, this invention, along with several existing semantic segmentation methods based on improved U-Net, achieves weak target segmentation on the public dataset NUDT-SIRST. Specific evaluation data can be found in [link to example]. Figure 5 .
[0060] like Figure 5 As shown, ten representative backgrounds from the NUST-SIRST dataset were selected for testing. Figure 5As can be seen, in background 1 (buildings), ResUnet missed detections, while other algorithms correctly segmented the target point. In background 2 (ground), U-Net had a few false positives, ResUnet and ResUnet++ had a large number of false positives, UCTransNet and U-Net3+ each had one false positive, and DeepLab3+ did not detect any target. The method of this invention correctly segmented the target. In background 3 (sky) with shallow clouds, U-Net had false positives, and ResUnet++ only detected one target point, with pixel descriptions differing significantly from the true target. ResUnet, UCTransNet, U-Net3+, DeepLab3+, and the method of this invention all correctly detected the target. In background 4 (grassland), due to the low contrast between the target and the background, all algorithms performed poorly. U-Net only detected one target with two false positives, ResUnet, ResUnet++, and DeepLab3+ did not detect the target, and UCTransNet and U-Net3+ only detected one target. The method of this invention achieved the best segmentation result, accurately segmenting both targets. Background 5 consists of water and sky, with minimal clutter, allowing all algorithms to correctly segment the target. The method of this invention achieves optimal results in target contour description.
Claims
1. A method for segmenting weak infrared objects based on an improved U-Net, characterized in that, Includes the following steps: A. Construct an improved U-Net network, which includes an encoding network and a decoding network; The encoding network includes five sets of multi-scale residual modules (MSRB) connected in sequence, with a downsampling module (HWD) between two adjacent sets of MSRB; the decoding network includes four upsampling operations, after which the signal is processed sequentially by a triple attention module (TRA) and a multi-scale residual module (MSRB). B. Train the improved U-Net network model to obtain the trained improved U-Net network model; C. Based on the trained deep neural network, the infrared image input encoding network sequentially passes through each group of multi-scale residual modules (MSRB) and downsampling modules (HWD). The processing procedure in the multi-scale residual module (MSRB) is as follows: The input result is first processed by a 3×3 convolution, and the resulting feature map is divided into four parts according to channels. The first part of the feature map is not processed and is used as a feature reuse layer. The second, third and fourth parts of the feature map are processed by depth convolution with dilation rates of 1, 3 and 5 respectively. The three dilated convolution results are concatenated with the first part of the feature map. Then, a 1×1 convolution unit is used to fuse the concatenated feature map between channels. At the same time, a residual connection is made with the input feature map to obtain the output result. D. The processing results of each multi-scale residual module (MSRB) are input into the decoding network, and the features are recovered by the decoding network to obtain the final prediction result.
2. The infrared weak eye segmentation method based on improved U-Net as described in claim 1, characterized in that: The processing procedure in the HWD downsampling module is as follows: The input results are divided into two paths. The first path is processed by a low-pass filter H0 and then downsampled by a factor of two to obtain the first downsampled result. The second path is processed by a high-pass filter H1 and then downsampled by a factor of two to obtain the second downsampled result. The first downsampling result is divided into two branches. The first branch is processed by a low-pass filter H0 to obtain the low-frequency component A, and the second branch is processed by a high-pass filter H1 to obtain the horizontal component H. The second downsampling result is divided into two branches. The first branch is processed by a low-pass filter H0 to obtain the vertical component V, and the second branch is processed by a high-pass filter H1 to obtain the diagonal component D. The low-frequency component A, horizontal component H, vertical component V, and diagonal component D are concatenated. The concatenated result is then processed by a 1×1 convolutional layer, batch normalization, and ReLU activation function to adjust the number of channels, resulting in the output.
3. The infrared weak eye segmentation method based on improved U-Net as described in claim 1, characterized in that: The processing procedure in the decoding network is as follows: The output of the fifth multi-scale residual module MSRB in the coding network is upsampled by two times and then concatenated and fused with the output of the fourth multi-scale residual module MSRB in the coding network. After being processed by the triple attention module TRA and the multi-scale residual module MSRB in sequence, the first TRA+MSRB processing result is obtained. The first TRA+MSRB processing result is upsampled by two times and then spliced and fused with the output of the third multi-scale residual module MSRB in the coding network. After being processed by the triple attention module TRA and the multi-scale residual module MSRB in sequence, the second TRA+MSRB processing result is obtained. The second TRA+MSRB processing result is upsampled by two times and then spliced and fused with the output of the second multi-scale residual module MSRB in the coding network. After being processed by the triple attention module TRA and the multi-scale residual module MSRB in sequence, the third TRA+MSRB processing result is obtained. The third TRA+MSRB processing result is upsampled by two times and then spliced and fused with the output of the first multi-scale residual module MSRB in the coding network. After being processed by the triple attention module TRA and the multi-scale residual module MSRB in sequence, the fourth TRA+MSRB processing result is obtained. The fourth TRA+MSRB processing result is fused between channels using 1×1 convolutional units to obtain the final prediction result.
4. The infrared weak eye segmentation method based on improved U-Net as described in claim 1, characterized in that: The processing procedure in the Triple Attention Module (TRA) is as follows: The input feature tensor is processed in three branches. The first branch rotates the tensor of size C×H×W counterclockwise by 90° around the H-axis, resulting in a tensor of size W×H×C. Z-pooling, 7×7 convolution, batch normalization, and Sigmoid activation function are performed sequentially to generate attention weights. These weights are then multiplied by the input result. Finally, the tensor is rotated 90° clockwise around the H-axis to restore its initial shape, resulting in the first branch result. After the second branch is rotated 90° counterclockwise around the W axis, the resulting tensor has a size of H×C×W. Z-pooling, 7×7 convolution, batch normalization, and Sigmoid activation function calculation are performed sequentially to generate attention weights. These weights are then multiplied by the input result. Finally, the tensor is rotated 90° clockwise around the W axis to restore it to its initial shape, thus obtaining the result of the second branch. The third branch sequentially performs Z-pool calculation, 7×7 convolution, batch normalization, and Sigmoid activation function calculation to generate attention weights. These weights are then multiplied by the input results to obtain the output results, which is the result of the third branch. Finally, the results of the first branch, the second branch, and the third branch are added together and the average is taken to obtain the output result.
Citation Information
Patent Citations
Cervical cell segmentation network training method, cervical cell segmentation method and device
CN114240949A
Single-frame image infrared weak and small target detection method based on deep U-shaped network
CN115311508A