Infrared small target detection method
By adopting branch fusion-based feature extraction module and attention mechanism-based feature fusion module in infrared small object detection network, the problem of insufficient feature extraction and fusion capabilities in the prior art is solved, and the detection accuracy and efficiency are significantly improved.
Patent Information
- Application Number
- CN202510060334.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-15
AI Technical Summary
The existing infrared small object detection technology has low detection accuracy due to weak feature extraction and feature fusion capabilities in complex backgrounds.
The feature extraction module based on branch fusion and the feature fusion module based on attention mechanism are used to build an infrared small object detection network model, and the feature extraction and fusion capabilities are improved through iterative training.
It effectively improves the accuracy and detection efficiency of infrared small object detection and optimizes the feature fusion effect.
Smart Images

Figure CN120070849A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and relates to an infrared small target detection method, which can be used in fields such as geographical target detection in aviation and intelligent traffic monitoring. Background Art
[0002] Infrared small targets usually refer to targets with a size of no more than 7×7 pixels in infrared images. These targets usually have few pixels and lack obvious shape, texture and other information, and can only be recognized through features such as gray level and position. In view of the above characteristics, the core of infrared small target detection lies in identifying small and low signal-to-noise ratio targets from complex infrared images. Therefore, how to perform target enhancement, background suppression and feature extraction in infrared small target detection are the key factors affecting the detection accuracy. There are mainly two ideas for infrared small target detection: detection-based and segmentation-based. Among them, the segmentation-based idea performs relatively well at the pixel index level.
[0003] For example, Xi'an Zhongke Lide Infrared Technology Co., Ltd. discloses an infrared small target detection method in its patent document "An Infrared Small Target Detection Method Based on Multi-Scale Feature Fusion" (Patent Application No.: CN202410894900.9, Publication No.: CN118918308A). This invention uses a path aggregation network framework to collect feature information at different levels into a collection module, and uses average pooling and bilinear interpolation operations to unify the scale of feature maps; through a fusion module, it fuses features at each level to generate global information, and uses an SE module to adaptively adjust feature weights, enhance important features, and suppress secondary features; through an injection module, it combines global information with local features to enhance the detection ability of small targets. However, its defect is that the feature fusion scheme adopted by this invention splices high-level feature information and low-level feature information after unified processing. The fusion method is relatively single, and a large amount of noise is introduced during fusion, lacking fine-grained modeling of target local information, thus resulting in a still low accuracy of small target detection by the model. Summary of the Invention
[0004] The purpose of the present invention is to propose an infrared small target detection method for solving the technical problem of low detection accuracy in the prior art due to weak feature extraction ability and feature fusion ability under complex backgrounds in view of the above deficiencies of the existing technology.
[0005] To achieve the above purpose, the technical solutions adopted by the present invention include the following steps:
[0006] (1) Obtain a training sample set and a test sample set:
[0007] Preprocess the M infrared small target images obtained, which include multiple scene categories, and label the small targets in each preprocessed infrared image. Then, form a training sample set with K preprocessed infrared images and their labels, and form a test sample set with the remaining M - K preprocessed infrared images and their labels, where
[0008] (2) Construct an infrared small target detection network model O:
[0009] Construct an infrared small target detection network model O including a cascaded input layer, an encoder Encoder, a decoder Decoder, and a detection head Head; where Encoder includes three feature extraction modules based on branch fusion in cascade, and a downsampling module based on spatial depth is loaded between adjacent feature extraction modules; Decoder includes two feature fusion modules based on attention mechanism with an upsampling module based on transposed convolution loaded at the input end in cascade; the output end of the last feature extraction module is cascaded with the input end of the first upsampling module, and the output ends of the first two feature extraction modules are also cascaded with the input ends of the feature fusion modules at the same level;
[0010] (3) Iteratively train the infrared small target detection network model O:
[0011] Iteratively train the infrared small target detection network model O through the training sample set to obtain a trained infrared small target detection network model O * ;
[0012] (4) Obtain the infrared small target detection results:
[0013] Use the test sample set as the input of the trained infrared small target detection network model O * to perform forward propagation and obtain the infrared small target detection results corresponding to M - K test samples.
[0014] Compared with the prior art, the present invention has the following advantages:
[0015] In the process of training the infrared small target detection model and obtaining the infrared small target detection results in the present invention: the downsampling module based on spatial depth in Encoder retains all feature information in the channel dimension while performing downsampling, and the feature extraction module based on branch fusion can effectively extract multi-level feature information of each sample without introducing new computational complexity; the feature fusion module based on attention mechanism in Decoder can strengthen the high-level feature information and low-level feature information respectively, fuse the feature expressions of different levels, optimize the feature fusion effect, and effectively improve the accuracy and detection efficiency of infrared small target detection. Description of the Drawings
[0016] Figure 1 This is the implementation flowchart of the present invention;
[0017] Figure 2 This is the structural schematic diagram of the infrared small target detection network of the present invention;
[0018] Figure 3 This is the structural schematic diagram of the feature extraction module in the present invention;
[0019] Figure 4 This is the structural schematic diagram of the feature fusion module in the present invention. Specific embodiments
[0020] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] Refer to Figure 1 , the present invention includes the following steps:
[0022] Step 1) Obtain a training sample set and a test sample set:
[0023] Obtain M infrared small target images including multiple scene categories in the NUAA-SIRST and IRSTD-1K datasets for preprocessing, and label the small targets in each preprocessed infrared image. Then, form a training sample set with K preprocessed infrared images and their labels, and form a test sample set with the remaining M - K preprocessed infrared images and their labels. In this embodiment, M = 1427 and K = 1141.
[0024] Among them, the specific steps for preprocessing the obtained M infrared small target images including multiple scene categories are as follows:
[0025] (1a) Use data augmentation techniques on the obtained dataset, perform random horizontal flipping on each original sample image with a probability of 50%, perform random equal-proportion scaling and Gaussian blur operations using bilinear interpolation, so as to increase the diversity of training data, improve the generalization ability of the model, and reduce the risk of model overfitting;
[0026] (1b) In order to meet the model input requirements, randomly crop the images after data augmentation to obtain images with a size of 512×512;
[0027] (1c) Use normalization techniques to transform the pixel value distribution of the cropped images into a distribution with a mean of 0 and a standard deviation of 1, so that the various features of the images are unified to the same scale, accelerating the model convergence speed. Among them, the means of the RGB three channels are 0.485, 0.456, and 0.406 respectively, and the standard deviations are 0.229, 0.224, and 0.225 respectively.
[0028] Step 2) Construct an infrared small target detection network model O, whose structure is as Figure 2 shown:
[0029] Construct an infrared small target detection network model O including a cascaded input layer, an encoder Encoder, a decoder Decoder, and a detection head Head; where the Encoder includes three cascaded feature extraction modules based on branch fusion, and a downsampling module based on spatial depth is loaded between adjacent feature extraction modules; the Decoder includes two feature fusion modules based on attention mechanism with a transposed convolutional layer-based upsampling module loaded at the input end of the cascade; the output end of the last feature extraction module is cascaded with the input end of the first upsampling module, and the output ends of the first two feature extraction modules are also cascaded with the input ends of the feature fusion modules at the same level, where:
[0030] The input layer includes a plurality of cascaded CBR modules, and each CBR module includes a stacked convolutional layer, a batch normalization layer, and a ReLU function layer.
[0031] The encoder Encoder, the structure of the feature extraction module in which is as Figure 3 shown, includes three parallel batch normalization layers and ReLU function layers cascaded therewith, where 3×3 convolutional layers and 1×1 convolutional layers are stacked in front of the first and second batch normalization layers respectively; the downsampling module includes a stacked spatial-depth layer and a non-strided convolutional layer.
[0032] The decoder Decoder, the upsampling module in which adopts a transposed convolutional layer structure; the structure of the feature fusion module is as Figure 4 shown, includes a channel attention block and a spatial attention block, and a computing unit loaded between the two modules, where the channel attention block includes parallel global average pooling layers, ReLU function layers, and Sigmoid function layers that are all stacked, and fully connected layers loaded between the global average pooling layer and the ReLU function layer and between the ReLU function layer and the Sigmoid function layer; the spatial attention block includes parallel average pooling layers and max pooling layers and a convolutional layer and a batch normalization layer cascaded therewith.
[0033] Step 3) Iteratively train the infrared small target detection network model O:
[0034] Iteratively train the infrared small target detection network model O with a training sample set to obtain a trained infrared small target detection network model O * , and the specific steps are as follows:
[0035] (3a) Initialize the iteration number as a, the maximum iteration number as A, A≥1000, and the weight and bias parameters in the infrared small target detection network model O at the a-th iteration are w a respectively.a , b a , and let a = 0. In this example, A = 1000;
[0036] (3b) Forward propagate P training samples randomly selected from the training sample set as the input of the infrared small target detection network model O to obtain the detection results of P infrared small targets. In this example, P = 16;
[0037] (3b1) The input size of the input layer is 3×512×512. Increase the number of channels of each training sample to 8 to obtain the initial feature map F 0p , where each channel can extract different local features, enabling the network to learn richer and more complex feature representations;
[0038] (3b2) The encoder Encoder performs primary feature extraction on each initial feature map, and performs secondary feature extraction on the feature map F containing primary features 1p after downsampling, and then performs tertiary feature extraction on the feature map F containing secondary features after downsampling, and obtains the feature map F containing tertiary features 2p ; 3p ;
[0039] Among them, the feature extraction adopts a structure with three parallel branches. The convolution kernel sizes in the first two branches are 3×3 convolutional layers and 1×1 convolutional layers respectively, and there is no convolution kernel in the third branch. The size of the convolution kernel is positively correlated with the size of the receptive field. The feature extraction module sets different sizes of convolution kernels, enabling the network to extract different fine-grained features under different receptive fields and strengthening the feature extraction ability.
[0040] When inferring the infrared small target detection model, the three branches in the feature extraction module can be fused into a branch including stacked 3×3 convolutional layers, batch normalization layers, and ReLU function layers according to the structural reparameterization technology, reducing the computational amount and improving the detection efficiency.
[0041] In addition, the downsampling operation on the feature map F containing primary features 1p is the same as the downsampling operation on the feature map F containing secondary features 2p , which is as follows:
[0042] (3b21) Label the pixels sequentially in the height and width dimensions of the feature map;
[0043] (3b22) Concatenate the pixels with odd labels in both the height dimension and the width dimension to form a vector c 1 ;
[0044] (3b23) Concatenate the pixels with odd labels in the height dimension and even labels in the width dimension to form a vector c2 ;
[0045] (3b24) Stitch the pixels with an even label in the height dimension and an odd label in the width dimension to form vector c 3 ;
[0046] (3b25) Stitch the pixels with even labels in both the height dimension and the width dimension to form vector c 4 ;
[0047] (3b26) For vector c 1 , c 2 , c 3 and c 4 Stitch them in the channel dimension and obtain the downsampling result through a non-strided convolution operation.
[0048] The downsampling operation of the present invention is different from the traditional pooling layer or strided convolution layer. It retains all the feature information in the feature map in the channel dimension, can effectively avoid the defect of losing image detail information in the traditional downsampling process, and is friendly enough to small targets with only weak detail features in infrared images.
[0049] (3b3) The decoder Decoder performs upsampling on the feature map F 3p and performs first-level feature fusion with the feature map F 2p to obtain a feature map Y containing second- and third-level features, and then performs upsampling on the feature map Y and performs second-level feature fusion with the feature map F 1p containing first-level features to obtain a feature map T containing features of each level p ;
[0050] Both the upsampling of the feature map F 3p and the upsampling of the feature map Y use transposed convolution;
[0051] During feature fusion, considering that high-level features contain more semantic information and low-level features contain more spatial location information, the present invention uses a channel attention mechanism to strengthen the semantic information in high-level features and a spatial attention mechanism to strengthen the spatial location information in low-level features, enhancing the feature fusion ability of the model.
[0052] During the process of performing upsampling on the feature map F 3p and performing first-level feature fusion with the feature map F 2p to obtain the feature map Y:
[0053] Y = L×L 1 + F 2p × H 1 + F 2p × H 1 × L 2
[0054] Among them, L is the feature map F 3p The result of upsampling, L 1 is the output of the low-level feature map of the channel attention block, L 2 is the output of the spatial attention block, H 1 is the output of the high-level feature map of the channel attention block.
[0055] After upsampling the feature map Y and the feature map F 1p Perform secondary feature fusion calculation in the same way as the above method. The input of the high-level feature map in the feature fusion module is the feature map F 2p , and the input of the low-level feature map is the feature map F 3p The result of upsampling to obtain the feature map T p ;
[0056] (3b4) The detection head Head changes the number of channels of the feature map T p to 1 through convolution operation to obtain the detection result;
[0057] (3c) Adopt the IoU loss function, and calculate the loss value of the network model through each infrared small target detection result and its corresponding true label Then adopt the gradient descent method, through The weights and bias parameters w a , b a are updated to obtain the infrared small target detection network model O of this iteration a ;
[0058] The loss value of the network model described The calculation formula is:
[0059]
[0060] Among them, X p represents the region of the model prediction target corresponding to the p-th sample of the training dataset, and Y p represents the region of the true target corresponding to the p-th sample of the training dataset. |·| represents the absolute value operation, ∩ and ∪ represent the intersection and union operations respectively, and ∑ represents the summation operation.
[0061] Calculate through the chain rule The partial derivatives of the weight parameter w a and the bias parameter b a and and And according to Update w a , b a :
[0062]
[0063] Among them, w a ′ represents the updated result of the bias parameter w a , b a ′ represents the updated result of b a , α represents the learning rate, represents taking the partial derivative of w a , represents taking the partial derivative of b a ;
[0064] (3d) Determine whether a≥A holds. If so, obtain the trained infrared small target detection network model O * , otherwise, set a = a + 1, O a = O, and execute step (3b).
[0065] Step 4) Obtain the infrared small target detection result:
[0066] Use the test sample set as the input of the trained infrared small target detection network model O * to perform forward propagation, and obtain the infrared small target detection results corresponding to M - K test samples.
[0067] Next, in combination with simulation experiments, the technical effects of the present invention will be described:
[0068] 1. Experimental conditions and content:
[0069] The hardware platform for the simulation experiment is: the processor is an Intel(R) Core i5 - 8400 CPU with a main frequency of 2.8GHz, the memory is 32GB, and the graphics card is an NVIDIA GeForce RTX 4060Ti. The software platform for the simulation experiment is: the Ubuntu20.04 operating system, the Python version is 3.7, and the Pytorch version is 1.7.1.
[0070] The present invention and two existing infrared small target detection methods are compared and simulated in terms of three indicators: the normalized intersection over union nIoU, the detection rate P d and the false alarm rate F a . The results are shown in Table 1.
[0071] The normalized intersection over union nIoU is used to describe the degree of overlap between the predicted value output by the algorithm and the true value output at the pixel level. Its calculation formula is:
[0072]
[0073] Among them, TP represents the number of positive samples predicted as positive, T represents the number of samples predicted correctly, and P represents the total number of positive samples.
[0074] Detection rate P d It represents the proportion of correctly discriminated samples in the discrimination of all positive and negative samples, and its calculation formula is:
[0075]
[0076] Among them, TP represents the number of positive samples predicted as positive, FP represents the number of negative samples predicted as positive, TN represents the number of negative samples predicted as negative, and FN represents the number of negative samples predicted as positive.
[0077] False alarm rate F a It represents the proportion of samples predicted as positive but actually negative, and its calculation formula is:
[0078]
[0079] Among them, TP represents the number of positive samples predicted as positive, and FP represents the number of negative samples predicted as positive.
[0080] 2. Analysis of experimental results:
[0081] Table 1 Simulation comparison results
[0082]
[0083] Referring to Table 1, compared with the prior art, in the NUAA - SIRST dataset and the IRSTD - 1K dataset, the normalized intersection over union nIoU and the detection rate P d indicators have both improved, and the false alarm rate F a indicator has decreased, indicating that the detection accuracy has been effectively improved.
Claims
1. A method for detecting small infrared targets, characterized in that: The steps include: (1) Obtain training sample set and test sample set: The acquired M infrared small target images including multiple scene categories are preprocessed, and the small targets in each preprocessed infrared image are labeled. Then, the K preprocessed infrared images and their labels form a training sample set, and the remaining MK preprocessed infrared images and their labels form a test sample set, where (2) Construct infrared small target detection network model O: An infrared small target detection network model O including a cascaded input layer, an encoder, a decoder and a detection head is constructed; wherein the encoder includes three cascaded feature extraction modules based on branch fusion, and a downsampling module based on spatial depth is loaded between adjacent feature extraction modules; the decoder includes two feature fusion modules based on attention mechanism loaded with a deconvolution-based upsampling module at the cascade input end; the output end of the last feature extraction module is cascaded with the input end of the first upsampling module, and the output ends of the first two feature extraction modules are also cascaded with the input end of the feature fusion module at the same level; (3) Iteratively train the infrared small target detection network model O: The infrared small target detection network model O is iteratively trained through the training sample set to obtain the trained infrared small target detection network model O * ; (4) Obtain infrared small target detection results: The test sample set is used as the trained infrared small target detection network model O * The input is forward propagated to obtain the infrared small target detection results corresponding to MK test samples.
2. The method according to claim 1, characterized in that The preprocessing of the acquired M infrared small target images including multiple scene categories in step (1) is implemented by: Data enhancement is performed on each infrared small target image, and each infrared small target image after data enhancement is randomly cropped. Then, each cropped infrared small target image is normalized to obtain M preprocessed infrared small target images.
3. The method according to claim 1, characterized in that The infrared small target detection network model O described in step (2), wherein: The input layer includes multiple cascaded CBR modules, each of which includes a stacked convolutional layer, a batch normalization layer, and a ReLU function layer; The feature extraction module of the encoder includes three parallel batch normalization layers and a ReLU function layer cascaded with them, where the first and second batch normalization layers are stacked with a 3×3 convolution layer and a 1×1 convolution layer respectively; the downsampling module includes a stacked space-depth layer and a non-strided convolution layer; The decoder includes an upsampling module with a transposed convolutional layer structure; the feature fusion module includes a channel attention block and a spatial attention block and a computing unit loaded between the two modules, wherein the channel attention block includes parallel and stacked global average pooling layers, ReLU function layers and Sigmoid function layers, and fully connected layers loaded between the global average pooling layer and the ReLU function layer and between the ReLU function layer and the Sigmoid function layer; the spatial attention block includes parallel average pooling layers and maximum pooling layers and convolutional layers and batch normalization layers cascaded therewith.
4. The method according to claim 3, characterized in that: The iterative training of the infrared small target detection network model O described in step (3) is implemented as follows: (3a) Initialize the number of iterations to a, the maximum number of iterations to A, A ≥ 1000, the infrared small target detection network model O of the ath iteration a The weight and bias parameters in are w a , b a , and let a=0; (3b) P training samples randomly selected from the training sample set are used as the input of the infrared small target detection network model O for forward propagation to obtain the detection results of P infrared small targets; (3c) The IoU loss function is used to calculate the network model loss value through each infrared small target detection result and its corresponding true label. Then, the gradient descent method is used, through For weight and bias parameters w a , b a Update to get the infrared small target detection network model O of this iteration a ; (3d) Determine whether a≥A holds. If so, obtain the trained infrared small target detection network model O * , otherwise, let a=a+1, O a =O, and execute step (3b).
5. The method according to claim 4, characterized in that The detection results of the P infrared small targets described in step (3b), wherein the detection result of the p-th infrared small target is obtained by: (3b1) The input layer adds channels to each training sample and performs preliminary feature extraction to obtain the initial feature map F of the pth training sample. 0p ; (3b2) Encoder Encodes each initial feature map F 0p Perform primary feature extraction and perform feature map F containing primary features 1p After downsampling, secondary feature extraction is performed, and then the feature map F containing the secondary features is 2p After downsampling, three-level feature extraction is performed to obtain a feature map F containing three-level features. 3p ; (3b3) Decoder Decodes the feature map F 3p After upsampling, the feature map F 2p Perform first-level feature fusion to obtain a feature map Y containing second- and third-level features, and then upsample the feature map Y and merge it with the feature map F 1p Perform secondary feature fusion to obtain a feature map T containing three levels of features p ; (3b4) Detection head Head to feature map T p Perform detection and obtain the detection result of the pth small infrared target.
6. The method according to claim 5, characterized in that Step (3b2) of the feature graph F containing the primary features 1p To perform downsampling, the implementation steps are: The feature map F containing the first-level features is respectively 1p The image is divided into equal intervals, and the four divided vectors are concatenated in sequence in the channel dimension and then non-stepped convolution is performed to obtain the downsampled feature map.
7. The method according to claim 5, characterized in that The feature graph Y containing the second-level and third-level features described in step (3b3) is obtained by: Y=L×L1+F 2p ×H1+F 2p ×H1×L2 Among them, L is the feature map F 3p As a result of upsampling, L1 and H1 are the low-level and high-level feature map outputs of the channel attention block, respectively, and L2 is the output of the spatial attention block.
8. The method according to claim 4, characterized in that The network model loss value described in step (3c) The calculation formula is: Among them, X p Indicates the region of the model prediction target corresponding to the pth sample in the training data set, Y p represents the region of the true target corresponding to the pth sample in the training data set, |·| represents the absolute value operation, ∩ and ∪ represent the intersection and union operations respectively, and ∑ represents the summation operation.
9. The method according to claim 4, characterized in that The weight and bias parameters w described in step (3c) a , b a To update, use the chain rule, and the update formulas are: Among them, w a ′ represents the bias parameter w a The updated result, b a ′ represents b a The update result, α represents the learning rate, express For w a Take the partial derivative, express For b a Take the partial derivative.
Citation Information
Patent Citations
Infrared small target detection method based on multi-scale feature fusion
CN118918308A
Small target detection method based on full-link multi-scale fusion network
CN117197441A
Small target detection method based on curvature attention mechanism
CN117218426A
Small target detection method based on graph attention network
CN117274744A
Medical image segmentation model establishment method based on harmonic attention and medical image segmentation method
CN118967714A