Infrared small target detection method based on segmentation rolling network
By introducing a slicing rolling network and multi-scale deep supervision fusion in infrared small object detection, the problem of lack of global and remote semantic information in the existing methods is solved, and higher detection accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510001896.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-13
AI Technical Summary
The existing infrared small-object detection methods lack clear global and remote semantic information, resulting in low detection accuracy in large field of view and long-distance infrared target monitoring occasions.
A method of infrared small object detection based on slicing rolling network is proposed. By constructing an SR-Unet network, combining multi-scale deep supervised fusion and far-near fusion modules, multi-directional remote dependencies are captured and local context information is integrated.
The accuracy and generalization ability of infrared small object detection is improved, and local and remote features in the image can be extracted and utilized more effectively, enhancing the detection ability of infrared small object.
Smart Images

Figure CN119992280A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of image processing, and in particular to an infrared small target detection method based on a segmentation rolling network. Background Art
[0002] Infrared target detection is widely used in civil and military fields because of its strong concealment, good portability, ability to detect blind spots, and insensitivity to environmental and lighting conditions. However, in large-field-of-view, long-distance infrared target monitoring scenarios, due to the imaging mechanism of the infrared detection system in remote sensing applications, the target itself is small in size, has a low signal-to-noise ratio, and lacks texture details, making it easily submerged in noise and complex backgrounds.
[0003] In the absence of a reliable high-level understanding of the overall scene, accurate detection of infrared small targets is very challenging. In recent years, infrared small target detection methods based on CNN, Faster R-CNN and YOLO series have been proposed, among which the U-type neural network structure performs well.
[0004] Dai et al. proposed an asymmetric context modulation (ACM) module, which adopted an asymmetric context feature fusion strategy to capture as many small target features as possible; in addition, they proposed an attention local contrast network (ALC-Net), which integrated deep networks with model-driven and introduced a modular design of local contrast measurement to enable it to more effectively extract and utilize local contrast information in images; Wu et al. proposed UIU-Net, which embedded a small UNet into UNet to learn the local information of the target and fuse multi-scale object features, thereby improving the performance of infrared small target detection; Li et al. proposed a DNA-Net nested network, which achieved effective fusion between high-level and low-level features through its unique nested architecture and progressive interaction mechanism; Yuan et al. proposed SCTransNet, which replaced the jump connection of UNet by introducing a spatial channel module, and used cross-spatial channel information to enhance the semantic difference between the target and the background.
[0005] Although the above methods have achieved satisfactory results, they still lack clear global and long-range semantic information. Summary of the invention
[0006] Aiming at the shortcomings of the existing methods, the present invention solves the problem that the existing methods lack clear global and long-range semantic information.
[0007] The technical solution adopted by the present invention is: a method for detecting small infrared targets based on a segmented rolling network comprises the following steps:
[0008] Step 1: Collect infrared images;
[0009] Step 2: Construct the SR-Unet network. The SR-Unet network replaces the structure of the fourth layer of the encoder of the Rolling-Unet network with a combination of the first FIB module and the first LSFM module, and replaces the convolution block of the bottleneck layer with a combination of the second FIB module, the second LSFM module and the third FIB module; the convolution block of the first layer of the decoder is replaced with a combination of the third LSFM module and the fourth FIB module;
[0010] As a preferred implementation of the present invention, the FIB module consists of Conv, GELU and LayerNorm.
[0011] As a preferred implementation of the present invention, the LSFM module includes: the feature maps are respectively input into the MSORMLP module and the LIEM module, and then the feature maps are input into the FC layer after the Concat operation is performed.
[0012] As a preferred embodiment of the present invention, the MSORMLP module includes: after the input feature map is respectively input into the first to fourth SORMLP modules and residually connected with the input feature map, the output feature maps of the four SORMLP modules are concat-operated, and then the LayerNorm and FC layers are performed and added to the input feature map.
[0013] As a preferred embodiment of the present invention, the construction process of the SORMLP module is:
[0014] Divide the channel C into N groups of equal size C according to width W and height H i , where i∈[1,N];
[0015] W=H=C i (10)
[0016] C=N×C i (11)
[0017] X=N×X i (12)
[0018] The divided groups are connected in cascade form, and the input X of each group i After ORMLP in the same direction, the sum is performed and the output of each sub-head is added to the input of the next head.
[0019] As a preferred embodiment of the present invention, the LIEM module includes: the feature maps are respectively input into two channels, the first channel is subjected to a 1x1 convolution and then a Chunk operation is performed, and 3x7 convolution and 7x3 convolution are respectively input, and the two convolutions are respectively input into the Relu function and then a Concat operation is performed, and then a 1x1 convolution is input to obtain the first channel feature map; the second channel is subjected to a depth convolution, a point-by-point convolution and a BN operation to obtain a second channel feature map, and the two channel feature maps are added.
[0020] Step 3: Use multi-scale deep supervision fusion to train the SR-Unet network;
[0021] As a preferred embodiment of the present invention, multi-scale deep supervision fusion includes:
[0022] Step 31: for each output O in the decoding stage i Use 1×1 convolution to obtain the saliency map M i ;M i =f 1×1 (O i )(i=1,2,3,4,5),f 1×1 Represents 1×1 convolution;
[0023] Step 32: saliency map M i Upsample to the original image size and fuse all saliency maps to obtain M0 = Sigmoid (f 1×1 [M1, B(M2), B(M3), B(M4), B(M5)]); where [·] is channel concatenation and B represents bilinear interpolation;
[0024] Step 33: Calculate the binary cross entropy loss to evaluate the difference between the overall saliency map and the ground truth, and use the total loss function to update the weights through back propagation.
[0025] As a preferred embodiment of the present invention, the formula of the total loss function is:
[0026] l i =L BCE (B(M i ),GT)(i=2,3,4,5) (4)
[0027] l0=L BCE (M0,GT) (5)
[0028]
[0029] Among them, λ i and λ0 are used to adjust the weight of the loss function, M i is the saliency map and GT is the ground truth.
[0030] As a preferred embodiment of the present invention, an infrared small target detection system based on a segmentation rolling network includes: a memory for storing instructions executable by a processor; and a processor for executing the instructions to implement an infrared small target detection method based on a segmentation rolling network.
[0031] As a preferred embodiment of the present invention, a computer readable medium stores a computer program code, and when the computer program code is executed by a processor, an infrared small target detection method based on a segmented rolling network is implemented.
[0032] Beneficial effects of the present invention:
[0033] 1. Aiming at the limitations of the prior art, the present invention proposes an infrared small target detection network algorithm SR-Unet that combines MLP with CNN on the basis of Rolling-Unet. In order to make full use of the advantages of features at different scales, multi-scale deep supervision fusion is used to extract features at each level of the decoding stage, and the BCE loss is calculated. The training weights are updated by combining the loss to improve the performance and generalization ability of the algorithm.
[0034] 2. To solve the problem of how to capture more complete long-range dependencies, a multi-directional segmentation orthogonal rolling MLP module is designed to capture long-range dependencies in four different diagonal directions. The inputs are grouped in each diagonal direction and connected in a cascade form, making the long-range dependencies clearer and improving the performance and accuracy of the algorithm.
[0035] 3. Design the LIEM module to integrate local information through convolution in different directions to improve the accuracy of the algorithm in local feature detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is the structural diagram of SR-Unet of the present invention;
[0037] Figure 2 It is the LSFM module structure diagram of the present invention;
[0038] Figure 3 It is a schematic diagram of the existing RMLP rolling direction;
[0039] Figure 4 is a schematic diagram of the existing RMLP in the width direction;
[0040] Figure 5 This is a schematic diagram of the existing single ORMLP scroll direction;
[0041] Figure 6 It is the existing ORMLP internal structure;
[0042] Figure 7It is a schematic diagram of the MSORMLP segmentation of the present invention;
[0043] Figure 8 is a schematic diagram of the SORMLP structure of the present invention;
[0044] Fig. 9 It is the ROC curve diagram of the present invention;
[0045] Fig.10 It is a visualization result diagram of the present invention. DETAILED DESCRIPTION
[0046] The present invention is further described below in conjunction with the accompanying drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner, and therefore it only shows the components related to the present invention.
[0047] like Figure 1 As shown, a method for detecting small infrared targets based on a segmentation rolling network includes the following steps:
[0048] The U-shaped neural network structure has achieved good results in previous infrared small target detection work, and the present invention also adopts this structure for detection.
[0049] An infrared image is input, and features are extracted through maximum pooling downsampling in the encoding stage. Then, bilinear interpolation upsampling is performed in the decoding stage to restore the original resolution output. Features of the same scale are fused through skip connections.
[0050] Specifically, the first three layers of encoder and decoder are composed of two standard 3×3 convolution blocks, BatchNorm and ReLU activation functions; the fourth layer of the encoder, the bottleneck layer, and the first layer of the decoder use a combination of feature incentive block FIB (FeatureIncentive Block) and long short fusion module LSFM (Long Short Fusion Module) to perform deeper learning and extract feature information; FIB is used to process feature channel compression and expansion, encode features and control the dimension and shape of feature output, and is composed of Conv, GELU and LayerNorm; the LSFM module contains modules for capturing long-range dependencies and local information; and the weights of the model are updated through multi-scale deep supervision fusion (DS) calculation.
[0051] Among them, multi-scale deep supervision fusion DS is an effective target detection method, which improves the accuracy and stability of the model by fusing features of different scales and utilizing deep supervision, and improves the efficiency of gradient propagation and feature representation;
[0052] Specifically, for each output O in the decoding stage i Use 1×1 convolution to obtain the saliency map Mi , recorded as:
[0053] M i =f 1×1 (O i )(i=1,2,3,4,5) (1)
[0054] In the formula, f 1×1 Represents a 1×1 convolution.
[0055] Next, the low-resolution saliency map is upsampled to the original image size and all saliency maps are fused to obtain:
[0056] M0=Sigmoid(f 1×1 [M1,B(M2),B(M3),B(M4),B(M5)]) (2)
[0057] Where [·] is channel concatenation, B represents bilinear interpolation, and M0 represents the sum of all saliency maps.
[0058] Finally, the binary cross entropy (BCE) loss is calculated to evaluate the difference between the overall saliency map and the ground truth GT (GroundTruth), and combined with the total loss function to update the weights of the model for training through the back propagation algorithm.
[0059] l1=L BCE (M1,GT) (3)
[0060] l i =L BCE (B(M i ),GT)(i=2,3,4,5) (4)
[0061] l0=L BCE (M0,GT) (5)
[0062]
[0063] In the formula, λ i (i=1,2,3,4,5) and λ0 are used to adjust the weights of different loss functions; λ i and λ0 are set to 1 according to experience; l i represents the BCE loss function of each layer, and L represents the total loss function.
[0064] The present invention places LSFM in the fourth layer, bottleneck layer and first layer of the encoder of the network structure respectively. LSFM is formed by combining a multi-directional segmentation orthogonal rolling MLP module (MSORMLP) and a local information extraction module (LIEM). Figure 2As shown in the figure, LSFM can be used to integrate local context information while capturing multi-directional long-range dependencies, thereby improving the detection accuracy of the algorithm.
[0065] LSFM(X)=fc[MSORMLP(X),LIEM(X)] (7)
[0066] In the formula, fc represents the fully connected layer.
[0067] like Figure 3 As shown in the existing rolling MLP (RMLP), for a feature matrix X∈H×W×C, the spatial resolution is H×W, and the number of channels is C; where w i (i∈[1,W]) represents the width index, h j (j∈[1,H]) represents the height index, c k (k∈[1,C]) represents the channel index; the feature map of each channel layer in the feature matrix is scrolled along the width or height direction, that is, moved and cropped; as the channel index c k For each additional layer, the shift step k increases by one bit; the positive or negative value of the shift step k represents the direction of scrolling. For the width direction, a positive value of k indicates scrolling from left to right, and a negative value indicates scrolling from right to left; for the height direction, a positive value of k indicates scrolling from top to bottom, and a negative value indicates scrolling from bottom to top.
[0068] Taking the RMLP in the width direction as an example, the c1 layer feature map of the channel layer is used as a reference. After the feature maps of other channel layers are moved by the corresponding step size, the redundant parts are cropped as the missing parts and spliced to the corresponding positions, such as Figure 4 As shown; during this operation, weight-sharing channel projection is performed to encode long-range dependencies. Through weight sharing, all channel projections share a set of parameters, reducing the sensitivity of MLP to input position information.
[0069] The original feature matrix has a fixed spatial index (w i ,h j ) has only one width feature W i , after RMLP operation, different channels have different width features. Encoding the width features of the entire image can be understood as a global, unidirectional, linear receptive field. Similarly, after operating in the height direction, the long-range dependency in the height direction can be captured; RMLP performs a cyclic operation of shifting and cropping feature maps, so that the position index order on each channel is not fixed. This further reduces the sensitivity of RMLP to position.
[0070] The existing orthogonal rolling MLP (ORMLP) and dual orthogonal rolling MLP (DORMLP) can capture the long-range dependencies in the width and height directions through RMLP, and obtain RMLPs in different shift directions by changing the positive and negative shift step size k, which are denoted as In order to capture long-range dependencies in other directions, RMLP is first applied along the width direction and then along the height direction, which is equivalent to performing synchronous shift operations on the feature map in two orthogonal directions to obtain a diagonal receptive field; Figure 5 Schematic diagram when the shift step lengths k of width and height are both positive.
[0071] For input X, a RMLP operation in the width direction, and then a RMLP operation in the height direction are performed, and the GELU activation function is used in the middle to form an ORMLP in the diagonal direction ( Figure 6 ); Rolling-Unet obtains long-range dependencies in more directions by combining two complementary ORMLPs into a dual-orthogonal rolling MLP (DORMLP).
[0072] The multi-directional orthogonal rolling MLP (MORMLP) of the present invention, although DORMLP shows better results, by combining RMLPs in different width and height directions (Table 1), the long-range dependencies in the originally ignored directions can be obtained, and four different diagonal directions can be operated at the same time.
[0073] Table 1 Combinations of RMLP in different directions
[0074]
[0075]
[0076] The ORMLPs in four diagonal directions are combined to obtain MORMLP, which is used to capture long-range dependencies in multiple directions and obtain diagonal receptive fields in multiple directions. In order to alleviate the problems of gradient vanishing and gradient exploding in deep neural networks and improve the stability and performance of the model, after the operation in each direction is completed, the residual is added to combine with the output, and the residual connection is added again after the multi-directional combination.
[0077] The present invention also constructs a multi-directional segmentation orthogonal rolling MLP (MSORMLP). Although RMLP and ORMLP can make different channels have different directional features through rolling operations, for rolling operations, with the channel c1 layer as the starting reference, the original feature matrix is shifted W times in the width direction or H times in the height direction, and then it can be indexed in a fixed space (w i ,h j ) to get all width or height features, and the channel index c at this time k=W=H; after continuing to shift once, the resulting channel c k+1 Layer has the same characteristics as layer c1, such as Figure 7 shown.
[0078] That is to say, after every W or H rolling operations, different features can be aggregated to the same spatial index, which is regarded as a cycle. After the cycle ends, the rolling operation returns to the starting point. In the overall structure, the feature matrix C of LSFM is a multiple of W and H. In order to avoid redundancy caused by repeated shift operations, each ORMLP in the proposed MORMLP is split; Figure 8 As shown, the channel C is divided into N groups of C with the same size according to the width W and height H. i , that is, divide the input X into N groups of the same size X i , where i∈[1,N];
[0079] W=H=C i (10)
[0080] C=N×C i (11)
[0081] X=N×X i (12)
[0082] The channels in each group will have different directional characteristics, while different groups have similar characteristics; the divided groups are connected in cascade form; the input X of each group i After ORMLP in the same direction, the output of each sub-head is added to the input of the next head, thereby further improving the capacity of the model and encouraging feature diversity. By obtaining four SORMLPs in different diagonal directions and replacing the ORMLPs in them, we can get MSORMLP, the formula is:
[0083] Y1=ORMLP1(X1) (13)
[0084] Y i =ORMLP1(Y i-1 +X i )(1<i≤N) (14)
[0085]
[0086] Local Information Extraction Module (LIEM), although the above module can capture global linear long-range dependencies in multiple directions, it lacks local context information; in order to better integrate local information and global dependencies, the present invention designs a LIEM module, which processes the input feature map differently in width and height directions through convolution kernels (3,7) and (7,3), and connects it in parallel with depth-wise separable convolution. The formula is:
[0087] X1,X2=Chunk(f 1×1 (X)) (17)
[0088] LIEM(X)=f 1×1 [ReLU(f 3×7 (X1)),ReLU(f 7×3 (X2))]+BN(DSC(X)) (18)
[0089] In the formula, Chunk means dividing the input into multiple parts, f 3×7 and f 7×3 They represent convolutions with kernels (3,7) and (7,3) respectively, and DSC stands for depthwise separable convolution.
[0090] Datasets and experimental indicators:
[0091] The experiment used the NUAA-SIRST and NUDT-SIRST datasets, which are widely used in the field of infrared small target detection. The NUAA-SIRST dataset consists of 427 images, including infrared small target images in various actual scenes. The small targets in these images have the characteristics of small size, varied shapes and complex backgrounds. The NUDT-SIRST dataset consists of 1,327 images with large-scale and detailed annotation information, and is an important benchmark for evaluating the performance of infrared small target detection algorithms.
[0092] The previous method of dividing the training and test sets of NUAA-SIRST and NUDT-SIRST is adopted, with 50% of the images used for training and the rest for testing; the input images are normalized, randomly cropped into 256×256 image blocks, and pre-processed by random flipping and rotation.
[0093] Evaluation Metrics:
[0094] IoU (intersection over union) is a pixel-level evaluation indicator used to evaluate the contour description ability of the algorithm. In order to more comprehensively evaluate the quality of the predicted box, normalization processing is performed on the basis of IoU to obtain nIoU (normalized intersection / union); and F1-measure (F1) is used to comprehensively evaluate the quality of the model, which are expressed as:
[0095]
[0096] Where N is the total number of samples, T and P represent the number of ground truth positive pixels and predicted positive pixels, TP represents the number of true positive pixels, FP represents the number of false positive pixels, and FN represents the number of false negative pixels. [i] represents the i-th sample.
[0097] We also introduce two target-level evaluation indicators, detection probability P d and false alarm rate F a ;P d It refers to the probability that the detection algorithm can correctly identify and detect the target; F a It refers to the probability that the detection algorithm mistakenly identifies and reports the existence of a target in an infrared image when the target does not exist. The formula is:
[0098]
[0099] Where N pred To correctly predict the number of targets, N all is the number of all targets, P false is the wrongly predicted target pixel, P all for all pixels in the image.
[0100] In order to reflect the computational efficiency of the algorithm, two indicators, parameter quantity and FLOPs, are introduced; among them, parameter quantity refers to the total number of all learnable parameters in the model, which is an important indicator reflecting the complexity and capacity of the algorithm. FLOPs is the number of floating-point operations, which is usually used to measure the computational complexity of the algorithm and can reflect the computational speed and resource requirements of the model.
[0101] Training environment and parameter configuration:
[0102] The experiment was carried out on a server with a GPU of Nvidia Geforce RTX 3090 and a video memory of 24GB. The Pytorch framework was used for testing and experiments, and the environment configuration was Python3.9.19+Pytorch2.0.1+CUDA11.8; the network of the present invention does not use pre-trained weights for training, each image is normalized and then randomly cropped into image blocks of 256×256, and the training data is enhanced by random flipping and rotation to avoid overfitting; the number of training rounds is set to 800 in the experiment, and the batchsize is set to 8; the Adam optimizer is used to train the model, the initial learning rate is 0.001, and the cosine annealing learning rate scheduler is used, and the minimum learning rate is 0.00001.
[0103] Experiment and result analysis:
[0104] In order to evaluate the performance of the model of the present invention, it is compared with five deep learning-based infrared small target detection algorithms, namely ACM, ALC-Net, UIU-Net, DNA-Net, and SCTransNet. The comparison process is carried out under the same data set and experimental conditions, and the parameters and configurations in their original papers are used for training.
[0105] Table 2 Comparative experiment
[0106]
[0107] Table 2 shows the IoU (%), nIoU (%), F1 (%), Pd (%) and Fa (10 -6 ) results; in the three indicators of IoU, nIoU and F1, the SR-Unet proposed in this paper is ahead of other algorithms on both public datasets; compared with the latest SCTransNet, IoU is improved by 3.38%, nIoU is improved by 2.19%, F1 is improved by 2.16%, Pd is improved by 0.38%, and Fa is reduced by 6.58×10 -6 ; This shows that the SR-Unet of the present invention has a strong ability to retain the target contour and can identify the pixel-level information difference between the target and the background; although the Pd and Fa of SR-Unet on NUAA-SIRST are slightly worse than those of DNA-Net, other indicators have excellent performance; This shows that the model of the present invention has achieved better overall performance.
[0108] The algorithm of the present invention is compared with the existing algorithms in terms of parameter quantity and FLOPs. The results show (Table 3) that the algorithm of the present invention has better parameter quantity and FLOPs while ensuring the detection effect.
[0109] Table 3 Parameters and FLOPs
[0110]
[0111] In order to comprehensively evaluate the model and get a more intuitive feeling effect, the ROC curve is used to show the trade-off between the true positive rate (TPR) and the false positive rate (FPR); Fig. 9 ROC curves of different algorithms on NUAA-SIRST (A) and NUDT-SIRST (B); Fig.10 The following are some visualization results of different algorithms on the data sets. The blue box represents the detection result, the red box represents the missed detection, and the yellow box represents the multiple detection. It can be seen that the performance of the algorithm of the present invention is better than other algorithms on both data sets, and it can detect the contour of small infrared targets more accurately, with a lower probability of missed detection and multiple detection.
[0112] Ablation experiment:
[0113] To prove the effectiveness of each module, seven groups of ablation experiments were conducted on SR-Unet on the NUAA-SIRST dataset. In each group of experiments, only one component was changed, the variables were strictly controlled, and the experiments were repeated many times to reduce the influence of accidental factors and ensure the reliability of the results.
[0114] Taking Rolling-Unet as the baseline, a combination of Rolling-Unet's core modules DORMLP and DSC was used in the first group of experiments; in order to explore the differences between the LIEM and DSC models proposed in the present invention in extracting local information, the DSC in the first group was replaced by the LIEM module as the second group of experiments; on the basis of the previous group, multi-scale deep supervision fusion (DS) was further added as the third group of experiments; in order to demonstrate the ability of different modules to capture long-range dependencies, the fourth and seventh groups of experiments replaced the original DORMLP with MORMLP and MSORMLP in turn on the basis of the third group of experiments; in order to reflect the impact of different modules on performance, the DS and LIEM modules of the seventh group of experiments were removed respectively as the fifth and sixth groups of experiments.
[0115] Table 4 Ablation experiment
[0116]
[0117] Table 4 shows the ablation results of IoU (%), F1 (%) and Pd (%) of SR-Unet on NUAA-SIRST; the seventh group of experiments using MSORMLP improved IoU by 1.97% and 1.5% respectively, F1 by 1.25% and 0.95% respectively, and Pd by 0.76% and 0.38% respectively compared with the third and fourth groups of experiments using DORMLP and MORMLP; this shows that MSORMLP can better capture long-range dependencies during detection; after removing the DS and LIEM modules in the fifth and sixth groups of experiments, all three indicators decreased; SR-Unet as a whole improved IoU by 3.88% compared with the baseline, and improved F1 and Pd by 2.49% and 2.28% respectively.
[0118] Aiming at the problem that it is difficult to extract and fuse local features and long-range dependencies in infrared small target detection, the present invention proposes an infrared small target detection network SR-Unet based on Rolling-Unet; by introducing multi-scale deep supervision fusion and constructing a near-far fusion module, the ability to detect infrared small targets is effectively improved; experimental results show that compared with existing methods, the method of the present invention can better capture long-range dependencies and fuse them with local context information. While ensuring excellent parameter quantity and FLOPs, the detection accuracy is higher than that of the current mainstream methods, and there are fewer missed detections and false alarms. The impact of different modules on performance is analyzed through ablation experiments. The correctness of the algorithm was verified through multiple groups of experiments, and the accuracy of small target detection in infrared images was improved.
[0119] Based on the above ideal embodiments of the present invention, the relevant staff can make various changes and modifications without departing from the technical concept of the present invention through the above description. The technical scope of the present invention is not limited to the contents of the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. A method for detecting small infrared targets based on segmented rolling networks, characterized in that: The following steps are involved: Step 1: Collect infrared images; Step 2: Construct the SR-Unet network. The SR-Unet network replaces the structure of the fourth layer of the encoder of the Rolling-Unet network with a combination of the first FIB module and the first LSFM module, and replaces the convolution block of the bottleneck layer with a combination of the second FIB module, the second LSFM module and the third FIB module; the convolution block of the first layer of the decoder is replaced with a combination of the third LSFM module and the fourth FIB module; Step 3: Use multi-scale deep supervision fusion to train the SR-Unet network.
2. The infrared small target detection method based on segmentation and rolling network according to claim 1 is characterized in that: The FIB module consists of Conv, GELU and LayerNorm.
3. The infrared small target detection method based on segmentation and rolling network according to claim 1 is characterized in that: The LSFM module includes: the feature maps are respectively input into the MSORMLP module and the LIEM module, and then the Concat operation is performed and then input into the FC layer.
4. The infrared small target detection method based on segmentation and rolling network according to claim 3 is characterized in that: The MSORMLP module includes: the input feature maps are respectively input into the first to fourth SORMLP modules and then residually connected with the input feature maps, the output feature maps of the four SORMLP modules are concat-operated, and then the LayerNorm and FC layers are performed and then added to the input feature maps.
5. The infrared small target detection method based on segmentation and rolling network according to claim 4 is characterized in that: The construction process of the SORMLP module is: Divide the channel C into N groups of equal size C according to width W and height H i , where i∈[1,N]; W=H=C i (10) C=N×C i (11) X=N×X i (12) The divided groups are connected in cascade form, and the input X of each group i After ORMLP in the same direction, the sum is performed and the output of each sub-head is added to the input of the next head.
6. The infrared small target detection method based on segmentation and rolling network according to claim 3 is characterized in that: The LIEM module includes: the feature maps are input into two channels respectively, the first channel is subjected to 1x1 convolution and then Chunk operation, and 3x7 convolution and 7x3 convolution are input respectively. The two convolutions are input into Relu function and then Concat operation is performed, and then 1x1 convolution is input to obtain the first channel feature map; the second channel is subjected to depth convolution, point-by-point convolution and BN operation to obtain the second channel feature map, and the two channel feature maps are added.
7. The infrared small target detection method based on segmentation and rolling network according to claim 1 is characterized in that: Multi-scale deep supervision fusion includes: Step 31: for each output O in the decoding stage i Use 1×1 convolution to obtain the saliency map M i ;M i =f 1×1 (O i )(i=1,2,3,4,5),f 1×1 Represents 1×1 convolution; Step 32: saliency map M i Upsample to the original image size and fuse all saliency maps to obtain M0 = Sigmoid (f 1×1 [M1, B(M2), B(M3), B(M4), B(M5)]); where [·] is channel concatenation and B represents bilinear interpolation; Step 33: Calculate the binary cross entropy loss to evaluate the difference between the overall saliency map and the ground truth, and use the total loss function to update the weights through back propagation.
8. The infrared small target detection method based on segmentation and rolling network according to claim 7 is characterized in that: The formula for the total loss function is: l i =L BCE (B(M i ),GT)(i=2,3,4,5) (4) l0=L BCE (M0,GT) (5) Among them, λ i and λ0 are used to adjust the weight of the loss function, M i is the saliency map and GT is the ground truth.
9. Infrared small target detection system based on segmentation rolling network, characterized in that: include: a memory for storing instructions executable by a processor; A processor, used for executing instructions to implement the infrared small target detection method based on segmentation rolling network as described in any one of claims 1 to 8.
10. A computer readable medium storing computer program code, characterized in that: When the computer program code is executed by a processor, the infrared small target detection method based on the segmentation rolling network as described in any one of claims 1 to 8 is implemented.