Infrared small target segmentation method and device based on gating unit and multi-scale convolutional network
By adopting a method based on gated units and multi-scale convolutional networks in the infrared small-objective segmentation technology, the problems of low segmentation accuracy and poor robustness in the existing technology are solved, and higher segmentation accuracy and lower false alarm rate are achieved.
Patent Information
- Application Number
- CN202510271279.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-08
- Publication Date
- 2025-06-20
AI Technical Summary
The existing infrared small-target segmentation technology has low accuracy when facing complex backgrounds and noise, insufficient fusion of multi-scale features, weak global context modeling capabilities, and unstable training process.
Using a method based on gating units and multi-scale convolutional networks, multi-scale features are extracted through pyramid visual Transformer, gated units are designed to dynamically adjust the contribution weight of cross-layer features, introduce multi-loss functions and deep supervision mechanisms, and optimize the gradient propagation path.
The segmentation accuracy and robustness of infrared small targets were significantly improved, with F1-score being increased by 5.3%, and the false alarm rate being reduced by 8.7%.
Smart Images

Figure CN120182293A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to an infrared small target segmentation method and device based on a gated unit and a multi-scale convolutional network. Background Art
[0002] Infrared small target segmentation is a key technology in infrared image processing, which plays an important role in traffic monitoring, maritime rescue, and target warning. The goal is to extract target regions with low signal-to-noise ratio and small size (usually less than 9×9 pixels) from complex backgrounds. Traditional methods mainly rely on manually designed features (such as local contrast and gradient statistics), but they have the following defects: 1) Manually designed features are sensitive to background clutter and noise and are difficult to adapt to complex and variable infrared scenes; 2) Small targets lack texture and shape information, resulting in missed detections or missegmentations easily occurring in methods based on thresholds or region growing; 3) The scales and shapes of infrared targets vary significantly in different scenes, making the detection problem quite challenging.
[0003] In recent years, methods based on convolutional neural networks (CNNs) have significantly improved the segmentation performance, but there are still the following bottlenecks: 1) Insufficient multi-scale feature fusion: Conventional U-shaped networks directly splice encoder and decoder features through skip connections, without fully considering the semantic differences of features at different levels, resulting in information redundancy or loss during the fusion process; 2) Weak global context modeling ability: The local receptive fields of CNNs limit the capture of long-range dependencies, while infrared small targets usually have similar radiation characteristics to the background and need to rely on global context for differentiation; 3) Unstable training process: Deep networks are vulnerable to the vanishing gradient problem. Especially when the proportion of target pixels is extremely low (usually less than 0.1%), the conventional cross-entropy loss function is difficult to effectively optimize the model. Summary of the Invention
[0004] To solve the above technical problems, the purpose of the present invention is to provide an infrared small target segmentation method based on a gated unit and a multi-scale convolutional network with good anti-background interference ability, improved small target segmentation accuracy, and robustness.
[0005] Another purpose of the present invention is to provide an infrared small target segmentation device based on a gated unit and a multi-scale convolutional network with good anti-background interference ability, improved small target segmentation accuracy, and robustness.
[0006] The technical solution adopted by the present invention is as follows: On the one hand, an infrared small target segmentation method based on a gated unit and a multi-scale convolutional network is provided, including the following steps: S1. Obtain an infrared image dataset and perform preprocessing, which includes size normalization, random cropping to H×W pixels, and data augmentation operations; S2. Construct a gated unit and multi-scale convolutional network; S3. Train the network using a deep supervision mechanism that includes multiple loss functions; S4. Deploy the trained model for target segmentation.
[0007] In step S2, the gated unit and multi-scale convolutional network uses an encoder-decoder structure, including: S2a. The encoder uses a Pyramid Vision Transformer (PVTv2) architecture to extract four-level multi-scale features {F1, F2, F3, F4}, and the four-level features are transmitted to the decoder through cross-level connections; S2b. The cross-level connection realizes cross-layer feature adaptive propagation through a gated unit; S2c. The decoder obtains multi-scale features through multi-scale parallel convolution.
[0008] The data augmentation in step S1 includes: S11. Geometric transformation: Random horizontal flipping (probability p = 0.5), central cropping (ratio ρ ∈ [0.8, 1]), fixed ratio scaling (ratio s ∈ 0.75, 1.25)); S12. Radiation transformation: Adding Gaussian noise with a standard deviation σ in the range of 0 to 0.1, Gaussian blur (kernel size k ∈ 3, 5, 7); S13. Elastic transformation: The affine matrix M ∈ R 2×3 Performs deformation with the constraint that |det(M)| ≥ 0.9.
[0009] The gated unit in step S2b specifically includes the following steps: S2b1. Perform adaptive average pooling (AAP) processing on the i-level feature map F i After 3×3 depthwise separable convolution (stride 1, padding 1), two-fold upsampling, and batch normalization (BN) processing, generate an intermediate feature map T i ; S2b2. Perform adaptive max pooling (AMP) processing on the j-level feature map F j After 3×3 depthwise separable convolution (stride 1, padding 1) and batch normalization processing, generate an intermediate feature map T j ; S2b3. Combine the intermediate feature map T i with T jPerform element-wise addition, and after processing through the ReLU6 activation function, perform 1×1 convolution operation, squeeze-and-excitation attention module (channel compression ratio of 16), batch normalization, and Sigmoid function mapping in sequence to generate the three-dimensional fusion weight Q; S2b4. Dynamically allocate the original feature map according to the weight Q: Multiply F j element-wise with (1 - Q) to obtain the optimized feature map Y j , and at the same time multiply F i element-wise with Q to obtain the optimized feature map Y i ; where: (i, j) ∈ (1, 2), (3, 4).
[0010] The multi-scale parallel decoder in step S2c specifically includes the following steps: where Y i is the optimized feature map output by the i-th level gating unit, i ∈ 1, 2, 3, 4; The initial input Y1 is downsampled by a factor of 2, followed by 3×3 convolution (stride 1, padding 1), batch normalization (BN), and ReLU6 activation, then followed by 3×3 convolution (stride 1, padding 1), BN, ReLU6 activation, and upsampling by a factor of 2, and concatenated with Y1 through channels to generate X1; X1 passes through the first multi-scale parallel convolution to generate Z1; Z1 is upsampled by a factor of 2 and concatenated with Y2 through channels to generate X2; X2 passes through the second multi-scale parallel convolution to generate Z2; Z2 is upsampled by a factor of 2 and concatenated with Y3 through channels to generate X3; X3 passes through the third multi-scale parallel convolution to generate Z3; Z3 is upsampled by a factor of 2 and concatenated with Y4 through channels to generate X4; X4 passes through the fourth multi-scale parallel convolution to generate Z4; The multi-scale parallel convolution specifically includes the following steps: (a) Large receptive field branch (Block1): Cascade two depthwise separable convolution blocks, and the convolution block includes: depthwise separable convolution with a 7×7 convolution kernel, BN layer, and ReLU6 activation function; (b) Medium receptive field branch (Block2): Cascade two depthwise separable convolution blocks, and the convolution block includes: depthwise separable convolution with a 3×3 convolution kernel, BN layer, and ReLU6 activation function; (c) Channel correction branch (Block3): 1×1 point convolution; (e) The final output feature satisfies: Z = (Block1(X) ⊙ Block2(X)) ⊙ Block3(X) where ⊙ represents element-wise multiplication, and the final feature dimension remains unchanged.
[0011] The deep supervision mechanism in step S4 includes: where Z iRepresents the output prediction map of each decoding stage; Adopt an adaptive weight adjustment strategy to fuse the output prediction maps of each decoding stage, and the final output is: Among them, ω i is a learnable hierarchical weight parameter, Up(·) represents bilinear upsampling to the original resolution, ω i is automatically optimized through backpropagation and satisfies The multi-loss function in the step S4 is the weighted sum of the binary cross-entropy loss and the confidence loss.
[0012] On the other hand, an infrared small target segmentation device is provided, including the following modules: An infrared sensor module for collecting the original infrared image; A preprocessing module configured to perform size normalization and cropping on the collected infrared image to a preprocessing of H×W pixels; A processor module loaded with any one of the above-mentioned trained methods based on the gated unit and the multi-scale convolutional network, including: a graphics processing unit, and the video memory capacity ≥ (total number of network parameters × 4 bytes) × 1.2; An output module for generating an infrared image with a target segmentation mask and a pixel-level confidence map.
[0013] The beneficial effects of the method of the present invention are: 1) Use the Pyramid Vision Transformer (pvtv2) as the encoder to capture multi-scale global context features through the self-attention mechanism; 2) Design a gated unit to dynamically adjust the cross-layer feature contribution weight and suppress the irrelevant background response; 3) Design a parallel multi-scale convolutional module to fully obtain the multi-scale features of the feature image and improve the multi-scale adaptation ability of the network; 4) Introduce a hybrid loss function and a deep supervision mechanism to optimize the gradient propagation path through adaptive weight allocation; Experiments show that compared with the existing technology MDvsFA on the mixed dataset of the public datasets NUAA-SIRST, NUDT-SIRST and IRSTD-1k, the F1-score of this method is increased by 5.3%, and the false alarm rate is reduced by 8.7%, significantly improving the segmentation accuracy and robustness of small targets.
[0014] The beneficial effects of the device of the present invention are as follows: First, an infrared image is collected by the infrared sensor module and then preprocessed. Then, the image is loaded into the neural network model trained in steps S1, S2, S3, and S4 to output the labeled image. The infrared image segmentation is performed using a gated unit and a multi-scale convolutional network, which significantly improves the segmentation accuracy and robustness of small targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings: Figure 1 It is an overall flowchart of an infrared small target segmentation method based on a gated unit and a multi-scale convolutional network; Figure 2 It is a schematic structural diagram of the network in the present invention Figure 3 It is a schematic structural diagram of the gated unit; Figure 4 It is a schematic structural diagram of the parallel convolution module; Figure 5 It is a system structure diagram of an embodiment of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0017] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.
[0018] Refer to Figures 1-4 , an infrared small target segmentation method based on a gated unit and a multi-scale convolutional network, includes the following steps: S1. Obtain an infrared image dataset and perform preprocessing, including size normalization, random cropping to H×W pixels, and data augmentation operations; specifically, in this embodiment, the infrared image dataset is from NUAA-SIRST, NUDT-SIRST, and IRSTD-1k. These datasets contain 427, 1327, and 1000 images respectively. Each image is normalized and randomly cropped into an image of 256×256 pixels; finally, 5648 pictures and masks are obtained for model training after selection, and they are divided into a training set and a test set according to a ratio of 8:2. S2. Construct a gated unit and a multi-scale convolutional network; S3. Train the network using a deep supervision mechanism that includes multiple loss functions; Specifically, in this embodiment, the training model uses the RAdam optimizer for training, with an initial learning rate of 0.001, and adopts a cosine annealing strategy to gradually reduce the learning rate to 1×10^-5. The batch size is set to 16, and it is trained for 200 epochs. Training on a single Nvidia GeForce 3090 GPU with 24G of video memory can complete the model training; S4. Deploy the trained model for target segmentation.
[0019] In step S2, the encoding-decoding structure is used based on the gated unit and the multi-scale convolutional network, including: S2a. The encoder uses the Pyramid Vision Transformer (PVTv2) architecture to extract four-level multi-scale features {F1, F2, F3, F4}, and the four-level features are transmitted to the decoder through cross-level connections; Specifically, in this embodiment, the pvt_v2_b2 version is used, and the number of channels of the four-level features it extracts are {64, 128, 320, 512} respectively. This article does not exclude that using other versions of the PVTv2 structure can achieve better results; S2b. The cross-level connection realizes cross-layer feature adaptive propagation through the gated unit; S2c. The decoder obtains multi-scale features through multi-scale parallel convolution.
[0020] The data augmentation in step S1 includes: S11. Geometric transformation: Random horizontal flipping (probability p = 0.5), central cropping (ratio ρ ∈ [0.8, 1]), fixed ratio scaling (ratio s ∈ 0.75, 1.25)); S12. Radiation transformation: Adding Gaussian noise with a standard deviation σ in the range of 0 to 0.1, Gaussian blur (kernel size k ∈ 3, 5, 7); S13. Elastic transformation: The affine matrix M ∈ R 2×3 Deforms, with the constraint condition |det(M)| ≥ 0.9.
[0021] The gated unit in step S2b specifically includes the following steps: S2b1. Perform adaptive average pooling on the i-th level feature map F i After performing 3×3 depthwise separable convolution (stride 1, padding 1), two-fold upsampling, and batch normalization processing, an intermediate feature map T is generated i ; S2b2. For the j-th level feature map F jPerform adaptive max pooling processing, and generate an intermediate feature map T after 3×3 depthwise separable convolution (stride 1, padding 1) and batch normalization processing j ; S2b3. Add the intermediate feature map T i and T j element-wise, and after processing through the ReLU6 activation function, perform 1×1 convolution operation, squeeze-and-excitation attention module (channel compression ratio 16), batch normalization, and Sigmoid function mapping in sequence to generate a three-dimensional fusion weight Q; S2b4. Perform dynamic allocation on the original feature map according to the weight Q: Multiply F j and (1 - Q) element-wise to obtain an optimized feature map Y j , and at the same time, multiply F i and Q element-wise to obtain an optimized feature map Y i ; where: (i, j) ∈ (1, 2), (3, 4).
[0022] The multi-scale parallel decoder in the step S2c specifically includes the following steps: where Y i is the optimized feature map output by the i-th level gating unit, i ∈ 1, 2, 3, 4; The initial input Y1 is downsampled by 2 times, followed by 3×3 convolution (stride 1, padding 1), batch normalization (BN), and ReLU6 activation, then followed by 3×3 convolution (stride 1, padding 1), BN, ReLU6 activation, and upsampling by 2 times, and is concatenated with Y1 through channels to generate X1; X1 passes through the first multi-scale parallel convolution to generate Z1; Z1 is upsampled by 2 times and concatenated with Y2 through channels to generate X2; X2 passes through the second multi-scale parallel convolution to generate Z2; Z2 is upsampled by 2 times and concatenated with Y3 through channels to generate X3; X3 passes through the third multi-scale parallel convolution to generate Z3; Z3 is upsampled by 2 times and concatenated with Y4 through channels to generate X4; X4 passes through the fourth multi-scale parallel convolution to generate Z4; The multi-scale parallel convolution specifically includes the following steps: (a) Large receptive field branch (Block1): Cascade two depthwise separable convolution blocks, and the convolution block includes: depthwise separable convolution with a convolution kernel of 7×7, a BN layer, and a ReLU6 activation function; (b) Medium receptive field branch (Block2): Cascade two depthwise separable convolution blocks, and the convolution block includes: depthwise separable convolution with a convolution kernel of 3×3, a BN layer, and a ReLU6 activation function; (c) Channel correction branch (Block3): 1×1 point convolution; (e) The final output feature satisfies: Z = (Block1(X) ⊙ Block2(X)) ⊙ Block3(X), where ⊙ represents element-wise multiplication, and the final feature dimension remains unchanged. Here, the multi-branch convolution is used in the decoder to enhance the network's segmentation effect on small targets. The deep supervision mechanism in step S4 includes: where Z i represents the output prediction map at each decoding stage; An adaptive weight adjustment strategy is adopted to fuse the output prediction maps at each decoding stage, and the final output is: where ω i is a learnable hierarchical weight parameter, Up(·) represents bilinear upsampling to the original resolution, and ω i is automatically optimized through backpropagation and satisfies Specifically, for image segmentation in this embodiment, the number of feature channels of the final output Z needs to be reduced to 1 and passed through the Sigmoid activation function before the loss function can be calculated.
[0023] The multi-loss function in step S4 is the weighted sum of the binary cross-entropy loss and the confidence loss. The formula for the loss function is: where and represent the binary cross-entropy and the confidence loss respectively, N is the total number of training data, i is the number of the training data, y i is the i-th true label, p i is the i-th predicted label, TP: true positive, FP: false positive, FN: false negative, ⊙ represents element-wise multiplication, and λ1 and λ2 represent the weights of and respectively; specifically, in the present invention, both λ1 and λ2 are set to 1.
[0024] Referring to Figure 5 , an infrared small target segmentation device includes the following modules: An infrared sensor module for collecting the original infrared image; A preprocessing module configured to perform size normalization and cropping on the collected infrared image to a preprocessing of H × W pixels; A processor module loaded with any one of the above-mentioned trained methods based on the gated unit and multi-scale convolutional network method, including: a graphics processing unit with a video memory capacity ≥ (total number of network parameters × 4 bytes) × 1.2; An output module for generating an infrared image with a target segmentation mask and a pixel-level confidence map.
[0025] It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or system comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or system comprising that element.
[0026] The serial numbers of the embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments. In the unit claims listing several devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not denote any order and these words may be interpreted as identifiers.
[0027] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A small infrared target segmentation method based on gated unit and multi-scale convolutional network, characterized in that: The following steps are involved: S1. Obtain infrared image dataset and perform preprocessing, including size normalization and random cropping into H×W pixels and data augmentation operations; S2, build a network based on gated units and multi-scale convolutional networks; S3, using a deep supervision mechanism including multiple loss functions to train the network; S4. Deploy the trained model for object segmentation.
2. The method according to claim 1, characterized in that In step S2, a coding and decoding structure is used based on a gating unit and a multi-scale convolutional network, including: S2a, the encoder uses the pyramid visual Transformer (PVTv2) architecture to extract four-level multi-scale features {F1, F2, F3, F4}, which are passed to the decoder through cross-level connections; S2b, cross-layer connections realize cross-layer feature adaptive propagation through gating units; S2c, the decoder obtains multi-scale features through multi-scale parallel convolution.
3. The method according to claim 1, characterized in that The data enhancement in step S1 includes: S11, geometric transformation: random horizontal flip (probability p = 0.5), center crop (ratio ρ∈[0.8,1]), fixed ratio scaling (ratio s∈(0.75,1.25)); S12, Radiative transformation: Add Gaussian noise and Gaussian blur with standard deviation σ in the range of 0 to 0.1 (kernel size k∈3,5,7); S13. Elastic transformation: affine matrix M∈R 2×3 Deformation is performed with the constraint |det(M)|≥0.
9.
4. The method according to claim 2, characterized in that: The gate control unit specifically includes the following steps: S2b1, for the i-th level feature map F i Perform adaptive average pooling (AAP) processing, generate the intermediate feature map T after 3×3 depth-separable convolution (step size 1, padding 1), two-fold upsampling and batch normalization (BN) processing i ; S2b2, for the j-th level feature map F j Perform adaptive max pooling (AMP) processing, generate the intermediate feature map T after 3×3 depthwise separable convolution (step 1, padding 1) and batch normalization j ; S2b3, the intermediate feature map T i With T j Element-by-element addition is performed, and after being processed by the ReLU6 activation function, 1×1 convolution operation, squeeze-excitation attention module (channel compression ratio 16), batch normalization and Sigmoid function mapping are performed in sequence to generate the three-dimensional fusion weight Q; S2b4, dynamically allocate the original feature map according to the weight Q: j Multiply element-wise with (1-Q) to obtain the optimized feature map Y j , and at the same time F i Multiply element-wise with Q to obtain the optimized feature map Y i ; Among them: (i,j)∈(1,2),(3,4).
5. The method according to claim 2, characterized in that: The multi-scale parallel decoder specifically comprises the following steps: where Y i is the optimized feature map output by the i-th level gating unit, i∈1,2,3,4; The initial input Y1 is downsampled by 2 times, 3×3 convolution (step size 1, padding 1), batch normalization (BN) and ReLU6 activated, and then 3×3 convolution (step size 1, padding 1), BN, ReLU6 activation and 2 times upsampling and concatenated with Y1 through channels to generate X1; X1 is convolved with the first multi-scale parallel convolution to generate Z1; Z1 is upsampled by 2 times and concatenated with Y2 through channels to generate X2; X2 is convolved with the second multi-scale parallel convolution to generate Z2; Z2 is upsampled by 2 times and concatenated with Y3 through channels to generate X3; X3 is convolved with the third multi-scale parallel convolution to generate Z3; Z3 is upsampled by 2 times and concatenated with Y4 through channels to generate X4; X4 is convolved with the fourth multi-scale parallel convolution to generate Z4; Among them, the multi-scale parallel convolution specifically includes the following steps: (a) Large receptive field branch (Block1): cascade two depth-wise separable convolution blocks, each of which includes a depth-wise separable convolution with a 7×7 kernel, a BN layer, and a ReLU6 activation function; (b) Middle receptive field branch (Block2): cascade two depthwise separable convolution blocks, each of which includes a depthwise separable convolution with a 3×3 kernel, a BN layer, and a ReLU6 activation function; (c) Channel correction branch (Block3): 1×1 point convolution; (e) The final output feature satisfies: Z = (Block1(X)⊙Block2(X))⊙Block3(X) where ⊙ represents element-by-element multiplication, and the final feature dimension remains unchanged.
6. The method according to claim 1, characterized in that The deep supervision mechanism in step S4 includes: Where Z i Represents the output prediction map of each decoding stage; The output prediction graph of each decoding stage is fused using an adaptive weight adjustment strategy, and the final output is: where ω i is a learnable layer weight parameter, Up(·) represents bilinear upsampling to the original resolution, ω i Automatically optimize through back propagation and satisfy 7. The method according to claim 1, characterized in that The multi-loss function in step S4 is a weighted sum of binary cross entropy loss and confidence loss.
8. An infrared small target segmentation device, characterized in that: Includes the following modules: Infrared sensor module, used to collect raw infrared images; a preprocessing module configured to preprocess the collected infrared images by normalizing their size and cropping them into H×W pixels; A processor module, loaded with the gated unit and multi-scale convolutional network model trained according to any one of claims 1 to 6, comprising: a graphics processor unit, a video memory capacity ≥ (total network parameters × 4 bytes) × 1.2; The output module generates an infrared image with a target segmentation mask and a pixel-level confidence map.
Citation Information
Cited By
Blurring processing method and device, equipment and storage medium
CN120833275A
Remote sensing image segmentation method based on multi-scale gating bottleneck convolution scanning
CN121640277A