Lightweight infrared small target segmentation system and method
By designing a lightweight infrared small target segmentation system and using a combination of depthwise separable convolution, dilated convolution, and asymmetric convolution to construct a DAAA module, the problems of high parameter and computational complexity in infrared small target segmentation are solved, achieving efficient detection and robustness on mobile devices.
Patent Information
- Application Number
- CN202310252896.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-15
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2043-03-15
AI Technical Summary
Existing technologies have problems with high parameter and computational complexity in infrared small target segmentation, making it difficult to effectively deploy on mobile terminals and improve detection accuracy and robustness.
A lightweight infrared small target segmentation system is designed. Through the combination of encoding module and decoding module, including input module, initial downsampling module, first downsampling module, convolution module and decoding module, the DAAA module is constructed by combining depthwise separable convolution, dilated convolution and asymmetric convolution to reduce network parameters and computational complexity while improving detection accuracy and robustness.
It improves the accuracy and robustness of infrared small target detection while reducing network parameters and computational complexity. It is suitable for mobile deployment and can achieve high-performance reasoning capabilities on embedded development boards.
Smart Images

Figure CN116416430B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of lightweight infrared small target segmentation, and in particular to a lightweight infrared small target segmentation system and method. Background Art
[0002] Quickly and accurately separating small infrared targets from complex backgrounds is an extremely challenging task. With the advancement of deep learning, many algorithms have adopted classic classification networks (such as VGG, U-Net, and ResNet) as backbones to effectively extract features of small infrared targets. Furthermore, various multi-scale feature fusion modules and channel / spatial attention mechanism fusion modules have been designed. While these designs have slightly improved the accuracy of small infrared target detection, they also significantly increase the number of parameters and floating-point operations (FLOPs). Summary of the Invention
[0003] The purpose of the present invention is to provide a lightweight infrared small target segmentation system and method, aiming to solve the problem of mobile terminal deployment. Through the design and combination of various modules of the lightweight infrared small target segmentation system, the detection accuracy and robustness are improved while reducing network parameters and computational complexity.
[0004] The present invention provides a lightweight infrared small target segmentation system, comprising:
[0005] An encoding module, the encoding module comprising: an input module, an initial downsampling module, a first downsampling module, and a convolution module;
[0006] An input module is used to input an infrared grayscale image and preprocess the infrared grayscale image to obtain a preprocessed image;
[0007] The initial downsampling module is used to reduce the image resolution and perform feature extraction on the preprocessed image to obtain low-level feature information of small infrared targets;
[0008] The first downsampling module is used to reduce the image resolution and extract the high-level semantic information of the infrared small target from the low-level feature information of the infrared small target;
[0009] The convolution module is used to further reduce the image resolution and obtain the infrared small target features from the extracted high-level semantic information of the infrared small target;
[0010] The decoding module is used to restore the resolution of the low-resolution image of the infrared small target features and complete the segmentation.
[0011] The present invention also provides a lightweight infrared small target segmentation method, comprising:
[0012] S1. Input an infrared grayscale image through an input module and preprocess the infrared grayscale image to obtain a preprocessed image;
[0013] S2, extracting features from the preprocessed image through the initial downsampling module to obtain low-level feature information of the infrared small target;
[0014] S3, reducing the image resolution through the first downsampling module to extract high-level semantic information of the infrared small target from the low-level feature information of the infrared small target;
[0015] S4, obtaining infrared small target features from the extracted high-level semantic information of infrared small targets through a convolution module;
[0016] S5. The low-resolution image of the infrared small target feature is restored to its resolution through the decoding module, and the segmentation is completed.
[0017] By adopting the embodiments of the present invention, lightweight infrared small target segmentation can be achieved. By designing and combining various modules of the lightweight infrared small target segmentation system, the detection accuracy and robustness can be improved while reducing network parameters and computational complexity.
[0018] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it is implemented in accordance with the contents of the specification, and in order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 is a schematic diagram of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0021] Figure 2 2. It is a schematic diagram of an initial downsampling module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0022] Figure 3 2. It is a schematic diagram of a downsampling module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0023] Figure 4 Schematic diagram of a conventional convolution module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0024] Figure 52. It is a schematic diagram of a depth-separable module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0025] Figure 6 2 is a schematic diagram of a dilated convolution module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0026] Figure 7 Schematic diagram of an asymmetric convolution module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0027] Figure 8 2 is a schematic diagram of a dilated convolution module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0028] Figure 9 2. It is a schematic diagram of the segmentation results of the lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0029] Figure 10 This is a flowchart of a lightweight infrared small target segmentation method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0030] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0031] System Example
[0032] According to an embodiment of the present invention, a lightweight infrared small target segmentation system is provided. Figure 1 Schematic diagram of a lightweight infrared small target segmentation system according to an embodiment of the present invention. Figure 1 As shown, specifically including:
[0033] An encoding module, the encoding module comprising: an input module, an initial downsampling module, a first downsampling module, and a convolution module;
[0034] An input module is used to input an infrared grayscale image and preprocess the infrared grayscale image to obtain a preprocessed image;
[0035] The initial downsampling module is used to reduce the image resolution and perform feature extraction on the preprocessed image to obtain low-level feature information of small infrared targets;
[0036] The first downsampling module is used to reduce the image resolution and extract the high-level semantic information of the infrared small target from the low-level feature information of the infrared small target;
[0037] The convolution module is used to further reduce the image resolution and obtain the infrared small target features from the extracted high-level semantic information of the infrared small target;
[0038] The decoding module is used to restore the resolution of the low-resolution image of the infrared small target features and complete the segmentation.
[0039] The input module is specifically used to: input infrared grayscale images, adjust the resolution of the infrared grayscale images and perform standardization on the RGB three channels.
[0040] The initial downsampling module performs parallel calculations through convolution and maximum pooling, and finally uses the concat function to concatenate the channels of the two branches for feature extraction.
[0041] The convolution branch has a convolution kernel of 3, a stride of 2, and a channel number of 5;
[0042] The maximum pooling branch convolution kernel is 3, the step size is 2, and the number of channels is 3.
[0043] The first downsampling module includes a first convolution branch, a maximum pooling branch and a first sum function.
[0044] The first convolution branch adopts an hourglass-shaped design structure.
[0045] 2×2 convolution, 3×3 convolution, 1×1 convolution and random neuron cropping are connected in sequence. The number of channels of the 2×2 convolution is 8, the number of channels of the 3×3 convolution is 2, the number of channels of the 1×1 convolution is 32, and the random neuron cropping is set to 0.01;
[0046] The maximum pooling branch has a convolution kernel of 2 and a stride of 2. The number of feature map channels is matched by padding with 0, so that the two branches are fused;
[0047] Finally, the first sum function is used to add the number of channels corresponding to the first convolution branch and the maximum pooling branch, and PRelu is used to activate the model parameters.
[0048] The convolution module includes four conventional convolution modules, a second downsampling module and four DAAA modules connected in sequence, wherein the conventional convolution module includes a second convolution branch, a shortcut branch and a second sum function;
[0049] The second convolution branch includes: 1×1 convolution, 3×3 convolution, 1×1 convolution and random neuron cropping connected in sequence, the number of channels of the 1×1 convolution is 32, the number of channels of the 3×3 convolution is 8, the number of channels of the 1×1 convolution is 32, and the random neuron cropping is set to 0.01;
[0050] The shortcut branch is the output result of the previous module and does not do any processing;
[0051] The second sum function is used to add the number of channels corresponding to the two branches, and PRelu activates the model parameters;
[0052] The second downsampling module has the same structure as the first downsampling module, but the number of output channels is 64;
[0053] The DAAA modules include: a first DAAA module, a second DAAA module, a third DAAA module and a fourth DAAA module, the four DAAA modules have the same structure but different parameter settings;
[0054] The first DAAA module includes:
[0055] Depth separation convolution module, the first hole convolution module, the asymmetric convolution module and the second hole convolution module. The depth separation convolution module is used to reduce the amount of network parameters. The depth separation convolution module adopts a spindle-shaped design structure. The number of input 1×1 convolution channels is 64, the number of middle layer 3×3 convolution channels is 128, and the number of output 1×1 convolution channels is 64. The first hole convolution module is used to increase the receptive field of infrared small target feature extraction. The first hole convolution module adopts an hourglass-shaped design structure. The number of input 1×1 convolution channels is 64, the middle layer 3×3 convolution channels is 128, and the output 1×1 convolution channels is 64. The number of channels of the middle layer 3×3 dilated convolution is 16, the dilation rate is 2, and the number of output 1×1 convolution channels is 64. The asymmetric convolution module is used to reduce the number of model parameters. The asymmetric convolution module adopts an hourglass design structure. The number of input 1×1 convolution channels is 64, the asymmetric convolution kernel size of the middle layer is 7, the number of channels is 16, and the number of output 1×1 convolution channels is 64. The second dilated convolution module is used to increase the receptive field for extracting features of small infrared targets. The second dilated convolution module has the same structure as the first dilated convolution module, but the dilation rate is 4;
[0056] The dilated convolution module in the second DAAA module has a dilated ratio of 8 and 16, respectively, and the random cropping of neurons in the first and second DAAA modules is set to 0.01;
[0057] The void ratio of the third DAAA module is the same as that of the first DAAA module, the void ratio of the fourth DAAA module is the same as that of the second DAAA module, and the random cropping of neurons in the third DAAA module and the fourth DAAA module is set to 0.1.
[0058] The image resolution feature decoding module includes: a first upsampling module, two conventional convolution modules, a second upsampling module, a conventional convolution module and a third upsampling module connected in sequence.
[0059] The output port of the initial downsampling module is added to the second upsampling output port of the decoding module, and the output port of the second downsampling module is added to the first upsampling output port of the decoding module. The addition is used to add the feature images of the same resolution corresponding to the encoding module and the decoding module to fully extract the features of the infrared small target.
[0060] The specific implementation methods are as follows:
[0061] Coding phase:
[0062] The input module is used to input an RGB infrared grayscale image and resize the image resolution to 256 x 256. The three RGB channels of the input image are normalized to the mean (0.485, 0.456, 0.406) and the variance (0.229, 0.224, 0.225). This is used to compress the input image resolution, reduce the computational complexity of the network, and accelerate the convergence of the model.
[0063] Figure 2 2. It is a schematic diagram of an initial downsampling module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0064] The convolution branch of the initial downsampling module has a convolution kernel of 3, a stride of 2, and a channel number of 5. The maximum pooling branch has a convolution kernel of 3, a stride of 2, and a channel number of 3.
[0065] Although the original image has been compressed to a size of 256×256, it still consumes a significant amount of computational resources. Therefore, in the initial downsampling module, the Inception V2 architecture is adopted. The convolution branch and the maximum pooling branch are computed in parallel. Finally, the concat function is used to combine the channels of the two branches. This allows for feature extraction of small infrared targets, while simultaneously reducing model parameters and accelerating inference time. The convolution branch uses a kernel size of 3, a channel size of 5, and a stride of 2; the maximum pooling branch uses a kernel size of 3, a channel size of 3, and a stride of 2. Finally, the concat function is used to combine the two branches.
[0066] Figure 3 2. It is a schematic diagram of a downsampling module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0067] To further reduce image resolution and extract high-level semantic information from small infrared targets, a downsampling module was designed. This module also consists of two branches. The convolution branch adopts an hourglass design, compressing the number of channels in the intermediate layers, effectively reducing the number of model parameters. The maximum pooling branch complements the convolution branch to prevent gradient divergence in deep networks, which can lead to training failures.
[0068] The convolution branch in the downsampling module adopts an hourglass design structure:
[0069] The first layer is 2×2 convolution with a stride of 2 and 8 channels;
[0070] The second layer is 3×3 convolution with a stride of 1 and the number of channels is compressed by 4 times to 2;
[0071] The third layer is 1×1 convolution, with a stride of 1 and 32 channels;
[0072] The fourth layer is pruned using Dropout (Dropout = 0.01);
[0073] The maximum pooling branch in the downsampling module matches the feature map resolution by padding with 0, so that the two branches can be fused.
[0074] Finally, the sum function is used to add the number of channels corresponding to the two branches, and PRelu is used to activate the model parameters.
[0075] Figure 4 Schematic diagram of a conventional convolution module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0076] The convolution branch of a conventional convolution module is:
[0077] The first layer is 1×1 convolution with a stride of 1 and 32 channels;
[0078] The second layer is a 3×3 convolution with a stride of 1 and a channel number compressed by 4 times to 8;
[0079] The third layer is 1×1 convolution with 32 channels;
[0080] Shortcut branch:
[0081] Output directly without changing the result obtained by the previous layer of network;
[0082] The conventional convolution module is executed four times in succession to further effectively extract the features of small infrared targets. At the same time, to prevent gradient divergence in deep networks, which can lead to training failures, a shortcut branch is added to this module.
[0083] In order to achieve network lightweight, the combination and arrangement of the four sub-modules of depthwise separable convolution, dilated convolution, and asymmetric convolution are fully utilized to form four DAAA modules to further extract features of infrared small targets.
[0084] The submodules in the DAAA module all utilize a shortcut connection design. The dilated-asymmetric convolution module employs an hourglass-shaped design, using 1×1 convolutions to compress the number of intermediate layer channels (by a factor of four), reducing the number of parameters. The depthwise separable module employs a spindle-shaped design, using 1×1 convolutions to expand the number of intermediate layer channels (by a factor of two), enhancing feature learning capabilities. The DAAA module is constructed using a combination of "depthwise separable-dilated-asymmetric-dilated."
[0085] The first DAAA module uses a combination of depthwise separable convolution-dilated convolution-asymmetric convolution-dilated convolution modules;
[0086] Figure 5 2. It is a schematic diagram of a depth-separable module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0087] A spindle-shaped structure design is adopted to modify the conventional convolution module, in which the number of channels in the intermediate layer of the convolution branch is expanded from 64 to 128, and channel-by-channel convolution is performed.
[0088] The depthwise separable convolution module effectively reduces the number of network parameters, but in order to fully extract the features of small infrared targets, the number of channels in the middle layer of the convolution branch is expanded.
[0089] Figure 6 2 is a schematic diagram of a dilated convolution module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0090] The conventional convolution module is modified, where the middle layer of the convolution branch is modified to a dilated convolution with a dilation ratio of 2. This increases the receptive field of the model for extracting features of small infrared targets.
[0091] Figure 7 Schematic diagram of an asymmetric convolution module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0092] The conventional convolution module is modified, where the middle layer of the convolution branch is modified to an asymmetric convolution with convolution kernels of 7×1 and 1×7.
[0093] Because the receptive field of a large convolution kernel (7×7) is larger than that of a small convolution kernel (3×3), but it also increases the number of parameters, an asymmetric convolution module is designed to expand the receptive field while reducing the number of model parameters.
[0094] Figure 8 2 is a schematic diagram of a dilated convolution module of a lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0095] The conventional convolution module is modified, where the middle layer of the convolution branch is modified to a dilated convolution with a dilation ratio of 4. This increases the receptive field of the model for extracting features of small infrared targets.
[0096] The dilated convolution module in the second DAAA module has a dilated ratio of 8 and 16, respectively, and the random cropping of neurons in the first and second DAAA modules is set to 0.01;
[0097] The void ratio of the third DAAA module is the same as that of the first DAAA module, the void ratio of the fourth DAAA module is the same as that of the second DAAA module, and the random cropping of neurons in the third and fourth DAAA modules is set to 0.1.
[0098] Decoding stage
[0099] In the decoding stage, the first upsampling module, two conventional convolution modules, the second upsampling module, one conventional convolution module and the third upsampling module are connected in sequence.
[0100] Among them, the first, second, and third upsampling modules restore the feature map size through deconvolution, also known as transposed convolution, where the convolution kernel size is 3 and the step size is 2.
[0101] Function: In order to ensure the effective segmentation of small infrared targets on the input image, the resolution of the low-resolution image rich in high-level semantic information must be restored.
[0102] Conventional convolutional module in the decoding stage
[0103] The design structure and function are consistent with the conventional convolution module in the encoding stage, the only difference is that the activation function is changed from PReLU to ReLU;
[0104] In order to achieve semantic interaction between the encoding and decoding stages and reduce complex fusion methods, the "sum" connection method is used between feature maps of the same resolution to add the number of channels to achieve feature fusion.
[0105] Figure 9 2. It is a schematic diagram of the segmentation results of the lightweight infrared small target segmentation system according to an embodiment of the present invention;
[0106] The input image is sent into the model, and the infrared small target probability map (heat map) is obtained through calculation; the threshold is set to 0, the infrared small target is segmented, and the final result is obtained.
[0107] This paper proposes a design strategy for a lightweight network for infrared small target segmentation:
[0108] In order to prevent the loss of small infrared targets in the feature extraction (encoding) stage, the depth of the backbone network is minimized.
[0109] In order to prevent the gradient from disappearing or exploding due to excessive network depth, we borrowed the idea of ResNet and added shortcut connections.
[0110] In order to achieve network lightweight, 1*1 convolution is used in conventional convolution, void convolution, and asymmetric convolution modules to implement an "hourglass" structure design to further reduce the network width (number of channels).
[0111] In order to improve the feature extraction capability of the network without increasing the number of network parameters, 1*1 convolution is used to implement the "spindle-shaped" structure design in the depth-wise separable convolution process, and the network width (number of channels) is appropriately increased.
[0112] In order to achieve network lightweight, the combination of depthwise separable convolution, dilated convolution, and asymmetric convolution is fully utilized to form the DAAA module, which further features infrared small target extraction and thus achieves network lightweight.
[0113] A lightweight infrared small target segmentation network model (LW-IRSTNet) was invented and designed. This model leverages a combination of depthwise separable convolution, dilated convolution, and asymmetric convolution to construct the DAAA module. Ablation experiments demonstrate the significant role of the DAAA module in improving segmentation accuracy and reducing parameters and FLOPs. To demonstrate the accuracy, robustness, and time-efficiency of LW-IRSTNet, a comparative analysis was conducted using 17 state-of-the-art algorithms as baselines. Experimental results show that the proposed LW-IRSTNet algorithm achieves comparable or superior performance on the mIOU, F1, and ROC metrics across five datasets, while also compressing parameters to 0.16M and FLOPs to 303M, significantly less than the baselines. Furthermore, LW-IRSTNet was deployed on the inexpensive Orange Pi 5 embedded development board. Leveraging the ONNX architecture, NPU acceleration, and CPU multithreading, it achieves high-performance inference, reaching 47 FPS.
[0114] Method Example
[0115] According to an embodiment of the present invention, a lightweight infrared small target segmentation method is provided. Figure 10 Flowchart of the lightweight infrared small target segmentation method according to an embodiment of the present invention. Figure 10 As shown, specifically including:
[0116] S1. Input an infrared grayscale image through an input module and preprocess the infrared grayscale image to obtain a preprocessed image;
[0117] S2, extracting features from the preprocessed image through the initial downsampling module to obtain low-level feature information of the infrared small target;
[0118] S3, reducing the image resolution through the first downsampling module to extract high-level semantic information of the infrared small target from the low-level feature information of the infrared small target;
[0119] S4, obtaining infrared small target features from the extracted high-level semantic information of infrared small targets through a convolution module;
[0120] S5. The low-resolution image of the infrared small target feature is restored to its resolution through the decoding module, and the segmentation is completed.
[0121] S1 specifically includes: inputting an infrared grayscale image, adjusting the resolution of the infrared grayscale image, and normalizing the RGB three channels.
[0122] S2 specifically includes: parallel calculation through two branches of convolution and maximum pooling, and finally using the concat function to splice the number of channels of the two branches for feature extraction.
[0123] The embodiment of the present invention is a system embodiment corresponding to the above-mentioned method embodiment. The specific operations of each module can be understood by referring to the description of the method embodiment, which will not be repeated here.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. These modifications or replacements of the technical solutions of the embodiments of the present invention do not cause the essence of the corresponding technical solutions to deviate from the scope of this solution.
Claims
1. A lightweight infrared small target segmentation system, characterized by: include, An encoding module, the encoding module comprising: an input module, an initial downsampling module, a first downsampling module, and a convolution module; An input module is used to input an infrared grayscale image and preprocess the infrared grayscale image to obtain a preprocessed image; The initial downsampling module is used to reduce the image resolution and perform feature extraction on the preprocessed image to obtain low-level feature information of small infrared targets; The first downsampling module is used to reduce the image resolution and extract the high-level semantic information of the infrared small target from the low-level feature information of the infrared small target; The convolution module is used to further reduce the image resolution and obtain the infrared small target features from the extracted high-level semantic information of the infrared small target; The decoding module is used to restore the resolution of the low-resolution image of the infrared small target features and complete the segmentation; The input module is specifically used to: input infrared grayscale images, adjust the resolution of the infrared grayscale images and perform standardization processing on the RGB three channels; The initial downsampling module performs parallel calculations through convolution and maximum pooling, and finally uses the concat function to concatenate the channels of the two branches for feature extraction; The convolution branch has a convolution kernel of 3, a stride of 2, and a channel number of 5; The maximum pooling branch convolution kernel is 3, the step size is 2, and the number of channels is 3; The first downsampling module includes a first convolution branch, a maximum pooling branch and a first sum function; The first convolution branch adopts an hourglass-shaped design structure; 2×2 convolution, 3×3 convolution, 1×1 convolution and random neuron cropping are connected in sequence. The number of channels of the 2×2 convolution is 8, the number of channels of the 3×3 convolution is 2, the number of channels of the 1×1 convolution is 32, and the random neuron cropping is set to 0.01; The maximum pooling branch has a convolution kernel size of 2 and a step size of 2. The number of feature map channels is matched by padding with 0, so that the two branches are fused; Finally, the first sum function is used to add the number of channels corresponding to the first convolution branch and the maximum pooling branch, and the PRelu function is used to activate the model parameters; The convolution module includes four conventional convolution modules, a second downsampling module and four DAAA modules connected in sequence, and the conventional convolution module includes a second convolution branch, a shortcut branch and a second sum function; The second convolution branch includes: 1×1 convolution, 3×3 convolution, 1×1 convolution and random neuron cropping connected in sequence, the number of channels of the 1×1 convolution is 32, the number of channels of the 3×3 convolution is 8, the number of channels of the 1×1 convolution is 32, and the random neuron cropping is set to 0.01; The shortcut branch is the output result of the previous module and does not undergo any processing; The second sum function adds the number of channels corresponding to the two branches and activates the model parameters using the PRelu function; The second downsampling module has the same structure as the first downsampling module, but the number of output channels is 64; The DAAA modules include: a first DAAA module, a second DAAA module, a third DAAA module and a fourth DAAA module. The four DAAA modules have the same structure but different parameter settings.
2. The system according to claim 1, wherein: The first DAAA module includes: Depth separation convolution module, the first hole convolution module, the asymmetric convolution module and the second hole convolution module. The depth separation convolution module is used to reduce the amount of network parameters. The depth separation convolution module adopts a spindle-shaped design structure. The number of input 1×1 convolution channels is 64, the number of middle layer 3×3 convolution channels is 128, and the number of output 1×1 convolution channels is 64. The first hole convolution module is used to increase the receptive field of infrared small target feature extraction. The first hole convolution module adopts an hourglass-shaped design structure. The number of input 1×1 convolution channels is 64, the middle layer 3×3 convolution channels is 128, and the output 1×1 convolution channels is 64. The number of channels of the middle layer 3×3 dilated convolution is 16, the dilation rate is 2, and the number of output 1×1 convolution channels is 64. The asymmetric convolution module is used to reduce the number of model parameters. The asymmetric convolution module adopts an hourglass design structure. The number of input 1×1 convolution channels is 64, the asymmetric convolution kernel size of the middle layer is 7, the number of channels is 16, and the number of output 1×1 convolution channels is 64. The second dilated convolution module is used to increase the receptive field for extracting features of small infrared targets. The second dilated convolution module has the same structure as the first dilated convolution module, but the dilation rate is 4; The dilated convolution module in the second DAAA module has a dilated ratio of 8 and 16, respectively, and the random cropping of neurons in the first and second DAAA modules is set to 0.01; The void ratio of the third DAAA module is the same as that of the first DAAA module, the void ratio of the fourth DAAA module is the same as that of the second DAAA module, and the random cropping of neurons in the third DAAA module and the fourth DAAA module is set to 0.
1.
3. The system according to claim 1, wherein: The image resolution feature decoding module includes: a first upsampling module, two conventional convolution modules, a second upsampling module, a conventional convolution module and a third upsampling module connected in sequence.
4. The system according to claim 3, characterized in that The output port of the initial downsampling module is added to the second upsampling output port of the decoding module, and the output port of the second downsampling module is added to the first upsampling output port of the decoding module. The addition is used to add the feature images of the same resolution corresponding to the encoding module and the decoding module to fully extract the features of the infrared small target.
5. A lightweight infrared small target segmentation method, applied to the lightweight infrared small target segmentation system according to any one of claims 1 to 4, characterized in that: include, S1. Input an infrared grayscale image through an input module and preprocess the infrared grayscale image to obtain a preprocessed image; S2, extracting features from the preprocessed image through the initial downsampling module to obtain low-level feature information of the infrared small target; S3, reducing the image resolution through the first downsampling module to extract high-level semantic information of the infrared small target from the low-level feature information of the infrared small target; S4, obtaining infrared small target features from the extracted high-level semantic information of infrared small targets through a convolution module; S5. The low-resolution image of the infrared small target feature is restored to its resolution through the decoding module, and the segmentation is completed.
6. The method according to claim 5, characterized in that Said S1 specifically includes: inputting an infrared grayscale image, adjusting the resolution of the infrared grayscale image and performing standardization processing on the RGB three channels.
7. The method according to claim 6, characterized in that The S2 specifically includes: performing parallel calculations on two branches, convolution and maximum pooling, and finally using the concat function to concatenate the channels of the two branches for feature extraction.
Citation Information
Patent Citations
Coal rock segmentation method based on multi-modal fusion
CN113159038A
Steel billet automatic semantic segmentation identification method based on deep learning
CN114612456A