Infrared weak and small target detection method and device
By improving the infrared weak object detection model of YOLOV8 network, using the channel transpose attention mechanism and super-segment coupling loss function, the problem of weak object detection accuracy and low efficiency in infrared remote sensing images is solved, and efficient infrared weak object detection is achieved.
Patent Information
- Application Number
- CN202510492262.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, the infrared weak target detection in infrared remote sensing images has problems such as low resolution, small signal-to-noise ratio, blurred edges, and unclear texture, resulting in low detection accuracy and low efficiency.
The infrared weak object detection model based on YOLOV8 is adopted, including the channel transpose attention mechanism, the improved backbone network of the fast fusion C2f module and the lightweight space pyramid pooling module, the improved neck network based on the infrared weak object detection module, and the improved detection head network of the infrared weak object detection head and three lightweight detection heads, and the model is trained using the super-segment coupling loss function.
The accuracy and efficiency of infrared weak target detection are improved, and by obtaining more channel and spatial feature information, the calculation amount is reduced, and the network expression ability is enhanced, so as to achieve efficient detection of infrared weak targets.
Smart Images

Figure CN120355901A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a method and device for detecting infrared dim and small targets. Background Art
[0002] Infrared remote sensing images are widely used in military surveillance, security monitoring, environmental supervision, autonomous navigation, scientific investigation and other fields because of their ability to observe the thermal radiation characteristics of different targets. Among them, for the infrared dim and small targets contained in infrared remote sensing images, due to their contrast less than 0.15, signal-to-noise ratio less than 1.5, and the total number of pixels less than 15% of the entire image, it is difficult to detect infrared dim and small targets. In recent years, with the development of artificial intelligence, the detection of dim and small targets in the field of infrared remote sensing has become an important research direction in the fields of computer vision and deep learning.
[0003] In the prior art, the object detection algorithms based on deep learning include: single-stage object detection algorithms represented by the YOLO series and two-stage object detection algorithms represented by R-CNN. Specifically, the single-stage object detection algorithm directly classifies and locates the entire image. Common algorithms include YOLO (you only look once), single short multi box detector (SSD), RetinaNet, etc. The two-stage object detection algorithm is divided into two main stages: candidate region proposal generation (Proposal Generation) and candidate region classification and bounding box regression (Proposal Classification). First, a series of candidate regions are generated in the input image, and then more accurate bounding box regression and object category prediction are performed on each candidate region generated in the first stage.
[0004] However, using the prior art, due to the problems of low resolution, low signal-to-noise ratio, blurred edges and unclear textures in the infrared dim and small target images, there is a lack of rich feature information, resulting in low detection accuracy and inability to effectively detect infrared dim and small targets. And due to the high complexity of the network structure and large amount of calculation of the detection model, the efficiency of detecting infrared dim and small targets is also reduced. Summary of the Invention
[0005] Based on this, it is necessary to provide a method and device for detecting infrared dim and small targets for the above technical problems.
[0006] In a first aspect, an embodiment of the present invention provides a method for detecting infrared dim and small targets, including:
[0007] Obtain an infrared image to be detected, where the infrared image to be detected contains at least one infrared dim and small target;
[0008] Input the infrared image to be detected into the improved infrared small and weak target detection model based on YOLOV8 to obtain the infrared small and weak target detection result of the infrared image to be detected;
[0009] Among them, the infrared small and weak target detection model includes: a backbone network improved based on the channel transposed attention mechanism, the fast fusion C2f module, and the lightweight spatial pyramid pooling module, a neck network improved based on the infrared small and weak target detection module, and a detection head network improved based on the infrared small and weak target detection head and three lightweight detection heads. The infrared small and weak target detection model is trained based on the super-resolution coupled loss function, and the super-resolution coupled loss function is calculated and determined according to the low-level features, high-level features, and super-resolution coupled network extracted by the backbone network.
[0010] In a second aspect, an embodiment of the present invention provides an infrared small and weak target detection device, including:
[0011] An infrared image to be detected acquisition module for acquiring an infrared image to be detected, where the infrared image to be detected includes at least one infrared small and weak target;
[0012] A detection module for inputting the infrared image to be detected into the improved infrared small and weak target detection model based on YOLOV8 to obtain the infrared small and weak target detection result of the infrared image to be detected;
[0013] Among them, the infrared small and weak target detection model includes: a backbone network improved based on the channel transposed attention mechanism, the fast fusion C2f module, and the lightweight spatial pyramid pooling module, a neck network improved based on the infrared small and weak target detection module, and a detection head network improved based on the infrared small and weak target detection head and three lightweight detection heads. The infrared small and weak target detection model is trained based on the super-resolution coupled loss function, and the super-resolution coupled loss function is calculated and determined according to the low-level features, high-level features, and super-resolution coupled network extracted by the backbone network.
[0014] The technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art:
[0015] An infrared dim and small target detection method provided by an embodiment of the present invention obtains a to-be-detected infrared image in this way. The to-be-detected infrared image is input into an infrared dim and small target detection model improved based on YOLOV8 to obtain the infrared dim and small target detection result of the to-be-detected infrared image. The infrared dim and small target detection model includes: a backbone network improved based on a channel transposed attention mechanism, a fast fusion C2f module, and a lightweight spatial pyramid pooling module; a neck network improved based on an infrared dim and small target detection module; and a detection head network improved based on an infrared dim and small target detection head and three lightweight detection heads. Since the channel transposed attention mechanism can improve the ability of the detection model to learn features and enhance the expression ability of the network, the fast fusion C2f module and the lightweight spatial pyramid pooling module can reduce the model calculation amount and obtain important feature information. The infrared dim and small target detection module can obtain more channel feature information and spatial feature information of infrared dim and small targets, and the infrared dim and small target detection head can detect infrared dim and small targets. Based on this, the use of the infrared dim and small target detection model improved based on YOLOV8 improves the accuracy and efficiency of detecting infrared dim and small targets. Further, the infrared dim and small target detection model is trained based on a super-resolution coupled loss function, and the super-resolution coupled loss function is determined according to the low-level features, high-level features, and super-resolution coupled network extracted by the backbone network. In this way, the method of calculating the super-resolution coupled loss using the super-resolution coupled network guides the training of the infrared dim and small target detection model, thereby further improving the accuracy of the infrared dim and small target detection model in detecting infrared dim and small targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a schematic flowchart of an infrared dim and small target detection method provided by an embodiment of the present invention;
[0019] Figure 2 It is a schematic structural diagram of a YOLOV8 network in the prior art provided by an embodiment of the present invention;
[0020] Figure 3 It is a schematic structural diagram of an infrared dim and small target detection model provided by an embodiment of the present invention;
[0021] Figure 4 Schematic diagrams of an existing C2f module and a fast fusion C2f module FL-C2f provided by an embodiment of the present invention;
[0022] Figure 5 Schematic diagram of a lightweight convolutional layer provided by an embodiment of the present invention;
[0023] Figure 6 Schematic diagrams of a spatial pyramid pooling module and a lightweight spatial pyramid pooling module provided by an embodiment of the present invention;
[0024] Figure 7 Schematic diagram of a channel transposed attention mechanism module provided by an embodiment of the present invention;
[0025] Figure 8 Schematic diagram of a detection head provided by an embodiment of the present invention;
[0026] Figure 9 Schematic diagram of a super-resolution coupled network provided by an embodiment of the present invention;
[0027] Figure 10 Schematic diagram of an infrared dim and small target detection device provided by an embodiment of the present invention. Detailed implementation manners
[0028] In order to more clearly understand the above objects, features and advantages of the present invention, the solutions of the present invention will be further described below. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0029] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention, but the present invention may be practiced in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present invention, rather than all the embodiments.
[0030] In one embodiment, as Figure 1 shown, Figure 1 Schematic flow diagram of an infrared dim and small target detection method provided by an embodiment of the present invention, specifically including the following steps:
[0031] S10: Obtain the infrared image to be detected.
[0032] Among them, the infrared image to be detected contains at least one infrared dim and small target. The infrared image to be detected can be collected by an infrared sensor.
[0033] S11: Input the infrared image to be detected into an infrared dim and small target detection model improved based on YOLOV8 to obtain the infrared dim and small target detection result of the infrared image to be detected.
[0034] Among them, the original YOLOV8 network is a lightweight network model, referring to Figure 2 As shown, the YOLOV8 network includes: a backbone network, a neck network, and a detection head network.
[0035] Optionally, based on the above embodiments, in some embodiments of the present invention, referring to Figure 3 As shown, the infrared small and weak target detection model improved based on YOLOV8 of the present invention includes: a backbone network 10 improved based on a channel transposed attention mechanism, a fast fusion C2f module, and a lightweight spatial pyramid pooling module, a neck network 20 improved based on an infrared small and weak target detection module, and a detection head network 30 improved based on an infrared small and weak target detection head and three lightweight detection heads.
[0036] Optionally, based on the above embodiments, in some embodiments of the present invention, continue to refer to Figure 3 As shown, the backbone network 10 improved based on a channel transposed attention mechanism, a fast fusion C2f module, and a lightweight spatial pyramid pooling module includes: a standard convolution module Conv, four CBS modules and four fast fusion C2f modules FL-C2f connected in sequence crosswise, a lightweight spatial pyramid pooling module SPPF, and a channel transposed attention mechanism module CAT.
[0037] It should be noted that the fast fusion C2f module FL-C2f is improved based on the C2f module in the backbone network included in the original YOLOV8 network. Referring to Figure 4 As shown, the fast fusion C2f module FL-C2f includes: a lightweight convolutional layer GSConv connected in sequence, two bottleneck layers Bottleneck, a concatenation layer Concat, and a lightweight convolutional layer GSConv. The bottleneck layer includes: a lightweight convolutional layer GSConv, two standard convolutional layers Conv, and a feature fusion layer Add connected in sequence, that is, replacing the standard convolutional layer included in the C2f module with a lightweight convolutional layer, and introducing a lightweight convolutional layer GSConv in the bottleneck layer Bottleneck.
[0038] Optionally, based on the above embodiments, in some embodiments of the present invention, continue to refer to Figure 5 As shown, the lightweight convolutional layer GSConv includes: a standard convolutional layer Conv, a linear transformation layer Cheap operation, and a concatenation layer Concat connected in sequence.
[0039] Specifically, when the lightweight convolutional layer performs a convolutional operation, first, it uses fewer convolutional kernels to perform a convolutional operation on the input feature map to obtain a convolutional feature map. Secondly, it performs a linear transformation on the convolutional feature map through a linear transformation layer to generate a linear convolutional feature map. Finally, it concatenates the convolutional feature map and the linear convolutional feature map through the concatenation layer Concat to obtain the finally output lightweight convolutional feature map, thereby reducing the model's computational amount and improving the efficiency of detecting infrared small and weak targets while retaining important feature information.
[0040] Optionally, based on the above embodiments, in some embodiments of the present invention, the lightweight spatial pyramid pooling module SPPF is improved based on the spatial pyramid pooling module SPPF in the backbone network included in the original YOLOV8 network. Refer to Figure 6 As shown, the lightweight spatial pyramid pooling module SPPF includes: a lightweight convolutional layer GSConv, three max pooling layers MaxPool2d, a concatenation layer Concat, and a lightweight convolutional layer GSConv connected in sequence. That is, the standard convolutional layer in the original spatial pyramid pooling module SPPF is replaced with a lightweight convolutional layer GSConv, which can reduce the model's computational amount and improve the efficiency of detecting infrared small and weak targets while retaining important feature information.
[0041] Optionally, based on the above embodiments, in some embodiments of the present invention, in order to further strengthen the spatial information and channel information of the extracted features and obtain more context position information and global detail features of infrared small and weak targets. A channel transpose attention mechanism module is introduced after the lightweight spatial pyramid pooling module SPPF included in the backbone network 10. Refer to Figure 7 As shown, the channel transpose attention mechanism module includes: a channel attention module, a depthwise separable spatial convolution module, two activation function layers Sigmoid, a fusion module, and a linear layer Linear.
[0042] Among them, the channel transpose attention mechanism module extracts features from the spatial pooling features input by the lightweight spatial pyramid pooling module to obtain channel transpose attention features.
[0043] Optionally, based on the above embodiments, in some embodiments of the present invention, an implementation manner for the channel transpose attention mechanism module to extract features from the spatial pooling features input by the lightweight spatial pyramid pooling module to obtain channel transpose attention features can be:
[0044] S20: Input the spatial pooling features into the channel attention module to obtain a channel attention feature map.
[0045] Among them, continue to refer to Figure 7As shown, the channel attention module includes: a channel self-attention layer Channel self-Attention and a channel attention layer Channel Projection connected in sequence. Based on this, one implementation of S20 can be:
[0046] S201: Input the spatial pooling feature into the channel self-attention layer to obtain the first channel attention feature map.
[0047] S202: Input the first channel attention feature map into the channel attention layer to obtain the channel attention feature map.
[0048] Specifically, input the spatial pooling feature input by the lightweight spatial pyramid pooling module into the channel self-attention layer, extract the channel attention feature map through the channel self-attention layer, obtain the first channel attention feature map containing channel attention information, and input the first channel attention feature map into the channel attention layer to obtain the channel attention feature map.
[0049] S21: Input the spatial pooling feature into the depthwise separable spatial convolution module to obtain the depthwise separable spatial convolution feature map.
[0050] Among them, continue to refer to Figure 7 As shown, the depthwise separable spatial convolution module includes: a standard convolution layer Conv, a depthwise separable convolution layer DWConv, and a spatial attention layer Spatial Projection connected in sequence. Based on this, one implementation of S21 can be:
[0051] S211: Input the spatial pooling feature into the standard convolution layer to obtain the first depthwise separable spatial convolution feature map.
[0052] S212: Input the first depthwise separable spatial convolution feature map into the depthwise separable convolution layer to obtain the second depthwise separable spatial convolution feature map.
[0053] Among them, the depthwise separable convolution layer includes: a depthwise convolution layer DepthwiseConv and a pointwise convolution layer PointwiseConv connected in sequence. In this way, when using the depthwise separable convolution layer to perform depthwise separable convolution operations on the input features, the depthwise convolution layer can first perform per-channel convolution to capture spatial feature information, and further perform channel-wise fusion through the pointwise convolution layer, thereby reducing the number of model parameters and memory access time, and improving the inference speed and performance of the model.
[0054] S213: Input the second depthwise separable spatial convolution feature map into the spatial attention layer to obtain the depthwise separable spatial convolution feature map.
[0055] Specifically, the spatially pooled features input to the lightweight spatial pyramid pooling module are input to a standard convolutional layer to obtain a first depthwise separable spatial convolutional feature map. The first depthwise separable spatial convolutional feature map is input to a depthwise separable convolutional layer for depthwise separable convolution operations to obtain a second depthwise separable spatial convolutional feature map. The second depthwise separable spatial convolutional feature map is input to a spatial attention layer for spatial attention feature extraction to obtain a depthwise separable spatial convolutional feature map.
[0056] S22: Input the channel attention feature map into the first activation function connected to the channel attention module to obtain a first activated channel attention feature map.
[0057] S23: Input the depthwise separable spatial convolutional feature map into the second activation function connected to the depthwise separable spatial convolution module to obtain a first activated depthwise separable spatial convolutional feature map.
[0058] S24: Input the first activated channel attention feature map into the second activation function to obtain a second activated channel attention feature map.
[0059] S25: Input the first activated depthwise separable spatial convolutional feature map into the first activation function to obtain a second activated depthwise separable spatial convolutional feature map.
[0060] S26: Multiply the channel attention feature map and the second activated depthwise separable spatial convolutional feature map element-wise to obtain a first fused feature.
[0061] S27: Multiply the depthwise separable spatial convolutional feature map and the second activated channel attention feature map element-wise to obtain a second fused feature.
[0062] S28: Input the first fused feature and the second fused feature into a fusion module to obtain a target fused feature.
[0063] S29: Input the target fused feature into a linear layer to obtain a channel transposed attention feature.
[0064] Exemplarily, continue to refer to Figure 7 As shown, the specific process of feature extraction for the input spatially pooled feature F x is as follows: First, input the spatially pooled feature F x into the channel attention module. Through the channel self-attention layer Channel self-Attention and the channel attention layer Channel Projection connected in sequence included in the channel attention module, obtain the channel attention feature map F1. Input the spatially pooled feature F xInput to the depthwise separable spatial convolution module, and obtain the depthwise separable spatial convolution feature map F2 through the standard convolution layer Conv, depthwise separable convolution layer DWConv, and spatial attention layer Spatial Projection that are sequentially connected in the depthwise separable spatial convolution module. The channel attention feature map F1 and the depthwise separable spatial convolution feature map F2 can be defined by the following expressions:
[0065]
[0066] Among them, CP(·) represents the operation of the channel attention layer, CsA(·) represents the operation of the channel self-attention layer, L in (·) represents the operation of the linear layer, Conv(·) represents the operation of the standard convolution layer, DWConv(·) represents the operation of the depthwise separable convolution layer, and SP(·) represents the operation of the spatial attention layer.
[0067] Secondly, the channel attention feature map F1 and the depthwise separable spatial convolution feature map F2 are respectively activated through the first activation function (Sigmoid function) and the second activation function (Sigmoid function) to obtain the first activated channel attention feature map F 11 , the first activated depthwise separable spatial convolution feature map F 22 . The first activated channel attention feature map F 11 is input into the second activation function to obtain the second activated channel attention feature map F 111 , and the first activated depthwise separable spatial convolution feature map F 22 is input into the first activation function to obtain the second activated depthwise separable spatial convolution feature map F 222 .
[0068] Finally, the channel attention feature map F1 and the second activated depthwise separable spatial convolution feature map F 222 are multiplied element-wise to obtain the first fusion feature. The depthwise separable spatial convolution feature map F2 and the second activated channel attention feature map F 111 are multiplied element-wise to obtain the second fusion feature. The first fusion feature and the second fusion feature are input into the fusion module to obtain the target fusion feature. The target fusion feature is input into the linear layer to obtain the channel transposed attention feature. The channel transposed attention feature can be defined by the following expression:
[0069] F out = L in (F1⊙F 222 ⊕ F2⊙F 111 )
[0070] Among them, ⊕ represents the element-wise addition feature fusion operation, ⊙ represents the element-wise multiplication operation, Lin (·) represents a linear layer.
[0071] In this way, by transposing and weighting the channels through the channel transposed attention mechanism module, the ability of the detection model to learn features can be improved, the expression ability of the network can be enhanced, and by weighting the channel attention, the performance of the visual task can be improved, and the detection accuracy of the detection model for detecting infrared small and weak targets can be improved.
[0072] Optionally, on the basis of the above embodiments, in some embodiments of the present invention, continue to refer to Figure 3 As shown, the neck network 20 improved based on the infrared small and weak target detection module includes: an infrared small and weak target detection module, a first neck sub-network, and a second neck sub-network;
[0073] Among them, the first neck sub-network is used to perform feature processing on the first backbone output feature map P3, the second backbone output feature map P4, and the channel attention feature map P5 to obtain a first fusion feature map.
[0074] It should be noted that the first backbone output feature map P3 is obtained by the output of the third fast fusion C2f module FL-C2f of the backbone network, and the second backbone output feature map P4 is obtained by the output of the fourth fast fusion C2f module FL-C2f of the backbone network.
[0075] The infrared small and weak target detection module is used to perform feature fusion on the first fusion feature map, the first backbone output feature map, and the third backbone output feature map P2 to obtain an infrared small and weak target fusion feature map and a second fusion feature map.
[0076] It should be noted that the third backbone output feature map P2 is obtained by the output of the second fast fusion C2f module FL-C2f of the backbone network.
[0077] Optionally, on the basis of the above embodiments, in some embodiments of the present invention, continue to refer to Figure 3 As shown, the infrared small and weak target detection module includes: a first C2f module, an upsampling layer Upsample, a first concatenation layer Concat, a second C2f module, a channel transposed attention mechanism module CAT, a standard convolutional layer Conv, and a second concatenation layer Concat. Based on this, an implementation manner for the infrared small and weak target detection module to perform feature fusion on the first fusion feature map and the first backbone output feature map to obtain an infrared small and weak target fusion feature map can be:
[0078] S30: Input the first fusion feature map into the first C2f module for feature processing to obtain a fourth fusion feature map.
[0079] S31: Input the fourth fused feature map into the upsampling layer for upsampling processing to obtain the fourth upsampled feature fused feature map.
[0080] S32: Input the fourth upsampled feature fused feature map and the third backbone output feature map P2 into the first splicing layer for processing to obtain the first spliced fused feature map.
[0081] S33: Input the first spliced fused feature map into the second C2f module to obtain the fifth fused feature map.
[0082] S34: Input the fifth fused feature map into the channel transposed attention mechanism module to obtain the infrared small and weak target fused feature map.
[0083] Among them, the infrared small and weak target fused feature map is used as the input of the infrared small and weak target detection module included in the detection head network, so that the infrared small and weak target detection module detects the infrared small and weak target fused feature map, thereby obtaining the infrared small and weak targets of the infrared image to be detected.
[0084] Optionally, based on the above embodiments, in some embodiments of the present invention, the second neck sub-network is used to perform feature fusion on the second fused feature map, the channel attention feature map, the second backbone output feature map P4, and the third fused feature map to obtain three target fused feature maps of different scales.
[0085] It should be noted that the third fused feature map is obtained by the output of the C2f module included in the first neck sub-network, and the second fused feature map is obtained by the output of the infrared small and weak target detection module.
[0086] In this way, in this embodiment, by introducing the infrared small and weak target detection module into the neck network included in the original YOLOV8 network, and the infrared small and weak target detection module introduces the channel transposed attention mechanism module, more channel feature information and spatial feature information of the infrared small and weak targets can be obtained, thereby improving the accuracy of detecting infrared small and weak targets.
[0087] Optionally, based on the above embodiments, in some embodiments of the present invention, continue to refer to Figure 3 As shown, the detection head network improved based on the infrared small and weak target detection head and three lightweight detection heads includes: an infrared small and weak target detection head, a first lightweight detection head, a second lightweight detection head, and a third lightweight detection head.
[0088] It should be noted that refer to Figure 8As shown, the infrared small and weak target detection head, the first lightweight detection head, the second lightweight detection head, and the third lightweight detection head all include: a classification detection head and a boundary detection head. Among them, the classification detection head includes: a lightweight convolutional layer GSConv, a standard convolutional layer Conv, a depthwise separable convolutional layer DWConv, and a standard convolutional layer Conv connected in sequence. The boundary detection head includes: a lightweight convolutional layer GSConv, a standard convolutional layer Conv, a depthwise separable convolutional layer DWConv, and a standard convolutional layer Conv connected in sequence.
[0089] Optionally, the infrared small and weak target detection head is used to detect the infrared small and weak target fusion feature map and obtain the infrared small and weak targets in the infrared image to be detected.
[0090] In this way, in this embodiment, by introducing an infrared small and weak target detection head into the detection head network included in the original YOLOV8 network, the detection of infrared small and weak targets is realized through the infrared small and weak target detection head. Furthermore, the three standard convolutional layers included in the infrared small and weak target detection head and the three detection heads in the original detection head network are replaced by the alternating action of lightweight convolution, depthwise separable convolution, and 1*1 standard convolutional layers, so as to be able to obtain more feature information of infrared small and weak targets and reduce the model calculation amount, thereby improving the detection accuracy and detection efficiency of infrared small and weak targets.
[0091] Optionally, on the basis of the above embodiment, in some embodiments of the present invention, the infrared small and weak target detection model is trained based on a super-resolution coupling loss function, and the super-resolution coupling loss function is determined according to the low-level features, high-level features, and super-resolution coupling network extracted by the backbone network.
[0092] Among them, the super-resolution coupling loss function is preset based on the super-resolution coupling network. Refer to Figure 9 As shown, the super-resolution coupling network includes: an encoding module and a decoding module. The encoding module includes: a deconvolution layer Deconv and a linear convolutional layer CSS connected in parallel, a fusion layer, and three linear convolutional layers CSS connected in sequence. The linear convolutional layer CSS includes: a convolutional layer, a spectral normalization layer SN, and a SeLU activation function layer connected in sequence. The decoding module includes: three deconvolution layers Deconv connected in sequence, or five deconvolution layers Deconv connected in sequence. The super-resolution coupling loss function can be defined by the following expression:
[0093]
[0094] Among them, i represents the i-th input infrared image, N represents the total number of multiple infrared images in the training set, and I SR represents the infrared image obtained by reconstructing the low-level features and high-level features through the super-resolution coupling network, and I xRepresents the input infrared image.
[0095] Optionally, continue to refer to Figure 3 As shown, the above-mentioned low-level features are the features output by the third CBS module of the backbone network. The low-level features are used to characterize the low-level information of each target included in the infrared image to be detected, such as contour information. The high-level features are the features output by the lightweight spatial pyramid pooling module SPPF in the backbone network. The high-level features are used to characterize the high-level semantic information of each target included in the infrared image to be detected.
[0096] Optionally, based on the above-mentioned embodiments, in some embodiments of the present invention, before executing S11, it further includes:
[0097] S01: Obtain an infrared image dataset, and perform data augmentation processing on the infrared image dataset to obtain an enhanced infrared image dataset.
[0098] Among them, the enhanced infrared image dataset includes multiple infrared images, and each infrared image includes multiple targets of different sizes. The data augmentation processing can be, for example, operations such as rotation and mirroring.
[0099] S02: Perform annotation processing on each infrared small target included in each infrared image to obtain the label data corresponding to each infrared image.
[0100] S03: Construct a training set according to multiple infrared images and the label data corresponding to the infrared images.
[0101] Optionally, based on the above-mentioned embodiments, in some embodiments of the present invention, an implementation manner of S01-S03 can be:
[0102] Obtain publicly available preset datasets, such as the SIRST dataset and the IRSTD-1k dataset. Among them, the SIRST dataset contains infrared images of 427 different scenarios and 480 targets. The size of the targets in multiple 300×300 infrared images in this dataset is approximately 4×4 pixels. The IRSTD-1k dataset includes 1000 real infrared images taken by an infrared camera with a size of 512×512. The 1000 real infrared images include 1000 targets with various different shapes and sizes, as well as a rich and cluttered background. Integrate the multiple infrared images included in the SIRST dataset and the IRSTD-1k dataset into an infrared image dataset, and augment the dataset through data augmentation operations such as rotation and mirroring to obtain an enhanced infrared image dataset. Further, use the LabelImg software to perform data annotation on multiple infrared small and weak targets included in each infrared image to obtain the label data corresponding to each infrared image. Finally, integrate multiple infrared images and the label data corresponding to the infrared images, and divide them according to 7:2:1 as the training set, test set, and validation set to obtain the training set. Among them, the validation set is used for verification after the model is trained.
[0103] Optionally, based on the above embodiments, in some embodiments of the present invention, before inputting the training set into the initial infrared small and weak target detection model for training, set the training parameters of the initial infrared small and weak target detection model. For example, the initial learning rate Ir = 0.005, the decay weight Weight_decay = 0.0005, the batch size Batch_size = 16, and the training batch Epoch = 300. However, it is not limited to this. The present invention does not specifically limit it, and those skilled in the art can set it according to the actual situation.
[0104] S04: Input the training set into the initial infrared small and weak target detection model. According to the low-level features and high-level features extracted by the backbone network, calculate the super-resolution coupling loss function through the super-resolution coupling network, and adjust the model parameters until the model converges to obtain the trained infrared small and weak target detection model.
[0105] Specifically, obtain the training set, perform initialization processing on the training parameters of the initial infrared small and weak target detection model, input the training set into the initial infrared small and weak target detection model. During the training process, input the low-level features and high-level features extracted by the backbone network into the super-resolution coupling network, calculate the super-resolution coupling loss function through the super-resolution coupling network, and adjust the model parameters until the model converges to obtain the trained infrared small and weak target detection model.
[0106] Optionally, based on the above embodiments, in some embodiments of the present invention, continue to refer to Figure 9As shown in the figure, during the training process, the low-level features and high-level features extracted by the backbone network are input into the super-resolution coupling network. One way to calculate the super-resolution coupling loss function through the super-resolution coupling network is as follows: The high-level features are input into the deconvolution layer Deconv of the encoding module for deconvolution feature extraction to obtain a deconvolution feature map. The low-level features are input into the linear convolution layer CSS of the encoding module for linear convolution feature extraction to obtain a linear convolution feature map. The deconvolution feature map and the linear convolution feature map are input into the fusion layer to obtain a fusion feature map. The fusion feature map is input into three consecutive linear convolution layers CSS for linear convolution feature extraction to obtain an encoded feature map. The obtained encoded feature map is input into the decoding module for decoding processing to obtain the infrared image I after encoding and decoding. SR .
[0107] In this way, in this embodiment, by setting the super-resolution coupling loss function and using the method of calculating the super-resolution coupling loss through the super-resolution coupling network, the training of the infrared small and weak target detection model is guided, thereby improving the ability of the infrared small and weak target detection model to extract the feature texture contours of low-resolution small targets, and thus improving the accuracy of the infrared small and weak target detection model in detecting infrared small and weak targets.
[0108] In this way, an infrared small and weak target detection method provided in this embodiment includes obtaining an infrared image to be detected, inputting the infrared image to be detected into an infrared small and weak target detection model improved based on YOLOV8 to obtain the detection result of the infrared small and weak target in the infrared image to be detected. The infrared small and weak target detection model includes a backbone network improved based on the channel transposed attention mechanism, the fast fusion C2f module, and the lightweight spatial pyramid pooling module, a neck network improved based on the infrared small and weak target detection module, and a detection head network improved based on the infrared small and weak target detection head and three lightweight detection heads. Since the channel transposed attention mechanism can improve the ability of the detection model to learn features and enhance the expression ability of the network, the fast fusion C2f module and the lightweight spatial pyramid pooling module can reduce the model calculation amount and obtain important feature information. The infrared small and weak target detection module can obtain more channel feature information and spatial feature information of infrared small and weak targets, and the infrared small and weak target detection head can detect infrared small and weak targets. Based on this, using the infrared small and weak target detection model improved based on YOLOV8 improves the accuracy and efficiency of detecting infrared small and weak targets. Further, the infrared small and weak target detection model is trained based on the super-resolution coupling loss function, and the super-resolution coupling loss function is determined according to the low-level features, high-level features extracted by the backbone network, and the super-resolution coupling network. In this way, using the method of calculating the super-resolution coupling loss through the super-resolution coupling network guides the training of the infrared small and weak target detection model, thereby further improving the accuracy of the infrared small and weak target detection model in detecting infrared small and weak targets.
[0109] It should be understood that although Figures 1 to 9 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figures 1 to 9 at least a part of the steps in
[0110] include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time and can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential either, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps. Figure 10 In one embodiment, as
[0111] shown, an infrared dim and small target detection device is provided, including: a to-be-detected infrared image acquisition module 10 and a detection module 11.
[0112] Among them, the to-be-detected infrared image acquisition module 10 is used to acquire a to-be-detected infrared image, where the to-be-detected infrared image contains at least one infrared dim and small target.
[0113] The detection module 11 is used to input the to-be-detected infrared image into an infrared dim and small target detection model improved based on YOLOV8 to obtain the infrared dim and small target detection result of the to-be-detected infrared image.
[0113] Among them, the infrared dim and small target detection model includes: a backbone network improved based on a channel transposed attention mechanism, a fast fusion C2f module, and a lightweight spatial pyramid pooling module, a neck network improved based on an infrared dim and small target detection module, and a detection head network improved based on an infrared dim and small target detection head and three lightweight detection heads. The infrared dim and small target detection model is trained based on a super-resolution coupled loss function, and the super-resolution coupled loss function is determined according to the low-level features, high-level features, and super-resolution coupled network extracted by the backbone network.
[0114] In this way, the present invention obtains the infrared image to be detected through the infrared image acquisition module to be detected, and the detection module inputs the infrared image to be detected into the infrared small target detection model improved based on YOLOV8 to obtain the infrared small target detection result of the infrared image to be detected. The infrared small target detection model includes: a backbone network improved based on a channel transposed attention mechanism, a fast fusion C2f module, and a lightweight spatial pyramid pooling module; a neck network improved based on an infrared small target detection module; and a detection head network improved based on an infrared small target detection head and three lightweight detection heads. Since the channel transposed attention mechanism can improve the ability of the detection model to learn features and enhance the expression ability of the network, the fast fusion C2f module and the lightweight spatial pyramid pooling module can reduce the model calculation amount and obtain important feature information. The infrared small target detection module can obtain more channel feature information and spatial feature information of infrared small targets, and the infrared small target detection head can detect infrared small targets. Based on this, the use of the infrared small target detection model improved based on YOLOV8 improves the accuracy and efficiency of detecting infrared small targets. Further, the infrared small target detection model is trained based on a super-resolution coupled loss function, and the super-resolution coupled loss function is determined according to the low-level features, high-level features, and super-resolution coupled network extracted by the backbone network. In this way, the method of calculating the super-resolution coupled loss using the super-resolution coupled network guides the training of the infrared small target detection model, thereby further improving the accuracy of the infrared small target detection model in detecting infrared small targets.
[0115] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium provided in the various embodiments of the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM).
[0116] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0117] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. An infrared dim and small target detection method, characterized in that Including: Obtain the infrared image to be detected, where the infrared image to be detected contains at least one small and weak infrared target; Input the infrared image to be detected into the infrared small and weak target detection model improved based on YOLOV8 to obtain the infrared small and weak target detection result of the infrared image to be detected; Among them, the infrared small and weak target detection model includes: a backbone network improved based on the channel transpose attention mechanism, a fast fusion C2f module, and a lightweight spatial pyramid pooling module, a neck network improved based on the infrared small and weak target detection module, and a detection head network improved based on the infrared small and weak target detection head and three lightweight detection heads. The infrared small and weak target detection model is trained based on the super-resolution coupling loss function, and the super-resolution coupling loss function is calculated and determined according to the low-level features, high-level features, and super-resolution coupling network extracted by the backbone network.
2. The method according to claim 1, wherein Before inputting the infrared image to be detected into the infrared small and weak target detection model improved based on YOLOV8 to obtain the infrared small and weak target detection result of the infrared image to be detected, it further includes: Obtain an infrared image dataset, and perform data augmentation processing on the infrared image dataset to obtain an enhanced infrared image dataset, where the enhanced infrared image dataset includes multiple infrared images, and each of the infrared images includes multiple targets of different sizes; Perform annotation processing on the multiple targets of different sizes included in each of the infrared images to obtain the label data corresponding to each of the infrared images; Construct a training set according to the multiple infrared images and the label data corresponding to the infrared images; Input the training set into the initial infrared small and weak target detection model, calculate the super-resolution coupling loss function through the super-resolution coupling network according to the low-level features and high-level features extracted by the backbone network, and adjust the model parameters until the model converges to obtain the trained infrared small and weak target detection model.
3. The method according to claim 2, wherein The super-resolution coupling network includes: an encoding module and a decoding module. The encoding module includes: a transposed convolution layer and a linear convolution layer connected in parallel, a fusion layer, and three linear convolution layers connected in sequence. The decoding module includes: three transposed convolution layers connected in sequence, or five transposed convolution layers connected in sequence; The super-resolution coupling loss function can be defined by the following expression: Among them, i represents the i-th input infrared image, N represents the total number of multiple infrared images in the training set, and I SR represents the infrared image obtained by reconstructing the low-level features and high-level features through the super-resolution coupling network, and I x represents the input infrared image.
4. The method according to claim 1, wherein The backbone network improved based on the channel transpose attention mechanism, a fast fusion C2f module, and a lightweight spatial pyramid pooling module includes: a standard convolution module, four CBS modules and four fast fusion C2f modules cross-connected in sequence, a lightweight spatial pyramid pooling module, and a channel transpose attention mechanism module; Among them, the fast fusion C2f module includes: a lightweight convolution layer, two bottleneck layers, a splicing layer, and a lightweight convolution layer connected in sequence. The bottleneck layer includes: a lightweight convolution layer, two standard convolution layers, and a feature fusion layer connected in sequence; The lightweight spatial pyramid pooling module includes: a lightweight convolution layer, three max pooling layers, a splicing layer, and a lightweight convolution layer connected in sequence; The channel transposed attention mechanism module includes: a channel attention module, a depthwise separable spatial convolution module, two activation function layers, a fusion module, and a linear layer. The channel transposed attention mechanism module extracts features from the spatial pooling features input by the lightweight spatial pyramid pooling module to obtain channel transposed attention features.
5. The method according to claim 4, characterized in that The channel transposed attention mechanism module extracts features from the spatial pooling features input by the lightweight spatial pyramid pooling module to obtain channel transposed attention features, including: Inputting the spatial pooling features into the channel attention module to obtain a channel attention feature map; Inputting the spatial pooling features into the depthwise separable spatial convolution module to obtain a depthwise separable spatial convolution feature map; Inputting the channel attention feature map into the first activation function connected to the channel attention module to obtain a first activated channel attention feature map; Inputting the depthwise separable spatial convolution feature map into the second activation function connected to the depthwise separable spatial convolution module to obtain a first activated depthwise separable spatial convolution feature map; Inputting the first activated channel attention feature map into the second activation function to obtain a second activated channel attention feature map; Inputting the first activated depthwise separable spatial convolution feature map into the first activation function to obtain a second activated depthwise separable spatial convolution feature map; Multiplying the channel attention feature map and the second activated depthwise separable spatial convolution feature map element-wise to obtain a first fusion feature; Multiplying the depthwise separable spatial convolution feature map and the second activated channel attention feature map element-wise to obtain a second fusion feature; Inputting the first fusion feature and the second fusion feature into the fusion module to obtain a target fusion feature; Inputting the target fusion feature into the linear layer to obtain channel transposed attention features.
6. The method according to claim 5, wherein The channel attention module includes: a channel self-attention layer and a channel attention layer connected in sequence; the step of inputting the spatial pooling features into the channel attention module to obtain a channel attention feature map includes: Inputting the spatial pooling features into the channel self-attention layer to obtain a first channel attention feature map; Inputting the first channel attention feature map into the channel attention layer to obtain a channel attention feature map; The depthwise separable spatial convolution module includes: a standard convolution layer, a depthwise separable convolution layer, and a spatial attention layer connected in sequence; the step of inputting the spatial pooling features into the depthwise separable spatial convolution module to obtain a depthwise separable spatial convolution feature map includes: Inputting the spatial pooling features into the standard convolution layer to obtain a first depthwise separable spatial convolution feature map; Inputting the first depthwise separable spatial convolution feature map into the depthwise separable convolution layer to obtain a second depthwise separable spatial convolution feature map; Inputting the second depthwise separable spatial convolution feature map into the spatial attention layer to obtain a depthwise separable spatial convolution feature map.
7. The method according to claim 6, characterized in that, The neck network improved based on the infrared small and weak target detection module includes: an infrared small and weak target detection module, a first neck sub-network, and a second neck sub-network; Among them, the first neck sub-network is used to perform feature processing on the first backbone output feature map, the second backbone output feature map, and the channel attention feature map to obtain a first fusion feature map. The first backbone output feature map is obtained by the output of the third fast fusion C2f module of the backbone network, and the second backbone output feature map is obtained by the output of the fourth fast fusion C2f module of the backbone network; The infrared small and weak target detection module is used to perform feature fusion on the first fusion feature map and the third backbone output feature map to obtain an infrared small and weak target fusion feature map. Among them, the third backbone output feature map is obtained by the output of the second fast fusion C2f module of the backbone network; The second neck sub-network is used to perform feature fusion on the second fusion feature map, the channel attention feature map, the second backbone output feature map, and the third fusion feature map to obtain three target fusion feature maps of different scales. Among them, the second fusion feature map is obtained by the output of the infrared small and weak target detection module.
8. The method according to claim 7, wherein The infrared small and weak target detection module includes: a first C2f module, an upsampling layer, a first splicing layer, a second C2f module, a channel transposed attention mechanism module, a standard convolutional layer, and a second splicing layer; the infrared small and weak target detection module is used to perform feature fusion on the first fusion feature map and the third backbone output feature map to obtain an infrared small and weak target fusion feature map, including: Input the first fusion feature map into the first C2f module for feature processing to obtain a fourth fusion feature map; Input the fourth fusion feature map into the upsampling layer for upsampling processing to obtain a fourth upsampled feature fusion feature map; Input the fourth upsampled feature fusion feature map and the third backbone output feature map into the first splicing layer for processing to obtain a first spliced fusion feature map; Input the first spliced fusion feature map into the second C2f module to obtain a fifth fusion feature map; Input the fifth fusion feature map into the channel transposed attention mechanism module to obtain an infrared small and weak target fusion feature map.
9. The method according to claim 8, characterized in that, The detection head network improved based on the infrared small and weak target detection head and three lightweight detection heads includes: an infrared small and weak target detection head, a first lightweight detection head, a second lightweight detection head, and a third lightweight detection head; the infrared small and weak target detection head, the first lightweight detection head, the second lightweight detection head, and the third lightweight detection head all include: a classification detection head and a boundary detection head; the classification detection head includes: a lightweight convolutional layer, a standard convolutional layer, a depthwise separable convolutional layer, and a standard convolutional layer connected in sequence; the boundary detection head includes: a lightweight convolutional layer, a standard convolutional layer, a depthwise separable convolutional layer, and a standard convolutional layer connected in sequence; Among them, the infrared small and weak target detection head is used to detect the infrared small and weak target fusion feature map to obtain the infrared small and weak target of the infrared image to be detected.
10. An infrared dim and small target detection device, characterized in that, Including: An infrared image to be detected acquisition module, used to acquire an infrared image to be detected, where the infrared image to be detected contains at least one infrared small and weak target; The detection module is used to input the infrared image to be detected into the improved infrared small and weak target detection model based on YOLOV8 to obtain the detection result of the infrared small and weak target in the infrared image to be detected; Among them, the infrared small and weak target detection model includes: a backbone network improved based on the channel transposed attention mechanism, the fast fusion C2f module, and the lightweight spatial pyramid pooling module, a neck network improved based on the infrared small and weak target detection module, and a detection head network improved based on the infrared small and weak target detection head and three lightweight detection heads. The infrared small and weak target detection model is trained based on the super-resolution coupled loss function, and the super-resolution coupled loss function is determined according to the low-level features, high-level features, and super-resolution coupled network extracted by the backbone network.
Citation Information
Cited By
Low-proportion infrared small target detection method and device
CN121767790A
Low-occupancy infrared small target detection method and device
CN121767790B