An Infrared Dim and Small Target Detection Method Based on Swin-Transformer and Multi-Scale Feature Fusion
By introducing the Swin-Transformer module and multi-scale feature fusion method in the Unet network, the detection difficulties of infrared weak target detection under complex background and low signal-to-noise ratio are solved, and higher detection accuracy and effectiveness are achieved, reducing the risk of target details loss.
Patent Information
- Application Number
- CN202310205449.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-06
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-03-06
AI Technical Summary
The existing infrared weak target detection methods have poor detection performance in complex backgrounds and low signal-to-noise ratio scenarios, making it difficult to accurately identify and match weak infrared targets, and the detection model is prone to lose target space details as the network deepens.
The Swin-Transformer module is used to replace the convolutional layer in the Unet network for feature extraction, and infrared image features are extracted layer by layer through multiple Swin-Transformer modules, and feature fusion is combined with multiple cross-layer feature fusion modules, and finally the target prediction results are output through the classifier.
It improves the accuracy and effectiveness of infrared weak target detection, can maintain detection performance in complex backgrounds and low signal-to-noise ratio scenarios, and reduces the risk of the detection model losing target details as the network deepens.
Smart Images

Figure CN116188944B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of infrared dim small target detection, and in particular to an infrared dim small target detection method based on Swin-Transformer and multi-scale feature fusion. Background Art
[0002] Infrared target detection and recognition is a key technology for distinguishing and locking the detected target based on the characteristic difference between the target thermal radiation and the background thermal radiation in the infrared imaging system. Due to its remarkable characteristics of working all-weather without being affected by lighting conditions, passive working, strong anti-interference ability, simple structure and small size for easy loading and concealment, infrared target detection and recognition technology has been widely used in various military and civilian technical fields such as military early warning reconnaissance, aviation guidance, long-range aircraft detection, and automatic driving. Among them, infrared target detection and recognition algorithms with high detection rate, low false alarm rate, and fast response conditions have always been important application requirements in many national defense and military fields, so they have very important research value and application prospects.
[0003] However, for actual infrared imaging systems, when the sensor is far away from the target to be detected and there are external factors such as scattering, diffraction, and atmospheric disturbances during the imaging process, the target often presents "small" scale and "weak" energy imaging characteristics in the image plane, that is, the target only occupies a very small number of pixels in the image and lacks obvious texture, shape, color and structural features. In addition, when performing infrared weak target detection and recognition in various complex application scenarios (such as sea surface, buildings, and continuous cloud scenes), it is often faced with the situation that the signal-to-noise ratio of the detected target is very low, and there is a large amount of structural noise interference in the background, which is difficult to detect. It can be seen that the imaging characteristics of infrared weak targets themselves and the complex and diverse background will bring great challenges to the detection and recognition tasks. Therefore, studying how to accurately, quickly, and stably detect infrared weak targets and perform fast matching has always been an important technical problem that needs to be solved.
[0004] The existing infrared dim and small target detection methods can be divided into two categories: traditional image processing methods based on physical models and data-driven methods based on deep learning. Among them, the target detection method based on a single-frame image dominates in traditional methods due to its advantages such as low complexity, high real-time performance, and easy hardware implementation. Specifically, it includes three image processing methods based on target features, background features, and morphological analysis, respectively considering means such as expanding the contrast between the target and the background, suppressing background interference, and using the morphological features of the target in the image to lock the target position area and complete the target detection and recognition process. However, the existing methods often can only obtain the local spatial features of infrared dim and small targets and lack the semantic distinguishability between the targets and other interfering backgrounds, resulting in poor detection performance of such methods in scenarios such as complex backgrounds and low signal-to-noise ratios, that is, the detection accuracy and effectiveness are not good. Therefore, how to design a method that can improve the detection accuracy and effectiveness of infrared dim and small targets is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] Aiming at the deficiencies of the above-mentioned existing technologies, the technical problem to be solved by the present invention is: how to provide an infrared dim and small target detection method based on Swin-Transformer and multi-scale feature fusion, which can ensure the detection performance of the detection model in scenarios such as complex backgrounds and low signal-to-noise ratios, and can reduce the risk of the detection model losing the spatial details of infrared dim and small targets as the network deepens, so as to improve the detection accuracy and effectiveness of infrared dim and small targets and provide a new idea for infrared dim and small target detection.
[0006] To solve the above technical problems, the present invention adopts the following technical solutions:
[0007] An infrared dim and small target detection method based on Swin-Transformer and multi-scale feature fusion, comprising:
[0008] S1: Introduce a Swin-Transformer module into the Unet network to replace the original convolutional layer for feature extraction to construct a target detection model;
[0009] S2: Obtain the infrared image to be detected;
[0010] S3: Input the infrared image to be detected into the trained target detection model and output the target prediction result;
[0011] The target detection model first extracts the feature information of the infrared image layer by layer through multiple Swin-Transformer modules to generate feature maps of multiple scales; then, through multiple cross-layer feature fusion modules, starting from the feature map of the highest scale, it fuses the feature maps of each scale in turn to generate corresponding multi-layer fused feature maps; finally, the multi-layer fused feature maps are input into the classifier for normalization processing and the corresponding target prediction results are output;
[0012] S4: Use the target prediction results output by the target detection model as the detection results of small and weak targets in the infrared image to be detected.
[0013] Preferably, the target detection model includes multiple Swin-Transformer modules connected end to end in sequence, and multiple cross-layer feature fusion modules connected end to end in sequence;
[0014] Among two adjacent Swin-Transformer modules, the output of the previous Swin-Transformer module is used as the input of the next Swin-Transformer module; among two adjacent cross-layer feature fusion modules, the output of the previous cross-layer feature fusion module is used as the input of the next cross-layer feature fusion module; the output of the last Swin-Transformer module is used as the input of the first cross-layer feature fusion module;
[0015] The first cross-layer feature fusion module uses the outputs of the last two Swin-Transformer modules as the input;
[0016] Except for the first cross-layer feature fusion module, other cross-layer feature fusion modules use the output of the previous cross-layer feature fusion module and the output of its corresponding Swin-Transformer module as the input.
[0017] Preferably, in the target detection model, multiple encoding layers connected end to end in sequence are arranged before the first Swin-Transformer module;
[0018] Among two adjacent encoding layers, the output of the previous encoding layer is used as the input of the next encoding layer; the input of the first encoding layer is the infrared image input by the target detection model, and the output of the last encoding layer is used as the input of the first Swin-Transformer module;
[0019] The cross-layer feature fusion module corresponding to the encoding layer uses the output of the previous cross-layer feature fusion module and the output of its corresponding encoding layer as the input; the output of the last cross-layer feature fusion module is the multi-layer fused feature map used as the input of the classifier.
[0020] Preferably, an input convolutional layer and an output convolutional layer are correspondingly arranged at the input end and the output end of the target detection model;
[0021] The input of the input convolutional layer is the infrared image input by the target detection model, which is used to increase the number of channels of the infrared image and use the infrared image with increased channels as the input of the first encoding layer;
[0022] The input of the output convolutional layer is the multi-layer fusion feature map output by the last encoding layer, which is used to restore the number of channels and the size of the multi-layer fusion feature map to be the same as those of the input infrared image, and use the output multi-layer fusion feature map as the input of the final classifier.
[0023] Preferably, in the Swin-Transformer module, when extracting features from a feature map I with an input size of M×N and a channel number of C, the following steps are included:
[0024] 1) Perform a layer normalization operation on the input feature map I, and perform normalization processing on the data in the channel dimension to obtain an output result of I LN ;
[0025] The formula description is:
[0026] I LN = LN(I);
[0027] 2) Calculate the feature weights based on the multi-head attention mechanism for the feature map I after layer normalization processing LN to obtain I Attention ;
[0028] The formula description is:
[0029] I Attention = MSA(I LN );
[0030] In the calculation of the multi-head attention mechanism, three weight matrices Q, K, and V with the same size as the input feature map I LN are respectively introduced;
[0031] Among them: Q = I LN P Q , K = I LN P K , V = I LN P V ;
[0032] In the formula: P Q , P K and P V are respectively the shared weight matrices under different local windows and are parameters that can be learned;
[0033] After calculating the weight matrices Q, K, and V, calculate I according to the attention mechanism calculation formula of the Transformer Attention ;
[0034] The formula is described as:
[0035]
[0036] In the formula: d represents the size of the input feature; b represents the learnable position encoding parameter;
[0037] 3) Perform a residual connection between the original input feature map I and the I obtained through the calculation based on the multi-head attention mechanism Attention to obtain the intermediate feature F, which is used as the input for the next layer structure;
[0038] The formula is described as:
[0039] F = I + I Attention ;
[0040] 4) Perform layer normalization LN operation on the obtained intermediate feature F, then adjust it with a multi-layer perceptron, and finally connect the adjusted result with the intermediate feature F through a residual network to obtain the output result S;
[0041] The formula is described as:
[0042] S = MLP(LN(F)) + F;
[0043] 5) Perform an image patch merging operation on the output result S, and use image patch stitching, layer normalization, and channel linear mapping operations to reduce its size by half, becoming the number of channels is doubled to 2C, and finally the corresponding feature map is output.
[0044] Preferably, the input of the cross-layer feature fusion module is two feature maps, where the relatively high-scale feature map is Y and the relatively low-scale feature map is X;
[0045] In the cross-layer feature fusion module, first upsample Y, and adjust the number of channels of the upsampled Y feature through pointwise convolution operation to generate the first feature map; then adjust the number of channels of X to be the same as that of the first feature Figure 1 through pointwise convolution operation, and then perform normalization processing using the Sigmoid activation function to generate the second feature map; then use the second feature map as the weight coefficient to perform multiplication operation with the first feature map to generate the first fusion map; finally, add the first fusion map and X to generate the corresponding fusion feature map.
[0046] Preferably, the formula for the cross-layer feature fusion module to generate the fusion feature map is described as follows:
[0047]
[0048] In the formula: Z represents the generated fused feature map; Y represents the feature map of a relatively high scale; X represents the feature map of a relatively low scale; PWConv represents the pointwise convolution operation; Sig represents the Sigmoid activation function operation; represents the pointwise addition operation of the feature maps corresponding to the same channels; represents the pointwise multiplication operation of the feature maps corresponding to the same channels; Up represents the image upsampling operation.
[0049] Preferably, during the sample data training stage of the object detection model, slice-assisted data augmentation operation is performed: First, each original sample image in the sample dataset is sliced into overlapping image patches; then, by fixing the aspect ratio of the image patches, the sizes of the sliced image patches are adjusted to scale them proportionally to the same size as the original sample data, so as to obtain new augmented sample images; finally, the new augmented sample images are added to the sample dataset to participate in the training and parameter optimization of the object detection model.
[0050] Preferably, during the inference of the object detection model, slice-assisted inference operation is performed: First, the infrared image is segmented into blocks using the slice segmentation method to obtain several images to be detected; then, under the condition of a fixed aspect ratio, the size of each image to be detected is adjusted to scale it proportionally to the same size as the original image; then, each image to be detected is input into the trained object detection model for object detection to obtain the predicted output results of the object at multiple different positions; finally, post-processing is performed on all the predicted output results, and the NMS non-maximum suppression strategy is used to filter the predicted output results at overlapping positions, and only the predicted result with the highest probability is retained at the same position.
[0051] Preferably, the object loss function during the training of the object detection model is as follows:
[0052]
[0053] In the formula: T represents the object loss function; (i, j) represents any coordinate position in the corresponding infrared image; p represents the prediction result finally output by the final network model; y represents the label of the corresponding infrared image; p i,j represents the predicted value output by the object detection model at the position (i, j) in the image, and its value ranges from (0, 1); y i,j represents the true normalized gray value at the position (i, j) in the image, representing the label result of the infrared image at the corresponding position.
[0054] Compared with the prior art, the infrared small and weak target detection method based on Swin-Transformer and multi-scale feature fusion in the present invention has the following beneficial effects:
[0055] In the present invention, a Swin-Transformer module is introduced into the Unet network to replace the original convolutional layer for feature extraction, and the feature information of the infrared image is extracted layer by layer through multiple Swin-Transformer modules to generate feature maps of multiple scales. On the one hand, through the Swin-Transformer module, the potential feature information of the target is fully mined under a larger receptive field, and the feature information of each scale of the target is extracted, which can meet the semantic discriminability between the target and other interference backgrounds, and further ensure the detection performance of the detection model in scenarios such as complex backgrounds and low signal-to-noise ratios. On the other hand, the present invention constructs a deeper network through multiple Swin-Transformer modules, which can provide better semantic features and understanding of the scene context, help to better solve the ambiguity problem caused by the target and background interference, and can adapt to the characteristics that the infrared small and weak target lacks semantic features and the target features are easily lost as the number of network layers increases, thereby improving the accuracy of infrared small and weak target detection.
[0056] On the basis of extracting feature maps of multiple scales through multiple Swin-Transformer modules, in order to better fuse the local spatial information and global semantic information of the infrared small and weak target, multiple bottom-up cross-layer feature fusion modules are used as the decoder of the target detection model to re-fuse the shallow local information and deep semantic information obtained at each scale, so as to retain the features of the infrared small and weak target from the complex background, reduce the risk that the detection model loses the spatial details of the infrared small and weak target as the network deepens, thereby improving the effectiveness of infrared small and weak target detection and providing a new idea for infrared small and weak target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to make the purpose, technical solution and advantages of the invention clearer, the present invention will be further described in detail below with reference to the drawings, where:
[0058] Figure 1 is the logic block diagram of the infrared small and weak target detection method based on Swin-Transformer and multi-scale feature fusion;
[0059] Figure 2 is the network structure diagram of the target detection model (UST-Net);
[0060] Figure 3 is the network structure diagram of the cross-layer feature fusion module (AFM);
[0061] Figure 4 It is a schematic diagram of the principle process for data augmentation;
[0062] Figure 5 It is a schematic diagram of the principle process for auxiliary reasoning;
[0063] Figure 6 They are some typical image scenes in the SIRST dataset;
[0064] Figure 7 They are the comparison of the detection results of various methods: Figure 7 (a) is the original infrared image, Figure 7 (b) is the detection result of MPCM, Figure 7 (c) is the detection result of NIPPS, Figure 7 (d) is the detection result of TBC-Net, Figure 7 (e) is the detection result of ALC-Net, Figure 7 (f) is the detection result of the object detection model (UST-Net). Specific implementation manners
[0065] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the drawings herein can be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present invention provided herein is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0066] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the figures, or the orientation or positional relationship in which the inventive product is customarily placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation of the present invention. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance. In addition, terms such as "horizontal" and "vertical" do not mean that the components are required to be absolutely horizontal or hanging vertically, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined. In the description of the present invention, it should also be noted that unless otherwise clearly specified and defined, the terms "set", "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0067] The following is a more detailed description through specific embodiments:
[0068] Embodiment:
[0069] This embodiment discloses an infrared small and weak target detection method based on Swin-Transformer and multi-scale feature fusion. As Figure 1 shown, the infrared small and weak target detection method based on Swin-Transformer and multi-scale feature fusion includes:
[0070] S1: Introduce a Swin-Transformer module into the Unet network to replace the original convolutional layer for feature extraction to form a target detection model (subsequently also referred to as UST-Net);
[0071] S2: Obtain the infrared image to be detected;
[0072] S3: Input the infrared image to be detected into the trained target detection model and output the target prediction result;
[0073] Combined with Figure 2As shown, the object detection model first extracts the feature information of the infrared image layer by layer through multiple Swin-Transformer modules, generating feature maps of multiple scales; then, through multiple cross-layer feature fusion modules (subsequently also referred to as AFM), starting from the feature map of the highest scale, it fuses the feature maps of each scale in turn to generate corresponding multi-layer fused feature maps; finally, the multi-layer fused feature maps are input into the classifier for normalization processing and the corresponding object prediction results are output;
[0074] S4: Use the object prediction result output by the object detection model as the detection result of the small and weak objects in the infrared image to be detected.
[0075] In the present invention, a Swin-Transformer module is introduced into the Unet network to replace the original convolutional layer for feature extraction, and the feature information of the infrared image is extracted layer by layer through multiple Swin-Transformer modules to generate feature maps of multiple scales. On the one hand, through the Swin-Transformer module, the present invention fully excavates the potential feature information of the object under a larger receptive field and extracts the feature information of each scale of the object, which can meet the semantic distinguishability between the object and other interfering backgrounds, and thus can ensure the detection performance of the detection model in scenarios such as complex backgrounds and low signal-to-noise ratios; on the other hand, the present invention constitutes a deeper network through multiple Swin-Transformer modules, which can provide better semantic features and understanding of the scene context, helps to better solve the ambiguity problem caused by object and background interference, and can adapt to the characteristics of the lack of semantic features of small and weak infrared objects and the easy loss of object features as the network depth increases, thereby improving the accuracy of small and weak infrared object detection.
[0076] On the basis of extracting feature maps of multiple scales through multiple Swin-Transformer modules, in order to better fuse the local spatial information and global semantic information of small and weak infrared objects, multiple cross-layer feature fusion modules from bottom to top are used as the decoder of the object detection model to re-fuse the shallow local information and deep semantic information obtained at each scale, so as to retain the features of small and weak infrared objects from the complex background, reduce the risk of the detection model losing the spatial details of small and weak infrared objects as the network deepens, thereby improving the effectiveness of small and weak infrared object detection and providing a new idea for small and weak infrared object detection.
[0077] Combined with Figure 2 As shown, the object detection model includes multiple Swin-Transformer modules connected end to end in sequence, and multiple cross-layer feature fusion modules connected end to end in sequence;
[0078] In this embodiment, a cross-layer feature fusion module is embedded in each decoding layer.
[0079] In two adjacent Swin-Transformer modules, the output of the previous Swin-Transformer module serves as the input of the subsequent Swin-Transformer module; in two adjacent cross-layer feature fusion modules, the output of the previous cross-layer feature fusion module serves as the input of the subsequent cross-layer feature fusion module; the output of the last Swin-Transformer module serves as the input of the first cross-layer feature fusion module.
[0080] The first cross-layer feature fusion module takes the outputs of the last two Swin-Transformer modules as inputs.
[0081] Except for the first cross-layer feature fusion module, other cross-layer feature fusion modules take the output of the previous cross-layer feature fusion module and the output of its corresponding Swin-Transformer module as inputs.
[0082] In this embodiment, the number of cross-layer feature fusion modules is one less than the number of Swin-Transformer modules, and the first cross-layer feature fusion module corresponds to the penultimate Swin-Transformer module, the second cross-layer feature fusion module corresponds to the third-to-last Swin-Transformer module, the third cross-layer feature fusion module corresponds to the fourth-to-last Swin-Transformer module, and so on.
[0083] In the present invention, the structural design of the Swin-Transformer modules and the cross-layer feature fusion modules enables multiple cross-layer feature fusion modules to re-fuse the shallow local information and deep semantic information obtained at each scale, thereby being able to retain the features of infrared small targets from complex backgrounds, reducing the risk that the detection model loses the spatial details of infrared small targets as the network deepens, and thus being able to further improve the effectiveness of infrared small target detection.
[0084] Specifically, in the target detection model, a plurality of encoding layers connected end to end in sequence are arranged before the first Swin-Transformer module;
[0085] In two adjacent encoding layers, the output of the previous encoding layer serves as the input of the subsequent encoding layer; the input of the first encoding layer is the infrared image input to the target detection model, and the output of the last encoding layer serves as the input of the first Swin-Transformer module.
[0086] The cross-layer feature fusion module corresponding to the encoding layer takes the output of the previous cross-layer feature fusion module and the output of its corresponding encoding layer as inputs; the output of the last cross-layer feature fusion module is the multi-layer fusion feature map used as the input to the classifier.
[0087] In the present invention, multiple encoding layers are arranged before the first Swin-Transformer module, and the convolutional layers of the encoding layers are used to extract the low-level feature maps of the infrared images. The convolutional layers have the characteristics of fast calculation speed and high efficiency, and can ensure the extraction effect of the low-level feature maps, thereby improving the overall detection efficiency of the target detection model.
[0088] Specifically, an input convolutional layer and an output convolutional layer are correspondingly arranged at the input end and the output end of the target detection model;
[0089] The input of the input convolutional layer is the infrared image input to the target detection model, which is used to increase the number of channels of the infrared image and use the infrared image with increased number of channels as the input to the first encoding layer;
[0090] The input of the output convolutional layer is the multi-layer fusion feature map output by the last encoding layer, which is used to restore the number of channels and the size of the multi-layer fusion feature map to be the same as those of the input infrared image, and use the output multi-layer fusion feature map as the input to the final classifier.
[0091] In the specific implementation process, in the Swin-Transformer module, when extracting features from the feature map I with an input size of M×N and a channel number of C, the following steps are included:
[0092] 1) Perform a layer normalization (Layer Norm, LN) operation on the input feature map I, and standardize the data in the channel dimension to obtain the output result as I LN ;
[0093] The formula description is:
[0094] I LN = LN(I);
[0095] 2) Calculate the feature weights based on the multi-head self-attention mechanism (Multi-head Self-Attention, MSA) for the feature map I after layer normalization processing LN to obtain I Attention ;
[0096] The formula description is:
[0097] I Attention = MSA(I LN );
[0098] In the calculation of the multi-head attention mechanism MSA, three weight matrices Q, K, and V with the same size as the input feature map I are respectively introduced; LN where: Q = I
[0099] P LN P Q , K = I LN P K , and V = I LN P V ;
[0100] In the formula: P Q , P K , and P V are the shared weight matrices under different local windows and are learnable parameters;
[0101] After calculating the weight matrices Q, K, and V, calculate I according to the attention mechanism calculation formula of the Transformer; Attention ;
[0102] The formula description is:
[0103]
[0104] In the formula: d represents the size of the input feature; b represents the learnable position encoding parameter;
[0105] 3) Perform a residual connection between the original input feature map I and I obtained through the calculation based on the multi-head attention mechanism to obtain the intermediate feature F as the input of the next layer structure; Attention The formula description is:
[0106] F = I + I
[0107] Attention ;
[0108] 4) Perform a layer normalization LN operation on the obtained intermediate feature F, then adjust it with a multi-layer perceptron (MLP, Multilayer Perceptron), and finally connect the adjusted result with the intermediate feature F through a residual network to obtain the output result S;
[0109] The formula description is:
[0110]
[0111] S = MLP(LN(F)) + F;
[0111] 5) Perform an image patch merging operation on the output result S, and use image patch stitching, layer normalization, and channel linear mapping operations to reduce its size by half to become and double the number of channels to become 2C, and finally output the corresponding feature map.
[0112] The Swin-Transformer module of the present invention can fully exploit the potential feature information of the target under a larger receptive field and extract the feature information of each scale of the target, which can meet the semantic discriminability between the target and other interfering backgrounds, thereby improving the detection performance of the detection model in scenarios such as complex backgrounds and low signal-to-noise ratios.
[0113] Combined with Figure 3 As shown, the input of the cross-layer feature fusion module is two feature maps, where the relatively high-scale feature map is Y and the relatively low-scale feature map is X;
[0114] In the cross-layer feature fusion module, first, Y is upsampled, and the number of feature channels of the upsampled Y is adjusted through pointwise convolution operation to generate the first feature map; then, the number of feature channels of X is adjusted to be the same as that of the first feature Figure 1 map through pointwise convolution operation, and then normalized using the Sigmoid activation function to generate the second feature map; then, the second feature map is used as the weight coefficient to perform multiplication operation with the first feature map to generate the first fusion map; finally, the first fusion map is added to X to generate the corresponding fusion feature map.
[0115] The formula for the cross-layer feature fusion module to generate the fusion feature map is described as follows:
[0116]
[0117] In the formula: Z represents the generated fusion feature map; Y represents the relatively high-scale feature map; X represents the relatively low-scale feature map; PWConv represents the pointwise convolution operation; Sig represents the Sigmoid activation function operation; represents the pointwise addition operation of the feature maps corresponding to the same channels; represents the pointwise multiplication operation of the feature maps corresponding to the same channels; Up represents the image upsampling operation.
[0118] The present invention uses multiple bottom-up cross-layer feature fusion modules as the decoder of the target detection model to re-fuse the shallow local information and deep semantic information obtained at each scale, thereby being able to retain the features of infrared small targets from complex backgrounds and reducing the risk that the detection model loses the spatial details of infrared small targets as the network deepens, thus further improving the effectiveness of infrared small target detection.
[0119] Combined with Figure 4As shown in the figure, a slicing-assisted data augmentation operation is performed during the sample data training phase of the target detection model: First, each original sample image in the sample dataset is sliced into overlapping image patches; then, by fixing the aspect ratio of the image patches, the sizes of the sliced image patches are adjusted so that they are scaled proportionally to the same size as the original sample data, thereby obtaining new augmented sample images; finally, the new augmented sample images are added to the sample dataset to participate in the training and parameter optimization of the target detection model.
[0120] In the present invention, slicing-assisted data augmentation is performed on the sample dataset during the sample data training phase of the target detection model, enabling an increase in the sample capacity of the sample dataset, improving the performance of the sample data, and facilitating the training of a target detection model with better performance.
[0121] Combined with Figure 5 As shown in the figure, a slicing-assisted inference operation is performed when the target detection model is performing inference: First, the infrared image is segmented into blocks using the slicing segmentation method to obtain several images to be detected; then, the size of each image to be detected is adjusted while fixing the aspect ratio so that its size is scaled proportionally to the same size as the original image; next, each image to be detected is separately input into the trained target detection model for target detection to obtain the predicted output results of the target at multiple different positions; finally, post-processing is performed on all the predicted output results, and the NMS (Non-Maximum Suppression) strategy is used to filter the predicted output results at overlapping positions, and only the predicted result with the highest probability is retained at the same position.
[0122] In the present invention, slicing-assisted inference is performed on the network model when the target detection model is performing inference, enabling an improvement in the final effect of target detection and facilitating the obtaining of a target detection model with better performance.
[0123] In the specific implementation process, in order to better handle the class imbalance problem between infrared small and weak targets and the background when optimizing the target detection model, this patent application constructs a loss function based on the Soft-IoU metric to handle such highly imbalanced segmentation tasks. The corresponding calculation formula for Soft-IoU is:
[0124]
[0125] During training, it is desired that the value of Soft-IoU is as large as possible. To unify the optimization form, the target loss function during the training of the target detection model in this patent application is as follows:
[0126]
[0127] Where: T represents the target loss function; (i,j) represents any coordinate position in the corresponding infrared image; p represents the prediction result finally output by the final network model; y represents the label of the corresponding infrared image; p i,j represents the predicted value output by the target detection model at the position (i,j) in the image, and its value ranges from (0,1). The larger the value, the greater the probability that the network model believes that there is a target at this point; y i,j represents the true normalized gray value at the position (i,j) in the image, indicating the label result of the infrared image at the corresponding position.
[0128] Through the above target loss function, the present invention can better handle the class imbalance problem between infrared small and weak targets and the background when optimizing the target detection model, and thus can better train the target detection model, which is beneficial to training a target detection model with better performance.
[0129] To better illustrate the advantages of the technical solution of this patent application, the following experiment is disclosed in this embodiment.
[0130] 1. Experimental design
[0131] To evaluate the performance of the infrared small and weak target detection model (UST-Net) proposed in this patent application, we tested this model on the public SIRST dataset (from DAI Y, WU Y, ZHOU F, et al. Attentional local contrast networks for infrared small target detection), and compared the test results with other typical infrared small and weak target detection methods. The SIRST dataset contains 427 representative images from hundreds of real worlds and 480 instances in different scenarios, such as Figure 6 shown. It can be seen that many infrared small targets are very dim and buried in a complex background with severe clutter. In addition, only 35% of the targets in this dataset contain the brightest pixels in the image. Therefore, methods purely based on the target saliency hypothesis or simply performing simple threshold processing on the original image may lead to poor detection effects.
[0132] The test environment corresponding to the method proposed in this experiment is ubuntu20.04, and the GPU model is NVIDIA GeForce GTX3080Ti 12G. When training the model, the Adam optimizer is used for training, the initial learning rate is set to 5e-4, the Batchsize is set to 16, and the value of the training round Epoch is set to 50. For the convenience of comparison, the training image size is uniformly fixed to a resolution of 512×512.
[0133] In terms of data augmentation, in addition to using the slicing assistance technique mentioned above, this experiment also used means such as flipping transformation, contrast adjustment, width-height distortion, and adding Gaussian noise to improve the generalization of training samples. Since this experiment uses a target segmentation-based method to predict the location of infrared small and weak targets, in order to more objectively and realistically evaluate the performance of this network model, this experiment selects the intersection over union (IoU) and normalized IoU (nIoU), which are commonly used in image segmentation evaluation, as the two indicators for algorithm evaluation. Their respective calculation expressions are as follows:
[0134]
[0135]
[0136] In the above formulas: N is the number of training samples, TP represents the targets correctly predicted by the model, T represents the true targets in the samples, and P represents all the targets predicted by the model. Using the above IoU and nIoU indicators to evaluate the UST-Net model can respectively reflect the segmentation effects of infrared small and weak targets with larger and smaller sizes.
[0137] 2. Target detection effect
[0138] To specifically verify the actual effect of UST-Net, this experiment compares this method with four other types of infrared small target detection and segmentation methods, and then calculates the IoU and nIoU results of each respectively. Among these four methods, there are two non-deep learning methods, MPCM (from WEI Y, YOU X, LI H. Multiscale patch-based contrast measure for small infrared target detection) and NIPPS (from DAI Y, WU Y, SONG Y, et al. Non-negative infrared patch-image model: Robust target-background separation via partial sum minimization of singular values), and two deep learning methods, TBC-Net (from ZHAO M, CHENG L, YANG X, et al. TBC-Net: A real-time detector for infrared small target detection using semantic Constraint) and ALC-Net (from DAI Y, WU Y, ZHOU F, et al. Attentional local contrast networks for infrared small target detection). The corresponding target detection and segmentation results are as Figure 7 shown.
[0139] In Figure 7 , we selected 5 different actual infrared scenes for experimental tests respectively. It can be seen that for general non-deep learning methods such as MPCM and NIPPS, there are obvious false detections or missed detections in the detection results; for deep learning training networks such as TBC-Net and ALC-Net, although there are no obvious false detections and missed detections in the finally obtained detection results, the target detection and segmentation results are still not fine enough. Due to the interference of image noise or background highlights during target detection and segmentation, there are broken or incomplete situations at some local positions of the target in the segmentation map. However, as can be seen from the results, the UST-Net proposed in this patent application can more completely reflect the overall diffusion characteristics of infrared small targets, and the detection and segmentation results are more complete and continuous.
[0140] 3. Index Performance Analysis
[0141] To further quantitatively evaluate the performance differences between the method proposed in this patent application and other related methods, the IoU and nIoU metrics are calculated respectively for the object detection results corresponding to each algorithm in Figure 7 and the processing frame rate FPS of each algorithm is compared and analyzed. Finally, the average values of the various metrics obtained in 5 scenarios are taken, and the final experimental analysis results are shown in Table 1 below.
[0142] Table 1 Comparison of test performance of different methods
[0143]
[0144] As can be seen from Table 1, compared with the other four infrared small and weak target detection methods, the UST-Net proposed in this patent application has a very large improvement in the two metrics of IoU and nIoU.
[0145] Specifically, compared with MPCM, NIPPS, TBC-Net, and ALC-Net, UST-Net improves the IoU metric by 123.7% (from 0.334 to 0.747), 75.8% (from 0.425 to 0.747), 11.2% (from 0.672 to 0.747), and 3.2% (from 0.724 to 0.747) respectively, and improves the nIoU metric by 89.4% (from 0.397 to 0.752), 31.2% (from 0.573 to 0.752), 6.2% (from 0.708 to 0.752), and 2.2% (from 0.736 to 0.752) respectively. Although UST-Net is not as fast as the other two deep learning methods in terms of algorithm speed, its frame rate can still reach 60 - 70 fps, which is faster than the other two non-deep learning methods, and can meet the real-time target detection and recognition process of the infrared detection sequence. Thus, the superiority of this method in various performances can be proved.
[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit the technical solutions. Those of ordinary skill in the art should understand that any modifications or equivalent replacements made to the technical solutions of the present invention without departing from the spirit and scope of the present technical solutions shall be covered by the scope of the claims of the present invention.
Claims
1. An infrared small and weak target detection method based on Swin-Transformer and multi-scale feature fusion, characterized in that, Including: S1: Introduce the Swin-Transformer module into the Unet network to replace the original convolutional layer for feature extraction to form an object detection model; S2: Obtain the infrared image to be detected; S3: Input the infrared image to be detected into the trained object detection model and output the object prediction result; The object detection model first extracts the feature information of the infrared image layer by layer through multiple Swin-Transformer modules to generate feature maps of multiple scales; then, starting from the feature map of the highest scale, it fuses the feature maps of each scale in turn through multiple cross-layer feature fusion modules to generate corresponding multi-layer fusion feature maps; finally, the multi-layer fusion feature maps are input into the classifier for normalization processing and the corresponding object prediction results are output; The input of the cross-layer feature fusion module is two feature maps, where the relatively high-scale feature map is Y and the relatively low-scale feature map is X; In the cross-layer feature fusion module, first perform upsampling on Y and adjust the number of feature channels of the upsampled Y through pointwise convolution operations to generate the first feature map; then adjust the number of feature channels of X to be the same as that of the first feature map through pointwise convolution operations, and then use the Sigmoid activation function for normalization processing to generate the second feature map; then use the second feature map as the weight coefficient to perform multiplication operations with the first feature map to generate the first fusion map; finally, add the first fusion map to X to generate the corresponding fusion feature map; The formula for the cross-layer feature fusion module to generate the fusion feature map is described as follows: Where: Z represents the generated fused feature map; Y represents the feature map of a relatively high scale; X represents the feature map of a relatively low scale; PWConv represents the pointwise convolution operation; Sig represents the Sigmoid activation function operation; represents the pointwise addition operation of the feature maps corresponding to the same channels; represents the pointwise multiplication operation of the feature maps corresponding to the same channels; Up represents the image upsampling operation; S4: Use the object prediction result output by the object detection model as the detection result of the small and weak objects in the infrared image to be detected.
2. The infrared dim and small target detection method based on Swin-Transformer and multi-scale feature fusion according to claim 1, characterized in that: The object detection model includes multiple Swin-Transformer modules connected end to end in sequence, and multiple cross-layer feature fusion modules connected end to end in sequence; Among adjacent two Swin-Transformer modules, the output of the previous Swin-Transformer module is used as the input of the next Swin-Transformer module; among adjacent two cross-layer feature fusion modules, the output of the previous cross-layer feature fusion module is used as the input of the next cross-layer feature fusion module; the output of the last Swin-Transformer module is used as the input of the first cross-layer feature fusion module; The first cross-layer feature fusion module uses the outputs of the last two Swin-Transformer modules as the input; Except for the first cross-layer feature fusion module, other cross-layer feature fusion modules use the output of the previous cross-layer feature fusion module and the output of its corresponding Swin-Transformer module as the input.
3. The infrared small and weak target detection method based on Swin-Transformer and multi-scale feature fusion according to claim 2, characterized in that: In the object detection model, multiple encoding layers connected end to end in sequence are set before the first Swin-Transformer module; In two adjacent encoding layers, the output of the previous encoding layer serves as the input of the subsequent encoding layer; the input of the first encoding layer is the infrared image input to the object detection model, and the output of the last encoding layer serves as the input of the first Swin-Transformer module; For the cross-layer feature fusion module corresponding to the encoding layer, the output of the previous cross-layer feature fusion module and the output of its corresponding encoding layer are used as inputs; the output of the last cross-layer feature fusion module is the multi-layer fusion feature map serving as the input of the classifier.
4. The infrared small and weak target detection method based on Swin-Transformer and multi-scale feature fusion according to claim 3, characterized in that: An input convolutional layer and an output convolutional layer are correspondingly arranged at the input end and the output end of the object detection model; The input of the input convolutional layer is the infrared image input to the object detection model, which is used to increase the number of channels of the infrared image, and the infrared image with increased channels is used as the input of the first encoding layer; The input of the output convolutional layer is the multi-layer fusion feature map output by the last encoding layer, which is used to restore the number of channels and the size of the multi-layer fusion feature map to be consistent with the input infrared image, and the output multi-layer fusion feature map is used as the input of the final classifier.
5. The infrared small and weak target detection method based on Swin-Transformer and multi-scale feature fusion according to claim 2, wherein, In the Swin-Transformer module, when performing feature extraction on the feature map I with an input size of M×N and a channel number of C, the following steps are included: 1) Perform layer normalization on the input feature map I, standardize the data in the channel dimension, and obtain the output result as I LN ; The formula is described as: I LN = LN(I); 2) Feature map I after layer regularization processing LN Calculate the feature weights based on the multi-head attention mechanism to obtain I Attention ; The formula is described as: I Attention = MSA(I LN ); In the calculation of the multi-head attention mechanism, three weight matrices Q, K, and V with the same size as the input feature map I are respectively introduced; LN Where: Q = I LN P Q , K = I LN P K , V = I LN P V ; Where: P Q , P K and P V are respectively the shared weight matrices under different local windows, which are parameters that can be learned; After calculating the weight matrices Q, K, and V, calculate I according to the attention mechanism calculation formula of the Transformer Attention ; The formula is described as: In the formula: d represents the size of the input feature; b represents the learnable position encoding parameter; 3) Residually connect the original input feature map I with I obtained through calculation based on the multi-head attention mechanism Attention to obtain the intermediate feature F, which serves as the input to the next layer structure; The formula is described as: F = I + I Attention ; 4) Perform layer normalization LN operation on the obtained intermediate feature F, then adjust it with a multi-layer perceptron, and finally connect the adjusted result with the intermediate feature F through a residual network to obtain the output result S; The formula is described as: S = MLP(LN(F)) + F; 5) Perform an image block merging operation on the output result S, and use image block stitching, layer regularization, and channel linear mapping operations to reduce its size by half, becoming doubling the number of channels to become 2C, and finally outputting the corresponding feature map.
6. The infrared small and weak target detection method based on Swin-Transformer and multi-scale feature fusion according to claim 1, characterized in that, During the training stage of the sample data of the object detection model, slice-assisted data augmentation operation is performed: First, each original sample image in the sample data set is sliced into overlapping image patches; then, by fixing the aspect ratio of the image patches, the sizes of the sliced image patches are adjusted to scale them proportionally to the same size as the original sample data, so as to obtain new enhanced sample images; finally, the new enhanced sample images are added to the sample data set to participate in the training and parameter optimization of the object detection model.
7. The infrared small and weak target detection method based on Swin-Transformer and multi-scale feature fusion according to claim 1, characterized in that During the inference of the object detection model, slice-assisted inference operation is performed: First, the infrared image is segmented into blocks using the slice segmentation method to obtain several images to be detected; then, under the condition of a fixed aspect ratio, the size of each image to be detected is adjusted to scale its size proportionally to the same size as the original image; then, each image to be detected is respectively input into the trained object detection model for object detection to obtain the predicted output results of the object at multiple different positions; finally, post-processing is performed on all the predicted output results, and the non-maximum suppression strategy NMS is used to filter the predicted output results at overlapping positions, and only the predicted result with the highest probability is retained at the same position.
8. The infrared small and weak target detection method based on Swin-Transformer and multi-scale feature fusion according to claim 1, characterized in that, The target loss function during the training of the object detection model is as follows: In the formula: T represents the target loss function; (i,j) represents an arbitrary coordinate position in the corresponding infrared image; p represents the prediction result finally output by the final network model; y represents the label of the corresponding infrared image; p i,j represents the predicted value output by the object detection model at the position (i, j) in the image, and its magnitude is in the range of (0, 1); y i,j represents the true normalized gray value at the position (i, j) in the image, indicating the label result of the infrared image at the corresponding position.
Citation Information
Patent Citations
Target detection method based on multi-scale feature fusion
CN114118284A
Aerial photography target detection method based on composite backbone network and multiple prediction heads
CN115035429A