Infrared small target detection method fusing long and short semantics and multivariate sensitive loss
Through the infrared small object detection method that combines long-term semantics and multivariate sensitive losses, combined with multi-scale feature extraction and global feature fusion network, the problem of large model parameters and calculations in the existing technology is solved, and efficient infrared small object detection is achieved.
Patent Information
- Application Number
- CN202510269398.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-04
AI Technical Summary
Among the existing infrared small object detection methods, the CNN-based model focuses on local semantics and ignores remote semantic connections, and the Transformer-based detection method focuses on the overall understanding of semantics and ignores local semantics, resulting in large amounts of model parameters and calculations, making it difficult to improve accuracy and reduce false alarm rates.
The infrared small object detection method that combines long-term semantics and multivariate sensitive losses is adopted. Through multi-scale feature extraction network, global feature fusion network and multi-scale wavelet fusion network, combined with lightweight design, long-distance semantic fusion and short-distance semantic fusion are carried out, and multi-scale sensitive loss function is used for training.
It effectively improves the accuracy of infrared small object detection, reduces the false alarm rate, and improves the generalization ability and detection efficiency of the model.
Smart Images

Figure CN120259820A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of infrared small target detection, and specifically relates to an infrared small target detection method that fuses long and short semantics with a multi-sensitive loss. Background Art
[0002] Currently, the methods for infrared small target detection are mainly divided into three types: model-driven, data-driven, and data-combined model hybrid-driven. The advantage of the model-driven method is that it relies on less data, but the performance of the model is generally lower than that of the data-driven model; when the hybrid-driven method faces problems with a more complex non-linear structure, it is difficult to decompose them into several sub-problems, and the implementation difficulty is high; the data-driven model relies on a large amount of data, which overcomes the performance deficiencies of many model-driven methods and also has excellent performance for complex non-linear structure problems. Therefore, the infrared target detection technology based on deep learning has gradually become the mainstream method and achieved good results. The methods of deep learning include CNN, Transformer, GAN, multi-modal fusion, self-supervised, and unsupervised methods. They play important roles in feature extraction, global semantic modeling, data augmentation, cross-modal information fusion, few-shot learning, etc. Among them, the CNN-based detection method mainly focuses on the research of methods aiming at segmenting the foreground and background of images. Starting from U-net, it has been improved by methods such as U-net++, Swim-U-net, and UIU-net. The Transformer-based detection method and its improved methods for images include ViT, DTER, Swim-Transformer, SAM, etc. The CNN-based detection method focuses on local semantics while ignoring long-range semantic connections, while the Transformer-based detection method focuses on the overall understanding of semantics while ignoring local semantics. Although models with a large number of parameters and a large amount of computation such as UIU-net have good accuracy and false alarm rates, the model parameters and the amount of computation are both large, and it is difficult to improve the accuracy of the algorithm due to limitations such as the number of parameters, the amount of computation, and the model algorithm structure of the other several models. Summary of the Invention
[0003] In view of the above deficiencies in the prior art, the present invention provides an infrared small target detection method that fuses long and short semantics with a multi-sensitive loss.
[0004] To achieve the above invention purpose, the technical solution adopted by the present invention is as follows:
[0005] An infrared small target detection method that fuses long and short semantics with a multi-sensitive loss, comprising the following steps:
[0006] Obtain an infrared image;
[0007] Construct an infrared small target detection network model; the infrared small target detection network model includes a multi-scale feature extraction network, a global feature fusion network, and a multi-scale wavelet fusion network;
[0008] Use the multi-scale feature extraction network to extract multi-level feature maps from the infrared image;
[0009] Use the global feature fusion network to perform long-range semantic fusion on the feature maps except the last-level feature map in the multi-level feature maps to obtain multi-level channel feature maps;
[0010] Use the multi-scale wavelet fusion network to decode the multi-level channel feature maps and the last-level feature map in the multi-level feature maps, and weighted sum the decoded feature maps of different sizes to obtain the final predicted image.
[0011] In some embodiments, when using the multi-scale feature extraction network to extract multi-level feature maps from the infrared image,
[0012] Extract multi-level feature maps of different sizes in sequence through the first feature extraction module, the second feature extraction module, the third feature extraction module, the fourth feature extraction module, and the fifth feature extraction module arranged in sequence; each feature extraction module includes two depthwise separable convolutional layers and a residual connection connecting the input and output.
[0013] In some embodiments, when using the global feature fusion network to perform long-range semantic fusion on the feature maps except the last-level feature map in the multi-level feature maps to obtain multi-level channel feature maps,
[0014] Perform embedding operations and layer normalization operations on the first-level feature map, the second-level feature map, the third-level feature map, and the fourth-level feature map in the multi-level feature maps respectively through the parallel first long-range semantic feature fusion channel, the second long-range semantic feature fusion channel, the third long-range semantic feature fusion channel, and the fourth long-range semantic feature fusion channel, then perform a linear transformation through a depthwise separable convolutional layer and flatten the last two dimensions of the feature vector to obtain a first flattened feature map; and splice the output feature maps of the embedding operations of each feature fusion channel in the channel dimension;
[0015] Perform a linear transformation on the spliced feature map through a depthwise separable convolutional layer and flatten the last two dimensions of the feature vector respectively through the parallel fifth long-range semantic feature fusion channel and the sixth long-range semantic feature fusion channel to obtain a second flattened feature map and a third flattened feature map;
[0016] The first flattened feature map is concatenated along the second dimension, then undergoes a discrete two-dimensional Fourier transform and performs a Hadamard product with the feature map obtained after the second flattened feature map undergoes a discrete two-dimensional Fourier transform. After dividing by the scaling factor, it performs the inverse discrete two-dimensional Fourier transform, then performs a Hadamard product with the third flattened feature map and undergoes a reconstruction operation. Finally, it adds the output feature map after the embedding operation with each feature fusion channel and performs a residual connection to obtain the final multi-level channel feature map.
[0017] In some embodiments, the first flattened feature map is concatenated along the second dimension, then undergoes a discrete two-dimensional Fourier transform and performs a Hadamard product with the feature map obtained after the second flattened feature map undergoes a discrete two-dimensional Fourier transform. After dividing by the scaling factor, it performs the inverse discrete two-dimensional Fourier transform, then performs a Hadamard product with the third flattened feature map and undergoes a reconstruction operation, specifically:
[0018]
[0019] where, T 4i is the output feature map of the reconstruction operation, Conv is the convolution operation, Reshape is the reconstruction operation, Softmax is the Softmax activation operation, Q is the first flattened feature map, K is the second flattened feature map, V is the third flattened feature map, d is the scaling factor, chunk is to re-divide the tensor into four different channel numbers along the channel dimension, IFFT is the inverse discrete two-dimensional Fourier transform, FFT is the discrete two-dimensional Fourier transform, and ⊙ is the Hadamard product.
[0020] In some embodiments, when using the multi-scale wavelet fusion network to decode the multi-level channel feature map and the last-level feature map in the multi-level feature map, and weighted summing the decoded feature maps of different sizes to obtain the final predicted image,
[0021] The multi-level channel feature map is subjected to a discrete two-dimensional wavelet transform through parallel first context semantic feature fusion channels, second context semantic feature fusion channels, third context semantic feature fusion channels, and fourth context semantic feature fusion channels to obtain the high and low frequency components of the image. Then, the low frequency component is feature-fused with the feature map output by the previous-level context semantic feature fusion channel, and then feature-fused with the high-level component. Finally, a channel attention feature map is extracted through a multi-scale channel attention feed-forward network module to obtain the output predicted feature map.
[0022] In some embodiments, when extracting the channel attention feature map through the multi-scale channel attention feed-forward network module,
[0023] First, the input feature map is passed through a layer normalization layer and a point convolution layer. Then, the feature map is evenly divided into three parts along the channel dimension. For each part of the feature map, depthwise separable convolution and an activation function layer are applied respectively. Then, the feature maps are concatenated along the channel dimension. Finally, after passing through a point convolution layer, the channel attention weights are calculated through a lightweight channel attention mechanism unit, and then the channel attention weights are multiplied with the input feature map using the Hadamard product to obtain the output channel attention feature map.
[0024] In some embodiments, when calculating the channel attention weights through the lightweight channel attention mechanism unit,
[0025] the input feature map is respectively input into the global average pooling layer and the global max pooling layer, and then two attention vector values are learned through the dynamic convolution module. The results obtained are added to get the channel attention vector, and then after passing through the activation function layer, it is multiplied with the input feature map using the Hadamard product to obtain the channel attention weights.
[0026] In some embodiments, when fusing the features of the low-frequency components with the feature map output from the previous-level context semantic feature fusion channel, an upsampling operation is performed through an efficient upsampling module. Specifically:
[0027] T 95 =Conv(Relu(BN(DWC(TransConv(T out )))))
[0028] where, T 95 is the feature map output by the upsampling operation, Relu is the Relu activation function, BN is batch normalization, DWC is depthwise separable convolution, TransConv is transposed convolution, and T out is the feature map output from the previous-level context semantic feature fusion channel.
[0029] In some embodiments, when training the infrared small target detection network model,
[0030] random flipping, random scaling, padding, and random cropping operations are performed on the original image and the mask image, and then a random copy-paste operation is carried out;
[0031] where the random copy-paste operation is specifically:
[0032] I=I t ·α+M sg
[0033] M=M t +M s
[0034] where, I is the image after the random copy-paste operation, It is the target image, α is a random variable, and M sg is the mask image corresponding to the original image M S is the mask image obtained after transformation, and M sg = G(x0, y0) ⊙ M S , where G(x0, y0) is a Gaussian convolution kernel centered at (x0, y0), and ⊙ is the Hadamard product. M t is the target image, is the corresponding mask image, and M is the mask image corresponding to the image after the random copy-paste operation.
[0035] In some embodiments, when training the infrared small target detection network model, a multi-scale sensitive loss function is adopted, specifically:
[0036] L Total = λ seg L Dice + λ loc L loc + λ dif L dif + λ freq L freq
[0037] Among them, L Total is the multi-scale sensitive loss, and L Dice , L loc , L dif , L freq are the segmentation loss, position sensitivity loss, diffusion sensitivity loss, and frequency domain sensitivity loss respectively. λ seg , λ loc , λ dif , λ freq are the coefficients of the four sensitivity losses respectively.
[0038] The present invention has the following beneficial effects:
[0039] (1) Through the long and short semantic fusion and the multi-scale sensitive loss function, the present invention can effectively improve the accuracy of infrared small target detection and reduce the false alarm rate at the same time.
[0040] (2) By performing data augmentation on the original image and the mask image with various operations (such as random flipping, scaling, padding, cropping, copy-pasting, etc.), the present invention can improve the generalization ability of the model and overcome the problems of small number of images and high similarity in the infrared dataset.
[0041] (3) By combining the multi-scale feature extraction network, the global feature fusion network, and the multi-scale wavelet fusion network, the present invention can not only obtain multi-level feature maps, but also achieve long-distance semantic fusion and short-distance semantic fusion, and fully utilize the local and global semantic information of the image in the decoding part, which helps to improve the detection effect.
[0042] (4) The present invention optimizes the loss function and adopts a multi-sensitive loss function to comprehensively consider the Dice loss, binary cross-entropy loss, and position sensitivity, which can calculate the loss more accurately, guide the model training, and make the adjustment of model parameters more reasonable.
[0043] (5) The present invention adopts a lightweight design in many places, which can avoid excessive parameter and computational amounts of the model due to establishing long semantic connections while ensuring performance, thereby ensuring the infrared small target detection efficiency. Description of the Drawings
[0044] Figure 1 It is a schematic flow diagram of an infrared small target detection method that fuses long and short semantics and multi-sensitive loss.
[0045] Figure 2 It is a schematic diagram of the structure of an infrared small target detection network model.
[0046] Figure 3 It is a schematic diagram of a data augmentation method.
[0047] Figure 4 It is a schematic diagram of the structure of a global feature fusion network.
[0048] Figure 5 It is a schematic diagram of the structure of a multi-scale wavelet fusion module.
[0049] Figure 6 It is a schematic diagram of the structures of a multi-scale channel attention feed-forward network module and a lightweight channel attention unit.
[0050] Figure 7 It is a schematic diagram of the structure of an efficient upsampling module.
[0051] Figure 8 It is a schematic diagram of a multi-scale sensitive loss function. Detailed Embodiments
[0052] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
[0053] As Figure 1 and Figure 2 shown, an infrared small target detection method that fuses long and short semantics and multi-sensitive loss provided by an embodiment of the present invention includes the following steps S1 to S5:
[0054] S1. Obtain an infrared image;
[0055] In an alternative embodiment of the present invention, for the infrared small target detection task, an infrared image is acquired.
[0056] Since there is a certain degree of confidentiality in the infrared dataset, usually the number of images in a single dataset is small and the similarity between different images is high. Therefore, in this embodiment, data augmentation is performed when preparing data for the training model to improve the generalization ability of the model.
[0057] As Figure 3 shown, when training the infrared small target detection network model in this embodiment, random flipping, random scaling, padding, and random cropping operations are performed on the original image and the mask image, and then random copy and paste operations are performed.
[0058] In this embodiment, since the dataset is a grayscale image and the pixel information of the RGB three channels is exactly the same, the channel information of the image is first compressed into a single channel. That is, for the input image tensor (B, C, H, W), generally C = 3, and the tensor slice with C = 1 is intercepted, so that the image can be compressed into a single channel without losing the information of the grayscale image. It should be clear that in this embodiment, the original image and the corresponding ground truth in the dataset are synchronously transformed, and the size of the example image is 256×256.
[0059] When performing the random flipping operation in this embodiment, a random number is generated using a uniform distribution between 0 and 1 to generate the probability of the flipping transformation. Among them, there is a 25% probability that the input image and the corresponding ground truth will be flipped left and right, and a 25% probability that they will be flipped up and down.
[0060] When performing the random scaling operation in this embodiment, first define the basic size base_size of the image, and then generate a random integer that follows a uniform distribution between 0.5 times and 2 times of base_size, named long_size. Next, obtain the length and width of the input image. If the height of the image is greater than the width, then scale the height to long_size and scale the width proportionally. Otherwise, scale the width to long_size and scale the height proportionally.
[0061] When performing the padding operation in this embodiment, padding is performed on the image and the mask. If it is detected that the shorter side of the image is less than the cropping size, it is necessary to determine whether the length and width meet the condition of being greater than the cropping size. Those that do not meet the condition need to be filled with 0, that is, filled with black.
[0062] In this embodiment, a random cropping operation is performed to randomly crop the image and the mask. First, the size (h, w) of the padded image and the cropping size crop_size are obtained. Random numbers are generated in the intervals (0, h - crop_size) and (0, w - crop_size) to generate the starting positions of the cropping, ensuring that the cropped part is within the image range.
[0063] When performing the random copy - paste operation in this embodiment, the mask of the original image and the target image are randomly selected by generating random numbers. Assume the target image is I t , and the corresponding mask is M t , and the mask corresponding to the original image is M s , and the pasted image I and its mask can be obtained by the following formula:
[0064] I = I t ·α + M sg
[0065] M = M t + M s
[0066] where I is the image after the random copy - paste operation, I t is the target image, α is a random variable uniformly distributed in the range (0.9, 1), M sg is the mask image corresponding to the original image M S is the mask image obtained after transformation, M sg = G(x0, y0) ⊙ M S , G(x0, y0) is a Gaussian convolution kernel centered at (x0, y0), (x0, y0) is the center coordinate of the connected region in M S , ⊙ is the Hadamard product, M t is the target image, is the corresponding mask image, and M is the mask image corresponding to the image after the random copy - paste operation. Since the size of the infrared target does not exceed 9×9, it can ensure that the target in the mask is covered and the purpose of smoothing its edges is achieved.
[0067] The above steps are all synchronous operations of the image and the mask, and finally data augmentation is realized.
[0068] S2. Construct an infrared small target detection network model; the infrared small target detection network model includes a multi - scale feature extraction network, a global feature fusion network, and a multi - scale wavelet fusion network;
[0069] S3. Use the multi - scale feature extraction network to extract multi - level feature maps from the infrared image;
[0070] In an alternative embodiment of the present invention, when using the multi - scale feature extraction network to extract multi - level feature maps from the infrared image,
[0071] The first feature extraction module, the second feature extraction module, the third feature extraction module, the fourth feature extraction module, and the fifth feature extraction module set in sequence extract multi-level feature maps of different sizes; each feature extraction module includes two depthwise separable convolutional layers and a residual connection connecting the input and the output.
[0072] In this embodiment, the input image passes through five lightweight-designed feature extraction modules, each of which is implemented by downsampling and residual connection, and includes two 3×3 depthwise separable convolutions (each convolution includes batch normalization and activation), obtaining five four-dimensional tensors with the shape of (B, C, H, W).
[0073] S4. Using the global feature fusion network to perform long-distance semantic fusion on the feature maps except the last-level feature map in the multi-level feature maps, obtaining multi-level channel feature maps;
[0074] In an alternative embodiment of the present invention, as Figure 4 shown, when using the global feature fusion network to perform long-distance semantic fusion on the feature maps except the last-level feature map in the multi-level feature maps to obtain multi-level channel feature maps,
[0075] Performing embedding operation and layer normalization operation on the first-level feature map, the second-level feature map, the third-level feature map, and the fourth-level feature map in the multi-level feature maps respectively through the parallel first long-distance semantic feature fusion channel, the second long-distance semantic feature fusion channel, the third long-distance semantic feature fusion channel, and the fourth long-distance semantic feature fusion channel, then performing linear transformation through a depthwise separable convolutional layer and flattening the last two dimensions of the feature vector to obtain a first flattened feature map; and splicing the output feature maps of the embedding operations of each feature fusion channel in the channel dimension;
[0076] Performing linear transformation on the spliced feature map through a depthwise separable convolutional layer and flattening the last two dimensions of the feature vector respectively through the parallel fifth long-distance semantic feature fusion channel and the sixth long-distance semantic feature fusion channel to obtain a second flattened feature map and a third flattened feature map;
[0077] Splicing the first flattened feature map according to the second dimension, then performing discrete two-dimensional Fourier transform and taking the Hadamard product with the feature map after discrete two-dimensional Fourier transform of the second flattened feature map, then performing inverse discrete two-dimensional Fourier transform after dividing by a scaling factor, then taking the Hadamard product with the third flattened feature map and performing reconstruction operation, and finally adding the output feature maps of the embedding operations of each feature fusion channel to perform residual connection to obtain the final multi-level channel feature maps.
[0078] In this embodiment, the global feature fusion network is used to establish the connection between long-range semantics and perform long-semantic fusion, that is, global feature fusion. In the example, the sizes of the 4 groups of four-dimensional tensors are (B, C, H, W) respectively, Taking one of the three-dimensional tensors (C, H, W) (named T) in the same batch as an example for illustration, the operation process of each three-dimensional tensor is the same. In the first step, an embedding operation is performed on T. Through a convolutional kernel with a size of 16×16 and a stride of 16, a series of three-dimensional tensors with a shape of (C, 16, 16) can be obtained (T 11 , T 12 , …, T 1B ) ∈ T1, where T 1i , i ∈ [1, B] is an integer, representing different three-dimensional tensors in the same batch after T undergoes one operation. The sizes of the feature maps output by the other 3 groups of feature extraction modules are 128×128, 64×64, and 32×32 respectively. The side length and stride of the embedding convolutional kernel are equal, and the convolutional kernel sizes are (8, 8), (4, 4), and (2, 2) respectively to ensure that the obtained T 1i shapes are all (C, 16, 16).
[0079] In the second step, T 1i is normalized using the following formula to accelerate the convergence of the model,
[0080]
[0081] where, T 2i is the normalized tensor, and are the mean and standard deviation of the original tensor respectively. ∈ is a small quantity close to 0 to prevent the denominator from being 0. In the example, ∈ = 1e-6.
[0082] In the third step, a learnable depthwise separable convolution is used for linear transformation and the last two dimensions of T 2i are flattened, that is,
[0083] T 3i = Flatten(DWC(T 2i ))
[0084] where DWC(·) is the depthwise separable convolution and Flatten(·) is the flattening operation, that is, T 2i → (C, 16, 16) becomes T 3i→(C, 256), so the four groups of tensors become (16, 4, 256), (16, 8, 256), (16, 16, 256), (16, 32, 256) respectively. A single sample in one batch of the four groups is named Q1, Q2, Q3, Q4 respectively, that is, Q1→(4, 256), Q2→(8, 256), Q3→(16, 256), Q4→(32, 256).
[0085] At the same time, the four groups of T 1i are concatenated in the channel dimension and passed through two parallel depthwise separable convolutions to obtain two three-dimensional tensors K and V with different sizes of (16, 60, 256).
[0086] Considering that the matrix multiplication calculation complexity of the self-attention mechanism is very high, if the channel and embedding block sizes are relatively large, it will bring a huge computational load to the computer. Here, a lightweight strategy is combined to optimize the problem of large computational amount of the self-attention mechanism. The computational complexity of self-attention mainly comes from the matrix multiplication of K and V. Assuming Y = KV, the elements Y in the j-th column of matrix Y j can be regarded as obtained by convolving the j-th column of matrix V with matrix K. Therefore, in this example, KV can be equivalent to a convolution operation. Combining with the Fourier transform, there is a conclusion that time-domain convolution is equal to frequency-domain product. Therefore, the following method steps of the fourth step are obtained.
[0087] Fourth step, concatenate Q i (i = 1, 2, 3, 4) along the second dimension to get Q→(60, 256), then perform the discrete two-dimensional Fourier transform FFT(·), and at the same time K also performs the discrete two-dimensional Fourier transform FFT(·), and calculate T through the Hadamard product and the reshape function 4i , and the specific calculation formula is as follows:
[0088]
[0089] Among them, d is the scaling factor, which is taken as 8 in this example. IFFT(·) represents the inverse operation of the Fourier transform, ⊙ represents the Hadamard product, Conv(·) represents the convolution operation with a convolution kernel of 1×1 and a stride of 1, and chunk(·) represents dividing the tensor into four different channel numbers that are the same as T 2i along the channel dimension. Reshape(·) represents the reconstruction operation of reshaping the last dimension (H*W) into (H, W).
[0090] Fifth step, add T 1i and T 4i to do a residual connection, and obtain a feature map T 5i with the same size as T through deconvolution upsampling. Finally, T and T 5iAdd and make residual connections to get 4 sets of feature maps T of different sizes 61 →(16,4,256,256),T 62 →(16,8,128,128),T 63 →(16,16,64,64),T 64 →(16,32,32,32). This is the end of the encoding part.
[0091] S5. Decode the multi-level channel feature map and the last-level feature map in the multi-level feature map by using the multi-scale wavelet fusion network, and perform weighted summation of decoded feature maps of different sizes to obtain a final predicted image.
[0092] In an optional embodiment of the present invention, step S5 uses the multi-scale wavelet fusion network to decode the multi-level channel feature map and the last level feature map in the multi-level feature map, and weighted sums the decoded feature maps of different sizes to obtain the final predicted image.
[0093] The multi-level channel feature map is discretely transformed into the high- and low-frequency components of the image through the parallel first context semantic feature fusion channel, the second context semantic feature fusion channel, the third context semantic feature fusion channel and the fourth context semantic feature fusion channel, and then the low-frequency component is feature fused with the feature map output by the previous level context semantic feature fusion channel, and then with the high-frequency component. Finally, the channel attention feature map is extracted through the multi-scale channel attention feedforward network module to obtain the output prediction feature map.
[0094] When step S5 extracts the channel attention feature map through the multi-scale channel attention feedforward network module,
[0095] First, the input feature map passes through a normalization layer and a point convolution layer, and then the feature map is divided into three parts according to the channel dimension. Each feature map passes through a depth-wise separable convolution and an activation function layer respectively, and then the feature maps are spliced according to the channel dimension. Finally, after passing through a point convolution layer, the channel attention weight is calculated through a lightweight channel attention mechanism unit, and then the channel attention weight is Hadamard multiplied with the input feature map to obtain the output channel attention feature map.
[0096] When step S5 calculates the channel attention weight through the lightweight channel attention mechanism unit,
[0097] The input feature map is input into the global average pooling layer and the global maximum pooling layer respectively, and then the values of the two attention vectors are learned through the dynamic convolution module respectively. The results are added together to obtain the channel attention vector, which is then passed through the activation function layer to perform Hadamard product with the input feature map to obtain the channel attention weight.
[0098] In this embodiment, the multi-scale wavelet fusion module is used to decode the multi-level channel feature map and the last-level feature map in the multi-level feature map, and the decoded feature maps of different sizes are weighted and summed to obtain the final predicted image. The decoder obtains the high-frequency and low-frequency components of the image through wavelet transform, and fuses the low-frequency components with the high-level features to take into account the context semantic information. Since small targets usually exist in the high-frequency information of the image, the high-frequency subbands can more easily extract high-quality attention maps through the multi-scale channel attention feedforward network module (MCFB), thereby improving the detection effect of infrared small targets.
[0099] As Figure 5 shown, taking T 65 →(16, 64, 16, 16) and T 64 →(16, 32, 32, 32) as inputs, in the first step, T 64 obtains four tensors of size (16, 32, 16, 16), namely {HH, HL, LH, LL}, through discrete two-dimensional wavelet transform. In the second step, the LL component generated by T 65 and T 64 is subjected to feature fusion through channel dimension concatenation and point convolution, and the output tensor size is (16, 32, 16, 16). At the same time, {HH, HL, LH} also undergoes feature fusion through channel dimension concatenation and point convolution to output a tensor size of (16, 32, 16, 16). The two are then concatenated along the channel dimension and named W out and input into the multi-scale channel attention feedforward network module (MCFB).
[0100] As Figure 6 shown, first, it passes through a layer normalization and a point convolution, and the feature map is evenly divided into three parts along the channel dimension. If the number of channels C cannot be divided evenly, take the decimal part. If the decimal part is greater than 0.5, then divide the difference between C and its integer part by 2. If the decimal part is less than 0.5, then divide the difference between C and the sum of its integer part plus 1 by 2, and then it can be divided into three integers with a difference less than or equal to 1. The three groups of feature maps are respectively passed through depthwise separable convolutions of sizes 3×3, 5×5, and 7×7 and activated through the Relu function, then concatenated along the channel dimension, and finally weighted by a 1×1 convolution to obtain the tensor T 74 →(16, 64, 16, 16). Input T 74 into the lightweight channel attention mechanism unit (LCAB), and perform global average pooling and global max pooling on T 74 respectively to extract the low-frequency information of the feature map and retain the target features. The pooling results are GAP(T 74 )→(16, 64, 1, 1) and GMP(T 74) → (16, 64, 1, 1). Considering that in the operation of the channel attention mechanism, there are not only differences in importance among different channels, but also certain relationships may exist between different channels, one-dimensional convolution is considered to capture the importance differences and relationships among channels. Also, due to the multi-layer structure of the model, the lengths of the attention vectors are different during the operation of the channel attention mechanism. Therefore, the dynamic convolution method is adopted to obtain a better model representation while reducing parameter redundancy. Next, the values of two attention vectors are learned through the dynamic convolution module. According to empirical conclusions, there is generally a relationship between the number of channels C and the convolution kernel size k as follows:
[0101] C = 2 2k-1
[0102] Therefore, the formula of the dynamic convolution module is expressed as follows:
[0103]
[0104] where k is the size of the convolution kernel, i.e., 1×k, and C in is the number of input channels, i.e., the length of the attention vector. Considering that with the change of the number of channels, it is desired to balance the convolution kernel size and the number of channels to obtain a more accurate learning representation, the dynamic convolution method is adopted. At the same time, to ensure that the number of channels remains the same before and after convolution, 0 elements with a length difference not greater than 1 will be padded at both ends of the attention vector before convolution. The specific implementation method is as follows: If k - 1 is even, then Otherwise, pad a length of at one end of the attention vector and
[0105] at the other end. The results are added to calculate the channel attention vector. After activation by Sigmoid, it is multiplied with T 74 using the Hadamard product to obtain T 84 → (16, 64, 16, 16), T 84 The residual connection with W out is the output of the entire multi-scale channel attention feed-forward network module (MCFB) and is named T out .
[0106] Second step, as Figure 7 shown, input T out into the efficient upsampling module (EUSB) and output T 95 :
[0107] T 95 = Conv(Relu(BN(DWC(TransConv(T out )))))
[0108] Among them, Conv(·) is point convolution, Relu(·) is the activation function, BN(·) is batch normalization, DWC(·) is a 3×3 depthwise separable convolution, TransConv(·) is deconvolution, and T 95 has a shape of (16, 32, 32, 32) and is the same as T 64 in shape.
[0109] Thus, one layer of decoding is completed. By analogy, four prediction maps with sizes of (16, 16), (32, 32), (64, 64), and (128, 128) can be obtained. Through deconvolution, the sizes are all expanded to (256, 256) and then concatenated. After being weighted by a learnable point convolution and threshold segmented (usually set to 0.5), the final predicted image is output.
[0110] In an optional embodiment of the present invention, when training the infrared small target detection network model, a multi-scale sensitive loss function is adopted, specifically:
[0111] L Total = λ seg L Dice + λ loc L loc + λ dif L dif + λ ferq L freq
[0112] Among them, L Dice , L loc , L dif , L greq respectively represent the segmentation loss, position sensitivity loss, diffusion sensitivity loss, and frequency domain sensitivity loss. λ seg , λ loc , λ dif , λ freq are the coefficients of the above four sensitivity losses respectively, belonging to hyperparameters, and are all set to 0.25 in this example. The advantage of designing this loss function is that through multi-scale constraints, the model training is comprehensively optimized from segmentation, position, diffusion, and frequency domain consistency, improving the infrared small target detection accuracy, and at the same time having strong anti-noise ability and wide adaptability.
[0113]
[0114] where i represents the position of each pixel in the image, P i represents the predicted value, Y i represents the ground truth, ∈ = 10 -6 to prevent division by zero and the values below ∈ have the same effect.
[0115] The algorithm for position sensitivity loss is as follows: Suppose there are k targets. Calculate the weighted center positions (first moments) of the prediction map and the ground truth map respectively. Define the coordinates of the non-zero elements of the n-th non-zero connected region in the image as (x ni , y ni ). The center of the n-th connected region of the prediction center is:
[0116]
[0117] where P ni is the i-th predicted value of the n-th connected region. The center coordinates of the n-th connected region of the ground truth are:
[0118]
[0119] where Y ni is the i-th ground truth value of the n-th connected region. The position loss uses the Euclidean distance to measure the center deviation:
[0120]
[0121] The diffusion sensitivity loss can be described by the variance of the target region. Calculate the weighted variances in the x and y directions for the n-th non-zero connected regions of the prediction and ground truth images respectively:
[0122]
[0123] The diffusion sensitivity loss is defined as the sum of the absolute values of the differences in variances between the prediction and the ground truth in the x and y directions:
[0124]
[0125] In fact, the above position sensitivity and diffusion sensitivity are both constraints on small targets in the spatial domain, and their weighted sum constitutes a constraint on the distribution consistency of small targets. Since infrared small target detection is easily interfered by high-frequency point noise, better results can be achieved by further constraining with frequency domain loss.
[0126] The frequency domain sensitivity loss mainly uses the Fourier transform to transform the prediction map and the ground truth map into the frequency domain and compare their frequency distributions. Let the two-dimensional discrete Fourier transform be Then there is
[0127]
[0128] Extract the amplitude spectrum:
[0129] A P = |F(P)|, A Y = |F(Y)|.
[0130] The frequency domain loss calculates the difference in amplitude spectra using the L2 distance:
[0131]
[0132] Finally, the above four sensitive scale losses are weighted and summed by the first formula of the loss function, and its operation structure diagram is as shown in Figure 8 shown, where ⊙ represents the Hadamard product of the image and the corresponding initial ground truth, and the obtained image Y is the ground truth participating in the operation of the loss function.
[0133] Through the above steps, the final loss can be calculated and the parameters can be modified by backpropagation in the model, and the model training is completed by iteration.
[0134] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a machine for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0135] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0136] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0137] In the present invention, specific embodiments are used to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only for helping to understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
[0138] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention according to these technical revelations disclosed by the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. An infrared small target detection method that fuses long and short semantics with a multi-sensitivity loss, characterized in that Including the following steps: Obtain an infrared image; Construct an infrared small target detection network model; the infrared small target detection network model includes a multi-scale feature extraction network, a global feature fusion network, and a multi-scale wavelet fusion network; Use the multi-scale feature extraction network to extract multi-level feature maps from the infrared image; Use the global feature fusion network to perform long-distance semantic fusion on the feature maps except the last-level feature map in the multi-level feature maps to obtain multi-level channel feature maps; Use the multi-scale wavelet fusion network to decode the multi-level channel feature maps and the last-level feature map in the multi-level feature maps, and perform weighted summation on the decoded feature maps of different sizes to obtain the final predicted image.
2. The infrared small target detection method integrating long and short semantics and multi - sensitive loss according to claim 1, wherein When using the multi-scale feature extraction network to extract multi-level feature maps from the infrared image, Extract multi-level feature maps of different sizes in sequence through the first feature extraction module, the second feature extraction module, the third feature extraction module, the fourth feature extraction module, and the fifth feature extraction module set in sequence; each feature extraction module includes two depthwise separable convolutional layers and a residual connection connecting the input and output.
3. The infrared small target detection method integrating long and short semantics and multi - sensitive loss according to claim 1, characterized in that, When using the global feature fusion network to perform long-distance semantic fusion on the feature maps except the last-level feature map in the multi-level feature maps to obtain multi-level channel feature maps, Perform embedding operations and layer normalization operations on the first-level feature map, the second-level feature map, the third-level feature map, and the fourth-level feature map in the multi-level feature maps respectively through the parallel first long-distance semantic feature fusion channel, the second long-distance semantic feature fusion channel, the third long-distance semantic feature fusion channel, and the fourth long-distance semantic feature fusion channel, then perform linear transformation through a depthwise separable convolutional layer and flatten the last two dimensions of the feature vector to obtain a first flattened feature map; and splice the output feature maps of the embedding operations of each feature fusion channel in the channel dimension; Perform linear transformation on the spliced feature map through a depthwise separable convolutional layer and flatten the last two dimensions of the feature vector respectively through the parallel fifth long-distance semantic feature fusion channel and the sixth long-distance semantic feature fusion channel to obtain a second flattened feature map and a third flattened feature map; Splice the first flattened feature map along the second dimension, then perform a discrete two-dimensional Fourier transform and take the Hadamard product with the feature map after the discrete two-dimensional Fourier transform of the second flattened feature map, then perform an inverse discrete two-dimensional Fourier transform after dividing by a scaling factor, then take the Hadamard product with the third flattened feature map and perform a reconstruction operation, and finally add the output feature maps of the embedding operations of each feature fusion channel and perform a residual connection to obtain the final multi-level channel feature maps.
4. An infrared small target detection method integrating long and short semantics and multi - sensitive loss according to claim 3, characterized in that, Splice the first flattened feature map along the second dimension, then perform a discrete two-dimensional Fourier transform and take the Hadamard product with the feature map after the discrete two-dimensional Fourier transform of the second flattened feature map, then perform an inverse discrete two-dimensional Fourier transform after dividing by a scaling factor, then take the Hadamard product with the third flattened feature map and perform a reconstruction operation, specifically: Among them, T 4i is the reconstructed operation output feature map, Conv is the convolution operation, Reshape is the reconstruction operation, Softmax is the Softmax activation operation, Q is the first flattened feature map, K is the second flattened feature map, V is the third flattened feature map, d is the scaling factor, chunk is to re-divide the tensor into four different channel numbers according to the channel dimension, IFFT is the inverse operation of the discrete two-dimensional Fourier transform, FFT is the discrete two-dimensional Fourier transform, and ⊙ is the Hadamard product.
5. An infrared small target detection method integrating long and short semantics and multi - sensitive loss according to claim 1, characterized in that, When using the multi-scale wavelet fusion network to decode the multi-level channel feature map and the last-level feature map in the multi-level feature map, and weighted summing the decoded feature maps of different sizes to obtain the final predicted image, the multi-level channel feature map is subjected to discrete two-dimensional wavelet transform through parallel first context semantic feature fusion channels, second context semantic feature fusion channels, third context semantic feature fusion channels, and fourth context semantic feature fusion channels to obtain the high-frequency and low-frequency components of the image. Then, the low-frequency component is feature-fused with the feature map output by the previous-level context semantic feature fusion channel, and then feature-fused with the high-level component. Finally, a channel attention feature map is extracted through a multi-scale channel attention feed-forward network module to obtain the output predicted feature map.
6. The infrared small target detection method integrating long and short semantics and multi - sensitive loss according to claim 5, wherein, When extracting the channel attention feature map through the multi-scale channel attention feed-forward network module, first, the input feature map passes through a layer normalization layer and a point convolution layer. Then, the feature map is evenly divided into three parts according to the channel dimension. Each part of the feature map passes through a depthwise separable convolution and an activation function layer respectively, and then the feature maps are concatenated according to the channel dimension. Finally, after passing through a point convolution layer, the channel attention weight is calculated through a lightweight channel attention mechanism unit, and then the channel attention weight is multiplied by the input feature map through Hadamard product to obtain the output channel attention feature map.
7. The infrared small target detection method integrating long and short semantics and multi - sensitive loss according to claim 6, characterized in that, When calculating the channel attention weight through the lightweight channel attention mechanism unit, the input feature map is respectively input into the global average pooling layer and the global maximum pooling layer, and then two attention vector values are learned through the dynamic convolution module respectively. The obtained results are added to obtain the channel attention vector, and then passed through the activation function layer and multiplied by the input feature map through Hadamard product to obtain the channel attention weight.
8. A method for infrared small target detection that fuses long and short semantics and multiple sensitive losses, characterized in that, When feature-fusing the low-frequency component with the feature map output by the previous-level context semantic feature fusion channel, an upsampling operation is performed through an efficient upsampling module. Specifically: T 95 = Conv(Relu(BN(DWC(TransConv(T out )))) Among them, T 95 is the feature map output by the upsampling operation, Relu is the Relu activation function, BN is batch normalization, DWC is depthwise separable convolution, TransConv is transposed convolution, and T out is the feature map output by the feature fusion channel of the previous-level context semantic features.
9. An infrared small target detection method integrating long and short semantics and multi - sensitive loss according to claim 1, characterized in that, When training the infrared small target detection network model, random flipping, random scaling, padding, and random cropping operations are performed on the original image and the mask image, and then a random copy-paste operation is performed; where the random copy-paste operation is specifically: I = I t ·α + M sg M = M g + M s Among them, I is the image after random copy-paste operation, I t is the target image, α is a random variable, M sg is the mask image corresponding to the original image M s is the mask image obtained after transformation, M sg = G(x0, y0) ⊙ M S , G(x0, y0) is a Gaussian convolution kernel centered at (x0, y0), ⊙ is the Hadamard product, M t is the target image, is the corresponding mask image, and M is the mask image corresponding to the image after random copy-paste operation.
10. An infrared small target detection method that fuses long and short semantics and multiple sensitive losses according to claim 1, characterized in that, When training the infrared small target detection network model, a multi-scale sensitive loss function is adopted. Specifically: L Total = λ seg L Dice + λ loc L loc + λ dif L dif + λ freq L freq Among them, L Total is the multi-scale sensitive loss, and L Dice , L loc , L dif , L freq are the segmentation loss, position sensitivity loss, diffusion sensitivity loss, and frequency domain sensitivity loss respectively. λ seg , λ loc , λ dif , λ freq are the coefficients of the four sensitivity losses respectively.
Citation Information
Cited By
Power scene defect small target detection method based on Gaussian mask supervision and cross-layer attention guidance
CN120707569A
A method for detecting small defects in power scenarios based on Gaussian mask supervision and cross-layer attention guidance
CN120707569B