An infrared small target detection method and system

By using hyper-fusion backbone network and convolutional attention fusion module for feature fusion and enhancement processing in infrared small object detection, the problem of poor infrared small object detection accuracy is solved, and more efficient infrared small object recognition and detection is achieved.

CN119888204BActive Publication Date: 2025-06-13XIHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510356410.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-13
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing infrared small-object detection methods have poor detection accuracy in complex infrared imaging environments, mainly due to the low distinction between background and target and the weak small target characteristics.

Method used

The hyper-converged backbone network is used to perform multi-level multi-scale feature fusion processing, and local and global feature enhancement processing is performed in combination with the convolutional attention fusion module to improve detection accuracy.

Benefits of technology

Through multi-level feature fusion and feature enhancement processing, the accuracy and recognition ability of infrared small object detection are significantly improved, and global and local information in the image can be captured more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888204B_ABST
    Figure CN119888204B_ABST
Patent Text Reader

Abstract

The present invention discloses an infrared small target detection method and system, belonging to the technical field of target detection. The present invention obtains an original image, performs scaling processing and normalization processing on the original image to obtain a preprocessed image F; inputs the preprocessed image F into a hyperfusion backbone network for multi-level multi-scale feature fusion processing to obtain a hyperfusion feature map; inputs the hyperfusion feature map into a convolutional attention fusion module for local feature enhancement processing and global feature enhancement processing to respectively obtain a local feature map RJ1 and a global feature map RQ1; splices and fuses the local feature map RJ1 and the global feature map RQ1, and then performs decoding processing to obtain a detection result. Overall, the present invention can effectively enhance the system's recognition ability of infrared targets, thereby improving the detection accuracy of infrared small targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to an infrared small target detection method and system. Background Art

[0002] Infrared small target detection is mainly used to identify and locate tiny objects in infrared images, and has broad application prospects and important value in many fields such as security, maritime rescue, and fire monitoring. However, in most cases, the infrared imaging environment is relatively complex, resulting in a low distinguishability between the background and the target, and at the same time, the features of small targets are also relatively weak. Traditional detection methods mainly include filter-based methods and human visual system-based methods, but the detection accuracy of these methods is usually low. The detection method based on convolutional neural network can achieve target detection by extracting the features of infrared images, but due to the problems of insufficient feature fusion and insufficient subsequent feature processing in the U-Net backbone network it adopts, the detection accuracy of the model is also poor. Summary of the Invention

[0003] The main object of the present invention is to provide an infrared small target detection method and system, aiming to solve the technical problem of poor detection accuracy of infrared small targets in related technologies.

[0004] To achieve the above object, the present invention provides an infrared small target detection method, and the method includes the following steps:

[0005] S1. Obtain the original image of the infrared small target, perform scaling processing and normalization processing on the original image of the infrared small target to obtain a preprocessed image F;

[0006] S2. Input the preprocessed image F into a superfusion backbone network, perform multi-level multi-scale feature fusion processing to obtain a superfusion feature map; wherein, the superfusion backbone network includes a plurality of fusion processing modules, the fusion processing module includes at least one non-linear convolution unit, and the non-linear convolution unit includes a first convolution layer, a second convolution layer, and a first non-linear attention layer; the multi-scale feature fusion processing includes splicing fusion processing and non-linear convolution processing, and the steps of the non-linear convolution processing specifically include:

[0007] A1. The feature map input into the non-linear convolution unit is first processed by the first convolution layer and then processed by the second convolution layer to obtain a feature map J1;

[0008] A2. Input the feature map J1 into the first non-linear attention layer for feature weighting processing to obtain a feature map J2;

[0009] A3. After splicing and fusing the feature map J2 with the feature map input into this non-linear convolution unit, output through an activation function to complete this non-linear convolution processing;

[0010] S3. Input the hyper-fused feature map into the convolutional attention fusion module for local feature enhancement and global feature enhancement, respectively obtaining the local feature map RJ1 and the global feature map RQ1;

[0011] S4. Concatenate and fuse the local feature map RJ1 and the global feature map RQ1, and then perform decoding processing to obtain the detection result.

[0012] Optionally, the hyper-fused backbone network includes five fusion processing modules, namely the first fusion processing module, the second fusion processing module, the third fusion processing module, the fourth fusion processing module, and the fifth fusion processing module;

[0013] The first fusion processing module includes five non-linear convolutional units with the same structure, namely the first non-linear convolutional unit (0,0), the second non-linear convolutional unit (1,0), the third non-linear convolutional unit (2,0), the fourth non-linear convolutional unit (3,0), and the fifth non-linear convolutional unit (4,0);

[0014] The second fusion processing module includes four non-linear convolutional units with the same structure, namely the first non-linear convolutional unit (0,1), the second non-linear convolutional unit (1,1), the third non-linear convolutional unit (2,1), and the fourth non-linear convolutional unit (3,1);

[0015] The third fusion processing module includes three non-linear convolutional units with the same structure, namely the first non-linear convolutional unit (0,2), the second non-linear convolutional unit (1,2), and the third non-linear convolutional unit (2,2);

[0016] The fourth fusion processing module includes two non-linear convolutional units with the same structure, namely the first non-linear convolutional unit (0,3) and the second non-linear convolutional unit (1,3);

[0017] The fifth fusion processing module includes one non-linear convolutional unit, which is the first non-linear convolutional unit (0,4);

[0018] S2 specifically includes:

[0019] Input the preprocessed image F into the first non-linear convolutional unit (0,0) for processing to obtain a 16-dimensional feature map P00;

[0020] After performing a two-fold downsampling on the feature map P00, input it into the second non-linear convolutional unit (1,0) for processing to obtain a 32-dimensional feature map P10;

[0021] After performing a two-fold downsampling on the feature map P10, input it into the third non-linear convolutional unit (2,0) for processing to obtain a 64-dimensional feature map P20;

[0022] After performing a two-fold downsampling on the feature map P20, it is then input into the fourth non-linear convolutional unit (3,0) for processing to obtain a 128-dimensional feature map P30;

[0023] After performing a two-fold downsampling on the feature map P30, it is then input into the fifth non-linear convolutional unit (4,0) for processing to obtain a 256-dimensional feature map P40;

[0024] The feature maps P00, P10, P20, P30, and P40 are concatenated and fused, and then input into the first non-linear convolutional unit (0,1) for processing to obtain a 16-dimensional feature map P01;

[0025] After performing a two-fold downsampling on the feature map P01, it is concatenated and fused with the feature maps P10, P20, P30, and P40, and then input into the second non-linear convolutional unit (1,1) for processing to obtain a 32-dimensional feature map P11;

[0026] After performing a two-fold downsampling on the feature map P11, it is concatenated and fused with the feature maps P00, P20, P30, and P40, and then input into the third non-linear convolutional unit (2,1) for processing to obtain a 64-dimensional feature map P21;

[0027] After performing a two-fold downsampling on the feature map P21, it is concatenated and fused with the feature maps P00, P10, P30, and P40, and then input into the fourth non-linear convolutional unit (3,1) for processing to obtain a 128-dimensional feature map P31;

[0028] The feature maps P01, P11, P21, P31, and P00 are concatenated and fused, and then input into the first non-linear convolutional unit (0,2) for processing to obtain a 16-dimensional feature map P02;

[0029] After performing a two-fold downsampling on the feature map P02, it is concatenated and fused with the feature maps P11, P21, P31, and P10, and then input into the second non-linear convolutional unit (1,2) for processing to obtain a 32-dimensional feature map P12;

[0030] After performing a two-fold downsampling on the feature map P12, it is concatenated and fused with the feature maps P01, P21, P31, and P20, and then input into the third non-linear convolutional unit (2,2) for processing to obtain a 64-dimensional feature map P22;

[0031] The feature maps P02, P12, P22, P00, and P01 are concatenated and fused, and then input into the first non-linear convolutional unit (0,3) for processing to obtain a 16-dimensional feature map P03;

[0032] After performing a two-fold downsampling on the feature map P03, it is concatenated and fused with the feature maps P12, P22, P10, and P11, and then input into the second non-linear convolutional unit (1,3) for processing to obtain a 32-dimensional feature map P13;

[0033] The feature maps P03, P13, P00, P01, and P02 are concatenated and fused, and then input into the first non-linear convolutional unit (0,4) for processing to obtain a 16-dimensional feature map P04;

[0034] The feature maps P40, P31, P22, P13, and P04 are concatenated and fused to obtain a 16-dimensional super-fused feature map.

[0035] Optionally, in A2, the steps of feature weighting processing specifically include:

[0036] A21, obtaining the channel attention weight w1 of the feature map J1 through a non-linear attention mechanism, and using the channel attention weight w1 to perform channel weighting processing on the feature map J1 to obtain the feature map J11;

[0037] A22, obtaining the spatial attention weight w2 of the feature map J11 through a spatial attention mechanism, and using the spatial attention weight w2 to perform spatial weighting processing on the feature map J11 to obtain the feature map J2.

[0038] Optionally, A21 specifically includes:

[0039] A21-1, using the ReLU activation function to perform activation processing on the input feature map J1 to enhance its non-linear features and obtain the activated feature map;

[0040] A21-2, using the channel attention mechanism to perform max-pooling processing and average-pooling processing on the activated feature map respectively to obtain the feature map i1 and the feature map i2;

[0041] A21-3, performing fully connected processing on the feature maps i1 and i2 respectively, mapping them to a feature space with a size of 1 / 16 of the input dimension, then using the ReLU activation function to perform activation processing on the fully connected feature map, and performing fully connected processing on the activated feature map again to restore it to the same dimension as the input feature map, obtaining the feature maps i11 and i21 respectively;

[0042] A21-4, fuse the feature map i11 and the feature map i21, and after fusion, use the Sigmoid logic function to output the channel attention weight w1;

[0043] A21-5, perform channel weighting on the feature map J1 using the channel attention weight w1 to obtain the feature map J11.

[0044] Optionally, A22 specifically includes:

[0045] A22-1, perform average pooling and max pooling on the feature channels of the feature map J11 respectively to obtain two single-channel feature maps;

[0046] A22-2, splice and fuse the two single-channel feature maps, and perform a convolution operation after fusion to obtain a single-channel spatial attention weight map;

[0047] A22-3, process the single-channel spatial attention weight map using the Sigmoid logic function to obtain the spatial attention weight w2;

[0048] A22-4, perform spatial weighting on the feature map J11 using the spatial attention weight w2 to obtain the feature map J2.

[0049] Optionally, the convolutional attention fusion module includes a second non-linear attention layer and a cascaded dilated convolution sub-module. The cascaded dilated convolution sub-module includes a first average pooling unit, a third convolutional layer, a cascaded dilated convolution channel unit, and a fourth convolutional layer. The structure of the second non-linear attention layer is the same as that of the first non-linear attention layer. S3 specifically includes:

[0050] Local feature enhancement processing:

[0051] Based on the second non-linear attention layer, perform feature weighting on the superfusion feature map to obtain the local feature map RJ1;

[0052] Global feature enhancement processing:

[0053] Based on the cascaded dilated convolution channel unit, use dilated convolutions with four different dilation rates to process the superfusion feature map respectively to obtain four feature maps with different scale receptive field information;

[0054] After performing average pooling and one convolution on the superfusion feature map in sequence using the first average pooling unit and the third convolutional layer, splice and fuse it with the four feature maps with different scale receptive field information, and then perform one convolution using the fourth convolutional layer to obtain the global feature map RQ1.

[0055] In addition, to achieve the above object, the present invention also provides an infrared small target detection system, which includes:

[0056] An image preprocessing module, configured to obtain an original image of an infrared small target, perform scaling processing and normalization processing on the original image of the infrared small target, and obtain a preprocessed image F;

[0057] A superfusion backbone network, configured to perform multi-level multi-scale feature fusion processing on the preprocessed image F to obtain a superfusion feature map; the superfusion backbone network includes five fusion processing modules, and the fusion processing module includes at least one non-linear convolution unit, configured to perform non-linear convolution processing on the input feature map;

[0058] A convolutional attention fusion module, configured to perform local feature enhancement processing and global feature enhancement processing on the superfusion feature map to respectively obtain a local feature map RJ1 and a global feature map RQ1;

[0059] A detection result output module, configured to splice and fuse the local feature map RJ1 and the global feature map RQ1, and then perform decoding processing to obtain a detection result.

[0060] Optionally, the non-linear convolution unit includes:

[0061] A first convolutional layer, configured to perform a first convolution processing on the feature map input to the non-linear convolution unit;

[0062] A second convolutional layer, configured to perform a second convolution processing on the feature map after the first convolution processing;

[0063] A first non-linear attention layer, configured to perform channel attention weighting processing on the feature map after the second convolution processing first, and then perform spatial attention weighting processing.

[0064] Optionally, the convolutional attention fusion module includes:

[0065] A second non-linear attention layer, configured to perform local feature enhancement processing on the superfusion feature map to obtain a local feature map RJ1;

[0066] A cascaded dilated convolution sub-module, configured to perform global feature enhancement processing on the superfusion feature map to obtain a global feature map RQ1.

[0067] Through the multi-level multi-scale feature fusion processing of the preprocessed image F by the superfusion backbone network of the present invention, the details of feature extraction can be greatly increased, and at the same time, the features of different dimensions are fully fused. On this basis, through the convolutional attention fusion module, global feature enhancement processing and local feature enhancement processing are simultaneously performed on the superfusion feature map, so as to improve the capture effect of the global information and local information in the image features by the system. Therefore, the present invention can effectively enhance the recognition ability of the system for infrared targets as a whole, thereby improving the detection accuracy for infrared small targets. Description of the Drawings

[0068] Figure 1 It is a schematic flowchart of the first embodiment of the infrared small target detection method of the present invention;

[0069] Figure 2 It is a schematic flowchart of the processing flow of the hyper-converged backbone network;

[0070] Figure 3 It is a schematic structural diagram of the non-linear convolution unit;

[0071] Figure 4 It is a schematic structural diagram of the first non-linear attention layer;

[0072] Figure 5 It is a schematic flowchart of the processing flow of the non-linear attention mechanism;

[0073] Figure 6 It is a schematic flowchart of the processing flow of the spatial attention mechanism;

[0074] Figure 7 It is a schematic structural diagram of the convolutional attention fusion module.

[0075] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments

[0076] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0077] The inventive concept of the present application will be further elaborated below in conjunction with some specific embodiments and specific implementation manners.

[0078] The embodiment of the present invention provides an infrared small target detection method, referring to Figure 1 , Figure 1 It is a schematic flowchart of the first embodiment of an infrared small target detection method of the present invention.

[0079] In this embodiment, the infrared small target detection method includes:

[0080] Step S1: Obtain the original image of the infrared small target, perform scaling processing and normalization processing on the original image of the infrared small target, and obtain the preprocessed image F.

[0081] Step S1 specifically includes:

[0082] Step S11: Perform scaling processing on the original image of the infrared small target to convert it into an image with a target resolution.

[0083] Step S12: Perform normalization processing on the R, G, and B channels of the image after scaling processing to obtain the preprocessed image F.

[0084] Specifically, the original image of the infrared small target is scaled and converted into an image with a resolution of 320*320 (target resolution), which can reduce the computational processing burden while retaining sufficient details.

[0085] Normalization processing is performed separately on each RGB channel of the scaled image to avoid the problem of large differences in image brightness in different regions of the infrared image caused by changes in illumination, enabling the system to focus more on the relative structural information of the image rather than the absolute brightness value, thereby improving the robustness of the system under different illumination conditions.

[0086] During the normalization process, for the red channel, a mean of 0.485 and a standard deviation of 0.229 are used; for the green channel, a mean of 0.456 and a standard deviation of 0.224 are used; for the blue channel, a mean of 0.406 and a standard deviation of 0.225 are used.

[0087] During this process, using specific means and standard deviations (such as the means of 0.485, 0.456, 0.406 and the standard deviations of 0.229, 0.224, 0.225) to perform image normalization on the RGB channels can unify the distribution of the color channels, making the numerical ranges of each channel more consistent when input; at the same time, it can reduce the influence of illumination and other external factors, thus better extracting the structural features of the image.

[0088] Step S2: Input the preprocessed image F into the hyperfusion backbone network for multi-level multi-scale feature fusion processing to obtain a hyperfusion feature map.

[0089] Figure 2 It is a schematic diagram of the processing flow of the hyperfusion backbone network. As Figure 2As shown in the figure, the hyper-converged backbone network includes five fusion processing modules, namely the first fusion processing module, the second fusion processing module, the third fusion processing module, the fourth fusion processing module, and the fifth fusion processing module; the first fusion processing module includes five non-linear convolution units with the same structure, namely the first non-linear convolution unit (0,0), the second non-linear convolution unit (1,0), the third non-linear convolution unit (2,0), the fourth non-linear convolution unit (3,0), and the fifth non-linear convolution unit (4,0); the second fusion processing module includes four non-linear convolution units with the same structure, namely the first non-linear convolution unit (0,1), the second non-linear convolution unit (1,1), the third non-linear convolution unit (2,1), and the fourth non-linear convolution unit (3,1); the third fusion processing module includes three non-linear convolution units with the same structure, namely the first non-linear convolution unit (0,2), the second non-linear convolution unit (1,2), and the third non-linear convolution unit (2,2); the fourth fusion processing module includes two non-linear convolution units with the same structure, namely the first non-linear convolution unit (0,3) and the second non-linear convolution unit (1,3); the fifth fusion processing module includes one non-linear convolution unit, which is the first non-linear convolution unit (0,4).

[0090] Step S2 specifically includes:

[0091] The processing process of the first fusion processing module: Input the preprocessed image F into the first non-linear convolution unit (0,0) for processing to obtain a 16-dimensional feature map P00; after performing a two-fold downsampling on the feature map P00, input it into the second non-linear convolution unit (1,0) for processing to obtain a 32-dimensional feature map P10; after performing a two-fold downsampling on the feature map P10, input it into the third non-linear convolution unit (2,0) for processing to obtain a 64-dimensional feature map P20; after performing a two-fold downsampling on the feature map P20, input it into the fourth non-linear convolution unit (3,0) for processing to obtain a 128-dimensional feature map P30; after performing a two-fold downsampling on the feature map P30, input it into the fifth non-linear convolution unit (4,0) for processing to obtain a 256-dimensional feature map P40.

[0092] Processing procedure of the second fusion processing module: Concatenate and fuse the feature maps P00, P10, P20, P30, and P40, and input them into the first non-linear convolution unit (0,1) for processing to obtain a 16-dimensional feature map P01; After performing a two-fold downsampling on the feature map P01, concatenate and fuse it with the feature maps P10, P20, P30, and P40, and then input it into the second non-linear convolution unit (1,1) for processing to obtain a 32-dimensional feature map P11; After performing a two-fold downsampling on the feature map P11, concatenate and fuse it with the feature maps P00, P20, P30, and P40, and then input it into the third non-linear convolution unit (2,1) for processing to obtain a 64-dimensional feature map P21; After performing a two-fold downsampling on the feature map P21, concatenate and fuse it with the feature maps P00, P10, P30, and P40, and then input it into the fourth non-linear convolution unit (3,1) for processing to obtain a 128-dimensional feature map P31.

[0093] Processing procedure of the third fusion processing module: Concatenate and fuse the feature maps P01, P11, P21, P31, and P00, and input them into the first non-linear convolution unit (0,2) for processing to obtain a 16-dimensional feature map P02; After performing a two-fold downsampling on the feature map P02, concatenate and fuse it with the feature maps P11, P21, P31, and P10, and then input it into the second non-linear convolution unit (1,2) for processing to obtain a 32-dimensional feature map P12; After performing a two-fold downsampling on the feature map P12, concatenate and fuse it with the feature maps P01, P21, P31, and P20, and then input it into the third non-linear convolution unit (2,2) for processing to obtain a 64-dimensional feature map P22.

[0094] Processing procedure of the fourth fusion processing module: Concatenate and fuse the feature maps P02, P12, P22, P00, and P01, and input them into the first non-linear convolution unit (0,3) for processing to obtain a 16-dimensional feature map P03; After performing a two-fold downsampling on the feature map P03, concatenate and fuse it with the feature maps P12, P22, P10, and P11, and then input it into the second non-linear convolution unit (1,3) for processing to obtain a 32-dimensional feature map P13.

[0095] Processing procedure of the fifth fusion processing module: Concatenate and fuse the feature maps P03, P13, P00, P01, and P02, and input them into the first non-linear convolution unit (0,4) for processing to obtain a 16-dimensional feature map P04.

[0096] Output process of the hyper-fusion backbone network: The feature maps P40, P31, P22, P13, and P04 are spliced and fused to obtain a 16-dimensional hyper-fusion feature map.

[0097] In this process, through five fusion processing modules, multi-scale feature fusion at multiple levels is performed on the preprocessed image F, which can greatly increase the details of feature extraction while ensuring the full fusion of features in different dimensions, thereby improving the detection performance of the system for infrared small targets.

[0098] Furthermore, the structures of each non-linear convolution unit in the above five fusion processing modules are the same. Here, taking the first non-linear convolution unit (0,0) in the first fusion processing module as an example, the processing process of the non-linear convolution unit will be described.

[0099] Figure 3 It is a schematic diagram of the structure of the non-linear convolution unit. As Figure 3 shown, the non-linear convolution unit includes a first convolution layer, a second convolution layer, and a first non-linear attention layer.

[0100] Among them, the processing process of the non-linear convolution unit specifically includes the following steps:

[0101] Step A1: The feature map input to the non-linear convolution unit is first processed by the first convolution layer and then by the second convolution layer to obtain the feature map J1.

[0102] Step A2: The feature map J1 is input into the first non-linear attention layer for feature weighting processing to obtain the feature map J2.

[0103] Step A3: After splicing and fusing the feature map J2 with the feature map input to this non-linear convolution unit, it is output through the activation function to complete this non-linear convolution processing.

[0104] That is to say, after the preprocessed image F is input into the first non-linear convolution unit (0,0), it is first subjected to convolution processing by the first convolution layer and the second convolution layer in sequence to extract the features of the preprocessed image F. After feature extraction, the first non-linear attention layer is used to perform feature weighting processing on the extracted features.

[0105] Furthermore, Figure 4 It is a schematic diagram of the structure of the first non-linear attention layer. As Figure 4 shown, the feature weighting processing of the first non-linear attention layer specifically includes the following steps:

[0106] Step A21: Obtain the channel attention weight w1 of the feature map J1 through the non-linear attention mechanism, and use the channel attention weight w1 to perform channel weighting processing on the feature map J1 to obtain the feature map J11.

[0107] Step A22: Obtain the spatial attention weight w2 of the feature map J11 through the spatial attention mechanism, and perform spatial weighting on the feature map J11 using the spatial attention weight w2 to obtain the feature map J2.

[0108] Furthermore, Figure 5 is a schematic diagram of the processing flow of the non-linear attention mechanism, as Figure 5 shown, step A21 specifically includes the following steps:

[0109] Step A21-1: Use the ReLU activation function to activate the input feature map J1 to enhance its non-linear features and obtain the activated feature map.

[0110] Step A21-2: Use the channel attention mechanism to perform max-pooling and average-pooling on the activated feature map respectively to obtain the feature map i1 and the feature map i2.

[0111] In this process, when performing max-pooling on the activated feature map and selecting the maximum value in the pooling window, it can focus on the most significant features. When performing average-pooling on the activated features and calculating the average value of the pooling window, it can obtain smoother and overall features.

[0112] Step A21-3: Perform fully connected processing on the feature maps i1 and i2 respectively, map them to a feature space with a size of 1 / 16 of the input dimension to reduce the dimension of the features and enhance the degree of information abstraction at the same time; then use the ReLU activation function to activate the fully connected feature map, and perform fully connected processing on the activated feature map again to restore it to the same dimension as the input feature map, and obtain the feature maps i11 and i21 respectively.

[0113] Step A21-4: Fuse the feature maps i11 and i21, and then use the Sigmoid logic function to output the channel attention weight w1;

[0114] Step A21-5: Perform channel weighting on the feature map J1 using the channel attention weight w1 to obtain the feature map J11.

[0115] In this process, using the non-linear attention mechanism to perform channel weighting on the feature map can strengthen the non-linear features of the feature map, enhance the system's recognition ability for complex features in the feature map, and thus improve the detection performance of the system.

[0116] Furthermore, Figure 6 is a schematic diagram of the processing flow of the spatial attention mechanism, as Figure 6 shown, step A22 specifically includes the following steps:

[0117] Step A22-1: Perform average pooling and max pooling on the feature channels of the feature map J11 respectively to obtain two single-channel feature maps;

[0118] Step A22-2: Concatenate and fuse the two single-channel feature maps, and perform a convolution operation after fusion to obtain a single-channel spatial attention weight map;

[0119] Step A22-3: Process the single-channel spatial attention weight map using the Sigmoid logic function to obtain the spatial attention weight w2;

[0120] Step A22-4: Perform spatial weighting on the feature map J11 using the spatial attention weight w2 to obtain the feature map J2.

[0121] That is to say, in the process taking the first non-linear convolution unit (0,0) in the first fusion processing module as an example, the first non-linear attention layer in the first non-linear convolution unit (0,0) outputs the feature map J2, and then the preprocessed image F input to the first non-linear convolution unit (0,0) and the feature map J2 are concatenated and fused through a residual connection, and then the feature map P00 is output through an activation function, which is the feature map output by the first non-linear convolution unit (0,0).

[0122] In this process, by concatenating and fusing the feature-weighted processed image and the input image through a residual connection, the original information of the input image can be retained as much as possible, avoiding the loss of useful information after passing through multiple attention mechanisms and convolution operations.

[0123] Step S3: Input the super fusion feature map into the convolutional attention fusion module for local feature enhancement processing and global feature enhancement processing to obtain the local feature map RJ1 and the global feature map RQ1 respectively.

[0124] Specifically, Figure 7 is the structural schematic diagram of the convolutional attention fusion module. As Figure 7 shown, the convolutional attention fusion module includes a second non-linear attention layer and a cascaded dilated convolution sub-module. The cascaded dilated convolution sub-module includes a first average pooling unit, a third convolutional layer, a cascaded dilated convolution channel unit, and a fourth convolutional layer. The structure of the second non-linear attention layer is the same as that of the first non-linear attention layer. Step S3 specifically includes:

[0125] Local feature enhancement processing:

[0126] Step S31: Based on the second non-linear attention layer, perform feature weighting on the super fusion feature map to obtain the local feature map RJ1.

[0127] Among them, the second non-linear attention layer performs the same processing steps as the first non-linear attention layer, that is, replacing the input from the feature map J1 with the hyper-fused feature map, and then the local feature map RJ1 can be obtained, which will not be elaborated here.

[0128] In this process, by using the second non-linear attention layer to weight the hyper-fused features through the non-linear attention mechanism and the spatial attention mechanism in sequence, the attention of the system to the key feature channels in the hyper-fused feature map can be improved, and thus the purpose of enhancing the local features can be achieved.

[0129] Global feature enhancement processing:

[0130] Step S32: Based on the cascaded dilated convolutional channel unit, use dilated convolutions with four different dilation rates (the dilation rates are 1, 4, 8, and 12 respectively) to process the hyper-fused feature map, and obtain feature maps with four different-scale receptive field information.

[0131] Step S33: After using the first average pooling unit and the third convolutional layer to perform average pooling processing and one convolutional processing on the hyper-fused feature map in sequence, splice and fuse it with the feature maps with four different-scale receptive field information, and then use the fourth convolutional layer to perform one convolutional processing to obtain the global feature map RQ1.

[0132] In this process, by performing dilated convolution processing with different dilation rates (such as 1, 4, 8, and 12), the receptive field can be expanded, capturing a larger range of context information, and improving the system's ability to capture global features. On this basis, fusing with the feature map after average pooling processing and convolutional processing can achieve a better balance between multi-scale context information and local details.

[0133] Step S4: Splice and fuse the local feature map RJ1 and the global feature map RQ1, and then perform decoding processing to obtain the detection result.

[0134] Specifically, splice and fuse the local feature map RJ1 and the global feature map RQ1 obtained in step S3, and input the fused image into the decoder for decoding processing. Among them, the decoder can specifically adopt a convolutional layer, and the number of channels of the feature map output by this convolutional layer is 1, which is used to represent the presence or location information of the target to obtain the detection result.

[0135] In this embodiment, by performing multi-level multi-scale feature fusion processing on the preprocessed image F through a hyper-fusion backbone network, the details of feature extraction can be greatly increased, while ensuring the full fusion of features in different dimensions. On this basis, through the convolutional attention fusion module, global feature enhancement processing and local feature enhancement processing are simultaneously performed on the hyper-fusion feature map, improving the system's capture effect of global information and local information in the image features. Therefore, the present invention can effectively enhance the system's ability to identify infrared targets as a whole, thereby improving the detection accuracy of infrared small targets.

[0136] Furthermore, to achieve the above object, the present invention also provides an infrared small target detection system, which may include an image preprocessing module, a hyper-fusion backbone network, a convolutional attention fusion module, and a detection result output module.

[0137] Among them, the image preprocessing module is used to obtain the original image of the infrared small target, perform scaling processing and normalization processing on the original image of the infrared small target, and obtain the preprocessed image F; the hyper-fusion backbone network is used to perform multi-level multi-scale feature fusion processing on the preprocessed image F to obtain a hyper-fusion feature map; the hyper-fusion backbone network may include five fusion processing modules, and the fusion processing module includes at least one non-linear convolutional unit for performing non-linear convolutional processing on the input feature map; the convolutional attention fusion module is used to perform local feature enhancement processing and global feature enhancement processing on the hyper-fusion feature map to obtain a local feature map RJ1 and a global feature map RQ1 respectively; the detection result output module is used to splice and fuse the local feature map RJ1 and the global feature map RQ1, and then perform decoding processing to obtain the detection result.

[0138] Furthermore, the non-linear convolutional unit may include a first convolutional layer, a second convolutional layer, and a first non-linear attention layer;

[0139] Among them, the first convolutional layer is used to perform the first convolutional processing on the feature map input to the non-linear convolutional unit; the second convolutional layer is used to perform the second convolutional processing on the feature map after the first convolutional processing; the first non-linear attention layer is used to perform channel attention weighting processing on the feature map after the second convolutional processing first, and then perform spatial attention weighting processing.

[0140] Furthermore, the convolutional attention fusion module may include a second non-linear attention layer and a cascaded dilated convolution sub-module.

[0141] Among them, the second non-linear attention layer is used to perform local feature enhancement processing on the hyper-fusion feature map to obtain a local feature map RJ1; the cascaded dilated convolution sub-module is used to perform global feature enhancement processing on the hyper-fusion feature map to obtain a global feature map RQ1.

[0142] It should be noted that the functions that can be realized by each module in the infrared small target detection system provided in this embodiment and the corresponding technical effects achieved can be referred to the descriptions of the specific implementation manners in the embodiments of the infrared small target detection method of the present invention. For the sake of brevity of the specification, they will not be elaborated here.

[0143] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0144] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for detecting small infrared targets, characterized in that: The method comprises: S1, obtaining an original image of a small infrared target, performing scaling and normalization processing on the original image of the small infrared target, and obtaining a preprocessed image F; S2, inputting the preprocessed image F into the hyper-fusion backbone network, performing multi-level multi-scale feature fusion processing, and obtaining a hyper-fusion feature map; The hyper-converged backbone network includes five fusion processing modules, namely a first fusion processing module, a second fusion processing module, a third fusion processing module, a fourth fusion processing module and a fifth fusion processing module; The first fusion processing module includes five nonlinear convolution units with the same structure, namely a first nonlinear convolution unit (0,0), a second nonlinear convolution unit (1,0), a third nonlinear convolution unit (2,0), a fourth nonlinear convolution unit (3,0) and a fifth nonlinear convolution unit (4,0); The second fusion processing module comprises four nonlinear convolution units with the same structure, namely a first nonlinear convolution unit (0, 1), a second nonlinear convolution unit (1, 1), a third nonlinear convolution unit (2, 1) and a fourth nonlinear convolution unit (3, 1); The third fusion processing module comprises three nonlinear convolution units with the same structure, namely a first nonlinear convolution unit (0, 2), a second nonlinear convolution unit (1, 2) and a third nonlinear convolution unit (2, 2); The fourth fusion processing module comprises two nonlinear convolution units with the same structure, namely a first nonlinear convolution unit (0, 3) and a second nonlinear convolution unit (1, 3); The fifth fusion processing module includes one nonlinear convolution unit, which is a first nonlinear convolution unit (0, 4); The S2 specifically includes: Input the preprocessed image F into the first nonlinear convolution unit (0,0) for processing to obtain a 16-dimensional feature map P00; After downsampling the feature map P00 by a factor of two, the feature map P00 is input into the second nonlinear convolution unit (1,0) for processing to obtain a 32-dimensional feature map P10; After downsampling the feature map P10 by a factor of two, the feature map P10 is input into the third nonlinear convolution unit (2,0) for processing to obtain a 64-dimensional feature map P20; After downsampling the feature map P20 by a factor of two, the feature map P20 is input into the fourth nonlinear convolution unit (3,0) for processing to obtain a 128-dimensional feature map P30; After downsampling the feature map P30 by two times, the feature map P30 is input into the fifth nonlinear convolution unit (4,0) for processing to obtain a 256-dimensional feature map P40; The feature map P00, the feature map P10, the feature map P20, the feature map P30 and the feature map P40 are concatenated and fused, and input into the first nonlinear convolution unit (0, 1) for processing to obtain a 16-dimensional feature map P01; After downsampling the feature map P01 by a factor of two, the feature map P01 is concatenated with the feature map P10, the feature map P20, the feature map P30 and the feature map P40, and then input into the second nonlinear convolution unit (1,1) for processing to obtain a 32-dimensional feature map P11; After downsampling the feature map P11 by a factor of two, the feature map P11 is concatenated with the feature map P00, the feature map P20, the feature map P30 and the feature map P40, and then input into the third nonlinear convolution unit (2,1) for processing to obtain a 64-dimensional feature map P21; After downsampling the feature map P21 by a factor of two, the feature map P21 is concatenated with the feature map P00, the feature map P10, the feature map P30 and the feature map P40, and then input into the fourth nonlinear convolution unit (3,1) for processing to obtain a 128-dimensional feature map P31; The feature map P01, the feature map P11, the feature map P21, the feature map P31 and the feature map P00 are concatenated and fused, and input into the first nonlinear convolution unit (0, 2) for processing to obtain a 16-dimensional feature map P02; After downsampling the feature map P02 by a factor of two, the feature map P02 is concatenated with the feature map P11, the feature map P21, the feature map P31 and the feature map P10, and then input into the second nonlinear convolution unit (1, 2) for processing to obtain a 32-dimensional feature map P12; After downsampling the feature map P12 by a factor of two, the feature map P12 is concatenated with the feature map P01, the feature map P21, the feature map P31 and the feature map P20, and then input into the third nonlinear convolution unit (2,2) for processing to obtain a 64-dimensional feature map P22; The feature map P02, the feature map P12, the feature map P22, the feature map P00 and the feature map P01 are concatenated and fused, and input into the first nonlinear convolution unit (0,3) for processing to obtain a 16-dimensional feature map P03; After downsampling the feature map P03 by a factor of two, the feature map P03 is concatenated and fused with the feature map P12, the feature map P22, the feature map P10 and the feature map P11, and then input into the second nonlinear convolution unit (1, 3) for processing to obtain a 32-dimensional feature map P13; The feature map P03, the feature map P13, the feature map P00, the feature map P01 and the feature map P02 are concatenated and fused, and input into the first nonlinear convolution unit (0, 4) for processing to obtain a 16-dimensional feature map P04; The feature map P40, the feature map P31, the feature map P22, the feature map P13 and the feature map P04 are spliced ​​and fused to obtain the 16-dimensional super-fused feature map; The nonlinear convolution unit includes a first convolution layer, a second convolution layer and a first nonlinear attention layer; the multi-scale feature fusion processing includes splicing fusion processing and nonlinear convolution processing, and the steps of the nonlinear convolution processing specifically include: A1, the feature map input to the nonlinear convolution unit is first processed by the first convolution layer and then by the second convolution layer to obtain the feature map J1; A2, input the feature map J1 into the first nonlinear attention layer for feature weighting processing to obtain the feature map J2; A3, after concatenating and fusing the feature map J2 with the feature map input to the nonlinear convolution unit, outputs it through the activation function to complete the nonlinear convolution processing; S3, inputting the super-fused feature map into a convolutional attention fusion module, performing local feature enhancement processing and global feature enhancement processing, and obtaining a local feature map RJ1 and a global feature map RQ1 respectively; S4, concatenate and fuse the local feature map RJ1 and the global feature map RQ1, and then perform decoding processing to obtain a detection result.

2. The infrared small target detection method according to claim 1, characterized in that: In A2, the feature weighting processing step specifically includes: A21, obtains the channel attention weight w1 of the feature map J1 through the nonlinear attention mechanism, and uses the channel attention weight w1 to perform channel weighted processing on the feature map J1 to obtain the feature map J11; A22, obtains the spatial attention weight w2 of the feature map J11 through the spatial attention mechanism, uses the spatial attention weight w2 to perform spatial weighted processing on the feature map J11, and obtains the feature map J2.

3. The infrared small target detection method according to claim 2, characterized in that: The A21 specifically includes: A21-1, uses the ReLU activation function to activate the input feature map J1, enhances its nonlinear characteristics, and obtains the activated feature map; A21-2, using the channel attention mechanism, performs maximum pooling and average pooling on the activated feature maps to obtain feature maps i1 and i2 respectively; A21-3, fully connect feature map i1 and feature map i2 respectively, map them to a feature space with a size of 1 / 16 of the input dimension, then use the ReLU activation function to activate the fully connected feature maps, and fully connect the activated feature maps again to restore them to the same dimension as the input feature map, and obtain feature maps i11 and i21 respectively; A21-4, fuse the feature map i11 and the feature map i21, and then use the Sigmoid logic function to output the channel attention weight w1; A21-5, use the channel attention weight w1 to perform channel weighted processing on the feature map J1 to obtain the feature map J11.

4. The infrared small target detection method according to claim 2, characterized in that: The A22 specifically includes: A22-1, perform average pooling and maximum pooling processing on the feature channels of feature map J11 respectively to obtain two single-channel feature maps; A22-2, concatenates and fuses the two single-channel feature maps, performs convolution operation after fusion, and obtains a single-channel spatial attention weight map; A22-3, use the Sigmoid logic function to process the single-channel spatial attention weight map to obtain the spatial attention weight w2; A22-4, use the spatial attention weight w2 to perform spatial weighted processing on the feature map J11 to obtain the feature map J2.

5. The infrared small target detection method according to claim 2, characterized in that: The convolutional attention fusion module includes a second nonlinear attention layer and a cascaded hole convolution submodule, the cascaded hole convolution submodule includes a first average pooling unit, a third convolution layer, a cascaded hole convolution channel unit and a fourth convolution layer, the second nonlinear attention layer has the same structure as the first nonlinear attention layer, and S3 specifically includes: Local feature enhancement processing: Based on the second nonlinear attention layer, performing the step of feature weighting processing on the super-fusion feature map to obtain the local feature map RJ1; Global feature enhancement processing: Based on the cascaded dilated convolution channel unit, the super-fused feature map is processed respectively using dilated convolutions with four different dilation rates to obtain feature maps of receptive field information of four different scales; The first average pooling unit and the third convolutional layer are used to perform average pooling processing and one convolution processing on the super-fused feature map in turn, and then the super-fused feature map is spliced ​​and fused with the feature maps of the four receptive field information of different scales, and then the fourth convolutional layer is used to perform one convolution processing to obtain the global feature map RQ1.

6. An infrared small target detection system, characterized in that: The system comprises: An image preprocessing module is used to obtain an original image of a small infrared target, perform scaling and normalization processing on the original image of the small infrared target, and obtain a preprocessed image F; A hyper-fusion backbone network, used for performing multi-level and multi-scale feature fusion processing on the pre-processed image F to obtain a hyper-fusion feature map; the hyper-fusion backbone network includes five fusion processing modules, and the fusion processing module includes at least one nonlinear convolution unit, used for performing nonlinear convolution processing on the input feature map; The five fusion processing modules are respectively a first fusion processing module, a second fusion processing module, a third fusion processing module, a fourth fusion processing module and a fifth fusion processing module; The first fusion processing module includes five nonlinear convolution units with the same structure, namely a first nonlinear convolution unit (0, 0), a second nonlinear convolution unit (1, 0), a third nonlinear convolution unit (2, 0), a fourth nonlinear convolution unit (3, 0) and a fifth nonlinear convolution unit (4, 0); the first nonlinear convolution unit (0, 0) is used to process the preprocessed image F to obtain a 16-dimensional feature map P00; the second nonlinear convolution unit (1, 0) is used to process the feature map P00 after two times downsampling to obtain a 32-dimensional feature map P10; the third nonlinear convolution unit (2, 0) is used to process the feature map P10 after two times downsampling to obtain a 64-dimensional feature map P20; the fourth nonlinear convolution unit (3, 0) is used to process the feature map P20 after two times downsampling to obtain a 128-dimensional feature map P30; the fifth nonlinear convolution unit (4, 0) is used to process the feature map P30 after two times downsampling to obtain a 256-dimensional feature map P40; The second fusion processing module comprises four nonlinear convolution units with the same structure, namely a first nonlinear convolution unit (0, 1), a second nonlinear convolution unit (1, 1), a third nonlinear convolution unit (2, 1) and a fourth nonlinear convolution unit (3, 1); the first nonlinear convolution unit (0, 1) is used to process the result of splicing and fusion of the feature map P00, the feature map P10, the feature map P20, the feature map P30 and the feature map P40 to obtain a 16-dimensional feature map P01; the second nonlinear convolution unit (1, 1) is used to process the result of splicing and fusion of the feature map P00, the feature map P10, the feature map P20, the feature map P30 and the feature map P40 after double downsampling, so as to obtain a 16-dimensional feature map P01; , the result of splicing and fusion of the feature map P30 and the feature map P40 is processed to obtain a 32-dimensional feature map P11; the third nonlinear convolution unit (2,1) is used to process the result of splicing and fusion of the feature map P11, the feature map P00, the feature map P20, the feature map P30 and the feature map P40 after two times downsampling, to obtain a 64-dimensional feature map P21; the fourth nonlinear convolution unit (3,1) processes the result of splicing and fusion of the feature map P21, the feature map P00, the feature map P10, the feature map P30 and the feature map P40 after two times downsampling, to obtain a 128-dimensional feature map P31; The third fusion processing module includes three nonlinear convolution units with the same structure, namely a first nonlinear convolution unit (0, 2), a second nonlinear convolution unit (1, 2) and a third nonlinear convolution unit (2, 2); the first nonlinear convolution unit (0, 2) is used to process the result of splicing and fusion of the feature map P01, the feature map P11, the feature map P21, the feature map P31 and the feature map P00 to obtain a 16-dimensional feature map P02; the second nonlinear convolution unit ( 1,2) is used to process the result of splicing and fusion of the feature map P02, the feature map P11, the feature map P21, the feature map P31 and the feature map P10 after two times downsampling, so as to obtain a 32-dimensional feature map P12; the third nonlinear convolution unit (2,2) is used to process the result of splicing and fusion of the feature map P12, the feature map P01, the feature map P21, the feature map P31 and the feature map P20 after two times downsampling, so as to obtain a 64-dimensional feature map P22; The fourth fusion processing module comprises two nonlinear convolution units with the same structure, namely a first nonlinear convolution unit (0, 3) and a second nonlinear convolution unit (1, 3); the first nonlinear convolution unit (0, 3) is used to process the result of splicing and fusion of the feature map P02, the feature map P12, the feature map P22, the feature map P00 and the feature map P01 to obtain a 16-dimensional feature map P03; the second nonlinear convolution unit (1, 3) is used to process the result of splicing and fusion of the feature map P03, the feature map P12, the feature map P22, the feature map P10 and the feature map P11 after two times downsampling to obtain a 32-dimensional feature map P13; The fifth fusion processing module includes a nonlinear convolution unit, which is a first nonlinear convolution unit (0, 4); the first nonlinear convolution unit (0, 4) is used to process the result of splicing and fusion of the feature map P03, the feature map P13, the feature map P00, the feature map P01 and the feature map P02 to obtain a 16-dimensional feature map P04; A convolutional attention fusion module is used to perform local feature enhancement processing and global feature enhancement processing on the super-fusion feature map to obtain a local feature map RJ1 and a global feature map RQ1 respectively; The detection result output module is used to splice and fuse the local feature map RJ1 and the global feature map RQ1, and then perform decoding processing to obtain the detection result.

7. The infrared small target detection system according to claim 6, characterized in that: The nonlinear convolution unit comprises: A first convolution layer, used for performing a first convolution process on the feature map input into the nonlinear convolution unit; A second convolution layer, used for performing a second convolution process on the feature map after the first convolution process; The first nonlinear attention layer is used to first perform channel attention weighted processing on the feature map after the second convolution processing, and then perform spatial attention weighted processing.

8. The infrared small target detection system according to claim 6, characterized in that: The convolutional attention fusion module includes: The second nonlinear attention layer is used to perform local feature enhancement processing on the super-fusion feature map to obtain a local feature map RJ1; The cascaded hole convolution sub-module is used to perform global feature enhancement processing on the super-fusion feature map to obtain the global feature map RQ1.