An Infrared Small Target Detection Method and System Based on Multi-Scale Fusion

Through the infrared small object detection method based on multi-scale fusion, the guided filter denoising, multi-scale image enhancement and multi-scale fusion object detection network are used to solve the problem of insufficient accuracy in complex backgrounds and multi-scale object detection, and higher detection accuracy and reliability are achieved.

CN120032188BActive Publication Date: 2025-07-01CHANGSHA CHAOCHUANG ELECTRONICS TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510503370.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-01
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing infrared small object detection method has low detection accuracy and insufficient recall and accuracy when dealing with complex backgrounds, noise interference and multi-scale targets.

Method used

The infrared small object detection method based on multi-scale fusion is adopted to improve detection accuracy through guided filtering denoising, multi-scale image enhancement and multi-scale fusion object detection network. Specific steps include: denoising the infrared image, image enhancement, training and application of the object detection network, and finally obtaining the final object detection box using the non-maximum suppression method.

Benefits of technology

It significantly improves the detection accuracy of small targets, improves the sensitivity of the detection network to small targets, reduces missed and missed detection, and enhances the reliability and applicability of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032188B_ABST
    Figure CN120032188B_ABST
Patent Text Reader

Abstract

The present invention discloses an infrared small target detection method and system based on multi-scale fusion, including: S1: Obtain a certain amount of infrared images to construct a training library; denoise the infrared images in the training library based on the guided filtering method; S2: Use the multi-scale image enhancement method to enhance the denoised infrared images; S3: Input the enhanced infrared images into the multi-scale fusion-based target detection network to obtain predicted detection boxes and confidence levels; S4: Use the predicted detection boxes to compare with the true detection boxes, calculate the overall loss function, and obtain the trained multi-scale fusion-based target detection network; S5: Use the trained multi-scale fusion-based target detection network to perform target detection on unlabeled infrared images, and use the non-maximum suppression method to obtain the final target detection boxes. The present invention can achieve more accurate target localization in multi-scale target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of small target detection, and in particular to an infrared small target detection method and system based on multi-scale fusion. Background Art

[0002] With the continuous development of intelligent technologies, the small target detection technology of infrared images has been widely applied in fields such as military, security, and traffic monitoring. Infrared images can obtain the thermal radiation information of objects in complex environments such as low light and bad weather. Therefore, they have important application values in target detection. However, due to the usually low resolution of infrared images, small target volume, and low contrast, the detection of small targets faces many challenges.

[0003] Traditional small target detection methods rely on manually designed features and simple image processing techniques such as filtering and histogram equalization. These methods often show great limitations when dealing with complex backgrounds, noise interference, and multi-scale targets. Traditional denoising methods such as Gaussian filtering often cause the loss of image detail information while removing noise, thus affecting the subsequent target detection effect. Existing image enhancement technologies show deficiencies in multi-scale target detection and cannot effectively improve the saliency of small targets, resulting in low recall rate and precision of target detection. In traditional target detection frameworks, the extraction and fusion of multi-scale features are often relatively simple and cannot make full use of the complementarity between different scale features, resulting in low detection accuracy. Summary of the Invention

[0004] In view of this, the present invention provides an infrared small target detection method based on multi-scale fusion, aiming to achieve more accurate target localization in multi-scale target detection and provide a more effective solution for the detection of infrared small targets.

[0005] To achieve the above object, an infrared small target detection method based on multi-scale fusion provided by the present invention includes the following steps:

[0006] S1: Obtain a certain amount of infrared images, perform manual annotation on each infrared image to obtain a true detection box corresponding to each infrared image; based on the infrared images and the corresponding true detection boxes, construct a training library; denoise the obtained infrared images based on the guided filtering method to obtain denoised infrared images;

[0007] S2: Use a multi-scale image enhancement method to enhance the denoised infrared images to obtain enhanced infrared images;

[0008] S3: Input the enhanced infrared images into a multi-scale fusion-based target detection network to obtain predicted detection boxes and confidence levels;

[0009] S4: Compare the obtained predicted detection boxes with the ground truth detection boxes, calculate the overall loss function, and update the parameters of the multi-scale fusion based object detection network to obtain the trained multi-scale fusion based object detection network;

[0010] S5: Use the trained multi-scale fusion based object detection network to perform object detection on unannotated infrared images, where the unannotated infrared images are infrared images without manually annotated ground truth detection boxes, and use non-maximum suppression method to obtain the final object detection boxes.

[0011] Optionally, S1 includes the following steps:

[0012] Construct a linear relationship within the sliding window based on the guided filter method to obtain the denoised infrared image, specifically:

[0013] The sliding window traverses each pixel position in the infrared image and extracts the pixels included in the window centered at the pixel position for the construction of the linear relationship, where is the window size, and the linear relationship is specifically:

[0014] ;

[0015] where is the pixel value of the denoised infrared image at the pixel position ; is the pixel value of the infrared image at the pixel position ; and are the slope and offset at the pixel position respectively, specifically:

[0016] ;

[0017] ;

[0018] where is the total number of pixels within the window; is the set of pixel positions included in the window centered at the pixel position ; is the pixel position included in ; is the pixel position included in is the set of pixel positions included in the window centered at the pixel position ; The set of pixel positions included within the window; is the variance of the pixel values corresponding to the pixel positions included in the regularization factor; is the mean of the pixel values corresponding to the pixel positions included in

[0019] Optionally, the S2 includes the following steps:

[0020] S21: Based on the denoised infrared image, perform multi-scale Gaussian blur:

[0021] Perform Gaussian blur on the denoised infrared image using Gaussian kernels of three different scales, with variances of the Gaussian kernels being 1, 2, and 4 respectively, to obtain the corresponding Gaussian-blurred infrared images , and ;

[0022] S22: Perform multi-scale image enhancement:

[0023] Optionally, the S22 includes the following steps:

[0024] S221: Extract multi-scale details based on the infrared images blurred by Gaussian kernels of different scales. The multi-scale details are specifically:

[0025] ;

[0026] Among them, , and are the optimal detail map, the medium detail map, and the rough detail map respectively;

[0027] S222: Perform multi-scale detail fusion to obtain the fused detail map :

[0028] ;

[0029] Among them, is the pixel value of the fused detail map at the pixel position ; , and are respectively the optimal detail map , the medium detail map and the rough detail map at the pixel position ; , and are detail control weights, respectively controlling the optimal detail map , medium-detail map and rough-detail map weights in the fused detail map; is the sign function, when is 1, 0, and -1 when greater than 0, equal to 0, and less than 0, respectively;

[0030] S223: Based on the fused detail map, enhance the denoised infrared image to obtain the enhanced infrared image:

[0031] ;

[0032] Among them, is the enhanced infrared image; is the bilateral filtering function;

[0033] Optionally, the S3 includes the following steps:

[0034] S31: Input the enhanced infrared image into the multi-scale fusion-based object detection network to obtain the multi-scale fused feature map:

[0035] The multi-scale fusion-based object detection network is the SSD object detection network. Input the enhanced infrared image into the SSD object detection network using the VGG-16 network as the backbone network, select three different-scale feature maps of the VGG-16 network, assign weights and perform multi-scale fusion to obtain the multi-scale fused feature map. The specific method is:

[0036] ;

[0037] Among them, is the multi-scale fused feature map; , and are the feature maps corresponding to conv3_3, conv_4_3, and conv_7_2 of the VGG-16 network, respectively; The function samples , and to a fixed length and width, which is ; , and are the fusion weights of , and respectively, and are obtained through the training of the network as part of the object detection network parameters; is the concatenation function concatenated by channel dimension;

[0038] S32: Perform object detection based on the feature map after multi-scale fusion:

[0039] Input the feature map after multi-scale fusion into the detector of the SSD object detection network to obtain the predicted detection boxes and confidence levels;

[0040] Optionally, the S4 includes the following steps:

[0041] S41: Compare the predicted detection boxes with the true detection boxes and calculate the overall loss function:

[0042] The loss value of the overall loss function of the SSD object detection network Consists of a confidence loss value and a localization loss value, specifically:

[0043] ;

[0044] Among them, Is the confidence loss value of the SSD object detection network; Is the improved localization loss value; Is the localization loss weight;

[0045] S42: Update the parameters of the object detection network based on multi-scale fusion:

[0046] Process each infrared image in the training library through steps S1 - S3 to obtain predicted detection boxes, and calculate the overall loss function using the corresponding true detection boxes;

[0047] Calculate the mean value of the overall loss functions corresponding to all infrared images in the training library, and update the parameters of the object detection network based on multi-scale fusion using the stochastic gradient descent method based on the mean value of the overall loss function;

[0048] Repeat step S42 until the set number of iterations is reached to obtain the trained object detection network based on multi-scale fusion.

[0049] Optionally, the S41 includes the following steps:

[0050] S411: Calculate the improved intersection over union:

[0051] Calculate the intersection over union between the predicted detection boxes and the true detection boxes, specifically:

[0052] ;

[0053] Among them, Is the true detection box; Is the predicted detection box;

[0054] Calculate the improved intersection over union based on the calculated intersection over union , specifically as follows:

[0055] ;

[0056] Among them, is the minimum bounding rectangle of the predicted detection box ; is and the difference set;

[0057] S412: Calculate the improved localization loss value:

[0058] .

[0059] Optionally, the S5 includes the following steps:

[0060] S51: Use the trained multi-scale fusion-based object detection network to perform object detection on the unlabeled infrared image to obtain candidate object detection boxes and confidence levels;

[0061] S52: Use the non-maximum suppression method to obtain the final object detection box:

[0062] Use the non-maximum suppression method to screen the candidate object detection boxes to obtain the final object detection box; specifically as follows:

[0063] S521: Screen the candidate object detection boxes:

[0064] Sort the candidate object detection boxes obtained in step S51 according to their confidence levels from high to low to generate a list of candidate object detection boxes;

[0065] S522: Select the box with the highest confidence level:

[0066] Select the candidate object detection box with the highest confidence level and use it as the initial object detection box;

[0067] S523: Suppress other overlapping boxes:

[0068] Calculate the improved intersection over union between the initial detection box and other candidate object detection boxes in the candidate object detection box list, and suppress the candidate object detection boxes with an improved intersection over union greater than the set threshold, that is, remove them from the candidate object detection list;

[0069] S524: Repeat the selection:

[0070] Repeat step S522 and step S523 until the candidate object detection box list is empty. All the finally retained candidate object detection boxes are the final object detection boxes.

[0071] The present invention also discloses an infrared small target detection system based on multi-scale fusion, including:

[0072] Network training library construction module: Obtain a certain amount of infrared images, manually annotate each infrared image, and construct a training library;

[0073] Denoising module: Denoise the infrared images in the training library based on the guided filtering method to obtain denoised infrared images;

[0074] Enhancement module: Use the multi-scale image enhancement method to enhance the denoised infrared images to obtain enhanced infrared images;

[0075] Network construction module: Input the enhanced infrared images into the multi-scale fusion-based target detection network to obtain predicted detection boxes and confidence levels;

[0076] Network training module: Compare the predicted detection boxes with the real detection boxes, calculate the overall loss function, and update the parameters of the multi-scale fusion-based target detection network to obtain the trained multi-scale fusion-based target detection network;

[0077] Network application module: Use the trained multi-scale fusion-based target detection network to perform target detection on unannotated infrared images, and use the non-maximum suppression method to obtain the final target detection boxes.

[0078] Beneficial effects:

[0079] The present invention effectively denoises infrared images by adopting the guided filtering method. Compared with traditional denoising methods, the denoising technology based on guided filtering can significantly reduce the noise in the image while retaining the image details, especially performing well in the high-noise background. Through the per-pixel processing of infrared images, the present invention not only improves the denoising effect but also provides a cleaner image input for subsequent image enhancement and target detection, thereby enhancing the robustness and accuracy of the entire detection process.

[0080] The present invention significantly improves the detection accuracy of small targets by introducing the multi-scale image enhancement and multi-scale fusion-based target detection network. In the image enhancement stage, the multi-scale Gaussian blur and detail fusion technologies are adopted to enhance the saliency of small targets in complex backgrounds. In the feature extraction stage, the fusion of multi-scale feature maps is used to effectively capture the target information at different scales, thereby improving the sensitivity of the detection network to small targets. Finally, the present invention can accurately detect small targets in various scenarios, reducing the cases of missed detection and false detection.

[0081] In the final stage of object detection, the present invention adopts an improved non-maximum suppression method, effectively reducing the overlap and redundancy between detection frames. By introducing an improved intersection over union (IoU) calculation method, the present invention can more accurately retain the true object detection frames when suppressing other overlapping frames, improving the reliability of the detection results. This method has stronger adaptability to background noise and complex scenes, ensuring the detection accuracy and robustness in diverse environments, and thus enabling the present invention to have a wider applicability in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 FIG. is a schematic flow chart of an infrared small target detection method based on multi-scale fusion according to an embodiment of the present invention;

[0083] Figure 2 FIG. is a diagram of the small target detection result according to an embodiment of the present invention, where Figure 2 (a) is an unannotated infrared image; Figure 2 (b) is an infrared image with the final object detection frames obtained. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0084] The present invention will be further described below with reference to the accompanying drawings, but the present invention is not limited in any way. Any transformation or replacement made based on the teachings of the present invention falls within the protection scope of the present invention.

[0085] Embodiment 1: An infrared small target detection method based on multi-scale fusion, as shown in Figure 1 and Figure 2 shown, includes the following steps:

[0086] S1: Obtain a certain amount of infrared images, perform manual annotation on each infrared image to obtain the true detection frames corresponding to each infrared image; based on the infrared images and the corresponding true detection frames, construct a training library; denoise the obtained infrared images based on the guided filter method to obtain the denoised infrared images:

[0087] Construct a linear relationship within a sliding window based on the guided filter method, and the sliding window traverses each pixel position in the infrared image in each pixel position , and extract the pixels contained within the window centered on the pixel position for the construction of the linear relationship, is the window size, and the linear relationship is specifically:

[0088] ;

[0089] wherein, is the denoised infrared image at the pixel position The pixel value at is the infrared image At the pixel position The pixel value at and are the slope and offset at the pixel positions Specifically:

[0090] ;

[0091] ;

[0092] Among them, is The total number of pixels within the window; is centered at the pixel position and Is the set of pixel positions included within the window; is The pixel position included in is centered at the pixel position and Is the set of pixel positions included within the window; is The variance of the pixel values corresponding to the pixel positions included in is the regularization factor, which is in this embodiment; is The mean of the pixel values corresponding to the pixel positions included in

[0093] Guided filtering is an edge-preserving filtering technique that can effectively remove noise while maximizing the retention of edges and detail information in the image. Compared with other denoising methods, it can reduce noise while avoiding image blurring, thus retaining the clear contour of the target object. This is very important for subsequent image enhancement and target detection, especially in infrared images, where small targets usually rely on weak edge information for identification.

[0094] S2: Use a multi-scale image enhancement method to enhance the denoised infrared image to obtain an enhanced infrared image:

[0095] S21: Based on the denoised infrared image, perform multi-scale Gaussian blur:

[0096] Perform Gaussian blur on the denoised infrared image using three different scales of Gaussian kernels, specifically:

[0097] ;

[0098] Among them, , and They are Gaussian kernels with variances of 1, 2, and 4 respectively. Different variances of Gaussian kernels correspond to different scales; , and Use , and Infrared image after Gaussian blur;

[0099] S22: Perform multi-scale image enhancement:

[0100] S221: Extracting multi-scale details based on the infrared image after Gaussian blurring at different scales, wherein the multi-scale details are specifically:

[0101] ;

[0102] in, , and They are optimal detail map, medium detail map and coarse detail map respectively;

[0103] S222: Multi-scale detail fusion:

[0104] The optimal detail image extracted in the fusion step S221 , Medium Detail and coarse-detail map , get the fused detail map :

[0105] ;

[0106] ;

[0107] in, Detailed image after fusion At pixel position The pixel value at ; , and The optimal detail map , Medium Detail and coarse-detail map At pixel position The pixel value at ; , and Control weights for details and control optimal detail maps , Medium Detail and coarse-detail map The weights in the fused detail image are 0.1, 0.8, and 0.2 in this embodiment;

[0108] S223: Perform image enhancement:

[0109] Inject the steps of S222 into the denoised infrared image, specifically:

[0110] ;

[0111] Among them, is the enhanced infrared image; is the bilateral filtering function;

[0112] By performing multi-scale blurring using Gaussian kernels with different variances, this method can extract detailed information of the image at different scales, including optimal details, medium details, and rough details. Multi-scale processing helps to retain the detailed features of various objects in the image, enabling the enhanced image to clearly show the structures and textures at various scales. Especially in infrared images, the details of target objects may be presented at different scales, and multi-scale detail fusion ensures that this information is not lost during the enhancement process.

[0113] S3: Input the enhanced infrared image into the multi-scale fusion-based object detection network to obtain predicted detection boxes and confidence levels:

[0114] S31: Perform feature extraction and fusion based on multi-scale fusion:

[0115] The multi-scale fusion-based object detection network is the SSD object detection network. Input the enhanced infrared image into the SSD object detection network using the VGG-16 network as the backbone network, select the feature maps of three different scales of the VGG-16 network, assign weights, and perform multi-scale fusion. The multi-scale fusion is as follows:

[0116] ;

[0117] Among them, is the feature map after multi-scale fusion; , and are the feature maps corresponding to conv3_3, conv_4_3, and conv_7_2 of the VGG-16 network respectively; The function samples , and to a fixed length and width, which is ; , and are respectively , and 's fusion weights, and are obtained through network training as part of the object detection network parameters; A splicing function spliced by channel dimension;

[0118] S32: Perform object detection:

[0119] Input the feature map after multi-scale fusion into the detector of the SSD object detection network to obtain predicted detection boxes and confidence levels;

[0120] By using the SSD object detection network with the VGG-16 network as the backbone network and performing weight allocation and multi-scale fusion on feature maps of different scales, this method can effectively extract and fuse multi-scale information in images. Different-level feature maps of the VGG-16 network, conv3_3, conv_4_3, conv_7_2, correspond to different receptive fields and abstraction levels respectively. Fusing these multi-scale features can significantly improve the accuracy of object detection, especially when dealing with infrared images with complex backgrounds and diverse object sizes, and can more accurately identify and locate objects.

[0121] S4: Use the predicted detection boxes obtained in step S3 to compare with the ground truth detection boxes, calculate the overall loss function, and update the parameters of the object detection network based on multi-scale fusion to obtain the trained object detection network based on multi-scale fusion:

[0122] S41: Calculate the overall loss function:

[0123] The loss value of the overall loss function of the SSD object detection network Consists of a confidence loss value and a localization loss value, specifically:

[0124] ;

[0125] Among them, Is the confidence loss value of the SSD object detection network; Is the improved localization loss value; Is the localization loss weight, which is 0.9 in this embodiment;

[0126] S411: Construct an improved intersection over union:

[0127] Calculate the intersection over union between the predicted detection box and the ground truth detection box, and the intersection over union Specifically:

[0128] ;

[0129] Among them, Is the ground truth detection box; Is the predicted detection box; Calculate the improved intersection over union based on the calculated intersection over union , and the improved intersection over union is specifically:

[0130] ;

[0131] wherein, is the minimum bounding rectangle of the predicted detection box ; is the difference set between and

[0132] S412: Calculate the improved localization loss value:

[0133] The specific improved localization loss value is:

[0134] ;

[0135] S42: Update the parameters of the multi-scale fusion based object detection network:

[0136] Process each infrared image in the training library through steps S1 - S3 to obtain the predicted detection boxes, and calculate the overall loss function using the corresponding ground truth detection boxes, where the ground truth detection boxes are obtained by manual pre-annotation;

[0137] Calculate the mean value of the overall loss functions corresponding to all infrared images in the training library, and update the parameters of the multi-scale fusion based object detection network using the stochastic gradient descent method based on the mean value of the overall loss function;

[0138] Repeat step S42 until the set number of iterations is reached to obtain the trained multi-scale fusion based object detection network. In this embodiment, the number of iterations is 10,000 times.

[0139] Based on the traditional SSD object detection network, the present invention introduces an improved localization loss function. By constructing an improved intersection over union, it effectively solves the deficiencies of the traditional intersection over union in object detection. The improved intersection over union not only considers the overlapping part between the detection box and the ground truth box, but also further optimizes the position accuracy of the detection box by introducing the concept of the bounding rectangle. This improvement helps the network to more accurately adjust the position of the predicted box, thereby improving the detection accuracy.

[0140] S5: Use the trained multi-scale fusion based object detection network to perform object detection on unannotated infrared images, as shown in Figure 2 (a), and obtain the final object detection boxes using the non-maximum suppression method, as shown in Figure 2 (b):

[0141] S51: Apply the multi-scale fusion based object detection network:

[0142] Use the trained multi-scale fusion-based object detection network to perform object detection on unannotated infrared images, obtaining candidate object detection boxes and confidence levels. The unannotated infrared images refer to infrared images without manually annotated true detection boxes.

[0143] S52: Use the non-maximum suppression method to obtain the final object detection boxes:

[0144] Use the non-maximum suppression method to screen the candidate object detection boxes. The non-maximum suppression is as follows:

[0145] S521: Screen the candidate object detection boxes:

[0146] Sort the candidate object detection boxes obtained in step S51 according to their confidence levels from high to low to generate a candidate object detection box list.

[0147] S522: Select the box with the highest confidence level:

[0148] Select the candidate object detection box with the highest confidence level and use it as the initial object detection box.

[0149] S523: Suppress other overlapping boxes:

[0150] Calculate the improved intersection over union (IoU) between the initial detection box and other candidate object detection boxes in the candidate object detection box list, and suppress the candidate object detection boxes with an improved IoU greater than the set threshold, that is, remove them from the candidate object detection box list.

[0151] S524: Repeat the selection:

[0152] Repeat step S522 and step S523 until the candidate object detection box list is empty. All the finally retained candidate object detection boxes are the final object detection boxes.

[0153] By using the trained multi-scale fusion-based object detection network, this step can accurately identify candidate object detection boxes from unannotated infrared images. Thanks to the multi-scale fusion strategy and the optimized network structure, this network can effectively process objects of different scales, ensuring relatively high positioning accuracy and confidence levels of the candidate detection boxes. This provides high-quality candidate boxes for the subsequent non-maximum suppression method, further improving the accuracy of the final object detection.

[0154] Embodiment 2: The present invention also discloses an infrared small target detection system based on multi-scale fusion, including the following five modules:

[0155] Network training library construction module: Obtain a certain amount of infrared images, perform manual annotation on each infrared image, and construct a training library.

[0156] Denoising module: Denoise the infrared images in the training library based on the guided filter method to obtain denoised infrared images;

[0157] Enhancement module: Enhance the denoised infrared images using a multi-scale image enhancement method to obtain enhanced infrared images;

[0158] Network construction module: Input the enhanced infrared images into an object detection network based on multi-scale fusion to obtain predicted detection boxes and confidence levels;

[0159] Network training module: Compare the predicted detection boxes with the ground truth detection boxes, calculate the overall loss function, and update the parameters of the object detection network based on multi-scale fusion to obtain a trained object detection network based on multi-scale fusion;

[0160] Network application module: Use the trained object detection network based on multi-scale fusion to perform object detection on unlabeled infrared images and use the non-maximum suppression method to obtain the final object detection boxes.

[0161] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments. And the term "including", "comprising" or any other variant thereof in this article is intended to cover a non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, device, article or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, device, article or method including that element.

[0162] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0163] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for detecting small infrared targets based on multi-scale fusion, characterized in that: The following steps are involved: S1: Obtain a certain amount of infrared images, manually annotate each infrared image, and obtain the real detection frame corresponding to each infrared image; build a training library based on the infrared images and the corresponding real detection frames; Based on the guided filtering method, the infrared image of the training library is denoised to obtain the denoised infrared image ; S2: Use multi-scale image enhancement method to enhance the denoised infrared image Performing enhancement to obtain an enhanced infrared image; S21: Based on denoised infrared image , perform multi-scale Gaussian blur: The denoised infrared image is Gaussian blurred using three different scales of Gaussian kernels, with variances of 1, 2, and 4, respectively, to obtain the corresponding Gaussian blurred infrared images. , and ; S22: perform multi-scale image enhancement; S221: Extracting multi-scale details based on the infrared image after Gaussian blurring at different scales, wherein the multi-scale details are specifically: ; in, , and They are optimal detail map, medium detail map and coarse detail map respectively; S222: Perform multi-scale detail fusion to obtain a fused detail map : ; in, Detailed image after fusion At pixel position The pixel value at ; , and The optimal detail map , Medium Detail and coarse-detail maps At pixel position The pixel value at ; , and Control weights for details and control optimal detail maps , Medium Detail and coarse-detail maps The weights in the fused detail map; is a symbolic function, when When greater than 0, equal to 0, and less than 0, they are 1, 0, and -1 respectively; S223: Based on the fused detail map, the denoised infrared image Perform enhancement to obtain the enhanced infrared image: ; in, is the enhanced infrared image; is the bilateral filtering function; S3: Input the enhanced infrared image into the target detection network based on multi-scale fusion to obtain the predicted detection box and confidence; S4: Use the predicted detection box to compare with the real detection box, calculate the overall loss function, and update the parameters of the multi-scale fusion-based target detection network to obtain the trained multi-scale fusion-based target detection network; S5: Use the trained multi-scale fusion-based target detection network to perform target detection on unlabeled infrared images, where the unlabeled infrared images are infrared images without manually labeled real detection frames, and use the non-maximum suppression method to obtain the final target detection frame.

2. The infrared small target detection method based on multi-scale fusion according to claim 1 is characterized in that: The S1 comprises the following steps: Based on the guided filtering method, the linear relationship within the sliding window is constructed to obtain the denoised infrared image, specifically: The sliding window traverses the infrared image Each pixel position , and extract the pixel position As the center, The pixels contained in the window of are used to construct the linear relationship. is the window size, and the linear relationship is specifically: ; in, The infrared image after denoising At pixel position The pixel value at ; For infrared images At pixel position The pixel value at ; and The pixel positions are The slope and offset at are: ; ; in, for The total number of pixels in the window; The pixel position As the center, The set of pixel positions contained in the window of ; for The pixel locations contained in ; The pixel position As the center, The set of pixel positions contained in the window of ; for The variance of the pixel values ​​corresponding to the pixel positions contained in ; is the normalization factor; for The mean of the pixel values ​​corresponding to the pixel positions contained in .

3. The infrared small target detection method based on multi-scale fusion according to claim 2 is characterized in that: The S3 comprises the following steps: S31: Input the enhanced infrared image into the target detection network based on multi-scale fusion to obtain the feature map after multi-scale fusion: The target detection network based on multi-scale fusion is an SSD target detection network. The enhanced infrared image is input into the SSD target detection network using the VGG-16 network as the backbone network. The feature maps of three different scales of the VGG-16 network are selected to assign weights and perform multi-scale fusion to obtain the feature map after multi-scale fusion. The specific method is as follows: ; in, It is the feature map after multi-scale fusion; , and They are the feature maps corresponding to conv3_3, conv_4_3 and conv_7_2 of the VGG-16 network respectively; The function will , and Sampling to a fixed length and width is ; , and They are , and The fusion weights are obtained through network training as a component of the target detection network parameters; is the concatenation function for concatenation according to the channel dimension; S32: Target detection based on multi-scale fused feature maps: The multi-scale fused feature map is input into the detector of the SSD target detection network to obtain the predicted detection box and confidence.

4. The infrared small target detection method based on multi-scale fusion according to claim 3 is characterized in that: The S4 comprises the following steps: S41: Compare the predicted detection box with the actual detection box and calculate the overall loss function: The loss value of the overall loss function of the SSD target detection network It is composed of confidence loss value and positioning loss value, specifically: ; in, It is the confidence loss value of the SSD target detection network; is the improved positioning loss value; is the positioning loss weight; S42: Update the parameters of the target detection network based on multi-scale fusion: Process each infrared image in the training library through steps S1-S3 to obtain a predicted detection frame, and use the corresponding real detection frame to calculate the overall loss function; Calculate the mean of the overall loss function corresponding to all infrared images in the training library, and use the stochastic gradient descent method based on the mean of the overall loss function to update the parameters of the target detection network based on multi-scale fusion; Repeat step S42 until the set number of iterations is reached to obtain a trained object detection network based on multi-scale fusion.

5. The infrared small target detection method based on multi-scale fusion according to claim 4 is characterized in that: The S41 comprises the following steps: S411: Calculate the improved intersection-over-union ratio: Calculate the intersection-over-union ratio between the predicted detection box and the actual detection box, specifically: ; in, is the real detection frame; is the predicted detection box; Based on the calculated IoU, calculate the improved IoU , specifically: ; in, is the predicted detection box The minimum enclosing rectangle of for and The difference of S412: Calculate the improved positioning loss value: 。 6. The infrared small target detection method based on multi-scale fusion according to claim 5 is characterized in that: The S5 comprises the following steps: S51: Use the trained multi-scale fusion-based target detection network to perform target detection on unlabeled infrared images and obtain candidate target detection boxes and confidence levels; S52: Use the non-maximum suppression method to obtain the final target detection box.

7. The infrared small target detection method based on multi-scale fusion according to claim 6 is characterized in that: The S52 comprises the following steps: Use the non-maximum suppression method to filter candidate target detection frames and obtain the final target detection frame; specifically: S521: Screening candidate target detection frames: According to the confidence levels of the candidate target detection frames obtained in step S51, the candidate target detection frames are sorted from high to low to generate a list of candidate target detection frames; S522: Select the highest confidence box: Select the candidate target detection box with the highest confidence and use it as the initial target detection box; S523: Suppress other overlapping frames: Calculate the improved intersection-over-union ratio between the initial detection frame and other candidate target detection frames in the candidate target detection frame list, and suppress the candidate target detection frames whose improved intersection-over-union ratio is greater than the set threshold, that is, remove them from the candidate target detection frame list; S524: Repeat selection: Repeat step S522 and step S523 until the list of candidate target detection frames is empty, and all the candidate target detection frames finally retained are the final target detection frames.

8. An infrared small target detection system based on multi-scale fusion, characterized in that: include: Network training library construction module: obtain a certain amount of infrared images, manually annotate each infrared image, and build a training library; Denoising module: Denoise the infrared image in the training library based on the guided filtering method to obtain the denoised infrared image; Enhancement module: Use multi-scale image enhancement method to enhance the denoised infrared image to obtain an enhanced infrared image; Network construction module: Input the enhanced infrared image into the target detection network based on multi-scale fusion to obtain the predicted detection box and confidence; Network training module: Use the predicted detection box to compare the real detection box, calculate the overall loss function, and update the parameters of the target detection network based on multi-scale fusion to obtain the trained target detection network based on multi-scale fusion; Network application module: Use the trained multi-scale fusion-based target detection network to detect targets in unlabeled infrared images, and use the non-maximum suppression method to obtain the final target detection frame; To realize an infrared small target detection method based on multi-scale fusion as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Infrared image enhancement method based on 3D filtering

    CN111899200A

  • Infrared image gas leakage and liquid leakage detection method and system based on deep learning

    CN114627052A

  • Infrared small target detection algorithm based on multi-scale differential contrast enhancement

    CN119323671A