Infrared small target detection method and system based on multi-scale fusion
Patent Information
- Application Number
- CN202510503370.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Small object detection of infrared images faces challenges of low resolution, noise interference and multi-scale object detection in complex environments, resulting in poor performance of traditional methods in recall and accuracy.
The infrared small object detection method based on multi-scale fusion is adopted to achieve effective processing and object detection of infrared images through guided filtering denoising, multi-scale image enhancement and multi-scale fusion object detection network.
It significantly improves the detection accuracy of small targets, enhances the robustness and accuracy of the detection network, reduces missed and missed detection, and achieves accurate detection of small targets in various scenarios.
Smart Images

Figure CN120032188A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of small target detection, and in particular to an infrared small target detection method and system based on multi-scale fusion. Background Art
[0002] With the continuous development of intelligent technology, infrared image small target detection technology has been widely used in military, security, traffic monitoring and other fields. Infrared images can obtain thermal radiation information of objects in complex environments such as low light and bad weather. Therefore, they have important application value in target detection. However, since the resolution of infrared images is usually low, the target size is small and the contrast is low, the detection of small targets faces many challenges.
[0003] Traditional small target detection methods rely on manually designed features and simple image processing techniques, such as filtering and histogram equalization. These methods often show great limitations when dealing with complex backgrounds, noise interference, and multi-scale targets. Traditional denoising methods such as Gaussian filtering often lead to the loss of image detail information while removing noise, thus affecting the subsequent target detection effect. Existing image enhancement technology is insufficient in multi-scale target detection and cannot effectively improve the saliency of small targets, resulting in low recall and accuracy of target detection. In the traditional target detection framework, the extraction and fusion of multi-scale features are often relatively simple, and the complementarity between features of different scales cannot be fully utilized, resulting in low detection accuracy. Summary of the invention
[0004] In view of this, the present invention provides an infrared small target detection method based on multi-scale fusion, aiming to achieve more accurate target positioning in multi-scale target detection and provide a more effective solution for the detection of infrared small targets.
[0005] To achieve the above object, the present invention provides an infrared small target detection method based on multi-scale fusion, comprising the following steps: S1: Obtain a certain amount of infrared images, manually annotate each infrared image, and obtain a real detection frame corresponding to each infrared image; construct a training library based on the infrared images and the corresponding real detection frames; denoise the acquired infrared images based on the guided filtering method to obtain denoised infrared images; S2: Use a multi-scale image enhancement method to enhance the denoised infrared image to obtain an enhanced infrared image; S3: Input the enhanced infrared image into the target detection network based on multi-scale fusion to obtain the predicted detection box and confidence; S4: Compare the obtained predicted detection frame with the real detection frame, calculate the overall loss function, and update the parameters of the target detection network based on multi-scale fusion to obtain the trained target detection network based on multi-scale fusion; S5: Use the trained multi-scale fusion-based target detection network to perform target detection on unlabeled infrared images, where the unlabeled infrared images are infrared images without manually labeled real detection frames, and use the non-maximum suppression method to obtain the final target detection frame.
[0006] Optionally, the S1 comprises the following steps: Based on the guided filtering method, the linear relationship within the sliding window is constructed to obtain the denoised infrared image, specifically: The sliding window traverses the infrared image Each pixel position , and extract the pixel position Centered on The pixels contained in the window of are used to construct the linear relationship. is the window size, and the linear relationship is specifically: ; in, The infrared image after denoising At pixel position The pixel value at ; For infrared images At pixel position The pixel value at ; and The pixel positions are The slope and offset at are: ; ; in, for The total number of pixels in the window; The pixel position Centered on The set of pixel positions contained in the window of ; for The pixel locations contained in ; The pixel position Centered on The set of pixel positions contained in the window of ; for The variance of the pixel values corresponding to the pixel positions contained in ; is the normalization factor; for The mean of the pixel values corresponding to the pixel positions contained in ; Optionally, S2 comprises the following steps: S21: Based on the denoised infrared image, perform multi-scale Gaussian blur: The denoised infrared image is Gaussian blurred using three different scales of Gaussian kernels, with variances of 1, 2, and 4, respectively, to obtain the corresponding Gaussian blurred infrared images. , and ; S22: Perform multi-scale image enhancement: Optionally, the S22 includes the following steps: S221: Extracting multi-scale details based on the infrared image after Gaussian blurring at different scales, wherein the multi-scale details are specifically: ; in, , and They are optimal detail map, medium detail map and coarse detail map respectively; S222: Perform multi-scale detail fusion to obtain a fused detail map : ; in, Detailed image after fusion At pixel position The pixel value at ; , and The optimal detail map , Medium Detail and coarse-detail maps At pixel position The pixel value at ; , and Control weights for details and control optimal detail maps , Medium Detail and coarse-detail maps The weights in the fused detail map; is a symbolic function, when When greater than 0, equal to 0, and less than 0, they are 1, 0, and -1 respectively; S223: Based on the fused detail image, the denoised infrared image is enhanced to obtain an enhanced infrared image: ; in, is the enhanced infrared image; is the bilateral filtering function; Optionally, S3 includes the following steps: S31: Input the enhanced infrared image into the target detection network based on multi-scale fusion to obtain the feature map after multi-scale fusion: The target detection network based on multi-scale fusion is an SSD target detection network. The enhanced infrared image is input into the SSD target detection network using the VGG-16 network as the backbone network. The feature maps of three different scales of the VGG-16 network are selected to assign weights and perform multi-scale fusion to obtain the feature map after multi-scale fusion. The specific method is as follows: ; in, It is the feature map after multi-scale fusion; , and They are the feature maps corresponding to conv3_3, conv_4_3 and conv_7_2 of the VGG-16 network respectively; The function will , and Sampling to a fixed length and width is ; , and They are , and The fusion weights are obtained through network training as a component of the target detection network parameters; is the concatenation function for concatenation according to the channel dimension; S32: Target detection based on multi-scale fused feature maps: The multi-scale fused feature map is input into the detector of the SSD target detection network to obtain the predicted detection box and confidence. Optionally, the S4 comprises the following steps: S41: Compare the predicted detection box with the actual detection box and calculate the overall loss function: The loss value of the overall loss function of the SSD target detection network It is composed of confidence loss value and positioning loss value, specifically: ; in, It is the confidence loss value of the SSD target detection network; is the improved positioning loss value; is the positioning loss weight; S42: Update the parameters of the target detection network based on multi-scale fusion: Process each infrared image in the training library through steps S1-S3 to obtain a predicted detection frame, and use the corresponding real detection frame to calculate the overall loss function; Calculate the mean of the overall loss function corresponding to all infrared images in the training library, and use the stochastic gradient descent method based on the mean of the overall loss function to update the parameters of the target detection network based on multi-scale fusion; Repeat step S42 until the set number of iterations is reached to obtain a trained object detection network based on multi-scale fusion.
[0007] Optionally, the S41 includes the following steps: S411: Calculate the improved intersection-over-union ratio: Calculate the intersection-over-union ratio between the predicted detection box and the actual detection box, specifically: ; in, is the real detection frame; is the predicted detection box; Calculate the improved IoU based on the calculated IoU , specifically: ; in, is the predicted detection box The minimum enclosing rectangle of for and The difference of S412: Calculate the improved positioning loss value: .
[0008] Optionally, the S5 comprises the following steps: S51: Use the trained multi-scale fusion-based target detection network to perform target detection on unlabeled infrared images and obtain candidate target detection boxes and confidence levels; S52: Use the non-maximum suppression method to obtain the final target detection frame: Use the non-maximum suppression method to filter candidate target detection frames and obtain the final target detection frame; specifically: S521: Screening candidate target detection frames: According to the confidence levels of the candidate target detection frames obtained in step S51, the candidate target detection frames are sorted from high to low to generate a list of candidate target detection frames; S522: Select the highest confidence box: Select the candidate target detection box with the highest confidence and use it as the initial target detection box; S523: Suppress other overlapping frames: Calculate the improved intersection-over-union ratio between the initial detection frame and other candidate target detection frames in the candidate target detection frame list, and suppress the candidate target detection frames whose improved intersection-over-union ratio is greater than the set threshold, that is, remove them from the candidate target detection list; S524: Repeat selection: Repeat step S522 and step S523 until the list of candidate target detection frames is empty, and all the candidate target detection frames finally retained are the final target detection frames.
[0009] The present invention also discloses an infrared small target detection system based on multi-scale fusion, comprising: Network training library construction module: obtain a certain amount of infrared images, manually annotate each infrared image, and build a training library; Denoising module: Denoise the infrared image in the training library based on the guided filtering method to obtain the denoised infrared image; Enhancement module: Use multi-scale image enhancement method to enhance the denoised infrared image to obtain an enhanced infrared image; Network construction module: Input the enhanced infrared image into the target detection network based on multi-scale fusion to obtain the predicted detection box and confidence; Network training module: Use the predicted detection box to compare the real detection box, calculate the overall loss function, and update the parameters of the target detection network based on multi-scale fusion to obtain the trained target detection network based on multi-scale fusion; Network application module: Use the trained multi-scale fusion-based target detection network to perform target detection on unlabeled infrared images, and use the non-maximum suppression method to obtain the final target detection frame.
[0010] Beneficial effects: The present invention effectively denoises infrared images by adopting a guided filtering method. Compared with traditional denoising methods, the denoising technology based on guided filtering can significantly reduce the noise in the image while retaining the image details, especially in a high-noise background. By processing the infrared image pixel by pixel, the present invention not only improves the denoising effect, but also provides a cleaner image input for subsequent image enhancement and target detection, thereby enhancing the robustness and accuracy of the entire detection process.
[0011] The present invention significantly improves the detection accuracy of small targets by introducing multi-scale image enhancement and multi-scale fusion target detection network. In the image enhancement stage, multi-scale Gaussian blur and detail fusion technology are used to enhance the saliency of small targets in complex backgrounds. In the feature extraction stage, the fusion of multi-scale feature maps is used to effectively capture target information of different scales, thereby improving the sensitivity of the detection network to small targets. Finally, the present invention can achieve accurate detection of small targets in a variety of scenarios, reducing missed detections and false detections.
[0012] In the final stage of target detection, the present invention adopts an improved non-maximum suppression method to effectively reduce the overlap and redundancy between detection frames. By introducing an improved intersection-over-union calculation method, the present invention can more accurately retain the true target detection frame when suppressing other overlapping frames, thereby improving the reliability of the detection results. The method is more adaptable to background noise and complex scenes, ensuring detection accuracy and robustness in diverse environments, thereby making the present invention more widely applicable in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 A schematic diagram of a flow chart of a method for detecting small infrared targets based on multi-scale fusion according to an embodiment of the present invention; Figure 2 is a diagram of small target detection results according to an embodiment of the present invention, wherein Figure 2 (a) is an unlabeled infrared image; Figure 2 (b) is the infrared image for obtaining the final target detection frame. DETAILED DESCRIPTION
[0014] The present invention is further described below in conjunction with the accompanying drawings, but the present invention is not limited in any way. Any changes or substitutions made based on the teachings of the present invention belong to the protection scope of the present invention.
[0015] Embodiment 1: A method for detecting small infrared targets based on multi-scale fusion, such as Figure 1 and Figure 2 As shown, the following steps are included: S1: Obtain a certain amount of infrared images, manually annotate each infrared image, and obtain the real detection frame corresponding to each infrared image; build a training library based on the infrared images and the corresponding real detection frames; denoise the acquired infrared images based on the guided filter method to obtain the denoised infrared images: Based on the guided filtering method, a linear relationship within a sliding window is constructed, and the sliding window traverses the infrared image. Each pixel position , and extract the pixel position As the center, The pixels contained in the window of are used to construct the linear relationship. is the window size, and the linear relationship is specifically: ; in, The infrared image after denoising At pixel position The pixel value at ; For infrared images At pixel position The pixel value at ; and The pixel positions are The slope and offset at are: ; ; in, for The total number of pixels in the window; The pixel position Centered on The set of pixel positions contained in the window of ; for The pixel locations contained in ; The pixel position Centered on The set of pixel positions contained in the window of ; for The variance of the pixel values corresponding to the pixel positions contained in ; is the regularization factor, which is ; for The mean of the pixel values corresponding to the pixel positions contained in ; Guided filtering is an edge-preserving filtering technique that can effectively remove noise while retaining the edge and detail information in the image to the maximum extent. Compared with other denoising methods, it can reduce noise while avoiding image blurring, thereby retaining the clear outline of the target object. This is very important for subsequent image enhancement and target detection, especially in infrared images, where small targets usually rely on weak edge information for identification.
[0016] S2: Use the multi-scale image enhancement method to enhance the denoised infrared image to obtain an enhanced infrared image: S21: Based on the denoised infrared image, perform multi-scale Gaussian blur: Gaussian blur is performed on the denoised infrared image using Gaussian kernels of three different scales, specifically: ; in, , and They are Gaussian kernels with variances of 1, 2, and 4 respectively. Different variances of Gaussian kernels correspond to different scales; , and Use , and Infrared image after Gaussian blur; S22: Perform multi-scale image enhancement: S221: Extracting multi-scale details based on the infrared image after Gaussian blurring at different scales, wherein the multi-scale details are specifically: ; in, , and They are optimal detail map, medium detail map and coarse detail map respectively; S222: Multi-scale detail fusion: The optimal detail image extracted in the fusion step S221 , Medium Detail and coarse-detail maps , get the fused detail map : ; ; in, Detailed image after fusion At pixel position The pixel value at ; , and The optimal detail map , Medium Detail and coarse-detail maps At pixel position The pixel value at ; , and Control weights for details and control optimal detail maps , Medium Detail and coarse-detail maps The weights in the fused detail image are 0.1, 0.8, and 0.2 in this embodiment; S223: Image enhancement: Step S222 is injected into the denoised infrared image, specifically: ; in, is the enhanced infrared image; is the bilateral filtering function; By using Gaussian kernels with different variances for multi-scale blurring, this method can extract image detail information at different scales, including optimal details, medium details, and coarse details. Multi-scale processing helps to preserve the detailed features of various targets in the image, so that the enhanced image can clearly show the structure and texture at various scales. Especially in infrared images, the details of the target object may appear at different scales, and multi-scale detail fusion ensures that this information will not be lost during the enhancement process.
[0017] S3: Input the enhanced infrared image into the target detection network based on multi-scale fusion to obtain the predicted detection box and confidence: S31: Feature extraction and fusion based on multi-scale fusion: The target detection network based on multi-scale fusion is an SSD target detection network. The enhanced infrared image is input into the SSD target detection network using the VGG-16 network as the backbone network. The feature maps of three different scales of the VGG-16 network are selected to assign weights and perform multi-scale fusion. The multi-scale fusion is: ; in, It is the feature map after multi-scale fusion; , and They are the feature maps corresponding to conv3_3, conv_4_3 and conv_7_2 of the VGG-16 network respectively; The function will , and Sampling to a fixed length and width is ; , and They are , and The fusion weights are obtained through network training as a component of the target detection network parameters; is the concatenation function for concatenation according to the channel dimension; S32: Target detection: The multi-scale fused feature map is input into the detector of the SSD target detection network to obtain the predicted detection box and confidence. By using the VGG-16 network as the backbone network of the SSD target detection network, and performing weight assignment and multi-scale fusion on feature maps of different scales, this method can effectively extract and fuse multi-scale information in the image. The different levels of feature maps of the VGG-16 network, conv3_3, conv_4_3, conv_7_2, correspond to different receptive fields and abstract levels, respectively. Fusion of these multi-scale features can significantly improve the accuracy of target detection, especially when processing infrared images with complex backgrounds and diverse target sizes, and can more accurately identify and locate targets.
[0018] S4: Use the predicted detection frame obtained in step S3 to compare with the real detection frame, calculate the overall loss function, and update the parameters of the target detection network based on multi-scale fusion to obtain the trained target detection network based on multi-scale fusion: S41: Calculate the overall loss function: The loss value of the overall loss function of the SSD target detection network It is composed of confidence loss value and positioning loss value, specifically: ; in, It is the confidence loss value of the SSD target detection network; is the improved positioning loss value; is the positioning loss weight, which is 0.9 in this embodiment; S411: Constructing an improved intersection-over-union ratio: Calculate the intersection-and-union ratio between the predicted detection box and the actual detection box. Specifically: ; in, is the real detection frame; is the predicted detection box; based on the calculated intersection-over-union ratio, the improved intersection-over-union ratio is calculated , the improved intersection-over-intersection ratio is specifically: ; in, is the predicted detection box The minimum enclosing rectangle of for and The difference of S412: Calculate the improved positioning loss value: The improved positioning loss value is specifically: ; S42: Update the parameters of the target detection network based on multi-scale fusion: Perform steps S1-S3 on each infrared image in the training library to obtain a predicted detection frame, and use the corresponding real detection frame to calculate the overall loss function, where the real detection frame is obtained by manual pre-annotation; Calculate the mean of the overall loss function corresponding to all infrared images in the training library, and use the stochastic gradient descent method based on the mean of the overall loss function to update the parameters of the target detection network based on multi-scale fusion; Repeat step S42 until the set number of iterations is reached to obtain a trained object detection network based on multi-scale fusion. In this embodiment, the number of iterations is 10,000.
[0019] Based on the traditional SSD target detection network, the present invention introduces an improved positioning loss function, and effectively solves the shortcomings of the traditional intersection-over-union when processing target detection by constructing an improved intersection-over-union. The improved intersection-over-union not only considers the overlapping part of the detection frame and the true frame, but also further optimizes the position accuracy of the detection frame by adding the concept of the circumscribed rectangle. This improvement helps the network to adjust the position of the prediction frame more accurately, thereby improving the detection accuracy.
[0020] S5: Use the trained multi-scale fusion-based target detection network to perform target detection on unlabeled infrared images, such as Figure 2 As shown in (a), the non-maximum suppression method is used to obtain the final target detection frame, as shown in Figure 2 (b) As shown: S51: Application of multi-scale fusion-based object detection network: Use the trained multi-scale fusion-based target detection network to perform target detection on unlabeled infrared images to obtain candidate target detection frames and confidence levels, where the unlabeled infrared images refer to infrared images without manually labeled real detection frames; S52: Use the non-maximum suppression method to obtain the final target detection frame: The candidate target detection box is screened using the non-maximum suppression method, where the non-maximum suppression is: S521: Screening candidate target detection frames: According to the confidence levels of the candidate target detection frames obtained in step S51, the candidate target detection frames are sorted from high to low to generate a list of candidate target detection frames; S522: Select the highest confidence box: Select the candidate target detection box with the highest confidence and use it as the initial target detection box; S523: Suppress other overlapping frames: Calculate the improved intersection-over-union ratio between the initial detection frame and other candidate target detection frames in the candidate target detection frame list, and suppress the candidate target detection frames whose improved intersection-over-union ratio is greater than the set threshold, that is, remove them from the candidate target detection frame list; S524: Repeat selection: Repeat step S522 and step S523 until the list of candidate target detection frames is empty, and all the candidate target detection frames finally retained are the final target detection frames.
[0021] By using the trained multi-scale fusion-based target detection network, this step can accurately identify candidate target detection boxes from unlabeled infrared images. Thanks to the multi-scale fusion strategy and optimized network structure, the network can effectively handle targets of different scales and ensure high positioning accuracy and confidence of candidate detection boxes. This provides high-quality candidate boxes for the subsequent non-maximum suppression method, further improving the accuracy of the final target detection.
[0022] Embodiment 2: The present invention also discloses an infrared small target detection system based on multi-scale fusion, comprising the following five modules: Network training library construction module: obtain a certain amount of infrared images, manually annotate each infrared image, and build a training library; Denoising module: Denoise the infrared image in the training library based on the guided filtering method to obtain the denoised infrared image; Enhancement module: Use multi-scale image enhancement method to enhance the denoised infrared image to obtain an enhanced infrared image; Network construction module: Input the enhanced infrared image into the target detection network based on multi-scale fusion to obtain the predicted detection box and confidence; Network training module: Use the predicted detection box to compare the real detection box, calculate the overall loss function, and update the parameters of the target detection network based on multi-scale fusion to obtain the trained target detection network based on multi-scale fusion; Network application module: Use the trained multi-scale fusion-based target detection network to perform target detection on unlabeled infrared images, and use the non-maximum suppression method to obtain the final target detection frame.
[0023] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. And the terms "including", "comprising" or any other variants thereof in this article are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "including a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0024] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0025] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for detecting small infrared targets based on multi-scale fusion, characterized in that: The following steps are involved: S1: Obtain a certain amount of infrared images, manually annotate each infrared image, and obtain the real detection frame corresponding to each infrared image; build a training library based on the infrared images and the corresponding real detection frames; Based on the guided filtering method, the infrared image of the training library is denoised to obtain the denoised infrared image; S2: Use a multi-scale image enhancement method to enhance the denoised infrared image to obtain an enhanced infrared image; S3: Input the enhanced infrared image into the target detection network based on multi-scale fusion to obtain the predicted detection box and confidence; S4: Use the predicted detection box to compare with the real detection box, calculate the overall loss function, and update the parameters of the multi-scale fusion-based target detection network to obtain the trained multi-scale fusion-based target detection network; S5: Use the trained multi-scale fusion-based target detection network to perform target detection on unlabeled infrared images, where the unlabeled infrared images are infrared images without manually labeled real detection frames, and use the non-maximum suppression method to obtain the final target detection frame.
2. The infrared small target detection method based on multi-scale fusion according to claim 1 is characterized in that: The S1 comprises the following steps: Based on the guided filtering method, the linear relationship within the sliding window is constructed to obtain the denoised infrared image, specifically: The sliding window traverses the infrared image Each pixel position , and extract the pixel position As the center, The pixels contained in the window of are used to construct the linear relationship. is the window size, and the linear relationship is specifically: ; in, The infrared image after denoising At pixel position The pixel value at ; For infrared images At pixel position The pixel value at ; and The pixel positions are The slope and offset at are: ; ; in, for The total number of pixels in the window; The pixel position As the center, The set of pixel positions contained in the window of ; for The pixel locations contained in ; The pixel position As the center, The set of pixel positions contained in the window of ; for The variance of the pixel values corresponding to the pixel positions contained in ; is the normalization factor; for The mean of the pixel values corresponding to the pixel positions contained in .
3. The infrared small target detection method based on multi-scale fusion according to claim 2 is characterized in that: The S2 comprises the following steps: S21: Based on the denoised infrared image, perform multi-scale Gaussian blur: The denoised infrared image is Gaussian blurred using three different scales of Gaussian kernels, with variances of 1, 2, and 4, respectively, to obtain the corresponding Gaussian blurred infrared images. , and ; S22: Perform multi-scale image enhancement.
4. The infrared small target detection method based on multi-scale fusion according to claim 3 is characterized in that: The S22 comprises the following steps: S221: Extracting multi-scale details based on the infrared image after Gaussian blurring at different scales, wherein the multi-scale details are specifically: ; in, , and They are optimal detail map, medium detail map and coarse detail map respectively; S222: Perform multi-scale detail fusion to obtain a fused detail map : ; in, Detailed image after fusion At pixel position The pixel value at ; , and The optimal detail map , Medium Detail and coarse-detail maps At pixel position The pixel value at ; , and Control weights for details and control optimal detail maps , Medium Detail and coarse-detail maps The weights in the fused detail map; is a symbolic function, when When greater than 0, equal to 0, and less than 0, they are 1, 0, and -1 respectively; S223: Based on the fused detail image, the denoised infrared image is enhanced to obtain an enhanced infrared image: ; in, is the enhanced infrared image; is the bilateral filter function.
5. The infrared small target detection method based on multi-scale fusion according to claim 4 is characterized in that: The S3 comprises the following steps: S31: Input the enhanced infrared image into the target detection network based on multi-scale fusion to obtain the feature map after multi-scale fusion: The target detection network based on multi-scale fusion is an SSD target detection network. The enhanced infrared image is input into the SSD target detection network using the VGG-16 network as the backbone network. The feature maps of three different scales of the VGG-16 network are selected to assign weights and perform multi-scale fusion to obtain the feature map after multi-scale fusion. The specific method is as follows: ; in, It is the feature map after multi-scale fusion; , and They are the feature maps corresponding to conv3_3, conv_4_3 and conv_7_2 of the VGG-16 network respectively; The function will , and Sampling to a fixed length and width is ; , and They are , and The fusion weights are obtained through network training as a component of the target detection network parameters; is the concatenation function for concatenation according to the channel dimension; S32: Target detection based on multi-scale fused feature maps: The multi-scale fused feature map is input into the detector of the SSD target detection network to obtain the predicted detection box and confidence.
6. The infrared small target detection method based on multi-scale fusion according to claim 5 is characterized in that: The S4 comprises the following steps: S41: Compare the predicted detection box with the actual detection box and calculate the overall loss function: The loss value of the overall loss function of the SSD target detection network It is composed of confidence loss value and positioning loss value, specifically: ; in, It is the confidence loss value of the SSD target detection network; is the improved positioning loss value; is the positioning loss weight; S42: Update the parameters of the target detection network based on multi-scale fusion: Process each infrared image in the training library through steps S1-S3 to obtain a predicted detection frame, and use the corresponding real detection frame to calculate the overall loss function; Calculate the mean of the overall loss function corresponding to all infrared images in the training library, and use the stochastic gradient descent method based on the mean of the overall loss function to update the parameters of the target detection network based on multi-scale fusion; Repeat step S42 until the set number of iterations is reached to obtain a trained object detection network based on multi-scale fusion.
7. The infrared small target detection method based on multi-scale fusion according to claim 5 is characterized in that: The S41 comprises the following steps: S411: Calculate the improved intersection-over-union ratio: Calculate the intersection-over-union ratio between the predicted detection box and the actual detection box, specifically: ; in, is the real detection frame; is the predicted detection box; Based on the calculated IoU, calculate the improved IoU , specifically: ; in, is the predicted detection box The minimum enclosing rectangle of for and The difference of S412: Calculate the improved positioning loss value: 。 8. The infrared small target detection method based on multi-scale fusion according to claim 5 is characterized in that: The S5 comprises the following steps: S51: Use the trained multi-scale fusion-based target detection network to perform target detection on unlabeled infrared images and obtain candidate target detection boxes and confidence levels; S52: Use the non-maximum suppression method to obtain the final target detection box.
9. The infrared small target detection method based on multi-scale fusion according to claim 5 is characterized in that: The S52 comprises the following steps: Use the non-maximum suppression method to filter candidate target detection frames and obtain the final target detection frame; specifically: S521: Screening candidate target detection frames: According to the confidence levels of the candidate target detection frames obtained in step S51, the candidate target detection frames are sorted from high to low to generate a list of candidate target detection frames; S522: Select the highest confidence box: Select the candidate target detection box with the highest confidence and use it as the initial target detection box; S523: Suppress other overlapping frames: Calculate the improved intersection-over-union ratio between the initial detection frame and other candidate target detection frames in the candidate target detection frame list, and suppress the candidate target detection frames whose improved intersection-over-union ratio is greater than the set threshold, that is, remove them from the candidate target detection frame list; S524: Repeat selection: Repeat step S522 and step S523 until the list of candidate target detection frames is empty, and all the candidate target detection frames finally retained are the final target detection frames.
10. An infrared small target detection system based on multi-scale fusion, characterized in that: include: Network training library construction module: obtain a certain amount of infrared images, manually annotate each infrared image, and build a training library; Denoising module: Denoise the infrared image in the training library based on the guided filtering method to obtain the denoised infrared image; Enhancement module: Use multi-scale image enhancement method to enhance the denoised infrared image to obtain an enhanced infrared image; Network construction module: Input the enhanced infrared image into the target detection network based on multi-scale fusion to obtain the predicted detection box and confidence; Network training module: Use the predicted detection box to compare the real detection box, calculate the overall loss function, and update the parameters of the target detection network based on multi-scale fusion to obtain the trained target detection network based on multi-scale fusion; Network application module: Use the trained multi-scale fusion-based target detection network to detect targets in unlabeled infrared images, and use the non-maximum suppression method to obtain the final target detection frame; To realize an infrared small target detection method based on multi-scale fusion as described in any one of claims 1-9.
Citation Information
Patent Citations
License plate positioning algorithm fusing affine invariant corner feature and visual color feature
CN105139017A
Earthen ruin crack detection method based on HVS and guide wave filter
CN105241886A
Infrared image detail enhancement and denoising method
CN110047055A
Infrared image enhancement method based on 3D filtering
CN111899200A
Infrared image gas leakage and liquid leakage detection method and system based on deep learning
CN114627052A
Cited By
Infrared small target detection supervision enhancement method
CN120580417A
A method for infrared small target detection supervision enhancement
CN120580417B