An infrared moving target detection system based on improved YOLOv8

CN118608995BActive Publication Date: 2026-09-29BEIJING INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410669458.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2026-09-29
Estimated Expiration
2044-05-28

AI Technical Summary

Technical Problem

然而,YOLO模型检测的准确性受到输入图像质量的影响,这会限制其识别小物体的能力

Benefits of technology

[0039]1、本发明提供一种基于改进YOLOv8的红外运动目标检测系统,提出了一种双输入目标检测模型,其输入可以分为单输入和多输入,其中多输入同时包括运动信息和原始帧信息;具体的,双输入处理模块利用检测置信度控制了光流处理模块的开启与否,并将来自主输入与副输入的特征有效的融合在了一起,双输入处理模块代替了YOLOv8模型的输入,这样能够高效的将主副输入融合在了一起,避免使用了两个YOLO网络处理特征,融合后使用一个YOLO网络即可;再者,本发明利用检测置信度控制了副输入的开启与否,检测置信度高的时候,仅单输入就可以很好的完成检测任务,这样就不需要副输入的介入,可以最大程度地节省计算能力以及保证检测的实时性;当检测的置信度低时,副输入辅助主输入进行目标检测,这样可以保证检测性能,使得本发明能够稳定、快速的处理复杂背景下快速运动的红外小物体的检测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118608995B_ABST
    Figure CN118608995B_ABST
Patent Text Reader

Abstract

The application provides an infrared moving target detection system based on improved YOLOv8, a double-input processing module controls whether a light flow processing module is turned on or not by using detection confidence, and effectively fuses features from main input and auxiliary input together, the double-input processing module replaces the input of the YOLOv8 model, so that the main and auxiliary inputs can be efficiently fused together, avoiding using two YOLO networks to process features, and after fusion, one YOLO network can be used; furthermore, the application controls whether the auxiliary input is turned on or not by using detection confidence, when the detection confidence is high, single input can well complete the detection task, so the intervention of the auxiliary input is not needed, and the calculation capacity can be saved to the maximum extent and the real-time detection is ensured; when the detection confidence is low, the auxiliary input assists the main input to perform target detection, so that the detection performance can be ensured, and the application can stably and quickly process the detection of infrared small objects moving fast in a complex background.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of machine vision and target detection technology, and particularly relates to an infrared moving target detection system based on an improved YOLOv8. Background Technology

[0002] In the field of aircraft target detection, the detection of small moving objects in infrared images faces significant challenges. These objects typically occupy only a small number of pixels in infrared images, resulting in limited feature information, significant feature loss, low recognition accuracy, and various challenges in single-frame detection. Currently, infrared small target detection methods can be mainly divided into two categories: single-frame detection methods and multi-frame detection methods. Single-frame detection methods are applicable to various situations and are currently the mainstream method type, using a single image for target detection. Single-frame detection methods can generally be further divided into three categories. The first category is morphology-based methods, and the second is filter-based methods. These two types of methods have relatively high computational efficiency, but it is difficult to find suitable templates or filters for different scenes. The third type of method is outlier-based methods. This method treats small targets as outliers in the saliency map, generates a saliency map of the infrared image using different techniques, and then filters out small targets as outliers from the saliency map using a certain threshold or rule. These methods still have a high false alarm rate even under point-like background features and noise interference.

[0003] Therefore, multi-frame detection methods have been extensively studied. Multi-frame detection methods use multiple frames of images for detection, achieving target detection by extracting the spatiotemporal motion information of the target from consecutive frames. They are primarily suitable for scenes with moving targets. Multi-frame detection algorithms heavily rely on image feature extraction. When the image feature extraction results fail to meet expectations, or when the target is stationary, the true target may be incorrectly identified as background and thus missed or misjudged, reducing the stability of the system.

[0004] Choosing the appropriate object detection algorithm for a specific scenario is crucial for the proper processing of input data. In recent years, Convolutional Neural Networks (CNNs) have achieved many breakthroughs in computer vision, offering significant advantages in detection accuracy and robustness compared to traditional methods. CNN-based detection algorithms boast superior computational performance compared to traditional algorithms, but their high computational cost makes real-time deployment on drone platforms difficult. The You Only Look Once (YOLO) model, a CNN-based object detection model, has become a fundamental real-time object detection model in robotics, autonomous vehicles, and video surveillance. Currently, the YOLOv8 model is the dominant object detection algorithm among YOLO models due to its speed, simplicity, and multi-scale detection capabilities. However, the accuracy of YOLO detection is affected by the quality of the input image, limiting its ability to identify small objects. Furthermore, YOLOv8 is a single-frame detection method, which can easily lead to detection failures when dealing with targets moving rapidly relative to complex backgrounds.

[0005] Due to the characteristics of infrared objects—small pixels, high speed, and complex backgrounds—the YOLOv8 model may experience false positives or false negatives in infrared object detection. First, a single input cannot carry motion information, which may lead to a decrease in the detection performance of large moving objects. Second, in complex backgrounds, the model may not be able to adequately focus on information about small objects. Third, in complex backgrounds, the interference from similar small objects is high, increasing the possibility of false positives. Furthermore, small, dark objects travel long distances, easily leading to missed detections. Finally, the features of small moving infrared objects are greatly affected by scene factors such as lighting and temperature, which may cause significant feature variations within the same scene, resulting in varying detection difficulties. In recent years, researchers have proposed various improved YOLOv8 models to address the challenges of infrared moving target detection. These improvements can be summarized in several aspects: First, innovations in the target detection model structure, including the use of multiple network stacks and the addition or removal of detection heads; second, the replacement and modification of modules such as attention mechanisms and convolutional kernels; and finally, innovations in loss functions, training methods, and hyperparameter settings. In addition, research has been gradually carried out on aspects such as fast model training, lightweight design, and input types. However, existing methods do not consider using more reasonable multi-input combinations to ensure that the object detection model maintains good detection performance while ensuring real-time efficiency. Moreover, the effective fusion of multi-frame motion information and original frame information, as well as the robustness of the model after feature extraction failure, have not been adequately considered.

[0006] Meanwhile, considering the limitations of single-frame detection methods, researchers utilize consecutive frames to extract effective spatiotemporal information to improve the accuracy and efficiency of object recognition. For example, Guimin J et al. used inertial navigation information to perform image registration between consecutive frames, which suppressed background pixels and highlighted moving targets; Yi X et al. used an adaptive local contrast module and a temporal correlation module after the visual saliency module to detect small infrared targets; Zhang W et al. combined target motion features extracted from multiple frames using Kalman filtering with saliency maps; Zhao F et al. used optical flow and the spatiotemporal consistency of trajectory points to detect small infrared targets; and Kwan C et al. used motion features extracted from multiple frames using optical flow to enhance the detection performance of small moving targets. All of the above studies utilize background suppression or optical flow methods to process multi-frame images. Although the input to the object detection network includes motion information, object information from the original frames may be lost. For complex backgrounds, the above detection models may incorrectly identify backgrounds with significant motion as objects. Li D et al. used image registration to extract multi-frame motion information and then used it along with the original images as input data for infrared object detection. However, when the original frame is fed into the object detection network along with the background-suppressed image, using additional input data may lead to computational waste when the original frame alone can achieve satisfactory detection results. Existing methods do not consider using a more reasonable dual-input combination to ensure that the object detection model maintains excellent detection performance while ensuring real-time efficiency. Furthermore, the effective fusion of multi-frame motion information and original frame information, as well as the model's robustness after the failure of background suppression or optical flow methods, have not been adequately considered. Summary of the Invention

[0007] To address the aforementioned issues, this invention provides an infrared moving target detection system based on an improved YOLOv8. A dual-input processing module determines whether to activate the optical flow processing module for the next frame based on the detection confidence values ​​output from the previous and current frames. According to this determination, the dual-input processing module processes the current frame and the optical flow-processed image separately, thereby thoroughly integrating and effectively utilizing information acquired from multiple inputs. This significantly improves the detection performance of the target detection model, enabling efficient and accurate detection of small moving infrared objects.

[0008] An infrared moving target detection system based on an improved YOLOv8 includes an optical flow processing module, a dual-input processing module, a target detection module, and an output module;

[0009] The optical flow processing module is used to process the received previous frame infrared image and current frame infrared image using an improved Horn-Schunck method when the detection confidence of the current iteration output by the output module is less than a set threshold, so as to obtain an optical flow processed image.

[0010] The dual-input processing module is used to extract features from the received image to obtain a fused feature map. When the detection confidence output by the output module is less than a set threshold, the image received by the dual-input processing module is the optical flow processed image and the current frame infrared image. When the detection confidence output by the output module for the current iteration is not less than the set threshold, the image received by the dual-input processing module is only the current frame infrared image.

[0011] The target detection module is used to detect infrared moving targets and the detection confidence of infrared moving targets in the current frame infrared image based on the aggregated feature map;

[0012] The output module is used to take the average of the detection confidence scores of the current frame infrared image and the previous frame infrared image as the detection confidence score for the next iteration.

[0013] Furthermore, the optical flow processing module uses an improved Horn-Schunck method to process the received previous frame infrared image and the current frame infrared image, resulting in the optical flow processed image as follows:

[0014] The same convolution operation is used to perform convolution operations on the grayscale images of the previous frame infrared image and the current frame infrared image respectively. The convolution results of the two frames infrared images are then fused to obtain the first convolution grayscale image.

[0015] Two opposite convolution operations are performed on the grayscale images of the previous frame infrared image and the current frame infrared image, respectively. The convolution results of the two frames infrared images are then fused to obtain the second convolution grayscale image.

[0016] Give the initial velocity component u of the optical flow in the u direction (0) and the initial velocity component of the optical flow in the v direction v (0) Set initial values;

[0017] The initial velocity component of the optical flow u (0) and the initial velocity component of optical flow v (0) After iterating the following optical flow equation a set number of times using the initial value as the initial value, the optical flow velocity component u is obtained. * and optical flow velocity component v * :

[0018]

[0019]

[0020] Where k represents the number of iterations, I x I represents the partial derivative of the image grayscale pixel (x,y) on the x-axis in the first convolution grayscale image. yI represents the partial derivative of the image grayscale pixel (x, y) with respect to y in the first convolutional grayscale image. t Let λ represent the partial derivative of the image grayscale pixel (x,y) with respect to t on the second convolutional grayscale image, λ represent the optical flow constant controlling the smoothness of the optical flow field, and ρ represent the momentum constant. This represents the local average optical flow velocity in the neighborhood of the current pixel in the u direction, obtained from the k-th iteration. u represents the local average optical flow velocity in the neighborhood of the current pixel in the v direction, obtained from the k-th iteration. (k-1) u k u (k+1) Let v represent the optical flow velocity components in the u direction obtained in the (k-1), k, and k+1 iterations, respectively. (k-1) v (k) v (k+1) Let ω1(x,y) represent the optical flow velocity components in the v direction obtained in the (k-1), k, and k+1 iterations, respectively, and let ω2(x,y) represent the first and second weights, respectively. Furthermore, we have:

[0021]

[0022] Where α1 is the weight for suppressing the magnitude of the optical flow vector, α2 is the ratio of the optical flow equation constraint to the compliance constraint, and T is the set threshold.

[0023] The optical flow velocity components u at each pixel * and optical flow velocity component v * The optical flow field is formed, and after the required visualization processing, an optical flow processed image is obtained.

[0024] Furthermore, letting α1 = α2, the first weight ω1(x,y) and the second weight ω2(x,y) are simplified to the following piecewise function ω(x,y):

[0025]

[0026] Where k' is the slope related to the values ​​of α1 and α2.

[0027] Furthermore, the dual-input processing module includes a first Focus submodule, a second Focus submodule, a Maxpool layer, an Avgpool layer, a first Concat submodule, a second Concat submodule, a first CBS convolution submodule, and a second CBS convolution submodule;

[0028] When the detection confidence level output by the output module is less than a set threshold, the first Focus submodule operates on the current frame infrared image to obtain a first Focus feature map; the second Focus submodule operates on the optical flow processed image to obtain a second Focus feature map; the second Focus feature map is pooled by a Maxpool layer and an Avgpool layer respectively to obtain a first pooled feature map and a second pooled feature map; the second Focus feature map, the first pooled feature map, and the second pooled feature map are superimposed by the second Concat submodule to obtain an optical flow feature map; the optical flow feature map is fused by the second CBS convolution submodule to obtain an optical flow fused feature map; the first Focus feature map and the optical flow fused feature map are superimposed by the first Concat submodule to obtain an intermediate fused feature map; the intermediate fused feature map is fused by the first CBS convolution submodule to obtain the final fused feature map;

[0029] When the detection confidence of the current iteration output by the output module is not less than the set threshold, the current frame infrared image is directly used as the final fused feature map.

[0030] Furthermore, the target detection module is an improved YOLOv8 network model, which includes a Backbone sub-network, a Neck sub-network, and a Head sub-network. Specifically, the C2f module and the Upsample module in the Neck sub-network are connected via the CBAM_α module, and the C2f module and the CBS module are connected via the CBAM_α module. Simultaneously, the SPPF module in the Backbone sub-network is connected to the first Upsample module in the Neck sub-network via the CBAM_α module.

[0031] The CBAM_α module includes a channel attention module, a spatial attention module, a first CBS module, and a second CBS module; the aggregated feature map F is processed by the channel attention module and the spatial attention module respectively to obtain the corresponding channel attention feature map M. c (F) and spatial attention feature map M s (F); Channel attention feature map M c (F) and spatial attention feature map M s (F) is multiplied element-wise with the aggregated feature map F to obtain the corresponding channel attention intermediate feature map. Intermediate feature map with spatial attention and After feature fusion is performed by the first CBS module and the second CBS module respectively, the feature map is multiplied element-wise with the aggregated feature map F to obtain the output feature map for the next level module.

[0032] Furthermore, the target detection module is an improved YOLOv8 network model, and the improved YOLOv8 network model includes a Backbone subnetwork, a Neck subnetwork, and a Head subnetwork;

[0033] The Neck subnetwork includes the following sequentially cascaded modules: CBAM_α module I, Upsample module I, Concat module I, C2f module I, CBAM_α module II, Upsample module II, Concat module II, C2f module II, CBAM_α module III, Upsample module III, Concat module III, C2f module III, CBAM_α module IV, CBS module I, Concat module IV, C2f module IV, CBAM_α module V, CBS module II, Concat module V, and C2f module V.

[0034] Specifically, the output of C2f module I is also connected to the input of Concat module V; the output of C2f module II is also connected to the input of Concat module IV; the output of C2f module III is also used as the input of the first level DetectI of the Head sub-network; the output of C2f module IV is also used as the input of the second level DetectII of the Head sub-network; and the output of C2f module V is also used as the input of the third level DetectIII of the Head sub-network. The feature map sizes of the outputs of the three levels are different.

[0035] The Backbone sub-network includes a first-level CBS module, a first-level C2f module, a second-level CBS module, a second-level C2f module, a third-level CBS module, a third-level C2f module, a fourth-level CBS module, a fourth-level C2f module, and an SPPF module connected in sequence.

[0036] The output of the first-level C2f module of the Backbone subnetwork is also connected to the input of Concat module III, the output of the second-level C2f module is also connected to the input of Concat module II, the output of the third-level C2f module is also connected to the input of Concat module I, and the output of the SPPF module is connected to the input of CBAM_α module I.

[0037] Furthermore, the WIoU v3 function is used as the bounding box regression loss function for training the improved YOLOv8 network model.

[0038] Beneficial effects:

[0039] 1. This invention provides an infrared moving target detection system based on an improved YOLOv8, proposing a dual-input target detection model. The input can be divided into single-input and multi-input, where the multi-input includes both motion information and original frame information. Specifically, the dual-input processing module uses detection confidence to control the activation of the optical flow processing module and effectively fuses features from the main input and secondary input. The dual-input processing module replaces the input of the YOLOv8 model, efficiently fusing the main and secondary inputs and avoiding the use of two YOLO networks to process features; only one YOLO network is needed after fusion. Furthermore, this invention uses detection confidence to control the activation of the secondary input. When the detection confidence is high, a single input can complete the detection task well, eliminating the need for the secondary input, thus maximizing computational efficiency and ensuring real-time detection. When the detection confidence is low, the secondary input assists the main input in target detection, ensuring detection performance. This allows the invention to stably and quickly process the detection of fast-moving small infrared objects in complex backgrounds.

[0040] 2. This invention provides an infrared moving target detection system based on an improved YOLOv8. The two inputs have different image features. The main input is the current frame of the original image, which contains a large number of pixels required for object localization, but also contains noise pixels and a large number of complex background pixels. The secondary input is an optical flow processed image, in which the part with a larger optical flow vector represents the position of the object. Since the main input and secondary input have different features, this invention uses different feature extraction methods to extract features from multiple input images to ensure full mining and utilization of information.

[0041] 3. This invention provides an infrared moving target detection system based on an improved YOLOv8. It adopts the method of adding a momentum term ρ to accelerate the calculation speed of the optical flow field, which can shorten the convergence time and ensure the high stability of the optical flow field. In addition, the weight function set by this invention can suppress the optical flow calculation of background pixels, thereby reducing the overall computational burden.

[0042] 4. This invention provides an infrared moving target detection system based on an improved YOLOv8. The Concat operation in the dual-input processing module only superimposes the features of the two inputs into channels, which can keep the features of the two inputs independent in the channel direction. Even if the secondary input fails, the features of the secondary input will not contaminate the features of the primary input, thus increasing the stability of the algorithm.

[0043] 5. This invention provides an infrared moving target detection system based on an improved YOLOv8, which replaces the "cascaded connection" of the traditional CBAM with "parallel connection" to improve the serial attention module of CBAM. This allows both attention modules of this invention to directly learn the original input feature map, reducing the impact of the disadvantages in the original input feature map, without paying attention to the spatial and channel attention order.

[0044] 6. This invention provides an infrared moving target detection system based on an improved YOLOv8. After performing element-wise multiplication of the channel attention and spatial attention with the original features and then performing a convolution operation, the system can extract and transform features more effectively, thereby enhancing the learning and generalization capabilities of the network. Attached Figure Description

[0045] Figure 1 The flowchart of the dual-input YOLOv8 algorithm provided by this invention;

[0046] Figure 2 A diagram of the dual-input YOLOv8 model provided by this invention;

[0047] Figure 3 A flowchart of the dual-input processing module provided by the present invention;

[0048] Figure 4 The network diagram of the dual-input YOLOv8 model provided by this invention;

[0049] Figure 5 The structure diagram of the traditional attention mechanism CBAM;

[0050] Figure 6 This is a diagram of the channel attention structure.

[0051] Figure 7 This is a diagram of the spatial attention structure.

[0052] Figure 8 The CBAM_α attention mechanism structure diagram provided by this invention;

[0053] Figure 9 This is a schematic diagram illustrating the visualization results of the ISD dataset and the Anti-UAV dataset provided by this invention. Detailed Implementation

[0054] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0055] Various CNN models have made significant progress in object detection. These methods can be broadly categorized into two groups: two-stage methods and end-to-end methods. The former includes Region-CNN (R-CNN), Faster-RCNN, and Cascade R-CNN, while the latter includes SSD, RetinaNet, YOLO, Yolov3, and Fcos. Two-stage methods can achieve higher accuracy than end-to-end methods, but are computationally slower. End-to-end methods, on the other hand, offer computational advantages at the cost of lower detection performance. End-to-end methods are suitable for object detection modules in aircraft.

[0056] The YOLO algorithm has undergone several extensive enhancements, particularly with the introduction of the YOLOv8 model. YOLOv8, the latest version in the YOLO model family, achieves a significant balance between detection accuracy and efficiency. Infrared small targets, compared to complex backgrounds, are characterized by their small size, low brightness, blurred shape, and fast movement. In infrared target detection, the main improvements to the YOLO model primarily involve attention mechanisms, small object detection structures, and loss functions. This significantly improves the ability to distinguish between infrared small objects and backgrounds, as well as the model's attention to moving objects.

[0057] To address the problems existing in the prior art, this invention takes the YOLO series as the research target and selects YOLOv8s as the baseline model. This invention optimizes this model in terms of loss function, attention mechanism, and small object detection layer. Furthermore, it proposes an infrared moving target detection network and system based on the improved YOLOv8.

[0058] Specifically, such as Figure 1 and Figure 2 As shown, an infrared moving target detection system based on an improved YOLOv8 includes an optical flow processing module, a dual-input processing module, a target detection module, and an output module.

[0059] The optical flow processing module is used to process the received previous frame infrared image and current frame infrared image using an improved Horn-Schunck method when the detection confidence of the current iteration output by the output module is less than a set threshold, so as to obtain an optical flow processed image.

[0060] The dual-input processing module is used to extract features from the received image to obtain a fused feature map. When the detection confidence output by the output module is less than a set threshold, the image received by the dual-input processing module is the optical flow processed image and the current frame infrared image. When the detection confidence output by the output module for the current iteration is not less than the set threshold, the image received by the dual-input processing module is only the current frame infrared image.

[0061] The target detection module is used to detect infrared moving targets and the detection confidence of infrared moving targets in the current frame infrared image based on the fused feature map;

[0062] The output module is used to output the detection confidence level corresponding to the current frame infrared image. Detection confidence level corresponding to the previous infrared image frame The mean of the values ​​is used as the detection confidence for the next iteration.

[0063] In other words, this invention compares the mean detection confidence score with a predefined detection confidence score threshold, such as 0.5. Based on the comparison result, it determines whether the optical flow processing module is activated. That is, the dual-input processing module uses the detection confidence score to control the activation of the optical flow processing module and effectively fuses the features from the main input and the secondary input. The dual-input processing module replaces the input of the YOLOv8 model. The advantages of this are: ① It efficiently fuses the main and secondary inputs, avoiding the use of two YOLO networks to process features; only one YOLO network is needed after fusion. ② It uses the detection confidence score to control the activation of the secondary input. When the detection confidence score is high, a single input can complete the detection task well, eliminating the need for the secondary input, thus saving computational power and ensuring real-time detection. When the detection confidence score is low, the secondary input assists the main input in target detection, ensuring detection performance. ③ This design creates a closed loop in the target detection model, allowing the model to be controlled between single and dual input states by adjusting the detection confidence score threshold.

[0064] It should be noted that the iterative equation of the traditional Horn-Schunck method is

[0065]

[0066] The optical flow calculated by the traditional Heterogeneous Optical Flow (HS) method may be unreasonable at certain points. Within the field of view of an aircraft, the number of pixels for small infrared objects is very small, and the outlines of moving object regions are relatively blurred. The HS method requires a certain number of iterations to meet the accuracy requirements, which not only reduces computational efficiency but also increases interference from complex backgrounds. This patent improves the iterative equation of the HS method and introduces a pyramid calculation method.

[0067] Specifically, the optical flow processing module of this invention uses an improved Horn-Schunck method to process the received previous frame infrared image and current frame infrared image to obtain the optical flow processed image as follows:

[0068] Initial velocity component of optical flow in the u direction u (0) and the initial velocity component of the optical flow in the v direction v (0)The initial value is usually set to 0, or an initial value is set according to the experimental requirements.

[0069] The same convolution operation is used to perform convolution operations on the grayscale images of the previous frame infrared image and the current frame infrared image respectively. The convolution results of the two frames infrared images are then fused to obtain the first convolution grayscale image.

[0070] Two opposite convolution operations are performed on the grayscale images of the previous frame infrared image and the current frame infrared image, respectively. The convolution results of the two frames infrared images are then fused to obtain the second convolution grayscale image.

[0071] Give the initial velocity component u of the optical flow in the u direction (0) and the initial velocity component of the optical flow in the v direction v (0) Set initial values;

[0072] The initial velocity component of the optical flow u (0) and the initial velocity component of optical flow v (0) After iterating the following optical flow equation a set number of times using the initial value as the initial value, the optical flow velocity component u is obtained. * and optical flow velocity component v * :

[0073]

[0074] Where k represents the number of iterations, I x I represents the partial derivative of the image grayscale pixel (x,y) on the x-axis in the first convolution grayscale image. y I represents the partial derivative of the image grayscale pixel (x, y) with respect to y in the first convolutional grayscale image. t Let λ represent the partial derivative of the image grayscale pixel (x,y) with respect to t on the second convolutional grayscale image, λ represent the optical flow constant controlling the smoothness of the optical flow field, and ρ represent the momentum constant. This represents the local average optical flow velocity in the neighborhood of the current pixel in the u direction, obtained from the k-th iteration. u represents the local average optical flow velocity in the neighborhood of the current pixel in the v direction, obtained from the k-th iteration. (k-1) u k u (k+1) Let v represent the optical flow velocity components in the u direction obtained in the (k-1), k, and k+1 iterations, respectively. (k-1) v (k) v (k+1) Let ω1(x,y) represent the optical flow velocity components in the v direction obtained in the (k-1), k, and k+1 iterations, respectively, and let ω2(x,y) represent the first and second weights, respectively. Furthermore, we have:

[0075]

[0076] Where α1 is the weight for suppressing the magnitude of the optical flow vector, α2 is the ratio of the optical flow equation constraint to the compliance constraint, α1 and α2 are generally taken as numbers greater than 20, and T is the set threshold.

[0077] The optical flow velocity components u at each pixel * and optical flow velocity component v * An optical flow field is formed, and an optical flow processing image is obtained through visualization processing.

[0078] Furthermore, for simplified calculation, this invention designs an approximate piecewise function to replace the power function. Let α1 = α2, then the first weight ω1(x,y) and the second weight ω2(x,y) are simplified to the following piecewise function ω(x,y):

[0079]

[0080] Here, k' is the slope related to the values ​​of α1 and α2. Setting different k' parameter values ​​can exclude certain pixels from the iterative calculation of the HS method and suppress their optical flow vectors.

[0081] It should be noted that the hyperparameter settings for the improved Horn-Schunck method are as follows:

[0082] 1. α1, α2, and k′ are the parameters of the weighting function. This invention suggests setting the values ​​of parameters α1 and α2 to [20, 80]. The larger the values ​​of α1 and α2, the fewer iterations are required. The parameter k′ can be adjusted according to specific experimental needs. It is suggested that the combination of α1, α2, and k′ be (20, 0.8) or (30, 0.9).

[0083] 2. The general setting of the T value depends on the spatial variation of pixel values, illumination variations, and the specific scene. For scenes with large moving targets or objects hovering near the edges of the image, a range of [15, 30] is recommended. For scenes with smaller moving targets, a range of [5, 15] is recommended.

[0084] 3. The momentum constant ρ should be in the range of [0,1], and it is recommended not to exceed 0.8 in order to obtain a stable optical flow field and reduce the number of iterations.

[0085] 4. The optical flow constant λ controls the smoothness of the optical flow field. A larger λ results in a smoother flow field, while a smaller λ results in more detailed flow field variations. The choice of λ value is influenced by data characteristics and the application scenario. For scenarios with complex backgrounds, high noise, small objects, large object motion, low image resolution, and low brightness, the λ value should be within the range of [50, 500]. For different datasets, the value of λ can be adjusted from zero.

[0086] 5. Based on many experimental results, the optimal number of iterations is in the range of [10, 50] or fixed at 30.

[0087] Furthermore, the two inputs have different image features. The primary input of this invention is the current frame infrared image, which serves as the original image. This input contains a large number of pixels needed for object localization, but also includes noise pixels and a large number of complex background pixels. The secondary input is the optical flow processed image, where the larger optical flow vectors represent the object's position. Because the primary and secondary inputs have different features, different feature extraction methods are needed to extract features from multiple input images separately to ensure full information mining and utilization.

[0088] Specifically, the dual-input processing module of the present invention includes a first Focus submodule, a second Focus submodule, a Maxpool layer, an Avgpool layer, a first Concat submodule, a second Concat submodule, a first CBS convolution submodule, and a second CBS convolution submodule.

[0089] When the detection confidence level output by the output module is less than a set threshold, the first Focus submodule operates on the current frame infrared image to obtain a first Focus feature map; the second Focus submodule operates on the optical flow processed image to obtain a second Focus feature map; the second Focus feature map is pooled by a Maxpool layer and an Avgpool layer respectively to obtain a first pooled feature map and a second pooled feature map; the second Focus feature map, the first pooled feature map, and the second pooled feature map are superimposed by the second Concat submodule to obtain an optical flow feature map; the optical flow feature map is fused by the second CBS convolution submodule to obtain an optical flow fused feature map; the first Focus feature map and the optical flow fused feature map are superimposed by the first Concat submodule to obtain an intermediate fused feature map; the intermediate fused feature map is fused by the first CBS convolution submodule to obtain the final fused feature map;

[0090] When the detection confidence of the current iteration output by the output module is not less than the set threshold, the current frame infrared image is directly used as the final fused feature map.

[0091] For example, taking a single-channel image of 1×256×256 as the main input, a Focus module with parameters "1,32" yields a feature dimension of 32×128×128. Feature extraction from the optical flow image is performed first using a Focus module with parameters "1,16", resulting in a feature C dimension of 16×128×128. Next, a 5×5 window Maxpooling and Avgpooling layers are used for processing. Finally, a Concat operation is used to stack the features of the optical flow image along the channel direction, resulting in a feature dimension of 48×128×128. A CBS convolutional unit is then used to reduce the dimensionality of the stacked features to 32×128×128. Finally, the features processed from the original image of the current frame and the features processed from the optical flow image are stacked using a Concat operation. The Concat operation only performs channel-wise concatenation of the features from both inputs, maintaining the independence of the two input features in the channel direction. Even if the secondary input fails, its features will not contaminate the features of the primary input, thus increasing the algorithm's stability. The final output dimension of the dual-input module is 64×128×128. The flowchart of the input processing module is as follows... Figure 3 As shown.

[0092] The following is about Figure 3 The modules involved will be explained.

[0093] CBS stands for Conv+BN+SiLU, meaning it consists of a Conv convolutional layer, a Batch Normalization (BN) layer, and a SiLU activation function layer. For a detailed explanation of the CBS structure, see [link to CBS documentation]. Figure 4 .

[0094] The concat operation joins two or more feature maps together along their dimensions to generate a larger feature map. In YOLOv8, feature maps from different layers are typically concatenated along the depth (number of channels) direction to form a richer feature representation. This cross-layer connection can occur at different layers, allowing the model to utilize feature information from both shallow and deep layers simultaneously.

[0095] The Focus module consists of a slice operation, a Concat operation, and a CBS unit. First, a slice operation is performed, slicing the input image into four sub-images by dividing it into pixels across its width and height, resulting in four sub-images. Next, a Concat operation is performed, concatenating these four sub-images along the channel dimension. This transfers the width and height information of the original image to the channel dimension, quadrupling the number of channels. Finally, convolution processing is performed, where the concatenated feature map is processed by a convolutional layer to extract features. The slice operation in Focus is similar to nearest-neighbor downsampling, reducing the input size while increasing the number of channels. Because the slice operation reduces the input size, the size and number of parameters in subsequent parts of the network decrease accordingly. Therefore, the slice operation in Focus reduces the overall computational cost of the network, thus speeding up the process. More importantly, the slicing process does not cause any loss of input information. For a detailed description of the Focus structure, see [link to Focus architecture]. Figure 4 .

[0096] Maxpooling divides the input feature map into several rectangular regions, then selects the maximum value from each region as the representative of that region, and outputs it to the next layer. This method helps extract salient features from the image and has some invariance to small changes in the input data. Avgpooling calculates the average of all elements within each rectangular region and outputs this average. Average pooling is often used to preserve background information or reduce the size of the feature map. The main difference between the two is that Maxpooling tends to retain the strongest feature signals, while Avgpooling provides an average representation of all features within a region.

[0097] It is particularly important to note that in this invention, the system uses only the original frame as input to detect the first frame and set... The initial value is 0.5. To avoid constant switching between single and multiple inputs, the input processing module combines the confidence value of the previous frame's output, enhancing system stability. This method ensures detection accuracy while saving computational resources, thereby reducing the time required for the model detection process.

[0098] Furthermore, the infrared moving target detection system of the present invention is an infrared moving small target detection model for aircraft based on the YOLOv8 model. The dual-input YOLOv8 model uses an optical flow processing module to process multi-frame motion information and compare it with the original frame. Figure 1 The input is fed into the input processing module, where it is processed and then used as input to the YOLOv8 object detection module for object detection. The dual-input YOLOv8 model network is as follows: Figure 4 As shown.

[0099] Specifically, such as Figure 4As shown, the target detection module of the present invention is an improved YOLOv8 network model, and the improved YOLOv8 network model includes a Backbone sub-network, a Neck sub-network, and a Head sub-network.

[0100] The Neck subnetwork includes the following sequentially cascaded modules: CBAM_α module I, Upsample module I, Concat module I, C2f module I, CBAM_α module II, Upsample module II, Concat module II, C2f module II, CBAM_α module III, Upsample module III, Concat module III, C2f module III, CBAM_α module IV, CBS module I, Concat module IV, C2f module IV, CBAM_α module V, CBS module II, Concat module V, and C2f module V.

[0101] Specifically, the output of C2f module I is also connected to the input of Concat module V; the output of C2f module II is also connected to the input of Concat module IV; the output of C2f module III is also used as the input of the first level DetectI of the Head sub-network; the output of C2f module IV is also used as the input of the second level DetectII of the Head sub-network; and the output of C2f module V is also used as the input of the third level DetectIII of the Head sub-network. The feature map sizes of the outputs of the three levels are different.

[0102] Each level of the Detect layer processes the output from the Neck sub-network and outputs the final detection result. However, with each pass through the CBS convolution kernel, the length and width of the feature map are reduced to half of their original size. Therefore, the first-level output DetectI receives the largest feature map, covering more comprehensive feature information, and has a smaller receptive field, making it suitable for detecting smaller targets. The third-level output DetectIII receives the smallest feature map, covering less comprehensive feature information, and has a larger receptive field, making it suitable for detecting larger targets. Thus, due to the different sizes of the received feature maps, the YOLOv8 network model can perform target detection at different scales. The third-level output DetectIII of the improved YOLOv8 network model of this invention is equivalent to the second-level output DetectII in the traditional YOLOv8 network model, while the first-level output DetectI has the ability to detect smaller targets. Therefore, the detection system of this invention can detect smaller infrared moving targets.

[0103] The Backbone sub-network includes a first-level CBS module, a first-level C2f module, a second-level CBS module, a second-level C2f module, a third-level CBS module, a third-level C2f module, a fourth-level CBS module, a fourth-level C2f module, and an SPPF module connected in sequence.

[0104] The output of the first-level C2f module of the Backbone subnetwork is also connected to the input of Concat module III, the output of the second-level C2f module is also connected to the input of Concat module II, the output of the third-level C2f module is also connected to the input of Concat module I, and the output of the SPPF module is connected to the input of CBAM_α module I.

[0105] In other words, the C2f module and the Upsample module in the Neck sub-network are connected through the CBAM_α module, the C2f module and the CBS module are connected through the CBAM_α module, and the SPPF module in the Backbone sub-network is connected to the first Upsample module in the Neck sub-network through the CBAM_α module.

[0106] It should be noted that, Figure 4 The "Submodule Structure Diagram" section shows the structure of each submodule used in the model. "Upsample" is an upsampling operation used to increase the resolution of the feature map, which is particularly important for improving the model's ability to detect small objects. Specifically, upsampling enlarges the size of the feature map using interpolation methods without adding new information. In YOLOv8, this is typically achieved using bilinear interpolation or nearest-neighbor interpolation. This makes the feature map more spatially dense, thus aiding in the detection of smaller objects.

[0107] Structurally, this invention uses a small target detection layer and a large target detection layer, which makes the target detection model more adaptable to the infrared small target detection environment.

[0108] It should be noted that this invention adds an attention mechanism CBAM_α after the C2f module in the Neck sub-network. The design concept of the attention mechanism CBAM_α is derived from CBAM. CABM is a lightweight attention mechanism module that can be easily embedded into existing popular convolutional neural network structures without any additional computation. To achieve the scaling operation on the original features, it uses two pooling methods: max pooling and average pooling. It generates weights based on two-dimensional (channel and spatial) information. Existing CBAM structures, such as... Figure 5 As shown.

[0109] Given an intermediate feature map F∈R C×H×W As input, CBAM sequentially infers the channel attention M. c ∈R C×1×1 Spatial attention M s ∈R 1×H×W The calculation formula is as follows:

[0110]

[0111] in, This represents element-wise multiplication. During multiplication, attention values ​​are copied accordingly. F2 is the final refined output.

[0112] The structure of channel attention, such as Figure 6 The calculation is as follows:

[0113]

[0114] in, and These represent the average pooling feature and the max pooling feature, respectively. The shared network consists of a multilayer perceptron (MLP) with one hidden layer. To reduce parameter overhead, the hidden activation size is set to R. C / r×1×1 , where r is the reduction ratio. σ represents the sigmoid function, W0∈R C / r×C W1∈R C×C / r The weights of the MLP are represented by W0 and W1, which are shared by the two inputs, and the ReLU activation function is followed by W0.

[0115] The structure of spatial attention, such as Figure 7 As shown, the calculation is as follows:

[0116]

[0117] in, and Let f represent the average pooling feature and max pooling feature of the entire channel, respectively. Then, they are concatenated and convolved through a standard convolutional layer to generate a 2D spatial attention map. σ represents the sigmoid function, f... 7×7 This represents a convolution operation with a filter size of 7×7.

[0118] However, regardless of whether spatial attention is enabled before channel attention, or vice versa, the weights ranked lower will be generated from the feature maps ranked higher. To some extent, the input of the attention features ranked lower is influenced by the preceding attention mechanism, which can cause interference and make the model unstable. This invention designs a new CBAM attention mechanism named CBAM_α. In this invention, the original "cascaded connection" is replaced with "parallel connection" to improve the serial attention module of CBAM. Therefore, both attention modules directly learn the original input feature maps without regard to the order of spatial and channel attention.

[0119] The detailed structure of CBAM_α is as follows: Figure 8As shown, the CBAM_α module includes a channel attention module, a spatial attention module, a first CBS module, and a second CBS module; the aggregated feature map F is processed by the channel attention module and the spatial attention module respectively to obtain the corresponding channel attention feature map M. c (F) and spatial attention feature map M s (F); Channel attention feature map M c (F) and spatial attention feature map M s (F) is multiplied element-wise with the aggregated feature map F to obtain the corresponding channel attention intermediate feature map. Intermediate feature map with spatial attention and After feature fusion is performed by the first CBS module and the second CBS module respectively, the feature map is multiplied element-wise with the aggregated feature map F to obtain the output feature map for the next level module.

[0120] It should be noted that in this invention, after the channel attention and spatial attention are multiplied element-wise with the original features, a convolution operation is performed. This is to extract and transform features more effectively, enhance the learning ability and generalization ability of the network, and then perform element-wise multiplication again to obtain the final output.

[0121] It's important to note that the loss function for bounding box regression considers various geometric factors associated with the bounding box. Generalized IoU (GIoU) ​​is proposed by incorporating the concept of a minimum closed region, which represents the smallest rectangular box that can enclose both the predicted and ground truth boxes. Distance-IoU (DIoU) is an extension of GIoU, designed by incorporating the distance between the centers of the predicted and ground truth boxes. Furthermore, complete-IoU (CIoU) is developed based on DIoU, taking into account the aspect ratio differences between the predicted and ground truth boxes. Efficient-IoU (EIoU) is developed by replacing aspect ratio consistency in CIoU with a separate aspect ratio consistency calculation. As a loss function for bounding box regression in YOLOv8, CIoU can hinder efficient model optimization when the predicted and ground truth boxes have the same aspect ratio but different width and height values; moreover, the calculation of CIoU is complex. In infrared detection tasks, the size, brightness, and distance of objects exhibit significant variations within the field of view. The sharpness and quality of images captured by infrared thermal imagers can vary considerably. Furthermore, most existing methods assume high-quality training data, thus focusing on improving the fitting ability of bounding box regression loss. However, indiscriminately strengthening bounding box regression for low-quality data samples degrades localization performance. These drawbacks can be addressed by Wise-IoU (WIoU), which uses a focusing mechanism to allocate small-quality gradient gains, allowing the bounding box regression loss to focus on anchor boxes of average quality. WIoU can be broadly categorized into three types: WIoU v1, which constructs an attention-based bounding box loss; WIoU v2 and WIoU v3, which build upon WIoU v1 by adding a focusing mechanism through gradient gain methods. WIoU v3 employs a reasonable gradient gain allocation strategy to dynamically optimize the weights of high-quality and low-quality anchor boxes in the loss, enabling the model to focus on average-quality samples and improving overall model performance.

[0122] In other words, the traditional YOLOv8 model uses distributed focus loss and CIoU to calculate the bounding box regression loss. However, CIoU has certain drawbacks. First, CIoU does not consider the balance of different samples. Second, CIoU introduces aspect ratio as a penalty factor into the loss function. However, if the aspect ratios of the ground truth and predicted boxes are the same, but their width and height values ​​are different, the penalty term cannot accurately distinguish between the two boxes. Third, the CIoU formula involves the calculation of inverse trigonometric functions, which increases the overall computational cost of the model. Based on this, this invention uses WIoU v3 as the bounding box regression loss function, the expression of which is shown below:

[0123]

[0124] The elements in Equation 11 are calculated using Equations 12-15.

[0125]

[0126] r = β / δα β-δ (14)

[0127]

[0128] IoU measures the degree of overlap between two regions by calculating the ratio between their intersection and their union; W g and H g Represents the width and height of the smallest bounding box formed by the true bounding box and the predicted bounding box; and δ and α represent the coordinates of the center points of the ground truth and predicted boxes, respectively; β is the outlier value that measures the quality of the anchor box. A non-monotonic focusing factor r is constructed based on β and applied to WIoU v1; δ and α are hyperparameters that can be adjusted according to different models. These are the coefficients of the constructed monotonic focusing factor; It is the average value of momentum m.

[0129] A smaller β value indicates a higher quality anchor frame, which is assigned a smaller r value to reduce the weight of high-quality anchor frames in a larger loss function. A larger β value indicates a low-quality anchor frame, which is assigned a smaller gradient gain to reduce the harmful gradients generated by low-quality anchor frames. Construct monotonic focusing coefficients. This effectively reduces the weight of simple instances in the loss value. However, considering the model training process... along with The decrease in the value leads to a slower convergence rate, therefore, the following is introduced: The average value is used to compare Normalize.

[0130] WIoU v3 employs a reasonable gradient gain allocation strategy to dynamically optimize the weights of high-quality and low-quality anchor boxes in the loss function, allowing the model to focus on average-quality samples and improving overall performance. On one hand, WIoU v3 combines some advantages of EIoU and SIoU, aligning with the design principles of an excellent loss function. On the other hand, WIoU v3 uses a dynamic non-monotonic mechanism to evaluate anchor box quality, making the model pay more attention to anchor boxes of average quality, thus improving its ability to locate targets. For infrared small target detection tasks, where small objects constitute a high proportion and increase detection difficulty, WIoU v3 can dynamically optimize the loss weights for small objects to improve the model's detection performance.

[0131] The improved dual-input YOLOv8 model is used for training and evaluation. During the training process, it is necessary to first select appropriate hyperparameters, optimizers and loss functions, and reasonably divide the training set, validation set and test set for model training and evaluation.

[0132] Analyze and evaluate the data after model training, adjust and optimize the model, retain the optimal weight model after training, and conduct experimental testing on the trained model. Visualization tools can be used to analyze the trends of changes in metrics such as loss and precision, and adjust and optimize the model to improve its performance. If the modified model shows improvements in loss, precision, recall, and AP... 50 AP 50:95 If a model demonstrates superior performance in terms of parameters compared to the original model, then the current training result is selected as the optimal model. Otherwise, if the model training is deemed unsatisfactory, the training parameters are adjusted further, and a new round of training is conducted. The optimal weight model after training is retained. The resulting optimal weight model, after selection, outperforms the original model in terms of precision, recall, and AP. 50 AP 50:95 To be promoted.

[0133] The experimental platform of this invention runs on the Windows 10 operating system, and the training platform uses an Intel Core i9-13900K processor and NVIDIA... TM The system uses an RTX 3090Ti GPU, runs on PyTorch 1.9.2, and is powered by Python 3.8 and CUDA 11.4.

[0134] In the model training process of this invention, stochastic gradient descent with a momentum of 0.937 was employed. The initial learning rate was set to 0.03, and the final learning rate was set to 0.0001. The batch size was 32, the training duration was 200 epochs, and the input image size was 256×256. The weight decay was 0.0005. Mosaic was used as a data augmentation strategy, but mosaic data augmentation was stopped in the last 10 epochs of training to accelerate convergence. During the testing phase, the confidence threshold for the object was 0.001, the non-maximum suppression IoU threshold was 0.6, and the batch size was 1.

[0135] To verify the effectiveness of the algorithm in this invention, comparative experiments were designed using currently mainstream single-stage object detection algorithms. The purpose of these comparative experiments is to evaluate the performance of the proposed dual-input YOLOv8 model and other mainstream models, including YOLOv3, YOLOv5s, YOLOv7-tiny, and YOLOv8s. These methods were chosen for comparison because they are representative or advanced methods in the field of infrared detection. These models differ in their backbone network, neck network, and prediction head, thus the comparative experiments will help evaluate the performance of each model in different scenarios. Specifically, the comparative experiments will evaluate precision, recall, and AP. 50 AP 50:95 We will conduct evaluations in various aspects to facilitate comparative analysis.

[0136] Table 1 Comparison of experimental results

[0137]

[0138] The experimental results show that the YOLOv8 model, as the latest algorithm in the YOLO series, has been proven to have higher detection performance compared to other existing algorithms. This is also the main motivation for using this model as the benchmark model in this study. The proposed two-input YOLOv8 model exhibits the highest average detection accuracy and the best overall detection performance among all models, with a real-time running speed of 77 FPS. Although the improved model has a slightly longer inference time compared to the YOLOv8 model, it can still achieve real-time detection, outperforming most detection algorithms.

[0139] To further verify the effectiveness of the algorithm performance of this invention, ablation experiments were conducted. In Table 2, the algorithms without WIoU v3 and CBAM_α have integrated the small target detection layer and the dual-input mode into the model.

[0140] Table 2 Ablation Experiment Results

[0141]

[0142] As shown in Table 2, the experimental results indicate that, firstly, compared to single-frame detection methods, multi-frame detection methods can capture more motion and object information, and the input processing module can effectively fuse features from multiple inputs, saving computational resources. Secondly, among single-frame detection models, YOLOv8s exhibits the best detection performance, therefore it was chosen as the benchmark model in this invention. WIoU v3 uses a more reasonable sample allocation strategy to improve the model's localization ability. Furthermore, CBAM_α enhances the model's noise suppression capability, making the model focus more on key information. Finally, the added small target detection layer improves the model's ability to detect small targets. In summary, when detecting small moving infrared objects in complex backgrounds, the dual-input YOLOv8 model maintains good performance, while the performance of other methods declines.

[0143] It should be noted that the dual-input YOLOv8 model was tested on the Infrared Shimmering Aircraft Detection and Tracking dataset (ISD) and compared with existing methods. Experimental results show that, compared with existing methods, the dual-input YOLOv8 model significantly improves detection performance in terms of accuracy, recall, AP50, and AP50:95, while maintaining real-time processing speed.

[0144] from Figure 9 The visualization results show that the two-input YOLOv8 model is highly effective for detecting infrared moving targets in complex backgrounds. The proposed model demonstrates excellent detection performance even with small, dimly lit objects and complex environments, achieving higher detection confidence and avoiding false positives and false negatives. It also achieves good performance on different datasets, proving the good generalization ability of the two-input YOLOv8 model.

[0145] Therefore, this invention addresses the challenges of current infrared moving target detection methods, such as the lack of obvious features and the presence of significant background motion in complex environments. It proposes an infrared moving target detection system based on an improved YOLOv8 model, named the Dual-Input YOLOv8 model. Specifically, the improvements of this invention are summarized as follows:

[0146] 1. The original YOLOv8 input layer is replaced by a feature map obtained by fusing the optical flow processed image with the original frame infrared image. Specifically, the dual-input YOLOv8 model uses the current frame and the optical flow processed image as input (the current frame is the main input and the optical flow processed image is the secondary input) for infrared moving small object detection. This model can obtain original feature information, object information and motion information from the input, thereby improving detection performance.

[0147] 2. The optical flow processing module in the dual-input YOLOv8 model uses the improved Horn-Schunck (HS) optical flow processing method;

[0148] 3. To optimize the utilization of multiple input information, the dual-input YOLOv8 model proposes an input processing module that determines the input composition based on the detection confidence. This enables the input processing module to process the feature information carried by each input more efficiently, allowing the model to switch between single-input and multi-input modes. It can effectively fuse features from multiple inputs and balance detection performance and real-time performance, thereby improving the model's detection performance and saving computing power.

[0149] 4. A small target detection layer was added to the traditional YOLOv8 model, and a detection layer for large targets was removed. This makes the model more focused on the detection of small targets and makes the model more lightweight, thereby achieving better detection performance in infrared small target detection and avoiding the decrease in real-time performance caused by increasing the number of model layers. The detection layer for large targets was removed to reduce the model size for later embedded deployment.

[0150] 5. Using Wise-IoU (WIoU)v3 instead of CIoU as the bounding box regression loss function of the model, experiments have verified that its performance is better than the existing bounding box regression loss function. It can effectively balance the gradient gain of high and low quality samples. At the same time, this improvement significantly enhances the model's localization ability and bounding box regression accuracy, while also increasing the flexibility of the target detection model in various scenarios and improving the model's ability to locate moving targets.

[0151] 6. The dual-input YOLOv8 model also uses an improved Convolutional Block Attention Module (CBAM) to enhance the feature representation capabilities of the convolutional neural network. CBAM focuses on important features of the image and suppresses unnecessary region responses, allowing the model to pay more attention to key information in the feature map and improving the model's detection performance.

[0152] Of course, the present invention may have other various embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.

Claims

1. An infrared moving target detection system based on an improved YOLOv8, characterized in that, It includes an optical flow processing module, a dual-input processing module, a target detection module, and an output module; The optical flow processing module is used to process the received previous frame infrared image and current frame infrared image using an improved Horn-Schunck method when the detection confidence of the current iteration output by the output module is less than a set threshold, to obtain an optical flow processed image. Specifically: The same convolution operation is used to perform convolution operations on the grayscale images of the previous frame infrared image and the current frame infrared image respectively, and the convolution results of the two frames infrared images are fused to obtain the first convolution grayscale image. Two opposite convolution operations are performed on the grayscale images of the previous frame infrared image and the current frame infrared image, respectively. The convolution results of the two frames infrared images are then fused to obtain the second convolution grayscale image. Give u Initial velocity component of optical flow in the direction u (0) and v Initial velocity component of optical flow in the direction v (0) Set initial values; The initial velocity component of optical flow u (0) and the initial velocity component of optical flow v (0) After iterating a set number of times using the initial values ​​as the initial values ​​for the following optical flow equation iterative formula, the optical flow velocity components are obtained. u and optical flow velocity components v : in, k Indicates the number of iterations. I x Represents the image grayscale pixels on the first convolutional grayscale image. exist x Partial derivatives on the axis, I y Represents the image grayscale pixels on the first convolutional grayscale image. exist y Partial derivatives on, I t Represents the image grayscale pixels on the second convolution grayscale image. right t The partial derivatives, The optical flow constant represents the smoothness of the optical flow field. ρ Represents the momentum constant. Indicates the first k The current pixel obtained in the next iteration is... u Set the local average value of the optical flow velocity in the neighborhood in the direction. Indicates the first k The current pixel obtained in the next iteration is... v Set the local average value of the optical flow velocity in the neighborhood in the direction. , , They represent the first k -1st time, the first k sequence k +1 iterations u The optical flow velocity component in the direction, , , They represent the first k -1st time, the first k sequence k +1 iterations v The optical flow velocity component in the direction, and Let represent the first weight and the second weight respectively, and we have: in, To suppress the weight of the optical flow vector magnitude, To adjust the ratio of optical flow equation constraints to compliance constraints, To set a threshold; The optical flow velocity components at each pixel u and optical flow velocity components v The optical flow field is formed, and after the required visualization processing, an optical flow processed image is obtained. The dual-input processing module is used to extract features from the received image to obtain a fused feature map. When the detection confidence output by the output module is less than a set threshold, the image received by the dual-input processing module is the optical flow processed image and the current frame infrared image. When the detection confidence output by the output module for the current iteration is not less than the set threshold, the image received by the dual-input processing module is only the current frame infrared image. The target detection module is used to detect infrared moving targets in the current frame infrared image and the detection confidence of the infrared moving targets based on the aggregated feature map; wherein, the target detection module is an improved YOLOv8 network model, and the improved YOLOv8 network model includes a Backbone sub-network, a Neck sub-network, and a Head sub-network; wherein, the C2f module and the Upsample module in the Neck sub-network communicate via CBAM_ α The modules are connected; the C2f module and the CBS module communicate via CBAM_ α The modules are connected, and simultaneously, the SPPF module in the Backbone subnetwork and the first Upsample module in the Neck subnetwork communicate via CBAM_ α The modules are connected; The CBAM_ α The module includes a channel attention module, a spatial attention module, a first CBS module, and a second CBS module; aggregated feature maps. After passing through the channel attention module and the spatial attention module respectively, the corresponding channel attention feature maps are obtained. Spatial attention feature map Channel attention feature map Spatial attention feature map Separately with aggregated feature maps After element-wise multiplication, the corresponding intermediate feature map of channel attention is obtained. Intermediate feature map with spatial attention ; and After feature fusion is performed by the first CBS module and the second CBS module respectively, the feature map is then aggregated. Element-wise multiplication is performed to obtain the output feature map that is output to the next level module; The output module is used to take the average of the detection confidence scores of the current frame infrared image and the previous frame infrared image as the detection confidence score for the next iteration.

2. The infrared moving target detection system based on improved YOLOv8 as described in claim 1, characterized in that, make = Then the first weight Second weight Simplify to the following piecewise function : in, To and , The slope related to the value.

3. The infrared moving target detection system based on improved YOLOv8 as described in claim 1, characterized in that, The dual-input processing module includes a first Focus submodule, a second Focus submodule, a Maxpool layer, an Avgpool layer, a first Concat submodule, a second Concat submodule, a first CBS convolution submodule, and a second CBS convolution submodule. When the detection confidence level output by the output module is less than a set threshold, the first Focus submodule operates on the current frame infrared image to obtain a first Focus feature map; the second Focus submodule operates on the optical flow processed image to obtain a second Focus feature map; the second Focus feature map is pooled by a Maxpool layer and an Avgpool layer respectively to obtain a first pooled feature map and a second pooled feature map; the second Focus feature map, the first pooled feature map, and the second pooled feature map are superimposed by the second Concat submodule to obtain an optical flow feature map; the optical flow feature map is fused by the second CBS convolution submodule to obtain an optical flow fused feature map; the first Focus feature map and the optical flow fused feature map are superimposed by the first Concat submodule to obtain an intermediate fused feature map; the intermediate fused feature map is fused by the first CBS convolution submodule to obtain the final fused feature map; When the detection confidence of the current iteration output by the output module is not less than the set threshold, the current frame infrared image is directly used as the final fused feature map.

4. The infrared moving target detection system based on improved YOLOv8 as described in claim 1, characterized in that, The target detection module is an improved YOLOv8 network model, and the improved YOLOv8 network model includes a Backbone subnetwork, a Neck subnetwork, and a Head subnetwork. The Neck subnetwork includes sequentially cascaded CBAM_ α Module I, Upsample Module I, Concat Module I, C2f Module I, CBAM_ α Module II, Upsample Module II, Concat Module II, C2f Module II, CBAM_ α Module III, Upsample Module III, Concat Module III, C2f Module III, CBAM_ α Module IV, CBS Module I, Concat Module IV, C2f Module IV, CBAM_ α Module V, CBS Module II, Concat Module V, C2f Module V; Specifically, the output of C2f module I is also connected to the input of Concat module V; the output of C2f module II is also connected to the input of Concat module IV; the output of C2f module III is also used as the input of the first level DetectI of the Head sub-network; the output of C2f module IV is also used as the input of the second level DetectII of the Head sub-network; and the output of C2f module V is also used as the input of the third level DetectIII of the Head sub-network. The feature map sizes of the outputs of the three levels are different. The Backbone sub-network includes a first-level CBS module, a first-level C2f module, a second-level CBS module, a second-level C2f module, a third-level CBS module, a third-level C2f module, a fourth-level CBS module, a fourth-level C2f module, and an SPPF module connected in sequence. In this subnetwork, the output of the first-level C2f module is also connected to the input of Concat module III; the output of the second-level C2f module is also connected to the input of Concat module II; the output of the third-level C2f module is also connected to the input of Concat module I; and the output of the SPPF module is connected to CBAM_ α Input to Module I.

5. An infrared moving target detection system based on an improved YOLOv8 as described in claim 1 or 4, characterized in that, The WIoU v3 function is used as the bounding box regression loss function for training the improved YOLOv8 network model.

Citation Information

Patent Citations

  • Single-frame infrared weak and small target detection method based on improved YOLOv7

    CN117611911A