A statistical filtering infrared small target detection and tracking method based on YOLOv5 and Deepsort

By combining YOLOv5 and Deepsort algorithms and performing data preprocessing and postprocessing, the problem of infrared small object detection and tracking in complex backgrounds is solved, and efficient and robust infrared small object detection and tracking is achieved.

CN114677554BActive Publication Date: 2025-05-09EAST CHINA UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210179351.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-05-09
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

The prior art is difficult to accurately detect and track small infrared targets in complex contexts, and has low robustness and detection rates.

Method used

The statistical filtered infrared small object detection and tracking method based on YOLOv5 and Deepsort is adopted, and the pre-processing and post-processing of data cleaning, enhancement and background removal is achieved through the target detection capability of YOLOv5 and the target tracking algorithm of Deepsort to real-time detection and tracking of infrared small objects.

Benefits of technology

The infrared small object detection and tracking in complex backgrounds is realized, which improves the robustness and detection rate, and can achieve real-time tracking effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114677554B_ABST
    Figure CN114677554B_ABST
Patent Text Reader

Abstract

The present invention relates to a statistical filtering infrared small target detection and tracking method based on YOLOv5 and Deepsort, comprising: step S1, using infrared imaging equipment to collect infrared small target images under complex backgrounds, and obtaining infrared small target image data sets; step S2, preprocessing the infrared small target image data sets, obtaining the preprocessed image data sets, and dividing the preprocessed image data sets into a training set and a verification set; step S3, training the YOLOv5s model in the YOLOv5 algorithm, and obtaining the training set of the Deepsort model; step S4, inputting the training set of the Deepsort model into the Deepsort model for training, and constructing an infrared small target detection and tracking identifier; step S5, using the infrared small target detection and tracking identifier, performing real-time detection and tracking of the infrared small target. The present invention can accurately and quickly detect infrared small targets under complex backgrounds, and improve robustness and detection rate. In addition, the present invention can achieve real-time tracking effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection and tracking, and more specifically to a statistical filtering infrared small target detection and tracking method based on YOLOv5 and Deepsort. Background Art

[0002] At present, target detection technology is widely used in various fields. Infrared small target detection has always been a hot topic in the field of infrared image processing, and its technical research is of great significance to military early warning, pattern recognition, image processing and other fields. However, in actual scenarios, the target imaging distance in infrared images is long, and usually appears in the form of points. In addition, the infrared signal is severely weakened by the air during propagation, resulting in it usually existing in the form of a Gaussian distributed point target, and it itself does not have significant shape and texture information. In addition, infrared small targets in complex backgrounds are often submerged by noise and clutter, resulting in a very low signal-to-noise ratio of infrared images containing small targets. These characteristics have brought great challenges to infrared small target detection technology.

[0003] Traditional target detection algorithms include region box selection, feature extraction, and classification. There are two problems with this algorithm: first, the region selection strategy is not targeted and has high time complexity; second, the robustness is poor. With the continuous development of deep learning in the field of computer vision, it has made major breakthroughs in target detection. Currently, it is mainly divided into two categories: one is the second-order algorithm based on detection box and classifier, such as the R-CNN (regional convolutional neural network) series. This type of algorithm has high accuracy, but the complex network structure leads to slow detection speed and is difficult to meet real-time target detection; the other is the first-order algorithm based on regression, such as the YOLO (You Only Look Once) series. This type of algorithm has fast reasoning speed and can meet real-time detection. The fifth generation version of the YOLO series, YOLOv5, has a lighter framework and faster reasoning speed, but it can only identify and detect targets and cannot track trajectories. In addition, YOLOv5 is very prone to missed detection under poor environmental conditions.

[0004] The Deepsort tracking algorithm uses Kalman filtering to predict and update the target, and uses the Hungarian algorithm to perform data association matching between the prediction box and the detection box of the target trajectory in the cascade matching, which can track the target well. However, the existing detection and tracking methods that combine YOLOv5 and Deepsort tracking algorithms are all targeted at large targets, such as pedestrians. Therefore, it is necessary to develop a detection and tracking method for small infrared targets in complex backgrounds. Summary of the invention

[0005] In order to solve the above problems in the prior art, the present invention provides a statistical filtering infrared small target detection and tracking method based on YOLOv5 and Deepsort, which can accurately and quickly realize background separation, detect and track infrared small targets in complex backgrounds, and at the same time improve robustness and detection rate.

[0006] The present invention provides a statistical filtering infrared small target detection and tracking method based on YOLOv5 and Deepsort, comprising:

[0007] Step S1, using an infrared imaging device to collect infrared small target images under a complex background to obtain an infrared small target image data set;

[0008] Step S2, preprocessing the infrared small target image data set to obtain a preprocessed image data set, and dividing the preprocessed image data set into a training set and a verification set;

[0009] Step S3, training the YOLOv5s model in the YOLOv5 algorithm according to the training set and the verification set to obtain a training set of the Deepsort model;

[0010] Step S4, inputting the training set of the Deepsort model into the Deepsort model for training, obtaining the weight parameters of the Deepsort model, and constructing an infrared small target detection and tracking recognizer;

[0011] Step S5, using the infrared small target detection and tracking identifier to perform real-time detection and tracking of the infrared small target.

[0012] Furthermore, the step S1 comprises:

[0013] Step S11, using an infrared imaging device to perform continuous frame shooting sampling on one or more aircraft under different complex backgrounds to obtain an initial image set;

[0014] Step S12, marking a real target detection frame in each image of the initial image set, constructing label information of the infrared small target, and the initial image set and the label information of the infrared small target together constitute an infrared small target image data set.

[0015] Furthermore, the method for preprocessing the infrared small target image data set in step S2 includes:

[0016] Step S21, performing data cleaning on the infrared small target image data set to obtain a cleaned image data set;

[0017] Step S22, randomly selecting a number of frames of images from the cleaned image data set, and performing data enhancement on the selected number of frames of images to obtain an enhanced image data set;

[0018] Step S23, extracting a background image of each frame from the cleaned image data set, and performing a mixing process on the extracted background image to obtain a mixed background image set;

[0019] Step S24, extracting noise points contained in each frame of the image from the infrared small target image dataset, and randomly pasting the extracted noise points into the cleaned image dataset, the enhanced image dataset and the mixed background image dataset;

[0020] Step S25, merging the cleaned image dataset, the enhanced image dataset and the mixed background image dataset with the randomly pasted noise in step S24 to obtain a final preprocessed image dataset.

[0021] Furthermore, the data enhancement method in step 22 is a pasting enhancement method or a mosaic enhancement method.

[0022] Furthermore, the step S3 comprises:

[0023] Step S31, initializing YOLOv5s model parameters, including: batch processing size, number of iterations, image resolution, intersection-over-union ratio threshold, and confidence threshold;

[0024] Step S32, inputting the training set into the YOLOv5s model, and outputting the predicted target detection box after each iteration according to the initialized YOLOv5s model parameters;

[0025] Step S33, performing IOU loss calculation on the predicted target detection box after each iteration and the real target detection box in the label information to obtain the learning weight after each iteration;

[0026] Step S34, for the learning weights after each iteration, according to the verification set, select the learning weight with the smallest test error on the verification set, and use the learning weight with the smallest test error as the weight parameter of the YOLOv5s model;

[0027] Step S35, according to the weight parameters of the YOLOv5s model, the target recognition candidate box in each frame image is obtained, and the target recognition candidate box in each frame image constitutes a training set of the Deepsort model.

[0028] Furthermore, the batch processing size is set to 64, the number of iterations is set to 2000, the image resolution is set to 640*640, the intersection-over-union ratio threshold is set to 0.4, and the confidence threshold is set to 0.6.

[0029] Furthermore, the step S4 comprises:

[0030] Step S41, determining the real position of the infrared small target in each frame of the image according to the training set of the Deepsort model, and using Kalman filtering to determine the predicted position of the infrared small target in the k-1th (k≥2)th frame of the image according to the real position of the infrared small target in the k-1th (k≥2)th frame of the image;

[0031] Step S42, using the Hungarian algorithm to perform cascade matching on the predicted position of the infrared small target in the k-th frame image and the actual position of the infrared small target in the k-th frame image, to obtain the result of the initial successful matching, the trajectory of the initial unmatched and the detection frame of the initial unmatched;

[0032] Step S43, performing IOU matching on the initially unmatched trajectory and the initially unmatched detection frame in step S42, to obtain a rematched successful result, a rematched trajectory, and a rematched detection frame;

[0033] Step S44, updating the parameters of the Kalman filter according to the result of the first successful matching and the result of the second successful matching;

[0034] Step S45, assign a new trajectory and a new ID to the detection frame that is not matched again, and extract the feature set of the target object in the detection frame through ReID; at the same time, determine whether the trajectory that is not matched again is in a determined state, retain the trajectory that is in a determined state and has a mismatch number of less than 30 frames, and repeat steps S41-step S44.

[0035] , step S5 comprises:

[0036] Step S51, setting parameters in the recognizer for infrared small target detection and tracking, and loading the weight parameters of the YOLOv5s model into the YOLOv5s model, and loading the weight parameters of the Deepsort model into the Deepsort model;

[0037] Step S52, using an infrared imaging device to collect a real-time image of the small infrared target, and performing contrast enhancement processing on the collected real-time image to obtain an enhanced real-time image;

[0038] Step S53, performing pixel average statistics on the first 100 frames of the enhanced real-time image, obtaining blind spots in the real-time image, and storing the position of each blind spot in a blind spot list;

[0039] Step S54, inputting the enhanced real-time image into the YOLOV5s model to generate a real-time target candidate frame set with a confidence level above 0.01;

[0040] Step S55, input the real-time target candidate frame set into the Deepsort model, obtain the filtered real-time target candidate frame set, and remove the target candidate frame containing the blind spot list in the filtered real-time target candidate frame set to generate the final target detection frame to achieve real-time tracking of small infrared targets.

[0041] Furthermore, the method for performing contrast enhancement processing on the collected real-time image in step S52 is: obtaining the histogram distribution of each frame of the image, and changing the histogram distribution of each frame of the image into an approximately uniform distribution histogram.

[0042] The present invention combines the YOLOv5 framework for target detection and the Deepsort framework for target tracking, and designs data cleaning and enhancement preprocessing and background removal postprocessing for the complex background of small infrared targets, which can overcome interference points and blind points in the background, so that small infrared targets in complex backgrounds can be accurately and quickly detected, and the robustness and detection rate can be improved. In addition, the present invention can achieve real-time tracking effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a flow chart of the statistical filtering infrared small target detection and tracking method based on YOLOv5 and Deepsort according to the present invention.

[0044] Figure 2 It is a schematic diagram of the label information of the marked infrared small target.

[0045] Figure 3 yes Figure 1 Flow chart of step S4 in FIG.

[0046] Figure 4(a)-Figure 4(h) The above performance indicator graphs are shown with a confidence threshold of 0.001.

[0047] Figure 5(a) is a PR curve diagram on the test set of the recognition detection model of the present invention; Figure 5(b) is a diagram showing the change of F1 score under different confidence thresholds; Figure 5(c) is a diagram showing the change of precision rate under different confidence thresholds; Figure 5(d) is a diagram showing the change of recall rate under different confidence thresholds.

[0048] Figure 6(a)-Figure 6(c) This is the final tracking prediction effect diagram. DETAILED DESCRIPTION

[0049] The preferred embodiments of the present invention are given below in conjunction with the accompanying drawings and described in detail.

[0050] like Figure 1 As shown, the present invention provides a statistical filtering infrared small target detection and tracking method based on YOLOv5 and Deepsort, comprising the following steps:

[0051] Step S1, using infrared imaging equipment to collect infrared small target images under complex backgrounds, and obtaining infrared small target image data sets. The complex background refers to backgrounds such as trees, sky with few clouds, buildings, continuously changing weather, complex clouds, sea surface and sea-sky.

[0052] Specifically, step S1 includes:

[0053] Step S11, use an infrared imaging device to perform continuous frame shooting sampling on one or more aircraft under different complex backgrounds to obtain an initial image set. The sampling frame rate can be 50ms, 100ms or other suitable frame rates. During the shooting process of the infrared imaging device, due to the influence of sensor delay, noise interference, etc., blurred images or missing images will be collected. For these blurred images or missing images, resampling processing is required to obtain a clear and complete initial image set.

[0054] Step S12, marking a real target detection frame in each image of the initial image set, constructing label information of the infrared small target, and the initial image set and the label information of the infrared small target together constitute an infrared small target image data set.

[0055] The method of marking the real target detection frame is manual marking, for example, using a labeling tool to obtain the coordinates (x, y) of the upper left corner of the target detection frame, the length h of the target detection frame, and the width w of the target detection frame, such as Figure 2 The x, y, h and w marked above constitute the label information of the infrared small target.

[0056] Since manual labeling may have human error factors, such as serious deviation from the real target or slight errors due to operational errors. Therefore, in order to minimize the error of manual behavior, enhance the information strength of the data set, and improve the accuracy of model training, it is necessary to use visual operations to determine whether the labeling is correct. The judgment process is as follows: convert the XML file format generated by labeling into a txt file format, and store the data information in the txt file in a list; use the opencv tool to read the image information in the initial image set; compare the data information stored in the list with the read image information to determine whether the target detection frame overlaps well with the area where the infrared small target is located, so as to obtain the accuracy of the label information. When the overlapping area of ​​the target detection frame and the area where the infrared small target is located occupies more than 90% of the total area of ​​the two areas, the target detection frame overlaps well with the area where the infrared small target is located, and the proportion of the overlapping area in the total area is the accuracy of the label information. If the accuracy of the label information is less than 90%, it is necessary to manually reset the label information and recalculate the accuracy.

[0057] Step S2, preprocessing the infrared small target image dataset, obtaining the preprocessed image dataset, and dividing the preprocessed image dataset into a training set and a verification set.

[0058] The methods for preprocessing infrared small target image datasets include:

[0059] Step S21, performing data cleaning on the infrared small target image data set to obtain a cleaned image data set.

[0060] Data cleaning refers to the use of statistical weighted average or machine learning methods to process default values ​​and abnormal values. During the image acquisition process, some images may be abnormal or missing due to noise interference, so cleaning is required. In this embodiment, the statistical weighted average method is used for data cleaning: detect whether there is a missing frame or an abnormal frame in the infrared small target image data set. If there is a missing frame, the previous and next frames of the frame are used to fill the average; if there is an abnormal frame (that is, there is a discontinuous association with the previous and next two frames), the frame is deleted and filled with the previous and next two frames.

[0061] Step S22, randomly selecting a number of frames of images from the cleaned image data set, performing data enhancement on the selected number of frames of images, and obtaining an enhanced image data set.

[0062] Data enhancement refers to randomly flipping, scaling, cutting, rotating, pasting, and piecing together the original data to expand additional data sets. In this embodiment, the paste enhancement method or the Mosaic enhancement method is used to perform data enhancement on several selected frames of images. The paste enhancement method is: after scaling the image by a large scale of 0.1 to 2 times, the infrared small targets in one frame of the image are randomly pasted to another frame of the image, so as to increase the number of infrared small targets and enhance the detection from the perspective of the number of infrared small targets. The Mosaic enhancement method is: randomly using 4 frames of pictures, splicing them in a random scaling, cropping, and sorting manner, so that the distribution of infrared small targets is more uniform, and there are multiple infrared small targets in one frame of the image, so as to enhance the detection from the perspective of the distribution and number of infrared small targets.

[0063] Step S23, extracting the background image of each frame from the cleaned image data set, and performing a mixing process on the extracted background image to obtain a mixed background image set. Obtaining a mixed background image set can enhance the model's recognition ability for complex backgrounds and reduce the false detection rate of treating background objects as targets. The method for extracting the background image is: using an image expansion and corrosion method, after corroding the target object and noise points, the retained image is the background image.

[0064] Step S24, extract the noise points contained in each frame of the infrared small target image data set, and randomly paste the extracted noise points into the above-mentioned cleaned image data set, enhanced image data set and mixed background image set to expand the data distribution of the noise points. The noise points are manually cut and extracted.

[0065] Step S25, merging the cleaned image dataset, the enhanced image dataset and the mixed background image dataset with randomly pasted noise in step S24 to obtain a final preprocessed image dataset.

[0066] After obtaining the preprocessed image dataset, the preprocessed image dataset is divided into a training set and a validation set in a ratio of 8:2 by random sampling. That is, 0.8 is the training set and 0.2 is the validation set.

[0067] is the validation set.

[0068] Step S3, training the YOLOv5s model in the YOLOv5 algorithm according to the training set and the validation set divided in step S2, and obtaining the training set of the Deepsort model. The present invention adopts the YOLOv5s model in order to achieve a faster detection effect.

[0069] The YOLOv5s model can be regarded as a two-stage target detection algorithm, including a baseline network layer based on ResNet, a Neck network layer with an FPN+PAN structure for outputting target detection results, an output layer, and an output end after non-maximum suppression processing. Among them, the baseline network layer outputs a feature mapping matrix. The Neck network layer adopts the FPN+PAN structure to improve the diversity and robustness of features and enhance the fusion ability of network features. Among them, FPN represents the feature pyramid network, which uses top-down upsampling to extract strong semantic features of the image (i.e., the shape of the target object); PAN represents the pixel aggregation network, which uses a bottom-up network to extract strong positioning features of the image (i.e., the position of the target object). The fusion of FPN and PAN can achieve the aggregation of shape and position features. The output layer uses GIoU_Loss as the loss function of the Bounding box and outputs the target detection result. After obtaining the target detection result, post-processing is performed, and non-maximum suppression is used to eliminate multiple boxes on the same target and the output bounding boxes stacked together.

[0070] Specifically, step S3 includes:

[0071] Step S31, initialize YOLOv5s model parameters, including: batch size, number of iterations (epoch), image resolution, intersection over union (IOU) threshold, and confidence threshold.

[0072] Due to the large image set, all images cannot be directly input together during model training, otherwise it will cause dimensionality explosion. Therefore, a batch method is adopted to output 64 frames of images each time. The YOLOv5s model network includes two processes: forward propagation and back propagation. When all images have undergone forward and back propagation once, it is called iteration 1. The iteration is to perform gradient optimization and find the weight when the loss function converges. In this embodiment, batchsize is set to 64, epoch is set to 2000, image resolution is set to 640*640, intersection-over-combination threshold is set to 0.4, and confidence threshold is set to 0.6.

[0073] Step S32, input the training set divided in step S2 into the YOLOv5s model, and output the predicted target detection frame after each iteration according to the initialized YOLOv5s model parameters. In the process of outputting the target detection frame at the output end of the YOLOv5s model, the detected target detection frame needs to be filtered according to the confidence, and the target detection frame with a confidence less than the set confidence threshold of 0.6 is deleted, and the retained target detection frame is the predicted target detection frame.

[0074] In step S33, the predicted target detection box after each iteration is calculated with the real target detection box in the label information for IOU loss to obtain the learning weight after each iteration.

[0075] Step S34, for the learning weights after each iteration, according to the verification set divided in step S2, select the learning weight with the smallest test error (i.e., the highest accuracy) on the verification set, and use the learning weight corresponding to the smallest test error as the weight parameter of the YOLOv5s model. The method for obtaining the test error is specifically as follows: input the verification set into the YOLOv5s model with the learning weight obtained in the current iteration as the weight parameter, obtain the prediction box including the annotation information and the confidence indicating whether the target object exists, and calculate the loss of the cross entropy of the IOU and the classification to obtain the test error.

[0076] Step S35, according to the weight parameters of the YOLOv5s model, target object detection is performed to obtain the target recognition candidate frame in each frame image, and the target recognition candidate frame in each frame image constitutes the training set of the Deepsort model. Specifically, the training set divided by step S2 is re-input into the YOLOv5s model with the weight parameters determined, and the target recognition candidate frame is obtained according to the set confidence threshold. For example, if there are two targets, the corresponding real target frame should also be two, but because the image data set is enhanced as described above, the anchor frame output by the YOLOv5s model may have 10, 15 or other numbers, and these anchor frames with a large number are the target recognition candidate frames. However, the number of these anchor frames is greater than 2, indicating that there is a false detection, and there may also be a missed detection, so these anchor frames need to be used as the training set of the Deepsort model and input into the Deepsort model for the next step of processing.

[0077] That is, step S4 is performed, the training set of the Deepsort model is input into the Deepsort model for training, the weight parameters of the Deepsort model are obtained, and an identifier for infrared small target detection and tracking is constructed.

[0078] It should be noted that before inputting the training set of the Deepsort model into the Deepsort model for training, the network structure of the Deepsort model needs to be modified to a network structure that can process small target images with a pixel width of about 1 to 20, or the target recognition candidate box needs to be preprocessed to a size that the Deepsort model can correctly identify.

[0079] The purpose of using the Deepsort model for training is to filter out false positives and supplement missed positives. Figure 3 As shown, step S4 includes:

[0080] Step S41, based on the training set of the Deepsort model, determine the real position of the small infrared target in each frame of the image, and based on the real position of the small infrared target in the k-1 (k ≥ 2) frame of the image, use Kalman filtering to determine the predicted position of the small infrared target in the k frame of the image. It should be noted that the real position referred to in this step is the position framed by the target recognition candidate box in each frame of the image, which is different from the position framed by the real target detection box in step S12. The real position is a determined value, which is used to distinguish between the determined value and the predicted value when the Kalman filter is used later.

[0081] Kalman filtering is used to predict trajectories and filter out false positives based on the predicted trajectory information. For example, if the predicted state information at the next moment is [2, 3, 4, 5], then the detection information of [10, 11, 12, 13] in the candidate box is a false positive, and this detection box needs to be filtered out. In other words, a rough filter is first performed based on the confidence threshold, and then a fine filter is performed based on the Kalman filter.

[0082] Kalman filtering mainly includes eight system states of the anchor frame: X = (x, y,, r,,, r,), x and y represent the horizontal and vertical coordinates of the center point of the anchor frame in the image, h represents the height of the anchor frame, and r represents the ratio of the height to the width of the anchor frame; x, y, r, are the derivatives of x, y,, in the current state, respectively, representing the state information of the target object in the previous frame. The Kalman filter algorithm predicts the state information at the next moment based on the current state information. The Kalman filter formula is:

[0083] x'=Ax

[0084] P'=APA T +Q

[0085] y=z-Hx'

[0086] K=P'H T (HP'H T +R) -1

[0087] P=(I-KH)P'

[0088] Where A is the system matrix, P is the covariance matrix, which is used to measure the error between the predicted value and the true value; Q is the process noise when the state is transferred; K is the gain matrix, which is used to minimize P; R is the measurement noise when the state is measured, which is generally Gaussian white noise; I is the unit matrix, and H is the measurement matrix.

[0089] Step S42, using the Hungarian algorithm to perform cascade matching on the predicted position of the infrared small target in the k-th frame image and the actual position of the infrared small target in the k-th frame image, to obtain the result of the initial successful matching, the trajectory that was not matched for the first time, and the detection frame that was not matched for the first time. Specifically, the cost matrix based on the cosine distance is calculated to delete the detection frame with a cosine matrix that is too large, and the Mahalanobis distance between the Kalman predicted trajectory and the actual detection frame is used as the cost matrix by the Hungarian algorithm to match the trajectory and the detection frame, and return the matching result. It should be noted that due to the delay of the Hungarian algorithm itself, and in order to overcome interference, the cascade matching of step S42 is performed between 30 frames of cyclic detection. If the match is still not successful after 30 frames, the match is abandoned.

[0090] Step S43, perform IOU matching on the trajectory that was not matched for the first time in step S42 and the detection frame that was not matched for the first time, and obtain the result of successful re-matching, the trajectory that was not matched again, and the detection frame that was not matched again. Specifically, for the trajectory that only matches one frame, calculate the IOU distance between it and the unmatched detection frame, delete the detection frame with too large IOU distance, and use the Hungarian algorithm to use the IOU distance between the Kalman predicted trajectory and the actual detection frame as the cost matrix to match the trajectory and the detection frame, and return the matching result. It should be noted that the IOU matching in step S43 is performed between each frame of the image, that is, if the match is not successful at the current moment due to interference or other reasons, the content of the current frame will not be matched in the next frame.

[0091] Step S44, based on the result of the first successful match and the result of the second successful match, update the parameters of the Kalman filter. Specifically, the covariance and mean of the Kalman filter parameters are updated using the above Kalman filter formula, where P is the covariance, is the predicted value of the trajectory, and its expectation is the mean.

[0092] Step S45, assign a new track and a new ID to the detection frame that is not matched again, and ReID Extract the feature set of the target object in the detection frame; at the same time, determine whether the trajectory that has not been matched again is in a determined state, retain the trajectory that is in a determined state and has a mismatch number of less than 30 frames, and repeat steps S41-S44. Specifically, if the detection frame does not match the existing tracking trajectory, the tracking trajectory of the detection frame is initialized (initialized using the cosine similarity measure matrix, with a default value of 0). The state of the new trajectory initialization is an undetermined state. Only when three consecutive frames are successfully matched can the undetermined state be converted to a determined state. If the trajectory in the determined state must mismatch the target object for more than 30 consecutive times, it will be deleted.

[0093] The Deepsort model uses the Hungarian algorithm to match moving targets, and uses a deep recognition network to identify target identifiers, tracks targets based on target identifiers, and stores target identifier information. When a target is lost for a period of time and then reappears, it can be re-identified based on the target identifier. The Hungarian algorithm is a method for solving the association between detection results and tracking prediction results. It uses the Mahalanobis distance between the Kalman prediction results of the motion state of an existing moving target and the detection results to associate running information.

[0094] The Mahalanobis distance formula is as follows:

[0095] d (1) (i,j)=(d i -y i ) T Si -1 (d i -y i )

[0096] d i represents the i-th detection position, y i represents the predicted position of the target by the ith tracker, S i represents the covariance matrix between the detected position and the average tracking position;

[0097] The Mahalanobis distance takes into account the uncertainty of the state measurement by calculating the standard deviation between the detected position and the average tracking position. If the Mahalanobis distance of a certain association is less than the specified threshold t (1) , then the motion state association is set successfully, and the function used is:

[0098]

[0099] When it is 1, it indicates that the association is successful. (1) Take 9.4877.

[0100] Step S5, using the infrared small target detection and tracking identifier, real-time detection and tracking of the infrared small target.

[0101] Specifically, step S5 includes:

[0102] Step S51, set the parameters in the recognizer for infrared small target detection and tracking, and load the weight parameters of the YOLOv5s model into the YOLOv5s model, and load the weight parameters of the Deepsort model into the Deepsort model, to obtain a serial model of the YOLOv5s model and the Deepsort model. The confidence threshold of the YOLOv5s model is set to 0.01, and the intersection-over-combination ratio is set to 0.4.

[0103] Step S52, using an infrared imaging device to collect a real-time image of the small infrared target, and performing contrast enhancement processing on the collected real-time image to obtain an enhanced real-time image to reduce the difficulty of feature recognition of the small infrared target.

[0104] The method for performing contrast enhancement processing on the collected real-time image is: obtaining the histogram distribution of each frame of the image, and changing the histogram distribution of each frame of the image into an approximately uniform distribution histogram to enhance the contrast of the image.

[0105] The method of changing the histogram distribution into an approximately uniform distribution histogram is to perform a monotone nonlinear mapping transformation f on each pixel of the original image, that is, D B =f(D A ), and keep the total number of pixels unchanged, as follows: Among them, H A (D) is the histogram distribution of the original image, H B (D) is uniform distribution, take A0 is the number of pixels, L is the grayscale depth, which is 256.

[0106] Step S53, average pixel statistics are performed on the first 100 frames of the enhanced real-time image to obtain blind spots in the real-time image, and the position of each blind spot is stored in a blind spot list. Specifically, the blind spot is obtained by averaging the grayscale of the pixels at the same point in the first 100 frames of the image. Since the average value of the blind spot pixels is much higher than the average value of the normal points, 0.7 to 0.9 times the average value of the blind spot pixels is intercepted as a threshold, and pixels exceeding the threshold are regarded as blind spots.

[0107] Step S54: input the enhanced real-time image into the YOLOV5s model to generate a real-time target candidate frame set with a confidence level above 0.01.

[0108] Step S55, input the real-time target candidate frame set into the Deepsort model, obtain the filtered real-time target candidate frame set, and remove the target candidate frame containing the blind spot list in the filtered real-time target candidate frame set to generate the final target detection frame to achieve real-time tracking of small infrared targets.

[0109] Real-time detection and tracking of small infrared targets can be performed, and the precision, recall, F1 score, mAP, etc. can be calculated on the verification set. Figure 4(a)-Figure 4(h) The above performance indicators are shown in the figure with a confidence threshold of 0.001. It can be seen from the figure that under the given threshold of 0.001, the precision rate can be roughly 0.94 and the recall rate can be 0.75. Figure 4(a) shows the error between the labeled box and the predicted box. The reason for the jump is that the training was interrupted and the label was adjusted. The low recall rate is because some parts of the training set given by the picture are blocked by the target background or submerged by the background light. These situations need to be processed by the prediction algorithm and the separation algorithm of the target background and the target.

[0110] FIG5(a) is a PR curve diagram of the recognition detection model test set of the present invention. The PR curve can intuitively reflect the performance of the model. In the PR curve, the precision is the ordinate and the recall is the abscissa. The larger the area under the PR curve, the better the performance of the model. From the figure, the performance of the model is good.

[0111] Figure 5(b) is a graph showing the change in F1 score under different confidence thresholds. Since both precision and recall reflect the quality of the model, but both quantities need to be measured at the same time, the F1 score curve is used, where the F1 score is the harmonic mean of precision and recall. When the confidence threshold is small, the F1 performance index increases as the confidence threshold increases; then, as the confidence threshold increases, the F1 performance index decreases.

[0112] Figure 5(c) is a graph showing the change in accuracy at different confidence thresholds. When the confidence level is 0.1, the prediction accuracy basically converges to 99%.

[0113] Figure 5(d) shows the corresponding recall rate changes under different confidence thresholds. As the confidence increases, the recall rate continues to decrease.

[0114] The final tracking prediction effect is as follows Figure 6(a)-Figure 6(c) As shown in the figure, there are single target and multiple target cases respectively. The number on the anchor box is the ID number of the target object, which is used for tracking and recording.

[0115] In addition, the confidence threshold changes are shown in Table 1:

[0116] Confidence Threshold Accuracy Recall mAP@0.5 mAP@[0.5:0.95] 0.001 0.937 0.756 0.786 0.635 0.01 0.937 0.756 0.771 0.625 0.1 0.998 0.5 0.505 0.434

[0117] The present invention combines the YOLOv5s model for target detection and the Deepsort model for target tracking, and designs preprocessing such as data cleaning and enhancement and post-processing such as background removal for the complex background in the infrared small target scene, which can accurately and quickly detect and track infrared small targets in complex backgrounds.

[0118] The above is only a preferred embodiment of the present invention, and is not intended to limit the scope of the present invention. The above embodiment of the present invention can also be modified in various ways. That is, all simple, equivalent changes and modifications made according to the claims and the description of the present invention fall within the scope of protection of the claims of the present invention. The contents not described in detail in the present invention are all conventional technical contents.

Claims

1. A statistical filtering infrared small target detection and tracking method based on YOLOv5 and Deepsort, characterized in that: include: Step S1, using an infrared imaging device to collect infrared small target images under a complex background to obtain an infrared small target image data set; Step S2, preprocessing the infrared small target image dataset, obtaining the preprocessed image dataset, and dividing the preprocessed image dataset into a training set and a verification set; wherein the method for preprocessing the infrared small target image dataset includes: Step S21, performing data cleaning on the infrared small target image data set to obtain a cleaned image data set; Step S22, randomly selecting a number of frames of images from the cleaned image data set, and performing data enhancement on the selected number of frames of images to obtain an enhanced image data set; Step S23, extracting a background image of each frame from the cleaned image data set, and performing a mixing process on the extracted background image to obtain a mixed background image set; Step S24, extracting noise points contained in each frame of the image from the infrared small target image dataset, and randomly pasting the extracted noise points into the cleaned image dataset, the enhanced image dataset and the mixed background image dataset; Step S25, merging the cleaned image dataset, the enhanced image dataset and the mixed background image dataset with the randomly pasted noise in step S24 to obtain a final preprocessed image dataset; Step S3, training the YOLOv5s model in the YOLOv5 algorithm according to the training set and the verification set to obtain a training set of the Deepsort model; including: Step S31, initializing YOLOv5s model parameters, including: batch processing size, number of iterations, image resolution, intersection-over-union ratio threshold, and confidence threshold; Step S32, inputting the training set into the YOLOv5s model, and outputting the predicted target detection box after each iteration according to the initialized YOLOv5s model parameters; Step S33, performing IOU loss calculation on the predicted target detection frame after each iteration and the real target detection frame in the label information to obtain the learning weight after each iteration; Step S34, for the learning weight after each iteration, according to the verification set, select the learning weight with the smallest test error on the verification set, and use the learning weight with the smallest test error as the weight parameter of the YOLOv5s model; Step S35, obtaining the target recognition candidate box in each frame image according to the weight parameters of the YOLOv5s model, and the target recognition candidate box in each frame image constitutes the training set of the Deepsort model; Step S4, inputting the training set of the Deepsort model into the Deepsort model for training, obtaining the weight parameters of the Deepsort model, and constructing an infrared small target detection and tracking recognizer; including: Step S41, determining the real position of the infrared small target in each frame of the image according to the training set of the Deepsort model, and using Kalman filtering to determine the predicted position of the infrared small target in the k-th frame of the image according to the real position of the infrared small target in the k-1th frame of the image, where k≥2; Step S42, using the Hungarian algorithm to perform cascade matching on the predicted position of the infrared small target in the k-th frame image and the actual position of the infrared small target in the k-th frame image, to obtain the result of the initial successful matching, the trajectory of the initial unmatched and the detection frame of the initial unmatched; Step S43, performing IOU matching on the initially unmatched trajectory and the initially unmatched detection frame in step S42, to obtain a rematched successful result, a rematched trajectory, and a rematched detection frame; Step S44, updating the parameters of the Kalman filter according to the result of the first successful matching and the result of the second successful matching; Step S45, assign a new track and a new ID to the detection frame that is not matched again, and ReID Extract the feature set of the target object in the detection frame; at the same time, determine whether the track that is not matched again is in a determined state, retain the track that is in a determined state and has a mismatch number of less than 30 frames, and repeat steps S41 to S44; Step S5, using the infrared small target detection and tracking identifier to perform real-time detection and tracking of the infrared small target.

2. The statistical filtering infrared small target detection and tracking method based on YOLOv5 and Deepsort according to claim 1, characterized in that: The step S1 comprises: Step S11, using an infrared imaging device to perform continuous frame shooting sampling on one or more aircraft under different complex backgrounds to obtain an initial image set; Step S12, marking a real target detection frame in each image of the initial image set, constructing label information of the infrared small target, and the initial image set and the label information of the infrared small target together constitute an infrared small target image data set.

3. The statistical filtering infrared small target detection and tracking method based on YOLOv5 and Deepsort according to claim 1, characterized in that: The data enhancement method in step 22 is a pasting enhancement method or a mosaic enhancement method.

4. The statistical filtering infrared small target detection and tracking method based on YOLOv5 and Deepsort according to claim 1, characterized in that: The batch size is set to 64, the number of iterations is set to 2000, the image resolution is set to 640*640, the intersection-over-union ratio threshold is set to 0.4, and the confidence threshold is set to 0.

6.

5. The statistical filtering infrared small target detection and tracking method based on YOLOv5 and Deepsort according to claim 1, characterized in that: The step S5 comprises: Step S51, setting parameters in the recognizer for infrared small target detection and tracking, and loading the weight parameters of the YOLOv5s model into the YOLOv5s model, and loading the weight parameters of the Deepsort model into the Deepsort model; Step S52, using an infrared imaging device to collect a real-time image of the small infrared target, and performing contrast enhancement processing on the collected real-time image to obtain an enhanced real-time image; Step S53, performing pixel average statistics on the first 100 frames of the enhanced real-time image, obtaining blind spots in the real-time image, and storing the position of each blind spot in a blind spot list; Step S54, inputting the enhanced real-time image into the YOLOV5s model to generate a real-time target candidate frame set with a confidence level above 0.01; Step S55, input the real-time target candidate frame set into the Deepsort model, obtain the filtered real-time target candidate frame set, and remove the target candidate frame containing the blind spot list in the filtered real-time target candidate frame set to generate the final target detection frame to achieve real-time tracking of small infrared targets.

6. The statistical filtering infrared small target detection and tracking method based on YOLOv5 and Deepsort according to claim 1, characterized in that: The method for performing contrast enhancement processing on the collected real-time image in step S52 is: obtaining the histogram distribution of each frame of the image, and changing the histogram distribution of each frame of the image into an approximately uniform distribution histogram.

Citation Information

Patent Citations

  • Ship multi-target tracking method based on YOLO V5 algorithm

    CN113269073A

  • Fingertip tracking method based on deep learning and K-curvature method

    CN113608663A