Video stream target tracking method, device and storage medium for unmanned aerial vehicle

By combining detection and tracking algorithms, using algorithms such as mobilenet-SSD and MOSSE or KCF, accurate tracking of drone video stream targets in a long and complex environment is achieved, solving the problems of limited computing resources and loss of target offsets, and improving the stability and accuracy of tracking.

CN115272393BActive Publication Date: 2025-05-20THE 54TH RESEARCH INSTITUTE OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210917866.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-01
Publication Date
2025-05-20
Estimated Expiration
2042-08-01

AI Technical Summary

Technical Problem

The prior art is difficult to achieve accurate tracking of drone video streaming targets under limited computing resources, especially how to rematch and track targets when targets are offset or lost.

Method used

Using an algorithm combining detection and tracking, target detection is performed through mobilenet-SSD, target position information is obtained and input into MOSSE or KCF tracker for initialization and tracking. At the same time, the strategies of IOU matching and deep feature matching are used to update the location information of the tracking target, and the stability and accuracy of tracking are improved through the feature sample library and self-feedback mechanism.

Benefits of technology

Real-time detection and tracking of drone video stream targets under limited computing resources is realized, target tracking accuracy and stability in complex environments is improved, target offsets and losses can be effectively handled, and average frame rate and tracking effect are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115272393B_ABST
    Figure CN115272393B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and storage medium for tracking a target in a video stream of an unmanned aerial vehicle, and relates to the technical field of target tracking in a video stream. The method of the present invention comprises: performing target detection on a real-time video stream of an unmanned aerial vehicle to obtain a detection target, position information and a detection target ID; selecting a detection target corresponding to the detection target ID as a tracking target; inputting the position information of the tracking target into a tracker for initialization; and tracking the target in a subsequent real-time video stream of the unmanned aerial vehicle through the tracker. The method of the present invention is simple and fast, and has a good tracking effect in a short time, and can effectively reduce the computational overhead and improve the average frame rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object tracking in video streams, and particularly to a method, device, and storage medium for object tracking in video streams for unmanned aerial vehicles (UAVs). Background Art

[0002] For the problem of object tracking in a video stream, it can generally be solved in two ways: One is to use a traditional filtering-based single-object tracker, which predicts the position where the object will appear in the next frame based on the features and position information of the object in the previous frame and the previous frames. This method is also the general single-object tracking algorithm, which is suitable for short-term tasks with relatively simple tracking. Its disadvantage is that there are not enough samples to fully learn the features of the object, so errors will gradually accumulate in continuous prediction, ultimately resulting in deviation and failure. At the same time, this kind of tracking failure generally cannot be corrected by the tracker itself; Another idea is to use a detector, which decomposes the video stream into frame-by-frame pictures and relatively independently uses detection algorithms on them to confirm the position where the object appears in the whole picture. This method can generally track the object accurately for a relatively long time, but its disadvantage is that the detector needs to be offline trained in advance, so it can only be used to track things that are known in advance and is generally unable to effectively distinguish similar objects due to reasons such as overfitting. At the same time, the tracking stability is poor and the speed is not high.

[0003] For the tracking process of a UAV for ground targets, there are often various visual environment factors interfering, and the tracking time is also relatively long. Generally speaking, it is a relatively complex tracking task. The method of frame-by-frame detection is often difficult to be deployed on devices with limited computing resources. Due to time and environment reasons, traditional tracking algorithms are also difficult to stably track the target. This technical solution is proposed in view of such technical problems and limitations. In actual engineering practice, a combined algorithm of detection and tracking can be adopted to jointly solve a long-term and complex object tracking task with complex tasks and environments. In such tasks, the most critical problems are: how to construct a feasible real-time detection and tracking scheme under the limitation of limited computing resources; how to accurately and stably track the specific target to be tracked among multiple potential targets, that is, the problem of object matching; how to re-match and track the target after a large deviation occurs during long-term tracking or when the target leaves the field of view and re-enters the field of view. Summary of the Invention

[0004] In view of this, the present invention provides a method, device, and storage medium for object tracking in video streams for UAVs, which can achieve feasible real-time object detection and tracking under the existing limitation of limited computing resources.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A video stream target tracking method for an unmanned aerial vehicle, comprising the following steps:

[0007] S1: Perform target detection on the image frames of the real-time video stream of the unmanned aerial vehicle to obtain the detected targets, the position information of the detected targets, and the target IDs;

[0008] S2: Select at least one detected target corresponding to a target ID as the tracking target;

[0009] S3: Input the position information of the tracking target into a tracker for initialization;

[0010] S4: Track the tracking target in the subsequent image frames of the real-time video stream of the unmanned aerial vehicle through the tracker.

[0011] Furthermore, step S1 further includes: displaying the target box and target ID of the detected target on the image frame;

[0012] The specific manner of step S2 is: select the target ID through an interactive method at the ground station, and use the detected target corresponding to the selected target ID as the tracking target.

[0013] Furthermore, after step S4, determine whether the number of tracked image frames has reached a preset number of frames. If so, stop tracking and perform the following steps:

[0014] S5: Re-perform target detection in the current image frame to obtain the position information of each detected target;

[0015] S6: Perform IOU matching on the position information of the re-detected targets and the position information of the previous tracking targets, and use the position information of the target with an intersection over union ratio greater than the IOU threshold and the largest intersection over union ratio as the updated position information of the tracking target;

[0016] S7: Repeat steps S3 to S6.

[0017] Furthermore, step S5 further includes: if the re-target detection fails, continue to use the tracker to track the previous tracking target, and perform step S5 on the next image frame.

[0018] Furthermore, step S6 further includes: if there is no position information of a target with an intersection over union ratio greater than the IOU threshold and the largest intersection over union ratio, continue to use the tracker to track the previous tracking target, and perform steps S5 and S6 on the next image frame.

[0019] Furthermore, in step S7, while performing step S3, it further includes: extracting the depth feature information of the picture corresponding to the target position in the image frame and storing it in the feature sample library;

[0020] The step S6 executed through step S7 further includes: determining whether the number of frames corresponding to the depth feature information in the feature sample library is greater than a threshold. If so, performing IOU matching on the newly detected target position information and the position information of the previously tracked target, and calculating the cosine distance matching between the depth feature information corresponding to the newly detected target position information and the depth feature information in the feature sample library to obtain a comprehensive metric C. Taking the target position information with the comprehensive metric C greater than a preset value and the largest intersection over union as the position information of the updated tracked target;

[0021] The calculation formula of the comprehensive metric C is: C = λd (1) +(1 - λ)d (2)

[0022] where d (1) is the degree of IOU matching, and d (2) is the degree of matching of depth features, and λ is the weight coefficient of the degree of IOU matching and the degree of matching of depth features.

[0023] Furthermore, the weight coefficient λ of the degree of IOU matching and the degree of matching of depth features decreases as the sample size of the depth feature information in the feature sample library increases, and λ satisfies:

[0024] λ = 1 - 0.02t

[0025] where t is the sample size of the depth feature information.

[0026] Furthermore, a preset number of invariant initial samples are stored in the feature sample library, and the sample number of the feature sample library has an upper limit. When the upper limit is exceeded, the newly successfully matched depth feature information sample replaces the depth feature information sample of the old non-initial sample.

[0027] A video stream target tracking device for a drone, which includes:

[0028] A target preliminary detection module, which is used to perform target detection on each frame of image in the real-time video stream of the drone by using a detection model of mobilenet-SSD to obtain the detection target of each frame of image, as well as the position information and target ID of the detection target;

[0029] A target determination module, which is used to select at least one detection target corresponding to a target ID as the tracked target;

[0030] An initialization module, which is used to input the position information of the tracked target into a MOSSE or KCF tracker for initialization;

[0031] A target tracking module, which is used to track the tracked target in the subsequent image frames of the real-time video stream of the drone through a MOSSE or KCF tracker.

[0032] A computer storage medium stores instructions, and when the instructions run, they execute a video stream target tracking method for an unmanned aerial vehicle as described in any one of the above.

[0033] The present invention has the following beneficial effects:

[0034] 1. Both the detection of the convolutional network and the extraction of deep features are parts with relatively large computational overheads. However, the correlation filtering-based tracking algorithms such as KCF and MOSSE adopted by the present invention are relatively simple and fast, and have good tracking effects in a short time. In addition, the strategy of multi-frame tracking difference frame detection adopted by the present invention can effectively reduce the computational overhead and improve the average frame rate.

[0035] 2. The accuracy of the position for target detection is generally higher than that for target tracking. The present invention updates the tracker through difference frame detection, which is less likely to deviate than simply using a tracking algorithm and is more suitable for long-term tracking.

[0036] 3. Usually, the tracking algorithm only searches for the region most similar to the previous one in the area around the target. Therefore, when the target is lost at a certain moment, it will be difficult to find a suitable target for tracking in the nearby area in the prediction of subsequent frames. The present invention combines detection and deep feature matching. Each detection will retrieve the optimal region globally again. Since the deep feature matching that focuses on the essential semantic information of the target accounts for a larger proportion than the IOU matching that focuses on position information, even if the target is lost for a short time, the original target can be retrieved more effectively when the target reappears subsequently.

[0037] 4. The present invention sets up self-feedback, making the tracking algorithm more inclined to believe the results of IOU matching when the feature samples are insufficient, preventing errors caused by insufficient samples in the early stage; and more inclined to the results of deep feature matching after the sample quantity is sufficient, making the tracking more stable.

[0038] 5. The present invention designs a sample update strategy, retaining an initial preset number of sample features, so that the sample library will not be completely contaminated due to subsequent tracking errors; and at the same time, updating the sample library even if, so that the sample library can always match the latest state of the target object. Description of the Drawings

[0039] Figure 1 is the flow chart of a video stream target tracking method for an unmanned aerial vehicle in Embodiment 1 of the present invention Figure 1 ;

[0040] Figure 2 is the flow chart of a video stream target tracking method for an unmanned aerial vehicle in Embodiment 1 of the present invention Figure 2 ;

[0041] Figure 3It is the flowchart of a video stream target tracking method for an unmanned aerial vehicle in the second embodiment of the present invention. Figure 1 ;

[0042] Figure 4 It is the flowchart of a video stream target tracking method for an unmanned aerial vehicle in the second embodiment of the present invention. Figure 2 ;

[0043] Figure 5 It is the structural block diagram of a video stream target tracking device for an unmanned aerial vehicle in the third embodiment of the present invention. Specific embodiments

[0044] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only the parts related to the present invention rather than all the structures are shown in the drawings.

[0045] Figure 1 It is the flowchart of a video stream target tracking method for an unmanned aerial vehicle provided in the first embodiment of the present invention. This embodiment is applicable to the tracking process of ground targets by an unmanned aerial vehicle. This method can be implemented by a computer device or a processor on the unmanned aerial vehicle and is used for the video stream target tracking method of the unmanned aerial vehicle. Specifically, it includes the following steps:

[0046] Step S1: Use the detection model of mobilenet-SSD to perform target detection on each frame of the real-time video stream of the unmanned aerial vehicle to obtain the detection target, the position information of the detection target, and the detection target ID of each frame of the image.

[0047] Among them, the real-time video stream of the unmanned aerial vehicle can be obtained in real time through a camera on the unmanned aerial vehicle.

[0048] Among them, the Mobilenet-SSD model uses the Mobilenet network as the backbone of the target detection network. It is a type of SSD target detection model and has the characteristic of fast running speed of the original SSD model. Compared with the original SSD model, Mobilenet-SSD greatly reduces the number of model parameters and the amount of calculation.

[0049] Mobilenet is mainly a lightweight deep network model proposed for use on mobile devices. It mainly uses depthwise separable convolution to decompose the standard convolution kernel for calculation, reducing the amount of calculation.

[0050] Two hyperparameters are introduced to reduce the number of parameters and the computational cost: Width Multiplier: reduces the number of input and output channels; Resolution Multiplier: reduces the size of the input and output feature maps.

[0051] Depthwise separable convolution divides a standard convolution kernel into a depthwise convolution kernel and a 1×1 pointwise convolution kernel. Assuming the input is a feature map with M channels, the convolution kernel size is DK×DK, and the number of output channels is N, then the number of parameters of the standard convolution kernel is M×DK×DK×N. For example, if the input feature map is m×n×16 and we want to output 32 channels, then the convolution kernel should be 16×3×3×32, which can be decomposed into a depthwise convolution: 16×3×3 to obtain a feature map with 16 channels. The pointwise convolution is 16×1×1×32. If using standard convolution, the computational cost is: m×n×16×3×3×32 = m×n×4608. After using depthwise separable convolution, the computational cost is m×n×16×3×3 + m×n×16×1×1×32 = m×n×656, and the number of parameters is greatly reduced. Step S2: Select at least one of the detected objects corresponding to the detected object IDs as the tracking object.

[0052] Preferably, after obtaining the detected objects, the position information of the detected objects, and the detected object IDs for each frame of the image in step S1, it further includes: displaying the object bounding boxes of the detected objects and the detected object IDs on the image.

[0053] Then step S2 further includes: using the detected objects corresponding to at least one of the detected object IDs sent by the ground station as the tracking objects.

[0054] Step S3: Input the position information of the tracking object into a MOSSE or KCF tracker for initialization.

[0055] Step S4: Track the tracking object in the video frames of the subsequent real-time video stream of the UAV through a MOSSE or KCF tracker.

[0056] Preferably, as Figure 2 shown, after step S4, it further includes in sequence: Step S5: Determine that a preset number of frames of the tracking object in the real-time video stream have been tracked, then stop tracking and execute step S1 for object detection to obtain the position information of the newly detected objects.

[0057] Among them, KCF is the Kernel Correlation Filter algorithm, which was proposed by Joao F. Henriques, Rui Caseiro, Pedro Martins, and Jorge Batista in 2014. It has excellent performance both in tracking effect and tracking speed. Its main principle is to use the given samples to train a discriminant classifier to determine whether the tracked object is the target or the surrounding background information.

[0058] Based on the correlation filtering method, MOSSE constructs an object tracker that can learn online. It uses grayscale images to construct a correlation filter for correlation calculation, and obtains the position of the object in the next frame image through the correlation operation between the image and the correlation filter. Since the grayscale image has only one channel and the features are very simple, the operation speed of MOSSE is very fast.

[0059] S6: Perform IOU matching on the position information of the re-detected object and the position information of the previously tracked object, and use the position information of the object with an IOU greater than the IOU threshold and the largest intersection-over-union ratio as the position information of the updated tracked object.

[0060] S7: Sequentially execute steps S3 to S6.

[0061] Preferably, in step S5, when performing object detection in step S1, it further includes: if the object detection of the current frame image fails when performing object detection in step S1, the current frame image uses the MOSSE or KCF tracker to continue tracking the tracked object, and re-perform step S1 on the next frame image for object detection.

[0062] Preferably, step S6 further includes: if there is no position information of the object with an IOU greater than the IOU threshold and the largest intersection-over-union ratio, the current frame image uses the MOSSE or KCF tracker to continue tracking the tracked object, re-perform step S1 on the next frame image for object detection, the re-detected position information of the object, and perform IOU matching on the re-detected position information of the object and the position information of the previously tracked object again. Preferably, in step S7, while performing step S3, it further includes: extracting the depth feature information of the image corresponding to the object position in the current frame and storing it in the feature sample library;

[0063] Step S6 further includes: if it is determined that the number of frames corresponding to the depth feature information in the feature sample library is greater than a preset threshold, then perform IOU matching between the newly detected target position information and the position information of the previously tracked target, and at the same time calculate and match the cosine distance between the depth feature information corresponding to the newly detected target position information and the depth feature information in the feature sample library to obtain a comprehensive metric C. If the comprehensive metric C is greater than the preset threshold and the target position information with the largest intersection over union is used as the position information of the updated tracked target, where the calculation formula for the comprehensive metric C is:

[0064] C = λd (1) +(1 - λ)d (2)

[0065] where d (1) is the degree of IOU matching, and d (2) is the degree of matching of depth features, and λ is the weight coefficient of the degree of IOU matching and the degree of matching of depth features. Among them, λ decreases as the sample size of the depth feature information in the feature sample library increases, and the weight coefficient λ of the degree of IOU matching and the degree of matching of depth features decreases as the sample size of the depth feature information in the feature sample library increases; λ satisfies:

[0066] λ = 1 - 0.02t

[0067] t is the sample size of the depth feature information. In this embodiment, the maximum value of the sample size of the depth feature information is 30. Of course, other maximum values can also be set separately.

[0068] In this embodiment, a preset number of invariant initial samples are stored in the feature sample library, and during the information update process of the feature sample library, the newly successfully matched depth feature information samples replace the old depth feature information samples.

[0069] The following is Embodiment 2, which specifically elaborates on the working principle of the video stream target tracking method for drones:

[0070] The flow of the entire detection and tracking task is as Figure 3 shown. This is a conceptual framework of an algorithm that combines a mobilenet-based SSD detector and a MOSSE or KCF tracker. We divide the entire tracking problem into two stages. The first stage is the retrieval stage, and the second stage is the stage of actual long-term tracking through the technical idea of combining detection and tracking. The following provides a detailed explanation of the division of labor in the two stages and the solution to the tracking problem.

[0071] The first stage is a pure detection stage. The detection model of mobilenet-SSD is adopted and starts running after starting to read the camera video stream. In this stage, each frame of the picture is detected. If a target (pedestrian) is detected, its position information box=(x,y,w,h) will be returned, and at the same time, the position of the target box and its ID number will be displayed on the image. This stage is the retrieval stage of the system and does not perform actual tracking. Its function is to continuously provide selectable tracking targets. After confirming the target to be tracked, the ground station can send the target ID number information and other methods to select the target area to be tracked, and then enter the second stage.

[0072] The second stage is a tracking algorithm based on IOU matching and deep feature matching. Its basic idea is as Figure 4 shown. After inputting the number information of the target in the first stage, the position information box corresponding to this number information is read. At the same time, the position information of this target is used as input to initialize the MOSSE or KCF tracker, and this target is continuously tracked in subsequent frames through the tracking algorithm. Compared with the Kalman filter used by deepsort to predict positions, the MOSSE or KCF tracker has higher accuracy and sufficient speed, and hardly occupies computing resources compared with deep learning algorithms. However, correspondingly, in a more complex environment, due to the accumulation of multi-frame prediction position deviations, it is easier to have offsets and misalignments of the predicted target positions. Therefore, in this stage, while continuously tracking, the mobilenet-SSD is re-run every fixed number of frames (10 frames) for target detection, and the position information of all detected targets is matched with the previous tracking target box information by IOU. The detection box with an IOU greater than the IOU threshold and the largest intersection-over-union ratio is used as the updated target position, and this position information is used to re-initialize the MOSSE or KCF tracker, and then the new target position is continuously tracked with a new tracker, and so on in a loop. Since there are not enough target samples to learn the target depth features at the beginning of the tracking stage, in this process, if the detection fails or cannot be matched with the tracking box, the detection result of this time is abandoned, and the tracking result of the MOSSE or KCF tracker is selected to continue tracking, and at the same time, the detection of the target is re-tried in the next frame and the tracker is updated. If this stage can complete the detection normally and the detected target position can be matched with the predicted target position of the tracker, while initializing the tracker with the new detection result, the depth features of the target box area picture of this frame will also be extracted and stored in the feature sample library.

[0073] As the number of successfully tracked video frames increases, the number of correctly matched feature samples obtained will gradually increase. A threshold (5 frames) is set. When enough threshold feature samples are obtained, deep feature matching is enabled. The deep feature is a high-dimensional vector obtained by the feature extraction network for extracting features from the target area image that has successfully completed the matching. The matching of deep features is to calculate the cosine distance between the vector sets of the current frame image and the sample library image. The smaller the distance, the higher the matching degree. The difference from the previous process is that when detecting after tracking 10 frames and matching the detection result with the tracker prediction result, it is necessary to consider both the IOU matching degree and the deep feature matching degree at the same time. The calculation formula for the comprehensive metric C is as follows: (1) Degree and deep feature matching degree (2) Degree, and the calculation formula for the comprehensive metric C is:

[0074] C = λd (1) +(1 - λ)d (2)

[0075] At the same time, self-feedback is set so that the proportion of deep feature matching increases as the number of sample features increases. To prevent the excessive sample number from increasing the additional calculation overhead, a maximum sample number of 30 is set. To enable the sample library to represent the latest state of the target and prevent the target from being unrecognizable after inversion or deformation, a sample library update strategy is set to replace the old samples with the newly successfully matched samples. At the same time, to prevent subsequent incorrect targets from contaminating the sample pool after tracking deviation, the initial 10 samples are retained without update, which can ensure the effectiveness of the sample library to the greatest extent.

[0076] In summary, through the above combined strategy of detection and tracking, the present invention can perform relatively stable target tracking with limited computing resources. After adding deep feature matching, the recognition ability and tracking ability of this tracking algorithm for fixed targets will be significantly enhanced over time. At the same time, after the proportion of deep feature matching gradually increases, it can also complete the re-discovery after the target is lost.

[0077] As Figure 5 Embodiment 3 of the present invention also provides a structural block diagram of a video stream target tracking device for an unmanned aerial vehicle, including a target preliminary detection module 310, a target determination module 320, an initialization module 330, and a target tracking module 340.

[0078] The target preliminary detection module 310 is used to perform target detection on each frame of the real-time video stream of the unmanned aerial vehicle by using the detection model of mobilenet-SSD, and obtain the detection target, the position information of the detection target, and the detection target ID of each frame of the image.

[0079] The target determination module 320 is used to select at least one of the detection targets corresponding to the detection target IDs as the tracking target.

[0080] An initialization module 330 is configured to input the position information of the tracked target into a MOSSE or KCF tracker for initialization.

[0081] A target tracking module 340 is configured to track the tracked target in video frames of the real-time video stream of the subsequent unmanned aerial vehicle by using the MOSSE or KCF tracker.

[0082] The video stream target tracking device for an unmanned aerial vehicle can also achieve corresponding technical effects of the video stream target tracking method for an unmanned aerial vehicle. Details have been described in the foregoing, and will not be repeated here.

[0083] Correspondingly, an embodiment of the present invention further provides a computer-readable storage medium. Instructions are stored in the storage medium, and when the instructions run, they execute any one of the video stream target tracking methods for an unmanned aerial vehicle provided in the foregoing embodiments. Therefore, corresponding technical effects can also be achieved. Details have been described in the foregoing, and will not be repeated here.

[0084] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0085] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A video stream target tracking method for unmanned aerial vehicles, characterized in that: The following steps are involved: S1: Target detection is performed on the image frames of the real-time video stream of the drone to obtain the detected target and its location information and target ID; S2: Select at least one detection target corresponding to the target ID as the tracking target; S3: inputting the position information of the tracking target into the tracker for initialization; S4: Tracking the tracking target in subsequent image frames of the real-time video stream of the drone through a tracker; Determine whether the tracked image frames have reached the preset frame number. If so, stop tracking and perform the following steps: S5: re-detecting the target in the current image frame to obtain the position information of each detected target; S6: Perform IOU matching on the re-detected target position information and the previously tracked target position information, and use the target position information with a value greater than the IOU threshold and the largest intersection-over-union ratio as the updated tracked target position information; S7: Repeating steps S3 to S6, wherein while executing step S3, the process further includes: extracting depth feature information of the image corresponding to the target position in the image frame, and storing the information in a feature sample library; Step S6 executed by step S7 also includes: determining whether the number of frames corresponding to the depth feature information in the feature sample library is greater than a threshold value, and if so, performing IOU matching on the re-detected target position information and the position information of the previously tracked target, and performing cosine distance matching calculation on the depth feature information corresponding to the re-detected target position information and the depth feature information of the feature sample library to obtain a comprehensive metric C, and taking the target position information whose comprehensive metric C is greater than a preset value and whose intersection-over-union ratio is the largest as the updated position information of the tracked target; The calculation formula of the comprehensive metric C is: in, is the IOU matching degree, is the matching degree of deep features, It is the weight coefficient of the IOU matching degree and the deep feature matching degree.

2. The video stream target tracking method for unmanned aerial vehicle according to claim 1, characterized in that: Step S1 also includes: displaying a target frame and a target ID of the detected target on the image frame; The specific method of step S2 is: select a target ID in an interactive manner at the ground station, and use the detection target corresponding to the selected target ID as the tracking target.

3. The video stream target tracking method for a drone according to claim 1, characterized in that: Step S5 also includes: if the re-target detection fails, continue to use the tracker to track the previous tracking target, and execute step S5 for the next image frame.

4. The video stream target tracking method for unmanned aerial vehicle according to claim 1, characterized in that: Step S6 also includes: if there is no target position information with an IOU greater than the IOU threshold and a maximum intersection-over-union ratio, continue to use the tracker to track the previous tracking target, and execute steps S5 and S6 for the next image frame.

5. The video stream target tracking method for unmanned aerial vehicle according to claim 1, characterized in that: The weight coefficient of the IOU matching degree and the deep feature matching degree As the number of deep feature information samples in the feature sample library increases, satisfy: = 1 - 0.02t t is the sample size of deep feature information.

6. The video stream target tracking method for unmanned aerial vehicle according to claim 1, characterized in that: The feature sample library stores a preset number of unchanged initial samples. The number of samples in the feature sample library has an upper limit. When the upper limit is exceeded, the newly successfully matched depth feature information samples replace the old non-initial sample's depth feature information samples.

7. A video stream target tracking device for a drone, characterized in that: include: The target initial detection module is used to use the mobilenet-SSD detection model to perform target detection on each frame of the drone's real-time video stream, and obtain the detection target of each frame as well as the location information and target ID of the detection target; A target determination module, used to select a detection target corresponding to at least one target ID as a tracking target; An initialization module is used to input the position information of the tracking target into the MOSSE or KCF tracker for initialization; The target tracking module is used to track the tracking target in the image frames of the real-time video stream of the subsequent drone through a MOSSE or KCF tracker.

8. A computer storage medium, characterized in that: The storage medium stores instructions, which, when executed, execute a video stream target tracking method for a drone as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Target tracking method and device, electronic equipment and readable storage medium

    CN110910422A

  • Multi-target tracking algorithm, electronic device and computer readable storage medium

    CN111428642A