Method, device, equipment and medium for tracking moving targets
By acquiring and combining the characteristic information of single-objective tracking frame, multi-objective tracking frame and multi-objective, the actual tracking frame of the moving target is determined, which solves the problem that it is difficult to accurately track the moving target under occlusion, and improves the accuracy of tracking.
Patent Information
- Application Number
- CN202210427115.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-22
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-04-22
AI Technical Summary
In the prior art, when tracking moving targets, especially when the target is frequently blocked, it is difficult to track accurately, which easily leads to the target being lost or wrong.
By obtaining the characteristic information of the single-objective tracking box, multi-objective tracking box and multi-objective in the video frame, and combining this information to determine the actual tracking box of the moving target to be tracked, thereby achieving accurate tracking of the moving target.
Improve the accuracy of determining the actual position of the moving target when there is an occlusion, avoiding the situation where it is difficult to determine with a single-target tracking box alone, and ensuring accurate tracking of the moving target.
Smart Images

Figure CN114820705B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method, device, equipment and medium for tracking a moving target. Background Art
[0002] Target tracking means first detecting the target of interest to the system in an image sequence, accurately locating the target, and then continuously updating the target's motion information as the target moves, thereby achieving continuous tracking of the target. Target tracking includes single target tracking, which only focuses on one target of interest. Its task is to design a motion model or appearance model to solve the influence of factors such as scale transformation, target occlusion, and illumination, and calibrate the image position corresponding to the target of interest frame by frame. However, when the target is frequently occluded, it is difficult to track the target, which usually causes the target to be lost or mistracked. Summary of the invention
[0003] The main purpose of the present invention is to provide a method, device, equipment and medium for tracking a moving target, aiming to solve the problem of how to accurately track the target to be tracked.
[0004] Get the video frame of the video to be tracked;
[0005] Determine a single target tracking frame, a multi-target tracking frame, and feature information of multiple targets in the video frame; wherein the multiple targets are multiple moving targets in the video frame, and the feature information is used to distinguish the multiple moving targets;
[0006] An actual tracking frame of the moving target to be tracked is determined according to the single target tracking frame, the multi-target tracking frame and the feature information of the multi-targets, so as to track the actual position of the moving target to be tracked in the video to be tracked according to the actual tracking frame.
[0007] In one embodiment, the step of determining the actual tracking frame of the moving target to be tracked according to the single target tracking frame, the multiple target tracking frame and the feature information of the multiple targets comprises:
[0008] Filtering the multi-target tracking frame according to the single-target tracking frame to obtain a filtered multi-target tracking frame;
[0009] According to the feature information of the multiple targets and the filtered multi-target tracking frame, an actual tracking frame of the moving target to be tracked is determined; wherein the actual tracking frame is a tracking frame in the filtered multi-target tracking frame in the current video frame whose similarity with the multi-target tracking frames in all video frames before the current video frame is greater than a preset threshold.
[0010] In one embodiment, the step of filtering the multi-target tracking frame according to the single-target tracking frame to obtain the filtered multi-target tracking frame includes:
[0011] Determine a first tracking frame and a second tracking frame according to the single target tracking frame and the multi-target tracking frame; wherein the first tracking frame is the multi-target tracking frame whose distance from the single target tracking frame in the previous video frame of the current video frame is less than a preset threshold; and the second tracking frame is the multi-target tracking frame whose intersection-over-union ratio with the single target tracking frame of the current video frame is less than a preset ratio;
[0012] A filtered multi-target tracking frame is determined according to the first tracking frame and the second tracking frame.
[0013] In one embodiment, the step of determining the first tracking frame and the second tracking frame according to the single-target tracking frame and the multi-target tracking frame includes:
[0014] Inputting the single target tracking frame and the multi-target tracking frame into a preset search area filter to obtain the first tracking frame; wherein the search area filter is used to filter out the multi-target tracking frame whose distance to the single target tracking frame in the previous video frame of the current video frame is greater than or equal to a preset threshold from the multi-target tracking frame;
[0015] The single target tracking frame and the multi-target tracking frame are input into a preset overlap filter to obtain the second tracking frame; wherein the overlap filter is used to filter out the multi-target tracking frame whose intersection-over-union ratio with the single target tracking frame of the current video frame is greater than or equal to a preset ratio from the multi-target tracking frame.
[0016] In one embodiment, determining the actual tracking frame of the moving target to be tracked according to the feature information of the multiple targets and the filtered multiple target tracking frame includes:
[0017] The feature information of the multiple targets and the filtered multi-target tracking frame are input into a preset ID filter to obtain the actual tracking frame; wherein the ID filter is used to filter out the tracking frames whose similarity with the multi-target tracking frames in all video frames before the current video frame is less than or equal to a preset threshold from the filtered multi-target tracking frame.
[0018] In one embodiment, the step of determining a single target tracking frame, a multi-target tracking frame, and feature information of multiple targets according to the video frame includes:
[0019] Inputting the video frame into a preset single target tracking model to obtain a single target tracking frame, wherein the single target tracking model is obtained by training a preset neural network model with a first training set, wherein the first training set includes training video frames with different resolutions;
[0020] The video frame is input into a preset multi-target tracking model to obtain a multi-target tracking frame and feature information of multiple targets, wherein the multi-target tracking model is obtained by training a preset network model with a second training set, and the second training set includes training video frames processed based on a preset motion blur algorithm.
[0021] In one embodiment, each set of training data in the first training set includes: a first training video frame with a preset first size, a preset first resolution, and clipped with the center position of the moving target; a second training video frame with a preset second size, a preset second resolution, and clipped with the tracking frame of the previous training video frame of the current training video frame as the center position; a third training video frame with a preset first size, a preset third resolution, and clipped with the center position of the moving target;
[0022] Among them, the neural network model includes a feature extraction network, a first convolutional layer, a second convolutional layer and a third convolutional layer; the first training video frame is used to be input into the feature extraction network so that the feature extraction network extracts the first position feature and the first category feature of the moving target, and the first position feature and the first category feature are input into the first convolutional layer; the second training video frame is used to be input into the feature extraction network so that the feature extraction network extracts the second position feature and the second category feature of the moving target, and the second position feature is input into the first convolutional layer and the second category feature is input into the second convolutional layer and the third convolutional layer; the third training video frame is used to be input into the third convolutional layer.
[0023] To achieve the above object, the present invention further provides a moving target tracking device, the moving target tracking device comprising:
[0024] An acquisition module, used for acquiring a video frame of a video to be tracked;
[0025] A determination module, used to determine a single target tracking frame, a multi-target tracking frame and feature information of multiple targets in the video frame; wherein the multiple targets are multiple moving targets in the video frame, and the feature information is used to distinguish the multiple moving targets;
[0026] A tracking module is used to determine an actual tracking frame of the moving target to be tracked based on the single-target tracking frame, the multi-target tracking frame and the feature information of the multi-targets, so as to track the actual position of the moving target to be tracked in the video to be tracked based on the actual tracking frame.
[0027] To achieve the above-mentioned purpose, the present invention also provides a moving target tracking device, which includes a memory, a processor, and a moving target tracking program stored in the memory and executable on the processor. When the moving target tracking program is executed by the processor, the various steps of the moving target tracking method described above are implemented.
[0028] To achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a moving target tracking program, and when the moving target tracking program is executed by a processor, the various steps of the moving target tracking method described above are implemented.
[0029] The present invention provides a method, device, equipment and medium for tracking a moving target, which obtains a video frame of a video to be tracked; determines a single target tracking frame, a multi-target tracking frame and feature information of the multi-targets in the video frame; wherein the multi-targets are multiple moving targets in the video frame, and the feature information is used to distinguish the multiple moving targets; determines an actual tracking frame of the moving target to be tracked according to the single target tracking frame, the multi-target tracking frame and the feature information of the multi-targets, so as to track the actual position of the moving target to be tracked in the video to be tracked according to the actual tracking frame. The actual tracking frame of the moving target to be tracked is determined by jointly determining the actual tracking frame of the moving target to be tracked by the single target tracking frame, the multi-target tracking frame and the feature information of the multi-targets. When the moving target to be tracked is blocked and it is difficult to determine the moving target to be tracked only by the single target tracking frame, the position of the actual tracking frame can be determined by combining the multi-target tracking frame and the feature information of the multi-targets, thereby improving the accuracy of obtaining the actual tracking frame used to determine the actual position of the moving target to be tracked, which is conducive to accurately tracking the target to be tracked. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A schematic diagram of the hardware structure of a moving target tracking device according to an embodiment of the present invention;
[0031] Figure 2 A schematic flow chart of a first embodiment of a method for tracking a moving target of the present invention;
[0032] Figure 3 It is a detailed flowchart diagram of step S30 of the second embodiment of the moving target tracking method of the present invention;
[0033] Figure 4 A schematic diagram of a single target tracking model, a multi-target tracking model and a filter in a moving target tracking method of the present invention;
[0034] Figure 5 It is a detailed flow chart of step S20 of the second embodiment of the moving target tracking method of the present invention;
[0035] Figure 6A schematic diagram of the structure of a single target tracking model of the moving target tracking method of the present invention;
[0036] Figure 7 The figure is a schematic diagram of the logical structure of the moving target tracking device of the present invention.
[0037] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0038] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0039] The main solution of the embodiment of the present invention is: obtaining a video frame of a video to be tracked; determining a single-target tracking frame, a multi-target tracking frame and feature information of multiple targets in the video frame; wherein the multiple targets are multiple moving targets in the video frame, and the feature information is used to distinguish the multiple moving targets; determining an actual tracking frame of the moving target to be tracked based on the single-target tracking frame, the multi-target tracking frame and the feature information of the multiple targets, so as to track the actual position of the moving target to be tracked in the video to be tracked based on the actual tracking frame.
[0040] The actual tracking frame of the moving target to be tracked is jointly determined by the single-target tracking frame, the multi-target tracking frame and the characteristic information of the multiple targets. When the moving target to be tracked is occluded and it is difficult to determine the moving target to be tracked by only the single-target tracking frame, the position of the actual tracking frame can be determined by combining the multi-target tracking frame and the characteristic information of the multiple targets, thereby improving the accuracy of the actual tracking frame used to determine the actual position of the moving target to be tracked, which is conducive to accurately tracking the target to be tracked.
[0041] As an implementation scheme, the moving target tracking device can be as follows Figure 1 shown.
[0042] The embodiment of the present invention relates to a moving target tracking device, which includes: a processor 101, such as a CPU, a memory 102, and a communication bus 103. The communication bus 103 is used to realize connection and communication between these components.
[0043] The memory 102 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Figure 1 As shown, the memory 102 as a computer-readable storage medium may include a moving target tracking program; and the processor 101 may be used to call the moving target tracking program stored in the memory 102 and perform the following operations:
[0044] Get the video frame of the video to be tracked;
[0045] Determine a single target tracking frame, a multi-target tracking frame, and feature information of multiple targets in the video frame; wherein the multiple targets are multiple moving targets in the video frame, and the feature information is used to distinguish the multiple moving targets;
[0046] An actual tracking frame of the moving target to be tracked is determined according to the single target tracking frame, the multi-target tracking frame and the feature information of the multi-targets, so as to track the actual position of the moving target to be tracked in the video to be tracked according to the actual tracking frame.
[0047] In one embodiment, the processor 101 may be used to call a moving target tracking program stored in the memory 102 and perform the following operations:
[0048] Filtering the multi-target tracking frame according to the single-target tracking frame to obtain a filtered multi-target tracking frame;
[0049] According to the feature information of the multiple targets and the filtered multi-target tracking frame, an actual tracking frame of the moving target to be tracked is determined; wherein the actual tracking frame is a tracking frame in the filtered multi-target tracking frame in the current video frame whose similarity with the multi-target tracking frames in all video frames before the current video frame is greater than a preset threshold.
[0050] In one embodiment, the processor 101 may be used to call a moving target tracking program stored in the memory 102 and perform the following operations:
[0051] Determine a first tracking frame and a second tracking frame according to the single target tracking frame and the multi-target tracking frame; wherein the first tracking frame is the multi-target tracking frame whose distance from the single target tracking frame in the previous video frame of the current video frame is less than a preset threshold; and the second tracking frame is the multi-target tracking frame whose intersection-over-union ratio with the single target tracking frame of the current video frame is less than a preset ratio;
[0052] A filtered multi-target tracking frame is determined according to the first tracking frame and the second tracking frame.
[0053] In one embodiment, the processor 101 may be used to call a moving target tracking program stored in the memory 102 and perform the following operations:
[0054] Inputting the single target tracking frame and the multi-target tracking frame into a preset search area filter to obtain the first tracking frame; wherein the search area filter is used to filter out the multi-target tracking frame whose distance to the single target tracking frame in the previous video frame of the current video frame is greater than or equal to a preset threshold from the multi-target tracking frame;
[0055] The single target tracking frame and the multi-target tracking frame are input into a preset overlap filter to obtain the second tracking frame; wherein the overlap filter is used to filter out the multi-target tracking frame whose intersection-over-union ratio with the single target tracking frame of the current video frame is greater than or equal to a preset ratio from the multi-target tracking frame.
[0056] In one embodiment, the processor 101 may be used to call a moving target tracking program stored in the memory 102 and perform the following operations:
[0057] The feature information of the multiple targets and the filtered multi-target tracking frame are input into a preset ID filter to obtain the actual tracking frame; wherein the ID filter is used to filter out the tracking frames whose similarity with the multi-target tracking frames in all video frames before the current video frame is less than or equal to a preset threshold from the filtered multi-target tracking frame.
[0058] In one embodiment, the processor 101 may be used to call a moving target tracking program stored in the memory 102 and perform the following operations:
[0059] Inputting the video frame into a preset single target tracking model to obtain a single target tracking frame, wherein the single target tracking model is obtained by training a preset neural network model with a first training set, wherein the first training set includes training video frames with different resolutions;
[0060] The video frame is input into a preset multi-target tracking model to obtain a multi-target tracking frame and feature information of multiple targets, wherein the multi-target tracking model is obtained by training a preset network model with a second training set, and the second training set includes training video frames processed based on a preset motion blur algorithm.
[0061] Based on the hardware architecture of the above-mentioned moving target tracking device, an embodiment of the moving target tracking method of the present invention is proposed.
[0062] Reference Figure 2 , Figure 2 The first embodiment of the method for tracking a moving target of the present invention comprises the following steps:
[0063] Step S10, obtaining a video frame of the video to be tracked.
[0064] Specifically, the video to be tracked may be a sports video, a variety show, or a road traffic video, etc., and a video frame of the video to be tracked is obtained, wherein the video frame includes a moving target to be tracked, and the moving target to be tracked may be an athlete in a sports video, a star of a variety show, or a target vehicle in a road traffic video. The video frames of the video to be tracked are obtained, and the video frames may be sorted in the video time sequence.
[0065] Step S20, determining a single target tracking frame, a multi-target tracking frame and feature information of multiple targets in the video frame; wherein the multiple targets are multiple moving targets in the video frame, and the feature information is used to distinguish the multiple moving targets.
[0066] Specifically, the single target tracking frame is the tracking frame of the moving target to be tracked in the video frame. However, since the moving target may be blocked by objects or people during its movement, the single target tracking frame cannot be accurately determined, resulting in the failure of tracking the target to be tracked. The multi-target tracking frame is the tracking frame of multiple moving targets in the video frame, and each moving target in the video frame corresponds to a tracking frame. The feature information of multiple targets allows users to distinguish multiple moving targets.
[0067] Determine the single target tracking frame, multi-target tracking frame and feature information of multiple targets in the video frame, wherein the feature information includes at least one of identity information, appearance information, tracking status, etc. of the moving target, and the feature information can be represented by a feature vector.
[0068] Step S30, determining an actual tracking frame of the moving target to be tracked according to the single target tracking frame, the multi-target tracking frame and the feature information of the multi-targets, so as to track the actual position of the moving target to be tracked in the video to be tracked according to the actual tracking frame.
[0069] Specifically, the actual tracking frame is the tracking frame of the moving target to be tracked. The actual position of the moving target to be tracked in the video to be tracked is tracked according to the actual tracking frame. The actual tracking frame can be displayed in the video to be tracked.
[0070] The actual tracking frame of the moving target to be tracked is jointly determined based on the single target tracking frame, the multi-target tracking frame and the characteristic information of the multi-targets, thereby avoiding the situation where the moving target to be tracked is blocked and the actual position of the moving target to be tracked is difficult to accurately determine with the single target tracking frame alone. Optionally, the single target tracking frame of the previous video frame in the current video frame is determined, the multi-target tracking frame in the current video frame is determined, the distance between the single target tracking frame and each multi-target tracking frame is determined, and the multi-target tracking frame whose distance is less than a preset threshold is determined as the actual tracking frame.
[0071] In the technical solution of this embodiment, a video frame of a video to be tracked is obtained; a single target tracking frame, a multi-target tracking frame, and feature information of multiple targets in the video frame are determined; wherein the multiple targets are multiple moving targets in the video frame, and the feature information is used to distinguish the multiple moving targets; an actual tracking frame of the moving target to be tracked is determined according to the single target tracking frame, the multi-target tracking frame, and the feature information of the multiple targets, so as to track the actual position of the moving target to be tracked in the video to be tracked according to the actual tracking frame. The actual tracking frame of the moving target to be tracked is determined by jointly determining the actual tracking frame of the moving target to be tracked by the single target tracking frame, the multi-target tracking frame, and the feature information of the multiple targets. When the moving target to be tracked is occluded and it is difficult to determine the moving target to be tracked only by the single target tracking frame, the position of the actual tracking frame can be determined by combining the multi-target tracking frame and the feature information of the multiple targets, thereby improving the accuracy of obtaining the actual tracking frame used to determine the actual position of the moving target to be tracked, which is conducive to accurately tracking the target to be tracked.
[0072] Reference Figure 3 , Figure 3 This is a second embodiment of the moving target tracking method of the present invention, based on the first embodiment, step S30 includes:
[0073] Step S31, filtering the multi-target tracking frame according to the single-target tracking frame to obtain a filtered multi-target tracking frame;
[0074] Step S32, determining an actual tracking frame of the moving target to be tracked based on the feature information of the multiple targets and the filtered multi-target tracking frame; wherein the actual tracking frame is a tracking frame in the filtered multi-target tracking frame in the current video frame whose similarity with the multi-target tracking frames in all video frames before the current video frame is greater than a preset threshold.
[0075] Specifically, the multi-target tracking frame is filtered according to the single-target tracking frame to obtain the filtered multi-target tracking frame. Optionally, the first tracking frame and the second tracking frame are determined according to the single-target tracking frame; wherein the first tracking frame is a multi-target tracking frame whose distance from the single-target tracking frame in the previous video frame of the current video frame is less than a preset threshold; the second tracking frame is a multi-target tracking frame whose intersection-over-intersection ratio with the single-target tracking frame of the current video frame is less than a preset ratio; the filtered multi-target tracking frame is determined according to the first tracking frame and the second tracking frame. Exemplarily, the same tracking frame in the first tracking frame and the second tracking frame is determined, and the same tracking frame is used as the filtered multi-target tracking frame.
[0076] like Figure 4As shown, according to the single-target tracking frame and the multi-target tracking frame, a first tracking frame and a second tracking frame are determined. Optionally, the single-target tracking frame and the multi-target tracking frame are input into a preset search area filter to obtain a first tracking frame; wherein the search area filter is used to filter out from the multi-target tracking frame a multi-target tracking frame whose distance from the single-target tracking frame in the previous video frame of the current video frame is greater than or equal to a preset threshold; the single-target tracking frame and the multi-target tracking frame are input into a preset overlap filter to obtain a second tracking frame; wherein the overlap filter is used to filter out from the multi-target tracking frame a multi-target tracking frame whose intersection-over-union ratio with the single-target tracking frame of the current video frame is greater than or equal to a preset ratio.
[0077] Among them, the post-filter PosFilter includes a search area filter and an overlap IoU filter. Optionally, the input of PosFilter is the multi-target tracking box and the single-target tracking box in the t-1 frame, and the input of the ID filter IDFilter is the feature information of the multiple targets in the t-th frame. The feature information of the multiple targets includes the tracking state representing the tracking state or the lost state, and the ID queue that maintains the historical appearance information of the target object.
[0078] The actual tracking frame is a tracking frame in the filtered multi-target tracking frame in the current video frame whose similarity with the multi-target tracking frame in all video frames before the current video frame is greater than a preset threshold. According to the feature information of the multi-target and the filtered multi-target tracking frame, the actual tracking frame of the moving target to be tracked is determined. Optionally, the feature information of the multi-target and the filtered multi-target tracking frame are input into a preset ID filter to obtain the actual tracking frame; wherein the ID filter is used to filter out the tracking frames whose similarity with the multi-target tracking frames in all video frames before the current video frame is less than or equal to the preset threshold from the filtered multi-target tracking frame.
[0079] Optionally, the ID filter focuses on the appearance information of the moving target, which can better utilize the IDFeature of the multi-target tracking model, that is, the feature information of the multi-target. Figure 4 As shown in the figure, in order to filter out the tracking frames whose similarity with the multi-target tracking frames in all video frames before the current video frame is less than or equal to the preset threshold from the filtered multi-target tracking frames, the ReIDFilter is proposed in the ID filter to calculate the distance matrix between the feature information of the multi-target and the ID queue of the historical appearance information of the moving target to be tracked. The ID filter filters out the multi-target tracking frames whose cosine distance with the ID queue of the historical appearance information of the moving target to be tracked is greater than the preset threshold. The filtered multi-target tracking frames are searched for the actual tracking frames through the multi-target matching module MOTMatch. If the actual tracking frame is not found, the range of the multi-target tracking frame is narrowed by a stricter ReIDFilter and a lower preset threshold.
[0080] Optionally, the output of the ID filter can be divided into four parts: ID features, actual tracking frames, initialization flags, and confidence scores. Among them, the ID features are used to update the ID queue, which is a queue that stores the appearance information of the moving target in the previous preset number of video frames. The actual tracking frame is used to indicate the position of the target to be tracked. The initialization flag Init_flag of the single target tracking model, when Init_flag = True, that is, the initialization indicates true, the single target tracking model is initialized. The confidence score Conf, where 0 indicates tracking and 1 indicates loss, is used to update the tracking state and control the parameters of the post-filter and the ID filter. When the previous frame is lost (tracking state>0) and the current frame is calibrated as a tracking state, the initialization flag is set to true, and the single target tracking model is initialized at the same time.
[0081] In the technical solution of this embodiment, the multi-target tracking frame is screened according to the single target tracking frame to obtain the screened multi-target tracking frame; the actual tracking frame of the moving target to be tracked is determined according to the feature information of the multi-targets and the screened multi-target tracking frame; wherein the actual tracking frame is a tracking frame in the screened multi-target tracking frame in the current video frame whose similarity with the multi-target tracking frames in all video frames before the current video frame is greater than a preset threshold. The multi-target tracking frame is screened according to the single target tracking frame to obtain the actual tracking frame, which avoids the situation where the moving target is blocked and it is difficult to determine it only by the single target tracking frame, improves the accuracy of the actual tracking frame obtained for determining the actual position of the moving target to be tracked, and is conducive to accurately tracking the target to be tracked.
[0082] Reference Figure 5 , Figure 5 The third embodiment of the method for tracking a moving target of the present invention is based on the first or second embodiment, and the step S20 includes:
[0083] Step S21, inputting the video frame into a preset single target tracking model to obtain a single target tracking frame, wherein the single target tracking model is obtained by training a preset neural network model with a first training set, wherein the first training set includes training video frames with different resolutions;
[0084] Step S22, input the video frame into a preset multi-target tracking model to obtain a multi-target tracking frame and feature information of multiple targets, wherein the multi-target tracking model is obtained by training a preset network model with a second training set, and the second training set includes training video frames processed based on a preset motion blur algorithm.
[0085] Specifically, the single target tracking model is used to identify a single target tracking frame in a video frame, and the single target tracking model is obtained by training a preset neural network model with a first training set, wherein the first training set includes training video frames with different resolutions.
[0086] Optionally, each set of training data in the first training set includes: a first training video frame with a preset first size, a preset first resolution, and cropped with the center position of the moving target; a second training video frame with a preset second size, a preset second resolution, and cropped with the tracking frame of the previous training video frame of the current training video frame as the center position; a third training video frame with a preset first size, a preset third resolution, and cropped with the center position of the moving target. For example, Figure 6 As shown, there are three sets of training video frames for input to the neural network model. The training video frame T is a picture cropped with the marked moving target as the center in the first frame; the training video frame S is a picture cropped with the tracking frame marked in the previous frame as the center; the training video frame C is a picture cropped with the marked moving target as the center in the first frame, wherein the resolution S>T>C.
[0087] Among them, Figure 6 As shown, the neural network model includes a feature extraction network, a first convolutional layer, a second convolutional layer and a third convolutional layer; the first training video frame is used to be input into the feature extraction network, so that the feature extraction network extracts the first position feature a1 and the first category feature b1 of the moving target, and the first position feature and the first category feature are input into the first convolutional layer; the second training video frame is used to be input into the feature extraction network, so that the feature extraction network extracts the second position feature a2 and the second category feature b2 of the moving target, and the second position feature a2 is input into the first convolutional layer and the second category feature b2 is input into the second convolutional layer and the third convolutional layer; the third training video frame is used to be input into the third convolutional layer.
[0088] like Figure 6 As shown, the single target tracking model includes a position branch, a score branch and a category branch, wherein the position branch is used to output the position of the single target tracking frame, the score branch is used to output the score heat map of the single target tracking frame, and the category branch is used to output the category of the single target tracking frame. Among them, the loss functions used for training the position branch, the score branch and the category branch are different. Optionally, the loss function corresponding to the position branch is shown in the following formula:
[0089]
[0090] Among them, L loc () is the IoU loss function, B is the true value, Represents the output of the neural network, Intersect is an intersection operation, and Union is a union operation.
[0091] Optionally, the loss function corresponding to the score branch is as shown in the following formula:
[0092]
[0093] Among them, L score () is the loss function of the score branch, N pos Represents the total number of pixels in the video frame, SCE() represents the Sigmoid Cross Entropy function; score i,j is the true value, represents the output of the score branch, i and j are the pixel positions of the score heat map.
[0094] Optionally, the loss function corresponding to the category branch is as shown in the following formula:
[0095]
[0096] Among them, L cls () is the loss function of the category branch; N pos represents the total number of pixels in the video frame, FOCAL() represents the Focal loss function; i and j are the pixel positions of the score heat map.
[0097] The multi-target tracking model is used to identify the multi-target tracking frame in the video frame. The multi-target tracking model adopts the multi-target tracking FairMOT network structure. The feature information of the multi-target can be the tracking state of the multi-target or the tracking state of the lost state, as well as the ID queue of the historical appearance information of the moving target.
[0098] The multi-target tracking model is obtained by training the preset network model with the second training set, including training video frames processed based on the preset motion blur algorithm, and the training video frames are of a preset size. The training video frames processed based on the preset motion blur algorithm are designed to adapt to the motion blur caused by high-speed athletes and high-speed cameras in sports scenes. Optionally, a linear point spread function is used to enhance the motion blur of the image. By convolving the point spread kernel function and the image, the motion blur phenomenon produced by the image can be simulated. In addition, enhancement methods such as image blurring, translation, and flipping are used to enable the network to obtain a stronger generalization ability.
[0099] In the technical solution of this embodiment, a video frame is input into a preset single target tracking model to obtain a single target tracking frame, the single target tracking model is obtained by training a preset neural network model with a first training set, and the first training set includes training video frames with different resolutions; a video frame is input into a preset multi-target tracking model to obtain a multi-target tracking frame and feature information of multiple targets, the multi-target tracking model is obtained by training a preset network model with a second training set, and the second training set includes training video frames processed based on a preset motion blur algorithm. The single target tracking frame is determined by the single target tracking model, and the feature information of the multi-target tracking frame and the multi-target tracking frame is determined by the multi-target tracking model, so as to improve the accuracy and efficiency of determining the single target tracking frame, the multi-target tracking frame and the feature information of the multi-target, so as to jointly determine the actual tracking frame of the moving target to be tracked, avoid the situation that when the moving target to be tracked is blocked, it is difficult to determine the moving target to be tracked only by the single target tracking frame, improve the accuracy of the actual tracking frame obtained for determining the actual position of the moving target to be tracked, and facilitate accurate tracking of the target to be tracked.
[0100] Reference Figure 7 The present invention also provides a moving target tracking device, the moving target tracking device comprising:
[0101] An acquisition module 100 is used to acquire a video frame of a video to be tracked;
[0102] A determination module 200, configured to determine a single target tracking frame, a multi-target tracking frame, and feature information of multiple targets in the video frame; wherein the multiple targets are multiple moving targets in the video frame, and the feature information is used to distinguish the multiple moving targets;
[0103] The tracking module 300 is used to determine the actual tracking frame of the moving target to be tracked according to the single-target tracking frame, the multi-target tracking frame and the feature information of the multi-targets, so as to track the actual position of the moving target to be tracked in the video to be tracked according to the actual tracking frame.
[0104] In one embodiment, in determining the actual tracking frame of the moving target to be tracked according to the single target tracking frame, the multi-target tracking frame and the feature information of the multi-targets, the tracking module 300 is specifically used to:
[0105] Filtering the multi-target tracking frame according to the single-target tracking frame to obtain a filtered multi-target tracking frame;
[0106] According to the feature information of the multiple targets and the filtered multi-target tracking frame, an actual tracking frame of the moving target to be tracked is determined; wherein the actual tracking frame is a tracking frame in the filtered multi-target tracking frame in the current video frame whose similarity with the multi-target tracking frames in all video frames before the current video frame is greater than a preset threshold.
[0107] In one embodiment, in terms of filtering the multi-target tracking frame according to the single-target tracking frame to obtain the filtered multi-target tracking frame, the tracking module 300 is specifically used to:
[0108] Determine a first tracking frame and a second tracking frame according to the single target tracking frame and the multi-target tracking frame; wherein the first tracking frame is the multi-target tracking frame whose distance from the single target tracking frame in the previous video frame of the current video frame is less than a preset threshold; and the second tracking frame is the multi-target tracking frame whose intersection-over-union ratio with the single target tracking frame of the current video frame is less than a preset ratio;
[0109] A filtered multi-target tracking frame is determined according to the first tracking frame and the second tracking frame.
[0110] In one embodiment, in determining the first tracking frame and the second tracking frame according to the single-target tracking frame and the multi-target tracking frame, the tracking module 300 is specifically configured to:
[0111] Inputting the single target tracking frame and the multi-target tracking frame into a preset search area filter to obtain the first tracking frame; wherein the search area filter is used to filter out the multi-target tracking frame whose distance to the single target tracking frame in the previous video frame of the current video frame is greater than or equal to a preset threshold from the multi-target tracking frame;
[0112] The single target tracking frame and the multi-target tracking frame are input into a preset overlap filter to obtain the second tracking frame; wherein the overlap filter is used to filter out the multi-target tracking frame whose intersection-over-union ratio with the single target tracking frame of the current video frame is greater than or equal to a preset ratio from the multi-target tracking frame.
[0113] In one embodiment, in determining the actual tracking frame of the moving target to be tracked according to the feature information of the multiple targets and the filtered multiple target tracking frame, the tracking module 300 is specifically used to:
[0114] The feature information of the multiple targets and the filtered multi-target tracking frame are input into a preset ID filter to obtain the actual tracking frame; wherein the ID filter is used to filter out the tracking frames whose similarity with the multi-target tracking frames in all video frames before the current video frame is less than or equal to a preset threshold from the filtered multi-target tracking frame.
[0115] In one embodiment, in determining the single-target tracking frame, the multi-target tracking frame, and the feature information of the multi-targets according to the video frame, the determination module 200 is specifically used to:
[0116] Inputting the video frame into a preset single target tracking model to obtain a single target tracking frame, wherein the single target tracking model is obtained by training a preset neural network model with a first training set, wherein the first training set includes training video frames with different resolutions;
[0117] The video frame is input into a preset multi-target tracking model to obtain a multi-target tracking frame and feature information of multiple targets, wherein the multi-target tracking model is obtained by training a preset network model with a second training set, and the second training set includes training video frames processed based on a preset motion blur algorithm.
[0118] The present invention also provides a moving target tracking device, which includes a memory, a processor, and a moving target tracking program stored in the memory and executable on the processor. When the moving target tracking program is executed by the processor, it implements the various steps of the moving target tracking method described in the above embodiment.
[0119] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a moving target tracking program, and when the moving target tracking program is executed by a processor, the various steps of the moving target tracking method described in the above embodiment are implemented.
[0120] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0121] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, system, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, system, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, system, article or device including the element.
[0122] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment system can be implemented by means of software plus a necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a computer-readable storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a parking management device, an air conditioner, or a network device, etc.) to execute the system described in each embodiment of the present invention.
[0123] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for tracking a moving target, characterized in that: The moving target tracking method comprises: Get the video frame of the video to be tracked; Determine a single target tracking frame, a multi-target tracking frame, and feature information of multiple targets in the video frame; wherein the multiple targets are multiple moving targets in the video frame, and the feature information is used to distinguish the multiple moving targets; Determine a first tracking frame and a second tracking frame according to the single target tracking frame and the multi-target tracking frame, wherein the first tracking frame is the multi-target tracking frame whose distance from the single target tracking frame in the previous video frame of the current video frame is less than a preset threshold; and the second tracking frame is the multi-target tracking frame whose intersection-over-union ratio with the single target tracking frame of the current video frame is less than a preset ratio; Determine a filtered multi-target tracking frame according to the first tracking frame and the second tracking frame; According to the feature information of the multiple targets and the filtered multi-target tracking frame, an actual tracking frame of the moving target to be tracked is determined, so as to track the actual position of the moving target to be tracked in the video to be tracked according to the actual tracking frame; wherein the actual tracking frame is a tracking frame in the filtered multi-target tracking frame in the current video frame whose similarity with the multi-target tracking frames in all video frames before the current video frame is greater than a preset threshold.
2. The method for tracking a moving target according to claim 1, wherein: The step of determining the first tracking frame and the second tracking frame according to the single-target tracking frame and the multi-target tracking frame comprises: Inputting the single target tracking frame and the multi-target tracking frame into a preset search area filter to obtain the first tracking frame; wherein the search area filter is used to filter out the multi-target tracking frame whose distance to the single target tracking frame in the previous video frame of the current video frame is greater than or equal to a preset threshold from the multi-target tracking frame; The single target tracking frame and the multi-target tracking frame are input into a preset overlap filter to obtain the second tracking frame; wherein the overlap filter is used to filter out the multi-target tracking frame whose intersection-over-union ratio with the single target tracking frame of the current video frame is greater than or equal to a preset ratio from the multi-target tracking frame.
3. The method for tracking a moving target according to claim 1, wherein: The step of determining the actual tracking frame of the moving target to be tracked according to the feature information of the multiple targets and the filtered multiple target tracking frame comprises: The feature information of the multiple targets and the filtered multi-target tracking frame are input into a preset ID filter to obtain the actual tracking frame; wherein the ID filter is used to filter out the tracking frames whose similarity with the multi-target tracking frames in all video frames before the current video frame is less than or equal to a preset threshold from the filtered multi-target tracking frame.
4. The method for tracking a moving target according to any one of claims 1 to 3, characterized in that: The step of determining a single target tracking frame, a multi-target tracking frame and feature information of multiple targets according to the video frame comprises: Inputting the video frame into a preset single target tracking model to obtain a single target tracking frame, wherein the single target tracking model is obtained by training a preset neural network model with a first training set, wherein the first training set includes training video frames with different resolutions; The video frame is input into a preset multi-target tracking model to obtain a multi-target tracking frame and feature information of multiple targets. The multi-target tracking model is obtained by training a preset network model with a second training set. The second training set includes training video frames processed based on a preset motion blur algorithm.
5. The method for tracking a moving target as claimed in claim 4, characterized in that: Each set of training data in the first training set includes: a first training video frame with a preset first size, a preset first resolution and cropped with the center position of the moving target; a second training video frame with a preset second size, a preset second resolution and cropped with the tracking frame of the previous training video frame of the current training video frame as the center position; a third training video frame with a preset first size, a preset third resolution and cropped with the center position of the moving target; Among them, the neural network model includes a feature extraction network, a first convolutional layer, a second convolutional layer and a third convolutional layer; the first training video frame is used to be input into the feature extraction network so that the feature extraction network extracts the first position feature and the first category feature of the moving target, and the first position feature and the first category feature are input into the first convolutional layer; the second training video frame is used to be input into the feature extraction network so that the feature extraction network extracts the second position feature and the second category feature of the moving target, and the second position feature is input into the first convolutional layer and the second category feature is input into the second convolutional layer and the third convolutional layer; the third training video frame is used to be input into the third convolutional layer.
6. A moving target tracking device, characterized in that: The moving target tracking device comprises: An acquisition module, used for acquiring a video frame of a video to be tracked; A determination module, used to determine a single target tracking frame, a multi-target tracking frame and feature information of multiple targets in the video frame; wherein the multiple targets are multiple moving targets in the video frame, and the feature information is used to distinguish the multiple moving targets; A tracking module is used to determine a first tracking frame and a second tracking frame according to the single-target tracking frame and the multi-target tracking frame, wherein the first tracking frame is the multi-target tracking frame whose distance from the single-target tracking frame in the previous video frame of the current video frame is less than a preset threshold; the second tracking frame is the multi-target tracking frame whose intersection-and-union ratio with the single-target tracking frame in the current video frame is less than a preset ratio; determine a filtered multi-target tracking frame according to the first tracking frame and the second tracking frame; determine an actual tracking frame of a moving target to be tracked according to feature information of the multi-targets and the filtered multi-target tracking frame, so as to track the actual position of the moving target to be tracked in the video to be tracked according to the actual tracking frame; wherein the actual tracking frame is a tracking frame in the filtered multi-target tracking frame in the current video frame whose similarity with the multi-target tracking frames in all video frames before the current video frame is greater than a preset threshold.
7. A moving target tracking device, characterized in that: The moving target tracking device includes a memory, a processor, and a moving target tracking program stored in the memory and executable on the processor. When the moving target tracking program is executed by the processor, the various steps of the moving target tracking method as described in any one of claims 1-5 are implemented.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a moving target tracking program, and when the moving target tracking program is executed by a processor, each step of the moving target tracking method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Identification-aided multi-target tracking method based on depth neural network
CN106097391A