Multi-target tracking method, device and equipment
By using the similarity calculation method of predicted tracking box and detection box in multi-objective tracking, the problem of low multi-objective tracking in the prior art is solved, and higher tracking accuracy is achieved.
Patent Information
- Application Number
- CN202311643714.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2025-06-03
AI Technical Summary
In the prior art, the accuracy of multi-objective tracking is low, especially under the influence of motion blur, lighting changes or occlusion, target matching errors are prone to occur.
By determining the predicted tracking box and detection box in the current video frame and obtaining the clarity of the historical detection box, the similarity between the detection box and the prediction tracking box is calculated based on the clarity, historical feature information and current feature information, thereby performing target tracking.
This method reduces the impact of image quality on similarity and reduces the probability of target tracking ID interchange, thereby improving the accuracy of multi-objective tracking.
Smart Images

Figure CN120088691A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a multi-object tracking method, device, and equipment. Background Art
[0002] With the development of computer vision technology, multi-object tracking is widely applied in fields such as autonomous driving, video surveillance, and behavior recognition. Therefore, how to perform multi-object tracking efficiently and accurately is an important topic in the industry at present.
[0003] In related technologies, in the process of multi-object tracking using video data, for the current video frame, first, computer vision technology is used to detect the objects in the current video frame, so as to obtain the actual positions of the objects in the current video frame, and the visual features of the objects in the current video frame are extracted. In addition, according to the historical trajectories of the objects in the previous video frames, the predicted positions of the objects in the current video frame can be predicted, and the visual features of the objects in the previous video frames are obtained. Thus, by determining the similarity between the actual position and the predicted position, and the similarity between the visual features of the objects in the current video frame and the visual features of the objects in the previous video frames, it is determined which objects belong to the same object, and then the object tracking is completed according to the actual position of the object in the current video frame.
[0004] However, in the above method, due to factors such as motion blur, illumination, or occlusion, the quality of the video frame image is poor, and at this time, the extracted visual features will be affected. Therefore, it is easy to have the situation of incorrect object matching, which affects the accuracy of multi-object tracking. Summary of the Invention
[0005] The present invention provides a multi-object tracking method, device, and equipment, which are used to solve the defect of low object tracking accuracy in the prior art and achieve the purpose of improving the object tracking accuracy.
[0006] The present invention provides a multi-object tracking method, including:
[0007] Determine at least one predicted tracking box in the current video frame and the detection box corresponding to each object; the historical feature information corresponding to the predicted tracking box is the feature information of the historical detection box in the first preset number of historical video frames before the current video frame, and the object represented by the historical detection box is the same as the object represented by the predicted tracking box;
[0008] For each of the predicted tracking boxes, obtain the clarity corresponding to each of the historical detection boxes associated with the predicted tracking box;
[0009] Determine the first similarity between each of the detection frames and the predicted tracking frame based on the clarity corresponding to each of the historical detection frames, the historical feature information corresponding to each of the historical detection frames, and the feature information corresponding to each of the detection frames;
[0010] Track each of the targets based on all the first similarities corresponding to each of the predicted tracking frames.
[0011] According to a multi-target tracking method provided by the present invention, the obtaining the clarity corresponding to each of the historical detection frames associated with the predicted tracking frame includes:
[0012] Determine the historical region where the target corresponding to each of the historical detection frames is located;
[0013] Determine the clarity of the historical region;
[0014] The determining the first similarity between each of the detection frames and the predicted tracking frame based on the clarity corresponding to each of the historical detection frames, the historical feature information corresponding to each of the historical detection frames, and the feature information corresponding to each of the detection frames includes:
[0015] For each of the historical detection frames, determine the feature weight of the historical detection frame based on the clarity of the historical region corresponding to the historical detection frame;
[0016] For each of the detection frames, determine the second similarity between the feature information corresponding to the detection frame and the historical feature information of each of the historical detection frames;
[0017] Based on the second similarity and the feature weight corresponding to each of the historical detection frames, determine the first similarity between the detection frame and the predicted tracking frame.
[0018] According to a multi-target tracking method provided by the present invention, the method further includes:
[0019] Determine the region occupied by the target corresponding to each of the detection frames in the current video frame;
[0020] Input the region into a feature extraction model, and divide the region into at least two sub-regions by the feature extraction model based on a preset ratio;
[0021] Extract the sub-feature information of each of the sub-regions through the feature extraction model;
[0022] Determine the feature information corresponding to the detection frame based on each of the sub-feature information.
[0023] According to a multi-target tracking method provided by the present invention, the tracking each of the targets based on all the first similarities corresponding to each of the predicted tracking frames includes:
[0024] For each of the predicted tracking boxes, when there is a first similarity greater than the similarity threshold among all the first similarities corresponding to the predicted tracking box, among the first similarities greater than the similarity threshold, the detection box corresponding to the maximum first similarity is determined as the detection box matching the predicted tracking box, and the target corresponding to the predicted tracking box is tracked based on the matching detection box.
[0025] According to a multi-object tracking method provided by the present invention, the method further includes:
[0026] When there is no first similarity greater than the similarity threshold among all the first similarities corresponding to the predicted tracking box, based on the position information of each detection box and the position information of the predicted tracking box, it is predicted whether each detection box and the predicted tracking box match;
[0027] When it is predicted that the target detection box and the predicted tracking box match, the first size information of the target detection box and the second size information of the target historical detection box corresponding to the predicted tracking box in the previous historical video frame are obtained;
[0028] Based on the first size information and the second size information, when it is determined that the change amplitude between the target detection box and the target historical detection box is greater than the preset amplitude, it is determined that the target detection box and the predicted tracking box do not match.
[0029] According to a multi-object tracking method provided by the present invention, the determination of the second similarity between the feature information corresponding to the detection box and the historical feature information of each historical detection box includes:
[0030] Based on the first position in the predicted tracking box, a first region including the predicted tracking box is determined, and the area of the first region is greater than the area of the predicted tracking box;
[0031] When it is determined that the detection box is within the first region, the second similarity between the feature information corresponding to the detection box and each historical feature information is determined;
[0032] When it is determined that the detection box is not within the first region, the second similarity between the feature information corresponding to the detection box and each historical feature information is determined as the preset similarity.
[0033] According to a multi-object tracking method provided by the present invention, the method further includes:
[0034] In the case where the prediction tracking boxes all fail to match in the second preset number of historical video frames before the current video frame, based on the first position and a preset value, expand the first region to obtain a second region, where the preset value is related to the second preset number;
[0035] Determine the second region as the new first region, and return to execute the step of determining the second similarity between the detection box and each piece of the historical feature information in the case where it is determined that the detection box is within the first region.
[0036] According to a multi-object tracking method provided by the present invention, each piece of historical feature information corresponding to the prediction tracking box is stored in a feature storage sequence;
[0037] The method further includes:
[0038] For each of the prediction tracking boxes, in the case where there is a first similarity greater than the similarity threshold among all the first similarities corresponding to the prediction tracking box, determine the minimum second similarity among the second similarities corresponding to the historical feature information of each of the historical detection boxes included in the feature storage sequence;
[0039] Delete the historical feature information corresponding to the minimum second similarity, and add the feature information of the detection box corresponding to the first similarity greater than the similarity threshold to the feature storage sequence.
[0040] The present invention also provides a multi-object tracking device, including:
[0041] A determination module, configured to determine at least one prediction tracking box in the current video frame and the detection box corresponding to each target; the historical feature information corresponding to the prediction tracking box is the feature information of the historical detection box in the first preset number of historical video frames before the current video frame, and the target represented by the historical detection box is the same as the target represented by the prediction tracking box;
[0042] An acquisition module, configured to, for each of the prediction tracking boxes, acquire the clarity corresponding to each of the historical detection boxes associated with the prediction tracking box;
[0043] The determination module is further configured to determine the first similarity between each detection box and the prediction tracking box based on the clarity corresponding to each historical detection box, the historical feature information corresponding to each historical detection box, and the feature information corresponding to each detection box;
[0044] A tracking module, configured to track each target based on all the first similarities corresponding to each prediction tracking box.
[0045] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the multi-object tracking method described in any one of the above is implemented.
[0046] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the multi-object tracking method described in any one of the above is implemented.
[0047] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the multi-object tracking method described in any one of the above is implemented.
[0048] The multi-object tracking method, device, and equipment provided by the present invention determine at least one predicted tracking box and a detection box corresponding to each target in the current video frame. The historical feature information corresponding to the predicted tracking box is the feature information of the historical detection boxes in the first preset number of historical video frames before the current video frame, and the target represented by the historical detection box is the same as the target represented by the predicted tracking box. For each predicted tracking box, the clarity corresponding to each historical detection box associated with the predicted tracking box can be obtained, and based on the clarity corresponding to each historical detection box, the historical feature information corresponding to each historical detection box, and the feature information corresponding to each detection box, the first similarity between each detection box and the predicted tracking box is determined. Then, based on all the first similarities corresponding to each predicted tracking box, each target is tracked. Since the clarity corresponding to each historical detection box can be used as the weight of the historical feature information to determine the first similarity between each detection box and the predicted tracking box, the influence of image quality on the first similarity can be reduced, thereby reducing the probability of the interchange of the tracking identity document (ID) of the target and improving the accuracy of multi-object tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 is a schematic flowchart of the multi-object tracking method provided by an embodiment of the present invention;
[0051] Figure 2 is a second schematic flowchart of the multi-object tracking method provided by an embodiment of the present invention;
[0052] Figure 3 is a schematic diagram of area division provided by an embodiment of the present invention;
[0053] Figure 4 It is a schematic structural diagram of a multi - target tracking device provided by an embodiment of the present invention;
[0054] Figure 5 It exemplifies a schematic physical structure diagram of an electronic device. Specific embodiments
[0055] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the protection scope of the present invention.
[0056] Multi - target tracking refers to a technology that continuously tracks multiple targets to be tracked in a video. In practical applications, especially in surveillance videos, video frames are often blurred due to reasons such as target movement, or the image quality of video frames is poor due to factors such as lighting or occlusion. Feature information extracted from video frames with poor image quality may have errors. Therefore, when performing feature comparison, there may be cases of incorrect feature comparison, resulting in the exchange of tracking IDs of targets, which affects the accuracy of multi - target tracking.
[0057] To address the above problems, an embodiment of the present invention provides a multi - target tracking method. In this method, the clarity corresponding to the historical detection box corresponding to the predicted tracking box can be determined. Since the clarity corresponding to the historical detection box can be used as the weight of the historical feature information to determine the first similarity between each detection box and the predicted tracking box, the interference of the feature information extracted from video frames with poor image quality on the matching result can be reduced, thereby improving the accuracy of multi - target tracking.
[0058] The following combines Figures 1-3 Describe the multi - target tracking method of the present invention. The multi - target tracking method provided by the embodiment of the present invention is applicable to target tracking in various application scenarios. For example, it can be applied to surveillance scenarios, autonomous driving, or the security field, etc. Among them, the targets can include pedestrians, vehicles, drones, or other objects, etc. The execution subject of this method can be an electronic device such as a mobile phone, a camera device, a terminal device, a computer, a server, a server cluster, or a specially designed multi - target tracking device, or it can be a multi - target tracking device set in the electronic device. The multi - target tracking device can be implemented through software, hardware, or a combination of both.
[0059] Figure 1It is a schematic flowchart of the multi-object tracking method provided by an embodiment of the present invention. As Figure 1 shown, the method includes:
[0060] Step 101: Determine at least one predicted tracking box in the current video frame and the detection box corresponding to each target.
[0061] Among them, the historical feature information corresponding to the predicted tracking box is the feature information of the historical detection box in the first preset number of historical video frames before the current video frame, and the target represented by the historical detection box is the same as the target represented by the predicted tracking box.
[0062] Specifically, the tracker includes a predicted tracking box in the current video frame. This predicted tracking box can be understood as the predicted position of the target in the current video frame, which is predicted by the tracker based on the motion information of the target represented by the historical detection box in the historical video frame. Therefore, the target represented by the historical detection box is the same as the target represented by the predicted tracking box, and the tracker includes the historical feature information of the target represented by the historical detection box corresponding to the predicted tracking box in each historical video frame. Among them, the historical feature information may include information such as motion information, position information, and appearance features extracted based on the historical detection box.
[0063] It should be understood that in each historical video frame, only one historical detection box corresponds to the predicted tracking box, that is, only the target represented by one historical detection box is the same as the target represented by the predicted tracking box.
[0064] In addition, the first preset number can be set according to experience or actual situation. For example, it can be set to 10 frames or 12 frames, etc. For the specific value of the first preset number, the embodiments of the present invention do not limit it here.
[0065] In addition, a pre-trained object detection model can be used to detect the detection boxes of each target object in the current video frame. Exemplarily, in the embodiments of the present invention, the yolov8 algorithm can be used to detect each target, so as to obtain the detection boxes corresponding to each target. Among them, yolov8 is the current state-of-the-art (SOTA) model in detection algorithms, supporting image detection and segmentation tasks. It uses effective training strategies of the previous yolov5, such as data augmentation means like mosaic and mixup, uses a network architecture based on the improved Cross Stage Partial Network (CSPNet), which can effectively reduce the computational amount and memory usage of the network, and uses a cross-stage hierarchical structure to reduce the redundancy of gradient backpropagation. In addition, yolov8 uses an anchor-free detection head and Generalized Focal Loss (GFL) to further improve the accuracy of the detection box when the target boundary is blurred. At the same time, for the segmentation task, yolov8 includes a segmentation head, which can not only output the detection box and category of the target, but also output the area occupied by the target, that is, perform instance segmentation on the current video frame image to segment the area where the target is located.
[0066] It should be noted that the above yolov8 is only for illustration, and other detection algorithms and segmentation models can also be used in the embodiments of the present invention to detect the detection boxes corresponding to each target and obtain the areas where each target is located. The embodiments of the present invention do not limit the object detection algorithm and segmentation model.
[0067] Step 102: For each predicted tracking box, obtain the clarity corresponding to each historical detection box associated with the predicted tracking box.
[0068] In this step, for each predicted tracking box in the current video frame, there are historical detection boxes corresponding to the predicted tracking box in the historical video frames, where the targets represented by these historical detection boxes and the target represented by the predicted tracking box are the same target.
[0069] After determining the historical detection boxes of each frame of historical video frames, the clarity of each historical detection box can be determined, where the clarity can represent the image quality in each historical detection box. Exemplarily, the clarity can be quantified and represented by a clarity score.
[0070] Step 103: Based on the clarity corresponding to each historical detection box, the historical feature information corresponding to each historical detection box, and the feature information corresponding to each detection box, determine the first similarity between each detection box and the predicted tracking box.
[0071] In this step, each detection box and the region where the target is located in the detection box can be input into the feature extraction model, so as to extract the feature information corresponding to each detection box. Among them, the feature information corresponding to each detection box can also be understood as the feature information corresponding to the region where the target is located in each detection box, that is, the feature information of the target in each detection box.
[0072] Exemplarily, when performing target tracking on the predicted tracking boxes in historical video frames before, the historical detection boxes in the historical video frames and the historical regions where the targets are located in the historical detection boxes can be input into the feature extraction model, so as to obtain the historical feature information corresponding to the historical detection boxes, and store the historical feature information corresponding to the historical detection boxes in the feature information library. In this way, when performing target tracking currently, the historical feature information corresponding to each historical detection box can be obtained from the feature information library. Of course, when performing target tracking currently, when inputting each detection box and the region where the target is located in the detection box into the feature extraction model, the historical detection boxes in the historical video frames and the historical regions where the targets are located in the historical detection boxes can also be input into the feature extraction model, so as to extract the feature information corresponding to each detection box and the historical feature information corresponding to each historical detection box.
[0073] Further, the feature weight can be determined based on the clarity corresponding to the historical detection box, and based on the feature weight, the historical feature information corresponding to each historical detection box, and the feature information corresponding to each detection box, the first similarity between each detection box and the predicted tracking box can be determined. The first similarity can represent the matching degree between each detection box and the predicted tracking box. The higher the first similarity, the higher the matching degree between the corresponding detection box and the predicted tracking box.
[0074] Exemplarily, the higher the clarity, the larger the feature weight can be set, and the lower the clarity, the smaller the feature weight can be set. By controlling the participation degree of images with different clarities in feature matching, the influence and interference on the first similarity when the image clarity is poor can be reduced.
[0075] Step 104: Track each target based on all the first similarities corresponding to each predicted tracking box.
[0076] Specifically, for each predicted tracking box, the first similarity between each detection box and the predicted tracking box can be determined based on the above steps, so as to determine whether there is a detection box in at least one detection box in the current video frame that matches the predicted tracking box based on all the first similarities corresponding to the predicted tracking box, and complete the tracking of the target represented by the predicted tracking box. In the case where there are multiple predicted tracking boxes in the current video frame, for each predicted tracking box, the target can be tracked according to the above method, so as to achieve the purpose of multi-target tracking.
[0077] Exemplarily, for the first similarity matrix formed by the first similarities calculated for each predicted tracking box and each detection box, the Hungarian matching algorithm can be used to obtain the matching relationship between each predicted tracking box and the detection box. For the detection box with a successful match, the coordinates and feature information of the detection box with a successful match are placed into the predicted tracking box, and the tracker is updated. For the detection box without a successful match, it is regarded as a newly emerged target, and a new tracker is created for continuous tracking; for the predicted tracking box without a successful match, according to its status, the corresponding tracker waits for the next round of matching or is deleted.
[0078] The multi-object tracking method provided by the embodiments of the present invention determines at least one predicted tracking box in the current video frame and the detection boxes corresponding to each target. The historical feature information corresponding to the predicted tracking box is the feature information of the historical detection boxes in the first preset number of historical video frames before the current video frame, and the targets represented by the historical detection boxes are the same as the targets represented by the predicted tracking box. For each predicted tracking box, the clarity corresponding to each historical detection box associated with the predicted tracking box can be obtained, and based on the clarity corresponding to each historical detection box, the historical feature information corresponding to each historical detection box, and the feature information corresponding to each detection box, the first similarity between each detection box and the predicted tracking box is determined, and then based on all the first similarities corresponding to each predicted tracking box, each target is tracked. Since the clarity corresponding to each historical detection box can be used as the weight of the historical feature information to determine the first similarity between each detection box and the predicted tracking box, the influence of the image quality on the first similarity can be reduced, thereby reducing the probability of the tracking ID interchange of the target and improving the accuracy of multi-object tracking.
[0079] Figure 2 This is the second flowchart of the multi-object tracking method provided by the embodiments of the present invention. As Figure 2 shown, this embodiment details the specific implementation manners of steps 102 and 103 in the Figure 1 embodiment shown. As Figure 2 shown, the method includes:
[0080] Step 201: Determine at least one predicted tracking box in the current video frame and the detection boxes corresponding to each target.
[0081] Among them, the historical feature information corresponding to the predicted tracking box is the feature information of the historical detection boxes in the first preset number of historical video frames before the current video frame, and the targets represented by the historical detection boxes are the same as the targets represented by the predicted tracking box.
[0082] Step 202: Determine the historical regions where the targets corresponding to the historical detection boxes are located.
[0083] Specifically, for each historical video frame, the historical video frame can be input into the target detection model to obtain the historical region in the historical detection box output by the target detection model. For example, it can be input into yolov8 described in step 101. The historical region can be understood as the region where the target is located in the historical detection box, or it can also be understood as the region that does not contain background information.
[0084] Step 203: Determine the clarity of the historical region.
[0085] Exemplarily, for each historical detection box, the Tenengrad algorithm can be used to determine the clarity of the historical region in the historical detection box, such as the clarity score. Generally, the clearer the image, the sharper the edges and details, so the image has a larger gradient value. The calculation method of the Tenengrad algorithm is shown in formulas (1) and (2):
[0086]
[0087] S = ∑ x,y T(x,y) * F(x,y) (2)
[0088] where T(x,y) represents the clarity of the pixel point (x,y) in the current video frame, I(x,y) represents the pixel value of the pixel point (x,y), G x represents the horizontal gradient template, and G y represents the vertical gradient template, and F(x,y) represents the target instance segmentation result output by the target detection model. If the pixel point (x,y) is within the historical region of the historical detection box, then F s (x,y) = 1. If the pixel point (x,y) is not within the historical region of the historical detection box, then F s (x,y) = 0. S represents the clarity score of the historical region.
[0089] In this step, in order to avoid background blurring and interfering with the clarity judgment, the instance segmentation result in the current video frame is used as a mask, and the clarity can be calculated only for the region where the target is located. For example, the clarity can be calculated only for the pedestrian region. Since the determined clarity does not contain background information, the interference of background information on the image quality of the historical video frame can be excluded.
[0090] Step 204: For each historical detection box, based on the clarity of the historical region corresponding to the historical detection box, determine the feature weight of the historical detection box.
[0091] In this step, after determining the clarity corresponding to each historical region, the feature weights corresponding to each clarity can be determined based on a preset mapping relationship. Exemplarily, the Softmax function can be used to convert each clarity score into a feature weight. The feature weight can be used to represent the proportion of feature information in determining the first similarity. It should be understood that the higher the clarity, the greater the feature weight, and the lower the clarity, the smaller the feature weight.
[0092] Step 205: For each detection box, determine the second similarity between the feature information corresponding to the detection box and the historical feature information of each historical detection box.
[0093] In this step, the current video frame includes at least one detection box, and each detection box corresponds to a target. The detection box represents the actual position of the corresponding target in the current video frame. Since the current video frame may include multiple detection boxes and it is necessary to determine the detection box that matches the predicted tracking box from multiple detection boxes, the second similarity between the feature information corresponding to each detection box and the historical feature information of each historical detection box will be determined respectively. The second similarity represents the matching degree between the feature information corresponding to the detection box and the historical feature information corresponding to the historical detection box. The higher the second similarity, the higher the matching degree between the feature information of the corresponding detection box and the historical feature information of the historical detection box.
[0094] Exemplarily, in practical applications, it often occurs that a partial area of a target is occluded. When a partial area of a target is occluded, if the feature information of the entire target is still extracted and the similarity is calculated using the feature information of the entire target for target tracking, it will cause a decrease in similarity and a problem of tracking failure. To solve this problem, in the embodiments of the present invention, when determining the feature information corresponding to a detection box, the area occupied by the target corresponding to each detection box in the current video frame can be determined, the area is input into a feature extraction model, the feature extraction model divides the area into at least two sub-areas based on a preset ratio, and the sub-feature information of each sub-area is extracted by the feature extraction model respectively, and the feature information corresponding to the detection box is determined based on the sub-feature information.
[0095] Specifically, after inputting the current video frame into the target detection model, the detection boxes where the targets in the current video frame are located can be detected. Using the instance segmentation algorithm, the segmentation result of the target can be obtained, that is, the area where the target corresponding to each detection box is located.
[0096] Figure 3 Schematic diagram of region division provided by the embodiments of the present invention, as Figure 3As shown, the obtained region can be input into the feature extraction model. In the feature extraction model, the region is divided into at least two sub-regions based on a preset ratio. For example, taking the target as a pedestrian, the region of the target can be divided into a pedestrian head region, an upper torso region, and a lower torso region. In another implementation, the detection frame and the region can also be input into the feature extraction model. After dividing the detection frame into at least two parts according to the preset ratio, the obtained parts are then intersected with the region, so as to obtain the sub-feature information of each more refined sub-region of the target. Among them, the preset ratio can be set according to experience, such as it can be set to 1:2:3, etc.
[0097] For each sub-region obtained after division, the sub-feature information in each sub-region can be extracted through the feature extraction model, and at least two obtained sub-feature information are spliced, so as to obtain the feature information corresponding to the detection frame, that is, the feature information of the target in the detection frame.
[0098] In this embodiment, the region is divided into at least two sub-regions based on a preset ratio through the feature extraction model, and the sub-feature information of each sub-region is extracted, so as to splice the obtained sub-regions to obtain the feature information corresponding to the detection frame. The above feature extraction model is more robust. In addition, by adopting the above method, the interference caused by objects in non-target regions in the detection frame can be reduced. Moreover, more refined features are extracted for each sub-region respectively, so as to be more adaptable to the problem of change in feature similarity caused by pose change during the tracking process, and effectively reduce the interference caused by background change to the features of the same target.
[0099] Step 206: Determine the first similarity between the detection frame and the predicted tracking frame based on the second similarities and feature weights corresponding to each historical detection frame.
[0100] In this step, for each historical detection frame, after determining the feature weight of the historical detection frame and the second similarity between the historical detection frame and the detection frame, the second similarity and the feature weight can be weighted to obtain the weighted second similarity of the historical detection frame.
[0101] Based on the determined weighted second similarities of each historical detection frame, determine the maximum weighted second similarity among all the weighted second similarities, and use the maximum weighted second similarity as the first similarity between the detection frame and the predicted tracking frame.
[0102] For example, during the movement of the target, the pose usually may change, such as turning to walk, turning around, or squatting, etc. Therefore, in order to ensure the accuracy of the tracking process, in the embodiments of the present invention, a feature storage sequence is created for each tracker, and the feature storage sequence includes historical feature information. For example, assume that the current video frame is the t-th frame. Denote the j-th predicted tracking box of the t-th video frame, and the feature storage sequence where N is the maximum value of the preset number of stored feature information Denote the feature storage sequence corresponding to the j-th predicted tracking box of the t-th video frame, b n Denote the historical detection box in the n-th historical video frame that matches the predicted tracking box, f(b n ) represents the historical feature information of the historical detection box b n . Additionally Denote the i-th detection box in the t-th video frame Denote the detection box 's feature information. When determining the first similarity between the detection box and the predicted tracking box , the second similarity between the feature information in the detection box and each historical feature information in will be determined first. For example, the cosine similarity can be determined
[0103] . Additionally, the feature weights of each historical detection box can be determined based on the clarity of the historical regions corresponding to each historical detection box corresponding to the predicted tracking box. For example, the clarity of the historical regions corresponding to each historical detection box corresponding to the predicted tracking box can be expressed by formula (3):
[0104]
[0105] where represents the clarity of the predicted tracking box , represents the clarity of the historical region in the historical detection box corresponding to the predicted tracking box in the m-th historical video frame, represents the clarity of the historical region in the historical detection box corresponding to the predicted tracking box in the (m + 1)-th historical video frame, represents the clarity of the historical region in the historical detection box corresponding to the predicted tracking box in the n-th historical video frame
[0106] . Additionally, the above clarity can be converted into the feature weights of each historical detection box according to formula (4):
[0107]
[0108] where represents the feature weight of the predicted tracking box , represents the historical video frame of the m-th frame that matches the predicted tracking box The feature weight of the corresponding historical detection box represents the feature weight of the historical detection box corresponding to the predicted tracking box in the (m + 1)-th historical video frame The feature weight of the corresponding historical detection box represents the feature weight of the historical detection box corresponding to the predicted tracking box in the n-th historical video frame The feature weight of the corresponding historical detection box
[0109] Furthermore, the second similarity and the feature weight corresponding to each historical detection box can be weighted to obtain the weighted similarity corresponding to each historical detection box, and then the maximum value among these weighted similarities is determined as the first similarity between the detection box and the predicted tracking box. Exemplarily, the first similarity between the detection box and the predicted tracking box can be determined according to formula (5):
[0110]
[0111] wherein represents the first similarity between the detection box and the predicted tracking box , y represents the weighted similarity of the historical detection box corresponding to the predicted tracking box in the n-th historical video frame. It should be understood that the historical detection box corresponding to the predicted tracking box in the n-th historical video frame and the historical detection box b are the same historical detection box. In addition, in the above formula (5), n traverses from 1 to N, so that the weighted similarities corresponding to each of the N historical detection boxes can be determined, and the maximum weighted similarity among the N weighted similarities is determined as the first similarity between the detection box and the predicted tracking box n is the same historical detection box. In addition, in the above formula (5), n traverses from 1 to N, so that the weighted similarities corresponding to each of the N historical detection boxes can be determined, and the maximum weighted similarity among the N weighted similarities is determined as the first similarity between the detection box and the predicted tracking box
[0112] After determining the first similarity between each detection box and the predicted tracking box, for the first similarity matrix of P detection boxes and Q predicted tracking boxes, it can be expressed as shown in formulas (6) and (7):
[0113] cosmatrix=(c ij )∈A P×Q (6)
[0114]
[0115] wherein, cosmatrix represents the first similarity matrix A
[0116] Step 207: Track each target based on all the first similarities corresponding to each predicted tracking box
[0117] Exemplarily, for each predicted tracking box, a first similarity between the predicted tracking box and each detection box is determined. Therefore, the first similarity can be compared with a similarity threshold. If there is a first similarity greater than the similarity threshold, it means that the detection box corresponding to the first similarity greater than the similarity threshold matches the predicted tracking box, that is, the target in the detection box and the target represented by the predicted tracking box are the same target, and thus the target can be tracked based on the detection box.
[0118] In this embodiment, the feature weight of the historical detection box can be determined based on the clarity of the historical area corresponding to the historical detection box, and thus the first similarity between the detection box and the predicted tracking box can be determined based on the feature weight. Since the clarity of the historical area part is used to calculate the first similarity, and clearer targets are given higher weights, it helps to perform high-quality feature matching, thereby effectively reducing the problem of feature matching errors caused by motion blur, occlusion, etc.
[0119] Exemplarily, based on the above embodiments, when tracking each target based on all the first similarities corresponding to each predicted tracking box, it can be performed in the following manner:
[0120] For each predicted tracking box, in the case where there is a first similarity greater than the similarity threshold among all the first similarities corresponding to the predicted tracking box, the detection box corresponding to the maximum first similarity among the first similarities greater than the similarity threshold is determined as the detection box matching the predicted tracking box, and the target corresponding to the predicted tracking box is tracked based on the matching detection box.
[0121] Specifically, when performing target tracking, a similarity threshold δ is usually set. For a certain predicted tracking box, if there is a first similarity greater than the similarity threshold δ among all the first similarities corresponding to the predicted tracking box, the detection box corresponding to the maximum first similarity among all the first similarities greater than the similarity threshold can be determined as the detection box matching the predicted tracking box, indicating that the matching detection box and the predicted tracking box are the same target. At this time, the position of the predicted tracking box can be updated based on the successfully matched detection box to track the target.
[0122] In this embodiment, the detection box corresponding to the maximum first similarity among the first similarities greater than the similarity threshold can be determined as the detection box matching the predicted tracking box, making the determined successfully matched detection box more accurate.
[0123] Exemplarily, in the case where there is no first similarity greater than the similarity threshold among all the first similarities corresponding to the predicted tracking box, the matching between each detection box and the predicted tracking box can be predicted based on the position information of each detection box and the position information of the predicted tracking box. In the case where it is predicted that the target detection box and the predicted tracking box match, the first size information of the target detection box and the second size information of the target historical detection box corresponding to the predicted tracking box in the previous historical video frame are obtained, and based on the first size information and the second size information, in the case where it is determined that the change range between the target detection box and the target historical detection box is greater than the preset range, it is determined that the target detection box and the predicted tracking box do not match.
[0124] Specifically, usually in the case where there is no first similarity greater than the similarity threshold among all the first similarities corresponding to the predicted tracking box, it is also determined that there is no detection box matching the predicted tracking box in the current video frame. However, the similarity threshold δ is usually set according to experience, and it is difficult to set a more appropriate value. When the similarity threshold δ is too small, it will cause incorrect ID tracking of the target; when the similarity threshold δ is too large, the tracking often interrupts, resulting in more IDs of the target. To solve the above problems and improve the tracking accuracy, the method of joint tracking is introduced in the embodiments of the present invention. In the case where there is no first similarity greater than the similarity threshold, tracking can be performed based on the spatial distance. Since the influence of feature tracking on the tracking effect in a dense scene is much smaller than that of spatial distance tracking, and the sensitivity of matching can be controlled by the similarity threshold δ, feature tracking is first performed based on the first similarity determined by using features. In the case where feature tracking fails, spatial distance tracking is then performed.
[0125] When performing feature tracking, a relatively large similarity threshold δ can be set, for example, it can be set to 0.7 or other values, so that the feature matching can be set more strictly, and the accuracy of feature matching can be guaranteed. For the detection box and the predicted tracking box with low first similarity, spatial distance matching is used for tracking.
[0126] When performing spatial distance matching, the center distance between each detection box and the predicted tracking box can be determined based on the position information of each detection box and the position information of the predicted tracking box, so as to construct a matching matrix, and based on this matching matrix, it is predicted whether each detection box and the predicted tracking box match. In a dense scene, since the occlusion between targets is very serious, the detection box and the predicted tracking box with the shortest center distance are not necessarily the same tracking target. Therefore, predicting only based on the above matching matrix will cause more incorrect matches.
[0127] To reduce the probability of the above-mentioned incorrect matching, in the embodiments of the present invention, it is possible to further determine whether the target detection box and the predicted tracking box are truly matched based on the change amplitude of the size of the target between two consecutive frames. It should be understood that in practical applications, for the same target, the change amplitude of its size in two consecutive frames is usually within a certain range. Therefore, when the target detection box and the predicted tracking box are matched in terms of spatial distance, the first size information of the target detection box and the second size information of the target historical detection box corresponding to the predicted tracking box in the previous historical video frame can be obtained. Then, based on the first size information and the second size information, it is determined whether the change amplitude between the target detection box and the target historical detection box is greater than the preset amplitude. When the change amplitude is greater than the preset amplitude, it indicates that the target corresponding to the target detection box and the target corresponding to the predicted tracking box may not be the same target, that is, there is an error in the matching of the target detection box and the predicted tracking box predicted based on the center distance.
[0128] Among them, the first size information includes the width, height or area of the target detection box, and the second size information includes the width, height or area of the target historical detection box. The target historical detection box is the target historical detection box that has been successfully matched with the predicted tracking box in the previous historical video frame of the current video frame.
[0129] For example, assume that the current video frame is the t-th frame image, and the width, height, and area of the target detection box that is successfully matched with the predicted tracking box in the t-th frame image are w t 、h t 、A t , and the width, height, and area of the target historical detection box in the previous historical video frame are w t-1 、h t-1 、A t-1 . If the width threshold is set to v w 、the height threshold is set to v h 、and the area threshold is set to δ A . Since the change amplitude of the target size between two consecutive frames is often within a certain limit, if any one of the following formulas (8) - (10) is satisfied, it is determined that the change amplitude between the target detection box and the target historical detection box is greater than the preset amplitude, and thus it can be further determined that the target detection box and the predicted tracking box do not match.
[0130]
[0131]
[0132]
[0133] In this embodiment, when the predicted target detection box and the predicted tracking box match, by obtaining the first size information of the target detection box and the second size information of the target historical detection box corresponding to the predicted tracking box in the previous historical video frame, and based on the first size information and the second size information, when it is determined that the change amplitude between the target detection box and the target historical detection box is greater than the preset amplitude, it is determined that the target detection box and the predicted tracking box do not match. Since the change amplitude is used to further verify whether the target detection box and the predicted tracking box match, the tracking error problem caused by the close distance between the target centers in a dense scene can be effectively filtered out.
[0134] Exemplarily, on the basis of the above embodiments, when determining the second similarity between the feature information corresponding to the detection box and the historical feature information of each historical detection box, the first region including the predicted tracking box can be determined based on the first position in the predicted tracking box, and the area of the first region is larger than the area of the predicted tracking box. When it is determined that the detection box is within the first region, the second similarity between the detection box and the historical feature information of each historical detection box is determined, and when it is determined that the detection box is not within the first region, the second similarity between the detection box and the historical feature information of each historical detection box is determined as the preset similarity.
[0135] Specifically, assume that during the process of tracking the target in the current video frame, the number of all detection boxes in the current video frame is P, the number of predicted tracking boxes is Q, and at most N features are stored in one predicted tracking box. Therefore, the computational complexity of determining the first similarity matrix is P * Q * N. When the number of targets in the scene is large, this computational complexity is very large, resulting in a long time consumption and poor real-time performance of target tracking. To solve this problem, in the embodiments of the present invention, since the moving distance between adjacent frames is limited when the target moves, for the predicted tracking box of a target, it is only necessary to calculate the first similarity with the detection boxes within a certain area around it, and the similarity outside the area is directly set to 0.
[0136] In practical applications, the first region can be determined based on the first position in the predicted tracking box, where the first region includes the predicted tracking box and the area of the first region is larger than the area of the predicted tracking box. For example, the center of the predicted tracking box can be determined as the first position, and taking this first position as the center of the first region, the width and height of the first region are set as the width and height of the predicted tracking box multiplied by a preset multiple, such as multiplying the width and height of the predicted tracking box by 2.5 as the width and height of the first region.
[0137] Among them, the first region can be understood as the region where the target corresponding to the predicted tracking box may be. Therefore, if the detection box is within the first region, it indicates that the target corresponding to the detection box may be the target corresponding to the predicted tracking box. Therefore, the second similarity between the feature information of the target in the detection box and each historical feature information will be calculated. If the detection box is not within the first region, it indicates that the probability that the target corresponding to the detection box and the target corresponding to the predicted tracking box are the same target is very small. Therefore, the second similarity between the feature information of the target corresponding to the detection box and each historical feature information can be directly set to a preset similarity. Among them, the preset similarity can be, for example, 0, or a value close to 0.
[0138] In the above method, since the second similarity between the feature information corresponding to the detection box and each historical feature information is determined only when the detection box is within the first region, the amount of calculation can be reduced, and the real-time performance of target tracking is improved. Moreover, the above method can reduce the phenomenon of obvious tracking errors caused by targets that are far away being considered the same target due to feature confusion, and improve the accuracy of target tracking.
[0139] Exemplarily, on the basis of the above embodiments, there often occurs a situation where a certain target is still moving in the picture but the tracking fails. This is usually because there is a deviation in the coordinates predicted by the predicted tracking box. Therefore, it is necessary to increase the first region to correct the prediction error and make the tracking more accurate. Specifically, when the predicted tracking box fails to match in the first second preset number of historical video frames of the current video frame, based on the first position and a preset value, the first region is expanded to obtain a second region. The preset value is related to the second preset number, and the second region is determined as the new first region, so as to re-determine whether the detection box is within the new first region. If it is within the new first region, the second similarity between the feature information in the detection box and each historical feature information will be determined. If it is not within the new first region, the second similarity between the feature information in the detection box and each historical feature information will be set to the preset similarity.
[0140] Among them, the second preset number k can be set according to experience or actual situation. For example, it can be set to 5 frames or 6 frames, etc. When the predicted tracking box fails to match in k consecutive historical video frames, that is, the same target is not matched, in order to avoid the situation of losing track of the target while the target is still moving in the video frame, the first region can be expanded based on the first position and a preset value, so as to obtain a second region with a larger area. Exemplarily, the first position can be the center point of the predicted tracking box, and the preset value can be 0.5k. That is, the center point of the predicted tracking box is used as the center point of the second region, and the width and height of the first region are respectively increased by 0.5k as the width and height of the second region. It should be noted that the width and height of the second region do not exceed the width and height of the current video frame at most.
[0141] For example, for a predicted tracking box if the width and height of the corresponding first region are increased by 2.5 times based on the width and height of the predicted tracking box, and no corresponding detection box is matched in k consecutive historical video frames, the corresponding second region can be determined as shown in formula (11):
[0142]
[0143] where x and y respectively represent the abscissa and ordinate of the center point of the predicted tracking box, x s and y s respectively represent the abscissa and ordinate of the center point of the second region, w represents the width of the predicted tracking box, h represents the height of the predicted tracking box, w s represents the width of the second region, h s represents the height of the second region.
[0144] After determining the second region, the second region is determined as the new first region, so as to re-determine whether the detection box is within the new first region. If it is within the new first region, the second similarity between the feature information in the detection box and each historical feature information is determined. If it is not within the new first region, the second similarity between the feature information in the detection box and each historical feature information is set to a preset similarity, such as set to 0, etc.
[0145] In this embodiment, when the predicted tracking box fails to match in the second preset number of historical video frames before the current video frame, the first region can be expanded based on the first position and the preset value, that is, the region is expanded to search for the target belonging to the same target as the predicted tracking box, so that the situation of target tracking failure can be reduced and the accuracy of target tracking can be improved. In addition, assume that P = 20, Q = 20, and N = 10 in the current video frame. According to the method of the prior art, 20×20×10 = 4000 times of matching are required. After adopting the above method of dividing the tracking region, assume that there are 4 detection boxes in the first region or the second region corresponding to a predicted tracking box, then 20×4×10 = 800 times of matching are required. Therefore, the calculation amount is only 20% of the prior art, which can greatly reduce the calculation amount, reduce the time consumption, and improve the real-time performance of target tracking.
[0146] Exemplarily, based on the above embodiments, each piece of historical feature information corresponding to the predicted tracking box is stored in the feature storage sequence. After the current matching is completed, if the matching is successful, the feature information of the target in the detected box with successful matching also needs to be stored in the feature storage sequence as the basis for subsequent target tracking. When storing the feature information, for each predicted tracking box, in the case that there is a first similarity greater than the similarity threshold among all the first similarities corresponding to the predicted tracking box, determine the minimum second similarity among the second similarities corresponding to the historical feature information of each historical detected box included in the feature storage sequence, delete the historical feature information corresponding to the minimum second similarity, and add the feature information of the detected box corresponding to the first similarity greater than the similarity threshold to the feature storage sequence.
[0147] Specifically, although during the tracking of the target, the target movement is often slow and regular, there may still be a situation where the front and rear features change greatly due to sudden changes in the target movement trajectory. Therefore, in order to improve the accuracy of feature matching and reduce the number of feature storage sequences to reduce the first similarity calculation amount, in the case that there is a first similarity greater than the similarity threshold among all the first similarities corresponding to the predicted tracking box, that is, in the case that there is a detected box with successful matching, the feature storage sequence of the predicted tracking box will be updated based on the feature information of the detected box with successful matching. When updating, the feature information of the target in the detected box with successful matching will be used to replace the historical feature information with the smallest second similarity in the feature storage sequence, rather than replacing the earliest saved historical feature information in the form of a queue.
[0148] Through the above method, since the historical feature information corresponding to the minimum second similarity in the feature storage sequence is deleted and the feature information of the detected box corresponding to the first similarity greater than the similarity threshold is added to the feature storage sequence, it is possible to reduce the number of historical feature information in the feature storage sequence while maintaining the tracking effect, thereby reducing the calculation amount and improving the real-time performance of target tracking.
[0149] Next, the multi-target tracking device provided by the present invention will be described. The multi-target tracking device described below can be correspondingly referred to the multi-target tracking method described above.
[0150] Figure 4 is a schematic structural diagram of the multi-target tracking device provided by an embodiment of the present invention. Refer to Figure 4 As shown, the multi-target tracking device 400 includes:
[0151] A determination module 401 is configured to determine at least one predicted tracking box in the current video frame and detection boxes corresponding to each target; the historical feature information corresponding to the predicted tracking box is the historical feature information of historical detection boxes in the first preset number of historical video frames before the current video frame, and the targets represented by the historical detection boxes are the same as the targets represented by the predicted tracking box.
[0152] An acquisition module 402 is configured to, for each of the predicted tracking boxes, acquire the clarity corresponding to each of the historical detection boxes associated with the predicted tracking box.
[0153] The determination module 401 is further configured to determine a first similarity between each of the detection boxes and the predicted tracking box based on the clarity corresponding to each of the historical detection boxes, the historical feature information corresponding to each of the historical detection boxes, and the feature information corresponding to each of the detection boxes.
[0154] A tracking module 403 is configured to track each of the targets based on all the first similarities corresponding to each of the predicted tracking boxes.
[0155] In an exemplary embodiment, the acquisition module 402 is specifically configured to:
[0156] Determine the historical region where the target corresponding to each of the historical detection boxes is located;
[0157] Determine the clarity of the historical region;
[0158] The determination module 401 is specifically configured to:
[0159] For each of the historical detection boxes, determine the feature weight of the historical detection box based on the clarity of the historical region corresponding to the historical detection box;
[0160] For each of the detection boxes, determine a second similarity between the feature information corresponding to the detection box and the historical feature information of each of the historical detection boxes;
[0161] Based on the second similarity and the feature weight corresponding to each of the historical detection boxes, determine the first similarity between the detection box and the predicted tracking box.
[0162] In an exemplary embodiment, the apparatus further includes a division module and an extraction module, where:
[0163] The determination module 401 is further configured to determine the region occupied by the target corresponding to each of the detection boxes in the current video frame;
[0164] The division module is configured to input the region into a feature extraction model, and divide the region into at least two sub-regions by the feature extraction model based on a preset ratio;
[0165] An extraction module, configured to extract sub - feature information of each of the sub - regions respectively through the feature extraction model;
[0166] A determination module 401, configured to determine the feature information corresponding to the detection box based on the sub - feature information of each of the sub - regions.
[0167] In an exemplary embodiment, the tracking module 403 is specifically configured to:
[0168] For each of the predicted tracking boxes, in the case where there is a first similarity greater than the similarity threshold among all the first similarities corresponding to the predicted tracking box, determine the detection box corresponding to the maximum first similarity among the first similarities greater than the similarity threshold as the detection box matching the predicted tracking box, and track the target corresponding to the predicted tracking box based on the matching detection box.
[0169] In an exemplary embodiment, the apparatus further includes a prediction module, where:
[0170] The prediction module is configured to, in the case where there is no first similarity greater than the similarity threshold among all the first similarities corresponding to the predicted tracking box, predict whether each of the detection boxes and the predicted tracking box match based on the position information of each of the detection boxes and the position information of the predicted tracking box;
[0171] An acquisition module 402, configured to, in the case where it is predicted that the target detection box and the predicted tracking box match, acquire the first size information of the target detection box and the second size information of the target historical detection box corresponding to the predicted tracking box in the previous historical video frame;
[0172] A determination module, configured to, based on the first size information and the second size information, determine that the target detection box and the predicted tracking box do not match in the case where it is determined that the change amplitude between the target detection box and the target historical detection box is greater than a preset amplitude.
[0173] In an exemplary embodiment, the determination module 401 is specifically configured to:
[0174] Based on the first position in the predicted tracking box, determine a first region including the predicted tracking box, and the area of the first region is greater than the area of the predicted tracking box;
[0175] In the case where it is determined that the detection box is within the first region, determine the second similarity between the feature information corresponding to the detection box and each of the historical feature information;
[0176] In the case where it is determined that the detection box is not within the first region, determine the second similarity between the feature information corresponding to the detection box and each of the historical feature information as a preset similarity.
[0177] In an exemplary embodiment, the apparatus further includes: an expansion module and a processing module, where:
[0178] The expansion module is configured to, when the prediction tracking boxes in the second preset number of historical video frames before the current video frame all fail to match, expand the first region based on the first position and a preset value to obtain a second region, where the preset value is related to the second preset number;
[0179] The processing module is configured to determine the second region as the new first region, and return to execute the step of determining the second similarity between the detection box and each piece of the historical feature information when it is determined that the detection box is within the first region.
[0180] In an exemplary embodiment, each piece of historical feature information corresponding to the prediction tracking box is stored in a feature storage sequence;
[0181] The determination module 401 is further configured to, for each prediction tracking box, when there is a first similarity greater than the similarity threshold among all the first similarities corresponding to the prediction tracking box, determine the minimum second similarity among the second similarities corresponding to the historical feature information of each historical detection box included in the feature storage sequence;
[0182] The deletion module is configured to delete the historical feature information corresponding to the minimum second similarity;
[0183] The addition module is configured to add the feature information of the detection box corresponding to the first similarity greater than the similarity threshold to the feature storage sequence.
[0184] The apparatus in this embodiment can be used to execute the method in any one of the embodiments on the multi-object tracking method side. The specific implementation process and technical effects are similar to those in the embodiments on the multi-object tracking method side. For details, reference can be made to the detailed introduction in the embodiments on the multi-object tracking method side, which will not be elaborated here.
[0185] Figure 5 An exemplary structural diagram of an electronic device is shown in Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute a multi-target tracking method, which includes: determining at least one predicted tracking box in the current video frame and detection boxes corresponding to each target; the historical feature information corresponding to the predicted tracking box is the historical feature information of the historical detection boxes in the first preset number of historical video frames before the current video frame, and the target represented by the historical detection box is the same as the target represented by the predicted tracking box; for each of the predicted tracking boxes, obtaining the clarity corresponding to each of the historical detection boxes associated with the predicted tracking box; based on the clarity corresponding to each of the historical detection boxes, the historical feature information corresponding to each of the historical detection boxes, and the feature information corresponding to each of the detection boxes, determining the first similarity between each of the detection boxes and the predicted tracking box; based on all the first similarities corresponding to each of the predicted tracking boxes, tracking each of the targets.
[0186] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0187] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-object tracking method provided by the above-mentioned various methods. The method includes: determining at least one predicted tracking box in the current video frame and detection boxes corresponding to each target; the historical feature information corresponding to the predicted tracking box is the historical feature information of historical detection boxes in the first preset number of historical video frames before the current video frame, and the target represented by the historical detection box is the same as the target represented by the predicted tracking box; for each of the predicted tracking boxes, obtaining the clarity corresponding to each of the historical detection boxes associated with the predicted tracking box; determining a first similarity between each of the detection boxes and the predicted tracking box based on the clarity corresponding to each of the historical detection boxes, the historical feature information corresponding to each of the historical detection boxes, and the feature information corresponding to each of the detection boxes; and tracking each of the targets based on all the first similarities corresponding to each of the predicted tracking boxes.
[0188] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the multi-object tracking method provided by the above-mentioned various methods. The method includes: determining at least one predicted tracking box in the current video frame and detection boxes corresponding to each target; the historical feature information corresponding to the predicted tracking box is the historical feature information of historical detection boxes in the first preset number of historical video frames before the current video frame, and the target represented by the historical detection box is the same as the target represented by the predicted tracking box; for each of the predicted tracking boxes, obtaining the clarity corresponding to each of the historical detection boxes associated with the predicted tracking box; determining a first similarity between each of the detection boxes and the predicted tracking box based on the clarity corresponding to each of the historical detection boxes, the historical feature information corresponding to each of the historical detection boxes, and the feature information corresponding to each of the detection boxes; and tracking each of the targets based on all the first similarities corresponding to each of the predicted tracking boxes.
[0189] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.
[0190] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0191] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A multi-object tracking method, characterized in that, comprising: Determine at least one predicted tracking box in the current video frame and the detection box corresponding to each target; the historical feature information corresponding to the predicted tracking box is the feature information of the historical detection boxes in the first preset number of historical video frames before the current video frame, and the targets represented by the historical detection boxes are the same as the targets represented by the predicted tracking box; For each of the predicted tracking boxes, obtain the clarity corresponding to each of the historical detection boxes associated with the predicted tracking box; Based on the clarity corresponding to each of the historical detection boxes, the historical feature information corresponding to each of the historical detection boxes, and the feature information corresponding to each of the detection boxes, determine the first similarity between each of the detection boxes and the predicted tracking box; Based on all the first similarities corresponding to each of the predicted tracking boxes, track each of the targets.
2. The multi-object tracking method according to claim 1, characterized in that, The obtaining the clarity corresponding to each of the historical detection boxes associated with the predicted tracking box includes: Determine the historical region where the target corresponding to each of the historical detection boxes is located; Determine the clarity of the historical region; Correspondingly, the determining the first similarity between each of the detection boxes and the predicted tracking box based on the clarity corresponding to each of the historical detection boxes, the historical feature information corresponding to each of the historical detection boxes, and the feature information corresponding to each of the detection boxes includes: For each of the historical detection boxes, determine the feature weight of the historical detection box based on the clarity of the historical region corresponding to the historical detection box; For each of the detection boxes, determine the second similarity between the feature information corresponding to the detection box and the historical feature information of each of the historical detection boxes; Based on the second similarity and the feature weight corresponding to each of the historical detection boxes, determine the first similarity between the detection box and the predicted tracking box.
3. The multi-object tracking method according to claim 1, characterized in that, The method further includes: Determine the region occupied by the target corresponding to each of the detection boxes in the current video frame; Input the region into a feature extraction model, and divide the region into at least two sub-regions by the feature extraction model based on a preset ratio; Extract the sub-feature information of each of the sub-regions through the feature extraction model; Based on each of the sub-feature information, determine the feature information corresponding to the detection box.
4. The multi-object tracking method according to any one of claims 1-3, characterized in that, The tracking each of the targets based on all the first similarities corresponding to each of the predicted tracking boxes includes: For each of the predicted tracking boxes, in the case that there is a first similarity greater than the similarity threshold among all the first similarities corresponding to the predicted tracking box, determine the detection box corresponding to the maximum first similarity among the first similarities greater than the similarity threshold as the detection box matching the predicted tracking box, and track the target corresponding to the predicted tracking box based on the matching detection box.
5. The multi-object tracking method according to claim 4, characterized in that, The method further includes: In the case that there is no first similarity greater than the similarity threshold among all the first similarities corresponding to the predicted tracking box, based on the position information of each of the detection boxes and the position information of the predicted tracking box, predict whether each of the detection boxes and the predicted tracking box match; In the case that it is predicted that the target detection box and the predicted tracking box match, obtain the first size information of the target detection box and the second size information of the target historical detection box corresponding to the predicted tracking box in the previous historical video frame; Based on the first size information and the second size information, in the case that it is determined that the change range between the target detection box and the target historical detection box is greater than the preset range, determine that the target detection box and the predicted tracking box do not match.
6. The multi-object tracking method according to claim 2, wherein, the determining the second similarity between the feature information corresponding to the detection box and the historical feature information of each of the historical detection boxes includes: Based on the first position in the predicted tracking box, determine a first region including the predicted tracking box, and the area of the first region is greater than the area of the predicted tracking box; In the case that it is determined that the detection box is within the first region, determine the second similarity between the feature information corresponding to the detection box and each of the historical feature information; In the case that it is determined that the detection box is not within the first region, determine that the second similarity between the feature information corresponding to the detection box and each of the historical feature information is a preset similarity.
7. The multi-object tracking method according to claim 6, wherein, the method further includes: In the case that the prediction tracking box fails to match in the previous second preset number of historical video frames of the current video frame, based on the first position and a preset value, expand the first region to obtain a second region, and the preset value is related to the second preset number; Determine the second region as the new first region, and return to execute the step of determining the second similarity between the detection box and each of the historical feature information in the case that it is determined that the detection box is within the first region.
8. The multi-object tracking method according to claim 2, wherein, each of the historical feature information corresponding to the predicted tracking box is stored in a feature storage sequence; the method further includes: For each of the predicted tracking boxes, in the case that there is a first similarity greater than the similarity threshold among all the first similarities corresponding to the predicted tracking box, determine the minimum second similarity among the second similarities corresponding to the historical feature information of each of the historical detection boxes included in the feature storage sequence; Delete the historical feature information corresponding to the minimum second similarity, and add the feature information of the detection box corresponding to the first similarity greater than the similarity threshold to the feature storage sequence.
9. A multi-object tracking device, wherein, comprises: a determination module, configured to determine at least one predicted tracking box in a current video frame and detection boxes corresponding to each target; The historical feature information corresponding to the predicted tracking box is the feature information of the historical detection boxes in the first preset number of historical video frames before the current video frame, and the targets represented by the historical detection boxes are the same as the targets represented by the predicted tracking box; An acquisition module, configured to acquire the clarity corresponding to each of the historical detection boxes associated with the predicted tracking box for each of the predicted tracking boxes; The determination module is further configured to determine a first similarity between each of the detection boxes and the predicted tracking box based on the clarity corresponding to each of the historical detection boxes, the historical feature information corresponding to each of the historical detection boxes, and the feature information corresponding to each of the detection boxes; A tracking module, configured to track each of the targets based on all the first similarities corresponding to each of the predicted tracking boxes.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, the multi-target tracking method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Target tracking method and device
CN120526406A
Target tracking method and apparatus
CN120526406B