View adaptive multi-target tracking method

Through the view adaptive multi-objective tracking method, the view type is identified using depth relationship cues and adaptively adjusting the matching strategy, the problem of poor tracking accuracy in different views is solved, and the multi-objective tracking effect with high accuracy and robustness is achieved.

CN120070917APending Publication Date: 2025-05-30GUIZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510043342.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When existing multi-objective tracking technologies process videos under different views, the tracking accuracy is poor, especially when a large number of low-score detection positions overlap, it is easy to cause false correlations.

Method used

A view adaptive multi-objective tracking method is proposed. Through the view category recognition method based on the depth relationship clue, the bounding box area and pseudo-depth of the detection target are calculated, and the degree of dispersion is compared to determine the view type. The trajectory matching strategy is adaptively adjusted according to the view type, and the depth interval division and cascade association are performed.

Benefits of technology

It improves the accuracy and robustness of multi-objective tracking, reduces interference between different pedestrians, improves matching accuracy and efficiency, adapts to data characteristics of different environments, and achieves uniform division of target sets in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070917A_ABST
    Figure CN120070917A_ABST
Patent Text Reader

Abstract

The invention discloses a view adaptive multi-target tracking method. The method comprises the following steps: obtaining a detection set Dk of pedestrian targets; the view category identification method based on the depth relation clues is utilized, according to coordinates of each detection target dk in a detection set Dk, the bounding box area ak and the pseudo depth pk of two corresponding depth relation clues are calculated, a bounding box area set Ak and a pseudo depth set Pk in a frame fk are obtained, and then the view category identification method based on the depth relation clues is obtained by comparing the discrete degree of the Ak and the discrete degree of the Pk. Obtaining a view type v corresponding to the frame fk; according to the obtained depth relation clue and the view type, depth interval division is carried out on a prediction track and a detection target according to a proposed track set and detection set division method of the self-adaptive view, and cascade association is completed; all frame sequences in the video V are processed by adopting the method, and multi-target identification tracking is carried out. The method has the advantages of being capable of achieving multi-target tracking, high in tracking accuracy and high in robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a view-adaptive multi-object tracking method. Background Art

[0002] Multi-object tracking technology (MOT), as a fundamental but important visual task, aims to maintain the identity of each object in a video scene and construct motion trajectories over time. MOT has the ability to perceive and understand environmental changes, can obtain the person ID and motion state in real time, and provides accurate environmental perception and autonomous decision-making support for intelligent systems.

[0003] MOT methods following the tracking-by-detection paradigm include two independent stages: detection and association. The first stage is used to detect the targets to be tracked in the scene frame by frame and obtain the detection set for each frame. The second stage is used to estimate the states of the existing trajectory set and associate the newly detected target set with the trajectory set. In the prior art, in order to accurately identify the targets in each frame and fully exploit the potential information in the detection boxes, ByteTrack proposed a two-stage matching method. This method performs tracking by associating almost every detection box. By using the similarity between the detection boxes and the tracking trajectories, while retaining the high-score detection results, it introduces a second association stage to recover the real targets from the low-score detection results, reducing the occurrence of real target loss and trajectory fragmentation and achieving an improvement in association performance. However, when a large number of low-score detection positions in the scene overlap, this two-stage matching method is prone to false associations. On this basis, SparseTrack starts from the perspective of decomposing low-score detections. According to the pseudo-depth of the targets obtained from the two-dimensional image, multiple depth intervals are evenly divided in the image. Then, the targets in them are associated in the order of the distance of the depth intervals, effectively alleviating the collision between trajectories with similar positions but different depths. However, due to the influence of factors such as the application field, shooting equipment, and camera angle, the views in different videos often show great differences, resulting in poor tracking accuracy. Summary of the Invention

[0004] The object of the present invention is to overcome the above-mentioned drawbacks and propose a view-adaptive multi-object tracking method with high tracking accuracy and strong robustness.

[0005] A view-adaptive multi-object tracking method of the present invention, wherein: the method includes the following steps:

[0006] Step 1: Input video V, and split video V into a separate frame sequence, with each frame denoted as f k , where k represents the serial number of the frame;

[0007] Step 2: For the current frame fk , use a target detector to process f k to obtain the detection set D of all pedestrian targets in frame f k ; k ;

[0008] Step 3: Calculate depth relationship clues and obtain view categories: Use the view category recognition method based on depth relationship clues (VTRM). This method calculates the bounding box area a k and pseudo-depth p k corresponding to each detected target d k in the detection set D according to the coordinates, and obtains the bounding box area set A k and pseudo-depth set P k in frame f k . Then, by comparing the dispersion degrees of A k and P k , obtain the view type v corresponding to frame f k . The specific steps include: k

[0009] Step 3.1: For each detected target d k in the detection set D k , represent its position in the image using coordinates (x k , y k , w k , h k ), where x k , y k represent the abscissa and ordinate of the upper left point of the detection box in the image plane, and w k and h k represent the width and height of the detection box respectively. Denote the image height as h img ; According to the coordinate information of the detection box, calculate the sizes of the two depth relationship clues corresponding to the detected target d k . The calculation formula is:

[0010]

[0011] where a k and p k represent the bounding box area and pseudo-depth of the detected target d k respectively. Then, repeat the above calculation process for each target in the detection set to obtain the bounding box area set A k and pseudo-depth set P k in frame f k , and construct depth relationship clues;

[0012] Step 3.2: For the bounding box area set A of each current frame f k ​​k and the pseudo-depth set P k , calculate their standard deviations as a measure of the degree of dispersion respectively. The calculation formula is as follows:

[0013]

[0014] where S a and S p represent the standard deviations of the bounding box area set A k and the pseudo-depth set P k respectively, n represents the number of elements in the set, and represent the averages of the bounding box area set A k and the pseudo-depth set P k respectively; compare with S p . When is greater than S p , it indicates that the degree of dispersion of the bounding box area set in the frame is greater than that of the pseudo-depth set. Set the view type v k of the current video frame to the front view; when S p is greater than , it indicates that the degree of dispersion of the pseudo-depth set in the current frame is greater than that of the bounding box area set. Set the view type v k of the current video frame to the top view;

[0015] Step 4: Adopt an adaptive view-based detection set partitioning method. For each target d k in the detection set D k , compare its detection confidence with a preset detection confidence threshold τ det ; if the detection confidence of the target d k is greater than or equal to the threshold τ det , then partition the target into the high-confidence target set D high ; if the detection confidence of the target d k is less than the threshold τ det , then partition the target into the low-confidence target set D low ; for the trajectory set T k-1 in the previous frame f k-1 , adopt an adaptive view-based trajectory set partitioning method. According to whether it is associated with the detection in the previous frame f k-1 , partition the successfully associated trajectories into the active trajectory set T active , and partition the unsuccessfully associated trajectories into the inactive trajectory set T unactive ;

[0016] Step 5: Association of Predicted Trajectory and Detected Target: Based on the obtained depth relation clues and view types, perform depth interval division on the predicted trajectory and the detected target and complete cascaded association. The specific steps are as follows:

[0017] Step 5.1: First Association Stage: Construct the spatial cost matrix C active between the active predicted trajectory set T high and the high-confidence detected target set D iou and the appearance cost matrix C reid ; Use the Hungarian algorithm to match between T iou and D reid based on C active and C high to determine the target trajectories without occlusion in the video frame; Keep the remaining trajectories as T remain and the remaining detected targets as D remain ;

[0018] Step 5.2: Second Association Stage: Perform a second match on the low-confidence detection set D low and the remaining trajectory set T remain . Based on the view-based adaptive partitioning method of trajectory set and detection set (VAPM), input the obtained view type v k , the low-confidence detection set D low and the remaining trajectory set T remain to obtain the evenly partitioned detection subset D sub and T sub . The calculation formula for the interval length is as follows:

[0019]

[0020] where represents the interval size in the top view, p i and p i+1 represent the sizes of the i-th and (i + 1)-th pseudo-depths in the pseudo-depth set respectively, p max represents the maximum pseudo-depth in the pseudo-depth set, p min represents the minimum pseudo-depth in the pseudo-depth set, n is the number of intervals, represents the interval size in the front view, a i and a i+1 represent the i-th and (i + 1)-th bounding box areas in the bounding box area set respectively, ρ is the proportionality coefficient; Perform multi-level cascaded matching on the detection subset D sub and T sub . Input each group of D i and T i into the IoU to calculate the spatial cost C i ; Take C iInput it into the Hungarian algorithm for matching to obtain the re-matching trajectory set T rematched ; Place the trajectories that fail to match into the unmatched trajectory set T unmatched , and place the detection targets that fail to match into D unmatched ;

[0021] Step 5.3: Remaining target matching: Construct the spatial cost matrix C of the inactive trajectories T unactive and D remain , calculate the cost based on the distance IoU between the detection box and the trajectory prediction box, and use the Hungarian algorithm to match between T iou and D i according to C unactive and D high ; The detection targets that fail to match will be initialized as new trajectories for tracking in the next frame;

[0022] Step 6: Loop through steps 2 - 5 until all frame sequences in the video V are processed to perform multi-target recognition and tracking.

[0023] In the above view-adaptive multi-target tracking method, where: in step 2, the target detector is the target detection model YOLOX.

[0024] In the above view-adaptive multi-target tracking method, where: in step 2, each detection target d k in the detection set D k records the coordinates of the upper left point, width, height, and detection confidence of its target detection box.

[0025] In the above view-adaptive multi-target tracking method, where: in step 4, the confidence threshold τ det is 0.6.

[0026] In the above view-adaptive multi-target tracking method, where: in step 4, the method for partitioning the trajectory set of the adaptive view: First, input each trajectory t k-1 into the Kalman filter for trajectory prediction, then input it into the global motion compensation module GMC for camera motion compensation, and finally, according to whether it is associated with the detection in the previous frame f k-1 , divide the successfully associated trajectories into the active trajectory set T active , and divide the unsuccessfully associated trajectories into the inactive trajectory set T unactive .

[0027] In the above view-adaptive multi-target tracking method, where: in step 5.1, the construction of the spatial cost matrix C between the active prediction trajectory set T active and the high-confidence detection target set D high ​iou , calculate the spatial cost based on the Intersection over Union (IoU) distance between the detection bounding box and the trajectory prediction bounding box.

[0028] The above-mentioned view-adaptive multi-object tracking method, wherein: in step 5.1, the construction of the active prediction trajectory set T active and the high-confidence detection target set D high the appearance cost matrix C between them reid , calculate the appearance cost using the appearance features of the detection bounding box obtained by the object re-identification model Re-ID.

[0029] Compared with the prior art, the present invention has obvious beneficial effects. As can be seen from the above solution, the present invention uses a view category recognition method based on depth relationship clues. According to each detection target d k in the set D k , the area a k of the two depth relationship clue bounding boxes and the pseudo-depth p k corresponding to it are calculated based on the coordinates, and the bounding box area set A k and the pseudo-depth set P k in the frame f k are obtained. Then, by comparing the dispersion of A k and P k , the view type v corresponding to the frame f k is obtained. Considering the law that the size of the detection bounding box changes with depth during the movement of the tracking target, using the coordinate data that can be directly obtained in the two-dimensional image, the bounding box area and pseudo-depth of the detection bounding box are calculated, and the depth relationship between the targets is reflected by comparing their sizes, constructing depth relationship clues, and assisting in the acquisition of the view type and the association of the trajectory and the detection target in the subsequent matching stage. According to the obtained depth relationship clues and view type, according to the proposed method for dividing the trajectory set and detection set of the adaptive view, the prediction trajectory and the detection target are divided into depth intervals and the cascade association is completed. Dividing the trajectory and detection set of the current scene can not only significantly reduce the interference between different pedestrians and improve the matching accuracy, but also reduce the number of negative samples in the matching process and improve the matching efficiency. And adopting different interval division methods according to the view type can adapt to the data characteristics of different environments and achieve uniform division of the target set in various complex scenes. Therefore, the present invention mainly has the following advantages:

[0030] 1) The present invention uses depth relationship clues to reveal the depth relationship between targets and adaptively adjusts the trajectory matching strategy based on the view type to achieve robust tracking of occlusion scenes under front views and top views.

[0031] 2) The present invention constructs depth relationship cues by using pseudo-depth and bounding box area, characterizes the depth relationship between objects based on the size relationship between them, and then realizes the accurate recognition of the current scene view type by using the difference in their dispersion under different views.

[0032] 3) Based on the difference in data distribution under different views, the present invention adaptively adjusts the interval division mode according to the view type to achieve uniform division of the trajectory set and detection set under different views.

[0033] 4) The present invention introduces the displacement information of the bounding box between two frames into the IoU calculation, enhances the ability to distinguish overlapping objects, and realizes the correct association of the detection box and the trajectory box.

[0034] In summary, the present invention has the characteristics of being able to achieve multi-object tracking, with high tracking accuracy and strong robustness.

[0035] The beneficial effects of the present invention are further described below through specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The specific embodiments, features and effects of a view-adaptive multi-object tracking method according to the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.

[0038] See Figure 1 , a view-adaptive multi-object tracking method of the present invention, wherein: the method includes the following steps:

[0039] Step 1: Input video V, and split video V into a separate frame sequence, with each frame denoted as f k , where k represents the frame number;

[0040] Step 2: For the current frame f k , use the object detector YOLOX to process f k to obtain the detection set D k of all pedestrian objects in frame f k , and each object d k in set D k records the upper left point coordinates, width and height of its object detection box and its detection confidence;

[0041] Step 3: Depth relationship cue calculation and view category acquisition: Use the view category recognition method based on depth relationship cues (VTRM), and according to each detection object d k in set D kThe area a of the two depth relationship clue bounding boxes corresponding to the coordinate calculation is obtained k and the pseudo-depth p k , and the bounding box area set A k in the frame f k and the pseudo-depth set P k are obtained. Then, by comparing the dispersion degree of A k and P k , the view type v corresponding to the frame f k is obtained. The specific steps are as follows:

[0042] Step 3.1: For each detection target d k in the detection set D k , its position in the image is represented by coordinates (x k , y k , w k , h k ), where x k , y k represent the abscissa and ordinate of the upper left point of the detection box in the image plane, and w k and h k represent the width and height of the detection box respectively. Denote the image height as h img . According to the coordinate information of the detection box, the sizes of the two depth relationship clues corresponding to the target d k are calculated. The calculation formula is:

[0043]

[0044] where a k and p k represent the bounding box area and pseudo-depth of the detection target d k respectively. Then, the above calculation process is repeated for each target in the detection set, and the bounding box area set A k in the frame f k and the pseudo-depth set P k are obtained to construct the depth relationship clues.

[0045] This embodiment is different from the usual depth acquisition method based on lidar. Considering the law that the size of the detection box changes with depth during the movement of the tracking target, the coordinate data that can be directly obtained in the two-dimensional image is used. According to the calculated bounding box area and pseudo-depth of the detection box, the depth relationship between targets is reflected by comparing their sizes, the depth relationship clues are constructed, and they are used to assist in obtaining the view type and associating the trajectory with the detection target in the subsequent matching stage.

[0046] Step 3.2: For the bounding box area set A k and the pseudo-depth set P k of each current frame f k, calculate their standard deviations respectively as a measure of the degree of dispersion, and the calculation formula is as follows:

[0047]

[0048] Among them, S a and S p represent the standard deviations of the bounding box area set A k and the pseudo-depth set P k respectively, n represents the number of elements in the set, and represent the averages of the bounding box area set A k and the pseudo-depth set P k respectively. Compare with S p . When is greater than S p , it indicates that the degree of dispersion of the bounding box area set in the frame is greater than that of the pseudo-depth set, and set the view type v k of the current video frame to the front view. When S p is greater than , it indicates that the degree of dispersion of the pseudo-depth set in the current frame is greater than that of the bounding box area set, and set the view type v k of the current video frame to the top view.

[0049] Step 4: Adopt an adaptive view detection set partitioning method. For each target d k in the detection set D k , compare its detection confidence with the preset detection confidence threshold τ det = 0.6; if the detection confidence of the target d k is greater than or equal to the threshold τ det , then divide the target into the high-confidence target set D high ; if the detection confidence of the target d k is less than the threshold τ det , then divide the target into the low-confidence target set D low ; for the trajectory set T k-1 in the previous frame f k-1 , adopt an adaptive view trajectory set partitioning method. This method first inputs each trajectory t k-1 into the Kalman filter for trajectory prediction, then inputs it into the global motion compensation module GMC for camera motion compensation, and finally, according to whether it is associated with the detection in the previous frame f k-1 , divide the successfully associated trajectories into the active trajectory set T active , and divide the unsuccessfully associated trajectories into the non-active trajectory set T unactive .

[0050] Step 5: Association of predicted trajectory and detected target: According to the obtained depth relation clues and view types, based on the proposed method for dividing trajectory set and detection set of adaptive view (VAPM), perform depth interval division on the predicted trajectory and detected target and complete cascaded association. The specific steps are as follows:

[0051] Step 5.1: First association stage: Construct the active predicted trajectory set T active and the high-confidence detected target set D high The spatial cost matrix C iou and the appearance cost matrix C reid between them. Calculate the spatial cost based on the intersection over union (IoU) distance between the detection box and the trajectory prediction box, and calculate the appearance cost using the appearance features of the detection box obtained by the Re-ID model; Use the Hungarian algorithm to match based on C iou and C reid between T active and D high to determine the target trajectories without occlusion in the video frame; Keep the remaining trajectories as T remain , and the remaining detected targets as D remain ;

[0052] Step 5.2: Second association stage: Perform a second match on the low-confidence detection set D low and the remaining trajectory set T remain . Based on the method for dividing trajectory set and detection set of adaptive view, input the obtained view type v k , the low-confidence detection set D low and the remaining trajectory set T remain to obtain the evenly divided detection subset D sub and T sub . The calculation formula for the interval length is as follows:

[0053]

[0054] where, represents the interval size in the top view, pi and pi +1 represent the sizes of the i-th and i + 1-th pseudo-depths in the pseudo-depth set respectively, p max represents the maximum pseudo-depth in the pseudo-depth set, p min represents the minimum pseudo-depth in the pseudo-depth set, n is the number of intervals, represents the interval size in the front view, a i and a i+1 represent the i-th and i + 1-th bounding box areas in the bounding box area set respectively, ρ is the proportionality coefficient. Perform multi-level cascaded matching on the detection subset D sub and T sub and match each group of Di and T i are input into IoU to calculate the spatial cost C i . C i is input into the Hungarian algorithm for matching to obtain the re-matched trajectory set T rematched . The trajectories that fail to be matched are placed into the unmatched trajectory set T unmatched , and the detection targets that fail to be matched are placed into D unmatched ;

[0055] Step 5.3: Remaining target matching: Construct the spatial cost matrix C unactive of T remain and D iou , calculate the cost based on the IoU distance between the detection box and the predicted trajectory box, and use the Hungarian algorithm to match according to C i between T unactive and D high to ensure that all possible targets are considered to avoid omission; the detection targets that fail to be matched are initialized as new trajectories for tracking in the next frame.

[0056] Distinguishing dense targets in the scene and separating overlapping trajectories as much as possible is an effective method to solve the occlusion phenomenon. Dividing the trajectories and detection sets of the current scene can not only significantly reduce the interference between different pedestrians, improve the matching accuracy, but also reduce the number of negative samples in the matching process and improve the matching efficiency. And adopting different interval division methods according to the view type can adapt to the data characteristics of different environments and achieve uniform division of the target set under various complex scenes.

[0057] Step 6: Loop through Steps 2 - 5 until all frame sequences in the video V are processed to perform multi-target recognition and tracking.

[0058] Performance analysis

[0059] On the Ubuntu 20.04.6 system, the present invention was verified using the PyTorch framework. In all subsequent tests, the publicly available YOLOX was used as the detector of the model, and the same detector parameters as ByteTrack were used. The detection confidence threshold τdet was set to 0.6, the number of intervals n was set to 8, and the proportionality coefficient ρ was set to 0.5. All experiments were run on a single NVIDIA A10 GPU with 24GB RAM.

[0060] MOT17 is a widely used standard benchmark in MOT, consisting of 7 training and test sequences captured by moving and static cameras, which contain complex pedestrian scenes under various views, with high human appearance similarity and variable movement routes.

[0061] In the comparison of the MOT17 dataset, the present invention is mainly compared with current SOTA online tracking methods, such as ByteTrack, ScoreMOT, OC-SORT, BOT-SORT, SparseTrack, etc.

[0062] Table 1 Comparative performance results of the method of the present invention and other methods on the MOT17 dataset

[0063]

[0064] Observing the data in Table 1, on the MOT17 test set, the three main indicators of the present invention, MOTA, IDF1, and HOTA, are 65.5, 81.7, and 80.7 respectively, achieving the best results among all the compared online tracking methods, and achieving the second-best result with 1206 in the IDS indicator. Compared with the second-place SparseTrack, the present invention has achieved significant improvements in all main indicators, with an increase of 0.4% in HOTA, 0.7% in MOTA, and 0.6% in IDF1 respectively.

[0065] In summary, in order to address the serious occlusion problem in complex scenarios, the present invention proposes a view-adaptive multi-object tracking method. This method differentiates different views and adopts specific association strategies for the scene differences between the front view and the top view. In the detection and processing stage, depth relationship clues are extracted from the bounding box position information by VTRM to characterize the depth relationship between objects, and the view type of the current scene is judged by comparing the standard deviation sizes of them. In the data association stage, first, the VAPM mode is adaptively adjusted based on the view type, and the detection set and the trajectory set under the front view and the top view are evenly divided respectively, so as to realize the sparse decomposition of the scene. Secondly, the IOU joint metric based on the object displacement and the position overlap degree is used to improve the accuracy of trajectory matching. When combined with the widely used appearance model Re-ID, the present invention shows higher performance and faster association speed.

[0066] The above is only a preferred embodiment of the present invention, and it does not impose any form of limitation on the present invention. Any simple modification, equivalent change, and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A view-adaptive multi-target tracking method, characterized by: The method comprises the following steps: Step 1: Input video V and divide it into separate frame sequences, each frame is denoted as f k , where k represents the frame number; Step 2: For the current frame f k , using the target detector to k Processing is performed to obtain frame f k The detection set D of all pedestrian targets in k ; Step 3: Calculate the depth relationship clues and obtain the view category: Use the view category recognition method based on the depth relationship clues (VTRM), which is based on the detection set D k Each detection target d k The coordinates of the two depth relationship clues are calculated to calculate the bounding box area a k and pseudo depth p k , get frame f k The bounding box area set A in k and the pseudo depth set P k , and then by comparing A k With P k The discrete degree of frame f k The corresponding view type v, the specific steps include: Step 3.1: Target the detection set D k Each detection target d k , using coordinates to represent its position in the image (x k ,y k ,w k ,h k ), where x k ,y k represents the horizontal and vertical coordinates of the upper left point of the detection box in the image plane, w k and h k Respectively represent the width and height of the detection frame, and the image height is h img ; According to the coordinate information of the detection box, calculate the detection target d k The corresponding sizes of the two depth relationship clues are calculated as follows: Among them, a k and p k Represent the detection target d k The bounding box area and pseudo depth of the frame are then repeated for each target in the detection set to obtain the frame f k The bounding box area set A in k and the pseudo depth set P k , build deep relationship clues; Step 3.2: For each current frame f k The bounding box area set A k and the pseudo depth set P k , respectively calculate their standard deviation as a measure of dispersion, the calculation formula is as follows: Among them, S a and S p Represents the bounding box area set A k and the pseudo depth set P k The standard deviation of n represents the number of elements in the set. and Represents the bounding box area set A k and the pseudo depth set P k The average value of With S p For comparison, Greater than S p When , it indicates that the discreteness of the bounding box area set in the frame is greater than the discreteness of the pseudo depth set, and the view type v of the current video frame is set k is the front view; when S p Greater than , indicating that the discreteness of the pseudo depth set in the current frame is greater than the discreteness of the bounding box area set, setting the view type v of the current video frame k It is a top view; Step 4: Adopt the detection set partitioning method of adaptive view. This method is used for the detection set D k Each target d k , and its detection confidence and the preset detection confidence threshold τ det Compare; if the target d k The detection confidence is greater than or equal to the threshold τ det , then the target is divided into the high confidence target set D high ; If the target d k The detection confidence is less than the threshold τ det , then the target is divided into the low confidence target set D low ; For the previous frame f k-1 The set of trajectories in k-1 , using the adaptive view trajectory set partitioning method, according to its previous frame f k-1 Whether to associate with the detection, and divide the successfully associated trajectories into the active trajectory set T active , divide the unsuccessfully associated trajectories into the inactive trajectory set T unactive ; Step 5: Associating the predicted trajectory with the detected target: According to the acquired depth relationship clues and view type, the predicted trajectory and the detected target are divided into depth intervals and cascaded association is completed. The specific steps are as follows: Step 5.1: First association phase: constructing the active prediction trajectory set T active And the high confidence detection target set D high The spatial cost matrix C between iou and the appearance cost matrix C reid ; Using the Hungarian algorithm based on C iou and C reid In T active and D high Match between them to determine the target track without occlusion in the video frame; keep the remaining track as T remain , the remaining detection targets are D remain ; Step 5.2: Second association phase: for low confidence detection set D low and the remaining trajectory set T remain Perform the second matching based on the adaptive view trajectory set and detection set partitioning method (VAPM), and input the acquired view type v k , low confidence detection set D low and the remaining trajectory set T remain , get the evenly divided detection subset D sub and T sub , the calculation formula of interval length is as follows: in, Indicates the size of the interval in the top view, p i and p i+1 They represent the size of the i-th and i+1-th pseudo depths in the pseudo depth set, respectively, and p max represents the maximum value of the pseudo depth in the pseudo depth set, p min Indicates the minimum value of the pseudo depth in the pseudo depth set, n is the number of intervals, Indicates the size of the interval under the front view, a i and a i+1 Respectively represent the i-th and i+1-th bounding box areas in the bounding box area set, ρ represents the proportional coefficient; for the detection subset D sub and T sub Perform multi-level cascade matching, and match each group of D i and T i Input into IoU to calculate the spatial cost C i ; C i Input into the Hungarian algorithm for matching, and obtain the re-matching trajectory set T rematched ; Place the unmatched trajectories into the unmatched trajectory set T unmatched , the unmatched detection targets are placed in D unmatched ; Step 5.3: Remaining target matching: constructing inactive tracks T unactive and D remain The spatial cost matrix C iou , based on the distance IoU between the detection box and the trajectory prediction box, the cost is calculated using the Hungarian algorithm based on C i In T unactive and D high The unmatched detection targets will be initialized as new tracks and tracked in the next frame. Step 6: Loop through steps 2 to 5 until all frame sequences in the video V are processed and perform multi-target recognition and tracking.

2. The view-adaptive multi-target tracking method according to claim 1, characterized in that: In step 2, the target detector is the target detection model YOLOX.

3. A view-adaptive multi-target tracking method according to claim 1 or 2, characterized in that: In step 2, the detection set D k Each detection target d in k The upper left point coordinates, width, height, and detection confidence of the target detection box are recorded.

4. The view-adaptive multi-target tracking method according to claim 1, characterized in that: In step 4, the confidence threshold τ det is 0.

6.

5. The view-adaptive multi-target tracking method according to claim 1, characterized in that: In step 4, the adaptive view trajectory set division method is as follows: first, each trajectory t k-1 It is input into the Kalman filter for trajectory prediction, and then input into the global motion compensation module GMC for camera motion compensation. Finally, according to its position in the previous frame f k-1 Whether to associate with the detection, and divide the successfully associated trajectories into the active trajectory set T active , divide the unsuccessfully associated trajectories into the inactive trajectory set T unactive .

6. The view-adaptive multi-target tracking method according to claim 1, characterized in that: In step 5.1, the active prediction trajectory set T is constructed active And the high confidence detection target set D high The spatial cost matrix C between iou ,The spatial cost is calculated based on the intersection over union (IoU) distance between the detection box and the trajectory prediction box.

7. A view-adaptive multi-target tracking method according to claim 1 or 6, characterized in that: In step 5.1, the active prediction trajectory set T is constructed active And the high confidence detection target set D high The appearance cost matrix C between reid ,The appearance cost is calculated using the appearance features of the detection box obtained by the target re-identification model Re-ID.

Citation Information

Cited By

  • Multi-target tracking method and device and medium

    CN120355749A