A sub-track based adaptive multi-target tracking method
Patent Information
- Application Number
- CN202311647353.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-12-04
AI Technical Summary
然而,使用外观信息进行目标关联时,容易受到外观变化的影响
[0016]Beneficial Effects: The video target tracking method provided by this invention embeds target detection and appearance feature extraction into the same network, outputting the target's appearance feature vector simultaneously with its location. This offers an advantage in inference speed. Furthermore, to address association errors when a target reappears after prolonged occlusion, an adaptive weighted association strategy based on trajectory segmentation is proposed. This strategy categorizes trajectories according to whether they are occluded and applies different association strategies to different types of trajectories. Particularly for occluded trajectories, the number of occluded frames is used as a weight reference for the position and appearance information during association, avoiding position prediction bias caused by the inability to update Kalman filter parameters due to prolonged occlusion, thus further improving tracking accuracy. The appearance feature update module uses the confidence of the associated detection box as the update weight, giving higher weight to high-quality appearance features to improve their discriminative power. This strategy not only improves tracking accuracy but also meets real-time operating speed requirements. Using the method of this invention, any target appearing in a video sequence can be tracked quickly and accurately, improving tracking accuracy, speed, and robustness.
Smart Images

Figure CN117974712B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video target tracking, and more specifically to an adaptive multi-target tracking method based on segmented trajectories. Background Technology
[0002] Visual object tracking is a crucial problem in computer vision, its main task being to automatically locate and track a pre-given target within a video sequence. Video object tracking technology is widely used in video surveillance, remote sensing satellites, and autonomous driving. Although visual object tracking has seen rapid development in recent years, many challenges remain to be addressed in practical applications, such as occlusion, rapid movement, target deformation, lighting conditions, scale variations, and cluttered backgrounds.
[0003] In multi-target tracking, occlusion refers to a situation where a target is partially or completely hidden during observation due to obstruction by other objects or the environment. Occlusion can be categorized into long occlusion and short occlusion based on its duration. Long occlusion occurs when the target is completely unobservable until the obstruction is removed or the target reappears. Short occlusion occurs when only a portion of the target is obscured. Both long and short occlusion are challenging in multi-target tracking tasks because they can lead to target loss or trajectory interruption.
[0004] Detection-based target tracking tasks detect the position and appearance information of targets in each frame and use Kalman filtering to predict the target's position in the next frame. The predicted position information and extracted appearance information are then correlated to form a trajectory. However, using position or appearance information alone has limitations. If only position information is used for correlation, the predicted target position must be obtained through Kalman filtering. As target density increases, long-term occlusion often occurs during multi-target tracking. Kalman filtering uses observations in state updates to prevent the posterior state estimate and covariance from deviating too far from the true value. When a target is lost for a long time due to occlusion, it cannot provide observations to the Kalman filter, preventing it from updating parameters. This leads to error accumulation as the number of occluded frames increases, resulting in inaccurate target position prediction. If only appearance information is relied upon for target correlation, the cosine distance between the appearance feature vectors of targets in adjacent frames needs to be calculated and used as an correlation matrix for matching. However, using appearance information for target correlation is easily affected by appearance changes. Factors such as changes in lighting, viewing angle, occlusion, and deformation can reduce the accuracy of cosine feature distance, thereby affecting the correct association of targets.
[0005] Therefore, in multi-target tracking, using only position or appearance information for target association between adjacent frames is not the optimal choice. Adaptively combining position and appearance information under different circumstances is crucial for improving multi-target tracking performance. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides an adaptive multi-object tracking algorithm based on split trajectory (STA-Track). This method employs a strategy of separating trajectories and adaptive weights to solve the target association problem under occlusion conditions and adapt to changes in target appearance during the tracking process. This results in better tracking performance and meets real-time tracking speed requirements.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A video target tracking method based on situation feedback and quality assessment includes the following steps: S1. Each frame of the entire video sequence is fed into a pre-trained object detection network for object detection and appearance feature extraction. Detection bounding box information of the target and target appearance feature vector ; S2. Classify the target detection box and trajectory, and obtain the predicted target position through Kalman filtering; S3, for the set of trajectories With the set of detection boxes Perform the first association; S4, For the set of trajectories With the set of detection boxes Perform a second association; S5, For the set of trajectories With the set of detection boxes Perform a third association; S6. Update the appearance features of successfully associated trajectories; S7. For the set of detection boxes Unassociated bounding boxes are initialized as trajectories for the trajectory set. Remove tracks that failed to associate successfully for 30 consecutive frames. Among them, steps S1-S2 are the target detection process, steps S3-S5 are the target association process, and steps S6-S7 are the subsequent processing process. By repeating steps S3-S7, a trajectory is gradually formed, and the entire target tracking is completed.
[0008] Furthermore, the object detection network described in step S1 has two main branches: a location detection branch and an appearance feature extraction branch. Both branches use ResNet50 for feature extraction and an FPN feature pyramid structure is added for feature fusion. The specific training steps are as follows: S1.1 Select training subsets of CityPerson, CalTech, MOT16, CUHK-SYSU, and PRW datasets to form a joint training set. Use the joint training set to train the model to prevent the problem of result bias when experimenting on small datasets. S1.2. Each frame of the video is fed into the network for feature extraction, resulting in three prediction heads of different scales. Each prediction head consists of several stacked convolutions and outputs the bounding box location information, target confidence, and target appearance feature vector. S1.3. If the IOU between the detection box and the ground truth is greater than 0.5, it is considered foreground; if it is less than 0.4, it is considered background. The detection branch includes foreground / background classification loss and bounding box regression loss. The cross-entropy loss function is used as the foreground / background classification loss and appearance feature discriminative loss. The SmoothL1Loss loss function is used as the bounding box regression loss. The learning objective of each prediction head can be modeled as a multi-task learning problem. The joint objective can be written as a weighted linear loss sum for each scale and each component. S1.4 Set the batch size to 64, the momentum and weight decay rates to 0.9 and 1*10-4 respectively, and use the stochastic gradient descent algorithm to iteratively train 30 times to optimize the network parameters and save the results of each iteration.
[0009] Furthermore, the specific implementation steps in step S2 are as follows: S2.1 Setting the confidence threshold For frames The confidence threshold is greater than Place the target detection box into In the set, for frames The confidence threshold is greater than Place the target detection box into In the collection.
[0010] S2.2 Add the successfully associated trajectories from the previous frame to the trajectory set. In the middle, the trajectories that were not successfully associated in the previous frame are placed into the set. middle; S2.3. Kalman filtering is used to predict the position of the trajectory in the next frame of the image, which helps to improve the accuracy of the multi-target tracking system.
[0011] S2.4 Calculate all trajectories in Prediction boxes in With frames Detection box in IOU distance and cosine characteristic distance ; Furthermore, the specific implementation steps for the first association in step S3 are as follows: S3.1, First time on the trajectory set With the set of detection boxes To establish a correlation, if the IOU distance is... Less than the IOU distance threshold Simultaneously, cosine characteristic distance Less than the cosine distance threshold Then set the cosine distance update value. for Otherwise, update the cosine distance value. Set to 1 to update the cosine distance value. Distance from IOU The minimum value is compared to the final value of the association cost matrix. Then, the Hungarian algorithm is used to... The trajectory in Associate the detection boxes in the data; S3.2. Update the successfully associated trajectories using Kalman filtering, and add the unassociated trajectories to a set. In the middle, detection boxes that were not successfully associated are placed into a collection. middle.
[0012] Furthermore, the specific implementation steps for the second association in step S4 are as follows: S4.1, Second pairing of trajectory sets With the set of detection boxes To perform the association, first calculate the trajectory set. The number of frames since the last successful association of all trajectories ,for The trajectory is selected based on the IOU distance between the trajectory and the detection box. As the final cost matrix. For The trajectory will result in the loss of frames. Multiply by weighting factor As IOU distance Cosine feature distance The weights between them, for IOU distance Cosine feature distance Weighted averages are used to form the final cost matrix. The trajectory is selected by choosing the cosine feature distance between the trajectory and the detection box. As the final cost matrix; S4.2, For The trajectory of the object, whose cost matrix is represented by a weighted average of appearance and location information, can be calculated using the following formula: ; S4.3. Update the successfully associated trajectories using Kalman filtering, and continue to add unassociated trajectories. In the process, detection boxes that were not successfully associated are initialized with new trajectories.
[0013] Furthermore, the specific implementation steps for the third association in step S5 are as follows: S5.1, Third time for the trajectory set With the set of detection boxes To perform association, select the cosine feature distance between the trajectory and the detection box. As the final cost matrix; S5.2. Update the successfully associated trajectories using Kalman filtering, and add the unassociated trajectories to the list. In the set, detection boxes that were not successfully associated are considered as background.
[0014] Furthermore, in step S6, to improve the discriminative power of the target appearance features, the appearance features of the trajectory need to be updated in real time. We designed a confidence-based appearance feature update module, the calculation formula of which is as follows: ,in It is the appearance feature vector of the currently matched and detected target. This is the updated trajectory appearance feature vector. The confidence score of the target being detected. Update the weight parameters for the features.
[0015] Furthermore, in the final step S7, we need to update the Kalman filter parameters and trajectory appearance features of the associated trajectories. For those that were not successfully associated... The trajectory in the set is put into the set. In the context of sets that failed to be successfully associated... The trajectory in the data is retained; if it is not successfully associated within 30 frames, it is deleted. For sets of unassociated data... The detection box in the image is used to treat the newly appearing target and initialize its trajectory.
[0016] Beneficial Effects: The video target tracking method provided by this invention embeds target detection and appearance feature extraction into the same network, outputting the target's appearance feature vector simultaneously with its location. This offers an advantage in inference speed. Furthermore, to address association errors when a target reappears after prolonged occlusion, an adaptive weighted association strategy based on trajectory segmentation is proposed. This strategy categorizes trajectories according to whether they are occluded and applies different association strategies to different types of trajectories. Particularly for occluded trajectories, the number of occluded frames is used as a weight reference for the position and appearance information during association, avoiding position prediction bias caused by the inability to update Kalman filter parameters due to prolonged occlusion, thus further improving tracking accuracy. The appearance feature update module uses the confidence of the associated detection box as the update weight, giving higher weight to high-quality appearance features to improve their discriminative power. This strategy not only improves tracking accuracy but also meets real-time operating speed requirements. Using the method of this invention, any target appearing in a video sequence can be tracked quickly and accurately, improving tracking accuracy, speed, and robustness. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the target detection network structure in this invention; Figure 2 This is a flowchart of the related parts; Figure 3 This is a comparison chart of the MOTA and IDF1 of the method of this invention (STA-Track) with other competitive methods in the MOT17 dataset comparison experiment; Figure 4 This is a comparison chart of the MT and ML methods of the method of this invention (STA-Track) with other competitive methods in the MOT17 dataset comparison experiment. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0019] An adaptive multi-target tracking method based on trajectory segmentation includes a trajectory segmentation association strategy and appearance feature updating, specifically comprising the following steps S1 to S7. In this invention, the trajectory set... Indicates unobstructed trajectory, Indicates the occlusion trajectory, The set of detection boxes represents the unoccluded trajectories for which the first association attempt failed. Indicates a high-confidence detection box, High-confidence detection boxes indicating the first failed association. This represents a low-confidence detection box.
[0020] S1. Each frame of the entire video sequence is fed into a pre-trained object detection network for object detection and appearance feature extraction. Detection bounding box information of the target and target appearance feature vector The specific pre-training methods are S2.1 to S1.3.
[0021] The object detection network described in step S1 has two main branches: a location detection branch and an appearance feature extraction branch. Both branches use ResNet50 for feature extraction and incorporate an FPN feature pyramid structure for feature fusion. The network structure is shown below. Figure 1 The specific training steps are as follows: S1.1 Select training subsets from CityPerson, CalTech, MOT16, CUHK-SYSU, and PRW datasets to form a joint training set. Use the joint training set to train the model to prevent bias in results when experimenting on small datasets.
[0022] S1.2. The input image is processed through a ResNet50 residual network to extract features, and then the features at different scales are fused using an FPN pyramidal structure. The output results are downsampled by 1 / 8, 1 / 16, and 1 / 32 times the original image, respectively, with an output dimension of [missing information]. Where A is the number of anchor boxes set, D represents the dimension of the appearance feature vector, and H and W are the height and width of the output feature map, respectively.
[0023] S1.3 The number of anchor frames is set to 12, that is, 12 anchor frames are allocated to each channel. The ratio of the anchor frames is set to 1:3 to better match the proportion of pedestrians. The anchor frames with an IOU > 0.5 with the ground truth are set as foreground and the anchor frames with an IOU < 0.4 are set as background.
[0024] The training loss function formula is as follows:
[0025] in It is a learnable parameter that represents the weights of different loss functions; This represents the foreground / background classification loss (Cross-Entrophy Loss); This represents the bounding box regression loss (Smooth-L1 Loss); This is represented as the appearance feature discriminative loss (Cross-Entrophy Loss).
[0026] The batch size was set to 64, the momentum and weight decay rates were 0.9 and 1*10-4 respectively, and the network parameters were optimized by iteratively training the network 30 times using the stochastic gradient descent algorithm and saving the results of each iteration. S2. Preparatory work before association: Classify the target detection box and trajectory, and obtain the predicted target position through Kalman filtering. The specific steps are as follows: S2.1 Setting the confidence threshold For frames The confidence threshold is greater than Place the target detection box into In the set, for frames The confidence threshold is greater than Place the target detection box into In the set; S2.2 Add the successfully associated trajectories from the previous frame to the trajectory set. In the middle, the trajectories that were not successfully associated in the previous frame are placed into the set. middle; S2.3. Use Kalman filtering to predict the position of the trajectory in the next frame of the image to help improve the accuracy of the multi-target tracking system. S2.4 Calculate all trajectories in Prediction boxes in With frames Detection box in IOU distance and cosine characteristic distance .
[0027] S3. Perform the first association (trajectory set) With the set of detection boxes (Associate).
[0028] S3.1, if the distance between IOUs Less than the IOU distance threshold Simultaneously, cosine characteristic distance Less than the cosine distance threshold Then set the cosine distance update value. for Otherwise, update the cosine distance value. Set to 1 to update the cosine distance value. Distance from IOU The minimum value is compared to the final value of the association cost matrix. Then, the Hungarian algorithm is used to... The trajectory in The detection boxes in the data are associated.
[0029] By setting IOU and cosine distance thresholds, we can exclude targets with low appearance similarity or large positional distances. In the cost matrix calculation, we choose the minimum of the IOU and cosine distances as the final cost value. This further reduces the likelihood of two targets failing to correlate in densely packed situations and enhances the robustness of the tracking process. The calculation formula is as follows:
[0030]
[0031] S3.2. Update the successfully associated trajectories using Kalman filtering, and add the unassociated trajectories to a set. In the middle, detection boxes that were not successfully associated are placed into a collection. middle.
[0032] S4. Perform a second association (trajectory set) With the set of detection boxes (Associate).
[0033] S4.1 First, calculate the trajectory set. The number of frames since the last successful association of all trajectories and according to Further classification of trajectories, for The trajectory is selected based on the IOU distance between the trajectory and the detection box. As the final cost matrix; for The trajectory will result in the loss of frames. Multiply by weighting factor As IOU distance Cosine feature distance The weights between them, for IOU distance Cosine feature distance Weighted averages are used to form the final cost matrix. The trajectory is selected by choosing the cosine feature distance between the trajectory and the detection box. As the final cost matrix.
[0034] Since the target's position cannot be observed when it is occluded, the Kalman filter parameters cannot be updated. Therefore, the position information used for association when the target reappears contains errors. How to reasonably utilize position and appearance information becomes crucial for successful association. Thus, the number of occluded frames is used as the weight between position and appearance information during association. The final cost matrix calculation formula can be expressed as: .
[0035] S4.2. Update the successfully associated trajectories using Kalman filtering, and continue to add unassociated trajectories. In the process, detection boxes that were not successfully associated are initialized with new trajectories.
[0036] S5. Perform the third association (trajectory set) With the set of detection boxes ).
[0037] S5.1, For the set of trajectories With the set of detection boxes To perform association, select the cosine feature distance between the trajectory and the detection box. As the final cost matrix ; S5.2. Update the successfully associated trajectories using Kalman filtering, and add the unassociated trajectories to the list. In the set, detection boxes that were not successfully associated are considered as background.
[0038] S6. Update the appearance features of successfully associated trajectories, using the following formula: ,in It is the appearance feature vector of the currently matched and detected target. This is the updated trajectory appearance feature vector. The confidence score of the target being detected. Update the weight parameters for the features.
[0039] The appearance feature update module addresses the issue of target appearance changes during tracking due to factors such as lighting, occlusion, and camera angle. To obtain reliable appearance information and maintain trajectory continuity, this module considers historical target appearance information and performs a weighted summation based on the quality of appearance features in each frame. This weighted summation method allows high-quality appearance features to account for a larger proportion during the update process, improving the correlation accuracy during tracking.
[0040] S7. For the detection box set Unassociated bounding boxes are initialized as trajectories for the trajectory set. Remove tracks that failed to associate successfully for 30 consecutive frames. S1 above is the target detection process, and S2-S7 are the target association processes. The two are combined to form a complete target tracking process. In the actual target tracking process, the entire target tracking is completed by repeating steps S2-S7. The bounding box information and appearance feature vector of the target tracking are obtained from step S1.
[0041] The effectiveness of the present invention is verified through simulation experiments. The simulation experiments use the MOTchallenge17 dataset and compare it with competitive open-source methods in the field of target tracking. Among them, STA-Track refers to the method of this invention, and the comparison methods used in the simulation experiments of this invention include the following seven: 1. TPM, Reference [1]. Peng J, Wang T, Lin W, et al. TPM: Multiple object tracking with tracklet-plane matching[J]. Pattern Recognition,2020, 107:107480. 2.HDTR, Reference[2]. Babaee M, Athar A, Rigoll G. Multiplepeopletracking using hierarchical deep tracklet re-identification[J]. arXivpreprintarXiv:1811.04091, 2018. 3. FAMNet, Reference[3]. Chu P, Ling H. Famnet: Joint learningoffeature, affinity and multi-dimensional assignment for online multipleobjecttracking[C] / / Proceedings of the IEEE / CVF International Conference onComputerVision. 2019: 6172-6181. 4. Tracktor++v2, Reference[4]. Bergmann P, MeinhardtT, Leal-Taixe L.Tracking without bells and whistles[C] / / Proceedings of the IEEE / CVFInternational Conference on Computer Vision. 2019: 941-951. 5. TADN, reference [5]. Psalta A, Tsironis V, Karantzalos K. Transformer-based assignment decision network for multiple object tracking [J]. arXivpreprint arXiv:2208.03571, 2022. Simulation results (see attached) Figure 3 and attached Figure 4 , Figure 3 This chart compares the tracking success rates of our method with other competitive algorithms on the MOT17 dataset. Figure 3 The vertical axis represents the algorithm's accuracy (MOTA). MOTA comprehensively considers factors such as FP, FN, and IDswitch, and is calculated by normalizing the values of these factors to the total number of targets. FP refers to the number of times the tracker incorrectly identifies background or other non-target objects as targets; FN refers to the number of times the tracker fails to correctly identify real targets as targets; and IDswitch represents the number of target ID switches that occur during tracking. MOTA measures the overall performance of the tracker. The horizontal axis represents the tracker's target identity preservation capability (IDF1). The IDF1 metric measures the accuracy and completeness of the tracker in correctly identifying target identities. Figure 4 The horizontal axis (ML) represents the proportion of targets that are lost for most of the tracking process. It measures the tracker's performance in maintaining target visibility. The vertical axis (MT) represents the proportion of targets that are successfully tracked for most of the tracking process. Combined... Figure 3 and Figure 4 As can be seen, on the MOT17 dataset, the accuracy and precision of the method proposed in this invention (STA-Track) are superior to the other algorithms included in the performance comparison. Furthermore, the tracking speed of this invention reaches a maximum of 27 FPS, meeting real-time requirements. In summary, this invention improves target tracking accuracy while maintaining tracking speed.
[0042] To address the issue of incorrect association after a target reappears due to prolonged occlusion, this invention employs a trajectory segmentation and adaptive weighting strategy. When the target in a video sequence experiences complex conditions such as prolonged occlusion, deformation, and changes in illumination, the occluded trajectory is separated from the unoccluded trajectory. The trajectories are then associated separately based on their characteristics, and adaptive weights are added during the association and appearance feature update process to accommodate changes in target appearance and tracking status. This strategy can adjust the ratio of appearance information to positional information during the association process based on the number of frames the trajectory is occluded. Simultaneously, confidence levels can be used to increase the proportion of high-quality features during appearance updates. This mitigates association failures caused by the inability to update Kalman filter parameters in a timely manner under occlusion conditions. It also avoids the situation where the discriminative power of appearance features is low when the target undergoes deformation, jitter, or changes in illumination in complex scenes. The method of this invention achieves better tracking performance while meeting real-time tracking speed requirements.
[0043] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. An adaptive multi-target tracking method based on trajectory segmentation, characterized in that, Includes the following steps: S1. Each frame of the entire video sequence is fed into a pre-trained object detection network for object detection and appearance feature extraction. Detection bounding box information of the target and target appearance feature vector ; S2. Classify the target detection box and trajectory, and obtain the predicted target position through Kalman filtering; S3, for the set of trajectories With the set of detection boxes Perform the first association; S4, For the set of trajectories With the set of detection boxes The second association is performed using the following steps: S4.1 First, calculate the trajectory set. The number of frames since the last successful association of all trajectories ,for The trajectory is selected based on the IOU distance between the trajectory and the detection box. As the final cost matrix; for The trajectory will result in the loss of frames. Multiply by weighting factor As IOU distance Cosine feature distance The weights between them, for IOU distance Cosine feature distance Weighted summation is performed to obtain the final cost matrix. The trajectory is selected by choosing the cosine feature distance between the trajectory and the detection box. As the final cost matrix; S4.2 After the association is completed, Kalman filtering is performed on the successfully associated trajectories to update them, and the unassociated trajectories are put back into the system. In the process, detection boxes that fail to be associated are initialized with new trajectories; S5, For the set of trajectories With the set of detection boxes Perform a third association; S6. Update the appearance features of successfully associated trajectories; S7. For the detection box set Unassociated bounding boxes are initialized as trajectories for the trajectory set. Remove tracks that failed to associate successfully for 30 consecutive frames. Among them, the set of trajectories Indicates unobstructed trajectory, Indicates the occlusion trajectory, The set of detection boxes represents the unoccluded trajectories for which the first association attempt failed. Indicates a high-confidence detection box, High-confidence detection boxes indicating the first failed association. This represents a low-confidence detection box; The above steps S1-S2 are the target detection process, steps S3-S5 are the target association process, and steps S6-S7 are the subsequent processing process. By repeating steps S3-S7, a trajectory is gradually formed, and the entire target tracking is completed.
2. The adaptive multi-target tracking method based on segmented trajectories as described in claim 1, characterized in that, The object detection network described in step S1 has two main branches: a location detection branch and an appearance feature extraction branch. Both branches use ResNet50 for feature extraction and an FPN feature pyramid structure is added for feature fusion. The specific training steps are as follows: S1.1 Select training subsets of CityPerson, CalTech, MOT16, CUHK-SYSU, and PRW datasets to form a joint training set. Use the joint training set to train the model to prevent the problem of result bias when experimenting on small datasets. S1.
2. Each frame of the video is fed into the network for feature extraction, resulting in three prediction heads of different scales. Each prediction head consists of several stacked convolutions and outputs the bounding box location information, target confidence, and target appearance feature vector. S1.
3. If the IOU between the detection box and the ground truth is greater than 0.5, it is considered foreground; if it is less than 0.4, it is considered background. The detection branch includes foreground / background classification loss and bounding box regression loss. The cross-entropy loss function is used as the foreground / background classification loss and appearance feature discriminative loss. The SmoothL1Loss loss function is used as the bounding box regression loss. The learning objective of each prediction head can be modeled as a multi-task learning problem. The joint objective can be written as a weighted linear loss sum for each scale and each component. S1.4 Set the batch size to 64, and the momentum and weight decay rates to 0.9 and 1*10, respectively. -4 The network parameters were optimized by iteratively training the network 30 times using the stochastic gradient descent algorithm, and the results of each iteration were saved.
3. The adaptive multi-target tracking method based on trajectory segmentation according to claim 1, characterized in that, The specific implementation steps for the preparatory work before association in step S2 are as follows: S2.1 Setting the confidence threshold For frames The confidence threshold is greater than Place the target detection box into In the set, for frames The confidence threshold is greater than Place the target detection box into In the set; S2.2 Add the successfully associated trajectories from the previous frame to the trajectory set. In the middle, the trajectories that were not successfully associated in the previous frame are placed into the set. middle; S2.
3. Use Kalman filtering to predict the position of the trajectory in the next frame of the image to help improve the accuracy of the multi-target tracking system. S2.4 Calculate all trajectories in Prediction boxes in With frames Detection box in IOU distance and cosine characteristic distance .
4. The adaptive multi-target tracking method based on segmented trajectories according to claim 1, characterized in that, In step S3, the trajectory set is... With the set of detection boxes The specific steps to establish the association are as follows: S3.1, If the IOU distance Less than the IOU distance threshold Simultaneously, cosine characteristic distance Less than the cosine distance threshold Then set the cosine distance update value. for Otherwise, update the cosine distance value. Set to 1 to update the cosine distance value. Distance from IOU The minimum value is compared to the final value of the association cost matrix. Then, the Hungarian algorithm is used to... The trajectory in Associate the detection boxes in the data; S3.
2. Update the successfully associated trajectories using Kalman filtering, and add the unassociated trajectories to a set. In the middle, detection boxes that were not successfully associated are placed into a collection. middle.
5. The adaptive multi-target tracking method based on segmented trajectories as described in claim 1, characterized in that, In step S5, the trajectory set is... With the set of detection boxes The specific steps to establish the association are as follows: S5.1, For the set of trajectories With the set of detection boxes To perform association, select the cosine feature distance between the trajectory and the detection box. As the final cost matrix; S5.
2. Update the successfully associated trajectories using Kalman filtering, and add the unassociated trajectories to the list. In the set, detection boxes that were not successfully associated are considered as background.
6. The adaptive multi-target tracking method based on segmented trajectories as described in claim 1, characterized in that, In step S6, the trajectory appearance features are updated in real time, and the calculation formula is as follows: ,in It is the appearance feature vector of the currently matched and detected target. This is the updated trajectory appearance feature vector. The confidence score of the target being detected. Update the weight parameters for the features.
7. The adaptive multi-target tracking method based on segmented trajectories as described in claim 1, characterized in that, In step S7, the Kalman filter parameters and trajectory appearance features are updated for the associated trajectories. For those that were not successfully associated... The trajectory in the set is put into the set. In the context of sets that failed to be successfully associated... The trajectory in the data is retained; if it is not successfully associated within 30 frames, the trajectory is deleted. For sets of unassociated data... The detection box in the image is used to treat the newly appearing target and initialize its trajectory.