Multi-target tracking method, device and equipment based on Transform and trajectory query group and medium
By using trajectory query groups in the multi-objective tracking method, multiple trajectory query features representing different occlusion levels are designed, which solves the problem that a single Track Query cannot adapt to the changes in complex scene characteristics, and improves tracking accuracy and stability.
Patent Information
- Application Number
- CN202510199521.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-02-24
AI Technical Summary
In complex scenarios, the existing multi-objective tracking method cannot adapt to feature changes by relying on a single Track Query, resulting in target loss and insufficient tracking accuracy.
Using a multi-objective tracking method based on Transformer, each tracked target is designed as multiple track query features that represent different occlusion levels using the Track Query Group (TG). It is associated and detected by multiple features within the trajectory query group, and dynamically adjusts the target tracking mechanism to adapt to target occlusion and state changes.
It improves the accuracy of multi-object tracking in complex scenarios, reduces target loss, enhances the perception of target state, and improves the stability and continuity of tracking.
Smart Images

Figure CN120279119A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to a multi-object tracking method, device, equipment and medium based on Transformer and trajectory query group. Background Art
[0002] In recent years, algorithms based on deep learning have been widely used to solve many problems in real life. Especially in the field of computer vision, tasks such as image classification, object detection and object tracking now rely on deep learning algorithms to achieve, and related technologies are still continuously and rapidly iterating. As a downstream task of object detection, object tracking is a basic but important computer vision task, and its core goal is to associate and match the detected objects in time series. Usually, object tracking technology is combined with object detection technology, which is not only used to integrate the motion trajectories of objects and predict their motion trends, but also can effectively filter out false detection results through trajectory matching.
[0003] Multi-Object Tracking (MOT), as a downstream task of object detection, aims to detect multiple interested objects in a scene, assign them unique IDs, and keep their identities unchanged in the video stream over time. Multi-object tracking technology has wide application value in fields such as security monitoring, intelligent transportation, military applications, and autonomous driving. However, although multi-object tracking algorithms have developed rapidly, they still face many challenges in practical applications, which affect the accuracy of tracking. For example, the occlusion problem between objects: Most multi-object tracking methods rely on image sequences collected by sensorless cameras, and in crowded scenes, the mutual occlusion of multiple objects will greatly interfere with the detection and tracking effects. In addition, problems such as the appearance and disappearance of objects in the scene, the frequent switching of IDs between objects, and the uncertainty of object motion trajectories also bring great difficulties to trajectory processing. Solving these problems has become a key direction for the development of multi-object tracking technology.
[0004] With DETR (a Transformer-based end-to-end object detector) bringing the Transformer architecture into the field of computer vision, Transformer-based algorithms provide a truly end-to-end solution for multi-target tracking tasks. In current Transformer-based tracking methods, a Track Query (track query feature) is usually associated with a tracked target. However, when the target is in a crowded scene, the movement of the target will often switch, resulting in the inability of the Track Query to update the content features in time, thereby losing the target. It can be seen that relying solely on a single TrackQuery for tracking in complex scenes is insufficient and cannot meet the challenges brought by feature transformation. Therefore, how to improve the tracking accuracy of multi-target tracking in complex scenes is a technical problem that needs to be solved urgently in the present invention. Summary of the invention
[0005] Based on the above technical problems, the present invention provides a multi-target tracking method, device, equipment and medium based on Transformer and trajectory query group, aiming to improve the tracking accuracy of multi-target tracking in complex scenes and avoid target loss due to the inability of a single TrackQuery to adapt to feature changes.
[0006] A first aspect of the present invention provides a multi-target tracking method based on Transformer and trajectory query group, the method comprising: Inputting the current image frame in the sample video stream into a multi-target tracking model to be trained, wherein the multi-target tracking model to be trained comprises at least: a backbone network, a Transformer encoder and a Transformer decoder; Processing the current image frame through the backbone network and the Transformer encoder to obtain current image coding features; In the case where a target is detected in an image frame previous to the current image frame, the current image coding feature, the detection query feature, and a trajectory query group corresponding to the tracked target in the current image frame are input into the Transformer decoder to obtain one or more prediction embedding results corresponding to multiple targets respectively; each tracked target corresponds to a trajectory query group, and the trajectory query group includes: multiple trajectory query features representing different occlusion levels; Determine an output embedding result corresponding to the multiple targets respectively from one or more predicted embedding results corresponding to the multiple targets respectively; Determine a sample multi-target tracking result corresponding to the current image frame based on the output embedding results corresponding to the multiple targets respectively; Based on the sample multi-object tracking results corresponding to each image frame in the sample video stream and the multi-object tracking result labels corresponding to each image frame, train the multi-object tracking model to be trained to obtain a trained multi-object tracking model; Input the video stream to be detected into the trained multi-object tracking model to obtain the multi-object tracking results corresponding to the video stream to be detected.
[0007] The second aspect of the present invention provides a multi-object tracking device based on Transformer and trajectory query groups. The device includes: An image input module for inputting the current image frame in the sample video stream into the multi-object tracking model to be trained, where the multi-object tracking model to be trained at least includes: a backbone network, a Transformer encoder, and a Transformer decoder; An image encoding module for processing the current image frame through the backbone network and the Transformer encoder to obtain the current image encoding features; An embedding prediction module for, when a target is detected in the previous image frame of the current image frame, inputting the current image encoding features, detection query features, and the trajectory query groups corresponding to the tracked targets in the current image frame into the Transformer decoder to obtain one or more predicted embedding results corresponding to each of the multiple targets; each tracked target corresponds to a trajectory query group, and the trajectory query group includes: multiple trajectory query features representing different occlusion levels; An embedding output module for determining one output embedding result corresponding to each of the multiple targets from one or more predicted embedding results corresponding to each of the multiple targets; A tracking prediction module for determining the sample multi-object tracking results corresponding to the current image frame based on the output embedding results corresponding to each of the multiple targets; A model training module for training the multi-object tracking model to be trained based on the sample multi-object tracking results corresponding to each image frame in the sample video stream and the multi-object tracking result labels corresponding to each image frame to obtain a trained multi-object tracking model; A tracking detection module for inputting the video stream to be detected into the trained multi-object tracking model to obtain the multi-object tracking results corresponding to the video stream to be detected.
[0008] The third aspect of the present invention provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the multi-object tracking method based on Transformer and trajectory query groups in the first aspect of the embodiments of the present invention.
[0009] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the multi-object tracking method based on Transformer and trajectory query group in the first aspect of the embodiments of the present invention.
[0010] In the multi-object tracking method based on Transformer and trajectory query group provided by the present invention, each image frame in the video stream is sequentially processed by a Transformer encoder and a Transformer decoder to perform tracking detection on multiple objects in the image frame. In the Transformer decoder, multiple objects are associated and detected through a trajectory query group (Track QueryGroup, TG) and a detection query feature (Detect Query), and the objects are tracked with the trajectory query group as the tracking unit. A trajectory query group is designed for each tracked object, and each trajectory query includes multiple trajectory query features (Track Query) representing different occlusion levels. Each trajectory query feature is responsible for identifying the state of the object at a specific occlusion level, so as to identify the same tracked object at different occlusion levels through the trajectory query group, so that the Track Query included in each trajectory query group can handle the association problem of the same object at different occlusion levels, thereby increasing the categorical representation of the object at different occlusion levels in a complex scene; the present invention proposes a mechanism capable of dynamically adjusting object tracking for the problems of object occlusion and state change, and designs a categorical representation method capable of adapting to different occlusion levels of the object to adapt to the frequent switching of different states of the object in a complex scene, timely update the content features of the object, and enhance the perception ability of multi-object tracking for the object state in a complex scene, thereby improving the tracking accuracy of multi-object tracking in a complex scene and avoiding the loss of the object caused by the inability of a single Track Query to adapt to feature changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 is a flowchart of the steps of a multi-object tracking method based on Transformer and trajectory query group shown in an embodiment of the present invention; Figure 2 is a schematic diagram of a trajectory query group shown in an embodiment of the present invention; Figure 3 Schematic diagram of a trajectory query group updater shown in an embodiment of the present invention; Figure 4 Schematic diagram of a position predictor shown in an embodiment of the present invention; Figure 5 Flowchart of a multi-object tracking model based on a trajectory query group shown in an embodiment of the present invention; Figure 6 Schematic diagram of the application of a multi-object tracking task in a street scene crowd shown in an embodiment of the present invention; Figure 7 Flowchart of a model training shown in an embodiment of the present invention; Figure 8 Flowchart of a model inference shown in an embodiment of the present invention; Figure 9 Schematic diagram of the visualization of street scene crowd tracking shown in an embodiment of the present invention; Figure 10 Schematic diagram of the application of a multi-object tracking task in a dance scene shown in an embodiment of the present invention; Figure 11 Schematic diagram of the visualization of tracking in a dance scene shown in an embodiment of the present invention; Figure 12 Block diagram of a multi-object tracking device based on Transformer and trajectory query group provided in an embodiment of the present invention. Detailed implementation manners
[0013] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0014] Please refer to Figure 1 , Figure 1 which is a step flowchart of a multi-object tracking method based on Transformer and trajectory query group shown in an embodiment of the present invention. As Figure 1 shown, the multi-object tracking method based on Transformer and trajectory query group provided in this embodiment at least includes the following steps: Step S11: Input the current image frame in the sample video stream into the multi-object tracking model to be trained, and the multi-object tracking model to be trained at least includes: a backbone network, a Transformer encoder, and a Transformer decoder.
[0015] In this embodiment, a sample data set is prepared for training the multi-object tracking model. The sample data set includes: a plurality of sample video streams, and the label corresponding to each sample video stream includes: the multi-object tracking result label corresponding to each image frame of the sample video stream. Among them, the sample video stream is a video stream used for training the multi-object tracking model.
[0016] In this embodiment, the current image frame in the sample video stream can be input into the multi-object tracking model to be trained for tracking detection processing, so as to input each image frame in the sample video stream into the multi-object tracking model to be trained.
[0017] Step S12: Process the current image frame through the backbone network and the Transformer encoder to obtain the current image encoded feature.
[0018] In this embodiment, after the current image frame is input into the multi-object tracking model to be trained, the current image frame is processed through the backbone network and the Transformer encoder to extract two-dimensional features, and the current image encoded feature is obtained as one of the inputs to the Transformer decoder. In a specific example, it can be to first input the current image frame into the backbone network for feature extraction to obtain the initial image feature, and then input the initial image feature into the Transformer encoder to obtain the current image encoded feature output by the Transformer encoder.
[0019] Step S13: When a target is detected in the previous image frame of the current image frame, input the current image encoded feature, the detection query feature, and the trajectory query group corresponding to the tracked target in the current image frame into the Transformer decoder to obtain one or more predicted embedding results corresponding to each target.
[0020] In this embodiment, when a target is detected in the previous image frame of the current image frame, that is, when there is a tracked target in the current image frame, the current image encoded feature, the detection query feature, and the trajectory query group corresponding to the tracked target in the current image frame can be input into the Transformer decoder to obtain one or more predicted embedding results corresponding to each target output by the Transformer decoder.
[0021] Among them, in this embodiment, a fixed number (which can be freely set) of detection query features (Detect Query) are used in each image frame and input into the Transformer decoder to detect newly emerging targets (i.e., new targets) in the current image frame. It can be understood that the detection query feature is a learnable embedding.
[0022] In the traditional method, each tracked target is only associated with one Track Query, which is effective when the occlusion level of the target is stable; however, in complex scenarios involving different occlusion levels, this method often fails because a single Track Query is difficult to handle this challenge, resulting in a large number of ID switches. Based on this, for each tracked target, each tracked target in this embodiment corresponds to a track query group, and each track query group includes: multiple Track Queries representing different occlusion levels, and multiple Track Queries in each track query group jointly predict the same target.
[0023] In this embodiment, the Transformer decoder can output corresponding predict embeddings for each of the detected multiple targets. Among them, a target among the multiple targets can correspond to one predict embedding result or multiple predict embedding results.
[0024] Step S14: Determine one output embedding result corresponding to each of the multiple targets from one or more predict embedding results corresponding to each of the multiple targets.
[0025] In this embodiment, after obtaining one or more predict embedding results corresponding to each of the multiple targets, one output embedding result corresponding to each of the multiple targets can be determined from one or more predict embedding results corresponding to each of the multiple targets. Among them, when a target corresponds to multiple predict embedding results, according to the multiple predict embedding results, the most suitable predict embedding result representing the current state of the target can be selected from the multiple predict embedding results as one output embedding result (output embeddings) corresponding to the target. When a target corresponds to one predict embedding result, directly use this predict embedding result as one output embedding result corresponding to the target.
[0026] Step S15: Based on the output embedding results corresponding to each of the multiple targets, determine the sample multi-target tracking result corresponding to the current image frame.
[0027] In this embodiment, the sample multi-target tracking result corresponding to the current image frame can be determined based on the output embedding results corresponding to each of the multiple targets. The sample multi-target tracking result is the multi-target tracking result output by the multi-target tracking model during training. The multi-target tracking result at least includes: the confidence levels and bounding boxes corresponding to each of the multiple targets. In a specific example, the Track Queries in the track query group are input into the Transformer decoder to obtain embeddings, and then processed by an MLP to predict the corresponding confidence scores and bounding boxes.
[0028] Step S16: Based on the sample multi-object tracking results corresponding to each image frame in the sample video stream and the multi-object tracking result labels corresponding to each image frame, train the multi-object tracking model to be trained to obtain a trained multi-object tracking model.
[0029] In this embodiment, each image frame in the sample video stream can be processed based on the above steps to obtain the sample multi-object tracking results corresponding to each image frame in the sample video stream. Then, based on the sample multi-object tracking results corresponding to each image frame in the sample video stream and the multi-object tracking result labels corresponding to each image frame in the sample video stream, train the multi-object tracking model to be trained to update the network parameters of the backbone network to be trained, the Transformer encoder to be trained, and the Transformer decoder to be trained, so as to obtain a trained multi-object tracking model.
[0030] Step S17: Input the video stream to be detected into the trained multi-object tracking model to obtain the multi-object tracking results corresponding to the video stream to be detected.
[0031] In this embodiment, the trained multi-object tracking model is used for multi-object tracking of a single image or a video stream. During the application of the trained multi-object tracking model, in the case of multi-object tracking of a video stream, the video stream to be detected can be input into the trained multi-object tracking model to obtain the multi-object tracking results corresponding to the video stream to be detected output by the trained multi-object tracking model. The multi-object tracking results corresponding to the video stream to be detected include: the multi-object tracking results corresponding to each image frame in the video stream to be detected. Among them, the video stream to be detected is the video stream that needs multi-object tracking.
[0032] In this embodiment, multiple targets are associated and detected by means of a trajectory query group and detection query features. Taking the trajectory query group as the tracking unit, a trajectory query group is designed for each tracked target. Each trajectory query includes multiple trajectory query features representing different occlusion levels, that is, different trajectory query features are designed for the target under different occlusion levels. Each trajectory query feature is responsible for identifying the state of the target under a specific occlusion level, so as to identify the same tracked target under different occlusion levels through the trajectory query group, so that the Track Queries included in each trajectory query group can handle the association problem of the same target under different occlusion levels, thereby increasing the categorical representation of the target under different occlusion levels in a complex scenario. Compared with the traditional method of using a single Track Query for tracking, the association using the trajectory query group in this embodiment is more robust, because the multi-object tracking model can further utilize various appearance features at different occlusion levels in the trajectory, and adapt to the change of the occlusion state of the target by selecting the most suitable predicted embedding result, thereby improving the accuracy of the tracking result of multi-object tracking in a complex scenario and alleviating the challenges brought by the complex scenario.
[0033] Combined with the above embodiments, in one implementation manner, the present invention further provides a multi-object tracking method based on a Transformer and a trajectory query group. In this method, in addition to the above steps, steps S21 to S23 may further be included: Step S21: In the case where no target is detected in the previous image frame of the current image frame, or, in the case where the current image frame is the first image frame, input the current image encoded feature and the detection query feature into the Transformer decoder to obtain a predicted embedding result corresponding to a new target in the current image frame.
[0034] In this embodiment, in the case where no target is detected in the previous image frame of the current image frame, or, in the case where the current image frame is the first image frame, that is to say, in the case where there is no tracked target in the current image frame, at this time, only the current image encoded feature and the detection query feature are input into the Transformer decoder to obtain a predicted embedding result corresponding to a new target in the current image frame output by the Transformer decoder. This is because the target detected based on the detection query feature (Detectquery) is a new target. Therefore, at this time, the output of the Transformer decoder is the predicted embedding result corresponding to the new target in the current image frame.
[0035] Step S22: Based on the predicted embedding result corresponding to the new target, determine the sample multi-object tracking result corresponding to the current image frame.
[0036] In this embodiment, each new target has only one predicted embedding result. Therefore, the output embedding result corresponding to the new target is the predicted embedding result corresponding to the new target. Thus, the sample multi-object tracking result corresponding to the current image frame can be directly determined based on the predicted embedding result corresponding to the new target.
[0037] Step S23: In the next image frame of the current image frame, use the new target as the tracked target, and initialize the trajectory query group corresponding to the tracked target in the next frame of image based on the predicted embedding result corresponding to the new target.
[0038] In this embodiment, for each newly detected target, that is, for each new target, a trajectory query group is initialized for this new target to be used as a dedicated Query set to track this new target. Specifically, after a new target is detected in the current image frame, in the next image frame of the current image frame, use the detected new target as the tracked target, and initialize the trajectory query group corresponding to the tracked target in the next frame of image based on the predicted embedding result corresponding to the new target in the current image frame. That is, effectively convert the predicted embedding result corresponding to the new target in the current image frame into the trajectory query feature in the trajectory query group for subsequent tracking tasks. In one example, the trajectory query group is initialized by combining the output content embedding and the position embedding.
[0039] Combining the above embodiments, in one implementation manner, the present invention further provides a multi-object tracking method based on Transformer and trajectory query group. In this method, the multiple targets include: tracked targets and new targets, or, multiple tracked targets. Among them, when the target is a tracked target, each target corresponds to multiple predicted embedding results, and each predicted embedding result is the predicted embedding result corresponding to the trajectory query features of different occlusion levels; when the target is a new target, each target corresponds to one predicted embedding result. And, the "determining one output embedding result corresponding to each of the multiple targets from the multiple predicted embedding results corresponding to each of the multiple targets" in step S14 above specifically includes step S31 and step S32: Step S31: Determine the loss between the multiple predicted embedding results corresponding to each target respectively and the label embedding result corresponding to each target.
[0040] In this embodiment, for the multiple predicted embedding results corresponding to each tracked target respectively, at this time, the loss between the multiple predicted embedding results corresponding to each target respectively and the label embedding result corresponding to each target can be determined.
[0041] Step S32: Determine the predicted embedding result with the minimum loss as one output embedding result corresponding to each target.
[0042] In this embodiment, the predicted embedding result corresponding to the calculated minimum loss is determined as an output embedding result corresponding to each tracked target.
[0043] For example, if each tracked target corresponds to trajectory query features with three different occlusion levels respectively, then each tracked target corresponds to predicted embedding results of trajectory query features with three different occlusion levels. At this time, the first loss value between the predicted embedding result of the trajectory query feature of the first occlusion level and the label embedding result corresponding to this target can be calculated, the first loss value between the predicted embedding result of the trajectory query feature of the second occlusion level and the label embedding result corresponding to this target can be calculated, and the third loss value between the predicted embedding result of the trajectory query feature of the third occlusion level and the label embedding result corresponding to this target can be calculated. Then, the minimum loss value is determined from the first loss value, the second loss value, and the third loss value, and then the predicted embedding result corresponding to the minimum loss value is determined as an output embedding result corresponding to this tracked target.
[0044] In an alternative example, when calculating the loss between multiple predicted embedding results corresponding to each target respectively and the label embedding result corresponding to each target, the total loss between the predicted embedding result and the label embedding result is calculated, where the total loss includes: classification loss and bounding box loss. However, since different occlusion levels are determined according to the confidence, it is not fair enough to simply perform equal-weight processing on the classification loss (calculating the loss based on the confidence) and the bounding box loss. Therefore, in order to reduce the influence of the classification loss, this embodiment can assign a higher weight to the bounding box loss to give priority to the trajectory query features (queries) that can predict more accurate bounding boxes. Further, in an alternative example, the bounding box loss can include: L1 loss and GIoU loss. Therefore, different weights can be assigned to each loss type. For example: L1 = 5, GIoU = 5, classification class = 1.
[0045] Combined with the above embodiments, in one implementation manner, the present invention also provides a multi-object tracking method based on Transformer and trajectory query groups. In this embodiment, the multi-object tracking model generates a set of trajectory query groups for each tracked target (i.e., trajectory), and the trajectory query groups store the content features of the tracked target at different occlusion levels, and the different occlusion levels include at least three occlusion levels.
[0046] In an optional example, this embodiment classifies different occlusion levels into three categories: fully exposed, slightly occluded, and severely occluded; the three occlusion levels respectively correspond to different occlusion thresholds. Since the occlusion level of the target changes continuously in a complex scene, the detection confidence of the trajectory query features in the trajectory query group also fluctuates accordingly. Generally, when the target is fully exposed, the confidence (the query confidence output by the Transformer decoder, that is, the query prediction confidence score) is relatively high; when slightly occluded, the confidence decreases; and when severely occluded, the confidence is the lowest. Therefore, this embodiment sets different occlusion thresholds for the three occlusion levels in the trajectory query group, that is, the three occlusion levels respectively correspond to different occlusion thresholds.
[0047] For example, for the three types of queries in the trajectory query group: fully exposed trajectory query features, slightly occluded trajectory query features, and severely occluded trajectory query features, they respectively correspond to occlusion thresholds 、 and . During the tracking process, according to the confidence score of the output embedding result corresponding to the target and the occlusion threshold, the output embedding result can be classified into a low-score embedding, a medium-score embedding, and a high-score embedding to respectively represent that the target is fully exposed, slightly occluded, and severely occluded. For example, the three occlusion thresholds are: = 0.85, = 0.7, = 0.5. When the confidence corresponding to the output embedding result of the new target is 0.75, <0.75< , then the new target is classified as slightly occluded, and the output embedding result corresponding to the new target is saved to the trajectory query features corresponding to slightly occluded in the trajectory query group corresponding to the target.
[0048] In this embodiment, during the process of the trained multi-object tracking model performing tracking detection on the to-be-detected video stream, when obtaining the predicted embedding results respectively corresponding to the three occlusion levels of the target, the predicted embedding result with the highest confidence is determined as the output embedding result corresponding to the target.
[0049] That is to say, during the model inference process of the trained multi-object tracking model, when the predicted embedding results corresponding to the three occlusion levels of the target are obtained, in this embodiment, the predicted embedding result with the highest confidence is preferentially selected as the output. That is, the predicted embedding result with the highest confidence is determined as the output embedding result corresponding to the target, and this output embedding result will be used for the update of the subsequent trajectory query group. When the confidence of the output embedding result falls within the occlusion range corresponding to other occlusion levels, it indicates that the occlusion level of the target has changed. In this case, if the new occlusion level is not yet represented in the trajectory query group, it will be added to the group; otherwise, the embedding previously existing for the new occlusion level will be replaced as the trajectory query feature for the new occlusion level.
[0050] For example, during the inference process, the confidence of the predicted embedding result corresponding to the fully exposed target is the highest, which is 0.75, and <0.75< , at this time it falls within the occlusion range corresponding to slight occlusion, which indicates that the occlusion level of the target has changed in the current image frame, from fully exposed in the previous image frame to slightly occluded in the current image frame. If it is determined that there is no trajectory query feature for slight occlusion in the trajectory query group of this target, the predicted embedding result will be added to the trajectory query group as the trajectory query feature for slight occlusion; if it is determined that there is a trajectory query feature for slight occlusion in the trajectory query group of this target, the predicted embedding result will replace the trajectory query feature for slight occlusion in the trajectory query group.
[0051] In one embodiment, as Figure 2 shown, Figure 2 is a schematic diagram of a trajectory query group shown in an embodiment of the present invention. In Figure 2 , the multi-object tracking model (TGFormer) generates a set of trajectory query features (trackqueries) for each target i, and the trajectory query group TG i stores the content features of the tracked target at different occlusion levels, and each group contains at most three Track Queries. Specifically, this method classifies the occlusion levels into three categories: fully exposed scene, slightly occluded scene, and heavily occluded scene, and each type of trajectory query feature is responsible for detecting the target at a specific occlusion level. Since the occlusion level of the target changes continuously in a complex scene, the detection confidence of the Queries within the group also fluctuates accordingly. Generally, when the target is fully exposed, the confidence is relatively high; when it is slightly occluded, the confidence decreases; and when it is heavily occluded, the confidence is the lowest.
[0052] In the related art, since DETR introduced the Transformer model into the field of computer vision, especially object detection, Query (the embedding used to detect objects in the Transformer Decoder in DETR) usually combines the content features and location information of objects in a unified manner. Although this combination is convenient, it may cause the model to be unable to fully utilize the independent information of different features to optimize the positioning accuracy. Based on this, in combination with the above embodiments, in one implementation, the present invention also provides a multi-object tracking method based on Transformer and trajectory query groups. In this embodiment, the content features and location information in the query features (including: detection query features and trajectory query features) are separated. Specifically, these trajectory query features share the same location information but have independent content embeddings adapted to different occlusion levels. That is to say, the trajectory query features include: query content embeddings and query location embeddings, and the output embedding results include: output content embeddings and output location embeddings; different query content embeddings corresponding to the same object share the same query location embedding.
[0053] Since there are multiple trajectory query features in the trajectory query group in this embodiment, the traditional method of updating trajectory query features cannot be directly used (in previous Transformer-based trackers, if an object is detected, the Detect Query of the previous frame will be passed as the Track Query to the next frame to gradually update the content features of the Query). Therefore, a special updater is set for the trajectory query group in this embodiment: the Trajectory Query Group Updater (TG Updater), to update the query content embeddings in the trajectory query group to further complete the update of the trajectory query group. Specifically, in this embodiment, in addition to the above steps, steps S41 to S44 may also be included: Step S41: For each object among multiple objects, input the output content embedding corresponding to the object in the current image frame, the query content embedding in the trajectory query group corresponding to the object in the current image frame, and the long-term memory embedding corresponding to the object into the trajectory query group updater to obtain the updated query content embedding corresponding to the object.
[0054] In this embodiment, for each object among the multiple objects detected in the current image frame, the output content embedding corresponding to the object in the current image frame, the query content embedding in the trajectory query group corresponding to the object in the current image frame, and the long-term memory embedding corresponding to the object can be input into the trajectory query group updater, so that the trajectory query group updater updates the query content embedding in the trajectory query group corresponding to the object to obtain the updated query content embedding corresponding to the object output by the trajectory query group updater.
[0055] Step S42: For each of the multiple targets, at least embed the output position corresponding to the target in the current image frame into the position predictor to obtain the predicted query position embedding corresponding to the target in the next image frame.
[0056] In this embodiment, for each of the multiple targets detected in the current image frame, at least the output position corresponding to the target in the current image frame can be embedded into the position predictor to obtain the predicted query position embedding corresponding to the target in the next image frame.
[0057] Step S43: Based on the updated query content embedding corresponding to the target and the predicted query position embedding corresponding to the target in the next image frame, obtain the updated trajectory query group corresponding to the target.
[0058] In this embodiment, after obtaining the updated query content embedding corresponding to the target and the predicted query position embedding corresponding to the target in the next image frame, the updated trajectory query features corresponding to the target can be obtained based on the updated query content embedding corresponding to the target and the predicted query position embedding corresponding to the target in the next image frame, so as to obtain the updated trajectory query group corresponding to the target.
[0059] Step S44: Use the updated trajectory query group as the trajectory query group corresponding to the tracked target in the next image frame.
[0060] In this embodiment, after obtaining the updated trajectory query group corresponding to the target, the updated trajectory query group corresponding to the target can be used as the trajectory query group corresponding to the target (i.e., the tracked target) in the next image frame.
[0061] In this embodiment, by independently processing the content feature and the position information, making full use of their independence, the tracking performance is further improved. The content feature and the position information in the Query are separated, and explicit position prior information is provided to the decoder to help determine the position of the target more quickly. This improvement makes the target localization process more explicit and efficient.
[0062] In an alternative embodiment, when the target is a tracked target, the long-term memory embedding corresponding to the target is the accumulation of the output content embeddings corresponding to the target in the current image frame and the previous image frames; when the target is a new target, the long-term memory embedding corresponding to the target is the output content embedding corresponding to the target in the current image frame.
[0063] Specifically, this method uses the long-term memory embedding to represent the stable features of the target to maintain the information of the tracked target for a longer time. When a new target is detected, the output content embedding of the new target is initialized as the long-term memory and apply the running average method with exponentially decaying weights to update this long-term memory: ; where λ represents the weight decay exponent, which can be set to 0.01, because it is assumed that the memory of the target in consecutive frames is smooth and changes stably.
[0064] Combined with the above embodiments, in one implementation manner, the present invention further provides a multi-object tracking method based on Transformer and trajectory query groups. In this embodiment, the "embedding the output content corresponding to the target in the current image frame, the query content in the trajectory query group corresponding to the target in the current image frame, and the long-term memory embedding corresponding to the target, and inputting them into the trajectory query group updater to obtain the updated query content embedding corresponding to the target" in step S41 above specifically includes steps S51 to S54: Step S51: Concatenate the query content embeddings of different occlusion levels in the trajectory query group corresponding to the target in the current image frame with the output content embedding corresponding to the target in the current image frame, and then input them into the short-term memory aggregator to obtain the query enhanced content embeddings of different occlusion levels corresponding to the target.
[0065] In this embodiment, the query content embedding in the trajectory query group corresponding to the target in the current image frame includes: the query content embeddings of different occlusion levels in the trajectory query group corresponding to the target in the current image frame. The query content embeddings of different occlusion levels in the trajectory query group corresponding to the target in the current image frame can be concatenated with the output content embedding corresponding to the target in the current image frame respectively to obtain the concatenation results corresponding to different occlusion levels. Then, input the concatenation results corresponding to different occlusion levels into the Short-Term Memory Aggregator for feature enhancement to obtain the query enhanced content embeddings of different occlusion levels output by the short-term memory aggregator. These embeddings serve the same target but contain unique features under different occlusion levels.
[0066] In an alternative implementation manner, the short-term memory aggregator consists of a MIP (multi-layer perceptron), and the query enhanced content embeddings of different occlusion levels corresponding to the target can be obtained through the following formula, aiming to enhance the target feature representation by concatenating multi-frame features, and at the same time combine the high-score, medium-score, and low-score embeddings to extract and fuse the target features in different scenarios.
[0067] ; where , , They are the fully-exposed query-enhanced content embedding, the slightly-occluded query-enhanced content embedding, and the severely-occluded query-enhanced content embedding respectively; MIP is the short-term memory aggregator, which is the output content embedding corresponding to the target in the current image frame, , , which are the fully-exposed query content embedding in the trajectory query group corresponding to the target in the current image frame, the slightly-occluded query content embedding in the trajectory query group corresponding to the target in the current image frame, and the severely-occluded query content embedding in the trajectory query group corresponding to the target in the current image frame respectively.
[0068] Step S52: Concatenate the query-enhanced content embeddings with different occlusion levels to obtain a concatenated content embedding. Use the concatenated content embedding as the key and value, and use the query-enhanced content embeddings with different occlusion levels as queries respectively to perform an occlusion-aware memory attention operation to obtain intermediate content embeddings corresponding to different occlusion levels.
[0069] In this embodiment, to distinguish these unique features ( , , ), a multi-head attention module called Occlusion-Aware Memory-Attention is designed. After obtaining the query-enhanced content embeddings with different occlusion levels, concatenate the query-enhanced content embeddings with different occlusion levels to obtain a concatenated content embedding , and use the concatenated content embedding as the key and value. Also, in this embodiment, use the query-enhanced content embeddings with different occlusion levels ( , , ) as queries respectively, and perform an occlusion-aware memory attention operation based on the key, value, and queries to obtain intermediate content embeddings corresponding to different occlusion levels. This intermediate content embedding represents the unique features of the target at different occlusion levels.
[0070] Step S53: Combine the intermediate content embeddings corresponding to different occlusion levels with the long-term memory embedding corresponding to the target respectively to obtain updated content embeddings corresponding to different occlusion levels.
[0071] In this embodiment, after obtaining the intermediate content embeddings corresponding to different occlusion levels, the intermediate content embeddings corresponding to different occlusion levels can be combined with the long-term memory embedding corresponding to the target respectively to obtain the updated content embeddings corresponding to different occlusion levels. Specifically, the intermediate content embeddings corresponding to different occlusion levels can be combined with the long-term memory embedding corresponding to the target respectively by using residual connections. For example, the updated content embeddings corresponding to different occlusion levels can be obtained according to the following formula: ; where , , are the updated content embeddings corresponding to full exposure, slight occlusion, and severe occlusion respectively; Attention is the occlusion-aware memory attention operation, is the concatenated content embedding, , , are the query-enhanced content embeddings for full exposure, slight occlusion, and severe occlusion respectively; is the long-term memory embedding corresponding to the target.
[0072] Step S54: Based on the updated content embeddings corresponding to different occlusion levels, obtain the updated query content embedding corresponding to the target.
[0073] In this embodiment, after obtaining the updated content embeddings corresponding to different occlusion levels, the updated query content embedding corresponding to the target can be obtained based on the updated content embeddings corresponding to different occlusion levels: .
[0074] Furthermore, the updated query content embeddings corresponding to the target share the same predicted query position embedding corresponding to the target in the next image frame, and together form an updated trajectory query group, which is input into the next image frame for tracking.
[0075] In this embodiment, through the above design of the trajectory query group updater, the association ability of the multi-target tracking model to the target is greatly improved in complex scenarios, enabling it to more accurately identify and track occluded or targets in different states. In addition, by assigning exclusive feature representations to the target at different occlusion levels, this design effectively reduces the ID switching problem caused by occlusion and appearance changes, thus significantly improving the stability and continuity of tracking.
[0076] In one embodiment, as Figure 3 shown, Figure 3 is a schematic diagram of the trajectory query group updater shown in an embodiment of the present invention. In Figure 3Among them, the content embeddings of object i include: the output content embeddings corresponding to the object in the current image frame , the query content embeddings in the trajectory query group corresponding to the object in the current image frame , and the long-term memory embeddings corresponding to the object ( is based on updated). For each occlusion type in , , are respectively concatenated with the output content embeddings corresponding to the object in the current image frame and input into the short-term memory aggregator to obtain , , , which are respectively: fully exposed query-enhanced content embeddings, slightly occluded query-enhanced content embeddings, and severely occluded query-enhanced content embeddings.
[0077] Then, , , are concatenated as , to be used as the key and value, and then the query-enhanced content embeddings of different occlusion levels ( , , ) are respectively used as the queries: Q1, Q2, and Q3, to promote feature interaction in the occlusion-aware memory attention, obtaining the intermediate content embeddings corresponding to different occlusion levels. Finally, the intermediate content embeddings corresponding to different occlusion levels are combined with the long-term memory embeddings corresponding to the object to obtain the updated content embeddings corresponding to different occlusion levels .
[0078] For a moving object in tracking, it is inaccurate to directly use the position information of the previous frame as the positioning prior for the next frame. To address this issue, in combination with the above embodiments, in one implementation, the present invention also provides a multi-object tracking method based on Transformer and a trajectory query group. In this method, there is a simple and efficient Position Predictor. The Position Predictor forms a Track Query by combining the predicted position with the content embeddings, enabling the multi-object tracking model to more accurately locate the object in a dynamic scenario. Specifically, in this embodiment, the above step S42 may specifically include step S61 and step S62: Step S61: Embed the output position corresponding to the target in the current image frame and the historical output position corresponding to the target in the previous image frame into the position predictor, to obtain the predicted query position embedding corresponding to the target in the next image frame and the historical output position embedding corresponding to the target in the current image frame.
[0079] In this embodiment, the output position embedding corresponding to the target in the current image frame and the historical output position embedding corresponding to the target in the previous image frame can be input into the position predictor to obtain the following outputs of the position predictor: the predicted query position embedding corresponding to the target in the next image frame and the historical output position embedding corresponding to the target in the current image frame. Among them, the historical output position corresponding to the target in the previous image frame is the set of output position embeddings corresponding to the target in the previous image frame and all previous image frames. In an alternative embodiment, the position predictor can be implemented by an LSTM network, an RNN (Recurrent Neural Network), or a Transformer.
[0080] Step S62: The historical output position embedding corresponding to the target in the current image frame is used to provide continuous temporal context to assist the position predictor in outputting subsequent predicted query position embeddings in the next image frame.
[0081] In this embodiment, the historical output position embedding corresponding to the target in the current image frame is the set of output position embeddings corresponding to the target in the current image frame and all previous image frames. This historical output position embedding corresponding to the target in the current image frame is used to provide continuous temporal context to assist the position predictor in outputting subsequent predicted query position embeddings in the next image frame.
[0082] In one embodiment, as Figure 4 shown, Figure 4 is a schematic diagram of a position predictor shown in an embodiment of the present invention. In Figure 4 , during the tracking process, the model continuously stores the target positions predicted each time. When the number of stored historical target positions exceeds a certain value (set arbitrarily, such as 5 frames, without limitation), the position predictor can start using this data to predict the movement trend of the target. Specifically, these position sequences are input into the position predictor. In Figure 6 , the position predictor is implemented by an LSTM network to predict the position information of the target in the next frame. The LSTM generates two outputs: and . Among them, is the predicted position information of the target in the next frame (i.e., the predicted query position embedding corresponding to the target in the next image frame), and stores the information implicit in the previous sequence to provide continuous temporal context, that is, Embed for the historical output position corresponding to the target in the current image frame. In the next frame, it will assist in object detection, and the output position embedding X corresponding to the next image frame t+1 are input into the LSTM together to continue the prediction. The advantage of this structural design is that it does not need to continuously store all the target positions, and only key information needs to be retained to complete the prediction task.
[0083] In this embodiment, through the position predictor, the position information of the target in the next frame is predicted using the historical trajectory and provided as a position prior to the Query to assist in improving the detection ability. The design of the position predictor significantly improves the positioning accuracy of moving targets during the tracking process, enabling the model to more effectively handle the dynamic changes of targets, especially showing higher accuracy and robustness in the tracking performance of moving targets in complex scenarios. By providing a more efficient position prior information update strategy, the position information of dynamic moving targets can timely reflect the actual motion state of the targets, reducing the decrease in tracking accuracy caused by lagging motion prediction, thereby optimizing the target positioning accuracy. Thus, directly using the position information of the previous frame as the positioning prior for the next frame is avoided because the tracked targets are usually dynamically moving, and this method is inaccurate in many cases, especially for fast-moving targets, where this prior information may lag behind the actual motion of the targets, resulting in a decrease in tracking accuracy.
[0084] In one embodiment, as Figure 5 shown, Figure 5 is a flowchart of a multi-object tracking model based on a trajectory query group shown in an embodiment of the present invention. In Figure 5 , for the frame in the video input stream, first, it is input into the backbone network Backbone and the Transformer encoder Encoder to extract two-dimensional features, which is one of the inputs to the decoder. The multi-object tracking model uses a fixed number of learnable embeddings, namely Detect Queries, to detect newly emerging targets in the current frame, and each tracked target corresponds to a trajectory query group , in each group, the Track Query jointly predicts the same target. Specifically, these queries share the same position information (position embeddings), but have independent content embeddings (content embeddings) adapted to different occlusion levels. At the same time, the multi-object tracking model predicts the position information of the target from the historical trajectory through the position predictor (Pos Predictor) and combines this information with the content embeddings to form a trajectory query group It is input into the decoder. The decoder generates embeddings for all Queries within the group (including DetectQueries and the track query group), and these embeddings are processed by the prediction network to obtain confidence scores and bounding boxes. The model selects the most appropriate prediction based on these results and uses the corresponding decoder output as the output embedding. In the post-processing stage, the output embedding, long-term memory, and the track query group are input into the TG Updater together to update the Queries within the track query group.
[0085] Among them, for long-term stable association, this embodiment proposes a novel updater TG Updater applicable to the track query group TG, and all Track Queries within the group will be continuously updated in each frame through the proposed TG Updater. For the TG Updater, the output content embeddings in the output embedding obtained from the current frame are respectively fused with the Queries within the group through the Short-Term Memory Aggregator. Since the target exhibits different characteristics at different occlusion levels, the Occlusion-Aware Memory Attention mechanism is applied to distinguish the characteristics of the target at each occlusion level, and a long-term memory is maintained for the target and connected to the output of the memory attention mechanism through a residual connection to help the model learn more discriminative representations. The track query group TG is continuously updated in each frame by the above method to ensure that the target is always accurately captured in complex scenarios. In addition, this embodiment introduces a PositionPredictor to enable the tracker to predict the motion trend and help the model more accurately locate the moving target.
[0086] Such as through I t-1 The updated content embedding output by the frame TG Updater and the position embedding output by the position predictor are used to update the track query features in the track query group of the target for input into the Transformer decoder in frame I t frame. Among them, Figure 5 For frame I t-1 frame, the process between the Transformer decoder and the TG Updater and the position predictor is omitted. In fact, it is the same as that in frame I tThe process between the Transformer decoder, the TG Updater, and the position predictor in the frame is the same. First, the predicted embedding results output by the decoder are obtained, and the output embedding results are filtered from the predicted embedding results. Then, the output content embedding in the output embedding results is input into the TG Updater to update the query content embedding, and the output position embedding in the output embedding results is input into the position predictor to output the corresponding predicted query position embedding in the next image frame.
[0087] In an alternative embodiment, multi-object tracking tasks play an important role in security monitoring applications. For example, detecting and tracking pedestrians on street scenes, in shopping malls, and on roads. As Figure 6 shown, Figure 6 is a schematic diagram of the application of a multi-object tracking task in a street scene crowd shown in an embodiment of the present invention. There are dense crowds distributed in the street scene. In security monitoring, it is necessary to detect and continuously track these pedestrians. To complete the security monitoring task, this embodiment needs to train a multi-object tracking model adapted to specific tasks. The process is as Figure 7 shown, Figure 7 is a flowchart of model training shown in an embodiment of the present invention. The detailed steps of this model training are as follows: 1. Construct a training data set. For street scene scenarios, the publicly available data sets MOT17 and MOT20 can usually be used as the training data set. At the same time, if monitoring is required in a specific scenario, the staff needs to construct the training data set by themselves. Use an RGB monocular camera to shoot the scene video for a certain period of time, and then label each task that appears in the video, including the bounding box and ID of the task, etc. The specific data format is: <id>,<bb_left>,<bb_top>,<bb_width>,<bb_height>,<trajectory_conf>,<trajectory_type>,<visibility_ratio> where it represents the frame number, <id>The trajectory ID represents the target, <bb_left> and <bb_top> represent the upper left coordinates of the bounding box, <bb_width> and <bb_height> represent the width and height of the bounding box, <trajectory_conf> represents the flag indicating whether the trajectory is considered (0 means ignored, 1 means activated), <trajectory_type> represents the category of the target, and <visibility_ratio> represents the visible ratio of the target, thus completing the construction of a specific dataset.
[0088] 2. Configure the training parameters. For example, the multi-object tracking model uses ResNet50 as the backbone network and DAB-Deformable-DETR (pre-trained on the COCO dataset) as the detector. It is recommended to train the model on a graphics card with 24G or more video memory, such as NVIDIA GeForce RTX 3090, NVIDIA GeForce RTX 4090, A100, A800, etc. The optimizer uses AdamW, and the initial learning rate is 2.0×10^(-4). During the training process, the model will filter out tracking targets with a score lower than the threshold of 0.5 and an IoU lower than the threshold of 0.5. In the model, each target corresponds to a Track Query Group. When a Detect Query detects a new target, a Track Query Group will be initialized, and the queries will be divided into high score embedding, midscore embedding, and low score embedding according to the confidence level, and the remaining embeddings are initially set to zero vectors. During the tracking process, the detection confidence of the output embedding may change. When the occlusion level of the target changes significantly, resulting in the detection confidence falling within the threshold range of the other two embeddings, the output embedding will be saved as the corresponding embedding. For example, the confidence thresholds are: high = 0.85, mid = 0.7, low = 0.5.
[0089] 3. Start training. Before training the model, the running environment needs to be configured. It is required to use the pytorch module, and the specific version requirements are: pytorch==1.13.1 torchvision==0.14.1 torchaudio==0.13.1 pytorch-cuda=11.7 4. Read the image data. Read the corresponding image data and its ground truth data from the training dataset. The batch size of the training image data loaded by each GPU is 1, and each batch contains a video clip with multiple frames. In each clip, the frames are sampled at random intervals (1 to 10 frames).
[0090] 5. Data augmentation processing. During training, to improve the robustness of the model, image data is augmented, such as cropping and horizontal flipping of the image. Finally, the RGB data is normalized and output as Tensor tensor data for model processing.
[0091] 6. Input the model for inference. The processed image data is input into the model for inference.
[0092] 7. Output the prediction results. The model inference obtains the tracking results, including the detection results of all targets in the image: detection confidence, bounding boxes, and target IDs.
[0093] 8. Calculate the loss and perform backpropagation. During training, calculate the loss between each query in the trajectory group and the corresponding ground truth, including classification loss, L1 loss, and GIoU loss. Select the query with the minimum loss as the final tracking result. To reduce the impact of the classification loss, different weights are assigned to each loss type: L1 = 5, GIoU = 5, class = 1.
[0094] 9. Save the model. After each round of training is completed, save the current model.
[0095] 10. End the training. Train for 130 epochs on the dataset, and reduce the learning rate by 10 times at the 120th epoch. At the 50th, 70th, 90th, and 120th epochs, the number of clipped frames gradually increases from the original 2 frames to 3, 4, 5, and 6 frames. End the training after all epochs are completed.
[0096] 11. Obtain the trained model. After training is completed, obtain the final multi-object tracking model, and this model can be verified on the validation set.
[0097] After training is completed, a saved multi-object tracking model is obtained, which can be put into practical applications. The inference process is as Figure 8 shown, Figure 8 is a model inference flowchart shown in an embodiment of the present invention. The specific inference process is as follows: 1. Obtain image data. In an actual security monitoring scenario, there are two ways to read data: one is to read previously stored historical video data from a database, and the other is to obtain real-time video captured by a camera (such as an RGB monocular camera) and return the image data as the model input.
[0098] 2. Process the image data. Convert the obtained image data into a tensor for model input.
[0099] 3. Load the model. Load the trained multi-object tracking model.
[0100] 4. Start inference. After the data and model are loaded, start tracking the targets in the scene.
[0101] 5. Input image data sequentially. Input the image data into the model frame by frame in chronological order.
[0102] 6. Model inference. After the image data passes through the model, the final prediction results are obtained, including the IDs of the targets and the bounding box information in each frame.
[0103] 7. Save the prediction results. Save the prediction results as a txt file.
[0104] 8. Visualize the tracking results. Draw the bounding box information in the prediction results on the image to mark the positions of the targets in the scene, and mark their IDs in the upper left corner. Give the visualized tracking results to the operator in real time. For Figure 6 the street view monitoring, the visualization result is as Figure 9 shown, Figure 9 which is a schematic diagram of the visualization of street view crowd tracking shown in an embodiment of the present invention.
[0105] 9. End inference. For the historical storage dataset, when all the data in the dataset has been inferred, it ends. For real-time monitoring, the operator needs to manually close the tracking program.
[0106] In addition, in another embodiment, the multi-target tracking task can be applied not only to security monitoring but also to some special tasks, such as tracking people in sports or dance performances. As Figure 10 shown, Figure 10 which is a schematic diagram of the application of a multi-target tracking task in a dance scene shown in an embodiment of the present invention. In such tasks, the people appearing in the scene are relatively fixed. Therefore, the difficulty of the task lies in that the appearances of the people in this scene are mostly similar, and the movement trends are difficult to estimate, which poses higher requirements on the tracker. The training and inference processes of such tasks are the same as those of the above embodiment, and the differences lie in some parameter settings and methods during training. The specific steps and differences are as follows: Regarding constructing the dataset: For the dance scene, the publicly available dataset DanceTrack can usually be used as the training dataset. At the same time, if tracking is required in a specific scene, the staff needs to construct the training dataset by themselves, and the method is the same as that of the above embodiment. Regarding inputting the model for inference: It is to input the processed image data into the model for inference.
[0107] Regarding ending the training: It is to train the model on the dataset for 18 epochs, and increase the number of clipped frames to 3, 4, and 5 frames at the 6th, 10th, and 14th epochs. Among them, for the tracking of the dance scene, the visualization result is as Figure 11 as shown Figure 11 It is a tracking visualization schematic diagram of a dance scene shown in an embodiment of the present invention.
[0108] In one embodiment, to verify the effectiveness of the multi-object tracking method based on Transformer and trajectory query group proposed in the embodiments of the present invention, experiments were conducted on the multi-object tracking public datasets MOT Challenge and DanceTrack in this embodiment, and compared with other techniques to verify the effectiveness of this method. For the verification metrics, CLEAR MOT Metrics and HOTA were used as the verification criteria in this embodiment. CLEAR MOT Metrics is widely used and includes MOTA, IDF1, ID-SW, MT, and ML. For HOTA, HOTA, AssA, and DetA were used in this project. MOTA is the multi-object tracking accuracy rate, which mainly focuses on the comprehensive error rate of tracking. It measures the overall accuracy of the tracking task by calculating the weighted sum of missed detections (Miss), false positives (False Positive), and ID switches (ID-SW). The higher the value, the better the overall performance of the algorithm in detection and tracking. IDF1 is used to measure the ID consistency of the target in the entire sequence. It reflects the stability and continuity of tracking by calculating the ID matching F1 score of the target in the tracking result. The higher the IDF1 score, the more accurately the target is tracked throughout the video and the better the ID consistency. The number of ID switches (ID-SW) represents the number of incorrect ID switches of the same target during the tracking process. The fewer the number of ID switches, the more stable the tracking algorithm is in target association. High-frequency ID-SW means that the algorithm is difficult to maintain ID consistency in a long time series. MT refers to the number of targets that are mostly correctly tracked, indicating the proportion of targets that are correctly tracked for at least 80% of the time in the tracking sequence. A higher MT value means that the algorithm performs better in tracking targets for a long time. ML refers to the number of targets that are mostly lost, indicating the proportion of targets whose tracking time is less than 20% in the entire sequence. The lower the ML value, the more stable the tracking effect of the algorithm between different frames and the less likely it is to lose targets. HOTA is a new multi-object tracking evaluation metric that comprehensively evaluates the algorithm's performance by balancing detection and association capabilities. While considering detection and association performance, HOTA also measures the quality of target detection in spatial and temporal dimensions. A higher HOTA value indicates that the algorithm has better balance in detection and ID association. AssA is the association accuracy sub-metric of HOTA, which focuses on evaluating the ID association ability of the algorithm and measures whether the same target is correctly maintained under the same ID in multiple frames. The higher AssA is, the better the ID consistency and association performance of the tracking. DetA is the detection accuracy sub-metric of HOTA, which focuses on the detection accuracy of the target. The higher DetA is, the more accurately the algorithm can detect and locate the target without being affected by association problems.
[0109] As shown in Table 1 and Table 2, "Ours" represents the multi-object tracking method proposed in the present invention based on Transformer and trajectory query groups. The detection results are compared with the latest Transformer-based multi-object tracking methods on the MOT17 and MOT20 test sets. It can be found that the method of the present invention has further improvements in HOTA and IDF1 compared with other technical methods, reduces ID switches, and significantly improves MOTA.
[0110] Table 1 Comparison table with other technologies on the MOT17 dataset
[0111] Table 2 Comparison table with other technologies on the MOT20 dataset
[0112] In the DanceTrack dataset, the motion states of the targets are more complex, which poses higher requirements for the tracking ability of the model. The method is tested on this dataset, and the results are shown in Table 3. The results show that compared with other models, the present invention ("Ours") has obvious advantages in detection and tracking performance.
[0113] Table 3 Comparison table with other technologies on the DanceTrack dataset
[0114] This indicates that the design of the method has been effectively verified in multiple datasets and different types of challenges, demonstrating that the trajectory query group design of the present invention effectively improves the detection ability of the tracking model in complex scenarios, can still maintain tracking in the context of frequent switching of target occlusion states, improves the robustness and fault tolerance rate of the model, and generally makes the model tracking performance more excellent.
[0115] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0116] Based on the same inventive concept, an embodiment of the present invention provides a multi-object tracking device based on Transformer and trajectory query groups. Refer to Figure 12 , Figure 12 It is a structural block diagram of a multi-object tracking device based on Transformer and trajectory query groups provided by an embodiment of the present invention. As Figure 12 shown, the multi-object tracking device based on Transformer and trajectory query groups in this embodiment may include: An image input module, configured to input the current image frame in the sample video stream into a multi-object tracking model to be trained, where the multi-object tracking model to be trained at least includes: a backbone network, a Transformer encoder, and a Transformer decoder; An image encoding module, configured to process the current image frame through the backbone network and the Transformer encoder to obtain a current image encoding feature; An embedding prediction module, configured to, in the case that a target is detected in the previous image frame of the current image frame, input the current image encoding feature, a detection query feature, and a trajectory query group corresponding to the tracked target in the current image frame into the Transformer decoder to obtain one or more predicted embedding results corresponding to each of the multiple targets; each tracked target corresponds to a trajectory query group, and the trajectory query group includes: multiple trajectory query features representing different occlusion levels; An embedding output module, configured to determine one output embedding result corresponding to each of the multiple targets from one or more predicted embedding results corresponding to each of the multiple targets; A tracking prediction module, configured to determine a sample multi-object tracking result corresponding to the current image frame based on the output embedding results corresponding to each of the multiple targets; A model training module, configured to train the multi-object tracking model to be trained based on the sample multi-object tracking result corresponding to each image frame in the sample video stream and the multi-object tracking result label corresponding to each image frame to obtain a trained multi-object tracking model; A tracking detection module, configured to input a video stream to be detected into the trained multi-object tracking model to obtain a multi-object tracking result corresponding to the video stream to be detected.
[0117] Optionally, the device further includes: A result prediction module, configured to, in the case that no target is detected in the previous image frame of the current image frame, or, in the case that the current image frame is the first image frame, input the current image encoding feature and the detection query feature into the Transformer decoder to obtain a predicted embedding result corresponding to a new target in the current image frame; A tracking determination module, configured to determine a sample multi-object tracking result corresponding to the current image frame based on the predicted embedding result corresponding to the new target; An initialization module, configured to use the new target as the tracked target in the next image frame of the current image frame, and initialize a trajectory query group corresponding to the tracked target in the next frame of image based on the predicted embedding result corresponding to the new target.
[0118] Optionally, the multiple targets include: a tracked target and a new target, or, multiple tracked targets; When the target is the tracked target, each target corresponds to multiple predicted embedding results, and each predicted embedding prediction result is the predicted embedding result corresponding to the trajectory query features of different occlusion levels respectively; When the target is the new target, each target corresponds to one predicted embedding result; The embedding output module includes: A loss calculation module, configured to determine the loss between the multiple predicted embedding results respectively corresponding to each target and the label embedding result corresponding to each target; An embedding result determination module, configured to determine the predicted embedding result with the minimum loss as an output embedding result corresponding to each target.
[0119] Optionally, the different occlusion levels include at least three occlusion levels, and the three occlusion levels respectively correspond to different occlusion thresholds; During the process of the trained multi-target tracking model performing tracking detection on the video stream to be detected, when the predicted embedding results corresponding to the three occlusion levels of the target are obtained, the predicted embedding result with the highest confidence is determined as the output embedding result corresponding to the target; When the confidence of the output embedding result falls within the occlusion range corresponding to other occlusion levels, it indicates that the occlusion level of the target has changed.
[0120] Optionally, the trajectory query features include: query content embedding and query position embedding, and the output embedding result includes: output content embedding and output position embedding; the apparatus further includes: A query content update module, configured to input, for each target among the multiple targets, the output content embedding corresponding to the target in the current image frame, the query content embedding in the trajectory query group corresponding to the target in the current image frame, and the long-term memory embedding corresponding to the target into a trajectory query group updater, to obtain an updated query content embedding corresponding to the target; A query position prediction module, configured to input, for each target among the multiple targets, at least the output position embedding corresponding to the target in the current image frame into a position predictor, to obtain a predicted query position embedding corresponding to the target in the next image frame; A query group update module, configured to obtain an updated trajectory query group corresponding to the target based on the updated query content embedding corresponding to the target and the predicted query position embedding corresponding to the target in the next image frame; A query group determination module, configured to use the updated trajectory query group as the trajectory query group corresponding to the tracked target in the next image frame; Wherein, different query content embeddings corresponding to the same target share the same query position embedding.
[0121] Optionally, the query content update module includes: A first input module, configured to splice the query content embeddings with different occlusion levels in the trajectory query group corresponding to the target in the current image frame and the output content embedding corresponding to the target in the current image frame, and then input them into the short-term memory aggregator to obtain query enhancement content embeddings with different occlusion levels corresponding to the target; An attention module, configured to splice the query enhancement content embeddings with different occlusion levels to obtain a spliced content embedding, use the spliced content embedding as keys and values, use the query enhancement content embeddings with different occlusion levels as queries respectively, perform occlusion-aware memory attention operations, and obtain intermediate content embeddings corresponding to different occlusion levels; A feature combination module, configured to combine the intermediate content embeddings corresponding to different occlusion levels with the long-term memory embedding corresponding to the target respectively to obtain updated content embeddings corresponding to different occlusion levels; A content update module, configured to obtain an updated query content embedding corresponding to the target based on the updated content embeddings corresponding to different occlusion levels.
[0122] Optionally, the query position prediction module includes: A first determination module, configured to input the output position embedding corresponding to the target in the current image frame and the historical output position embedding corresponding to the target in the previous image frame into the position predictor to obtain a predicted query position embedding corresponding to the target in the next image frame and the historical output position embedding corresponding to the target in the current image frame; The historical output position embedding corresponding to the target in the current image frame is used to provide continuous temporal context to assist the position predictor in outputting subsequent predicted query position embeddings in the next image frame.
[0123] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the multi-target tracking method based on Transformer and trajectory query groups as described in any one of the above embodiments of the present invention.
[0124] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the steps in the multi-object tracking method based on Transformer and trajectory query groups described in any of the above embodiments of the present invention.
[0125] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiment.
[0126] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0127] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention can take the form of completely hardware embodiments, completely software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0128] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing terminal devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0129] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0130] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide for implementing the process Figure 1 one process or multiple processes and / or blocks Figure 1 steps for the functions specified in one block or multiple blocks.
[0131] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
[0132] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.
[0133] The above has introduced in detail a multi-object tracking method, device, equipment and medium based on Transformer and trajectory query groups provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.< / id> < / id>
Claims
1. A multi-object tracking method based on Transformer and trajectory query groups, characterized in that The method includes: Inputting the current image frame in the sample video stream into a multi-object tracking model to be trained, where the multi-object tracking model to be trained at least includes: a backbone network, a Transformer encoder, and a Transformer decoder; Processing the current image frame through the backbone network and the Transformer encoder to obtain the current image encoded feature; When a target is detected in the previous image frame of the current image frame, inputting the current image encoded feature, the detection query feature, and the trajectory query group corresponding to the tracked target in the current image frame into the Transformer decoder to obtain one or more predicted embedding results corresponding to each of the multiple targets; each tracked target corresponds to a trajectory query group, and the trajectory query group includes: multiple trajectory query features representing different occlusion levels; Determining one output embedding result corresponding to each of the multiple targets from one or more predicted embedding results corresponding to each of the multiple targets; Based on the output embedding results corresponding to each of the multiple targets, determining the sample multi-object tracking result corresponding to the current image frame; Training the multi-object tracking model to be trained based on the sample multi-object tracking result corresponding to each image frame in the sample video stream and the multi-object tracking result label corresponding to each image frame to obtain a trained multi-object tracking model; Inputting the video stream to be detected into the trained multi-object tracking model to obtain the multi-object tracking result corresponding to the video stream to be detected.
2. The multi-object tracking method based on Transformer and trajectory query group according to claim 1, wherein The method further includes: When a target is not detected in the previous image frame of the current image frame, or when the current image frame is the first image frame, inputting the current image encoded feature and the detection query feature into the Transformer decoder to obtain the predicted embedding result corresponding to the new target in the current image frame; Based on the predicted embedding result corresponding to the new target, determining the sample multi-object tracking result corresponding to the current image frame; In the next image frame of the current image frame, using the new target as the tracked target, and initializing the trajectory query group corresponding to the tracked target in the next frame of image based on the predicted embedding result corresponding to the new target.
3. The multi-object tracking method based on Transformer and trajectory query group according to claim 1, characterized in that, The multiple targets include: tracked targets and new targets, or, multiple tracked targets; When the target is a tracked target, each target corresponds to multiple predicted embedding results, and each predicted embedding result is the predicted embedding result corresponding to the trajectory query features of different occlusion levels respectively; When the target is a new target, each target corresponds to one predicted embedding result; Determining one output embedding result corresponding to each of the multiple targets from the multiple predicted embedding results corresponding to each of the multiple targets includes: Determining the loss between the multiple predicted embedding results corresponding to each target respectively and the label embedding result corresponding to each target; Determining the predicted embedding result with the minimum loss as one output embedding result corresponding to each target.
4. The multi-object tracking method based on Transformer and trajectory query group according to claim 1, wherein The different occlusion levels at least include three occlusion levels, and the three occlusion levels respectively correspond to different occlusion thresholds; During the process of the trained multi-object tracking model performing tracking detection on the video stream to be detected, when the predicted embedding results corresponding to the three occlusion levels of the target are obtained, the predicted embedding result with the highest confidence is determined as the output embedding result corresponding to the target; When the confidence of the output embedding result falls within the occlusion range corresponding to other occlusion levels, it indicates that the occlusion level of the target has changed.
5. The multi-object tracking method based on Transformer and trajectory query group according to any one of claims 1 to 4, characterized in that, The trajectory query feature includes: query content embedding and query position embedding, and the output embedding result includes: output content embedding and output position embedding; the method further includes: For each target among multiple targets, the output content embedding corresponding to the target in the current image frame, the query content embedding in the trajectory query group corresponding to the target in the current image frame, and the long-term memory embedding corresponding to the target are input into the trajectory query group updater to obtain the updated query content embedding corresponding to the target; For each target among multiple targets, at least the output position embedding corresponding to the target in the current image frame is input into the position predictor to obtain the predicted query position embedding corresponding to the target in the next image frame; Based on the updated query content embedding corresponding to the target and the predicted query position embedding corresponding to the target in the next image frame, the updated trajectory query group corresponding to the target is obtained; The updated trajectory query group is used as the trajectory query group corresponding to the tracked target in the next image frame; Among them, different query content embeddings corresponding to the same target share the same query position embedding.
6. The multi-object tracking method based on Transformer and trajectory query group according to claim 5, characterized in that, Inputting the output content embedding corresponding to the target in the current image frame, the query content embedding in the trajectory query group corresponding to the target in the current image frame, and the long-term memory embedding corresponding to the target into the trajectory query group updater to obtain the updated query content embedding corresponding to the target includes: Respectively, the query content embeddings of different occlusion levels in the trajectory query group corresponding to the target in the current image frame are concatenated with the output content embedding corresponding to the target in the current image frame and then input into the short-term memory aggregator to obtain the query enhanced content embeddings of different occlusion levels corresponding to the target; The query enhanced content embeddings of different occlusion levels are concatenated to obtain a concatenated content embedding. The concatenated content embedding is used as the key and value, and the query enhanced content embeddings of different occlusion levels are respectively used as queries to perform occlusion-aware memory attention operations to obtain the intermediate content embeddings corresponding to different occlusion levels; Respectively, the intermediate content embeddings corresponding to different occlusion levels are combined with the long-term memory embedding corresponding to the target to obtain the updated content embeddings corresponding to different occlusion levels; Based on the updated content embeddings corresponding to different occlusion levels, the updated query content embedding corresponding to the target is obtained.
7. The multi-object tracking method based on Transformer and trajectory query group according to claim 5, wherein Inputting at least the output position embedding corresponding to the target in the current image frame into the position predictor to obtain the predicted query position embedding corresponding to the target in the next image frame includes: Embed the output position corresponding to the target in the current image frame, and the historical output position embedding corresponding to the target in the previous image frame into the position predictor, to obtain the predicted query position embedding corresponding to the target in the next image frame and the historical output position embedding corresponding to the target in the current image frame; The historical output position embedding corresponding to the target in the current image frame is used to provide continuous temporal context to assist the position predictor in outputting subsequent predicted query position embeddings in the next image frame.
8. A multi-object tracking device based on Transformer and a trajectory query group, characterized in that, The device includes: An image input module, configured to input the current image frame in the sample video stream into the multi-object tracking model to be trained, where the multi-object tracking model to be trained at least includes: a backbone network, a Transformer encoder, and a Transformer decoder; An image encoding module, configured to process the current image frame through the backbone network and the Transformer encoder to obtain the current image encoding features; An embedding prediction module, configured to, when a target is detected in the previous image frame of the current image frame, input the current image encoding features, detection query features, and the trajectory query group corresponding to the tracked targets in the current image frame into the Transformer decoder to obtain one or more predicted embedding results corresponding to each of the multiple targets; each tracked target corresponds to a trajectory query group, and the trajectory query group includes: multiple trajectory query features representing different occlusion levels; An embedding output module, configured to determine one output embedding result corresponding to each of the multiple targets from one or more predicted embedding results corresponding to each of the multiple targets; A tracking prediction module, configured to determine the sample multi-object tracking result corresponding to the current image frame based on the output embedding results corresponding to each of the multiple targets; A model training module, configured to train the multi-object tracking model to be trained based on the sample multi-object tracking result corresponding to each image frame in the sample video stream and the multi-object tracking result label corresponding to each image frame, to obtain a trained multi-object tracking model; A tracking detection module, configured to input the video stream to be detected into the trained multi-object tracking model to obtain the multi-object tracking result corresponding to the video stream to be detected.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the multi-object tracking method based on Transformer and trajectory query group according to any one of claims 1 to 7.
10. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the multi-object tracking method based on Transformer and trajectory query group according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-target tracking method based on selective query collection and improved data association
CN117609528A
Transform decoder and attention-based vehicle-mounted image small target tracking method and application
CN118115858A
Multi-target tracking method based on lightweight network
CN118968031A
Deep learning method for multiple object tracking from video
US20240144489A1
Cited By
Multi-target tracking method used in complex scene
CN121639739A