An unsupervised domain adaptation method for multi-target tracking based on pseudo-label correction

By adopting the unsupervised domain adaptation method of pseudo-label correction and adversarial training in the multi-objective tracking model, the model's performance degradation on the unlabeled target data domain is solved, and more efficient model generalization and manual annotation cost reduction are achieved.

CN114693979BActive Publication Date: 2025-05-06ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210368119.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-08
Publication Date
2025-05-06
Estimated Expiration
2042-04-08

AI Technical Summary

Technical Problem

The existing multi-objective tracking model has degraded performance and requires manual complex annotation, which is costly when migrating from a supervised information-rich source data domain to an unlabeled target data domain.

Method used

The unsupervised domain adaptation method based on pseudo-label correction is adopted, and the source domain data is migrated to the target domain through the image style conversion model, and the gradient inversion layer and domain classifier are combined for adversarial training. The pseudo-label correction module is used to continuously improve the pseudo-label and improve the generalization performance of the model.

Benefits of technology

It significantly reduces the manual marking cost of multi-objective tracking models in practical application scenarios, improves the generalization performance of the model in an open environment, and can better adapt to differences between data domains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114693979B_ABST
    Figure CN114693979B_ABST
Patent Text Reader

Abstract

The present invention discloses an unsupervised domain adaptation method for multi-target tracking based on pseudo-label correction. The method first uses a generative adversarial network to perform style conversion on an image in a source data domain to a target data domain; then, in the process of training a domain adaptation model, after each round of training, the current model is used to generate pseudo-labels in the target data domain, and the pseudo-labels are corrected and added to the domain adaptation training; finally, the training is completed to obtain the final tracking network. The method can add domain adaptation training by continuously tuning and correcting the pseudo-label supervision information of the target domain, so that the tracking model can better learn features with domain-invariant properties, and obtain performance close to supervised learning in the target data domain without supervised information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and intelligent recognition technology, and in particular relates to an unsupervised domain adaptation method for multi-target tracking based on pseudo-label correction. Background Art

[0002] As an important basic task in the field of computer vision, multi-target tracking completes the positioning and identification of the target of interest by analyzing the visual, motion and other information of the video frame image sequence. It has a wide range of applications in engineering, such as autonomous driving, video surveillance, behavior prediction, traffic management, etc. With the rapid development of deep learning technology, thanks to the powerful feature representation ability of deep networks for image targets, multi-target tracking technology has also been further developed.

[0003] The multi-target tracking system usually includes two parts: target detection and data association. The target detection part can locate the target appearing in the image, and the data association part assigns an identity ID to the detected target based on the continuity of the target trajectory in terms of shape, motion and other information, and completes the data association between the trajectory and the detection result. Based on the various implementation forms of this process, the multi-target tracking method can be divided into a tracking method based on appearance feature distance, a tracking method based on motion information, a tracking method based on video clips, etc. The unsupervised domain adaptation method described in the present invention takes the more popular tracking framework based on appearance feature distance as an example, but is not limited to this form, but is applicable to all multi-target tracking paradigms based on deep networks.

[0004] By training on a deep network using a dataset containing supervised information, the multi-target tracking model can fit a relatively good performance on the current dataset (source data domain), but when applied to data from other scenarios (target data domain), the model cannot perform well due to differences in data distribution (such as seasonal changes, virtual synthesis and real shooting, different camera parameters, etc.). Since the identity-level supervised information labeling required for multi-target tracking is a complex and time-consuming task, most application scenarios in real applications do not have true value labels. If the performance of the deep tracking model trained on a dataset with supervised information can be better transferred to an unlabeled application scenario dataset by using an unsupervised domain adaptation method, it will greatly save the cost of complex manual labeling and improve work efficiency.

[0005] After long-term research and development, the unsupervised domain adaptation method has achieved certain results in the fields of image classification, target detection, pedestrian re-identification, etc., but there has been no related work in the field of multi-target tracking; the unsupervised multi-target tracking method that is more related to the present invention aims to perform multi-target tracking on a data set with no supervised information at all. Its problem setting and implementation method are different from those of unsupervised domain adaptation tracking and cannot be directly applied to the current problem. Summary of the invention

[0006] The purpose of the present invention is to provide a multi-target tracking unsupervised domain adaptation method based on pseudo-label correction in view of the deficiencies in the prior art.

[0007] The objective of the present invention is achieved through the following technical solution: a multi-target tracking unsupervised domain adaptation method based on pseudo-label correction. It includes the following steps:

[0008] To achieve the above purpose, the present invention adopts the following technical solution: first input the source domain data D containing complete supervision information s {x s ,(Box s , ID s )}, target domain data D without labels t {x t}, perform the following steps:

[0009] Step 1: Use the image style transfer model G to the source domain data D s {x s ,(Box s , ID s )} to the target domain data D t {x t} style migration to obtain the converted dataset After transformation, data x s After style transfer, it becomes x s′ , but the label information (Box s , ID s ) remains unchanged. Merge D′ s and D t Form a domain adaptation training dataset.

[0010] Step 2: Randomly sample the data set so that each training batch contains an equal number of D′ s and D t Source data;

[0011] Step 3: Add the modules required for domain adaptation training to the original multi-target tracking model. The specific changes are: add a gradient reversal layer GRL and a domain classifier D to the feature extraction deep network F in the model. The gradient reversal layer is responsible for negating the gradient and then returning it during training. The domain classifier is responsible for classifying the domains of the feature extraction network (source domain, target domain). The goal of domain adaptation training is:

[0012]

[0013] Among them, x represents the input data, err represents the error probability of the domain classifier for feature classification. The optimization of the min-max problem is achieved by adding a gradient reversal layer between the feature extraction network and the domain classifier and performing adversarial training;

[0014] Step 4: Perform domain adaptation training of a multi-target tracking model in one stage to obtain the model M of the current training stage curr ;

[0015] Step 5: Use the tracking model M in the current training stage curr and the target domain dataset D t {x t}, get a rough pseudo label (Box p , ID p );

[0016] Step 5: Transform the target domain dataset D t {x t} and rough pseudo-labels (Box p , ID p ) is sent to the pseudo-label correction module, and the target trajectory is completed by predicting the single target tracker of the forward and backward traversal frame sequence to obtain a corrected more accurate pseudo-label (Box p′ , ID p′ ), the specific steps are as follows:

[0017] Step 5.1: Forward traverse the frame image of the target domain data, using the rough pseudo label (Box p , ID p ) In each frame image, for each newly appearing target in each image, a single target tracker based on visual information is established, and the position of subsequent frames is predicted based on the continuity of visual features between frames;

[0018] Step 5.2: For the target for which the single target tracker has been established, the overlap coverage of the prediction result and the pseudo label position is matched in each subsequent frame. If a match is achieved, tracking continues. If a match cannot be achieved, the label is temporarily lost and the single target tracker is continued to run in subsequent frames to try to match;

[0019] Step 5.3: For the target whose mark is temporarily lost in step 5.2, if the matching fails for more than a certain number of frames, it is considered that the target has left the frame image field of view, the single target tracker is deleted, and the matching is stopped. If the target whose mark is temporarily lost is matched in the subsequent frames, the position predicted by the single target tracker is used as the result to complete the target trajectory of the frames where the matching fails in the middle, and add the pseudo label {Box p , ID p}middle;

[0020] Step 5.4: traverse the frame images of the target domain data in reverse order and repeat the label correction steps from step 5.1 to step 5.3;

[0021] Step 5.5: Fuse the pseudo-label results obtained by the forward traversal and the backward traversal, take the union of the target trajectory, and output the corrected pseudo-label (Box p′ , ID p′ ).

[0022] Step 6: Use the corrected pseudo-label as the supervision information of the target domain data and combine it with the current target domain data D t {x t ,(Box p′ , ID p′ )} and the source domain data D′ after style conversion s {x s′ ,(Box s , ID s )} to form a new data set, repeat steps 2 to 6 until the training of the domain adaptation model converges, and then end the training to obtain the final domain adaptation model M out ;

[0023] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention is an unsupervised domain adaptation method for multi-target tracking based on pseudo-label correction, which first uses an image style conversion model to narrow the distribution difference between source domain data and target domain data in the original image, and then uses adversarial training based on adding gradient reversal layers and domain classifiers to learn the model domain adaptation ability, and finally continuously adds the corrected target domain data pseudo-labels during the training process, thereby improving the generalization performance of the multi-target tracking model in the target domain. The present invention can greatly reduce the manual labeling cost of the multi-target tracking model in actual application scenarios, improve the generalization performance of the model in data domains containing differences in an open environment, and better apply it to actual scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a flow chart of unsupervised domain adaptation training based on pseudo-label correction in the present invention;

[0025] Figure 2 This is a pseudo-label correction flow chart of the present invention;

[0026] Figure 3 This is a network architecture diagram of the domain adaptation training model of the present invention. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. On the contrary, the present invention encompasses any substitutions, modifications, equivalent methods and schemes made on the essence and scope of the present invention as defined by the claims. Further, in order to enable the public to have a better understanding of the present invention, some specific details are described in detail in the detailed description of the present invention below. For those skilled in the art, the present invention can be fully understood without the description of these details.

[0028] refer to Figure 1 , shown is a flowchart of the steps of the unsupervised domain adaptation method for multi-target tracking based on pseudo-label correction according to an embodiment of the present invention.

[0029] Input source domain data D containing complete supervision information s {x s ,(Box s , ID s )}, target domain data D without labels t {x t}, where x s 、x t Represents a video frame image sequence, Box s Indicates the bounding box annotation information of the tracked target, ID s Indicates the identity information of each tracking target. Perform the following steps:

[0030] Step 1: Training with source domain data D s {x s ,(Box s , ID s )} to the target domain data D t {x t} style transfer model G. Model G uses the CycleGAN model, whose training does not require a one-to-one correspondence between source domain data and target domain data. The model loss function L G Set to:

[0031]

[0032] Where D s is the source domain data, D t is the target domain data, L GANThe loss for adversarial generation includes the loss of conversion from the source domain to the target domain and the loss of conversion from the target domain to the source domain; L cyc Represents cycle consistency loss; They represent the generators from source domain to target domain and from target domain to source domain respectively; B s , B t Denotes the discriminator of the above two. Taking the loss of converting from source domain to target domain as an example, it is:

[0033]

[0034] in Indicates that D s and D t Expectations on. cyc It represents the cycle consistency loss, which indicates the consistency loss of the conversion between source domain data and target domain data. It is specifically expressed as:

[0035]

[0036] Use the trained model G to perform style conversion on the source domain data and obtain the converted dataset Merge D′ s and D t Form a domain adaptation training dataset.

[0037] Step 2: Randomly sample the data set so that each training batch contains an equal number of D′ s and D t Source data;

[0038] Step 3: Reference Figure 3 Multi-target tracking module. The present invention uses a more popular end-to-end multi-target tracking model based on target detection and re-identification feature reid matching. After the image is input into the model, it first passes through the feature extraction networks F1 and F2 to obtain a visual feature map. After feature combination enhancement, the input prediction branch can obtain classification regression information for target positioning and appearance feature vectors for data association.

[0039] The loss function L of the multi-target tracking network track Contains the foreground classification prediction loss l cls , the bounding box regression prediction loss l reg and the classification loss l for re-ID feature prediction reid , the specific expression is:

[0040] l track = l cls +l reg +l rei d

[0041]

[0042] l reg =1-IoU(box gt ,box pred )

[0043]

[0044] Among them, l cls is the cross entropy loss of the classification branch, H and W are the width and height of the feature map, and p i is the true value probability of a point on the feature map being the foreground, is the predicted probability that a point on the feature map is a foreground point; l reg is the IoU (overlap coverage) loss of the regression branch, where box gt is the true value bounding box position, box pred To predict the location of the bounding box; l reid is the classification cross entropy loss of the re-identification feature, where N is the number of targets in the current image, and e i is the true classification vector of a target, is the predicted classification vector for a target.

[0045] refer to Figure 3 The domain adaptation module makes the following specific changes to the multi-target tracking model: add a gradient reversal layer GRL and domain classifiers D1 and D2 to the output of the feature extraction deep networks F1 and F2 in the model, respectively. The gradient reversal layer is responsible for negating the gradient and then transmitting it back during training, and the domain classifier is responsible for classifying the domains of the feature extraction network (source domain, target domain). The input of the domain classifier D1 is the low-level visual features passed in by the shallow network, which is the domain classification at the pixel level; the input of the domain classifier D2 is the pooled high-level features, which is the overall domain classification of the features. The goals of the domain adaptation part training are:

[0046]

[0047] Where x represents the input image, err represents the error probability of the domain classifier for feature classification. The optimization of the min-max problem is achieved by adding a gradient reversal layer between the feature extraction network and the domain classifier and performing adversarial training.

[0048] Step 4: Perform one epoch of multi-target tracking domain adaptation training to obtain the model Mc of the current training stage urr ;

[0049] Step 5: Use the tracking model M in the current training stage curr and the target domain dataset D t {x t}, get a rough pseudo label (Boxp , ID p );

[0050] Step 5: Transform the target domain dataset D t {x t} and rough pseudo-labels (Box p , ID p ) is sent to the pseudo-label correction module, and the target trajectory is completed by predicting the single target tracker of the forward and backward traversal frame sequence to obtain a corrected more accurate pseudo-label (Box p ′,ID p ') The specific steps are as follows:

[0051] Step 5.1: Forward traverse the frame image of the target domain data, using the rough pseudo label (Box p , ID p ) In each frame image, for each newly appearing target in each image, a single target tracker based on a twin network is established, and the position of subsequent frames is predicted based on the continuity of visual features between frames;

[0052] Step 5.2: For the target for which the single target tracker has been established, the overlap coverage of the prediction result and the pseudo label position is matched in each subsequent frame. If a match is achieved, tracking continues. If a match cannot be achieved, the label is temporarily lost and the single target tracker is continued to run in subsequent frames to try to match;

[0053] Step 5.3: For the target whose mark is temporarily lost in step 5.2, if the match fails for 10 frames continuously, it is considered that the target has left the frame image field of view, the single target tracker is deleted, and the matching is stopped. If the target whose mark is temporarily lost is matched in the subsequent frames, the position predicted by the single target tracker is used as the result to complete the target trajectory of the frames where the match failed in the middle, and add the pseudo label {Box p , ID p}, get the positive completion trajectory label {Box fwd , ID fwd};

[0054] Step 5.4: Reversely traverse the frame image of the target domain data, repeat the label correction steps from step 5.1 to step 5.3, and obtain the reverse completion trajectory label {Box bkd , ID bkd}

[0055] Step 5.5: Obtain the forward completion trajectory label by combining the pseudo-label results obtained by forward traversal and backward traversal {Box fwd , ID fwd}、{Box bkd , ID bkd} Take the union and output the corrected pseudo-label (Box p′ , ID p′ ).

[0056] Step 6: Use the corrected pseudo-label as the supervision information of the target domain data and combine it with the current target domain data D t {x t ,(Box p′ , ID p′ )} and the source domain data D′ after style conversion s {x s ,,(Box s , ID s )} form a new data set, repeat steps 2 to 6, and train the domain adaptation model until the loss function of the model is lower than the fixed threshold. End the training and obtain the final domain adaptation model M out ;

[0057] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A multi-target tracking unsupervised domain adaptation method based on pseudo-label correction, characterized in that: The specific steps include: (1) Forming a domain adaptation training dataset: Using the image style transfer model G to transform the source domain data D s {x s ,(Box s , ID s )} to the target domain data D t {X t } style migration to obtain the converted dataset Merge D′ s and D t Form a domain adaptation training dataset; (2) Training to obtain tracking model: using merged D′ s and D t The training data set is formed, and a domain adaptation training of a multi-target tracking model is performed to obtain the tracking model M in the current training stage. curr ; (3) Obtaining rough pseudo labels: Using the tracking model M in the current training phase curr and the target domain dataset D t {X t }, get a rough pseudo label (Box p , ID p ); (4) Correct the rough pseudo-labels: t {x t } and rough pseudo-labels (Box p , ID p ) is sent to the pseudo-label correction module to obtain a more accurate pseudo-label after correction (Box p′ , ID p′ ); (5) Combined data: The corrected pseudo-labels are used as the supervision information of the target domain data and combined with the current target domain data D t {x t ,(Box p′ , ID p′ )} and the source domain data D′ after style conversion s {x s′ ,(Box s , ID s )} form a new data set; (6) Repetition and convergence: Repeat steps (2) to (5) to train the domain adaptation model until the model loss function is lower than a fixed threshold. End the training and obtain the final domain adaptation model M. out .

2. The unsupervised domain adaptation method for multi-target tracking based on pseudo-label correction according to claim 1, characterized in that: The image style transfer model G in step (1) uses a generative adversarial network to narrow the data distribution distance between the source domain and the target domain.

3. The unsupervised domain adaptation method for multi-target tracking based on pseudo-label correction according to claim 1, characterized in that: The multi-target tracking model described in step (2) is composed of a target detection part and a data association part, wherein the target detection part obtains the target position of the current frame, and the data association part uses visual feature correlation information to complete the identity matching of the trajectory and the detection result.

4. The unsupervised domain adaptation method for multi-target tracking based on pseudo-label correction according to claim 1, characterized in that: The domain adaptation training of the multi-target tracking model described in step (2) adopts the following technology: adding a gradient reversal layer GRL and a domain classifier D to the feature extraction deep network F in the model, wherein the gradient reversal layer is responsible for negating the gradient and then returning it during the training process, and the domain classifier is responsible for classifying the source domain and the target domain of the feature extraction network. The objective function of the domain adaptation training is: Among them, x represents the input data, err represents the error probability of the domain classifier for feature classification, and the optimization of the min-max problem is achieved by adding a gradient reversal layer between the feature extraction network and the domain classifier and performing adversarial training.

5. The unsupervised domain adaptation method for multi-target tracking based on pseudo-label correction according to claim 1, characterized in that: The pseudo-label correction step described in step (4) is as follows: (1) Forward traverse the frame images of the target domain data, for each image, the rough pseudo label (Box p ,ID p ) result, if it is a new target, the target is modeled based on visual information for position prediction in subsequent frames; (2) If the visual model predicts the location of a target that does not exist in the rough pseudo-label (Box p ,ID p ) and the prediction result is reasonable, the prediction result of the visual model is used to complete the pseudo label (Box p ,ID p ); (3) Reversely traverse the frame image of the target domain data, repeat steps (1) to (2), and obtain the pseudo-label completion result of the reverse traversal; (4) Take the union of the pseudo-label results obtained by the forward traversal and the backward traversal, and output the corrected pseudo-label (Box p ′,ID p′ ).

Citation Information

Patent Citations

  • Visual tracking method combining classification and domain adaptation

    CN109840518A

  • Unsupervised domain adaptation method based on adversarial learning loss function

    CN110837850A