A video multi-target tracking method based on deep learning and temporal feature enhancement
By introducing timing feature enhancement module and dual-frame output loss calculation in the multi-objective tracking model, the IDSwitch problem in the case of inter-frame occlusion and the training efficiency reduction caused by the sharing of object detection and tracking task parameters in the existing model is solved, and more efficient multi-objective tracking performance is achieved.
Patent Information
- Application Number
- CN202210632698.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-06
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-06-06
AI Technical Summary
The existing multi-objective tracking model fails to effectively utilize timing information, resulting in frequent IDSwitch phenomena under inter-frame occlusion, reducing the accuracy of the model. At the same time, object detection and tracking tasks share too many parameters during training, resulting in reduced training efficiency and poor performance.
A video multi-objective tracking method based on deep learning and timing feature enhancement is proposed. By splitting the detection and feature generation structure of the model and adding a feature enhancement module based on timing information on the ReID branch, the model's ability to discriminate ReID information is improved. At the same time, the single-frame output is changed to a double-frame output, and the double-frame loss calculation is performed to improve training efficiency and bias.
By utilizing timing information and feature enhancement modules, the multi-objective tracking performance of the model on the drone video sequence is significantly improved, the IDSwitch phenomenon is reduced, and the detection accuracy and training efficiency are improved.
Smart Images

Figure CN115035159B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a video multi-target tracking method based on deep learning and temporal feature enhancement. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, computer vision technology has also penetrated into various fields, and its application scenarios are becoming more and more extensive. Making full use of computer vision technology can greatly improve the efficiency of environmental monitoring and national defense security monitoring. And using computer vision technology to complete multi-target detection and tracking tasks has gradually become one of the research hotspots in recent years. Multi-target tracking refers to the detection of all target objects in each frame of a continuous frame sequence of a video, obtaining the position, bounding box size, speed and other attributes of each target, and assigning the same ID to the same individual in each frame, thereby completing the detection and tracking tasks of multiple targets in a video sequence.
[0003] The main idea of tracking based on detection is to complete the subsequent tracking and matching based on the results of target detection. The mainstream bounding box-based target detection algorithms include the YOLO series of algorithms, Faster RCNN, and RetinaNet. In recent years, many anchor-free algorithms for target detection based on center points have also appeared, such as CenterNet, FCOS, CenterPoint, and other target detection algorithms. In the TBD algorithm system, after target detection, data association and matching between frames are required. SORT and DeepSORT algorithms are relatively classic target matching algorithms, which can often be combined with simple target detection algorithms to complete multi-target tracking tasks. The task of multi-target tracking can also be achieved based on the anchor-free framework. In the JDT framework structure, such as RetinaTrack and FairMOT, the two methods combine detection and tracking, simplifying the model structure and improving the real-time performance of the calculation.
[0004] Multi-target tracking is a cross-frame video interpretation task, and most models in current research do not make good use of time series information. Relying only on the image information of the current frame has certain limitations, and the target lacks the connection between frames. For example, if an object is occluded in a certain frame, if data association is performed only based on single-frame information, the same object representation information will often be different, which will lead to IDSwitch, thereby reducing the accuracy of the model. Therefore, how to make good use of time series information can greatly improve the performance of the model. In addition, although the JDT paradigm model jointly trains detection and data association to achieve end-to-end multi-target tracking, target detection and tracking are often two different visual tasks. Target detection requires distinguishing objects of multiple categories, maximizing the distance between objects of different categories, and minimizing the distance between objects of the same category, so as to improve the accuracy of target detection. However, target tracking requires maximizing the distance between all objects of the same category. Therefore, if the two subtasks share more parameters during training, the training efficiency of the model may be reduced, and the performance of the trained model may deteriorate in some cases. Summary of the invention
[0005] In view of the above problems, the present invention proposes a video multi-target tracking method based on deep learning and temporal feature enhancement, taking FairMOT as the original benchmark structure. For the overall model structure, the structures of the original model detection and feature generation are first split, and a feature enhancement module based on temporal information is added to the ReID branch to improve the model's ability to discriminate ReID information. In the calculation process of the model loss, compared with the original single-frame loss calculation, we perform dual-frame output in the output part of the model detection, and perform loss calculation on the outputs of adjacent frames at the same time, thereby improving the training efficiency and bias of the model.
[0006] In order to achieve the above object, the present invention provides a video multi-target tracking method based on deep learning and temporal feature enhancement, comprising the following steps:
[0007] S1. Prepare and process the data set and use the processed data as input data for model training and testing;
[0008] S2, separate the object detection and ReID tasks in the model structure;
[0009] S3, using time series information to build a ReID task module to improve the model structure;
[0010] S4. Post-processing reasoning of the model: applying the improved model structure to the data association matching process of multi-target tracking.
[0011] Preferably, the step S1 specifically includes the following steps:
[0012] S11, collect a set of drone video sequences as a dataset;
[0013] S12, marking the data set into a coco format, wherein the coco format can provide the frame number, target ID, coordinates of the upper left vertex of the bounding box, width and height of the bounding box, whether the target is blocked, and whether the target needs to be ignored;
[0014] S13, counting IDs of the data set according to categories;
[0015] S14, rotating and scaling each image in the data set.
[0016] Preferably, the step S2 specifically includes the following steps:
[0017] S21, the decoder of the backbone network on the model is changed to two decoders with the same structure for target detection and ReID tasks respectively;
[0018] S22, the model input is changed to double-frame input and the parameters of the two frames are shared and then feature extraction is performed through the encoder;
[0019] S23, input the extracted features into the two decoders with the same structure to perform target detection and ReID tasks respectively.
[0020] Preferably, the step S23 is specifically as follows: in the target detection part, first, a multi-layer convolution is performed after the feature of the previous frame obtained by the decoder, and the feature map is spliced with the feature of the current frame obtained by the decoder, and finally the output of the target detection branch is obtained through the heat map branch; in the ReID task part, a feature enhancement module is added, and the adjacent frame features obtained by the decoder and the heat map of the previous frame are used as input information of the feature module, and the output of the ReID task branch is obtained after the information of the module is integrated.
[0021] Preferably, the step S3 is specifically divided into a training phase and an inference phase.
[0022] Preferably, the training phase specifically includes the following steps:
[0023] S311, obtaining the features of the corresponding position of the feature map in the previous frame through the annotation information of the data set, and calculating the similarity between the features and the current feature map to obtain the feature distance between each object in the previous frame and each point in the current frame;
[0024] S312, after obtaining the position information of the pairwise correspondence between the feature map in the previous frame and the current feature map, perform feature fusion.
[0025] Preferably, the reasoning stage specifically includes the following steps:
[0026] S321, using the heat map to obtain the number of targets that may exist in the previous frame, and using the ReID feature information of the corresponding positions of these targets as one of the inputs to the feature module;
[0027] S322, setting a threshold. If the distance between the center point of the previous frame and the center point of the matched current frame exceeds the threshold, the matched point is considered unreliable and ignored, and only the matching point with high credibility is retained for feature fusion with the current feature map.
[0028] Preferably, the step S3 specifically comprises performing feature fusion on the heat map of the previous frame, the feature map of the previous frame and the feature map of the current frame.
[0029] Preferably, the step S4 specifically includes the following steps:
[0030] S41, taking three frames as a round, normalize and standardize the heat map and ReID features obtained by the model in the first frame, perform non-maximum suppression on the heat map, screen out possible objects according to the set threshold, and assign IDs to the objects in the first frame;
[0031] S42, the second frame repeats the operation of the first frame, and after obtaining the possible objects, matches the bounding box iou with the objects in the first frame, retains the detections that meet the expectations, assigns the same ID, and retains the unmatched objects;
[0032] S43, the third frame adds ReID features based on the second frame, calculates the cosine distance of the detected targets in adjacent frames using the ReID features, performs motion prediction using Kalman filtering, and associates data based on appearance and motion features;
[0033] S44, calculate the iou of the unmatched objects in the third frame and the objects in the previous frame. If it is less than a fixed threshold, it is regarded as a new target and assigned a new ID. Repeat the above steps for each subsequent frame to complete the post-processing steps of video multi-target tracking.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] A video multi-target tracking method based on deep learning and temporal feature enhancement provided by the present invention can realize end-to-end detection and tracking of multi-category objects in drone video sequences. Firstly, the problem that there may be certain conflicts between target detection and ReID tasks during training is improved, and the target detection and ReID branches are separated to make the two structures more independent and improve the detection accuracy. In addition, the temporal information is utilized, the center point features of the historical frames are combined and a feature enhancement module is added to improve the multi-target tracking performance of the model on the drone video sequence. On the visdrone multi-target tracking dataset, compared with other multi-target tracking models for drone videos, the use of this method can achieve better multi-category multi-target tracking effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is the overall structure of the algorithm model of the present invention;
[0037] Figure 2 This is the first improved model structure diagram of the ReID task module constructed by using time series information in the present invention;
[0038] Figure 3 This is the second improved model structure diagram of the ReID task module constructed by utilizing time series information in the present invention. DETAILED DESCRIPTION
[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0040] In view of the problems and shortcomings in the prior art, the present invention provides a video multi-target tracking method based on deep learning and temporal feature enhancement. The present invention has made improvements on the structure of the model: 1. The single-frame output of the original model is changed to double-frame output to improve the training efficiency and bias of the model; 2. The conflict problem during target detection and ReID task training under the JDT paradigm is improved; 3. A feature enhancement module based on temporal information is constructed to improve the representation of ReID information, thereby improving the training efficiency of the ReID branch of the model and ensuring the accuracy of tracking and matching.
[0041] The present invention proposes a video multi-target tracking method based on deep learning and temporal feature enhancement, comprising the following steps:
[0042] S1. Prepare and process the data set and use the processed data as input data for model training and testing;
[0043] S2, separate the object detection and ReID tasks in the model structure;
[0044] S3, using time series information to build a ReID task module to improve the model structure;
[0045] S4. Post-processing reasoning of the model: applying the improved model structure to the data association matching process of multi-target tracking.
[0046] The following is a detailed description of each step.
[0047] Step S1: prepare and process the data set, and use the processed data as input data for model training and testing, that is, prepare and process the data set, process the image information and annotation information, obtain the required training data for the multi-target tracking task, and use it as input data for model training and testing. The main steps are:
[0048] S11. Collect a set of drone video sequences as a data set; that is, select a corresponding set of drone video sequences, the data set includes a training set and a validation set, and the training set and the validation set include multiple drone video sequences. Each video is an optical image, containing different scenes and targets. The size and format of each frame in the video are consistent, and different video sequences have different image sizes and shooting methods. The data set can mainly contain multiple target categories, such as pedestrians, vehicles, trucks, etc.
[0049] S12. Annotate the dataset into coco format (the COCO format of the dataset consists of a JSON file, which contains all the details of the image, such as size, annotations (i.e. bounding box coordinates), labels corresponding to its bounding box, etc.). Compared with the target detection dataset, the dataset is annotated with an additional ID. Therefore, the original format of the annotation can provide information mainly including the frame number, target ID, coordinates of the upper left vertex of the bounding box, width and height of the bounding box, whether the target is occluded and whether it needs to be ignored, etc.
[0050] S13, counting IDs of the data set according to categories;
[0051] Compared with the single-category multi-target tracking task, multi-category multi-target tracking requires additional processing of the data set. For the original annotation, since it is multi-category data information, it is necessary to process multi-category IDs when constructing training data. The main method is to count the IDs according to the category, that is, the ID of each category starts counting from 0.
[0052] S14, rotating and scaling each image in the data set;
[0053] Compared with single-category tasks, it is necessary to count the number of target IDs of each category that appear in the entire video sequence, and use them as classification categories when training the ReID branch of the model. The model performs data augmentation during training, and rotates and scales each image imported into the model, which improves the training effect of the model.
[0054] Step S2: Separate the object detection and ReID tasks in the model structure.
[0055] The present invention uses FairMOT as the original benchmark structure. FairMOT is an end-to-end anchor-free multi-target tracking framework built based on CenterNet. The feature extraction part of the model uses DLA34 as the backbone network to extract features from two-dimensional video images, and then generates multiple branch heads according to different visual tasks, namely heat map branches, offset branches, wh branches and ReID branches. These branches share the feature map after feature extraction. Compared with the feature extraction part of the original FairMOT network framework, it is considered that there is a certain degree of conflict between target detection and ReID tasks during training, that is, the target detection task is to maximize the distance between objects of different categories and minimize objects of the same category, while the ReID task is to maximize the distance between different individuals of the same category, so as to achieve the goal of accurate re-identification. Therefore, for the two tasks of target detection and ReID, conflicts often occur in some aspects during model training, so it is necessary to adjust the model structure for this problem. The main steps are:
[0056] S21, the decoder of the backbone network on the model is changed to two decoders with the same structure for target detection and ReID tasks respectively;
[0057] In order to solve the conflict between ReID and object detection, the model of the present invention mainly separates the two branches. The functions of the two structures on the original FairMOT are realized by four branches, and these branches share the same encoder and decoder. The structural adjustment of the model of the present invention in this regard is mainly to improve the decoder part of the backbone network DLA34, and change it into two decoders with the same structure, which are used for object detection and ReID respectively. However, the two do not share parameters in the decoder part, thus reducing the mutual influence between parameters during model training. Figure 1 The above is the overall structure of the algorithm model of the present invention.
[0058] S22, the model input is changed to double-frame input and the parameters of the two frames are shared and then feature extraction is performed through the encoder;
[0059] Compared with the single-frame input of FairMOT, the input of the model of the present invention is changed to the input of adjacent frames. In the test phase, if it is the first frame of the video sequence, two first frame images are input. In the input part, the two frames are processed by parameter sharing, and feature extraction is performed through the DLA34 encoder.
[0060] S23, inputting the extracted features into the two decoders with the same structure to perform target detection and ReID tasks respectively;
[0061] The obtained features are input into two decoders at the same time, and information processing for target detection and ReID is performed respectively. In the target detection part, compared with the original FairMOT structure, the structure is adjusted on the branch of the heat map. First, a multi-layer convolution is added after the features of the previous frame obtained by the decoder, and a central feature map with a thickness of 1 is added. The feature map is spliced with the features of the current frame obtained by the decoder, and finally the output of the branch is obtained through the heat map branch. The outputs of the model of the present invention in these two branches are the prediction results of adjacent frames, and the loss is calculated for these two prediction results at the same time during training. In the ReID branch, compared with the original FairMOT, this branch adds a feature enhancement module, and the adjacent frame features obtained by decoder B and the heat map of the previous frame are used as input information of the feature module. After the information integration of the module, the final output of the ReID branch is obtained.
[0062] Step S3: Use the time series information to construct a ReID task module to improve the model structure. The present invention provides two methods of using the time series information to construct a ReID task module to improve the model structure, such as Figure 2 The figure shows the first improved model structure diagram of the ReID task module constructed by utilizing time series information in the present invention.
[0063] The first method uses time series information to build a ReID task module to improve the model structure, which is divided into a training phase and an inference phase. The training phase specifically includes the following steps:
[0064] S311, obtaining the features of the corresponding position of the feature map in the previous frame through the annotation information of the data set, and calculating the similarity between the features and the current feature map to obtain the feature distance between each object in the previous frame and each point in the current frame;
[0065] In the training phase, auxiliary training is performed by inputting annotation information, that is, the number of objects that exist simultaneously in two adjacent frames and the corresponding position index. This information is used to obtain the features of the corresponding positions of the feature map in the previous frame, and the similarity is calculated between it and the current feature map, so as to obtain the feature distance between each object in the previous frame and each point in the current frame, and retain the point with the smallest distance as the possible position of the object in the previous frame in the current frame image.
[0066] S312, after obtaining the position information of the pairwise correspondence between the feature map in the previous frame and the current feature map, perform feature fusion;
[0067] After obtaining the pairwise corresponding position information, feature fusion is performed. The fusion method selected in this step is to add and average the corresponding feature matrices.
[0068] The reasoning phase specifically includes the following steps:
[0069] S321, using the heat map to obtain the number of targets that may exist in the previous frame, and using the ReID feature information of the corresponding positions of these targets as one of the inputs to the feature module;
[0070] In the inference stage of the model, since there is no annotation information, the heat map of the previous frame obtained by the model is needed as auxiliary information. The number of targets that may have existed in the previous frame is obtained from the heat map, and the ReID feature information of the corresponding positions of these targets is used as one of the inputs to the feature module.
[0071] S322, setting a threshold value. If the distance between the center point of the previous frame and the center point of the matched current frame exceeds the threshold value, the matched point is considered unreliable and ignored, and only the matching point with high credibility is retained to perform feature fusion with the current feature map.
[0072] Since an object that appeared in the previous frame may disappear in the current frame during the inference stage, a distance constraint needs to be added and a threshold needs to be set. That is, if the center point of the previous frame is far away from the center point of the matched current frame and exceeds the threshold, the matched point is considered unreliable and is ignored. Only matching points with high credibility are retained, and feature fusion is performed in the same way as in the training stage as the final output of the ReID branch.
[0073] The second method uses time series information to build a ReID task module to improve the model structure, specifically by fusing the heat map of the previous frame, the feature map of the previous frame, and the feature map of the current frame; Figure 3 The figure shows the structure diagram of the second improved model of the ReID task module constructed by utilizing time series information in the present invention.
[0074] The second method of using temporal information mainly inputs three parts, namely the heat map of the previous frame and the feature map information of the adjacent frames. After the three are channel-joined, they are passed through the multi-layer convolution of the ReID branch as the final output on this branch.
[0075] In the training phase of the model, the heat map input of the previous frame is also provided as annotation information, and in the inference phase, the heat map obtained by the model detection part is used as the input of the ReID branch. The following formula represents the specific loss function of the training model algorithm in this method:
[0076]
[0077]
[0078]
[0079]
[0080] Formula (1) represents the training loss function of the heat map branch, (2) and (3) represent the loss functions of the detection box width and height and ReID itself respectively, and formula (4) represents the integration of the loss function of the final model training and the corresponding weights of each part.
[0081] Step S4: Post-processing reasoning of the model, applying the improved model structure to the data association matching process of multi-target tracking. The main steps are:
[0082] The post-processing part of the model is roughly the same as the original FairMOT, and data association is mainly completed through DeepSORT. Compared with single-category multi-target tracking, this method adjusts the post-processing part on multiple categories. Unlike the training stage, the post-processing stage does not assign IDs to each category, that is, multiple categories are assigned IDs together according to the order of object detection.
[0083] S41, taking three frames as a round, normalize and standardize the heat map and ReID features obtained by the model in the first frame, perform non-maximum suppression on the heat map, screen out possible objects according to the set threshold, and assign IDs to the objects in the first frame;
[0084] S42, the second frame repeats the operation of the first frame, and after obtaining the possible objects, matches the bounding box iou with the objects in the first frame, retains the detections that meet the expectations, assigns the same ID, and retains the unmatched objects;
[0085] The reasoning part uses DeepSORT as the main process framework, with three frames as a round. In the first frame, the heat map and ReID features obtained by the model are normalized and standardized, and the heat map is subjected to non-maximum suppression. Possible objects are screened out according to the set threshold, and the objects in the first frame are assigned an ID. The second frame repeats the operation of the first frame, and after obtaining the possible objects, the bounding box iou is matched with the objects in the first frame. The detections that meet the expectations are retained, assigned the same ID, and those unmatched objects are retained.
[0086] S43, the third frame adds ReID features based on the second frame, calculates the cosine distance of the detected targets in adjacent frames using the ReID features, performs motion prediction using Kalman filtering, and associates data based on appearance and motion features;
[0087] S44, calculate the iou of the unmatched objects in the third frame and the objects in the previous frame. If it is less than a fixed threshold, it is regarded as a new target and assigned a new ID. Repeat the above steps for each subsequent frame to complete the post-processing steps of video multi-target tracking.
[0088] Although the present invention is described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. It should therefore be understood that many modifications may be made to the exemplary embodiments and that other arrangements may be devised without departing from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the various dependent claims and features described herein may be combined in a manner different from that described in the original claims. It should also be understood that features described in conjunction with individual embodiments may be used in other described embodiments.
Claims
1. A video multi-target tracking method based on deep learning and temporal feature enhancement, It is characterized in that The following steps are involved: S1. Prepare and process the data set and use the processed data as input data for model training and testing; S2, separate the object detection and ReID tasks in the model structure; S3, using time series information to build a ReID task module to improve the model structure; S4, post-processing reasoning of the model, applying the improved model structure to the data association matching process of multi-target tracking; The step S3 is specifically divided into a training phase and an inference phase; The training phase specifically includes the following steps: S311, obtaining the features of the corresponding position of the feature map in the previous frame through the annotation information of the data set, and calculating the similarity between the features and the current feature map to obtain the feature distance between each object in the previous frame and each point in the current frame; S312, after obtaining the position information of the pairwise correspondence between the feature map in the previous frame and the current feature map, perform feature fusion; The reasoning stage specifically includes the following steps: S321, using the heat map to obtain the number of targets that may exist in the previous frame, and using the ReID feature information of the corresponding positions of these targets as one of the inputs to the feature module; S322, setting a threshold. If the distance between the center point of the previous frame and the center point of the matched current frame exceeds the threshold, the matched point is considered unreliable and ignored, and only the matching point with high credibility is retained for feature fusion with the current feature map.
2. According to claim 1, a video multi-target tracking method based on deep learning and temporal feature enhancement, It is characterized in that The step S1 specifically includes the following steps: S11, collect a set of drone video sequences as a dataset; S12, marking the data set into a coco format, wherein the coco format can provide the frame number, target ID, coordinates of the upper left vertex of the bounding box, width and height of the bounding box, whether the target is blocked, and whether the target needs to be ignored; S13, counting IDs of the data set according to categories; S14, rotating and scaling each image in the data set.
3. According to claim 1, a video multi-target tracking method based on deep learning and temporal feature enhancement, It is characterized in that The step S2 specifically includes the following steps: S21, the decoder of the backbone network on the model is changed to two decoders with the same structure for target detection and ReID tasks respectively; S22, the model input is changed to double-frame input and the parameters of the two frames are shared and then feature extraction is performed through the encoder; S23, input the extracted features into the two decoders with the same structure to perform target detection and ReID tasks respectively.
4. According to claim 3, a video multi-target tracking method based on deep learning and temporal feature enhancement, It is characterized in that The step S23 is specifically as follows: in the target detection part, first, a multi-layer convolution is performed after the feature of the previous frame obtained by the decoder, and the feature map is spliced with the feature of the current frame obtained by the decoder, and finally the output of the target detection branch is obtained through the heat map branch; in the ReID task part, a feature enhancement module is added, and the adjacent frame features obtained by the decoder and the heat map of the previous frame are used as input information of the feature module, and the output of the ReID task branch is obtained after the information of the module is integrated.
5. According to claim 1, a video multi-target tracking method based on deep learning and temporal feature enhancement, It is characterized in that The step S4 specifically comprises the following steps: S41, taking three frames as a round, normalize and standardize the heat map and ReID features obtained by the model in the first frame, perform non-maximum suppression on the heat map, screen out possible objects according to the set threshold, and assign IDs to the objects in the first frame; S42, the second frame repeats the operation of the first frame, and after obtaining the possible objects, matches the bounding box iou with the objects in the first frame, retains the detections that meet the expectations, assigns the same ID, and retains the unmatched objects; S43, the third frame adds ReID features based on the second frame, calculates the cosine distance of the detected targets in adjacent frames using the ReID features, performs motion prediction using Kalman filtering, and associates data based on appearance and motion features; S44, calculate the IOU of the unmatched object in the third frame with the object in the previous frame. If it is less than a fixed threshold, it is regarded as a new target and assigned a new ID. Repeat the above steps for each subsequent frame to complete the post-processing steps of video multi-target tracking.
Citation Information
Patent Citations
Pedestrian multi-target tracking method based on multivariate difference fusion
CN113221787A
Online multi-target tracking method of unified target motion perception and re-identification network
CN113313736A