Object Tracking Method, Device and Medium Based on Single-Stage Object Tracking Model
By combining feature extraction, detection and pedestrian re-identification branches based on a single-stage target tracking model, the problem of slow target tracking speed and low accuracy in the prior art is solved, and efficient pedestrian tracking effect is achieved.
Patent Information
- Application Number
- CN202211287178.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-20
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-10-20
AI Technical Summary
In the prior art, the target tracking algorithm based on the pedestrian re-identification network has the problem that the inference time is long, the target tracking speed is reduced, and the tracking accuracy is low.
A single-stage target tracking model is used to process each frame of the image in the video data using the feature extraction module. The target box regression and low-dimensional feature vector extraction are performed respectively by detecting branches and pedestrian recognizing branches, and the target box and low-dimensional feature vector of the tracking object are tracked.
While ensuring the tracking speed, the accuracy and robustness of target tracking are improved and the accuracy of tracking is improved.
Smart Images

Figure CN115661198B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to an object tracking method, device, and medium based on a single-stage object tracking model. Background Art
[0002] Pedestrian detection is to use computer vision technology to identify whether there are pedestrians in an image or video stream and give precise positioning. This technology has a wide range of application fields. It can be combined with technologies such as pedestrian tracking and pedestrian re-identification, and can be well applied to real-world scenario fields such as artificial intelligence systems, vehicle assisted driving systems, intelligent video surveillance, human behavior analysis, and intelligent transportation.
[0003] Currently, in the object tracking method based on the pedestrian re-identification network, each frame of image is extracted from the video stream, and then the image is input into the pedestrian re-identification network for object detection. According to the results of object detection, the position of the pedestrian in the image is determined. The current object detection methods are mainly divided into single-stage object detection methods and two-stage object detection methods. Among them, the single-stage object detection method has a shorter inference time and a faster tracking speed, but the tracking accuracy is lower. Although the two-stage object detection method has a certain accuracy, the inference time of the two-stage object detection algorithm is longer, which reduces the speed of object tracking. Summary of the Invention
[0004] In view of this, the embodiments of this application provide an object tracking method, device, and medium based on a single-stage object tracking model to solve the problems in the prior art that the inference time of the object tracking algorithm is long, reducing the object tracking speed, and the tracking accuracy is low.
[0005] In the first aspect of the embodiments of this application, an object tracking method based on a single-stage object tracking model is provided, including: obtaining video data containing a tracking object and inputting the video data into a predetermined single-stage object tracking model; using a feature extraction module in the single-stage object tracking model to process each frame of image in the video data to obtain a feature map; inputting the feature map into a detection branch in the single-stage object tracking model, and using the detection branch to regress the position of the tracking object in the image to obtain a target box corresponding to the tracking object, where the detection branch uses an Anchor-Free based detection module; inputting the feature map into a pedestrian re-identification branch in the single-stage object tracking model, and using the pedestrian re-identification branch to extract a low-dimensional feature vector of the feature map to obtain a pedestrian re-identification low-dimensional feature vector; based on the target box corresponding to the tracking object and the pedestrian re-identification low-dimensional feature vector, tracking the trajectory generated by the tracking object in the video data.
[0006] In the second aspect of the embodiments of the present application, a target tracking device based on a single-stage target tracking model is provided, including: an input module configured to obtain video data containing a tracking object and input the video data into a predetermined single-stage target tracking model; a processing module configured to process each frame of the image in the video data by using a feature extraction module in the single-stage target tracking model to obtain a feature map; a regression module configured to input the feature map into a detection branch in the single-stage target tracking model and use the detection branch to regress the position of the tracking object in the image to obtain a target box corresponding to the tracking object, where the detection branch uses an Anchor-Free based detection module; an extraction module configured to input the feature map into a person re-identification branch in the single-stage target tracking model and use the person re-identification branch to extract a low-dimensional feature vector of the feature map to obtain a person re-identification low-dimensional feature vector; a tracking module configured to track the trajectory generated by the tracking object in the video data based on the target box corresponding to the tracking object and the person re-identification low-dimensional feature vector.
[0007] In the third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above method are implemented.
[0008] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0009] The above at least one technical solution adopted in the embodiments of the present application can achieve the following beneficial effects:
[0010] By obtaining video data containing a tracking object, inputting the video data into a predetermined single-stage target tracking model; processing each frame of the image in the video data by using a feature extraction module in the single-stage target tracking model to obtain a feature map; inputting the feature map into a detection branch in the single-stage target tracking model and using the detection branch to regress the position of the tracking object in the image to obtain a target box corresponding to the tracking object, where the detection branch uses an Anchor-Free based detection module; inputting the feature map into a person re-identification branch in the single-stage target tracking model and using the person re-identification branch to extract a low-dimensional feature vector of the feature map to obtain a person re-identification low-dimensional feature vector; tracking the trajectory generated by the tracking object in the video data based on the target box corresponding to the tracking object and the person re-identification low-dimensional feature vector. The single-stage target tracking model based on Anchor-Free in the present application improves the accuracy of target tracking and the robustness of tracking while ensuring the tracking speed. Description of the Drawings
[0011] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0012] Figure 1 is a schematic structural diagram of a single-stage object tracking model provided by an embodiment of the present application;
[0013] Figure 2 is a schematic flowchart of an object tracking method based on a single-stage object tracking model provided by an embodiment of the present application;
[0014] Figure 3 is a schematic diagram of the principle of data augmentation using the MixUp augmentation algorithm provided by an embodiment of the present application;
[0015] Figure 4 is a schematic structural diagram of a detection branch based on Anchor-Free provided by an embodiment of the present application;
[0016] Figure 5 is a schematic structural diagram of a Re-ID branch provided by an embodiment of the present application;
[0017] Figure 6 is a schematic structural diagram of an object tracking device based on a single-stage object tracking model provided by an embodiment of the present application;
[0018] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0019] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0020] Multi-object tracking (MOT) has always been a long-term goal in computer vision. The goal is to estimate the trajectories of multiple objects in a video. The successful solution of this task will benefit many applications, such as action recognition, sports video analysis, elderly care, and human-computer interaction. Most of the existing state-of-the-art (SOTA) methods adopt a two-step approach for object detection. Although with the development of object detection algorithms and Re-ID in recent years, the two-step method has shown significant performance improvement in object tracking, it does not share the feature maps of the detection algorithm and Re-ID, so it is very slow and difficult to perform inference at video rate. Therefore, there is an urgent need to provide a fast object tracking algorithm.
[0021] Current object detection methods are mainly divided into single-stage object detection methods and two-stage object detection methods. Among them, although the two-stage object detection method has a certain accuracy, the inference time of the two-stage object detection algorithm is relatively long, which reduces the speed of object tracking. With the maturity of the two-step object tracking algorithm, more researchers have started to study the one-shot algorithm that simultaneously detects objects and learns Re-ID features. After the feature maps are shared between object detection and Re-ID, the inference time can be greatly reduced, but the accuracy will be much lower than that of the two-step method. Therefore, there is a need to provide a single-stage object tracking algorithm that can ensure both tracking speed and tracking accuracy.
[0022] In view of this, to solve the above problems, the embodiments of the present application provide a single-stage object tracking algorithm based on Anchor-Free. First, the feature extraction module is used to process each frame of the image to obtain a feature map. Then, the feature map is input into the detection branch and the person re-identification branch respectively. The detection branch is used to regress the position of the tracking object in the image to determine the target box of the tracking object. The person re-identification branch is used to extract the low-dimensional feature vector in the feature map. Finally, based on the target box corresponding to the tracking object and the person re-identification low-dimensional feature vector, single-stage object tracking of pedestrians is realized. The content of the technical solution of the present application will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0023] Figure 1 It is a schematic structural diagram of the single-stage object tracking model provided by the embodiments of the present application. As Figure 1 shown, the single-stage object tracking model may specifically include:
[0024] The single-stage object tracking model of the embodiment of the present application includes an input module Input, a feature extraction module Backbone, a detection branch Detection Head, and a person re-identification branch Re-ID Head; among them, the input module Input is used to process video data into a series of image frames, and sequentially input each frame of image into the single-stage object tracking model for object detection; the feature extraction module Backbone is used to extract feature maps from images, and a deformable convolutional network DCN is inserted into the feature extraction module Backbone; the detection branch Detection Head adopts an Anchor-Free based detection module, which is used to regress the specific position of the tracking object (such as a pedestrian) from the feature map to obtain the target box corresponding to the tracking object; the person re-identification branch Re-ID Head is used to extract low-dimensional feature vectors from the feature map and input Re-ID feature information (i.e., low-dimensional vector information).
[0025] Next, based on Figure 1 the structure of the single-stage object tracking model shown, the implementation process of the object tracking method based on the single-stage object tracking model of the present application will be described in detail.
[0026] Figure 2 is a schematic flowchart of the object tracking method based on the single-stage object tracking model provided by the embodiment of the present application. Figure 2 The object tracking method based on the single-stage object tracking model can be executed by a server. As Figure 2 shown, the object tracking method based on the single-stage object tracking model may specifically include:
[0027] S201, obtain video data containing a tracking object, and input the video data into a predetermined single-stage object tracking model;
[0028] S202, use the feature extraction module in the single-stage object tracking model to process each frame of image in the video data to obtain a feature map;
[0029] S203, input the feature map into the detection branch in the single-stage object tracking model, and use the detection branch to regress the position of the tracking object in the image to obtain the target box corresponding to the tracking object, where the detection branch adopts an Anchor-Free based detection module;
[0030] S204, input the feature map into the person re-identification branch in the single-stage object tracking model, and use the person re-identification branch to extract the low-dimensional feature vector of the feature map to obtain the person re-identification low-dimensional feature vector;
[0031] S205. Track the trajectory generated by the tracking object in the video data based on the target box corresponding to the tracking object and the low-dimensional feature vector of person re-identification.
[0032] Specifically, when tracking the trajectory of a person over a period of time, what is often obtained is a video stream. The video stream is input into the model for detection. However, the object processed by the single-stage object tracking model is an image. The single-stage object tracking model actually performs separate object detection processing on each frame of the video stream. Therefore, the embodiments of the present application can either obtain video data containing the tracking object or directly obtain image data and use the image data as the input of the single-stage object tracking model. The type of input data obtained does not limit the technical solution of the present application.
[0033] Furthermore, the single-stage object tracking model of the embodiments of the present application adopts a detection branch based on Anchor-Free. Compared with the Anchor-Based object detection algorithm, a large number of detection boxes need to be generated in the image, and one box closest to the target position is selected from the large number of detection boxes as the detection result. However, with Anchor-Free, the position of the pedestrian in the image can be directly regressed without generating a large number of detection boxes, thereby accelerating the detection speed of the object detection algorithm.
[0034] In some embodiments, the pre-training process of the single-stage object tracking model includes: obtaining pre-configured detection data, performing data augmentation on the detection data using the MixUp augmentation algorithm to obtain new detection data, and using the new detection data to train the detection branch in the single-stage object tracking model; after training the detection branch, using the pre-configured tracking data to train the complete single-stage object tracking model to obtain the trained single-stage object tracking model, where the tracking data includes target boxes and the object identifier corresponding to each target box.
[0035] Specifically, before using the single-stage object tracking model for object tracking, the single-stage object tracking model needs to be pre-trained first. During the pre-training process, first perform data augmentation on the detection data using the MixUp augmentation algorithm, and use the new detection data formed after data augmentation to train the Anchor-Free-based detection branch. When training the detection branch, the person re-identification branch can be removed from the model first; after training the detection branch, use the tracking data to train the complete single-stage object tracking model.
[0036] It should be noted that the detection data may only contain the target bounding boxes without the object identifiers corresponding to the target bounding boxes, while the tracking data contains both the target bounding boxes and the object identifiers corresponding to each target bounding box. Here, the object identifier can be regarded as the label of the person re-identification branch (Re-ID branch). The object identifier of a target bounding box refers to which tracking object (such as which pedestrian) the target bounding box corresponds to.
[0037] In some embodiments, the MixUp augmentation algorithm is used to augment the detection data to obtain new detection data, including: fusing any two original images in the detection data so as to fuse the two original images into one image to obtain the fused image, superimposing the target bounding boxes in the original images in the fused image, and generating new detection data based on the original images and the fused image.
[0038] Specifically, MixUp is an algorithm for image cross-class augmentation used in computer vision. It can mix images between different classes to expand the training data set. The implementation process and principle of the MixUp augmentation algorithm will be described below with reference to the accompanying drawings. Figure 3 is a schematic diagram of the principle of data augmentation using the MixUp augmentation algorithm provided by an embodiment of the present application. As Figure 3 shown, the process of data augmentation using the MixUp augmentation algorithm can specifically include:
[0039] In the pre-training stage of the Detection Head in the detection branch, in order to improve the performance of the detection algorithm, the MixUp augmentation algorithm is used for data augmentation in this embodiment of the present application. The principle of this augmentation algorithm is to fuse two groups of pictures into one group, that is, to pairwise fuse any two groups of original pictures in the detection data to obtain new pictures, and use the obtained new pictures as the fused new pictures. At the same time, the target bounding boxes in the original pictures will also be superimposed in the new pictures. Finally, the original pictures and the fused new pictures are jointly used to generate new detection data, and the new detection data is used to train the Detection Head in the detection branch.
[0040] In some embodiments, a deformable convolutional network is inserted into the feature extraction module. The deformable convolutional network uses a ResNet34 residual network. There are 2 3×3 convolutional layers in the stem stage of the ResNet34 residual network. The strides of the first two convolutional layers in the residual branch of the original residual module are exchanged, and the 1×1 convolutional layer with a stride of 2 in the original shortcut branch is replaced with an average pooling layer with a stride of 2 and a 1×1 convolutional layer.
[0041] Specifically, a deformable convolutional network (DCN) is inserted into the feature extraction module Backbone of the embodiments of the present application. In practical applications, the deformable convolutional network DCN can adopt the ResNet34 network. Compared with the existing ResNet34 network, the following improvements are made to the original ResNet34 network in the present application:
[0042] In the feature extraction stage, in order to extract richer features, the embodiments of the present application improve the original ResNet34 network. The stem stage in the original ResNet Backbone includes a convolutional layer with a size of 7×7 and a stride of 2. Since the computational amount of convolution is quadratic with respect to the length and width, the 7×7 convolution has a computational amount 5.4 times larger than that of the 3×3 convolution. Therefore, the embodiments of the present application replace the 7×7 convolution kernel here with three traditional 3×3 convolution kernels, thereby reducing the convolution computational amount while maintaining the same receptive field as the 7×7 convolution kernel. Then, the embodiments of the present application exchange the strides of the first two convolutional layers in the residual branch of the residual module to avoid information loss caused by the 1×1 convolution with a stride of 2. Similarly, the 1×1 convolution with a stride of 2 in the shortcut branch is replaced with an average pooling layer with a stride of 2 and a 1×1 convolution, which can further reduce the information loss to the feature map in the residual module.
[0043] In some embodiments, the detection branch includes a heatmap branch, a center offset branch, and a box branch. The heatmap branch, the center offset branch, and the box branch are all composed of 2 3×3 convolutional layers and 1 1×1 convolutional layer. Among them, the heatmap branch is used to output a heatmap with a size of (1, H, W), the center offset branch is used to output a center offset with a size of (2, H, W), and the box branch is used to output box coordinates with a size of (2, H, W).
[0044] Specifically, currently, Anchor-Based trackers will simultaneously extract the Re-ID information (low-dimensional vector information) of the detection box and the object within the detection box. However, the anchors generated by Anchor-Based detectors are not suitable for learning appropriate Re-ID information because an object may be responsible for detection by multiple anchors, but the differences between these anchors will be very large, and the extracted Re-ID differences will also be particularly large. Therefore, this will lead to serious network ambiguity, and thus the quality of the Re-ID features is severely affected by the quality of the detection box. The framework of detecting first and then re-identifying makes the network unable to fairly learn the Re-ID branch. In view of the problems existing in the existing Anchor-Based trackers, the embodiments of the present application provide an Anchor-Free detection branch.
[0045] The structure of the detection branch based on Anchor-Free provided in the embodiments of the present application will be described below in conjunction with the accompanying drawings. Figure 4 It is a schematic diagram of the structure of the detection branch based on Anchor-Free provided in the embodiments of the present application. As Figure 4 shown, the detection branch based on Anchor-Free may specifically include:
[0046] The detection branch structure includes the following three branches, namely the heatmap branch, the center offset branch, and the box branch. Each branch is composed of 2 3x3 convolutional layers and 1 1x1 convolutional layer. The output size of the heatmap branch is (1, H, W), the output size of the center offset branch is (2, H, W), and the output size of the box branch is (2, H, W), where H is the height and W is the width.
[0047] In some embodiments, the detection branch is used to regress the position of the tracking object in the image to obtain the target box corresponding to the tracking object, including: regressing the position of the tracking object in the image based on the heatmap, the center offset, and the box coordinates to obtain the target box corresponding to the tracking object in the image, and determining the position of the target box in the image.
[0048] Specifically, the output of the detection branch based on Anchor-Free contains three types of data, namely the heatmap, the center offset, and the box coordinates. Based on these three types of data, the detection branch based on Anchor-Free can regress the specific position of the tracking object (such as a pedestrian) in the image (i.e., the position of the target box), so as to accurately determine the position of the pedestrian in the image.
[0049] In some embodiments, the pedestrian re-identification branch adopts the Re-ID branch. The Re-ID branch contains 2 3×3 convolutional layers and 1 1×1 convolutional layer. The Re-ID branch outputs a low-dimensional feature vector with a size of (128, H, W).
[0050] Specifically, the pedestrian re-identification branch (Re-ID branch) is used to efficiently extract low-dimensional feature vectors. The feature vectors output in the Re-ID task usually have a high dimension, and training high-dimensional Re-ID features requires a large amount of training data, which is not available for the one-shot tracking algorithm. Previous two-step methods are less affected by this problem because they can utilize rich reid datasets that provide cropped human bodies. However, one-shot tracking algorithms cannot use them because they require original uncropped images. One solution is to reduce the dimensionality of the reid features to reduce their dependence on data. Therefore, using lower-dimensional feature vectors is more friendly to one-shot. Learning lower-dimensional feature vectors can reduce the risk of overfitting and improve the robustness of tracking.
[0051] The structure of the Re-ID branch provided in the embodiments of the present application will be described below in conjunction with the accompanying drawings. Figure 5 is a schematic structural diagram of the Re-ID branch provided in the embodiments of the present application. As Figure 5 shown, the Re-ID branch may specifically include:
[0052] The Re-ID branch contains one branch, which consists of 2 3×3 convolutional layers and 1 1×1 convolutional layer, and its output size is (128, H, W), where H is the height and W is the width. In practical applications, the output of the Re-ID branch is a low-dimensional feature vector.
[0053] According to the technical solution provided in the embodiments of the present application, the present application proposes a single-stage object tracking algorithm based on Anchor-Free, which can ensure the tracking speed and tracking accuracy; the present application improves on the basis of the traditional feature extraction module Backbone, and at the same time introduces DCN to extract more sufficient information, which is beneficial to the learning of the subsequent Re-ID branch; in addition, the present application designs a detection branch based on Anchor-Free and an efficient Re-ID branch in the structure of the single-stage object tracking model, achieving the effect of improving the tracking accuracy and the robustness of the tracking while ensuring the tracking speed.
[0054] The following is an embodiment of the device of the present application, which can be used to execute the method embodiment of the present application. For the details not disclosed in the embodiment of the device of the present application, please refer to the method embodiment of the present application.
[0055] Figure 6 is a schematic structural diagram of the object tracking device based on the single-stage object tracking model provided in the embodiments of the present application. As Figure 6 shown, the object tracking device based on the single-stage object tracking model includes:
[0056] An input module 601, configured to obtain video data containing a tracking object and input the video data into a predetermined single-stage object tracking model;
[0057] A processing module 602, configured to process each frame of image in the video data by using the feature extraction module in the single-stage object tracking model to obtain a feature map;
[0058] A regression module 603, configured to input the feature map into the detection branch in the single-stage object tracking model, and use the detection branch to regress the position of the tracking object in the image to obtain a target box corresponding to the tracking object, where the detection branch uses a detection module based on Anchor-Free;
[0059] An extraction module 604 is configured to input a feature map into a person re-identification branch in a single-stage object tracking model, and use the person re-identification branch to extract a low-dimensional feature vector of the feature map, obtaining a person re-identification low-dimensional feature vector;
[0060] A tracking module 605 is configured to track the trajectory generated by the tracking object in video data based on the target box corresponding to the tracking object and the person re-identification low-dimensional feature vector.
[0061] In some embodiments, Figure 6 A pre-training module 606 obtains pre-configured detection data, uses the MixUp augmentation algorithm to perform data augmentation on the detection data, obtaining new detection data, and uses the new detection data to train the detection branch in the single-stage object tracking model; after training the detection branch, uses the pre-configured tracking data to train the complete single-stage object tracking model, obtaining a trained single-stage object tracking model, where the tracking data includes target boxes and object identifiers corresponding to each target box.
[0062] In some embodiments, Figure 6 The pre-training module 606 of fuses any two original images in the detection data, so as to fuse the two original images into one image, obtaining a fused image, and superimposes the target boxes in the original images in the fused image, generating new detection data based on the original images and the fused image.
[0063] In some embodiments, a deformable convolutional network is inserted into the feature extraction module. The deformable convolutional network adopts a ResNet34 residual network. There are 2 3×3 convolutional layers in the stem stage of the ResNet34 residual network. The strides of the first two convolutional layers in the residual branch of the original residual module are exchanged, and the 1×1 convolutional layer with a stride of 2 in the original shortcut branch is replaced by an average pooling layer with a stride of 2 and a 1×1 convolutional layer.
[0064] In some embodiments, the detection branch includes a heatmap branch, a center offset branch, and a box branch. The heatmap branch, the center offset branch, and the box branch are all composed of 2 3×3 convolutional layers and 1 1×1 convolutional layer; wherein, the heatmap branch is used to output a heatmap with a size of (1, H, W), the center offset branch is used to output a center offset with a size of (2, H, W), and the box branch is used to output box coordinates with a size of (2, H, W).
[0065] In some embodiments, Figure 6 A regression module 603 of regresses the position of the tracking object in the image based on the heatmap, the center offset, and the box coordinates, obtaining a target box corresponding to the tracking object in the image, and determining the position of the target box in the image.
[0066] In some embodiments, the pedestrian re-identification branch adopts a Re-ID branch, which includes two 3×3 convolutional layers and one 1×1 convolutional layer, and outputs a low-dimensional feature vector of size (128, H, W).
[0067] It should be understood that the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0068] Figure 7 It is a schematic structural diagram of the electronic device 7 provided by the embodiment of the present application. As Figure 7 shown, the electronic device 7 of this embodiment includes: a processor 701, a memory 702, and a computer program 703 stored in the memory 702 and executable on the processor 701. When the processor 701 executes the computer program 703, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 701 executes the computer program 703, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0069] Exemplarily, the computer program 703 can be divided into one or more modules / units. One or more modules / units are stored in the memory 702 and executed by the processor 701 to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 703 in the electronic device 7.
[0070] The electronic device 7 can be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 7 may include, but is not limited to, the processor 701 and the memory 702. Those skilled in the art can understand that Figure 7 merely being examples of the electronic device 7 does not constitute a limitation to the electronic device 7. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.
[0071] The processor 701 may be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0072] The memory 702 may be an internal storage unit of the electronic device 7. For example, the hard disk or memory of the electronic device 7. The memory 702 may also be an external storage device of the electronic device 7. For example, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 7. Further, the memory 702 may also include both the internal storage unit and the external storage device of the electronic device 7. The memory 702 is used to store computer programs and other programs and data required by the electronic device. The memory 702 may also be used to temporarily store data that has been output or is to be output.
[0073] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0074] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0075] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0076] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer device and method can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the apparatus or unit can be in an electrical, mechanical or other form.
[0077] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0078] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0079] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0080] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A target tracking method based on a single-stage target tracking model, characterized in that, Including: Obtain video data including a tracking object, and input the video data into a predetermined single-stage object tracking model; Process each frame image in the video data by using a feature extraction module in the single-stage object tracking model to obtain a feature map; Input the feature map into a detection branch in the single-stage object tracking model, and use the detection branch to regress the position of the tracking object in the image to obtain a target box corresponding to the tracking object, where the detection branch adopts an Anchor-Free detection module; Input the feature map into a person re-identification branch in the single-stage object tracking model, and use the person re-identification branch to extract a low-dimensional feature vector of the feature map to obtain a person re-identification low-dimensional feature vector; Track the trajectory generated by the tracking object in the video data based on the target box corresponding to the tracking object and the person re-identification low-dimensional feature vector; Wherein, a deformable convolutional network is inserted in the feature extraction module, the deformable convolutional network adopts a ResNet34 residual network, 2 3×3 convolutional layers are provided in the stem stage of the ResNet34 residual network, the strides of the first two convolutional layers of the residual branch in the original residual module are exchanged, and the 1×1 convolutional layer with a stride of 2 in the original short-circuit branch is replaced by an average pooling layer with a stride of 2 and a 1×1 convolutional layer; The detection branch includes a heat map branch, a center offset branch and a box branch, and the heat map branch, the center offset branch and the box branch are all composed of 2 3×3 convolutional layers and 1 1×1 convolutional layer; wherein, the heat map branch is used to output a heat map with a size of (1, H, W), the center offset branch is used to output a center offset with a size of (2, H, W), and the box branch is used to output box coordinates with a size of (2, H, W); The using the detection branch to regress the position of the tracking object in the image to obtain a target box corresponding to the tracking object includes: regressing the position of the tracking object in the image based on the heat map, the center offset and the box coordinates to obtain a target box corresponding to the tracking object in the image, and determining the position of the target box in the image; The person re-identification branch adopts a Re-ID branch, the Re-ID branch includes 2 3×3 convolutional layers and 1 1×1 convolutional layer, and the Re-ID branch outputs a low-dimensional feature vector with a size of (128, H, W).
2. The method according to claim 1, wherein The pre-training process of the single-stage object tracking model includes: Obtain pre-configured detection data, perform data augmentation on the detection data by using the MixUp augmentation algorithm to obtain new detection data, and use the new detection data to train the detection branch in the single-stage object tracking model; After training the detection branch, use pre-configured tracking data to train the complete single-stage object tracking model to obtain a trained single-stage object tracking model, where the tracking data includes target boxes and object identifiers corresponding to each target box.
3. The method according to claim 2, wherein Performing data augmentation on the detection data by using the MixUp augmentation algorithm to obtain new detection data includes: Fusing any two original images in the detection data so as to fuse the two original images into one image to obtain a fused image, superimposing the target boxes in the original images in the fused image, and generating the new detection data based on the original images and the fused image.
4. An object tracking device based on a single-stage object tracking model, characterized in that, Including: An input module configured to acquire video data including a tracking object and input the video data into a predetermined single-stage object tracking model; A processing module configured to process each frame image in the video data by using a feature extraction module in the single-stage object tracking model to obtain a feature map; A regression module configured to input the feature map into a detection branch in the single-stage object tracking model, and use the detection branch to regress the position of the tracking object in the image to obtain a target box corresponding to the tracking object, where the detection branch adopts an Anchor-Free detection module; An extraction module configured to input the feature map into a person re-identification branch in the single-stage object tracking model, and use the person re-identification branch to extract a low-dimensional feature vector of the feature map to obtain a person re-identification low-dimensional feature vector; A tracking module configured to track the trajectory generated by the tracking object in the video data based on the target box corresponding to the tracking object and the person re-identification low-dimensional feature vector; Wherein, a deformable convolutional network is inserted into the feature extraction module, the deformable convolutional network adopts a ResNet34 residual network, 2 3×3 convolutional layers are provided in the stem stage of the ResNet34 residual network, the strides of the first two convolutional layers in the residual branch of the original residual module are exchanged, and the 1×1 convolutional layer with a stride of 2 in the original short-circuit branch is replaced with an average pooling layer with a stride of 2 and a 1×1 convolutional layer; The detection branch includes a heat map branch, a center offset branch, and a box branch, and the heat map branch, the center offset branch, and the box branch are all composed of 2 3×3 convolutional layers and 1 1×1 convolutional layer; wherein, the heat map branch is used to output a heat map with a size of (1, H, W), the center offset branch is used to output a center offset with a size of (2, H, W), and the box branch is used to output box coordinates with a size of (2, H, W); The regression module regresses the position of the tracking object in the image based on the heat map, the center offset, and the box coordinates to obtain a target box corresponding to the tracking object in the image, and determines the position of the target box in the image; The person re-identification branch adopts a Re-ID branch, the Re-ID branch includes 2 3×3 convolutional layers and 1 1×1 convolutional layer, and the Re-ID branch outputs a low-dimensional feature vector with a size of (128, H, W).
5. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 3 is implemented.
6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Target detection method, network, device, terminal equipment and storage medium
CN112633299A
Anchor-free multi-target tracking method for enhancing ID re-identification
CN113971688A