Single target tracking method, device, electronic device and storage medium

Through the two-dimensional dynamic convolution operation implemented in parallel using CUDA in TensorRT inference framework, the problems of slow detection speed and low accuracy in the prior art are solved, fast and accurate single-object tracking is achieved, and the inference speed and real-time performance are improved.

CN114155276BActive Publication Date: 2025-06-06NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111371170.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-18
Publication Date
2025-06-06
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

In the prior art, the model detection speed is slow and the time cost is high, and it cannot meet the needs of fast and accurate single-target tracking.

Method used

The branch network is detected by inputting the current frame image and the tracking coordinates of the target to be tracked, and using the TensorRT inference framework and CUDA in parallel to perform two-dimensional dynamic convolution operations, generating the target coordinate offset vector and confidence vector, and post-processing is performed to obtain the tracking coordinates.

Benefits of technology

It improves the speed and accuracy of target detection, meets the needs of fast and accurate single-target tracking, and ensures tracking accuracy while improving inference speed and real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155276B_ABST
    Figure CN114155276B_ABST
Patent Text Reader

Abstract

The present invention provides a single target tracking method, device, electronic device and storage medium, including: inputting the tracking coordinates of the current frame image and the target to be tracked in the previous frame image into the detection branch network, obtaining the coordinate offset features and classification features of the target search area under the TensorRT reasoning framework; using the convolution kernel template as the target tracking dynamic convolution kernel, using the coordinate offset features and classification features as the features to be convolved, performing a two-dimensional dynamic convolution operation based on CUDA parallel implementation, generating a target coordinate offset vector and a target confidence vector; performing post-processing on the target coordinate offset vector and the target confidence vector, and obtaining the tracking coordinates of the target to be tracked on the current frame image. The present invention provides a feasible method for applying the TensorRT reasoning framework to dynamically variable convolution kernels to accelerate end-to-end single target tracking, which ensures tracking accuracy while improving reasoning speed and real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision tracking technology, and in particular to a single target tracking method, device, electronic equipment and storage medium. Background Art

[0002] The most basic way to track the target in the video image is to locate and classify the target in each frame according to the target detection results. As an important branch of artificial intelligence, computer vision has become increasingly mature, especially the single target tracking algorithm based on deep convolutional neural networks (CNN) has made breakthrough progress in the field of computer vision, and the accuracy in many applications has gradually reached the standards of industrialization and productization, which is mainly due to its deep network structure and the ability to be pre-trained through a large amount of training data.

[0003] However, it is precisely because of these two characteristics of CNN that the deep network model has a large storage capacity and runs slowly during forward reasoning, resulting in the detection speed not keeping up, and the detection accuracy of the tracked target is also relatively low. This is one of the biggest obstacles in the industrialization and productization of deep learning results.

[0004] In view of this, there is an urgent need to further improve the inference speed and accuracy of existing target detection models. Summary of the invention

[0005] The present invention provides a single target tracking method, device, electronic device and storage medium, which are used to solve the defects of the prior art that the model detection speed is slow, the time cost is high, and the demand for fast and accurate single target tracking cannot be met.

[0006] In a first aspect, the present invention provides a single target tracking method, comprising:

[0007] Step 1: Input the current frame image and the tracking coordinates of the target to be tracked in the previous frame image into the detection branch network, and obtain the coordinate offset feature Freg and classification feature Fcls of the target search area output by the detection branch network under the TensorRT reasoning framework;

[0008] Step 2, using the convolution kernel template as the target tracking dynamic convolution kernel, using the coordinate offset feature Freg and the classification feature Fcls as the features to be convolved, performing a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation, and generating a target coordinate offset vector delta_offsets and a target confidence vector scores;

[0009] Step 3: perform post-processing on the target coordinate offset vector delta_offsets and the target confidence vector scores to obtain the tracking coordinates of the target to be tracked on the current frame image.

[0010] According to a single target tracking method provided by the present invention, before inputting the current frame image and the tracking coordinates of the target to be tracked in the previous frame image into the detection branch network, the method further includes:

[0011] Input any initial image and the known coordinates of the target to be tracked on the initial image into the template branch network, and obtain the coordinate offset convolution kernel template Kreg and the classification convolution kernel template Kcls output by the template branch network under the TenorRT reasoning framework;

[0012] The convolution kernel template is composed according to the coordinate offset convolution kernel template Kreg and the classification convolution kernel template Kcls.

[0013] According to a single target tracking method provided by the present invention, the template branch network and the detection branch network are both generated by converting the ONNX model saved after training the target tracking algorithm into a network structure model of the TensorRT framework;

[0014] The template branch network and the detection branch network are a pair of weight-shared twin network structure models.

[0015] According to a single target tracking method provided by the present invention, a convolution kernel template is used as a target tracking dynamic convolution kernel, the coordinate offset feature Freg and the classification feature Fcls are used as features to be convolved, a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation is performed, and a target coordinate offset vector delta_offsets and a target confidence vector scores are generated, including:

[0016] Using the coordinate offset feature Freg, the classification feature Fcls and the convolution kernel template, a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation is constructed;

[0017] The two-dimensional dynamic convolution operation cuConv includes the coordinate offset dynamic convolution operation and the classification dynamic convolution operation;

[0018] Under CUDA parallel implementation, the coordinate offset feature Freg and the classification feature Fcls are used as the features to be convolved, and the coordinate offset dynamic convolution operation and the classification dynamic convolution operation are performed to generate a plurality of target coordinate offset vectors delta_offsets and target confidence vectors scores.

[0019] According to a single target tracking method provided by the present invention, the coordinate offset feature Freg, the classification feature Fcls and the convolution kernel template are used to construct a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation, including:

[0020] Constructing the coordinate shift dynamic convolution operation based on the coordinate shift feature Freg and the first weight of the coordinate shift convolution kernel template Kreg on the convolution kernel template;

[0021] Constructing the classified dynamic convolution operation based on the classification feature Fcls and the second weight of the classification convolution kernel template Kcls on the convolution kernel template;

[0022] The first weight and the second weight are saved after training the target tracking algorithm.

[0023] According to a single target tracking method provided by the present invention, the post-processing of the target coordinate offset vector delta_offsets and the target confidence vector scores is performed to obtain the tracking coordinates of the target to be tracked on the current frame image, including:

[0024] Performing a modified dynamic convolution operation on the target coordinate offset vector delta_offsets to generate a corresponding tracking coordinate offset vector deltas;

[0025] Constructing a Hamming window and a scale change penalty corresponding to the size of the feature to be convolved;

[0026] According to the Hamming window and scale change penalty, determine the optimal tracking coordinate offset vector deltas corresponding to the maximum confidence among all confidence scores max ;

[0027] According to the optimal tracking coordinate offset vector deltas max and the tracking coordinates of the target to be tracked in the previous frame image, to determine the tracking coordinates of the target to be tracked in the current frame image.

[0028] According to a single target tracking method provided by the present invention, after obtaining the tracking coordinates of the target to be tracked on the current frame image, the method further includes:

[0029] Iterate step 1 to step 3 until a preset iteration stop condition is reached, and output the running trajectory of the target to be tracked.

[0030] In a second aspect, the present invention further provides a single target tracking device, comprising:

[0031] The feature extraction unit is mainly used to input the tracking coordinates of the current frame image and the target to be tracked in the previous frame image into the detection branch network, and obtain the coordinate offset feature Freg and all classification features Fcls of the target search area output by the detection branch network under the TensorRT reasoning framework;

[0032] The offset operation unit is mainly used to use the convolution kernel template as the target tracking dynamic convolution kernel, the coordinate offset feature Freg and the classification feature Fcls as the features to be convolved, perform a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation, and generate a target coordinate offset vector delta_offsets and a target confidence vector scores;

[0033] The coordinate output unit is mainly used to perform post-processing on the target coordinate offset vector delta_offsets and the target confidence vector scores to obtain the tracking coordinates of the target to be tracked on the current frame image.

[0034] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any of the single target tracking methods described above are implemented.

[0035] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the single target tracking methods described above.

[0036] The single target tracking method, device, electronic device and storage medium provided by the present invention provide a feasible method for applying the TensorRT reasoning framework to dynamically variable convolution kernels to accelerate the end-to-end single target tracking algorithm, which ensures the tracking accuracy while improving the reasoning speed and real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0038] Figure 1 It is a flow chart of the single target tracking method provided by the present invention;

[0039] Figure 2 It is a schematic diagram of a two-dimensional dynamic convolution operation cuConv operation based on CUDA acceleration provided by the present invention;

[0040] Figure 3 It is a flow chart of realizing target tracking based on the TensorRT end-to-end dynamic convolution kernel acceleration method provided by the present invention;

[0041] Figure 4 A time diagram of the forward reasoning of each frame image algorithm provided by the present invention;

[0042] Figure 5 A schematic diagram of the score of the forward reasoning algorithm for each frame of the image provided by the present invention;

[0043] Figure 6 It is a structural schematic diagram of a single target tracking device provided by the present invention;

[0044] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] It should be noted that in the description of the embodiments of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "include one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements. The orientation or position relationship indicated by the terms "upper", "lower" and the like is based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. Unless otherwise clearly specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, or it can be a connection between two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0047] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0048] TensorRT is a high-performance deep learning inference optimizer that can provide low-latency, high-throughput deployment inference for deep learning applications. However, TensorRT's deep neural network inference optimization engine can only accelerate the optimization of neural network structures with unchanged convolution kernel weights, that is, it can only accelerate neural network structures with only one branch to achieve end-to-end forward reasoning.

[0049] The parallel computing platform (Compute Unified Device Architecture, CUDA) launched by NVIDIA enables the graphics processing unit (Graphics Processing Unit, GPU) to have the ability to solve complex computing problems and can help users program the processor for different tasks.

[0050] The single target tracking method, device, electronic device and storage medium provided by the present invention combine TensorRT and GPU through CUDA, and can perform fast and efficient deployment reasoning in almost all network model frameworks.

[0051] Combine the following Figure 1-Figure 7 The single target tracking method, device, electronic device and storage medium provided by the embodiments of the present invention are described.

[0052] Figure 1 It is a flow chart of the single target tracking method provided by the present invention, such as Figure 1 As shown, including but not limited to the following steps:

[0053] Step 1: Input the current frame image and the tracking coordinates of the target to be tracked in the previous frame image into the detection branch network, and obtain the coordinate offset feature Freg and all classification features Fcls of the target search area output by the detection branch network under the TensorRT reasoning framework.

[0054] The current frame image is a frame image taken in real time by a drone, and the target to be tracked can be a target marked in advance by the user in the initial image, which is generally an active target or a stationary target. When the target to be tracked is an active target, the single target tracking method provided by the present invention can obtain the moving trajectory of the target to be tracked by continuously iteratively classifying and locating the target to be tracked in adjacent frame images, so as to achieve tracking thereof.

[0055] The previous frame image may be an image acquired in a sampling period before the current frame image when sampling is performed according to a preset sampling period, and the tracking coordinates of the target to be tracked are predetermined on the image. The present invention can realize the classification, identification and positioning of the target to be tracked in the current frame image through the tracking coordinates of the target to be tracked in the previous frame image.

[0056] Specifically, the current frame image and the tracking coordinates of the target to be tracked in the previous frame image are used as inputs of the detection branch network DNet, so that the detection branch network DNet extracts features of the current frame image under the TensorRT reasoning framework, and outputs the relevant position coordinate offset features Freg and classification features Fcls in the target search area where the target to be tracked is located.

[0057] Among them, the detection branch network DNet can adopt a general convolutional neural network model for realizing target detection, and the present invention does not make specific limitations on this.

[0058] Among them, the coordinate offset feature Freg is mainly used to reflect the position feature of the target to be tracked in the current frame image, and the classification feature Fclc is mainly used to reflect the classification recognition feature of the target to be tracked in the current frame image.

[0059] As an optional embodiment, step 1 may include but is not limited to the following steps:

[0060] Step 1.1, construct the detection branch network DNet related to the target tracking algorithm, which is a network model based on the TensorRT framework and converted from the Open Neural Network Exchange (ONNX) model, where the ONNX model is a trained target detection model.

[0061] Step 1.2: Input the latest acquired current frame image frame in the video image sequence and the tracking coordinates roi[x, y, w, h] of the target to be tracked obtained after recognizing the previous frame image into the detection branch network DNet to obtain the position coordinate offset feature Freg and classification feature Fcls of the target tracking area output by the detection branch network DNet. The specific process is as follows:

[0062] Freg,Fcls=DNet(frame,roi).

[0063] Step 2: Using the convolution kernel template as the target tracking dynamic convolution kernel, using the coordinate offset feature Freg and the classification feature Fcls as the features to be convolved, performing a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation, and generating a target coordinate offset vector delta_offsets and a target confidence vector scores.

[0064] Among them, the convolution kernel template is determined according to the initial image, and the convolution kernel template mainly includes the coordinate offset convolution kernel template Kreg and the classification convolution kernel template Kcls.

[0065] The present invention is based on CUDA accelerated two-dimensional dynamic convolution operation cuConv, realizes the output of target coordinate offset vector delta_offsets and target confidence vector scores, and realizes the positioning of the target to be tracked in the current frame image, specifically including but not limited to the following steps:

[0066] The convolution kernel template constructed by the coordinate offset convolution kernel template Kreg and the classification convolution kernel template Kcls is used as the target tracking dynamic convolution kernel, and each coordinate offset feature Freg and the classification feature Fcls are used as the convolution features to be performed in the two-dimensional dynamic convolution operation cuConv. The coordinate offset dynamic convolution operation based on CUDA parallelism is implemented to generate the corresponding target coordinate offset vector delta_offsets and target confidence vector scores. The specific process is as follows:

[0067] delta_offsets, scores=cuConv(Kreg,Kcls,Freg,Fcls).

[0068] Step 3: perform post-processing on the target coordinate offset vector delta_offsets and the target confidence vector scores to obtain the tracking coordinates of the target to be tracked on the current frame image.

[0069] First, according to the trained modified convolution kernel, a biased dynamic convolution operation can be performed on the target coordinate offset vector delta_offsets and its target confidence vector scores to generate the corresponding tracking coordinate offset vector deltas, so as to further improve the positioning accuracy of the target to be tracked in the current frame image.

[0070] Furthermore, since multiple target coordinate offset vectors delta_offsets with confidence scores are output in step 2, it is necessary to determine an optimal coordinate offset from all target coordinate offset vectors delta_offsets as the final positioning of the target to be tracked in the current frame image.

[0071] Furthermore, after the target coordinate offset vector delta_offsets is corrected and the corresponding tracking coordinate offset vector deltas is obtained, the tracking coordinate offset vector deltas and its corresponding confidence scores may be post-processed.

[0072] Therefore, the present invention performs post-processing on all tracking coordinate offset vectors deltas (or target coordinate offset vectors delta_offsets) and their corresponding confidence scores, so as to achieve effective tracking of the target to be tracked in the current frame image.

[0073] The single target tracking method provided by the present invention provides a feasible method for applying the TensorRT reasoning framework to dynamically variable convolution kernels to accelerate end-to-end single target tracking, which ensures tracking accuracy while improving reasoning speed and real-time performance.

[0074] Based on the content of the above embodiment, as an optional embodiment, before the current frame image and the tracking coordinates of the target to be tracked in the previous frame image are input into the detection branch network, the method further includes:

[0075] Input any initial image and the known coordinates of the target to be tracked on the initial image into the template branch network, and obtain the coordinate offset convolution kernel template Kreg and the classification convolution kernel template Kcls output by the template branch network under the TenorRT reasoning framework;

[0076] The convolution kernel template is composed according to the coordinate offset convolution kernel template Kreg and the classification convolution kernel template Kcls.

[0077] Among them, any initial image can be any frame initial image manually selected by the user from the video image sequence, and the target search area where the target to be selected is located is determined in the initial image by manual annotation, and the position of the target search area is used as the known coordinates.

[0078] As an optional embodiment, the above step of creating the convolution kernel template includes but is not limited to the following steps:

[0079] First, a template branch network TNet is constructed, where the template branch network TNet can be a convolutional neural network model consisting of 5 convolutional layers.

[0080] Then, the branch neural network ONNX model in the twin network model saved by Pytorch is converted to the TensorRT reasoning framework network model.

[0081] The initial image containing the pre-set target to be tracked and the known coordinates initRoi[x, y, w, h] of the target to be tracked on the initial image are input into the template branch network TNet, and the dynamically changeable coordinate offset convolution kernel template Kreg and classification convolution kernel template Kcls of the target to be tracked are output. Then, a dynamic convolution kernel template can be set for the two-dimensional dynamic convolution operation cuConv operation of the detection branch network DNet.

[0082] It should be noted that for any target to be tracked set by the user in the initial image (including the known coordinates of the target position), it is equivalent to determining the coordinate offset convolution kernel template Kreg and the classification convolution kernel template Kcls associated with a convolution kernel template, and in the subsequent entire target tracking process, the convolution kernel template remains unchanged.

[0083] The user can locate and track different targets to be tracked according to the tracking needs, that is, only the initial image needs to be re-determined, or the new target to be tracked and its known coordinates are re-determined in the initial image, and the subsequent tracking of different targets to be tracked can be completed.

[0084] Based on the content of the above embodiment, as an optional embodiment, the template branch network and the detection branch network are both generated by converting the ONNX model saved after training the target tracking algorithm into a network structure model of the TensorRT framework;

[0085] The template branch network and the detection branch network are a pair of weight-shared twin network structure models.

[0086] It should be noted that the template branch network TNet and the detection branch network DNet can be the same convolutional neural network model, or they can adopt neural network models with different structures. That is, when the twin network structure model is adopted, the two can share weights.

[0087] Optionally, both can also use pseudo-twin neural networks, for example, one is an LSTM model and the other is a CNN model, in which case the weights of the two will not be shared.

[0088] The present invention selects two branches of the twin network structure model as the template branch network TNet and the detection branch network DNet, which can effectively improve the training efficiency of the network model, and can improve the accuracy of target recognition and positioning performed by the subsequent detection branch network DNet according to the convolution kernel template determined by the template branch network TNet.

[0089] Based on the content of the above embodiment, as an optional embodiment, a convolution kernel template is used as a target tracking dynamic convolution kernel, the coordinate offset feature Freg and the classification feature Fcls are used as features to be convolved, and a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation is performed to generate a target coordinate offset vector delta_offsets and a target confidence vector scores, including:

[0090] Using the coordinate offset feature Freg, the classification feature Fcls and the convolution kernel template, a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation is constructed;

[0091] The two-dimensional dynamic convolution operation cuConv includes the coordinate offset dynamic convolution operation and the classification dynamic convolution operation;

[0092] Under CUDA parallel implementation, the coordinate offset feature Freg and the classification feature Fcls are used as the features to be convolved, and the coordinate offset dynamic convolution operation and the classification dynamic convolution operation are performed to generate a plurality of target coordinate offset vectors delta_offsets and target confidence vectors scores.

[0093] The specific calculation process of constructing the two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation by using the coordinate offset feature Freg, the classification feature Fcls and the convolution kernel template is as follows:

[0094] delta_offsets, scores=cuConv(Kreg,Kcls,Freg,Fcls).

[0095] Figure 2 is a schematic diagram of a two-dimensional dynamic convolution operation cuConv operation based on CUDA acceleration provided by the present invention, such as Figure 2 As shown, the present invention constructs a two-dimensional dynamic convolution operation cuConv based on CUDA acceleration, including:

[0096] Constructing the coordinate shift dynamic convolution operation based on the coordinate shift feature Freg and the first weight of the coordinate shift convolution kernel template Kreg on the convolution kernel template;

[0097] Constructing the classified dynamic convolution operation based on the classification feature Fcls and the second weight of the classification convolution kernel template Kcls on the convolution kernel template;

[0098] The first weight and the second weight are saved after training the target tracking algorithm.

[0099] Among them, the first weight, also known as the coordinate offset convolution kernel weight Wreg, and the second weight, also known as the classification convolution kernel weight Wcls, are the weights and biases of the corresponding layers determined and saved after training the target detection model.

[0100] The single target tracking method provided by the present invention adopts an end-to-end dynamic convolution kernel single target tracking algorithm implemented based on TensorRT and CUDA, which provides a feasible method for applying the TensorRT reasoning framework to the dynamic variable convolution kernel acceleration end-to-end single target tracking algorithm, and at the same time realizes the end-to-end acceleration of the multi-branch variable convolution kernel neural network with universal applicability, meeting the requirements of improving the efficiency and real-time performance of the target tracking algorithm.

[0101] Based on the content of the above embodiment, as an optional embodiment,

[0102] Post-processing the target coordinate offset vector delta_offsets and the target confidence vector scores is performed to obtain the tracking coordinates of the target to be tracked on the current frame image, including:

[0103] Performing a modified dynamic convolution operation on the target coordinate offset vector delta_offsets to generate a corresponding tracking coordinate offset vector deltas;

[0104] Constructing a Hamming window and a scale change penalty corresponding to the size of the feature to be convolved;

[0105] According to the Hamming window and scale change penalty, determine the optimal tracking coordinate offset vector deltas corresponding to the maximum confidence among all confidence scores max ;

[0106] According to the optimal tracking coordinate offset vector deltas max and the tracking coordinates of the target to be tracked in the previous frame image, to determine the tracking coordinates of the target to be tracked in the current frame image.

[0107] First, a modified dynamic convolution operation may be performed on each of the target coordinate offset vectors delta_offsets and the target confidence vector scores to generate a corresponding tracking coordinate offset vector deltas and determine the confidence vector scores of each tracking coordinate offset vector deltas. The specific calculation process is as follows:

[0108] deltas=cuConv(Wdelta,Bias,delta_offsets).

[0109] Furthermore, since the target coordinate offset vector delta_offsets with the confidence vector scores will be output in step 2, it is necessary to determine an optimal coordinate offset from all tracking coordinate offset vectors deltas as the final positioning of the target to be tracked in the current frame image. Therefore, the post-processing method based on the combination of Hamming window and scale change penalty in the present invention can further post-process all tracking coordinate offset deltas and their corresponding confidence scores to achieve tracking of the target to be tracked in the current frame image, specifically including but not limited to the following steps:

[0110] First, construct the Hamming window hanning and scale change penalty penalty corresponding to the feature map size;

[0111] Sort all confidence scores to get the maximum confidence scores max ;

[0112] The tracking coordinate offset vector deltas and the maximum confidence scores max A tracking coordinate offset vector with the same score index is the target coordinate offset value and the optimal tracking coordinate offset vector deltas max ;

[0113] According to the tracking coordinates of the target to be tracked in the previous frame image, the optimal tracking coordinate offset vector deltas max By performing coordinate transformation, the tracking coordinates of the target to be tracked on the current frame image can be obtained. The expression of coordinate transformation can be expressed as:

[0114] score=max(score[i]*penalty[i]*hanning[i]);

[0115] roi.x=anchor.x+deltas.x*anchor.w;

[0116] roi.y=anchor.y+deltas.y*anchor.h;

[0117] roi.w=anchor.w*exp(delta.w);

[0118] roi.h=anchor.h*exp(delta.h);

[0119]

[0120] 0≤i≤j,j∈{sizea,sizeb},PI=3.14159;

[0121] Among them, anchor[x,y,w,h] is the tracking coordinate of the feature map corresponding to the preset target in the previous frame image; roi[x,y,w,h] is the tracking coordinate obtained after coordinate transformation; deltas[x,y,w,h] is the optimal tracking coordinate offset vector; {sizea,sizeb} is the feature map size.

[0122] The single target tracking method provided by the present invention uses a Hamming window and a scale change penalty to optimize the tracking coordinate offset vector, which can further improve the accuracy of target tracking.

[0123] Based on the content of the above embodiment, as an optional embodiment, after obtaining the tracking coordinates of the target to be tracked on the current frame image, the method further includes:

[0124] Iterate step 1 to step 3 until a preset iteration stop condition is reached, and output the running trajectory of the target to be tracked.

[0125] Figure 3 This is a flow chart of the target tracking method based on the TensorRT end-to-end dynamic convolution kernel acceleration method provided by the present invention, such as Figure 3 As shown, the overall operation includes but is not limited to the following steps:

[0126] The first step is to create a convolution kernel template. Based on the initialization of the trained tracker or target detection model template convolution weights, a template branch network TNet is constructed. In the present invention, the branch neural network onnx model in the twin network template saved by Pytorch is converted into a TensorRT reasoning framework network model, where the template branch network TNet can contain 5 convolution layers.

[0127] The initial image frame containing the target to be tracked in the video sequence and the known coordinates initRoi[x, y, w, h] of the target to be tracked on the initial image frame are input into the template branch network TNet, and the dynamically changeable coordinate offset convolution kernel template Kreg and classification convolution kernel template Kcls of the target to be tracked are output to set the convolution kernel template for the two-dimensional dynamic convolution operation cuConv of the detection branch model DNet.

[0128] The second step is real-time image detection. Construct a detection branch network DNet. The present invention converts the branch neural network ONNX model of the twin network template saved by Pytorch into a TensorRT reasoning framework network model, wherein the detection branch network DNet can also be composed of 5 convolutional layers, which is a twin network with the template branch network TNet, and the weights of the two are shared.

[0129] The current frame image frame in the video image sequence and the tracking coordinates roi′[x, y, w, h] of the target to be tracked in the previous frame image are input into the detection branch network DNet to realize target detection, and the coordinate offset feature Freg and classification feature Fcls related to the target tracking area position are output.

[0130] Furthermore, a two-dimensional dynamic convolution operation cuConv based on CUDA acceleration is constructed, that is, Kreg and Kcls are used as targets to track the dynamic convolution kernel, and the position coordinate offset feature Freg and the classification feature Fcls are used as the features to be convolved in the two-dimensional dynamic convolution operation cuConv to realize dynamic variable convolution under CUDA acceleration, which specifically includes two aspects:

[0131] Constructing a coordinate shift dynamic convolution operation based on the coordinate shift feature Freg and the coordinate shift convolution kernel weight of the coordinate shift convolution kernel template Kreg on the convolution kernel template;

[0132] Based on the classification feature Fcls and the classification convolution kernel weight of the classification convolution kernel template Kcls on the convolution kernel template, a classification dynamic convolution operation is constructed.

[0133] Furthermore, the convolution kernel template is used as the target tracking dynamic convolution kernel, the coordinate offset feature Freg and the classification feature Fcls are used as the features to be convolved, and a two-dimensional dynamic convolution operation is performed under the TensorRT reasoning framework to generate multiple target coordinate offsets delta_offsets and corresponding confidence vectors scores to realize the convolution of the coordinate offset feature.

[0134] Furthermore, a modified dynamic convolution operation can be performed on each of the target coordinate offset vectors delta_offsets and target confidence vectors scores to generate a corresponding tracking coordinate offset vector deltas to achieve convolution of classification features. The convolution kernel weights Wcls and bias Bias of the classification dynamic convolution operation in this step are the weights and biases of the corresponding layers saved after the training algorithm.

[0135] Finally, post-processing of all output tracking coordinate offset vector deltas and related confidence scores is performed to obtain the optimal tracking coordinate offset vector deltas max Then, according to the optimal tracking coordinate offset vector deltas max and the tracking coordinates of the target to be tracked in the previous frame image, to determine the tracking coordinates of the target to be tracked in the current frame image.

[0136] Iterate all the steps in the above step 2, identify and locate the target to be tracked for each current image frame in the video image sequence in turn, and fit the coordinate values ​​in each frame of the image into the running trajectory of the target to be tracked in the video image sequence, complete the generation of the historical running trajectory of the target to be tracked, and at the same time, locate the position coordinates of the target to be tracked in the newly acquired image in real time until the rapid tracking of the target is completed.

[0137] Among them, the preset iteration stop condition can be defined according to the actual detection needs. For example, if the user inputs the initial image or re-specifies a new target to be tracked in the initial image, then after re-executing the above steps one and two, step two is iteratively executed again until the detection of the entire video image sequence is achieved and the tracking task is completed.

[0138] In order to further illustrate the feasibility and effectiveness of the single target tracking method provided by the present invention, the UAV123 unmanned aerial vehicle single target vehicle tracking dataset (https: / / ivul.kaust.edu.sa / Pages / Dataset-UAV123.aspx) is selected for verification, and the car6 data subset is selected for verification (the car6 data subset has a total of 4861 frames of images).

[0139] In this experiment, the template branch network TNet is executed regularly to update the position coordinate offset convolution kernel template Kreg and the classification convolution kernel template Kcls of the target to be tracked, and the detection branch network DNet is executed for each frame image and the subsequent processing and tracking target is performed.

[0140] Figure 4 The time diagram of the forward reasoning of each frame image algorithm provided by the present invention is as follows: Figure 4As shown, it has been verified that the single target tracking method provided by the present invention can stably track the target vehicle, and the average execution time of the algorithm after acceleration is 0.00643ms, reaching 160FPS.

[0141] Figure 5 The schematic diagram of the score of the forward reasoning algorithm for each frame of the image provided by the present invention is as follows: Figure 5 As shown, the tracking confidence of the single target tracking method provided by the present invention is between 0.44 and 0.99.

[0142] In summary, the present invention provides a feasible method for applying the TensorRT reasoning framework to dynamically variable convolution kernels to accelerate the end-to-end single target tracking algorithm, which meets the requirements of improving the efficiency and real-time performance of the target tracking algorithm.

[0143] Figure 6 is a schematic diagram of the structure of the single target tracking device provided by the present invention, such as Figure 6 As shown, it mainly includes a feature extraction unit 61, an offset operation unit 62, a classification operation unit 63 and a coordinate output unit 64, wherein:

[0144] The feature extraction unit 61 is mainly used to input the tracking coordinates of the current frame image and the target to be tracked in the previous frame image into the detection branch network, and obtain the coordinate offset feature Freg and the classification feature Fcls of the target search area output by the detection branch network under the TensorRT reasoning framework;

[0145] The offset operation unit 62 is mainly used to use the convolution kernel template as the target tracking dynamic convolution kernel, use the coordinate offset feature Freg and the classification feature Fcls as the features to be convolved, perform a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation, and generate a target coordinate offset vector delta_offsets and a target confidence vector scores;

[0146] The coordinate output unit 64 is mainly used to perform post-processing on the target coordinate offset vector delta_offsets and the target confidence vector scores to obtain the tracking coordinates of the target to be tracked on the current frame image.

[0147] It should be noted that the single target tracking device provided in the embodiment of the present invention can execute the single target tracking method described in any of the above embodiments during specific operation, which will not be described in detail in this embodiment.

[0148] The single target tracking device provided by the present invention provides a feasible method for applying the TensorRT reasoning framework to dynamically variable convolution kernels to accelerate end-to-end single target tracking, which ensures tracking accuracy while improving reasoning speed and real-time performance.

[0149] Figure 7 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 7 As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730 and a communication bus 740, wherein the processor 710, the communication interface 720 and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logic instructions in the memory 730 to execute the single target tracking method, which includes the following steps:

[0150] Step 1: Input the current frame image and the tracking coordinates of the target to be tracked in the previous frame image into the detection branch network, and obtain the coordinate offset feature Freg and classification feature Fcls of the target search area output by the detection branch network under the TensorRT reasoning framework;

[0151] Step 2, using the convolution kernel template as the target tracking dynamic convolution kernel, using the coordinate offset feature Freg and the classification feature Fcls as the features to be convolved, performing a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation, and generating a target coordinate offset vector delta_offsets and a target confidence vector scores;

[0152] Step 3: perform post-processing on the target coordinate offset vector delta_offsets and the target confidence vector scores to obtain the tracking coordinates of the target to be tracked on the current frame image.

[0153] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0154] On the other hand, the present invention further provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, when the program instructions are executed by a computer, the computer can execute the single target tracking method provided by the above methods, the method comprising the following steps:

[0155] Step 1: Input the current frame image and the tracking coordinates of the target to be tracked in the previous frame image into the detection branch network, and obtain the coordinate offset feature Freg and classification feature Fcls of the target search area output by the detection branch network under the TensorRT reasoning framework;

[0156] Step 2, using the convolution kernel template as the target tracking dynamic convolution kernel, using the coordinate offset feature Freg and the classification feature Fcls as the features to be convolved, performing a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation, and generating a target coordinate offset vector delta_offsets and a target confidence vector scores;

[0157] Step 3: perform post-processing on the target coordinate offset vector delta_offsets and the target confidence vector scores to obtain the tracking coordinates of the target to be tracked on the current frame image.

[0158] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the single target tracking method provided in the above embodiments is implemented, and the method includes the following steps:

[0159] Step 1: Input the current frame image and the tracking coordinates of the target to be tracked in the previous frame image into the detection branch network, and obtain the coordinate offset feature Freg and classification feature Fcls of the target search area output by the detection branch network under the TensorRT reasoning framework;

[0160] Step 2, using the convolution kernel template as the target tracking dynamic convolution kernel, using the coordinate offset feature Freg and the classification feature Fcls as the features to be convolved, performing a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation, and generating a target coordinate offset vector delta_offsets and a target confidence vector scores;

[0161] Step 3: perform post-processing on the target coordinate offset vector delta_offsets and the target confidence vector scores to obtain the tracking coordinates of the target to be tracked on the current frame image.

[0162] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0163] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A single target tracking method, It is characterized in that include: Step 1: Input the current frame image and the tracking coordinates of the target to be tracked in the previous frame image into the detection branch network, and obtain the coordinate offset feature Freg and classification feature Fcls of the target search area output by the detection branch network under the TensorRT reasoning framework; Step 2, using the convolution kernel template as the target tracking dynamic convolution kernel, using the coordinate offset feature Freg and the classification feature Fcls as the features to be convolved, performing a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation, and generating a target coordinate offset vector delta_offsets and a target confidence vector scores; Step 3: perform post-processing on the target coordinate offset vector delta_offsets and the target confidence vector scores to obtain the tracking coordinates of the target to be tracked on the current frame image.

2. The single target tracking method according to claim 1, It is characterized in that Before inputting the current frame image and the tracking coordinates of the target to be tracked in the previous frame image into the detection branch network, it also includes: Input any initial image and the known coordinates of the target to be tracked on the initial image into the template branch network, and obtain the coordinate offset convolution kernel template Kreg and the classification convolution kernel template Kcls output by the template branch network under the TenorRT reasoning framework; The convolution kernel template is composed according to the coordinate offset convolution kernel template Kreg and the classification convolution kernel template Kcls.

3. The single target tracking method according to claim 2, It is characterized in that The template branch network and the detection branch network are both generated by converting the ONNX model saved after training the target tracking algorithm into a network structure model of the TensorRT framework; The template branch network and the detection branch network are a pair of weight-shared twin network structure models.

4. The single target tracking method according to claim 3, It is characterized in that The convolution kernel template is used as the target tracking dynamic convolution kernel, the coordinate offset feature Freg and the classification feature Fcls are used as the features to be convolved, and a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation is performed to generate a target coordinate offset vector delta_offsets and a target confidence vector scores, including: Using the coordinate offset feature Freg, the classification feature Fcls and the convolution kernel template, a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation is constructed; The two-dimensional dynamic convolution operation cuConv includes the coordinate offset dynamic convolution operation and the classification dynamic convolution operation; Under CUDA parallel implementation, the coordinate offset feature Freg and the classification feature Fcls are used as the features to be convolved, and the coordinate offset dynamic convolution operation and the classification dynamic convolution operation are performed to generate a plurality of target coordinate offset vectors delta_offsets and target confidence vectors scores.

5. The single target tracking method according to claim 4, It is characterized in that The method of using the coordinate offset feature Freg, the classification feature Fcls and the convolution kernel template to construct a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation includes: Constructing the coordinate shift dynamic convolution operation based on the coordinate shift feature Freg and the first weight of the coordinate shift convolution kernel template Kreg on the convolution kernel template; Constructing the classified dynamic convolution operation based on the classification feature Fcls and the second weight of the classification convolution kernel template Kcls on the convolution kernel template; The first weight and the second weight are saved after training the target tracking algorithm.

6. The single target tracking method according to claim 1, It is characterized in that The performing post-processing of the target coordinate offset vector delta_offsets and the target confidence vector scores to obtain the tracking coordinates of the target to be tracked on the current frame image includes: Performing a modified dynamic convolution operation on the target coordinate offset vector delta_offsets to generate a corresponding tracking coordinate offset vector deltas; Constructing a Hamming window and a scale change penalty corresponding to the size of the feature to be convolved; According to the Hamming window and scale change penalty, determine the optimal tracking coordinate offset vector deltas corresponding to the maximum confidence among all confidence scores max ; According to the optimal tracking coordinate offset vector deltas max and the tracking coordinates of the target to be tracked in the previous frame image, to determine the tracking coordinates of the target to be tracked in the current frame image.

7. The single target tracking method according to claim 1, It is characterized in that After obtaining the tracking coordinates of the target to be tracked on the current frame image, the method further includes: Iterate step 1 to step 3 until a preset iteration stop condition is reached, and output the running trajectory of the target to be tracked.

8. A single target tracking device, It is characterized in that include: A feature extraction unit is used to input the tracking coordinates of the current frame image and the target to be tracked in the previous frame image into the detection branch network, and obtain the coordinate offset feature Freg and the classification feature Fcls of the target search area output by the detection branch network under the TensorRT reasoning framework; An offset operation unit is used to use the convolution kernel template as a target tracking dynamic convolution kernel, use the coordinate offset feature Freg and the classification feature Fcls as features to be convolved, perform a two-dimensional dynamic convolution operation cuConv based on CUDA parallel implementation, and generate a target coordinate offset vector delta_offsets and a target confidence vector scores; The coordinate output unit is used to perform post-processing on the target coordinate offset vector delta_offsets and the target confidence vector scores to obtain the tracking coordinates of the target to be tracked on the current frame image.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the steps of the single target tracking method according to any one of claims 1 to 7 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the single target tracking method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Visual multi-target tracking method and device based on deep learning

    CN111161311A

  • Unmanned aerial vehicle non-cooperative target tracking system based on binocular vision

    CN113467500A