A lightweight object recognition and tracking system and method based on Transformer
By splitting and transforming models on edge devices and combining multi-input convolution plug-in, the deployment problem of Transformer target tracking algorithm on edge devices is solved, efficient target tracking and detection is achieved, and real-time requirements are met.
Patent Information
- Application Number
- CN202310922594.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-07-26
AI Technical Summary
The existing Transformer-based target tracking algorithm is difficult to deploy on edge embedded devices, mainly due to the computing volume and resource limitations, and it is impossible to deal with challenges in complex scenarios in real time, such as occlusion and rotation.
采用多进程设计和指控平台,通过拆分模型为多个部分在ONNX和TensorRT框架下转换,结合多输入卷积插件,实现模型在边缘设备上的加速和部署。
It improves real-time and performance on edge devices, realizes efficient tracking and detection of targets on edge devices, and meets real-time requirements.
Smart Images

Figure CN116862951B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target recognition and tracking, and particularly relates to a lightweight target recognition and tracking system and method based on Transformer. Background Art
[0002] Target tracking is widely applied in actual scenarios and is usually used as a component in large visual systems, having important applications in fields such as unmanned driving, scene perception detection, intelligent transportation, and robots. With the continuous upgrading of hardware, deep learning methods based on massive data have been increasingly studied and applied. Relying on complex model designs and big data training, target tracking algorithms based on convolutional neural networks have great advantages in tracking accuracy compared with traditional tracking algorithms. However, due to the huge amount of computation, it is still difficult to run them in real time on edge devices. Similarly, challenges in target tracking such as target occlusion, rotation and deformation of the target, background clutter, interference from similar targets, and background changes still exist, and it is extremely easy to cause the loss of the target.
[0003] The prior art FEAR (FEAR: Fast, Efficient, Accurate and Robust Visual Tracker) proposed a fast, efficient, accurate, and robust tracker, and proposed two lightweight models, a dual-template module and a pixel-wise fusion block. The former integrates temporal information using a learnable parameter, while the latter encodes more discriminative features with fewer parameters. The prior art LightTrack (LightTrack: Finding Lightweight Neural Networks for Object Tracking via One-Shot Architecture Search) uses neural architecture search (NAS) to design a more lightweight and efficient target tracker. Its paper experiments show that compared with manually designed SOTA trackers (such as SiamRPN++ and Ocean), LightTrack can achieve superior performance while requiring much less computation and fewer parameters. However, the above two lightweight models still lack robustness to challenges such as similar targets and occlusion in real scenarios and are difficult to be used in real scenarios with great challenges.
[0004] Two direct methods to address complexity and efficiency issues are model compression and compact model design. Existing compression techniques, such as pruning and quantization, can reduce model complexity, but they inevitably lead to a significant performance degradation due to information loss. On the other hand, manually designing new compact and efficient models requires high engineering costs and heavily relies on human expertise and experience. The object tracking network obtained through neural network search greatly reduces the model parameters for feature extraction in the algorithm and improves the computing speed, but it is still difficult to handle various challenges in actual scenarios and difficult to deploy on edge computing devices. Generally speaking, the existing tracking algorithms have the following disadvantages when actually used: the tracking models are becoming more and more complex, the training costs are increasing, more computing resources and training data are required; the state-of-the-art trackers require large model parameters and computing amounts and are difficult to deploy on edge devices; the tracking algorithms based on Tranformer perform better in issues such as long-range dependencies and occlusions, but their computing amounts and training costs are also high and difficult to deploy in actual applications; the current model compression and compact model design methods can reduce model complexity, but they will cause performance degradation or require higher engineering costs, and the trackers obtained through neural network search are still difficult to handle various challenges in actual applications and difficult to deploy on edge computing devices. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a method and system for real-time visual detection and tracking at the embedded end of an edge computing device, which is used to solve the problem that the object tracking algorithm based on deep learning is difficult to be deployed on edge embedded devices in real-world scenarios to cope with the actual challenges in real-world scenarios. These challenges include limited computing resources, power consumption limitations, real-time requirements, etc.
[0006] The first aspect of the present invention discloses a method for lightweight object tracking based on Transformer, also known as a method for real-time tracking at the embedded end of an edge computing device. The method includes: designing tracking solutions for onnx and TensorRT based on ET.Track to meet the deployment requirements of different tracking scenarios; selecting different tracking models based on the position of the detected object and the acceleration flag to track the object to be tracked, and obtaining the tracking information and detection results of the object to be tracked; encapsulating the detection results and the tracking information to obtain the encapsulated tracking information, and sending the encapsulated tracking information in the form of a tracking topic for the unmanned vehicle chassis control node to obtain the detection results and the tracking information.
[0007] The method according to the first aspect of the present invention for tracking and detecting a target to be tracked based on position and an acceleration identifier includes: if the acceleration identifier indicates that no acceleration processing is required, tracking and detecting the target to be tracked based on a first model to obtain tracking information and a detection result of the target to be tracked; or if the acceleration identifier indicates that acceleration processing is required, based on the acceleration type indicated by the acceleration identifier, performing acceleration processing based on a second model or a third model to obtain the tracking information and the detection result of the target to be tracked.
[0008] The first model of the method according to the first aspect of the present invention includes a first input end, a second input end, a backbone network module, an MCBN module, a first CCBR module, a second CCBR module, a third CCBR module, and a fourth CCBR module. Tracking and detecting the target to be tracked based on the first model includes: receiving a reference frame image of the target to be tracked through the first input end, and receiving a video frame to be detected in the video stream through the second input end; inputting the reference frame image and the video frame to be detected into the backbone network module, wherein the backbone network module extracts a first feature of the reference frame image and a second feature of the video frame to be detected; inputting the first feature and the second feature into the MCBN module, wherein the MCBN module performs pixel-level fusion processing on the first feature to obtain a third feature, performs pixel-level fusion processing on the second feature to obtain a fourth feature, subjects the third feature to convolution processing, batch normalization, and RELU activation function processing to obtain a fifth feature, and subjects the fourth feature to convolution processing, batch normalization, and RELU activation function processing to obtain a sixth feature; inputting the fifth feature into the first CCBR module and inputting the sixth feature into the third CCBR module, wherein the first CCBR module performs pixel-level correlation fusion processing on the fifth feature to obtain a seventh feature, and the third CCBR module performs pixel-level correlation fusion processing on the sixth feature to obtain an eighth feature; inputting the seventh feature into the second CCBR module to obtain a classification score of the target to be tracked, inputting the eighth feature into the fourth CCBR module to obtain coordinates of the target to be tracked, and performing matching and screening on the classification score of the target to be tracked and the coordinates of the target to be tracked to obtain the tracking information and the detection result of the target to be tracked.
[0009] According to the method of the first aspect of the present invention, the second model includes a first reference frame inference branch, a first search frame inference branch, and a first classification and regression inference branch. Accelerating the processing based on the second model to obtain the tracking information and detection result of the target to be tracked includes: receiving the reference frame of the target to be tracked through the first reference frame inference branch, extracting and normalizing the reference frame to obtain a normalized reference frame feature; receiving the subsequent image frame of the reference frame in the video stream through the first search frame inference branch, taking the subsequent image frame as a search frame, extracting and normalizing the search frame to obtain a search frame depth feature, wherein the first search frame inference branch shares the weight parameters for extraction and normalization with the first reference frame inference branch; receiving the normalized reference frame feature and the search frame depth feature through the first classification and regression inference branch, and performing pixel-level fusion processing on the normalized reference frame feature and the search frame depth feature respectively to obtain the classification score of the target to be tracked and the coordinates of the target to be tracked, and performing matching and screening on the classification score of the target to be tracked and the coordinates of the target to be tracked to obtain the tracking information and detection result of the target to be tracked.
[0010] According to the method of the first aspect of the present invention, the third model includes a second reference frame inference branch, a second search frame inference branch, and a second classification and regression inference branch. Among them, the reference frame inference branch includes a fourth backbone network module and a batch normalization module. The second search frame inference branch includes a fifth backbone network module and a batch normalization module. The second classification and regression inference branch includes a pixel-level feature fusion module, an Exemplar attention module, and a CCBR module. The CCBR module includes a ninth CCBR module, a tenth CCBR module, and an eleventh CCBR module. The fifth backbone network module shares weight parameters with the fourth backbone network module. Accelerating processing based on the third model to obtain the tracking information and detection results of the target to be tracked includes: the fourth backbone network module receives the reference frame of the target to be tracked and extracts the depth features of the target to be tracked in the reference frame; the batch normalization module normalizes the depth features of the target to be tracked to obtain the normalized reference frame depth features of the target to be tracked; the fifth backbone network module receives the subsequent image frame of the reference frame and extracts the depth features of the target to be tracked in the subsequent image frame; the batch normalization module performs batch normalization on the depth features of the target to be tracked in the subsequent image frame to obtain the normalized search frame depth features; the pixel-level feature fusion module processes the normalized reference frame depth features and the normalized search frame depth features to obtain the normalized reference frame enhanced depth features and the normalized search frame enhanced depth features, and inputs the normalized reference frame enhanced depth features and the normalized search frame enhanced depth features into the ninth CCBR module and the eleventh CCBR module. Among them, the ninth CCBR module performs two convolutions, one batch normalization, and one ReLU activation processing on the normalized reference frame enhanced depth features, and inputs the processed features into the Exemplar attention module to obtain the tenth feature; the eleventh CCBR module performs two convolutions, one batch normalization, and one ReLU activation processing on the normalized search frame enhanced depth features, and inputs the processed features into the Exemplar attention module to obtain the eleventh feature; input the tenth feature into the tenth CCBR module to obtain the classification score of the target to be tracked, input the eleventh feature into the twelfth CCBR module to obtain the coordinates of the target to be tracked, and perform matching and screening on the classification score of the target to be tracked and the coordinates of the target to be tracked to obtain the tracking information and detection results of the target to be tracked.
[0011] The second aspect of the present invention discloses a lightweight object recognition and tracking system based on Transformer, also known as a real-time visual detection and tracking system for the embedded end of edge computing devices. The system includes: an acquisition unit configured to acquire a video stream, obtain a target to be tracked from the video stream, and generate a task instruction for tracking the target to be tracked; an analysis unit configured to analyze the task instruction, where the task instruction includes the position of the target to be tracked and an acceleration flag, and the acceleration flag is used to indicate whether the tracking algorithm needs to be accelerated; a tracking and detection unit configured to track and detect the target to be tracked based on the position and the acceleration flag to obtain the tracking information and detection result of the target to be tracked; and a packaging unit configured to package the detection result and the tracking information to obtain the packaged tracking information, and send the packaged tracking information in the form of a tracking topic for the unmanned vehicle chassis control node to obtain the detection result and the tracking information.
[0012] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. When the processor executes the computer program stored in the memory, the steps in the method for real-time visual detection and tracking of the embedded end of the edge computing device in any one of the first aspects of the present disclosure are implemented.
[0013] The fourth aspect of the present invention discloses a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps in the method for real-time visual detection and tracking of the embedded end of an edge computing device in any one of the first aspects of the present invention are implemented.
[0014] In summary, the solution proposed by the present invention has the following technical effects: By adopting a multi-process design and introducing a command and control platform, the generation of task instructions and the selection of target tracking are realized. Through the command and control platform, the target to be tracked can be specified according to the detected target information, and the information on whether to accelerate can be obtained according to the task instruction, and the acceleration method can be selected to accelerate the inference of the tracking algorithm, thereby improving the real-time performance and performance of the system. This method can flexibly generate task instructions and select target tracking according to actual needs, improving the intelligence and flexibility of the system.
[0015] The present invention splits the tracking algorithm model that cannot be directly converted according to the splitting method of this patent to convert it into a model that can be converted under ONNX and TensorRT, and performs inference acceleration and deployment through the corresponding framework. To solve the abnormal conversion of the algorithm model under TensorRT, this patent designs a multi-input convolution plugin, packages it into a dynamic library using C++ under ubuntu, calls the dynamic library under torch2trt, and designs a converter to enable the model to be normally converted and the confidence after inference to be normal. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 FIG. is a flowchart of a method for real-time visual detection and tracking of an embedded end of an edge computing device according to an embodiment of the present invention;
[0018] Figure 2 FIG. is a schematic structural diagram of a method for real-time visual detection and tracking of an embedded end of an edge computing device according to an embodiment of the present invention;
[0019] Figure 3 FIG. is a schematic diagram of a first model architecture according to an embodiment of the present invention;
[0020] Figure 4 FIG. is a schematic diagram of an Exemplar module according to an embodiment of the present invention;
[0021] Figures 5(A)-5(C) FIG. is a schematic diagram of each branch of a second model according to an embodiment of the present invention;
[0022] Figure 6 FIG. is a schematic diagram of a third module according to an embodiment of the present invention;
[0023] Figures 7(A)-7(B) FIG. is an effect diagram of a method for real-time visual detection and tracking of an embedded end of an edge computing device according to an embodiment of the present invention;
[0024] Figure 8 FIG. is a structural diagram of a system for real-time visual detection and tracking of an embedded end of an edge computing device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0026] The first aspect of the present invention discloses a lightweight object recognition and tracking method based on Transformer, also known as a real-time visual detection and tracking method for the embedded end of edge computing devices. As Figures 1-2 shown, the method includes:
[0027] Step S1: Obtain a video stream, obtain the target to be tracked from the video stream, and generate a task instruction for tracking the target;
[0028] Step S2: Parse the task instruction, obtain the position of the target to be tracked and an acceleration flag, where the acceleration flag is used to indicate whether the tracking algorithm needs to be accelerated;
[0029] Step S3: If the acceleration flag indicates that no acceleration is required, perform tracking and detection on the target to be tracked based on the first model to obtain the tracking information and detection result of the target to be tracked; otherwise, perform acceleration processing based on the acceleration type indicated by the acceleration flag, based on the second model or the third model, to obtain the tracking information and detection result of the target to be tracked; encapsulate the detection result and the tracking information to obtain the encapsulated tracking information, and send the encapsulated tracking information in the form of a tracking topic for the unmanned vehicle chassis control node to obtain the detection and tracking results.
[0030] Step S4: Encapsulate the detection result and the tracking information to obtain the encapsulated tracking information, and send the encapsulated tracking information in the form of a tracking topic for the unmanned vehicle chassis control node to obtain the detection result and the tracking information.
[0031] Step S1: Obtain a video stream, obtain the target to be tracked from the video stream, and generate a task instruction for tracking the target, including:
[0032] In this embodiment, the video stream is obtained by using the topic subscription or rtsp protocol. The video stream first enters the detection algorithm module, and the detection algorithm module outputs the information of the detected target and sends the information of the target to the command and control platform. After receiving the information of the target, the command and control platform generates a task instruction, and the task instruction indicates whether the target is selected, and the selected target is determined according to the information of the detected target. Then, the task instruction is parsed by the parsing module to obtain the target to be tracked.
[0033] Step S2: Parse the task instruction, obtain the position of the target to be tracked and an acceleration flag, where the acceleration flag is used to indicate whether the tracking algorithm needs to be accelerated. In this embodiment, it is determined whether the target is selected according to the task instruction. If no target is selected, the command and control platform is accessed regularly. If a target is selected, the relevant information of the selected target is used as the initialization information of the target to be tracked, and it is used to initialize the template of the target to be tracked in the target tracking algorithm of the Tracker.
[0034] AsFigure 3 As shown, the first model is used to implement the target tracking algorithm. The first model includes a first input end, a second input end, a backbone network module, an MCBN module, a first CCBR module, a second CCBR module, a third CCBR module, and a fourth CCBR module. The first input end is used to receive and process a reference frame image including the target to be tracked. The second input end is used to receive and process a video frame to be detected in the video stream. The processed reference frame image of the target to be tracked and the processed video frame to be detected are input into the backbone network module. The backbone network module is a feature extraction network, which is used to extract the first feature of the processed reference frame image of the target to be tracked and the second feature of the processed video frame to be detected. The first feature and the second feature are respectively input into the MCBN module, and the MCBN module performs pixel-level correlation calculation and fusion processing on the first feature and the second feature respectively to obtain a third feature and a fourth feature corresponding to the first feature and the second feature respectively. The third feature is processed through convolution, batch normalization, and the RELU activation function to obtain a fifth feature. The fourth feature is processed through convolution, batch normalization, and the RELU activation function to obtain a sixth feature. The fifth feature is input into the first CCBR module. The first CCBR module performs pixel-level correlation fusion processing on the fifth feature and inputs the fused feature into the classification branch. The classification branch includes four sequentially connected examplar attention modules. The examplar attention modules in the classification branch are used to perform attention calculation on the classification features. The obtained seventh feature is input into the second CCBR module to obtain the classification score of the target to be tracked. The sixth feature is input into the third CCBR module. The third CCBR module performs pixel-level correlation fusion processing and inputs the fused feature into the regression branch. The regression branch includes six sequentially connected examplar attention modules. The examplar attention modules in the regression branch are used to perform attention calculation on the regression features. The obtained eighth feature is input into the fourth CCBR module to obtain the coordinates of the target to be tracked. The obtained classification score of the target to be tracked and the coordinates of the target to be tracked are subjected to matching and screening to obtain the position where the target to be tracked is located.
[0035] The target tracking algorithm implemented by the first model is ET.Track, namely Exemplaar Transformer tracking. The architecture of the first model is based on the lightweighting of the Transformer tracking model, that is, a lightweight backbone network for tracking obtained through NAS search in the middle is adopted, and the first dimensions of K and W in the Exemplar are set to 4. At the same time, the depthwise separable convolution technology is combined to reduce the computational amount of the entire network to achieve the purpose of lightweighting. In this embodiment, the target tracking algorithm implemented by the first model is ET.Track, which serves as the basis of the real-time visual detection and tracking system. The two input ends respectively receive a reference frame and a search frame. The reference frame refers to the first frame of reference frame image, which is cropped to twice the size of the target to be tracked, and the search frame is the next video frame to be tracked. Both frames pass through the backbone network to obtain the feature reference frame feature of the target in the reference frame and the search frame feature of the frame to be tracked. Then, both features are processed by the MCBN module, and pixel-level correlation operations are performed on the obtained features. The purpose of the correlation operation is to enable the target information in the reference frame to be fused into the feature of the video frame to be tracked, so as to enhance the feature of the video frame to be tracked. Finally, after the fused features pass through two convolutions, they are processed by batch normalization and ReLU activation, and the processing results are sent into the exemplar attention module, and then after passing through the CCBR module, classification and regression predictions are made. The classification branch uses 4 exemplar attention modules, and the regression branch uses 6 exemplar attention modules. The specific exemplar attention module is as Figure 4 shown. This module is a conventional Transformer attention module. Among them, K and W are respectively learnable parameters of 4x256 and 4x1, which are used for linear projection of the spatial position of the input X. This exemplar module can learn spatial correlation as a self-attention layer, but each feature needs to be calculated for attention with other features. By learning a set of typical representations inherent in a group of data, it is used to express the difference between the target feature information to be tracked and the context. In this embodiment, after having the feature of the target in the reference frame, when the subsequent video frame arrives, the search frame feature can be obtained according to the process of the search branch, and then fused with the reference frame feature, and then the target tracking in this frame is carried out.
[0036] Based on the second model, acceleration processing is performed to obtain the tracking information and detection results of the target to be tracked. Among them, the second model uses the ONNX framework to implement the target tracking algorithm, and the third model uses TensorRT to implement the target tracking algorithm. In this embodiment, the target tracking algorithm needs to be deployed on an embedded device. The inference speed of the original algorithm on Quotra p6000 is 30FPS, which far cannot meet the deployment requirements of an embedded platform with less computing power. It is accelerated by converting it into the ONNX and TesnsorRT formats.
[0037] The second model implements the object tracking algorithm using the ONNX framework and accelerates it in a way of stage splitting, so as to achieve the acceleration of the visual detection and tracking system. The second model includes a first reference frame inference branch, a first search frame inference branch, and a first classification and regression inference branch. Among them, the first reference frame inference branch includes a first backbone network module and a first processing module. The first backbone network module uses the first frame of the selected target to be tracked in the video stream as the reference frame; the first backbone network module extracts the features of the target to be tracked in the reference frame, and the first processing module is used to process the reference frame features, batch-normalize the reference frame features, and obtain the normalized reference frame features; the first search frame inference branch includes a second backbone network module and a second processing module. The second backbone network module shares weight parameters with the first backbone network module. The second backbone network module receives all subsequent image frames of the reference frame in the video stream, and uses each subsequent image frame as a search frame. The second backbone network module extracts the features of the search frame; the second processing module is used to process the search frame features, batch-normalize the reference frame features, and obtain the normalized reference frame features. The first search frame inference branch obtains the depth features of each search frame; the first classification and regression inference branch receives the depth features of the target to be tracked and the depth features of each search frame, performs pixel-level correlation fusion processing on the depth features of the target to be tracked and the depth features of each search frame respectively to obtain the ninth feature, and inputs the ninth feature into the fifth CCBR module. The fifth CCBR module performs pixel-level correlation fusion processing on the ninth feature and inputs the fused feature into the classification branch. The classification branch includes four exemplar attention modules connected in sequence. The exemplar attention modules in the classification branch are used to calculate the attention of the classification features, input the obtained tenth feature into the second CCBR module to obtain the target classification score of the target to be tracked; input the ninth feature into the seventh CCBR module. The seventh CCBR module performs pixel-level correlation fusion processing and inputs the fused feature into the regression branch. The regression branch includes six exemplar attention modules connected in sequence. The exemplar attention modules in the regression branch are used to calculate the attention of the regression features, input the obtained tenth feature into the fourth CCBR module to obtain the coordinates of the target to be tracked; perform matching and screening on the obtained target classification score of the target to be tracked and the coordinates of the target to be tracked to obtain the position where the target to be tracked is located.
[0038] In this embodiment, ONNX is an open format for representing machine learning models. It defines a set of general operators that can be used to express the building blocks of machine learning and deep learning models, supports model optimization and model deployment, and can help users develop, optimize, and deploy deep learning models more efficiently. Due to its good compatibility, the present invention uses it to optimize and deploy the tracking algorithm model.
[0039] Since the ET.Track tracking algorithm model is a siamese network, the input sizes of different branches are different, and the training and inference phases are very different, so it cannot be converted as a whole network in the way of forward inference during training. Therefore, before model conversion, the model needs to be split first, and then reconverted into the ONNX format for inference.
[0040] In this embodiment, the feature extraction model of the reference frame is the reference frame backbone network, and the input data is image data with a size of 1x3x128x128, and the output is the depth tensor feature normalized reference frame feature of 1x96x8x8; the feature extraction model of the search frame is the search frame backbone network, and the input size of the model image is 1x3x256x256, and the output is the depth tensor feature normalized search frame feature of 1x96x16x16; the feature fusion, classification and regression inference branch model is shown in Fig. 5(c), and its input is the output of the reference frame inference branch and the search frame inference branch. After pixel-level feature fusion, the search frame tracking feature fusion feature with the target information of the reference frame is obtained, and then it is respectively input into the classification and regression branches for the inference of the target position and the inference of the bounding box. The target position is the position with the largest classification score in the search video frame, and the bounding box is the coordinate of the target rectangle box.
[0041] Acceleration processing is performed based on the third model to obtain the tracking information and detection results of the target to be tracked. The third model uses TensorRT to implement the target tracking algorithm. The third model converts the target tracking algorithm into the TensorRT format and accelerates it by splitting it in stages and designing the multi-input convolution operator MultiInputConv, so as to realize the acceleration of the ET.Track tracking algorithm based on TensorRT.
[0042] The third model includes a second reference frame inference branch, a second search frame inference branch, and a second classification and regression inference branch. The reference frame inference branch includes a fourth backbone network module and a batch normalization module. The fourth backbone network module uses the first frame of the video stream in which the target to be tracked is selected as the reference frame. The fourth backbone network module extracts the depth features of the target to be tracked in the reference frame. The normalization module is used to normalize the depth features to obtain the normalized depth features of the target to be tracked. The second search frame inference branch includes a fifth backbone network module and a batch normalization module. The fifth backbone network module shares weight parameters with the fourth backbone network module. The fifth backbone network module receives all subsequent image frames of the reference frame in the video stream and uses each subsequent image frame as the reference frame. The fifth backbone network module extracts the features of the reference frame. The batch normalization module is used to process the search frame features and batch-normalize the reference frame features to obtain the normalized reference frame features. The second search frame inference branch obtains the normalized depth features of each search frame. The second classification and regression inference branch includes a pixel-level feature fusion module, an Exemplar attention module, and a CCBR module. The pixel-level feature fusion module receives the normalized depth features of the target to be tracked and the normalized depth features of each search frame. After being processed by the pixel-level feature fusion module, the enhanced depth features of the target to be tracked and the enhanced normalized depth features of each search frame are obtained.
[0043] The features after pixel-level fusion are input into the ninth CCBR module. The ninth CCBR module performs two convolutions, one batch normalization, and one ReLU activation on the pixel-level related fusion features, and inputs the processed features into the Exemplar attention module. The Exemplar attention module includes four exemplar attention layers connected in sequence. The exemplar attention module in the classification branch is used to calculate the attention of the classification features, and inputs the obtained tenth feature into the tenth CCBR module to obtain the target classification score to be tracked. The ninth feature is input into the eleventh CCBR module. The eleventh CCBR module performs two convolutions, one batch normalization, and one ReLU activation on the pixel-level related fusion features, inputs the processed features into the Exemplar attention module, and inputs the fusion-processed features into the regression branch. The regression branch includes six exemplar attention layers connected in sequence. The exemplar attention module in the regression branch is used to calculate the attention of the regression features, and inputs the obtained eleventh feature into the twelfth CCBR module to obtain the coordinates of the target to be tracked. The obtained target classification score to be tracked and the coordinates of the target to be tracked are subjected to matching and screening to obtain the position where the target to be tracked is located.
[0044] As Figure 6As shown, in this embodiment, TensorRT is a library introduced by NVIDIA for high-performance deep learning inference, which can accelerate deep learning inference on NVIDIA GPUs. It has a highly optimized engine that can optimize a trained deep learning model into an efficient inference model to improve inference performance and throughput. TensorRT supports multiple deep learning frameworks and can convert the models trained by these frameworks into a format usable by TensorRT. In addition to model optimization, TensorRT also provides some tools and APIs for model deployment and integration into applications, and it is widely used in various deep learning applications.
[0045] The reason for using TensorRT for model deployment is that it has an advantage over ONNX in terms of model inference speed. However, the operators it supports are not as rich as those of ONNX. Therefore, when converting and inferring a model, it is often necessary to develop custom operators and then embed them into the model conversion framework for conversion.
[0046] The splitting scheme of TensorRT is roughly the same as that of ONNX. The reference frame backbone network and the search frame backbone network adopt the same scheme as (a) and (b) in ONNX. For (c), since it cannot be directly split, this patent reuses its basic operators to reconstruct the Exemplar layer and adopts the scheme as Figure 6 shown for conversion.
[0047] Since the ET.Track model used depends on a multi-input dynamic convolution operator. This operator requires the input weights to be the tensor features generated in the previous step, that is, the dynamic convolution kernel. However, no plugin for this operator is provided in TensorRT. Therefore, when converting the torch model of ET.Track into a TensorRT model, the model conversion will fail or the behavior of the converted model will be abnormal. The abnormalities are manifested as follows: when directly converting to a TensorRT model using torch2trt, neither the Gaussian response map nor the regression branch of the tracking model can be accurately predicted. Specifically, the confidence in the Gaussian response map drops sharply. During the movement of the target, the confidence drops from a score of more than 0.9 to below 0.1. Although the tracked target can still be normally followed, the scale does not change with the change of the target.
[0048] Furthermore, multiple TensorRT plugins are configured for the third model. The TensorRT plugins receive the dependent convolution parameters and are used to simulate the operators that cannot be accelerated in the Exemplar layer. The TensorRT plugins are called by the Exemplar attention layer during the inference of the tracking algorithm.
[0049] In this embodiment, a TensorRT plugin multi-input convolution plugin is developed, and a dynamic library file is generated for it under Ubuntu. Then, under the torch2trt framework, this operator is called to create a multi-input dynamic convolution operator, and a converter is written using torch.nn.functional.conv2d. This enables the automatic capture of inputs according to the called multi-input function and hook during the conversion of the ET.Track tracking model and introduces them into the custom operator for calculation. Thus, problems such as the inability to convert the model or abnormal converted models are avoided.
[0050] The defined operator is called MultiInputConv Operator (multi-input convolution operation), and the defined plugin is called MultiInputConvPlugin (multi-input convolution plugin). Two classes need to be implemented for this operator, namely MultiInputConvPlugin and MultiInputConvCreator, which inherit from nvinfer1::IPluginV2DynamicExt and public nvinfer1::IPluginCreator. The MultiInputConvPlugin class is responsible for serializing, calculating, and deserializing this operator. In this operator, CUDNN is used to implement forward inference, and the convolution mode is CUDNN_CROSS_CORRELATION mode. MultiInputConvCreator is responsible for providing an interface to create this plugin. By calling the corresponding constructor of MultiInputConvPlugin and passing in the given parameters, a MultiInputConv plugin is created. Finally, the convolution operator is registered through REGISTER_TENSORRT_PLUGIN(MultiInputConvPlugin) for future calls. When the creator of torch2trt calls TensorRT's get_plugin_registry() to return the registration class and calls get_plugin_creator("MultiInputConv", '1', ”) to obtain the Creator in the multi-input convolution dynamic library. Then, the passed-in initialization information, namely the four parameters of strides, pads, dilations, and groups, is collected through PluginFieldCollection, and a plugin is created by the creator's create_plugin according to the given layer name and the collected information for the converter to call.
[0051] This application adopts a multi - process design, introduces a command and control platform, and realizes the generation of task instructions and the selection of target tracking. Through the command and control platform, the target to be tracked can be specified according to the detected target information, and the information on whether to accelerate can be obtained according to the task instructions, and the acceleration method is selected to accelerate the inference of the tracking algorithm, thereby improving the real - time performance and performance of the system. This method can flexibly generate task instructions and select target tracking according to actual needs, improving the intelligence and flexibility of the system.
[0052] The tracking algorithm model that cannot be directly converted in this application is split according to the splitting method of this patent, so that it is converted into a model that can be converted under ONNX and TensorRT, and inference acceleration and deployment are carried out through the corresponding framework. To solve the abnormal conversion of the algorithm model under TensorRT, this patent designs a multi - input convolution plugin, encapsulates it into a dynamic library using C++ under ubuntu, calls this dynamic library under torch2trt, and designs a converter so that the model can be normally converted and the confidence after inference is normal.
[0053] This application adopts a multi - process parallel method, combines deep - learning lightweight object detection, a command and control platform, and ROS to design a lightweight object recognition and tracking system based on Transformer; this application designs a model conversion and splitting scheme in ONNX and TensorRT to enable the ET.Track model to be accelerated and deployed to an embedded platform; this application designs a multi - input dynamic convolution operator plugin and generates a dynamic library file for use in the conversion of the Exemplar layer.
[0054] As Figures 7(A)-7(B) shown, in the current hardware environment of Quatro RTX P6000, the inference speed of the original model is 30FPS. After the above conversion, the inference speed under onnx is increased to 85FPS +. When the model is converted to TensorRT, its inference speed under fp16 precision reaches a speed of 195FPS +, and the inference speed on the edge - computing device Orion can reach 40FPS +, meeting the real - time requirements of deployment. At the same time, when running this algorithm under ONNX and TensorRT, its accuracy does not decrease significantly. Using OTB100 for testing, the results are not much different from the original method.
[0055] Figures 7(A)-7(B)Among them, the precision plot refers to the percentage of video frames where the distance between the predicted center position and the ground truth center position is less than 20 pixels, and the success plot represents the percentage of video frames successfully tracked under multiple overlapping thresholds. ETTrack_ori represents the effect of the original ET.Track tracking model, ETTrack_onnx represents the tracking model result of the onnx-accelerated model, and ETTrack represents the effect of the tracking model after TensorRT acceleration. As can be seen from the above figure, compared with the original ET.Track model, the speed of the accelerated model has increased several times, but the accuracy can still reach the level of the pre-acceleration model.
[0056] In this embodiment, for the tracking model splitting scheme, there may be other splitting schemes, such as splitting according to a single functional module. However, such a splitting method increases the burden of IO and reduces the inference performance. In addition, for the acceleration scheme, mobile acceleration methods such as MNN, NCNN, TNN, etc. can also be used, but the hardware they rely on cannot match the speed of the current edge-side hardware.
[0057] Furthermore, the method of embedding real-time visual detection and tracking in edge computing devices is applied to the unmanned vehicle chassis control system. The unmanned vehicle chassis control system uses the topic communication mechanism. Here, the topic name is Track_topic, receives the encapsulated tracking information, parses the position and direction information of the tracked target from the encapsulated tracking information, performs coordinate calculation on the position and direction information of the tracked target to obtain the moving direction and distance of the tracked target, and then the unmanned vehicle chassis control module controls the moving direction and distance of its own unmanned vehicle according to the obtained moving direction and distance of the tracked target, so as to track the tracked target.
[0058] In this embodiment, the unmanned vehicle chassis control system is an unmanned vehicle chassis control node, subscribes to the tracking information from the tracking topic node, performs coordinate calculation on the tracked target information it obtains, and transmits the obtained moving direction and distance to the unmanned vehicle chassis control module. The unmanned vehicle chassis control module controls the unmanned vehicle chassis according to the settlement information.
[0059] This application uses the lightweight object tracking algorithm ET.Track based on Transformer: The lightweight object tracking algorithm ET.Track based on the Transformer architecture is adopted. While ensuring high tracking accuracy, the algorithm reduces the computational complexity, enabling it to run in real time on edge embedded devices.
[0060] This application splits the backbone network and the prediction head for ONNX and TensorRT model conversion: To address the issue that it is difficult to directly perform ONNX and TensorRT model conversion for the ET.Track algorithm, it is proposed to split it into two major parts, the backbone network and the prediction head, for separate conversion to achieve more efficient model deployment.
[0061] This application converts the ET.Track tracking model to TensorRT for inference: To solve problems such as insufficient real-time performance of the ONNX-converted model on edge devices, the model is converted to TensorRT for inference to further improve the running efficiency of the model on edge devices.
[0062] This application uses CUDNN to write the MultiInputConv operator as a TensorRT plugin: To address the problem that the ET.Track tracking model after conversion to TensorRT does not support the multi-input dynamic convolution in the Exemplar layer at the TensorRT bottom layer, CUDNN is used to write the MultiInputConv operator as a TensorRT plugin and compile it into a dynamic library for calling in torch2trt.
[0063] This application writes a converter in torch2trt to directly convert the tracking model from torch to TensorRT for inference, further improving the running efficiency on edge devices.
[0064] This application embeds the accelerated tracking algorithm into the target tracking system based on the ROS operating system: To achieve the wide application of the algorithm in actual application scenarios, the accelerated tracking algorithm is embedded as a node into the target tracking framework based on ROS (Robot Operating System);
[0065] This application is applied to various tasks of small unmanned vehicles. The accelerated ET.Track algorithm can be used in tasks such as target tracking, target following, target searching, and visual grasping assistance for small unmanned vehicles, providing a more efficient and reliable target tracking solution for actual application scenarios.
[0066] Through the above solutions, the present invention has successfully achieved the deployment of edge embedded devices based on the ET.Track target tracking algorithm, providing an effective solution to the actual challenges in real-world scenarios.
[0067] This application transmits the results of object detection to the accusation platform, providing a basis for issuing task instructions to it. Combined with a lightweight tracking method, it realizes the tracking and following of specified targets. It is embedded in ROS and communicates with the unmanned vehicle chassis node through the topic method to control the movement of the chassis. Despite using a deep learning solution based on Transformer, this system can still run in real time on an embedded platform.
[0068] Compared with other Transformer-based object tracking algorithms deployed on the cloud platform, this application splits the ET.Track tracking algorithm that cannot run in real time on the embedded side into three segments, enabling the model that cannot be directly converted to be deployed in the corresponding inference optimization framework and meeting the real-time requirements.
[0069] After converting the TensorRT model based on the torch2trt framework and loading it through TRModule for use, neither the Gaussian response map nor the regression branch of the tracking model can accurately predict. Specifically, the confidence in the Gaussian response map drops sharply, and the bbox regression branch cannot accurately predict the scale of the target being tracked. Through various experimental schemes, it is located that the problem lies in that TensorRT does not support the multi-input dynamic convolution operator, and this operator is required in the Exemplar layer to perform convolution operations on the input using dynamically generated weights. Therefore, this application redesigned the Exemplar layer and designed a multi-input convolution plugin MultiInputConvPlugin for the unsupported operator. Finally, the problem that the Exemplar layer cannot be accelerated in TensorRT is solved.
[0070] The second aspect of the present invention discloses a lightweight object recognition and tracking system based on Transformer, also known as a real-time visual detection and tracking system for the embedded end of an edge computing device. Figure 8 It is a structural diagram of a real-time visual detection and tracking system (800) for the embedded end of an edge computing device according to an embodiment of the present invention; the system includes:
[0071] An acquisition unit 802, configured to acquire a video stream, obtain a target to be tracked from the video stream, and generate a task instruction for tracking the target to be tracked;
[0072] An analysis unit 804, configured to analyze the task instruction, where the task instruction includes the position of the target to be tracked and an acceleration flag, and the acceleration flag is used to indicate whether the tracking algorithm needs to be accelerated;
[0073] A tracking and detection unit 806, configured to track and detect the target to be tracked based on the position and the acceleration flag, and obtain the tracking information and detection result of the target to be tracked;
[0074] The encapsulation unit 808 is configured to encapsulate the detection result and the tracking information to obtain the encapsulated tracking information, and send out the encapsulated tracking information in the form of a tracking topic, so that the unmanned vehicle chassis control node can obtain the detection result and the tracking information.
[0075] A third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps in the lightweight target recognition and tracking method based on Transformer developed by the present disclosure are implemented.
[0076] The electronic device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, near field communication (NFC), or other technologies. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, a touchpad, or a mouse, etc.
[0077] Those skilled in the art can understand that the electronic device is only a structural diagram of a part related to the technical solution of the present disclosure, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0078] A fourth aspect of the present invention discloses a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps in the lightweight target recognition and tracking method based on Transformer developed by the present disclosure are implemented.
[0079] Please note that the technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification. The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A lightweight object recognition and tracking system based on Transformer, characterized in that, Including: A target recognition and tracking node and an intelligent unmanned chassis control node; The target recognition and tracking node includes: An acquisition unit, configured to acquire a video stream, obtain a target to be tracked from the video stream, generate a task instruction for tracking the target to be tracked, and send the task instruction to an analysis unit; The analysis unit is configured to analyze a task instruction including the position of the target to be tracked and an acceleration identifier, determine whether to select the target according to the task instruction, if not, return to the command and control platform, if so, use the position of the selected target and the acceleration identifier as the initialization information of the target to be tracked; The tracking and detection unit is configured to track and detect the target to be tracked based on the position and the acceleration identifier, determine whether acceleration processing is required according to the indication of the acceleration identifier, if so, track and detect the target to be tracked based on a first model to obtain the tracking information and detection result of the target to be tracked; if not, based on the acceleration type indicated by the acceleration identifier, perform acceleration processing based on a second model or a third model to obtain the tracking information and detection result of the target to be tracked; Send the tracking information and the detection result to an encapsulation unit; The encapsulation unit is configured to encapsulate the detection result and the tracking information to obtain the encapsulated tracking information, and send the encapsulated tracking information to the intelligent unmanned chassis control node in the form of a tracking topic; The intelligent unmanned chassis control node includes: A tracking topic acquisition module, which acquires a tracking topic from the tracking information; A tracking information acquisition module, which subscribes to the tracking topic to obtain the information of the target to be tracked and sends it to a coordinate calculation module; The coordinate calculation module transmits the moving direction and distance information obtained by calculating the coordinates of the information of the target to be tracked to the intelligent unmanned chassis control module; The intelligent unmanned chassis control module receives the moving direction and distance information to control the intelligent unmanned chassis.
2. The lightweight target recognition and tracking system according to claim 1, characterized in that The acquisition unit includes: An algorithm detection sub-module, which acquires a video stream by using the method of topic subscription or rtsp protocol, detects target information based on the video stream, and outputs it to the command and control platform; The command and control platform, upon receiving the target information, generates a task instruction for tracking the target to be tracked, and sends the task instruction to the analysis unit; Wherein, the task instruction indicates whether to select the target, and determines the selected target according to the detected target information.
3. The lightweight target recognition and tracking system according to claim 2, wherein the first model includes a first input end, a second input end, a backbone network module, an MCBN module, a first CCBR module, a second CCBR module, a third CCBR module, and a fourth CCBR module, characterized in that Tracking and detecting the target to be tracked based on the first model includes: Receiving a reference frame image of the target to be tracked through a first input end, and receiving a video frame to be detected in the video stream through a second input end; Inputting the reference frame image and the video frame to be detected into a backbone network module, wherein the backbone network module extracts a first feature of the reference frame image and a second feature of the video frame to be detected; Input the first feature and the second feature into the MCBN module. Among them, the MCBN module performs pixel-level fusion processing on the first feature to obtain a third feature, performs pixel-level fusion processing on the second feature to obtain a fourth feature, subjects the third feature to convolution processing, batch normalization, and RELU activation function processing to obtain a fifth feature, and subjects the fourth feature to convolution processing, batch normalization, and RELU activation function processing to obtain a sixth feature; Input the fifth feature into the first CCBR module and input the sixth feature into the third CCBR module. Among them, the first CCBR module performs pixel-level correlation fusion processing on the fifth feature to obtain a seventh feature, and the third CCBR module performs pixel-level correlation fusion processing on the sixth feature to obtain an eighth feature; Input the seventh feature into the second CCBR module to obtain the classification score of the target to be tracked, input the eighth feature into the fourth CCBR module to obtain the coordinates of the target to be tracked, and perform matching screening on the classification score of the target to be tracked and the coordinates of the target to be tracked to obtain the tracking information and detection result of the target to be tracked.
4. The lightweight target recognition and tracking system according to claim 3, wherein the second model includes a first reference frame inference branch, a first search frame inference branch, and a first classification and regression inference branch, characterized in that Performing acceleration processing based on the second model to obtain the tracking information and detection result of the target to be tracked includes: Receiving the reference frame of the target to be tracked through the first reference frame inference branch, and extracting and normalizing the reference frame to obtain the normalized reference frame feature; Receiving the subsequent image frames of the reference frame in the video stream through the first search frame inference branch, using the subsequent image frames as search frames to extract and normalize the search frames to obtain search frame depth features, where the first search frame inference branch shares the weight parameters for extraction and normalization with the first reference frame inference branch; Receiving the normalized reference frame feature and the search frame depth feature through the first classification and regression inference branch, and respectively performing pixel-level fusion processing on the normalized reference frame feature and the search frame depth feature to obtain the classification score of the target to be tracked and the coordinates of the target to be tracked, and performing matching screening on the classification score of the target to be tracked and the coordinates of the target to be tracked to obtain the tracking information and detection result of the target to be tracked.
5. The lightweight target recognition and tracking system according to claim 4, wherein the third model includes a second reference frame inference branch, a second search frame inference branch, and a second classification and regression inference branch, where The reference frame inference branch includes a fourth backbone network module and a batch normalization module, the second search frame inference branch includes a fifth backbone network module and a batch normalization module, the second classification and regression inference branch includes a pixel-level feature fusion module, an Exemplar attention module, and a CCBR module, the CCBR module includes a ninth CCBR module, a tenth CCBR module, and an eleventh CCBR module, the fifth backbone network module shares weight parameters with the fourth backbone network module. It is characterized in that performing acceleration processing based on the third model to obtain the tracking information and detection result of the target to be tracked includes: The fourth backbone network module receives the reference frame of the target to be tracked and extracts the depth features of the target to be tracked in the reference frame; the batch normalization module normalizes the depth features of the target to be tracked to obtain the normalized reference frame depth features of the target to be tracked; The fifth backbone network module receives the subsequent image frames of the reference frame and extracts the depth features of the target to be tracked in the subsequent image frames; the batch normalization module performs batch normalization on the depth features of the target to be tracked in the subsequent image frames to obtain the normalized search frame depth features; The pixel-level feature fusion module processes the normalized reference frame depth features and the normalized search frame depth features to obtain the normalized reference frame enhanced depth features and the normalized search frame enhanced depth features, and inputs the normalized reference frame enhanced depth features and the normalized search frame enhanced depth features into the ninth CCBR module and the eleventh CCBR module. Among them, the ninth CCBR module performs two convolutions, one batch normalization, and one ReLU activation processing on the normalized reference frame enhanced depth features, and inputs the processed features into the Exemplar attention module to obtain the tenth feature; the eleventh CCBR module performs two convolutions, one batch normalization, and one ReLU activation processing on the normalized search frame enhanced depth features, and inputs the processed features into the Exemplar attention module to obtain the eleventh feature; the tenth feature is input into the tenth CCBR module to obtain the classification score of the target to be tracked, the eleventh feature is input into the twelfth CCBR module to obtain the coordinates of the target to be tracked, and the classification score of the target to be tracked and the coordinates of the target to be tracked are matched and screened to obtain the tracking information and detection results of the target to be tracked.
6. The lightweight target recognition and tracking system according to claim 5, wherein the third model includes a plurality of TensorRT plugins, and the plurality of TensorRT plugins generate dynamic library files. The TensorRT plugin receives the dependent convolution parameters and is used to simulate the tracking algorithm that cannot be accelerated in the Exemplar attention module. The TensorRT plugin is called by the Exemplar attention module during the inference of the tracking algorithm.
7. A lightweight object recognition and tracking method based on Transformer, characterized in that The method includes: Obtain a video stream, obtain a target to be tracked from the video stream, and generate a task instruction for tracking the target to be tracked; Parse the task instruction, where the task instruction includes the position and acceleration flag of the target to be tracked. Determine whether the target is selected according to the task instruction. If not, return to the command platform. If so, use the position and acceleration flag of the selected target as the initialization information of the target to be tracked; Track and detect the target to be tracked based on the position and the acceleration identification, and judge whether acceleration processing is required according to the indication of the acceleration identification. If so, track and detect the target to be tracked based on the first model to obtain the tracking information and detection results of the target to be tracked; if not, based on the acceleration type indicated by the acceleration identification, perform acceleration processing based on the second model or the third model to obtain the tracking information and detection results of the target to be tracked; Package the detection results and the tracking information to obtain the packaged tracking information, and send the packaged tracking information to the intelligent unmanned chassis control node in the form of a tracking topic; The intelligent unmanned chassis control node obtains the tracking topic from the tracking information, determines the information of the target to be tracked, and performs coordinate calculation on the information of the target to be tracked to obtain the target movement direction and distance information, and controls the intelligent unmanned chassis based on the movement direction and distance information.
Citation Information
Patent Citations
Multi-target tracking acceleration method based on time-space optimization in edge calculation environment
CN113723279A
Target tracking method and system based on dual attention feature fusion network
CN116030097A