Real-time target tracking method based on RKNN model efficient reasoning
Through the RKNN model and optimization algorithm, combined with the Siamese network and multi-scale feature fusion, the deployment difficulties of deep learning models on resource-constrained devices are solved, efficient and accurate real-time target tracking is achieved, and the target appearance changes and occlusions are adapted, which improves the computing efficiency and tracking stability of the device.
Patent Information
- Application Number
- CN202510555240.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-09-19
AI Technical Summary
Deep learning models are difficult to deploy on resource-constrained devices, require large computing resources, take long to train, and are not adaptable enough to changes in target appearance and occlusion.
The RKNN model and optimized tracking algorithm are used to extract template features through the Siamese network, combine multi-scale feature fusion and template matching, and combine historical trajectory information to track the target, and then optimize and adapt it on the NPU of the RK platform.
It achieves efficient and accurate real-time target tracking on resource-constrained devices, reduces computing and storage resource requirements, and improves adaptability, stability, and accuracy to target appearance changes and occlusions.
Smart Images

Figure CN120672793A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target tracking technology, and in particular to a real-time target tracking method based on efficient reasoning of an RKNN model. Background Art
[0002] Object tracking plays an important role in the field of computer vision and is widely used in security monitoring, intelligent transportation, autonomous driving, human-computer interaction, and other fields. Traditional object tracking methods mainly include feature-based methods, model-based methods, and region-based methods.
[0003] With the development of deep learning technology, neural network-based object tracking algorithms have made significant progress. Deep learning models can automatically learn the characteristic representation of the target, and have greater adaptability and accuracy. However, deep learning models typically require large amounts of computing resources and are complex, making them difficult to deploy on resource-constrained devices such as embedded devices and mobile terminals. Furthermore, the model training process requires a large amount of labeled data and computing resources, resulting in a long training time. Summary of the Invention
[0004] The present invention is made in view of the above problems, and its purpose is to provide a real-time target tracking method based on efficient reasoning of the RKNN model. Through the optimized RKNN model and carefully designed tracking algorithm, it can more accurately locate the target, reduce tracking errors, have stronger adaptability to the target's appearance changes, occlusions, etc., and improve the stability and accuracy of tracking. The characteristics of the RKNN model make the model more efficient in resource utilization, reduce the demand for computing resources and storage resources, and can be deployed on resource-constrained devices.
[0005] Specifically, a first aspect of the present invention provides a real-time target tracking method based on efficient reasoning of an RKNN model, comprising the following steps:
[0006] Step 1: Initialize the prediction model, obtain the image frame containing the target, crop the target area and perform template feature extraction based on the Siamese network;
[0007] Step 2: Extract and fuse multi-scale features of the target area image;
[0008] Step 3: Match the template features and multi-scale fusion features based on template matching, calculate the similarity, and obtain the similarity score matrix;
[0009] Step 4: Obtain the preliminary predicted position of the target in the current frame based on the similarity score matrix;
[0010] Step 5: Combine the target's historical trajectory information to construct an objective function, calibrate the target's predicted position, and evaluate its reliability based on the confidence score;
[0011] Step 6: Optimize and adapt the prediction model based on the NPU of the RK platform and convert it into an RKNN model.
[0012] During the initialization phase, when the target tracking system is started or a new tracking task is received, the initial frame image and target frame information are acquired. An image frame containing the target is obtained from an image source (such as a camera or video file). If manual annotation is used, the user uses a graphical interface tool to select the target's initial position on the image and determine the target's bounding box (upper left and lower right corner coordinates). If automatic target detection is used, the system detects the target in the image based on a preset target detection algorithm and acquires its bounding box information. Based on the target bounding box, the target area is cropped from the original image to obtain a subwindow image. This subwindow will serve as the primary object for subsequent processing, extracting target feature information and performing template matching operations.
[0013] Furthermore, the cropping of the target area is to manually mark the target bounding box through a graphical interface tool or automatically obtain the target bounding box through a target detection algorithm.
[0014] Furthermore, the template features include geometric features and texture features.
[0015] When extracting template features, it is necessary to build a lightweight deep neural network and select a model based on the Siamese network architecture. The convolutional neural network (CNN) part of the model contains L convolutional layers, and the convolution kernel size of the lth (l<=L) convolutional layer is k l ×k l , the number of convolution kernels is n l , the step size is s l , then the size of the feature map output by the lth layer is AH l ,AW l Satisfies the formula:
[0016]
[0017] Among them: AH l The height of the output feature map of the lth convolutional layer;
[0018] AW l The width of the output feature map of the l-th convolutional layer;
[0019] H l The height of the input feature map for the lth convolutional layer;
[0020] W l The width of the input feature map of the lth convolutional layer;
[0021] k lis the size of the convolution kernel of the lth convolution layer, that is, the height or width, the height is equal to the width;
[0022] s l is the step size of the lth convolutional layer;
[0023] When building a lightweight deep neural network for template feature extraction, this formula is used to calculate the height and width of the output feature map of the lth convolution layer. It can also accurately derive the size of the output feature map after each convolution operation based on the size of the input feature map, the convolution kernel size, and the step size. This is crucial for network structure design and feature extraction process, and affects the subsequent effective capture of target features.
[0024] By optimizing the number and size of convolution kernels and adjusting the connection mode of network layers, the complexity of the model can be significantly reduced while maintaining the ability to effectively extract target features. The model parameter P can be approximately calculated as:
[0025]
[0026] Where: P is the model parameter;
[0027] L is the total number of convolutional layers contained in the convolutional neural network (CNN) part;
[0028] n l is the number of convolution kernels in the lth convolution layer;
[0029] k l is the size of the convolution kernel of the lth convolution layer, i.e., height or width;
[0030] n l-1 is the number of convolution kernels in the l-1th convolution layer;
[0031] The cropped target sub-window image is input into the lightweight deep neural network. After forward propagation calculation, the initial features of the target are extracted. The feature extraction process can be expressed as: F T =M(I R ), expressed in the form of vectors or feature maps, which can represent key information such as the appearance and texture of the target.
[0032] In the model initialization stage, the target sub-window image I is cropped from the original image R Input into the lightweight deep neural network for template feature extraction, the extracted template feature F T It is stored in the memory as a reference for target matching and position prediction in subsequent frames. M represents a lightweight deep neural network model. The formula reflects the process of feature extraction of the target sub-window image through the network. The extracted template feature F TIt will serve as an important reference benchmark for target matching and position prediction in subsequent frames.
[0033] Furthermore, the step 2 includes the following steps:
[0034] Step 2.1: Using the pyramid feature extraction method, perform multi-scale cropping operations on the cropped target area image to form a multi-resolution image set;
[0035] Multi-scale feature extraction uses pyramid feature extraction method, for the sub-window image I where the target is located R , perform multi-scale cropping operations, according to the preset scale factor α={α1,α2,…,α m}, generate m sub-window images with different resolutions Its size satisfy Form a multi-resolution feature input set.
[0036] Through this multi-scale cropping, sub-window images at different resolutions can be obtained, providing rich input for the subsequent use of a lightweight feature pyramid network (FPN) to extract multi-scale deep features, and enhancing the algorithm's adaptability to target deformation and size changes.
[0037] Step 2.2: Extract features from multi-resolution images using a lightweight feature pyramid network (FPN).
[0038] A lightweight Feature Pyramid Network (FPN) is used to extract features from sub-window images at multiple resolutions. Through a top-down and lateral connection structure, FPN performs convolution operations on feature maps at different scales, extracting multi-scale deep features with rich semantic information. Lower-resolution feature maps capture the overall structural information of the target, while higher-resolution feature maps capture detailed features, enhancing the algorithm's adaptability to target deformation and size changes.
[0039] Step 2.3: Perform weighted fusion processing on the different scale features extracted from the multi-scale feature pyramid to obtain multi-scale fusion features.
[0040] The feature fusion method is to design a weighted fusion mechanism to extract different scale features from the multi-scale feature pyramid. Perform fusion processing. Assign a weight coefficient ω to each scale feature i The weight coefficient is determined based on the importance and confidence of the scale feature in the target representation. For example, the weight coefficient can be dynamically adjusted by analyzing the contribution of different scale features to target positioning and recognition.
[0041] Multiply each scale feature by its corresponding weight coefficient and then perform the accumulation operation to obtain a unified feature representation F fused , the formula is as follows:
[0042]
[0043] Among them: F fused It is the unified feature representation after fusion;
[0044] m is the total number of scales;
[0045] is the feature of the i-th scale;
[0046] ω i is the weight coefficient of the feature of the i-th scale;
[0047] This fusion method can effectively integrate local detail features and global structural features. i The weight coefficient assigned to each scale feature is determined based on the importance and confidence of the scale feature in the target representation. By multiplying each scale feature with its corresponding weight coefficient and then accumulating them, it can effectively integrate local detail features and global structural features, making the final feature representation more representative and robust, providing more accurate information for subsequent target matching and position prediction.
[0048] Furthermore, the step 2.2 includes: capturing the overall structural information of the target on a feature map with lower resolution through a lightweight feature pyramid network FPN; and obtaining detailed features of the target on a feature map with higher resolution.
[0049] In the present invention, a higher resolution means that the number of pixels per frame is greater than or equal to 1920×1080, and a lower resolution means that the number of pixels per frame is less than 1920×1080.
[0050] Furthermore, the step three includes: using a template matching-based strategy to match the multi-scale fusion features extracted from the current frame with the target template features, calculating the correlation between the two features in the feature space through a similarity measurement method, and obtaining a similarity score matrix.
[0051] In the process of target matching and position prediction, a template matching-based strategy is adopted to combine the multi-scale fusion features F extracted from the current frame into the fused and the stored target template feature F T Perform matching calculations and calculate the similarity between the two. The similarity measurement method can use efficient lightweight cross-correlation operations. By calculating the correlation between the template features and the current frame features in the feature space, a similarity score matrix S is obtained. The calculation formula is:
[0052] sij =F fused (i,j)·F T ;
[0053] Where: s ij is an element in the similarity score matrix S, representing the template feature F T The degree of similarity with the feature at the corresponding position (i, j) in the current frame;
[0054] F fused (i, j) is the feature of position (i, j) in the multi-scale fusion feature;
[0055] F T is the template feature;
[0056] It is a vector dot product operation;
[0057] Each element in the similarity score matrix S represents the degree of similarity between the template feature and the feature at the corresponding position in the current frame. By calculating the similarity score matrix, the correlation between the template feature and the features at each position in the current frame can be measured in the feature space, thereby determining the possible position of the target in the current frame. The area with a higher similarity score is more likely to be the location of the target.
[0058] Furthermore, the step four includes: performing non-maximum suppression processing on the similarity score matrix, removing areas with high local similarity but not global optimal, and extracting the candidate area corresponding to the highest score, which is the preliminary predicted position of the target in the current frame.
[0059] Let the threshold of non-maximum suppression be t NMS , then the candidate region after non-maximum suppression (NMS) processing:
[0060]
[0061] Among them: B pred is the candidate area range;
[0062] is the minimum value of the x-axis of the candidate area;
[0063] is the minimum value of the y-axis of the candidate area;
[0064] is the maximum value of the x-axis of the candidate region;
[0065] is the maximum value of the y-axis of the candidate region.
[0066] Furthermore, the step five comprises the following steps:
[0067] Step 5.1: Based on the preliminary predicted position of the target, further output the center position and size of the target;
[0068] Central Location:
[0069]
[0070] size:
[0071] S pred =(ω pred ,h pred );
[0072]
[0073] Where: C pred is the center position coordinate of the target;
[0074] is the median value of the x-axis of the candidate region;
[0075] is the median value of the y-axis of the candidate region;
[0076] is the minimum value of the x-axis of the candidate area;
[0077] is the minimum value of the y-axis of the candidate area;
[0078] is the maximum value of the x-axis of the candidate region;
[0079] is the maximum value of the y-axis of the candidate area;
[0080] S pred is the size of the target;
[0081] ω pred is the width of the candidate region;
[0082] h pred is the height of the candidate area;
[0083] Step 5.2: Integrate the historical trajectory information of the target, smoothly adjust the target predicted position of the current frame, and correct the predicted position;
[0084] The target's historical trajectory information is integrated to smoothly adjust the target's predicted position in the current frame. This historical trajectory information includes the target's position, velocity, acceleration, and other motion information from past frames. By analyzing the target's historical motion patterns and predicting its reasonable position range in the current frame, the network's predicted position is corrected, improving the accuracy and stability of the target's position prediction.
[0085] Step 5.3: Based on the nonlinear optimization method, the historical trajectory information of the target is combined with the feature information of the current frame to construct the objective function;
[0086] Using an optimized nonlinear update method, the historical trajectory information of the target is combined with the feature information of the current frame to construct the objective function E, which is formulated as follows:
[0087]
[0088] Where: E is the objective function;
[0089] N is the dimension of the feature vector;
[0090] F obs (i) is the i-th observed feature of the current frame, obtained by feature extraction from the image data of the current frame;
[0091] F pred (i) is the predicted value of the i-th feature based on the predicted position, obtained by mapping the predicted position back to the feature space;
[0092] Step 5.4: Obtain the optimal solution of the objective function and further correct the target predicted position;
[0093] By solving the optimal solution of the objective function, the target position is further corrected so that the target position update is more consistent with the actual motion situation and the jitter and drift during the tracking process are reduced.
[0094] Step 5.5: Calculate the confidence score of the current predicted position. If the confidence score is higher than the threshold, it means that the result is more reliable. If it is lower than the threshold, reinitialize the model for prediction.
[0095] The scoring mechanism of this invention evaluates the stability and reliability of target tracking based on the confidence score output by the network. The confidence score is a numerical indicator generated by the network while predicting the target position. It reflects the network's confidence in the current prediction result. The formula is as follows:
[0096] s conf =ω m ×m+ω c ×c+ω f ×f;
[0097] Where: s conf Score confidence level;
[0098] m is the matching degree of the target feature;
[0099] ω m is the weight coefficient corresponding to the matching degree of the target feature;
[0100] c is the consistency of target position;
[0101] ω c is the weight coefficient corresponding to the consistency of the target position;
[0102] f is the stability of the characteristic;
[0103] ω f is the weight coefficient corresponding to the stability of the feature;
[0104] Furthermore, the historical trajectory information includes the target's motion state information in the past several frames, including position, speed, and acceleration.
[0105] Furthermore, step six includes converting the model into ONNX format, optimizing and accelerating model reasoning using multi-platform hardware acceleration technology NPU during model deployment, adjusting the model parameters and calculation process, and converting the model into RKNN format.
[0106] The model is optimized and, after training, converted to the ONNX format, which offers excellent cross-platform performance and efficient inference. During model deployment, the multi-platform hardware acceleration technology, NPU, is used to accelerate model inference. Based on the NPU's hardware architecture, the model is optimized for specific purposes, such as optimizing data layout and adjusting calculation order, fully leveraging the NPU's computing advantages.
[0107] In embedded devices, a memory sharing strategy is employed to enable different modules or frames to share memory space for model parameters and intermediate calculation results, reducing repeated memory allocation and deallocation, and lowering memory usage. Furthermore, a batch inference strategy is employed to combine data from multiple image frames for batch processing, fully utilizing hardware computing resources, improving computational efficiency, and significantly reducing inference time.
[0108] The NPU platform is adapted to the Rockchip RK3588 platform, leveraging its integrated NPU for efficient inference. The model is optimized and adapted to the NPU characteristics of the RK3588 platform, adjusting the model's parameters and computational flow, and converting the model to the RKNN format to achieve optimal inference performance. This optimization achieves real-time performance of up to 120 FPS on this platform, meeting the requirements of real-time target tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0109] In order to more clearly illustrate the embodiments of the present drawings or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present drawings. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0110] Figure 1 is a flow chart of the steps of the present invention;
[0111] Figure 2 This is a schematic diagram of the structure of the template feature extraction network of the present invention;
[0112] Figure 3 A schematic diagram of manually marking target boundaries according to the present invention;
[0113] Figure 4 Schematic diagram of the processing time and frame rate of each frame of the model in an embodiment of the present invention;
[0114] Figure 5 This is a schematic diagram of the interface during target tracking according to an embodiment of the present invention;
[0115] Figure 6 Schematic diagram of the IOU values of some frames and tracking accuracy according to an embodiment of the present invention;
[0116] Figure 7 Schematic diagram of the total number of frames of tracking video and the total tracking accuracy rate according to an embodiment of the present invention;
[0117] The purpose, features and advantages of this drawing will be further described with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION
[0118] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention. Based on the embodiments provided by the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work are within the scope of protection of the present invention.
[0119] Obviously, the drawings described below are merely examples or embodiments of the present invention. Those skilled in the art can apply the present invention to other similar scenarios based on these drawings without inventive effort. Furthermore, it is understood that while the effort involved in such a development process may be complex and lengthy, for those skilled in the art related to the disclosure of the present invention, any design, manufacturing, or production changes based on the technical content disclosed in the present invention are merely conventional technical means and should not be construed as an inadequacy of the disclosure of the present invention.
[0120] Unless otherwise specified, all embodiments and optional embodiments of the present invention can be combined with each other to form new technical solutions.
[0121] Unless otherwise specified, all technical features and optional technical features of the present invention can be combined with each other to form a new technical solution.
[0122] Unless otherwise specified, all steps of the present invention may be performed sequentially or randomly, preferably sequentially. For example, the method includes steps (a) and (b), which means that the method may include steps (a) and (b) performed sequentially, or may include steps (b) and (a) performed sequentially. For example, the method may further include step (c), which means that step (c) may be added to the method in any order, for example, the method may include steps (a), (b) and (c), or may include steps (a), (c) and (b), or may include steps (c), (a) and (b), etc.
[0123] Unless otherwise specified, the terms "include" and "comprising" used in the present invention may be open-ended or closed-ended. For example, "include" and "comprising" may mean that other components not listed may also be included or that only the listed components are included.
[0124] Unless otherwise specified, the term "or" is inclusive in this disclosure. For example, the phrase "A or B" means "A, B, or both A and B." More specifically, the condition "A or B" is satisfied by any of the following conditions: A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); or both A and B are true (or exist).
[0125] In order to better understand the solutions of the embodiments of the present invention, some relevant terms and concepts that may be involved in the embodiments of the present invention are first introduced below.
[0126] (1) Artificial intelligence (AI), also known as intelligent machines or machine intelligence, refers to machines created by humans that can display intelligence. Generally, AI refers to the technology that displays human intelligence through ordinary computer programs.
[0127] (2) Machine learning (ML). Machine learning is the core of artificial intelligence. The theory of machine learning mainly involves designing and analyzing algorithms that allow computers to learn automatically. Machine learning algorithms are a type of algorithm that automatically analyzes data to obtain patterns and uses these patterns to predict unknown data. Therefore, the core of machine learning is data, algorithms (models), and computing power (computer computing power). The application areas of machine learning are very broad, including data mining, data classification, computer vision, natural language processing (NLP), biometric recognition, search engines, medical diagnosis, credit card fraud detection, securities market analysis, DNA sequencing, speech and handwriting recognition, strategic games, and robot applications. Machine learning is to design an algorithm model to process data and output the results that users want. Users can continuously tune the algorithm model to form more accurate data processing capabilities.
[0128] (3) Convolutional neural network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor consisting of a convolution layer and a subsampling layer, which can be regarded as a filter. A convolution layer refers to a neuron layer in a convolutional neural network that performs convolution processing on the input signal. In the convolution layer of a convolutional neural network, a neuron can only be connected to some neurons in the adjacent layer. A convolution layer usually contains several feature planes, and each feature plane can be composed of a number of rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weights here are convolution kernels. Shared weights can be understood as a way of extracting image information that is independent of position. The convolution kernel can be initialized in the form of a matrix of random size, and during the training process of the convolutional neural network, the convolution kernel can obtain reasonable weights through learning. In addition, the direct benefit of shared weights is to reduce the connections between the layers of the convolutional neural network, while reducing the risk of overfitting.
[0129] (5) RKNN (Rockchip Neural Network) is a deep learning model format designed by Rockchip for its own NPU. It is used to efficiently deploy neural networks on edge devices. It is deeply optimized for Rockchip NPU architectures (such as RK3399Pro and RK3588), improves inference speed, supports model conversion from frameworks such as TensorFlow and PyTorch (requires the RKNN-Toolkit tool chain), and is suitable for embedded devices.
[0130] In this embodiment, Figure 1As shown in FIG, a real-time target tracking method based on efficient reasoning of the RKNN model includes the following steps:
[0131] Step 1: Initialize the prediction model, obtain the image frame containing the target, crop the target area and perform template feature extraction based on the Siamese network;
[0132] The Siamese network is a two-branch network structure that uses shared weights to learn the similarity measure of input samples.
[0133] Step 2: Extract and fuse multi-scale features of the target area image;
[0134] Step 3: Match the template features and multi-scale fusion features based on template matching, calculate the similarity, and obtain the similarity score matrix;
[0135] Step 4: Obtain the preliminary predicted position of the target in the current frame based on the similarity score matrix;
[0136] Step 5: Combine the target's historical trajectory information to construct an objective function, calibrate the target's predicted position, and evaluate its reliability based on the confidence score;
[0137] Step 6: Optimize and adapt the prediction model based on the NPU of the RK platform and convert it into an RKNN model.
[0138] During the initialization phase, when the target tracking system is started or a new tracking task is received, the initial frame image and target frame information are acquired. An image frame containing the target is obtained from an image source (such as a camera or video file). If manual annotation is used, the user uses a graphical interface tool to select the target's initial position on the image and determine the target's bounding box (upper left and lower right corner coordinates). If automatic target detection is used, the system detects the target in the image based on a preset target detection algorithm and acquires its bounding box information. Based on the target bounding box, the target area is cropped from the original image to obtain a subwindow image. This subwindow will serve as the primary object for subsequent processing, extracting target feature information and performing template matching operations.
[0139] NPU (Neural Processing Unit) is a processor designed specifically for neural network computing. It focuses on parallel computing and energy efficiency, supports matrix multiplication and addition operation acceleration, and reduces power consumption several times compared to GPU.
[0140] Furthermore, the target area is cropped out, and the target bounding box is manually annotated using a graphical interface tool or automatically obtained using a target detection algorithm.
[0141] Furthermore, template features include geometric features and texture features.
[0142] When extracting template features, it is necessary to build a lightweight deep neural network and select a model based on the Siamese network architecture. The convolutional neural network (CNN) part of the model contains L convolutional layers, and the convolution kernel size of the lth (l<=L) convolutional layer is k l ×k l , the number of convolution kernels is n l , the step size is s l , then the size of the feature map output by the lth layer is AH l ,AW l Satisfies the formula:
[0143]
[0144] When building a lightweight deep neural network for template feature extraction, this formula is used to calculate the height and width of the output feature map of the lth convolution layer. It can also accurately derive the size of the output feature map after each convolution operation based on the size of the input feature map, the convolution kernel size, and the step size. This is crucial for network structure design and feature extraction process, and affects the subsequent effective capture of target features.
[0145] In this embodiment, the total number of convolutional layers is L=3. In the first convolutional layer, the input feature map size is H1=127, W1=127, and the convolution kernel size is k. l ×k l =4×4, the number of convolution kernels is n l =48, step length s l =16, the size of the output feature map of the first layer is calculated to be AH1=8, AW1=8, the total size of the output feature map of the first layer is 8×8×48, the total size of the output feature map of the second layer is 3×3×128, and the total size of the output feature map of the third layer is 1×1×256.
[0146] By optimizing the number and size of convolution kernels and adjusting the connection mode of network layers, the complexity of the model can be significantly reduced while maintaining the ability to effectively extract target features. The model parameter P can be approximately calculated as:
[0147]
[0148] The number of parameters in the first layer is 48×4 2 ×3=2304; the number of parameters in the second layer is 128×3 2 ×48=55296; the number of parameters in the third layer is 256×3 2 ×128=294912.
[0149] The cropped target sub-window image is input into the lightweight deep neural network. After forward propagation calculation, the initial features of the target are extracted. The feature extraction process can be expressed as: F T =M(I R ), expressed in the form of vectors or feature maps, which can represent key information such as the appearance and texture of the target.
[0150] In the model initialization stage, the target sub-window image I is cropped from the original image R Input into the lightweight deep neural network for template feature extraction, the extracted template feature F T It is stored in the memory as a reference for target matching and position prediction in subsequent frames. M represents a lightweight deep neural network model. The formula reflects the process of feature extraction of the target sub-window image through the network. The extracted template feature F T It will serve as an important reference benchmark for target matching and position prediction in subsequent frames.
[0151] Furthermore, step 2 includes the following steps:
[0152] Step 2.1: Using the pyramid feature extraction method, perform multi-scale cropping operations on the cropped target area image to form a multi-resolution image set;
[0153] Multi-scale feature extraction uses pyramid feature extraction method, for the sub-window image I where the target is located R , perform multi-scale cropping operations, according to the preset scale factor α={α1,α2,…,α m}, generate m sub-window images with different resolutions Its size satisfy Form a multi-resolution feature input set.
[0154] In this embodiment, a sub-window image I with a size of 127×127 is selected. R Perform multi-scale cropping operation, preset scale factor α={0.5,1}(ie m=2), according to the formula Here ω = 127, h = 127, i = 1, α1 = 0.5: (rounded), Get the subwindow image The size is 63×63. When i=2, α2=1: That is, the original sub-window image I R , with a size of 127×127. This forms a multi-resolution feature input set Similarly, for the 255×255 image in the figure, α is preset to {0.6, 1}, α1 = 0.6: get α2=1, keep the original size of 255×255
[0155] These sub-window images of different resolutions are input into the subsequent network (such as the convolution part in the figure). After convolution operations (such as layer processing with parameters such as c=48 and step size s=16), and then through pixel-by-pixel correlation matching, they can adapt to the feature extraction of targets of different sizes and improve the accuracy of capturing target features. For example, it can better handle targets of different scales in the image (such as vehicles) and ensure that target features can be effectively analyzed at multiple scales.
[0156] Through this multi-scale cropping, sub-window images at different resolutions can be obtained, providing rich input for the subsequent use of a lightweight feature pyramid network (FPN) to extract multi-scale deep features, and enhancing the algorithm's adaptability to target deformation and size changes.
[0157] Step 2.2: Extract features from multi-resolution images using a lightweight feature pyramid network (FPN).
[0158] A lightweight Feature Pyramid Network (FPN) is used to extract features from sub-window images at multiple resolutions. Through a top-down and lateral connection structure, FPN performs convolution operations on feature maps at different scales, extracting multi-scale deep features with rich semantic information. Lower-resolution feature maps capture the overall structural information of the target, while higher-resolution feature maps capture detailed features, enhancing the algorithm's adaptability to target deformation and size changes.
[0159] Step 2.3: Perform weighted fusion processing on the different scale features extracted from the multi-scale feature pyramid to obtain multi-scale fusion features.
[0160] The feature fusion method is to design a weighted fusion mechanism to extract different scale features from the multi-scale feature pyramid. Perform fusion processing. Assign a weight coefficient ω to each scale feature i The weight coefficient is determined based on the importance and confidence of the scale feature in the target representation. For example, the weight coefficient can be dynamically adjusted by analyzing the contribution of different scale features to target positioning and recognition.
[0161] Multiply each scale feature by its corresponding weight coefficient and then perform the accumulation operation to obtain a unified feature representation F fused , the formula is as follows:
[0162]
[0163] This fusion method can effectively integrate local detail features and global structural features. i The weight coefficient assigned to each scale feature is determined based on the importance and confidence of the scale feature in the target representation. By multiplying each scale feature with its corresponding weight coefficient and then accumulating them, it can effectively integrate local detail features and global structural features, making the final feature representation more representative and robust, providing more accurate information for subsequent target matching and position prediction.
[0164] Furthermore, step 2.2 includes: capturing the overall structural information of the target on a feature map with lower resolution through a lightweight feature pyramid network FPN; and obtaining the detailed features of the target on a feature map with higher resolution.
[0165] In the present invention, a higher resolution means that the number of pixels per frame is greater than or equal to 1920×1080, and a lower resolution means that the number of pixels per frame is less than 1920×1080.
[0166] Furthermore, step three includes: using a template matching-based strategy to match the multi-scale fusion features extracted from the current frame with the target template features, calculating the correlation between the two features in the feature space through a similarity measurement method, and obtaining a similarity score matrix.
[0167] In the process of target matching and position prediction, a template matching-based strategy is adopted to combine the multi-scale fusion features F extracted from the current frame into the fused and the stored target template feature F T Perform matching calculations and calculate the similarity between the two. The similarity measurement method can use efficient lightweight cross-correlation operations. By calculating the correlation between the template features and the current frame features in the feature space, a similarity score matrix S is obtained. The calculation formula is:
[0168] s ij =F fused (i,j)·F T ;
[0169] Each element in the similarity score matrix S represents the degree of similarity between the template feature and the feature at the corresponding position in the current frame. By calculating the similarity score matrix, the correlation between the template feature and the features at each position in the current frame can be measured in the feature space, thereby determining the possible position of the target in the current frame. The area with a higher similarity score is more likely to be the location of the target.
[0170] Furthermore, step four includes: performing non-maximum suppression processing on the similarity score matrix, removing areas with high local similarity but not global optimal, and extracting the candidate area corresponding to the highest score. The candidate area is the preliminary predicted position of the target in the current frame.
[0171] Let the threshold of non-maximum suppression be t NMS , then the candidate region after non-maximum suppression (NMS) processing:
[0172]
[0173] In this embodiment, the selected NMS algorithm is: the candidate area meets the conditional formula of the selected NMS algorithm:
[0174] Furthermore, step five includes the following steps:
[0175] Step 5.1: Based on the preliminary predicted position of the target, further output the center position and size of the target;
[0176] Central Location:
[0177]
[0178] size:
[0179] S pred =(ω pred ,h pred );
[0180]
[0181] In this embodiment, in a certain frame image, three candidate regions are obtained through correlation calculation, and their coordinates and similarity scores are: region B1 = (10, 10, 50, 50), similarity score is 0.9; region B2 = (20, 20, 60, 60), similarity score is 0.85; region B3 = (100, 100, 140, 140), similarity score is 0.8. Set the non-maximum suppression (NMS) threshold t NMS =0.5. Calculate the overlap ratio of B1 and B2 (such as the intersection over union (IoU)). Assuming its value is greater than 0.5, B2 is suppressed because B1 has a higher score. Calculate the overlap ratio of B1 and B3. Assuming its value is less than 0.5, B3 is retained. Finally, the candidate region B after NMS processing is pred =(10,10,50,50), which is the initial predicted position of the target in the current frame.
[0182] Then calculate the center position and size of the target:
[0183] Central Location:
[0184] Therefore, C pred =(30,30).
[0185] size:
[0186] Therefore, S pred =(40,40).
[0187] Step 5.2: Integrate the historical trajectory information of the target, smoothly adjust the target predicted position of the current frame, and correct the predicted position;
[0188] The target's historical trajectory information is integrated to smoothly adjust the target's predicted position in the current frame. This historical trajectory information includes the target's position, velocity, acceleration, and other motion information from past frames. By analyzing the target's historical motion patterns and predicting its reasonable position range in the current frame, the network's predicted position is corrected, improving the accuracy and stability of the target's position prediction.
[0189] Step 5.3: Based on the nonlinear optimization method, the historical trajectory information of the target is combined with the feature information of the current frame to construct the objective function;
[0190] Using an optimized nonlinear update method, the historical trajectory information of the target is combined with the feature information of the current frame to construct the objective function E, which is formulated as follows:
[0191]
[0192] In this embodiment, the target has been tracked for N=3 frames, and the current frame is the 3rd frame. We have the observation feature F of each frame obs and the predicted feature F pred as follows:
[0193] Frame 1: Observed feature F obs (1) = [10, 20], prediction feature F pred (1) = [12, 18];
[0194] Frame 2: Observed feature F obs (2) = [15, 25], prediction feature F pred (2) = [16, 23];
[0195] Frame 3: Observed feature F obs (3) = [20, 30], prediction feature F pred (3) = [22, 28];
[0196] According to the objective function Here the features are vectors, and the sum of the squares of the differences of each element is calculated:
[0197] Frame 1: (F obs (1)-F pred (1) 2 =(10-12) 2 +(20-18) 2 =4+4=8;
[0198] Frame 2: (F obs (2)-F pred (2) 2 =(15-16) 2 +(25-23) 2 =1+4=5;
[0199] Frame 3: (F obs (3)-F pred (3) 2 =(20-22) 2 +(30-28) 2 =4+4=8;
[0200] Then the objective function E=8+5+8=21.
[0201] In this embodiment, the optimized nonlinear update method used is a least squares optimization algorithm based on the Ceres optimization library.
[0202] Step 5.4: Obtain the optimal solution of the objective function and further correct the target predicted position;
[0203] By solving the optimal solution of the objective function, the target position is further corrected so that the target position update is more consistent with the actual motion situation and the jitter and drift during the tracking process are reduced.
[0204] The least squares optimization algorithm based on the Ceres optimization library is used to find the optimal solution for the objective function E. In actual applications, the Ceres optimization library will iteratively adjust the parameters based on the objective function and initial parameters we provide, so that the objective function E is minimized. In this embodiment, after the Ceres optimization, a new set of parameters is obtained, and the target prediction position is further corrected based on these parameters. The center of the target position initially predicted in a certain frame is After correction, it becomes
[0205] Step 5.5: Calculate the confidence score of the current predicted position. If the confidence score is higher than the threshold, it means that the result is more reliable. If it is lower than the threshold, reinitialize the model for prediction.
[0206] The scoring mechanism of this invention evaluates the stability and reliability of target tracking based on the confidence score output by the network. The confidence score is a numerical indicator generated by the network while predicting the target position. It reflects the network's confidence in the current prediction result. The formula is as follows:
[0207] s conf =ω m ×m+ω c ×c+ω f ×f;
[0208] Assume that the weight coefficients are ω m =0.3,ω c =0.4,ω f =0.3, and: m represents the smoothness score of the target motion, where m=0.8; c represents the matching score of the target feature, where c=0.7; f represents the clarity score of the current frame image, where f=0.6.
[0209] According to the confidence scoring formula s conf =ω m ×m+ω c ×c+ω f ×f, we can get:
[0210] s conf =0.3×0.8+0.4×0.7+0.3×0.6=0.7;
[0211] In this embodiment, the confidence threshold is set to 0.6, because s conf =0.7>0.6, indicating that the current prediction result is relatively reliable and there is no need to reinitialize the model for prediction. conf If the value is lower than the threshold, the model needs to be reinitialized and the target prediction needs to be made again.
[0212] Furthermore, the historical trajectory information includes the target's motion state information in the past several frames, including position, speed, and acceleration.
[0213] Furthermore, step six includes converting the model into ONNX format. During the model deployment process, the multi-platform hardware acceleration technology NPU is used to optimize and accelerate the model reasoning, adjust the model parameters and calculation process, and convert the model into RKNN format.
[0214] The model is optimized and, after training, converted to the ONNX format, which offers excellent cross-platform performance and efficient inference. During model deployment, the multi-platform hardware acceleration technology, NPU, is used to accelerate model inference. Based on the NPU's hardware architecture, the model is optimized for specific purposes, such as optimizing data layout and adjusting calculation order, fully leveraging the NPU's computing advantages.
[0215] In embedded devices, a memory sharing strategy is employed to enable different modules or frames to share memory space for model parameters and intermediate calculation results, reducing repeated memory allocation and deallocation, and lowering memory usage. Furthermore, a batch inference strategy is employed to combine data from multiple image frames for batch processing, fully utilizing hardware computing resources, improving computational efficiency, and significantly reducing inference time.
[0216] The NPU platform is adapted to the Rockchip RK3588 platform, leveraging its integrated NPU for efficient inference. The model is optimized and adapted to the NPU characteristics of the RK3588 platform, adjusting the model's parameters and computational flow, and converting the model to the RKNN format to achieve optimal inference performance. This optimization achieves real-time performance of up to 120 FPS on this platform, meeting the requirements of real-time target tracking.
[0217] In this embodiment, the structural diagram of the template feature extraction network is as follows: Figure 2 As shown in the figure, the image input part consists of two branches, one for the template image and the other for the search image. The template image has a size of 127×127, and the other for 255×255. The template image represents the initial region of the target being tracked, while the search image represents the region in the current frame that may contain the target. The images from both branches are fed into a weight-shared convolutional neural network for feature extraction. The CNN consists of multiple convolutional layers, which gradually extract high-level features from the image. The output feature map of the template branch has a size of 127×127×48, while the output feature map of the search branch has a size of 255×255×48. The extracted template features are fused with the search features. A pixel-by-pixel correlation operation is performed on the features extracted from the template and search branches, and their similarity is calculated to generate a correlation matrix, which is used to determine the location of the target in the search image. A classification branch determines whether a candidate region is a target. After a series of convolution and pooling operations on the input features, a probability value indicating the presence or absence of the target is output. The regression branch accurately predicts the target's location and size. The regression branch adjusts the coordinates of the candidate region to obtain a more accurate target bounding box.
[0218] The schematic diagram of manually marking the target boundary in this embodiment is as follows Figure 3 As shown in the blue box, the white vehicle in the figure is the tracking target of this embodiment;
[0219] The schematic diagram of the processing time and frame rate of each frame of the tracking video by the model in the embodiment of the present invention is as follows: Figure 4 As shown, FPS is the frame rate;
[0220] The interface diagram of the target tracking embodiment of the present invention is as follows: Figure 5 As shown, Figure 5 The red box in the middle indicates the target position predicted by the tracking method of the present invention in the current frame. This is the specific position of the target in the current frame obtained by the algorithm based on comprehensive calculations such as template matching, multi-scale feature extraction and fusion, and historical trajectory information. The green box indicates the actual position of the target (i.e., the annotation box), which is the accurate target position determined by manually annotating the video or image sequence. It is used to compare with the red box in the evaluation stage to measure the accuracy of the tracking algorithm.
[0221] The IOU values of some frames and the tracking accuracy of the embodiment of the present invention are shown in the following figure: Figure 6 As shown, the IOU value is Figure 5 The overlap degree of the red box and the green box, Frame is the current frame number, Accurate is whether the tracking is accurate, and the judgment condition is that the IOU value > 0.5 is judged to be accurate, that is, the value of Accurate is Yes;
[0222] The total number of frames of the tracking video and the total tracking accuracy of the embodiment of the present invention are shown in the figure below: Figure 7 As shown, specifically:
[0223] Total frames: indicates the total number of frames in the video sequence used to evaluate the tracking algorithm, which is 247 frames in this embodiment. This means that the algorithm performs tracking processing on every frame in the entire video sequence.
[0224] Accurate frames: This indicates the number of frames in the total number of frames where the tracking results meet the accuracy criterion (i.e., IOU > 0.5). As can be seen from the figure, the number of accurate frames is also 247, which is equal to the total number of frames.
[0225] Accuracy: This represents the ratio of accurate frames to the total number of frames, expressed as a percentage. An accuracy of 100% here means that during the tracking process of all 247 frames, the tracking results of each frame met the preset accuracy standard (IOU>0.5). This fully demonstrates that the tracking method of the present invention has extremely high stability and accuracy throughout the entire video sequence, and can reliably locate the target in every frame, providing a strong technical support for practical applications (such as security monitoring and intelligent transportation).
[0226] It should be noted that the present invention is not limited to the above-mentioned embodiments. The above-mentioned embodiments are merely examples, and any embodiments having substantially the same structure and effect as the technical concept within the scope of the technical solution of the present invention are all included in the technical scope of the present invention. In addition, without departing from the scope of the present invention, other embodiments that can be conceived by those skilled in the art and that combine some of the constituent elements in the embodiments are also included in the scope of the present invention.
Claims
1. A real-time target tracking method based on efficient reasoning of RKNN model, characterized in that: The following steps are involved: Step 1: Initialize the prediction model, obtain the image frame containing the target, crop the target area and perform template feature extraction based on the Siamese network; Step 2: Extract and fuse multi-scale features of the target area image; Step 3: Match the template features and multi-scale fusion features based on template matching, calculate the similarity, and obtain the similarity score matrix; Step 4: Obtain the preliminary predicted position of the target in the current frame based on the similarity score matrix; Step 5: Combine the target's historical trajectory information to construct an objective function, calibrate the target's predicted position, and evaluate its reliability based on the confidence score; Step 6: Optimize and adapt the prediction model based on the NPU of the RK platform and convert it into an RKNN model.
2. A real-time target tracking method based on efficient reasoning of RKNN model according to claim 1, characterized in that: The target area is cropped by manually marking the target bounding box through a graphical interface tool or automatically obtaining the target bounding box through a target detection algorithm.
3. The real-time target tracking method based on efficient reasoning of the RKNN model according to claim 1 is characterized in that: The template features include geometric features and texture features.
4. The real-time target tracking method based on efficient reasoning of the RKNN model according to claim 1, characterized in that: The second step comprises the following steps: Step 2.1: Using the pyramid feature extraction method, perform multi-scale cropping operations on the cropped target area image to form a multi-resolution image set; Step 2.2: Extract features from multi-resolution images using a lightweight feature pyramid network (FPN). Step 2.3: Perform weighted fusion processing on the different scale features extracted from the multi-scale feature pyramid to obtain multi-scale fusion features.
5. The real-time target tracking method based on efficient reasoning of the RKNN model according to claim 4 is characterized in that: The step 2.2 includes: capturing the overall structural information of the target on a feature map with a lower resolution through a lightweight feature pyramid network (FPN); and obtaining the detailed features of the target on a feature map with a higher resolution.
6. The real-time target tracking method based on efficient reasoning of the RKNN model according to claim 1, characterized in that: The step three includes: using a template matching-based strategy to match the multi-scale fusion features extracted from the current frame with the target template features, and calculating the correlation between the two features in the feature space using a similarity measurement method to obtain a similarity score matrix.
7. The real-time target tracking method based on efficient reasoning of RKNN model according to claim 1, characterized in that: The fourth step includes: performing non-maximum suppression processing on the similarity score matrix, removing areas with high local similarity but not global optimal, and extracting the candidate area corresponding to the highest score. The candidate area is the preliminary predicted position of the target in the current frame.
8. The real-time target tracking method based on efficient reasoning of RKNN model according to claim 1 is characterized in that: The step five comprises the following steps: Step 5.1: Based on the preliminary predicted position of the target, further output the center position and size of the target; Step 5.2: Integrate the historical trajectory information of the target, smoothly adjust the target predicted position of the current frame, and correct the predicted position; Step 5.3: Based on the nonlinear optimization method, the historical trajectory information of the target is combined with the feature information of the current frame to construct the objective function; Step 5.4: Obtain the optimal solution of the objective function and further correct the target predicted position; Step 5.5: Calculate the confidence score of the current predicted position. If the confidence score is higher than the threshold, it means that the result is more reliable. If it is lower than the threshold, reinitialize the model for prediction.
9. The real-time target tracking method based on efficient reasoning of the RKNN model according to claim 8, characterized in that: The historical trajectory information includes the target's motion state information in the past several frames, including position, speed, and acceleration.
10. The real-time target tracking method based on efficient reasoning of RKNN model according to claim 1, characterized in that: The step six includes converting the model into ONNX format. During the model deployment process, the multi-platform hardware acceleration technology NPU is used to optimize and accelerate the model reasoning, adjust the model parameters and calculation process, and convert the model into RKNN format.