A single-object tracking method applicable to the edge side
By using feature fusion networks and lightweight MobileNetV4 backbone networks on end-side devices, combining channel attention and spatial pyramid pooling layer, the shortcomings of existing tracking algorithms in low-light and end-side devices are solved, and stable single-target tracking and real-time effects are achieved.
Patent Information
- Application Number
- CN202510479961.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Existing tracking algorithms are susceptible to interference from factors such as lighting changes, target occlusion, and scale changes in complex scenarios, especially in low-light situations, and the low-computing equipment on the opposite side is unfriendly, making real-time tracking impossible.
The y-components of visible and infrared tiles are weighted through the feature fusion network, combined with the lightweight MobileNetV4 backbone network and heavy parameter structure, a channel attention mechanism and spatial pyramid pooling layer are introduced, the tracking model is optimized to adapt to the end-side devices, and target loss in low light is treated through Kalman filtering.
It realizes stable single-target tracking under low light conditions, reduces the computing power demand of the opposite side equipment, improves the feature extraction capability, and ensures real-time tracking effect.
Smart Images

Figure CN119991760B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision processing, and particularly to a single-object tracking method applicable to the edge side. Background Art
[0002] Traditional tracking algorithms (such as KCF) are vulnerable to factors such as illumination changes, target occlusion, and scale changes in complex scenarios. Deep learning Siamese network tracking algorithms (such as SiamFC) also have good tracking effects under the influence of factors such as target occlusion and scale changes, but they perform poorly under illumination changes, especially in low-light conditions, and are not friendly to low-computing-power devices on the edge side. The edge side refers to the device side, usually devices directly used by users, such as smartphones, tablets, Internet of Things (IoT) devices, etc. However, these devices usually have limited computing power and energy consumption, and mostly perform data transmission and processing through wireless connection networks. Due to their own computing power and other defects, they cannot achieve real-time tracking in actual operation. Summary of the Invention
[0003] The purpose of the present invention is to overcome the problem that the current tracking algorithms are not friendly to the devices on the edge side, and provide a single-object tracking method applicable to the edge side.
[0004] The purpose of the present invention is achieved by the following technical solutions:
[0005] A single-object tracking method applicable to the edge side, the steps include:
[0006] Step 1. Obtain the target position through target detection, intercept the regional patch of the target position, and perform weighted fusion on the y components of the visible light and infrared of the registered patch coordinates through a feature fusion network to obtain a fused search region patch, strengthening the feature representation of low light (illumination intensity less than 50 lux);
[0007] Step 2. Input the patch into the tracking model. The tracking model selects nanotrack_v2 in the baseline. The backbone network of the tracking model introduces the lightweight model MobileNetV4, and introduces a reparameterized structure to replace the original convolutional layer; obtain a new target position and confidence through the output of the tracking model;
[0008] Step 3. Set the range of the confidence p:
[0009] a) High threshold (p > 0.95): Continuously track the target;
[0010] b) Medium threshold (0.8 ≤ p ≤ 0.95): Update the feature vector of the tracking target;
[0011] c) Low threshold (p < 0.8): Determine that the target is lost, and use Kalman filtering to predict the target position in the next frame of the target for matching the target position in the next frame.
[0012] Furthermore, in step 2, a channel attention mechanism and a spatial pyramid pooling layer are introduced into the tracking model, and the tracking model is trained.
[0013] Furthermore, in step 1, the target position is obtained through target detection to obtain the original target position coordinates (cx, cy, h, w), which are the center point coordinates and the width and height of the target respectively. The width and height z_l of the target area are obtained according to the following formula:
[0014] ;
[0015] where w is the width of the target position coordinates and h is the height of the target position coordinates.
[0016] Furthermore, the target area tile (cx, cy, z_l, z_l) is intercepted, and the tile coordinates of the registered infrared image are obtained. After interception, it is adjusted (resized) to (z_l, z_l). Then, the y component of the visible light and infrared target area tiles is weighted and fused through a feature fusion network to obtain the fused target area yuv tile, and the yuv tile is input into the tracking model.
[0017] Furthermore, after the yuv tile is input into the tracking model, in the second frame after initialization, (cx, cy, 2*z_l, 2*z_l) is intercepted to obtain the tile coordinates of the registered infrared image. After interception, it is adjusted to (2*z_l, 2*z_l). Then, the y component of the visible light and infrared target area tiles is weighted and fused through a feature fusion network to obtain the fused target area yuv tile. The yuv tile is adjusted and then input into the tracking model, and the tracking model is used for inference to output the new target position (cx_, cy_, h_, w_) and the confidence level;
[0018] where (cx_, cy_, h_, w_) are the coordinates of the target position in the updated next frame.
[0019] Furthermore, after obtaining the confidence level, option selection for target tracking is performed according to the confidence level:
[0020] When the confidence level is higher than the high threshold (p > 0.95), continuous tracking of the target is performed: (cx_, cy_, h_, w_) is updated to replace (cx, cy, h, w), and the new (cx, cy, 2*z_l, 2*z_l) is calculated, and step 2 is continuously executed;
[0021] When the confidence level is at the medium threshold (0.8 ≤ p ≤ 0.95), update the feature vector of the tracked target: update and replace (cx, cy, h, w) with (cx_, cy_, h_, w_), complete the tracking initialization with the new coordinates, and continuously execute step 2 as described above;
[0022] When the confidence level is at the low threshold (p < 0.8): determine that the target is lost, use Kalman filtering to predict the target position in the next frame, obtain the coordinates (cx_, cy_, h_, w_) of the target position in the next frame, then perform the matching of the target coordinate position in the next frame, update and replace (cx, cy, h, w) with (cx_, cy_, h_, w_), and continuously execute step 2 as described above;
[0023] When the number of frames in the target lost state reaches the lost frame threshold, exit the tracking state.
[0024] Furthermore, during the training of the tracking model, data is collected for scenes with low light intensity less than 50 lux under the same optical axis and the same field of view, the tiles are cropped and adjusted according to the registered infrared and visible light data, and data augmentation is performed through dynamic scale.
[0025] Furthermore, during the training of the tracking model, a compensation mechanism for the lack of visible light information in low light environments is established. The weight of the compensation mechanism is calculated as the image entropy of the registered infrared image and the pixel average value of the visible light, and target data is made according to the compensation mechanism through infrared and visible light data. The weights are as follows:
[0026] ;
[0027] where q is the fusion weight of the visible light y component of the compensation mechanism, i entropy is the image entropy of the registered infrared, and v average is the pixel average value of the visible light.
[0028] Furthermore, during the training of the tracking model, the pixel matching is supervised through the loss function. Specifically, by calculating the pixel difference between the model output and the target image, the model is supervised to learn the pixel-level matching of the target features to ensure that the output is consistent with the true label. The loss function is:
[0029] where A(i, j) is the pixel value of the tracking model output image at the (i, j) coordinate, B(i, j) is the pixel value of the target image at the (i, j) coordinate, and m, n are the width and height of the model output and the target image.
[0030] Furthermore, during the training of the tracking model, the tracking model is fine-tuned through the fusion dataset combined with the existing tracking dataset.
[0031] The beneficial effects of the present invention are as follows:
[0032] (1) By combining the lightweight MobileNetV4 backbone network with a reparameterization structure, channel attention, and dynamic multi-scale feature fusion, the effects of model lightweighting and improved feature extraction efficiency are achieved;
[0033] (2) After fusing the data of infrared and visible light, the features of the tracking target are strengthened, and the tracking model is made more lightweight, improving the ability to extract features, thereby enhancing the stable tracking in low-light conditions and being more friendly to the computing power requirements on the edge side. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flowchart of the steps of a single-object tracking method applicable to the edge side;
[0035] Figure 2 It is a schematic diagram of the lightweight tracking network architecture and feature processing based on MobileNetV4. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0037] Embodiment 1
[0038] Refer to Figure 1 , and a single-object tracking method applicable to the edge side is provided, and the steps include:
[0039] Step 1. Obtain the target position through object detection, intercept the regional tile of the target position, and perform weighted fusion on the y components of the registered tile coordinates of visible light and infrared through a feature fusion network to obtain the fused search region tile, so as to strengthen the feature representation of low light;
[0040] Step 2. Input the tile into the tracking model. The tracking model selects nanotrack_v2 in the baseline. The backbone network of the tracking model introduces the lightweight model MobileNetV4, and a reparameterization structure is introduced to replace the original convolutional layer; the new target position and confidence are obtained through the output of the tracking model;
[0041] Step 3. Set the range of the confidence p:
[0042] a) High threshold (p > 0.95): Continuously track the target;
[0043] b) Medium threshold (0.8 ≤ p ≤ 0.95): Update the feature vector of the tracking target;
[0044] c) Low threshold (p < 0.8): Determine that the target is lost, and use Kalman filtering to predict the target position in the next frame to perform the matching of the target position in the next frame.
[0045] Based on the target obtained by the customer's click (or target detection or other means), obtain the original position coordinates (cx, cy, h, w) of the target, which are the center point coordinates and the width and height of the target respectively. Obtain the width and height z_l of the target area according to the following formula:
[0046] ;
[0047] where w is the width of the target position coordinates and h is the height of the target position coordinates.
[0048] Crop the target area patch with coordinates (cx, cy, z_l, z_l), and obtain the patch coordinates of the registered infrared image. For example, at this time, the visible light resolution is 1080*1920 and the infrared resolution is 512*640. The registered coordinates will obtain the corresponding mapping. After cropping the coordinates, resize to (z_l, z_l), and then use the feature fusion network to perform weighted fusion on the y components of the visible light and infrared target area patches to obtain the fused target area yuv patch. Resize this patch to (127*127) and send it to input0 of the updated tracking model.
[0049] Among them, the tracking model is trained in the following way:
[0050] Collect data on the scene under low-light conditions according to the same optical axis and the same field of view, and crop the registered infrared and visible light data pairs into patches of 127*127 to 255*255. Here, dynamic scale can also be used for data augmentation, such as cropping a 100*100 patch and then scaling it up proportionally to 150*150.
[0051] Establish a compensation mechanism for the lack of visible light information in the low-illumination environment. The weight of the compensation mechanism is calculated from the image entropy of the registered infrared image and the pixel average value of the visible light, and make target data through the infrared and visible light data pairs according to the compensation mechanism. The weights are as follows:
[0052] ;
[0053] where q is the fusion weight of the visible light y component of the compensation mechanism, i entropy is the image entropy of the registered infrared, and v average is the pixel average value of the visible light.
[0054] Then, the tracking model is trained through a loss function. The loss function supervises pixel matching, specifically by calculating the pixel difference between the model output and the target image, and supervising the model to learn pixel-level matching of target features to ensure that the output is consistent with the true label. The loss function is as follows:
[0055] ;
[0056] where A(i, j) is the pixel value of the model output image at the (i, j) coordinate, B(i, j) is the pixel value of the target image at the (i, j) coordinate, and m, n are the width and height of the model output and the target image respectively.
[0057] Then, the tracking model is fine-tuned based on the data fusion dataset combined with the existing tracking dataset. Stable tracking in low-light scenarios is achieved through multi-modal data fusion and lightweight network design: The input module receives the registered visible light and the y-component tile of the infrared image, and generates a fused feature map through a feature fusion network (dynamically calculating the weight q based on the infrared image entropy and the visible light mean); The tracking model uses MobileNetV4 as the backbone network, combines a reparameterized structure, a channel attention mechanism, and a spatial pyramid pooling layer with dynamic weights to output classification confidence and regression coordinates; The multi-threshold decision module controls the target state according to the confidence (high threshold directly updates, medium threshold re-initializes, low threshold triggers Kalman filter prediction); Data processing includes the calculation of the target area (z_l formula) during initialization and the expansion of the search area (2×z_l) during tracking. During the training phase, the model performance is optimized through data augmentation (dynamic cropping), a compensation mechanism, and a specific loss function (combining pixel error and logarithmic adjustment term), and finally, lightweight, real-time, and stable single-object tracking in low-light scenarios is achieved.
[0058] Refer to Figure 2 , the processing flow of the tracking network based on MobileNetV4: The input is two fused feature images with sizes of 127×127 and 255×255 respectively. After the data is input, it is processed by a network improved based on MobileNetV4, and features with an output channel number of 48 (marked "c = 48") and a stride of 16 (marked "s = 16") are generated, and feature maps with sizes of 8 and 16 are generated. Subsequently, the features are fused through a "PointwiseCorrelation" operation, and the fused features are respectively input into the classification (cls) and regression (reg) branches: The classification branch processes features with a channel number of 128 (marked "c = 128"), and finally outputs a result with a size of 16; The regression branch processes features with a channel number of 96 (marked "c = 96"), and also outputs a result with a size of 16, realizing the dual-task output of target classification and position regression.
[0059] Continuously track the target through the trained tracking model:
[0060] In the second frame after initialization, intercept the coordinates (cx, cy, 2*z_l, 2*z_l), and obtain the tile coordinates of the registered infrared image. After interception, resize it to the coordinates (2*z_l, 2*z_l). Then, use the feature fusion network to perform weighted fusion on the y component of the visible light and infrared target area tiles to obtain the fused target area yuv tile. Resize this tile to (255*255) and send it to input1 of the updated tracking model (the update method is the same as input0), and use the model for inference to obtain the target position coordinates (cx_, cy_, h_, w_) of the new next frame, as well as the confidence level.
[0061] Among them, (cx_, cy_, h_, w_) are the coordinates of the target position in the updated next frame.
[0062] When the confidence level is higher than the high threshold (p > 0.95), continuously track the target: update and replace (cx, cy, h, w) with the new target position coordinates (cx_, cy_, h_, w_), calculate the new coordinates (cx, cy, 2*z_l, 2*z_l), and continuously execute step 2 above.
[0063] When the confidence level is the medium threshold (0.8 ≤ p ≤ 0.95), update the feature vector of the tracking target: update and replace (cx, cy, h, w) with (cx_, cy_, h_, w_), complete the tracking initialization with the new coordinates, and continuously execute step 2 above.
[0064] When the confidence level is the low threshold (p < 0.8): determine that the target is lost, use Kalman filtering to predict the target position in the next frame, obtain the coordinates (cx_, cy_, h_, w_) of the target position in the next frame, then perform matching of the target coordinate position in the next frame, update and replace (cx, cy, h, w) with (cx_, cy_, h_, w_), and continuously execute step 2 above.
[0065] Continuously track the target through the above steps. When the number of frames in the target lost state reaches the lost frame threshold, exit the tracking state.
[0066] Embodiment 2
[0067] Step 1. Spatiotemporal feature fusion initialization: Target acquisition: Obtain the target center point (cx, cy) and width and height (w, h) through the improved CenterNet target detection algorithm, and record the timestamp of the initial frame. Area calculation is based on the formula:
[0068] ;
[0069] This is used to calculate the size of the target area. Spatiotemporal feature construction: intercept the (cx, cy, z_l, z_l) area of the visible light and infrared images, and resize to (z_l, z_l). Perform optical flow calculation on the multimodal data (YUV+infrared) of the current frame and the previous two frames to generate an optical flow field map containing motion information. Channel-join the optical flow field and the multimodal feature map to form a spatiotemporal fusion feature (127×127×(3+1+2)), and input it into the tracking model input0 to complete initialization.
[0070] Step 2. Tracking model improvement: The backbone network uses MobileNetV4-DeformableConv, and deformable convolution is introduced in the convolution layer to adaptively capture target deformation. The Transformer Encoder module is embedded to perform global dependency modeling on spatiotemporal fusion features and enhance the use of temporal information. Dynamic template management: Maintain a template library to store the target feature templates of the last 5 frames. When inferring each frame, multiple templates are dynamically selected for weighted fusion according to the confidence and motion speed of the current frame to generate the final matching template. Feature pyramid matching: Extract multi-scale feature pyramids for the 255×255 search area. Through the cross-scale attention mechanism, the template features are matched with the search area features at multiple scales, and the target position (cx_, cy_, h_, w_) and confidence p are output.
[0071] Step 3. Adaptive threshold decision and template update confidence dynamic threshold:
[0072] High confidence (p>0.95): Directly update the target coordinates and add the current frame features to the template library (replacing the oldest template);
[0073] Medium confidence (0.8≤p≤0.95): triggers template weight update, and adjusts the weight of each template in the template library according to the current matching result. Use online hard example mining (OHEM) to select samples with difficult matching to fine-tune the model;
[0074] Low confidence (p<0.8): Enable dual template matching: Use the best template in history and the current predicted template for secondary matching. Combine IMU sensor data (such as acceleration and angular velocity) to correct the Kalman filter prediction frame and improve motion prediction accuracy.
[0075] Step 4. Adversarial training and self-supervised learning optimization:
[0076] Data Augmentation Upgrade: Use StyleGAN2 to generate virtual multi-modal data in low-light scenarios to enhance data diversity. Introduce CutMix and MixUp techniques to improve the model's robustness to occlusion and blur. Self-supervised pre-training pre-trains the model on unlabeled data and learns the general representation of multi-modal features through Masked Image Modeling (MIM). Maximize the feature similarity of the same target in different modalities and minimize the feature similarity of different targets. The loss function can be further expressed as:
[0077] ;
[0078] where, is the classification loss, which is used to supervise the model's classification ability of the target and the background, ensuring that the model can accurately distinguish whether the current search area contains the target; is the regression loss, which is used to supervise the prediction accuracy of the model for the target position. By calculating the coordinate differences (such as the center point coordinates, width, and height) between the predicted box and the ground truth box, the accuracy of target localization is optimized; is the optical flow prediction loss, is the contrastive loss, which is used to enhance the feature discrimination between the target and the background.
[0079] The above solution combines the optical flow field and multi-modal data to strengthen the target representation in dynamic scenarios; adapts to the target appearance changes through the template library and weighted fusion. Introduce IMU data through multiple sensors to improve the motion prediction accuracy and enhance the environmental adaptability of edge devices. And through adversarial self-supervised training, use the generative model and self-supervised learning to solve the problem of insufficient low-light data and improve the model's generalization ability. Achieve 32fps real-time tracking on edge devices (such as RK3588), improve the tracking accuracy in low-light, fast-moving, and occluded scenarios, and significantly enhance the model's anti-interference ability.
[0080] By combining the lightweight MobileNetV4 backbone network with the reparameterization structure, channel attention, and dynamic multi-scale feature fusion, the effect of model lightweight and improved feature extraction efficiency is achieved; after fusing the infrared and visible light data, the features of the tracking target are strengthened, and the tracking model is made more lightweight, improving the feature extraction ability, thus enhancing the stable tracking in low-light conditions and being more friendly to the edge computing power requirements.
[0081] The above is only the preferred implementation manner of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And the changes and alterations made by those skilled in the art that do not depart from the spirit and scope of the present invention should all be within the protection scope of the appended claims of the present invention.
Claims
1. A single target tracking method applicable to a terminal side, characterized in that the steps include: Step 1. Obtain the target position through target detection, intercept the area block of the target position, and perform weighted fusion of the visible light and infrared y components of the aligned block coordinates through the feature fusion network to obtain the fused search area block; Step 2. Input the image block into the tracking model. Select nanotrack_v2 in the baseline as the tracking model. The lightweight model MobileNetV4 is introduced into the backbone network of the tracking model, and the re-parameter structure is introduced to replace the original convolutional layer. The new target position and confidence are obtained through the output of the tracking model. Step 3. Set the range of confidence level p: a) High threshold (p>0.95): continuous tracking of the target; b) Medium threshold (0.8≤p≤0.95): Update the feature vector of the tracked target; c) Low threshold (p<0.8): The target is judged as lost, and the Kalman filter is used to predict the target position in the next frame to match the target position in the next frame; In the step 2, a channel attention mechanism and a spatial pyramid pooling layer are introduced into the tracking model, and the tracking model is trained; When training the tracking model, data is collected for scenes in low-light conditions with an illumination intensity of less than 50 lux on the same optical axis and field of view. The image blocks are cropped and adjusted based on the infrared and visible light data obtained after registration, and data enhancement is performed through dynamic scaling. In the training of the tracking model, a compensation mechanism for the missing visible light information is established. The weight of the compensation mechanism is to calculate the image entropy of the registered infrared image and the pixel average value of the visible light. The target data is generated by the infrared and visible light data according to the compensation mechanism. The weights are as follows: ; Among them, q is the fusion weight of the visible light y component of the compensation mechanism, i entropy is the entropy of the infrared image after registration, v average is the average value of visible light pixels.
2. A single target tracking method applicable to a terminal side according to claim 1, characterized in that: In step 1, the target position is obtained through target detection, and the original position coordinates of the target (cx, cy, h, w) are obtained, which are the center point coordinates and width and height of the target respectively. The width and height z_l of the target area are obtained according to the following formula: ; Among them, w is the width of the target position coordinates, and h is the height of the target position coordinates.
3. The single target tracking method applicable to the terminal side according to claim 2, characterized in that: The (cx, cy, z_l, z_l) target area tile is intercepted and the tile coordinates of the registered infrared image are obtained. After interception, it is adjusted to (z_l, z_l). Then, the y components of the visible light and infrared target area tiles are weightedly fused through the feature fusion network to obtain the fused target area YUV tile, and the YUV tile is input into the tracking model.
4. The single target tracking method applicable to the terminal side according to claim 3, characterized in that: After feeding the YUV patch into the tracking model, intercept (cx, cy, 2 z_l, 2 z_l), get the block coordinates of the registered infrared image, cut and adjust to (2 z_l, 2 z_l), and then use the feature fusion network to weightedly fuse the y components of the visible light and infrared target area tiles to obtain the fused yuv tile of the target area, adjust the yuv tile and input it into the tracking model, and use the tracking model for reasoning to output the new target position (cx_, cy_, h_, w_) and confidence; Among them, (cx_, cy_, h_, w_) are the coordinates of the target position of the next frame to be updated.
5. The single target tracking method applicable to the terminal side according to claim 4, characterized in that: After obtaining the confidence level, select the option to track the target based on the confidence level: When the confidence is above a high threshold (p>0.95), the target is continuously tracked: (cx_, cy_, h_, w_) is updated to replace (cx, cy, h, w), and a new (cx, cy, 2 z_l, 2 z_1), and continue to execute step 2; When the confidence is the medium threshold (0.8≤p≤0.95), update the feature vector of the tracked target: update (cx_, cy_, h_, w_) to replace (cx, cy, h, w), complete the tracking initialization with the new coordinates, and continue to execute step 2; When the confidence is a low threshold (p<0.8): the target is determined to be lost, and the Kalman filter is used to predict the target position of the next frame, and the coordinates of the target position of the next frame (cx_, cy_, h_, w_) are obtained. Then, the coordinate position of the target in the next frame is matched, and (cx_, cy_, h_, w_) is updated to replace (cx, cy, h, w), and step 2 is continuously executed; When the number of frames in the target lost state reaches the lost frame threshold, the tracking state is exited.
6. The single target tracking method applicable to the terminal side according to claim 1, characterized in that: When training the tracking model, a loss function is set to supervise pixel matching. The loss function is: ; Among them, A(i, j) is the pixel value of the tracking model output image at the (i, j) coordinate, B(i, j) is the pixel value of the target image at the (i, j) coordinate, and m, n are the width and height of the model output and target images.
7. The single target tracking method applicable to the terminal side according to claim 1, characterized in that: When training the tracking model, the tracking model is fine-tuned by combining the fusion dataset with the existing tracking dataset.
Citation Information
Patent Citations
Lightweight infrared and visible light image fusion method based on convolutional neural network
CN116681636A
Lightweight unmanned aerial vehicle target tracking method and system based on dual cross-correlation fusion
CN119251522A