Single target tracking method suitable for end side

By using feature fusion networks and lightweight MobileNetV4 backbone networks on end-side devices, combining channel attention mechanisms and spatial pyramid pooling layer, the problem of poor tracking in low-light and complex scenarios is solved, stable single-target tracking is achieved and the demand for device computing power is reduced.

CN119991760AActive Publication Date: 2025-05-13CHENGDU HAOFU TECH CO LTD

Patent Information

Application Number
CN202510479961.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Existing tracking algorithms are susceptible to interference from factors such as lighting changes, target occlusion, and scale changes in complex scenarios, especially in low-light situations, and the low-computing equipment on the opposite side is unfriendly, making real-time tracking impossible.

Method used

The y-components of visible and infrared tiles are weighted through the feature fusion network, combined with the lightweight MobileNetV4 backbone network and the heavy parameter structure, a channel attention mechanism and spatial pyramid pooling layer are introduced, the tracking model is optimized to adapt to the end-side devices, and target loss in low light is treated through Kalman filtering.

Benefits of technology

It realizes stable single-target tracking under low-light conditions, reduces the computing power requirement of the opposite side equipment, and improves the lightweight and feature extraction efficiency of the tracking model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991760A_ABST
    Figure CN119991760A_ABST
Patent Text Reader

Abstract

The invention discloses a single target tracking method suitable for an end side, and belongs to the field of computer vision processing. And dynamic weighted fusion is carried out on the y components of the visible light and the infrared image through the feature fusion network, and target feature representation under weak light is enhanced. The tracking model takes MobileNetV4 as a backbone network and combines a heavy parameter structure, a channel attention mechanism and dynamic multi-scale feature fusion, so that the feature extraction efficiency is improved while the weight is reduced. And designing a multi-threshold decision strategy, and realizing continuous tracking, feature updating or Kalman filtering prediction according to the confidence coefficient. In the training stage, the model performance is optimized through dynamic data enhancement, visible light information compensation and a specific loss function. According to the method, feature robustness is enhanced through multi-modal fusion, network adaptation end-side computing power is lightened, tracking stability is improved through a multi-threshold strategy, the method is suitable for low-power-consumption scenes such as security monitoring and mobile equipment, and efficient and stable single-target tracking in a weak light environment is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision processing, and in particular to a single target tracking method applicable to a terminal side. Background Art

[0002] Traditional tracking algorithms (such as KCF) are easily disturbed by factors such as illumination changes, target occlusion, and scale changes in complex scenes. Deep learning twin network tracking algorithms (such as SiamFC) also have good tracking effects under the influence of factors such as target occlusion and scale changes, but they perform poorly under illumination changes, especially in low-light conditions, and are not friendly to devices with lower computing power on the edge side. The edge side (EdgeSide) is the device side, usually the device directly used by users, such as smartphones, tablets, and Internet of Things (IoT) devices. However, these devices usually have limited computing power and energy consumption, and most of them transmit and process data through wireless connection networks. Due to their own computing power and other defects, they cannot achieve real-time tracking in actual operation. Summary of the invention

[0003] The purpose of the present invention is to overcome the problem that the current tracking algorithm is not friendly to the equipment on the terminal side, and to provide a single target tracking method suitable for the terminal side.

[0004] The objective of the present invention is achieved through the following technical solutions:

[0005] A single target tracking method applicable to a terminal side includes the following steps:

[0006] Step 1. Obtain the target position through target detection, intercept the area tile of the target position, and perform weighted fusion of the visible light and infrared y components of the aligned tile coordinates through the feature fusion network to obtain the fused search area tile, and strengthen the feature representation of weak light (light intensity less than 50 lux);

[0007] Step 2. Input the image block into the tracking model. Select nanotrack_v2 in the baseline as the tracking model. The lightweight model MobileNetV4 is introduced into the backbone network of the tracking model, and the re-parameter structure is introduced to replace the original convolutional layer. The new target position and confidence are obtained through the output of the tracking model.

[0008] Step 3. Set the range of confidence level p:

[0009] a) High threshold (p>0.95): continuous tracking of the target;

[0010] b) Medium threshold (0.8≤p≤0.95): Update the feature vector of the tracked target;

[0011] c) Low threshold (p<0.8): The target is judged as lost, and the Kalman filter is used to predict the target position in the next frame to match the target position in the next frame.

[0012] Furthermore, in step 2, a channel attention mechanism and a spatial pyramid pooling layer are introduced into the tracking model, and the tracking model is trained.

[0013] Furthermore, in step 1, the target position is obtained through target detection, and the target original position coordinates (cx, cy, h, w) are obtained, which are the center point coordinates and width and height of the target respectively. The width and height z_l of the target area are obtained according to the following formula:

[0014] ;

[0015] Among them, w is the width of the target position coordinates, and h is the height of the target position coordinates.

[0016] Furthermore, the (cx, cy, z_l, z_l) target area tile is intercepted, and the tile coordinates of the registered infrared image are obtained. After interception, it is resized to (z_l, z_l), and then the y components of the visible light and infrared target area tiles are weighted fused through the feature fusion network to obtain the fused target area YUV tile, and the YUV tile is input into the tracking model.

[0017] Furthermore, after the YUV block is input into the tracking model, the second frame of the initialization is completed, and (cx, cy, 2*z_l, 2*z_l) is intercepted to obtain the block coordinates of the registered infrared image, which are then adjusted to (2*z_l, 2*z_l). The y components of the visible light and infrared target area blocks are weightedly fused through the feature fusion network to obtain the YUV block of the fused target area. The adjusted YUV block is input into the tracking model, and the tracking model is used for reasoning to output the new target position (cx_, cy_, h_, w_) and confidence.

[0018] Among them, (cx_, cy_, h_, w_) are the coordinates of the target position of the next frame to be updated.

[0019] Furthermore, after obtaining the confidence level, the option of target tracking is selected according to the confidence level:

[0020] When the confidence is higher than the high threshold (p>0.95), the target is continuously tracked: (cx_, cy_, h_, w_) is updated to replace (cx, cy, h, w), and a new (cx, cy, 2*z_l, 2*z_l) is calculated, and step 2 is continuously executed;

[0021] When the confidence is the medium threshold (0.8≤p≤0.95), update the feature vector of the tracked target: update (cx_, cy_, h_, w_) to replace (cx, cy, h, w), complete the tracking initialization with the new coordinates, and continue to execute step 2;

[0022] When the confidence is a low threshold (p<0.8): the target is determined to be lost, and the Kalman filter is used to predict the target position of the next frame, and the coordinates of the target position of the next frame (cx_, cy_, h_, w_) are obtained. Then, the coordinate position of the target in the next frame is matched, and (cx_, cy_, h_, w_) is updated to replace (cx, cy, h, w), and step 2 is continuously executed;

[0023] When the number of frames in the target lost state reaches the lost frame threshold, the tracking state is exited.

[0024] Furthermore, during the training of the tracking model, data is collected for scenes under weak light conditions with an illumination intensity of less than 50 lux through the same optical axis and field of view, and the blocks are cropped and adjusted according to the infrared and visible light data obtained after registration, and data enhancement is performed through dynamic scaling.

[0025] Furthermore, in the training of the tracking model, a compensation mechanism for the missing visible light information in a low-light environment is established. The weight of the compensation mechanism is to calculate the image entropy of the registered infrared image and the pixel average value of the visible light, and the target data is produced by the infrared and visible light data according to the compensation mechanism. The weights are as follows:

[0026] ;

[0027] Among them, q is the fusion weight of the visible light y component of the compensation mechanism, i entropy is the entropy of the infrared image after registration, v average is the average value of visible light pixels.

[0028] Furthermore, during the training of the tracking model, the pixel matching is supervised by the loss function. Specifically, the pixel difference between the model output and the target image is calculated to supervise the model to learn the pixel-level matching of the target features and ensure that the output is consistent with the true label. The loss function is:

[0029] Among them, A(i, j) is the pixel value of the tracking model output image at the (i, j) coordinate, B(i, j) is the pixel value of the target image at the (i, j) coordinate, and m, n are the width and height of the model output and target images.

[0030] Furthermore, during the training of the tracking model, the tracking model is fine-tuned by combining the fusion data set with the existing tracking data set.

[0031] The beneficial effects of the present invention are:

[0032] (1) By combining a lightweight MobileNetV4 backbone network with a heavy parameter structure, channel attention, and dynamic multi-scale feature fusion, the model is lightweight and feature extraction efficiency is improved;

[0033] (2) After the infrared and visible light data are fused, the characteristics of the tracked target are enhanced, and the tracking model is made lighter, which improves the ability to extract features, thereby improving stable tracking in low-light conditions and being more friendly to the computing power requirements on the end side. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A flowchart of a single target tracking method applicable to the terminal side;

[0035] Figure 2 Schematic diagram of the lightweight tracking network architecture and feature processing based on MobileNetV4. DETAILED DESCRIPTION

[0036] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0037] Example 1

[0038] See also Figure 1 , a single target tracking method applicable to the terminal side is provided, and the steps include:

[0039] Step 1. Obtain the target position through target detection, intercept the area tile of the target position, and perform weighted fusion of the visible light and infrared y components of the aligned tile coordinates through the feature fusion network to obtain the fused search area tile to enhance the feature representation of weak light;

[0040] Step 2. Input the image block into the tracking model. Select nanotrack_v2 in the baseline as the tracking model. The lightweight model MobileNetV4 is introduced into the backbone network of the tracking model, and the re-parameter structure is introduced to replace the original convolutional layer. The new target position and confidence are obtained through the output of the tracking model.

[0041] Step 3. Set the range of confidence level p:

[0042] a) High threshold (p>0.95): continuous tracking of the target;

[0043] b) Medium threshold (0.8≤p≤0.95): Update the feature vector of the tracked target;

[0044] c) Low threshold (p<0.8): The target is judged as lost, and the Kalman filter is used to predict the target position in the next frame to match the target position in the next frame.

[0045] According to the target clicked by the customer (or obtained by target detection or other methods), the original position coordinates (cx, cy, h, w) of the target are obtained, which are the center point coordinates and width and height of the target respectively. The width and height z_l of the target area are obtained according to the following formula:

[0046] ;

[0047] Among them, w is the width of the target position coordinates, and h is the height of the target position coordinates.

[0048] The target area tile with coordinates (cx, cy, z_l, z_l) is intercepted, and the tile coordinates of the registered infrared image are obtained. For example, the visible light resolution is 1080*1920 and the infrared resolution is 512*640. The registered coordinates will be mapped accordingly. After intercepting the coordinates, resize to (z_l, z_l), and then use the feature fusion network to weightedly fuse the y components of the visible light and infrared target area tiles to obtain the fused target area yuv tile. The tile is resized to (127*127) and sent to the input0 of the updated tracking model.

[0049] Among them, the tracking model is trained in the following ways:

[0050] Data is collected for scenes under low-light conditions based on the same optical axis and field of view, and the registered infrared and visible light data are cropped into square blocks of 127*127 to 255*255. Dynamic scale can also be used for data enhancement, such as cropping 100*100 blocks and then scaling them up to 150*150.

[0051] A compensation mechanism for the lack of visible light information in low-light environments is established. The weight of the compensation mechanism is calculated by the image entropy of the registered infrared image and the average pixel value of the visible light. The target data is produced by the infrared and visible light data pairs according to the compensation mechanism. The weights are as follows:

[0052] ;

[0053] Among them, q is the fusion weight of the visible light y component of the compensation mechanism, i entropy is the entropy of the infrared image after registration, v average is the average value of visible light pixels.

[0054] The tracking model is then trained through the loss function, and pixel matching is supervised by the loss function. Specifically, the pixel difference between the model output and the target image is calculated to supervise the model to learn pixel-level matching of target features to ensure that the output is consistent with the true label. The loss function is:

[0055] ;

[0056] Among them, A(i, j) is the pixel value of the model output image at the (i, j) coordinate, B(i, j) is the pixel value of the target image at the (i, j) coordinate, and m, n are the width and height of the model output and target images.

[0057] Then, the tracking model is fine-tuned and trained based on the data fusion dataset combined with the existing tracking dataset. Stable tracking in low-light scenes is achieved through multimodal data fusion and lightweight network design: the input module receives the y component blocks of the registered visible light and infrared images, and generates a fused feature map through the feature fusion network (dynamically calculating the weight q based on the infrared image entropy and the visible light mean); the tracking model uses MobileNetV4 as the backbone network, combined with the heavy parameter structure, channel attention mechanism and spatial pyramid pooling layer with dynamic weights, to output classification confidence and regression coordinates; the multi-threshold decision module controls the target state according to the confidence (high threshold direct update, medium threshold reinitialization, low threshold trigger Kalman filter prediction); data processing includes target area calculation during initialization (z_l formula) and search area expansion during tracking (2×z_l). In the training stage, the model performance is optimized through data enhancement (dynamic cropping), compensation mechanism and specific loss function (combining pixel error and logarithmic adjustment term), and finally a lightweight, real-time and stable single target tracking in low-light scenes is achieved.

[0058] See also Figure 2 ,The processing flow of the tracking network based on MobileNetV4: The input is two feature fusion images with sizes of 127×127 and 255×255 respectively. After the data is input, it is processed by the improved network based on MobileNetV4, and the features with 48 channels (marked "c=48") and 16 steps (marked "s=16") are output, generating feature maps with sizes of 8 and 16. The features are then fused through the "PointwiseCorrelation" operation, and the fused features are input into the classification (cls) and regression (reg) branches respectively: the classification branch processes the features with 128 channels (marked "c=128") and finally outputs the result with a size of 16; the regression branch processes the features with 96 channels (marked "c=96") and also outputs the result with a size of 16, realizing the dual-task output of target classification and position regression.

[0059] Through the trained tracking model, the target is continuously tracked:

[0060] After the second frame is initialized, the coordinates (cx, cy, 2*z_l, 2*z_l) are intercepted, and the coordinates of the tile of the registered infrared image are obtained. After interception, it is resized to the coordinates (2*z_l, 2*z_l). Then, the feature fusion network is used to weightedly fuse the y components of the visible light and infrared target area tiles to obtain the fused target area yuv tile. The tile is resized to (255*255) and sent to the input1 of the updated tracking model (the update method is the same as input0), and the model is used for reasoning to obtain the new target position coordinates (cx_, cy_, h_, w_) and confidence of the next frame.

[0061] Among them, (cx_, cy_, h_, w_) are the coordinates of the target position of the next frame to be updated.

[0062] When the confidence is higher than the high threshold (p>0.95), the target is continuously tracked: the new target position coordinates (cx_, cy_, h_, w_) are updated and replaced with (cx, cy, h, w), and the new coordinates (cx, cy, 2*z_l, 2*z_l) are calculated, and step 2 is continuously executed;

[0063] When the confidence is the medium threshold (0.8≤p≤0.95), update the feature vector of the tracked target: update (cx_, cy_, h_, w_) to replace (cx, cy, h, w), complete the tracking initialization with the new coordinates, and continue to execute step 2;

[0064] When the confidence is a low threshold (p<0.8): the target is determined to be lost, and the Kalman filter is used to predict the target position of the next frame, and the coordinates of the target position of the next frame (cx_, cy_, h_, w_) are obtained. Then, the coordinate position of the target in the next frame is matched, and (cx_, cy_, h_, w_) is updated to replace (cx, cy, h, w), and step 2 is continuously executed;

[0065] The target is continuously tracked through the above steps, and the tracking state is exited when the number of frames in the target lost state reaches the lost frame threshold.

[0066] Example 2

[0067] Step 1. Initialization of spatiotemporal feature fusion: Target acquisition: The target center point (cx, cy) and width and height (w, h) are obtained through the improved CenterNet target detection algorithm, and the timestamp of the initial frame is recorded. The area calculation is based on the formula:

[0068] ;

[0069] This is used to calculate the size of the target area. Spatiotemporal feature construction: intercept the (cx, cy, z_l, z_l) area of ​​the visible light and infrared images, and resize to (z_l, z_l). Perform optical flow calculation on the multimodal data (YUV+infrared) of the current frame and the previous two frames to generate an optical flow field map containing motion information. Channel-join the optical flow field and the multimodal feature map to form a spatiotemporal fusion feature (127×127×(3+1+2)), and input it into the tracking model input0 to complete initialization.

[0070] Step 2. Tracking model improvement: The backbone network uses MobileNetV4-DeformableConv, and deformable convolution is introduced in the convolution layer to adaptively capture target deformation. The Transformer Encoder module is embedded to perform global dependency modeling on spatiotemporal fusion features and enhance the use of temporal information. Dynamic template management: Maintain a template library to store the target feature templates of the last 5 frames. When inferring each frame, multiple templates are dynamically selected for weighted fusion according to the confidence and motion speed of the current frame to generate the final matching template. Feature pyramid matching: Extract multi-scale feature pyramids for the 255×255 search area. Through the cross-scale attention mechanism, the template features are matched with the search area features at multiple scales, and the target position (cx_, cy_, h_, w_) and confidence p are output.

[0071] Step 3. Adaptive threshold decision and template update confidence dynamic threshold:

[0072] High confidence (p>0.95): Directly update the target coordinates and add the current frame features to the template library (replacing the oldest template);

[0073] Medium confidence (0.8≤p≤0.95): triggers template weight update, and adjusts the weight of each template in the template library according to the current matching result. Use online hard example mining (OHEM) to select samples with difficult matching to fine-tune the model;

[0074] Low confidence (p<0.8): Enable dual template matching: Use the best template in history and the current predicted template for secondary matching. Combine IMU sensor data (such as acceleration and angular velocity) to correct the Kalman filter prediction frame and improve motion prediction accuracy.

[0075] Step 4. Adversarial training and self-supervised learning optimization:

[0076] Data enhancement upgrade: Use StyleGAN2 to generate virtual multimodal data in low-light scenes to enhance data diversity. Introduce CutMix and MixUp technologies to improve the robustness of the model to occlusion and blur. Self-supervised pre-training pre-trains the model on unlabeled data and learns the universal representation of multimodal features through mask image modeling (MIM). Maximize the feature similarity of the same target in different modalities and minimize the feature similarity of different targets. The loss function can be further expressed as:

[0077] ;

[0078] in, It is the classification loss, which is used to supervise the model's ability to classify targets and backgrounds, ensuring that the model can accurately distinguish whether the current search area contains targets; It is a regression loss used to supervise the model’s prediction accuracy of the target location. It optimizes the accuracy of target positioning by calculating the coordinate difference between the predicted box and the real box (such as the center point coordinates, width and height). is the optical flow prediction loss, It is a contrast loss used to enhance the feature distinction between the target and the background.

[0079] The above solution combines the optical flow field with multimodal data to enhance the target representation in dynamic scenes; through the template library and weighted fusion, it adapts to the changes in the target appearance. The introduction of IMU data through multiple sensors improves the accuracy of motion prediction and enhances the environmental adaptability of the end-side device. And through adversarial self-supervised training, the generative model and self-supervised learning are used to solve the problem of insufficient low-light data and improve the generalization ability of the model. 32fps real-time tracking is achieved on the end-side device (such as RK3588), and the tracking accuracy is improved in low-light, fast-moving, and occluded scenes, and the model's anti-interference ability is significantly enhanced.

[0080] By combining the lightweight MobileNetV4 backbone network with the heavy parameter structure, channel attention and dynamic multi-scale feature fusion, the model is lightweight and the feature extraction efficiency is improved. After the infrared and visible light data are fused, the characteristics of the tracked target are enhanced, and the tracking model is made lighter, the feature extraction capability is improved, thereby improving the stable tracking in low-light conditions and being more friendly to the computing power requirements on the end side.

[0081] The above is only a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not deviate from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.

Claims

1. A single target tracking method applicable to a terminal side, characterized in that the steps include: Step 1. Obtain the target position through target detection, intercept the area block of the target position, and perform weighted fusion of the visible light and infrared y components of the aligned block coordinates through the feature fusion network to obtain the fused search area block; Step 2. Input the image block into the tracking model. Select nanotrack_v2 in the baseline as the tracking model. The lightweight model MobileNetV4 is introduced into the backbone network of the tracking model, and the re-parameter structure is introduced to replace the original convolutional layer. The new target position and confidence are obtained through the output of the tracking model. Step 3. Set the range of confidence level p: a) High threshold (p>0.95): continuous tracking of the target; b) Medium threshold (0.8≤p≤0.95): Update the feature vector of the tracked target; c) Low threshold (p<0.8): The target is judged as lost, and the Kalman filter is used to predict the target position in the next frame to match the target position in the next frame.

2. A single target tracking method applicable to a terminal side according to claim 1, characterized in that: In the step 2, a channel attention mechanism and a spatial pyramid pooling layer are introduced into the tracking model, and the tracking model is trained.

3. The single target tracking method applicable to the terminal side according to claim 1, characterized in that: In step 1, the target position is obtained through target detection, and the original position coordinates of the target (cx, cy, h, w) are obtained, which are the center point coordinates and width and height of the target respectively. The width and height z_l of the target area are obtained according to the following formula: ; Among them, w is the width of the target position coordinates, and h is the height of the target position coordinates.

4. The single target tracking method applicable to the terminal side according to claim 3, characterized in that: The (cx, cy, z_l, z_l) target area tile is intercepted and the tile coordinates of the registered infrared image are obtained. After interception, it is adjusted to (z_l, z_l). Then, the y components of the visible light and infrared target area tiles are weightedly fused through the feature fusion network to obtain the fused target area YUV tile, and the YUV tile is input into the tracking model.

5. The single target tracking method applicable to the terminal side according to claim 4, characterized in that: After the YUV block is input into the tracking model, (cx, cy, 2*z_l, 2*z_l) is intercepted to obtain the block coordinates of the registered infrared image, which is then adjusted to (2*z_l, 2*z_l). The Y components of the visible light and infrared target area blocks are weighted fused through the feature fusion network to obtain the YUV block of the fused target area. The adjusted YUV block is input into the tracking model, and the tracking model is used for reasoning to output the new target position (cx_, cy_, h_, w_) and confidence. Among them, (cx_, cy_, h_, w_) are the coordinates of the target position of the next frame to be updated.

6. A single target tracking method applicable to a terminal side according to claim 5, characterized in that: After obtaining the confidence level, select the option to track the target based on the confidence level: When the confidence is higher than the high threshold (p>0.95), the target is continuously tracked: (cx_, cy_, h_, w_) is updated to replace (cx, cy, h, w), and a new (cx, cy, 2*z_l, 2*z_l) is calculated, and step 2 is continuously executed; When the confidence is the medium threshold (0.8≤p≤0.95), update the feature vector of the tracked target: update (cx_, cy_, h_, w_) to replace (cx, cy, h, w), complete the tracking initialization with the new coordinates, and continue to execute step 2; When the confidence is a low threshold (p<0.8): the target is determined to be lost, and the Kalman filter is used to predict the target position of the next frame, and the coordinates of the target position of the next frame (cx_, cy_, h_, w_) are obtained. Then, the coordinate position of the target in the next frame is matched, and (cx_, cy_, h_, w_) is updated to replace (cx, cy, h, w), and step 2 is continuously executed; When the number of frames in the target lost state reaches the lost frame threshold, the tracking state is exited.

7. The single target tracking method applicable to the terminal side according to claim 2, characterized in that: When training the tracking model, data is collected for scenes under weak light conditions with an illumination intensity of less than 50 lux on the same optical axis and field of view. The image blocks are cropped and adjusted based on the infrared and visible light data obtained after registration, and data enhancement is performed through dynamic scaling.

8. The single target tracking method applicable to the terminal side according to claim 7, characterized in that: When training the tracking model, a compensation mechanism for the missing visible light information is established. The weight of the compensation mechanism is to calculate the image entropy of the registered infrared image and the average pixel value of the visible light. The target data is generated by the infrared and visible light data according to the compensation mechanism. The weights are as follows: ; Among them, q is the fusion weight of the visible light y component of the compensation mechanism, i entropy is the entropy of the infrared image after registration, v average is the average value of visible light pixels.

9. The single target tracking method applicable to the terminal side according to claim 2, characterized in that: When training the tracking model, a loss function is set to supervise pixel matching. The loss function is: ; Among them, A(i, j) is the pixel value of the tracking model output image at the (i, j) coordinate, B(i, j) is the pixel value of the target image at the (i, j) coordinate, and m, n are the width and height of the model output and target images.

10. The single target tracking method applicable to the terminal side according to claim 2, characterized in that: When training the tracking model, the tracking model is fine-tuned by combining the fusion dataset with the existing tracking dataset.

Citation Information

Patent Citations

  • RGBT target tracking method based on twin network structure and anchor frame adaptive thought

    CN116563343A

  • Lightweight infrared and visible light image fusion method based on convolutional neural network

    CN116681636A

  • RGBT target tracking method based on target perception enhancement fusion structure

    CN117474957A

  • Adaptive dark light enhanced target tracking network structure and target tracking method

    CN118365676A

  • Real-time target tracking method based on twin network and embedded device

    CN118674752A

Cited By

  • Sea surface target tracking method and system based on infrared and visible light image fusion

    CN120259372A

  • Multi-modal fusion gas leakage detection system based on super-division graph reconstruction

    CN121564495A

  • Multimodal fusion gas leak detection system based on super-resolution image reconstruction

    CN121564495B

  • Homomorphic filtering enhanced fusion tracking method and system in low-light infrared scene

    CN121788361A

  • Motion robot double-light fusion furnace bottom non-stop temperature measurement method and system

    CN121977701A