Visible light infrared target tracking method and system based on edge information dynamic enhancement

By building an edge information adaptation module and a dynamic routing mechanism, the performance degradation problem caused by illumination changes and background confusion in RGB-T target tracking is solved, a more stable target tracking effect is achieved, and the utilization efficiency of multimodal features is improved.

CN120807561APending Publication Date: 2025-10-17XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510813891.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing RGB-T target tracking methods are susceptible to illumination changes and background confusion in complex scenes, resulting in degraded tracking performance. In particular, in low-light conditions at night or when the target is camouflaged, edge information is insufficiently utilized, leading to blurred positioning boundaries or tracking drift.

Method used

An edge information adaptation module that can be embedded in the backbone network is constructed. The edge information of multimodal images is extracted through a local gradient edge detector. The hybrid multi-expert dynamic selection unit and the edge adapter are used for adaptive dynamic routing to suppress invalid information interference and achieve efficient fusion of edge features and appearance features.

Benefits of technology

It improves the multimodal tracking performance, suppresses the invalid information and multimodal feature noise introduced by the confusion between target and background, enhances the target representation ability, and improves the tracking accuracy and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807561A_ABST
    Figure CN120807561A_ABST
Patent Text Reader

Abstract

The invention discloses a visible light infrared target tracking method and system based on edge information dynamic enhancement, and mainly solves the problems of insufficient utilization of target structure information and equal utilization of modal information in the prior art. According to the scheme, the method comprises the following steps: intercepting a multi-mode video into multiple frames of images according to time, and preprocessing by taking a GTOT data set as a training set; a tracking model comprising a local gradient edge detector, a hybrid multi-expert dynamic selection unit, an edge adapter and a backbone network is constructed, and the backbone network is in parallel connection with the local gradient edge detector, the edge adapter and the hybrid multi-expert dynamic selection unit which are connected in series and is used for extracting multi-modal image features; inputting the preprocessed training set into a tracking model for iterative training; and inputting the test set RGBT234 into the trained tracking model to output the position information of the target in each frame of the video. The method can efficiently utilize the edge information of the target to enhance the representation capability of the whole network to the target, improves the tracking performance, and can be used for infrared warning, electric power, medical treatment and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of target tracking, and particularly relates to a visible light-infrared target tracking method and system, which can be used for infrared warning, power, medical treatment and fire warning. BACKGROUND

[0002] An infrared thermal image is a special image. An infrared thermal imager converts the distribution of the temperature of a target surface into an image through detection of the radiation of the target object, photoelectric conversion and signal processing. Due to its particularity, the infrared image is applied to many fields such as infrared warning, power, medical treatment and fire warning. The infrared thermal imaging technology is used to realize temperature monitoring in a certain scene due to its high precision, non-contact and wide temperature measurement range. In addition, the infrared image also shows more and more strong vitality in material detection, production monitoring, building energy-saving evaluation, disaster reduction and prevention and many other aspects. At the same time, various monitoring and identification systems derived from the research and application of infrared image recognition technology have been widely applied in more and more industries.

[0003] Single target tracking, as a basic task of video processing, generally refers to predicting the position of a target in subsequent frames of a video given the position of the target in the first frame of the video, wherein the position of the target is represented by a bounding box. In general, the single target tracking task does not give the category of the target to be tracked, and requires the tracker to have a category generalization capability. In target tracking, visible light images often face some challenges such as illumination changes, shadows and view angle changes, which can all lead to the failure of tracking algorithms. However, thermal infrared images can make up for the limitations of visible light images. Thermal infrared images are not affected by lighting conditions and can effectively detect the thermal features of the target, so they can still maintain stable target recognition and tracking ability at night or in bad weather. In order to make full use of the advantages of the two modalities, visible light images and thermal infrared images can be combined to form a robust multi-modal target tracker. This tracker can improve the accuracy and stability of tracking by fusing the information of the two image modalities. Specifically, visible light images provide visual information such as the shape and color of the target, while thermal infrared images provide thermal feature information of the target, including temperature differences under different lighting conditions. By comprehensively utilizing these two kinds of information, more reliable target detection and tracking can be achieved in various complex scenes.

[0004] Currently, the thermal infrared multi-modal target tracking is regarded as an effective tracking algorithm that makes full use of the complementary characteristics of multi-modal. According to the different utilization methods, it can be divided into multi-modal feature fusion tracking algorithm and modal prompt based tracking algorithm. The multi-modal feature fusion tracking algorithm mainly focuses on multi-modal feature fusion, which is mainly to obtain accurate modal heterogeneity features to maximize the complementary advantages of visible light and thermal infrared data. Some current works use simple feature summation and splicing operations to directly fuse the features of the two modalities, and some works consider balancing the contributions of visible light and thermal infrared information in the fused features and using attention mechanisms to adaptively learn multi-modal information fusion. Some methods also explore the influence of pixel-level fusion, feature-level fusion and decision-level fusion strategies on performance to improve tracking performance. Recently, prompt tuning has become a mainstream paradigm in the field of natural language processing (NLP), which adapts the base model to different tasks by adding text prompts to the model input. Some methods transplant this paradigm to the field of computer vision by adding learnable visual prompts to the frozen base model, and research shows that visual prompt learning is expected to become an alternative method to full fine-tuning. The idea of this kind of algorithm is to freeze the weights of the RGB large model and only fine-tune a small part of the network parameters to adapt the frozen base network to different downstream tasks.

[0005] In practical applications, visible light images often encounter problems of light changes and viewing angles in target tracking, which leads to the risk of tracking failure. In contrast, thermal infrared images are not affected by light and can stably identify the thermal features of the target, making them particularly suitable for use in night and harsh weather environments. In the field of multi-modal thermal infrared target tracking, common ordinary target tracking algorithms are still the main method. Because in essence, multi-modal thermal infrared target tracking is no different from single target tracking in natural scenes, both of which are given the location of the target in the first frame of the video and track the target in the subsequent frames. However, since the input of the network is two modal images, compared to single modal input, the network needs to process two modal images at the same time, and how to effectively utilize the complementary information between modalities needs to be optimized accordingly.

[0006] The patent document with application number CN202411122743.6 discloses an infrared tracking method based on feature local correction and multi-modal channel sparse selection prompt. It uses an embedding network to extract feature information from visible light and thermal infrared modalities respectively, then sends the two modal feature information into a modal space correction module to obtain spatially aligned and enhanced thermal infrared image features. After that, the visible light features and the aligned and enhanced thermal infrared features are input into a backbone network and a channel selection prompter. The final features are obtained through multi-level interaction of the prompter and the backbone network. Finally, a prediction head is used for classification and regression. This method considers the internal relationship and complementarity of different modal features from two angles in the semantic space through the prompt module, which can improve the effectiveness of the prompt information and solve the feature deviation problem caused by spatial misalignment. However, this method mainly focuses on the interaction and alignment of modal information, and still relies on apparent features such as color, texture, and temperature distribution, so it lacks the use of essential geometric structure information.

[0007] The patent document with application number CN202411910627.0 discloses an RGB-T target tracking method combining asymmetric enhancement and interactive fusion. It uses a backbone network to extract feature information from visible light and thermal infrared modalities respectively, then sends the feature information into a visible light image enhancement module and a thermal infrared image enhancement module for processing in parallel, outputs two modal features, and finally uses a prediction head for classification and regression. This method achieves more accurate target recognition and positioning. However, since it mainly focuses on the fusion and enhancement of different modalities, it does not consider the imbalance of different modal image contributions to target representation, which will introduce noise information in the modalities and affect the tracking performance.

[0008] In summary, existing RGB-T target tracking methods mainly focus on the design of inter-modal interaction and fusion mechanisms. Although they can improve the robustness of basic scenes through cross-modal feature complementarity, their core still relies on the fusion strategy of apparent features such as color, texture, and temperature distribution. When the target and background have similar colors in the visible light modality and similar temperature distributions in the thermal infrared modality, such as complex scenes like low light at night and target camouflage, existing methods are easily disturbed by invalid information due to excessive reliance on apparent feature fusion, leading to blurred positioning boundaries or even tracking drift, exposing the defect of insufficient use of essential geometric structure information. At the same time, the above existing methods have significant limitations in utilizing multi-modal features, mainly in two aspects: one is that the visible light edge is easily disturbed by dynamic light changes such as strong reflection and shadow mutations, leading to edge breakage or artifact generation; the other is that the thermal infrared edge may become blurred or even disappear under isothermal environments or thermal radiation interference such as heat source crossing. SUMMARY

[0009] The present application aims at the deficiencies of the prior art, and provides a visible light infrared target tracking method and system based on edge information dynamic enhancement, so as to suppress invalid information introduced due to target and background confusion and noise information introduced due to equal use of multi-modal features, and improve multi-modal tracking performance.

[0010] To achieve the above object, the key technology of the present application is: by constructing an edge information adaptive module that can be embedded in a backbone network, efficient fusion of edge features and appearance features is realized to suppress invalid information introduced due to target and background confusion; by establishing a visible light, thermal infrared and cross-modal edge expert collaborative framework, real-time distribution of edge modal weights is realized by using an adaptive dynamic routing mechanism to suppress noise information introduced due to equal use of multi-modal features.

[0011] The above two technologies are combined to ultimately improve the performance of multi-modal tracking. The implementation scheme includes:

[0012] 1. A visible light infrared target tracking method based on edge information dynamic enhancement, characterized in that it comprises:

[0013] (1) A multi-modal video is cut into multiple frames of still images in chronological order, and a GTOT dataset is used as a training set and an RGBT234 dataset is used as a test set, and the training set is preprocessed by cropping and scaling according to the target location and size;

[0014] (2) A tracking model including a local gradient edge detector, a mixed multi-expert dynamic selection unit, an edge adapter and a backbone network is constructed:

[0015] The local gradient edge detector includes depth change calculation, translation transformation and feature fusion, and is used to extract edge information of multi-modal images;

[0016] The mixed multi-expert dynamic selection unit includes a router network and a multi-expert network, and is used to intelligently and dynamically enhance the edge information of thermal infrared and visible light to generate more excellent target representation;

[0017] The edge adapter includes a multi-layer fully connected network, and is used to compress and reconstruct the edge features of the target in the multi-modal image;

[0018] The backbone network is connected in parallel with the locally connected gradient edge detector, the edge adapter and the mixed multi-expert dynamic selection unit connected in series, and is used to extract image features after multi-modal fusion;

[0019] (3) The preprocessed training set is input into the tracking model for iterative training;

[0020] (4) The test set is input into the trained tracking model, and the position information of the target in each frame of the video is output.

[0021] Further, the router network comprises a linear layer network, a noise superposition operation and a Top-K selection operation, the linear layer network generates an assignment probability for each expert through an input edge feature, and then the optimal expert network is selected for each edge feature through superposition of noise and using the Top-K selection operation.

[0022] Further, the multi-expert network comprises a plurality of multi-layer perceptrons, each expert network is a multi-layer perceptron, and the multi-layer perceptron is used for specific processing of the edge feature.

[0023] Further, the multi-layer fully connected network comprises compression dimension reduction, hidden layer transformation and reconstruction optimization, that is, the dynamic enhanced multi-modal edge feature is subjected to compression dimension reduction first, then the multi-modal edge feature after dimension reduction is subjected to hidden layer transformation, and finally the output after transformation is subjected to reconstruction optimization to obtain the edge feature suitable for backbone network fusion.

[0024] 2. A visible light infrared target tracking system based on dynamic enhancement of edge information, comprising:

[0025] A local gradient edge detection module is configured to extract edge information of a multi-modal image.

[0026] A hybrid multi-expert dynamic selection module is configured to intelligently and dynamically enhance edge information of thermal infrared and visible light to generate more excellent target representation.

[0027] An edge adaptation module is configured to compress and reconstruct the edge feature of the target in the multi-modal image.

[0028] A backbone network module is configured to extract image features after multi-modal fusion.

[0029] A tracking model construction module is configured to connect the backbone network module in parallel with the locally connected local gradient edge detection module, the edge adaptation module and the hybrid multi-expert dynamic selection module in series to construct a tracking model.

[0030] A training module is configured to input a preprocessed training set into the tracking model for iterative training.

[0031] A test module is configured to input a test set into the trained tracking model to output position information of the target in each frame of the video.

[0032] Compared with the prior art, the present application has the following advantages:

[0033] First, the hybrid multi-expert dynamic selection module designed in the present invention can not only intelligently and dynamically enhance the edge information of thermal infrared and visible light to generate a better target representation, but also suppress the noise information introduced by the equal use of multimodal features, thereby improving the multimodal tracking performance;

[0034] Secondly, due to the design of the edge adaptation module, the present invention can compress and reconstruct the edge features of the target in the optimized multimodal image, obtain the edge features fused by the adaptive backbone network, and at the same time suppress the invalid information introduced by the confusion between the target and the background, thereby improving the multimodal tracking performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 This is a flow chart of the visible light infrared target tracking method based on dynamic enhancement of edge information of the present invention;

[0036] Figure 2 It is a schematic diagram of the tracking model structure in the method of the present invention;

[0037] Figure 3 This is a block diagram of the visible light infrared target tracking system based on dynamic enhancement of edge information of the present invention. DETAILED DESCRIPTION

[0038] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0039] Embodiment 1: A visible light infrared target tracking method based on dynamic enhancement of edge information.

[0040] Reference Figure 1 , the implementation steps of the embodiment of the present invention include the following:

[0041] Step 1: Obtain the training set from the GTOT dataset and preprocess the training set images to obtain the template images and search images of the two modalities.

[0042] 1.1) Obtain the GTOT dataset, which consists of time-aligned infrared and visible light video images. Images are captured from these two videos in chronological order to form multiple still images, which serve as the training set.

[0043] 1.2) For the two modal images in the training set, infrared and visible light images are first expanded according to the size of the target bounding box, the area of ​​the expanded region is calculated, and the square root of the expanded region is calculated to obtain the side length to be cropped, and the region is cropped into a square region.

[0044] 1.3) The cropped square area is scaled to different scales, adjusted to 128×128 pixels to obtain the template image of the image, and then resized to 256×256 pixels to obtain the search image of the image, completing the preprocessing of the training set.

[0045] Step 2: Build a tracking model.

[0046] Reference Figure 2 , this step includes:

[0047] 2.1) Select a local gradient edge detector to extract edge information from the search graphs of the two modalities:

[0048] 2.1.1) Using sliding window to obtain images All window areas X patches :

[0049]

[0050] Among them, the window size win size =3;

[0051] 2.1.2) Calculate the maximum value of each window area and obtain its maximum depth map The calculation formula is as follows:

[0052] D max =max(X patches ),

[0053] Among them, max(·) represents the maximum value operation.

[0054] 2.1.3) Use the affine transformation matrix to translate the image in four directions: up, down, left, and right. The translation matrix T is:

[0055]

[0056] in, is the horizontal translation, is the vertical translation;

[0057] 2.1.4) Obtain the translated depth map D according to the translation matrix T shifted :

[0058] D shifted =T×D max ,

[0059] 2.1.5) Calculate the difference between the depth map before and after translation to obtain the gradient information G, and remove the unchanged area after translation through mask operation to obtain the edge information G final :

[0060] G = D max -D shifted

[0061] G final = G x M

[0062] wherein the mask M is a binary matrix indicating the region where the depth changes;

[0063] 2.1.6) The gradient information of the four directions is spliced to obtain the final gradient feature G global , and then the maximum value G max of the edge information is further calculated as the edge information of the multi-modal search graph:

[0064] G max = max(G global );

[0065] 2.1.7) The multi-modal search graph edge information obtained is processed using feature embedding to obtain the edge features of the search graphs of the two modalities:

[0066] First, the maximum gradient information G max is processed using the feature embedding PatchEmb(·) to obtain the edge feature

[0067] EdgeEmb = PatchEmb(G max ), wherein dim = 128

[0068] Next, the maximum gradient information G max is processed using the feature embedding PatchEmb(·) to obtain the edge feature of the visible light search graph: EdgeEmb RGB = PatchEmb(G max );

[0069] Next, the maximum gradient information G max is processed using the feature embedding PatchEmb(·) to obtain the edge feature of the infrared search graph: EdgeEmb TIR = PatchEmb(G max );

[0070] 2.1.8) The edge feature EdgeEmb RGB of the visible light search graph and the edge feature EdgeEmb TIR of the infrared search graph are spliced to obtain the fused edge feature

[0071] EdgeEmb fusion = Concat(EdgeEmb RGB,EdgeEmb TIR );

[0072] 2.2) Establish a mixed multi-expert dynamic selection unit:

[0073] 2.2.1) Establish a router network composed of a linear layer of edge features, for each edge feature to be assigned to different experts:

[0074] First, the router network converts the input edge features into assignment probabilities logits of experts through a linear layer Fc route (·) :

[0075] logits = Fc route (EdgeEmb fusion )

[0076] Next, the most important experts are selected according to the probability logits to process the edge features;

[0077] 2.2.2) Establish a multi-expert network composed of multiple linear layers, for each edge feature to be assigned to the corresponding expert for processing, each expert network contains a multi-layer perceptron responsible for processing the edge features assigned to it, to obtain the output y expert of the expert:

[0078] y expert = MLP(x expert ),

[0079] where x expert is the edge feature assigned to the expert.

[0080] 2.2.3) Weighted fusion of all expert outputs to generate optimized edge features Edge final :

[0081]

[0082] where i represents the i-th expert network, and k = 3.

[0083] 2.3) Establish an edge adapter:

[0084] 2.3.1) A plurality of linear layers are connected in series to form an edge adapter, which is used to compress edge features, hidden layer changes and reconstruction optimization.

[0085] 2.3.2) Dimensionality reduction compression is performed on the edge features to obtain compressed edge features Edge down :

[0086] Edge down = Fc down (Edgefinal ),

[0087] where Fc down (·) denotes a linear layer compressing edge features, Edge final represents edge features;

[0088] 2.3.3) applying a hidden layer transformation to the compressed edge features to obtain hidden layer edge features Edge hidden :

[0089] Edge hidden = Fc hidden (Edge down )),

[0090] where Fc hidden (·) denotes a linear layer transforming edge features in hidden layer, and σ(·) represents a ReLU activation function;

[0091] 2.3.4) reconstructing and optimizing the hidden layer edge features to obtain reconstructed and optimized visible light edge features Edge RGB and reconstructed and optimized infrared edge features Edge TIR :

[0092] Edge RGB ,Edge TIR = Fc reconstructs (Edge hidden )

[0093] where Fc reconstructs (·) denotes a linear layer reconstructing and optimizing edge features;

[0094] 2.3.5) padding the reconstructed and optimized visible light edge features Edge RGB and the reconstructed and optimized infrared edge features Edge TIR to obtain visible light modality edge features AptEdgeEmb RGB adapted to the backbone network and infrared modality edge features AptEdgeEmb TIR adapted to the backbone network:

[0095] AptEdgeEmb RGB = Concat(Edge RGB , Pad),

[0096] AptEdgeEmb TIR = Concat(Edge TIR , Pad),

[0097] where Pad is a padding vector of all zeros;

[0098] 2.4) Establish a backbone network composed of 12 layers of self-attention encoders in series to obtain edge-enhanced image features:

[0099] 2.4.1) Fuse the visible light edge features AptEdgeEmb RGB and the infrared edge features AptEdgeEmb TIR respectively with the output of the i-th layer self-attention encoder to obtain two inputs of the next layer self-attention encoder and

[0100]

[0101] where Encoder(·) represents the self-attention encoder, i = 1, 2,..., 12;

[0102] 2.4.2) Corresponding interaction is performed between each layer of the multi-layer self-attention encoder and each layer of the multi-layer edge adapter until the last layer is interacted, to obtain the response feature map X Fuse :

[0103]

[0104] 2.4.3) Input the two modal fusion response feature maps X Fuse into the target tracking head for tracking to obtain the tracking result (x, y, w, h):

[0105] (x, y, w, h) = TrackHead(X Fuse ),

[0106] where TrackHead(·) is the target tracking head, and (x, y, w, h) represents the center horizontal coordinate, center vertical coordinate, width and height of the tracking result;

[0107] 2.5) Connect the above local gradient edge detector, edge adapter and mixed multi-expert dynamic selection unit in series, and then connect them in parallel with the backbone network to build a tracking model.

[0108] Step 3, input the preprocessed training set into the tracking model for iterative training.

[0109] 3.1) Calculate the distribution difference loss value of the response feature map and the real feature map, the center distance loss value between the predicted box and the real box, and the loss value of the IOU between the predicted box and the real box:

[0110] 3.1.1) Use the Focal loss function to calculate the loss value Loss Focal between the response feature map and the high-frequency feature map generated by the real label:

[0111]

[0112] where FL(p t ) = -(1 - p t ) γ log(p t ) is the cross-entropy value between the response feature map and the high-temperature feature map generated by the real label, is the probability that the model predicts correctly, w and h are the width and height of the response map respectively, y is the high-temperature feature map generated by the real label, and γ is a constant;

[0113] 3.1.2) Calculate the loss value Loss L1 between the predicted target information (x, y, w, h) and the real target information (gtx, gty, gtw, gth) using the L1 loss function

[0114]

[0115] where n is the number of total samples, where (x, y, w, h) contains the center coordinates (x, y) and the width and height (w, h) of the predicted box, and (gtx, gty, gtw, gth) contains the center coordinates (gtx, gty) and the width and height (gtw, gth) of the real box;

[0116] 3.1.3) Calculate the loss value Loss GIoU between the predicted target information (x, y, w, h) and the real target information (gtx, gty, gtw, gth) using the GIoU loss function

[0117]

[0118] where A represents the predicted box (x, y, w, h), B represents the real box (gtx, gty, gtw, gth), C represents the minimum bounding box of A and B, \ represents the difference set, and |·| represents the area of the region;

[0119] 3.2) Backpropagate the above three loss values Loss Focal , Loss L1 and Loss GIoU , calculate the gradient of each layer parameter of the model, and iteratively update the model weight through gradient descent method until the maximum training times are reached, obtain the trained tracking model, and save the optimal model parameters.

[0120] Step 4, input the test set into the trained tracking model, and output the position information of the target in each frame of the video.

[0121] 4.1) The first frame image of the multi-modal video RGBT234 to be tracked is preprocessed using the same cropping and scaling method as the training set;

[0122] 4.2) The preprocessed multi-modal image is used as the template image and search image for this tracking, and is input into the trained tracking model to output the target features enhanced by edge information, obtaining the tracking result of the current frame of the multi-modal video;

[0123] 4.3) After each frame of tracking is completed, the next frame of video image is cropped and scaled with the current predicted position as the center until the last frame of the video is tracked, i.e. the tracking of the entire video is completed.

[0124] Embodiment two, visible and infrared target tracking system based on dynamic enhancement of edge information.

[0125] Reference Figure 3 The present example includes: a local gradient edge detection module 1, a mixed multi-expert dynamic selection module 2, an edge adaptation module 3, a backbone network module 4, a tracking model construction module 5, a training module 6 and a test module 7. The working principle of the whole system is as follows:

[0126] The local gradient edge detection module 1 is used to extract the edge information of the multi-modal image and input the edge information into the mixed multi-expert dynamic selection module 2;

[0127] The mixed multi-expert dynamic selection module 2 is used to intelligently and dynamically enhance the edge information output by module 1, generate more excellent edge features, and input the edge features to the edge adaptation module 3;

[0128] The edge adaptation module 3 is used to compress and optimize the reconstruction of the edge features output by module 2, to obtain edge features suitable for the backbone network, and input them to the backbone network module 4;

[0129] The backbone network module 4 is used to fuse and interact the edge features output by module 3, output the multi-modal image features enhanced by edges, and obtain the tracking result of the multi-modal image;

[0130] The tracking model construction module 5 is used to connect the local gradient edge detection module 1, the mixed multi-expert dynamic selection module 2 and the edge adaptation module 3 in series, and connect the backbone network module 4 in parallel, to construct a tracking model;

[0131] The training module 6 is used to input the preprocessed training set into the tracking model for iterative training;

[0132] The test module 7 is used to input the test set into the trained tracking model to output the position information of the target in each frame of the video.

[0133] It should be noted that the above-mentioned functional modules can be realized by software, hardware, firmware or any combination thereof, in whole or in part. When realized by software, it can be realized in the form of program instruction product in whole or in part. The program instruction product includes one or a group of program instructions. When the program instructions are loaded and executed on a computer, the flow or function is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The program instructions can be stored in a computer readable and writable storage medium or transferred from one computer readable and writable storage medium to another.

[0134] The direct coupling or communication connection between the modules shown or discussed in the embodiments can be realized by indirect coupling or communication connection of some interfaces, devices or modules. The functional modules and sub-modules in the embodiments can be dynamically in one processing component, or each module can be physically present alone, or two or more modules can be dynamically in one processing component. When the above dynamic components are realized in the form of software functional modules and sold or used as independent products, they can also be stored in a computer readable and writable storage medium. The storage medium can be a memory, a magnetic disk or an optical disk, etc.

[0135] The effects of the present application can be further illustrated by the following simulation results:

[0136] I. Simulation conditions

[0137] The simulation experiment uses GTOT dataset and RGBT234 dataset, GTOT dataset is used for model training, and RGBT234 dataset is used for model testing.

[0138] The model training environment and algorithm test environment of the simulation experiment are unified, the computing device used is RTX206012GB, the operating system used is Ubuntu 20.04.7LST, the CUDA version is 11.1, the Python version is 3.8.12, the Torch version is 1.8.0, the training round is 60, the optimizer weight decay is 10 -4 , and the initial learning rate is set to 4x10 -5 .

[0139] II. Simulation content and results

[0140] Under the above conditions, the present application and the existing nine target tracking methods DAFNet, DAPNet, CAT, ADRNet, MANet++, HMFT, MIRNet, LANet and Vipt are used to perform tracking simulation experiments on different 234 video segments on the RGBT234 dataset, and the precision and success rate evaluation indexes of the tracking simulation experiments are calculated, and the results are shown in Table 1.

[0141] Table 1 Simulation results of the present application and the prior art on RGBT234 dataset

[0142]

[0143] The calculation method of the above-mentioned precision is the proportion of the frame number whose Euclidean distance between the center coordinates of the predicted target frame and the center coordinates of the real target frame is less than the threshold value 20 in the total frame number, and the calculation formula is as follows:

[0144]

[0145] Where total_frames represents the total frame number of the video, represents whether the Euclidean distance between the center coordinates of the target frame and the center coordinates of the real target frame is less than the threshold value 20, represents the Euclidean distance between the center coordinates of the target frame and the center coordinates of the real target frame, (x pre ,y pre ) represents the center coordinates of the predicted target frame, and (x gt ,y gt ) represents the center coordinates of the real target frame.

[0146] The calculation method of the above-mentioned success rate is the proportion of the frame number whose IOU between the predicted target frame and the real target frame is greater than 0.5 in the total frame number. Let the predicted bounding box of the target tracking algorithm be bbox_predict, the real target frame be bbox_gt, and the total frame number of the video be total_frames , Then the tracking accuracy of a frame of video is calculated as follows:

[0147]

[0148] Where, represents whether the IOU between the target frame and the real target frame is greater than 0.5, wherein

[0149] The source of the above-mentioned prior art is as follows:

[0150] DAFNet, Deep adaptive fusion network for high performance rgbt tracking, published in International Conference on Computer Vision Workshop in 2019.

[0151] DAPNet, Dense feature aggregation and pruning for rgbt tracking, published in ACM International Conference on Multimedia, 2019.

[0152] CAT, Challenge-aware rgbt tracking, published in European Conference on Computer Vision, 2020.

[0153] ADRNet, Learning adaptive attribute-driven representation for real-time rgb-t tracking, published in International Journal of Computer Vision, 2021.

[0154] MANet++, Rgbt tracking via multi-adapter network with hierarchical divergence loss, published in IEEE Transactions on Image Processing, 2020.

[0155] HMFT, Visible-thermal uav tracking: A large-scale benchmark and new baseline, published in IEEE Conference on Computer Vision and Pattern Recognition, 2022.

[0156] MIRNet, MIRNet: A Robust RGBT Tracking Jointly with Multi-Modal Interaction and Refinement, published in IEEE International Conference on Multimedia and Expo, 2022.

[0157] LANet, Learning adaptive attribute-driven representation for real-time rgb-t tracking, International Journal of Computer Vision, 2021.

[0158] Vipt, Rgb-t tracking via multi-modal mutual prompt learning, IEEE Conference on Computer Vision and Pattern Recognition, 2023.

[0159] As can be seen from Table 1, the tracking method of the present application has improved in both success rate and precision rate compared with the Vipt tracking method, and has improved by 6.9% and 7.9% respectively on the RGBT234 dataset.

[0160] The simulation results show that the tracking model can enhance the representation ability of the target by using the multi-expert dynamic selection unit and the edge adapter, thereby improving the tracking performance.

[0161] The above description is only one specific example of the present application and does not constitute any limitation on the present application. Obviously, for those skilled in the art, after understanding the content and principles of the present application, various modifications and changes in form and details can be made without departing from the principles and structures of the present application. However, these modifications and changes based on the idea of the present application are still within the protection scope of the claims of the present application.

[0162] It should be noted that the step numbers in the specification and claims of the present application are only for the clear description of the embodiments of the present application, and are not limited in sequence.

Claims

1. A visible light infrared target tracking method based on dynamic enhancement of edge information, characterized in that: include: (1) The multimodal video is cut into multiple still images in chronological order, and the GTOT dataset is used as a training set. The training set is preprocessed by cropping and scaling according to the location and size of the target; (2) Construct a tracking model including a local gradient edge detector, a hybrid multi-expert dynamic selection unit, an edge adapter, and a backbone network: The local gradient edge detector includes depth change calculation, translation transformation and feature fusion, and is used to extract edge information of multimodal images; The hybrid multi-expert dynamic selection unit includes a router network and a multi-expert network, which is used to intelligently and dynamically enhance the edge information of thermal infrared and visible light to generate a better target representation; The edge adapter includes a multi-layer fully connected network for compressing and reconstructing edge features of targets in multimodal images; The backbone network is connected in parallel with the local gradient edge detector, the edge adapter and the hybrid multi-expert dynamic selection unit connected in series, and is used to extract the image features after multimodal fusion; (3) Input the preprocessed training set into the tracking model for iterative training; (4) Input the test set RGBT234 into the trained tracking model and output the target location information in each frame of the video.

2. The method according to claim 1, characterized in that The pre-processing of cropping and scaling the training set according to the location and size of the target in (1) includes: For two modal images, first expand them according to the size of the target bounding box, calculate the area of ​​the expanded region, take the square root of the expanded region, and get the side length of the square region to be cropped, and then crop it into a square region; The cropped square area is scaled, that is, for the template image, it is adjusted to 128×128 pixels, and for the search image, it is adjusted to 256×256 pixels.

3. The method according to claim 1, characterized in that The local gradient edge detector in (2) extracts edge information of the multimodal image, and its implementation includes: 2a) Obtain the maximum depth value of each position in the image through the depth change calculation to provide a depth map for subsequent translation transformation. The specific steps are as follows: First, use the sliding window method to obtain the image All window areas X patches : Among them, the window size win size =3; Then calculate the maximum value of each window area to obtain its maximum depth map The calculation formula is as follows: D max =max(X patches ), Among them, max(·) represents the maximum value operation. 2b) Obtain a multi-directionally translated depth map through the translation transformation to provide an offset depth map for gradient calculation in subsequent feature fusion. The specific steps are as follows: Use the affine transformation matrix to translate the image in four directions: up, down, left, and right. The translation matrix T is: in, is the horizontal translation, is the vertical translation; Then obtain the translated depth map D according to the translation matrix T shifted : D shifted =T×D max 2c) Through the feature fusion, the specific steps are as follows: Calculate the difference between the depth map before and after translation to obtain the gradient information G, and remove the unchanged area after translation through mask operation to obtain the edge information G final : G=D max -D shifted G final =G×M The mask M is a binary matrix that indicates the area where the depth changes. In order to obtain the global edge information, the gradient information in the four directions is spliced ​​to obtain the final gradient feature G global , and then further calculate the maximum value G of the edge information max As edge information of a multimodal search graph: G max =max(G global ) 2d) Finally, the calculated maximum gradient information G max As multimodal edge information.

4. The method according to claim 1, wherein The router network in (2) includes a linear layer network, a noise superposition operation and a Top-K selection operation. The linear layer network generates the assignment probability for each expert through the input edge features, and then selects the optimal expert network for each edge feature by superimposing noise and using the Top-K selection operation.

5. The method according to claim 1, characterized in that The multi-expert network in (2) includes multiple multi-layer perceptrons, each of which is a multi-layer perceptron used to perform specialized processing on edge features. Finally, the features processed by each expert are weighted and summed to obtain dynamically enhanced multimodal edge features.

6. The method according to claim 1, characterized in that The multi-layer fully connected network in (2) includes compression dimensionality reduction, hidden layer transformation and reconstruction optimization, that is, the multimodal edge features after dynamic enhancement are first compressed and reduced in dimension, and then the multimodal edge features after dimensionality reduction are transformed in the hidden layer, and finally the transformed output is reconstructed and optimized to obtain edge features that are adapted to the backbone network fusion.

7. The method according to claim 1, characterized in that The backbone network includes multiple layers of self-attention encoders. The output of each layer of encoders is fused with the output of its corresponding edge adapter. After multiple layers of fusion interaction, the target representation enhanced with multimodal edge information is extracted.

8. The method according to claim 1, characterized in that The (3) step of inputting the pre-processed training set into the tracking model for iterative training includes: 3a) Set the model training hyperparameters and the optimizer weight decay exponent to 10 -4 , the number of training rounds is 60, and the initial learning rate is 4×10 -5 . 3b) Perform weighted summation of the Focal loss function, L1 loss function, and GIoU loss function as the loss function of the tracking model, and calculate the loss of each iteration of the model; 3c) Backpropagation is used to calculate the gradient of the loss with respect to the parameters of each layer of the model. The model weights are iteratively updated using the gradient descent method until the maximum number of training times is reached. This results in a trained tracking model and saves the optimal model parameters.

9. The method according to claim 1, characterized in that In (4), the test set is input into the trained tracking model, and the position information of the target in each frame of the video is output. The implementation includes: 4a) Preprocessing the first frame of the multimodal RGBT234 video to be tracked using the same cropping and scaling method as used for the training set; 4b) The pre-processed multimodal image is used as the template image and search image for this tracking, and is input into the trained tracking model. The target features after edge information enhancement are output to obtain the tracking result of the current frame of the multimodal video; 4c) After each frame of tracking is completed, the next frame of the video image is cropped and scaled with the current predicted position as the center until the last frame of the video is tracked, thus completing the tracking of the entire video.

10. A visible light infrared target tracking system based on dynamic enhancement of edge information, characterized in that: include: Local gradient edge detection module, used to extract edge information of multimodal images; A hybrid multi-expert dynamic selection module is used to intelligently and dynamically enhance thermal infrared and visible light edge information to generate better target representations; Edge adaptation module, used to compress and reconstruct the edge features of targets in multimodal images; Backbone network module, used to extract image features after multimodal fusion; A tracking model building module is used to connect the backbone network module with the local gradient edge detection module, the edge adaptation module and the hybrid multi-expert dynamic selection module connected in series in parallel to build a tracking model; The training module is used to input the preprocessed training set into the tracking model for iterative training; The testing module is used to input the test set into the trained tracking model and output the location information of the target in each frame of the video.

Citation Information

Patent Citations

  • Infrared tracking method based on feature local correction and multi-modal channel sparse selection prompt

    CN119068217A

  • RGBT target tracking method combining asymmetric enhancement and interactive fusion

    CN120013990A