Multi-modal target tracking method and device, electronic equipment and computer program product

By fusing the feature modal domains of visible light and infrared light image frames in RGBT target tracking, and utilizing the cross-attention mechanism for difference injection and cross-modal interaction, a more robust feature enhancement domain is generated, solving the problem of tracking accuracy degradation caused by single-modal failure and achieving more stable and accurate target tracking.

CN121883532APending Publication Date: 2026-04-17CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing RGBT target tracking algorithms fail to fully utilize intermodal difference information when single-mode failure occurs in extreme environments, resulting in decreased tracking accuracy, limited noise suppression capabilities, and insufficient robustness.

Method used

By acquiring the feature modal domains of visible light and infrared light image frames, feature fusion and template fusion are performed. A cross-attention mechanism is used for difference injection and cross-modal interaction to generate a more robust feature enhancement domain, thereby enabling target tracking in visible light and infrared light image frames.

Benefits of technology

Improve tracking accuracy in complex scenarios such as low light or occlusion, avoid tracking drift and accuracy degradation caused by single-mode failure, and achieve more stable and accurate target tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883532A_ABST
    Figure CN121883532A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode target tracking method and device, electronic equipment and a computer program product. The method comprises: in at least one image frame group of a to-be-tracked video stream, obtaining a feature modal domain of the to-be-tracked image frame group, the feature modal domain comprising a first search domain and a first template domain corresponding to a visible light image frame, and a second search domain and a second template domain corresponding to an infrared light image frame; obtaining a feature fusion domain according to the first search domain and the second search domain, and obtaining a template fusion domain according to the first template domain and the second template domain; determining a feature enhancement domain according to the first search domain, the second search domain, the feature fusion domain and the template fusion domain; the target object is tracked in the visible light image frame according to the first enhancement domain, and the target object is tracked in the infrared light image frame according to the second enhancement domain. According to the invention, the technical problem of low tracking precision of the existing target tracking method is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, and more specifically, to a method, apparatus, electronic device, and computer program product for multimodal target tracking. Background Technology

[0002] In the field of RGBT target tracking technology, current multimodal tracking algorithms aim to fuse visible light images and thermal infrared images to improve target tracking performance in complex environments. However, the two modalities tracked in RGBT have different tracking difficulties. When a modality completely fails due to environmental interference, traditional fusion strategies still introduce noise features into the system, leading to a severe decline in tracking performance. Existing modal fusion mechanisms mostly use feature addition or attention mechanisms to achieve modal interaction, but these algorithms overemphasize the common features of the two modalities while ignoring the complementary value of the differences, failing to build a dynamic complementary association model, and resulting in insufficient information interaction capabilities between the two modal images.

[0003] Existing target tracking methods often rely on visible light features, which are prone to feature degradation under low light, strong interference, or severe weather conditions. Infrared features, on the other hand, can maintain contour feature stability even under extreme conditions. Therefore, combining complementary information from visible and infrared modes can overcome the limitations of single-mode tracking and achieve more robust tracking results. However, existing cross-modal fusion strategies are insufficient in their dynamic response to differences between modes, leading to a sharp drop in tracking accuracy when single-mode failures occur under extreme lighting conditions.

[0004] The TBSI model, as an advanced solution in this field, achieves accurate cross-modal information exchange through the Vision Transformer (ViT) architecture and a template bridging search region interaction approach, aiming to overcome the limitations of coarse modal fusion and insufficient context modeling in traditional methods. However, even with the innovations made by the TBSI model in multimodal fusion, existing technologies still face the following key challenges:

[0005] 1. Insufficient Utilization of Modal Difference Complementarity: When fusing information from two modalities, RGBT tracking algorithms often overemphasize the common features between them while neglecting the differences between the modalities. These differences are crucial for enhancing tracking robustness in complex scenarios such as insufficient lighting, excessive lighting, or thermal crossover. Existing technologies have failed to effectively construct dynamically complementary association models and fully utilize the differences between the two modal images, thus limiting their effectiveness in addressing single-modal failures.

[0006] 2. Limited noise suppression capability: In low-light or occluded scenarios at night, RGB images may contain a large amount of noise, and traditional modal fusion methods are not effective in handling this noise. Similarly, TIR images may also be affected by high-temperature environments, and existing technologies have failed to effectively remove noise in these situations, leading to a decline in tracking performance.

[0007] 3. Tracking robustness under modal failure: In extreme environments where single modal information is completely lost, such as in the dark at night or in high temperature environments, existing RGBT tracking systems can hardly rely on another mode for accurate tracking, because their fusion strategies often fail to fully consider the possibility of modal failure, resulting in a serious decline in tracking performance.

[0008] In summary, the current RGBT tracking algorithm is insufficient in utilizing the complementary information between modes, which leads to a significant decline in system tracking accuracy when a single mode fails.

[0009] There is currently no effective solution to the problem of low tracking accuracy in the existing target tracking methods. Summary of the Invention

[0010] This invention provides a method, apparatus, electronic device, and computer program product for multimodal target tracking, to at least solve the technical problem of low tracking accuracy in existing target tracking methods.

[0011] According to one aspect of the present invention, a method for multimodal target tracking is provided, comprising: acquiring a feature modal domain of at least one group of image frames in a video stream to be tracked, wherein each group of image frames includes: a visible light image frame and an infrared light image frame acquired at the same time for the same target object, the feature modal domain including: a first modal domain corresponding to the visible light image frame and a second modal domain corresponding to the infrared light image frame, the first modal domain including: a first search domain and a first template domain, the second modal domain including: a second search domain and a second template domain, the first template domain and the second template domain including features for tracking the target object; performing feature fusion on the first modal domain and the second modal domain to obtain a modal fusion domain, wherein the modal fusion domain includes The method comprises: a feature fusion domain obtained based on the first search domain and the second search domain, and a template fusion domain obtained based on the first template domain and the second template domain; a feature enhancement domain determined based on the first search domain, the second search domain, the feature fusion domain, and the template fusion domain, wherein the feature fusion domain is used for difference injection based on the differences between the first search domain and the second search domain, the template fusion domain is used for cross-modal interaction between the first search domain and the second search domain, and the feature enhancement domain includes at least: a first enhancement domain corresponding to the first search domain and a second enhancement domain corresponding to the second search domain; the first enhancement domain is used to track the target object in the visible light image frame, and the second enhancement domain is used to track the target object in the infrared light image frame.

[0012] Optionally, feature fusion of the first modal domain and the second modal domain to obtain a modal fusion domain includes: channel concatenation of the first modal domain and the second modal domain to obtain a spliced ​​feature domain. The first modal domain includes: visible light feature sets corresponding to multiple spatial locations in the visible light image frame, each visible light feature set including: visible light features extracted from the same spatial location using multiple first feature channels. The second modal domain includes: infrared light feature sets corresponding to the same multiple spatial locations in the infrared light image frame, each infrared light feature set including: infrared light features extracted from the same spatial location using multiple second feature channels. The number of channels in the first feature channels and the number of channels in the second feature channels are... The same quantity is used in the splicing feature domain, which includes: multiple splicing feature sets corresponding to the spatial positions respectively. Each splicing feature set includes: splicing features based on multiple fourth feature channels at the same spatial position. The fourth feature channel is either the first feature channel or the second feature channel. The splicing feature is either the visible light feature or the infrared light feature. The splicing feature domain is transformed in the channel dimension using a first linear layer to obtain the modal fusion domain. The modal fusion domain includes: multiple fusion feature sets corresponding to the spatial positions respectively. Each fusion feature set includes: fusion features based on multiple third feature channels at the same spatial position. The number of third feature channels is the same as that of the first feature channel.

[0013] Optionally, after concatenating the channels of the first modal domain and the second modal domain to obtain a spliced ​​feature domain, the method further includes: processing the spliced ​​feature domain using a channel attention mechanism to obtain a channel weight matrix, wherein the channel weight matrix includes: channel weights corresponding to multiple fourth feature channels, each channel weight representing the importance of each fourth feature channel; determining a channel feature domain based on the channel weight matrix and the spliced ​​feature domain, wherein the channel feature domain includes: a set of channel features corresponding to multiple spatial locations, each set of channel features including: channel features corresponding to multiple fourth feature channels at the same spatial location, each channel feature being the product of the channel weight corresponding to the same fourth feature channel and the spliced ​​feature; processing the channel feature domain using a spatial attention mechanism to obtain a spatial weight matrix. The system comprises: a spatial weight matrix, wherein the spatial weight matrix includes: position weights corresponding to multiple spatial locations, each position weight representing the importance of each spatial location; a position feature domain is determined based on the spatial weight matrix and the channel feature domain, wherein the position feature domain includes: a set of position features corresponding to multiple spatial locations, each set of position features including: position features corresponding to multiple fourth feature channels at the same spatial location, each position feature being the product of the position weight and the channel feature corresponding to the same spatial location; a second linear layer is used to perform a weighted summation of the spliced ​​feature domain and the position feature domain to obtain the modal fusion domain, wherein the fusion feature in the modal fusion domain is: the result of weighted merging of the spliced ​​features in the spliced ​​feature domain and the position features in the position feature domain by channel dimension.

[0014] Optionally, determining the feature enhancement domain based on the first search domain, the second search domain, the feature fusion domain, and the template fusion domain includes: processing the first search domain, the second search domain, and the feature fusion domain using a cross-attention mechanism to obtain a modality complementation domain, wherein the cross-attention mechanism is used to generate a cross-modal similarity domain based on the first search domain, the second search domain, and the feature fusion domain, the cross-modal similarity domain including: a first modal similarity domain representing the difference between the first search domain and the second search domain, and a second modal similarity domain representing the difference between the second search domain and the first search domain, the modality complementation domain including: a first complementation domain obtained by injecting a difference into the first search domain based on the first modal similarity domain, and a second complementation domain obtained by injecting a difference into the second search domain based on the second modal similarity domain; processing the first complementation domain, the second complementation domain, and the template fusion domain using a cross-attention mechanism to obtain the feature enhancement domain, wherein the cross-attention mechanism is used to perform bidirectional cross-modal interaction between the first complementation domain and the second complementation domain using the template fusion domain as a medium, the feature enhancement domain including at least: a first enhancement domain corresponding to the first complementation domain, and a second enhancement domain corresponding to the first complementation domain.

[0015] Optionally, processing the first search domain, the second search domain, and the feature fusion domain using a cross-attention mechanism to obtain the modality supplementation domain includes at least the following: using the cross-attention mechanism, taking the first search domain as a query vector, the second search domain as a key vector, and the feature fusion domain as a value vector to obtain a first modality similarity domain; determining a first modality difference domain based on the difference between the feature fusion domain and the first modality similarity domain; performing layer normalization processing after superimposing the first search domain and the first modality difference domain to obtain a first modality correction domain; and performing layer normalization processing after superimposing the first modality correction domain using a multilayer perceptron to obtain the first supplementation domain.

[0016] Optionally, processing the first supplementary domain, the second supplementary domain, and the template fusion domain using a cross-attention mechanism to obtain a feature enhancement domain includes at least: using the cross-attention mechanism, using the template fusion domain as a query vector and the first supplementary domain as a key vector and a value vector to obtain first search context information based on the first supplementary domain; enriching the first search context information into the template fusion domain to obtain a first template enhancement domain; and using the cross-attention mechanism, using the first template enhancement domain as a key vector and a value vector and the second supplementary domain as a query vector to obtain a second enhancement domain.

[0017] Optionally, the feature enhancement domain further includes: a template update domain configured for the next image frame group of the image frame group to be tracked, wherein the template update domain includes: an updated first template domain and an updated second template domain; processing the first supplementary domain, the second supplementary domain, and the template fusion domain using a cross-attention mechanism to obtain the feature enhancement domain includes at least: using the cross-attention mechanism, using the first template domain as a query vector, using the template fusion domain as a key vector and a value vector to obtain first template context information, and enriching the first template context information into the first template domain to obtain an updated first template domain; using the cross-attention mechanism, using the second template domain as a query vector, using the template fusion domain as a key vector and a value vector to obtain second template context information, and enriching the second template context information into the second template domain to obtain an updated second template domain.

[0018] According to another aspect of the present invention, a multimodal target tracking device is also provided, comprising: an acquisition module, configured to acquire a feature modal domain of at least one image frame group of a video stream to be tracked, wherein each image frame group includes: a visible light image frame and an infrared light image frame acquired at the same time for the same target object, the feature modal domain including: a first modal domain corresponding to the visible light image frame and a second modal domain corresponding to the infrared light image frame, the first modal domain including: a first search domain and a first template domain, the second modal domain including: a second search domain and a second template domain, the first template domain and the second template domain including features for tracking the target object; and a feature fusion module, configured to perform feature fusion on the first modal domain and the second modal domain to obtain a modal fusion domain, wherein the modal fusion domain includes The system comprises: a feature fusion domain obtained based on the first search domain and the second search domain, and a template fusion domain obtained based on the first template domain and the second template domain; a feature enhancement module, configured to determine a feature enhancement domain based on the first search domain, the second search domain, the feature fusion domain, and the template fusion domain, wherein the feature fusion domain is used to perform difference injection by combining the differences between the first search domain and the second search domain, and the template fusion domain is used to perform cross-modal interaction between the first search domain and the second search domain, and the feature enhancement domain includes at least: a first enhancement domain corresponding to the first search domain and a second enhancement domain corresponding to the second search domain; and a tracking module, configured to track the target object in the visible light image frame based on the first enhancement domain and to track the target object in the infrared light image frame based on the second enhancement domain.

[0019] According to another aspect of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to perform the above-described multimodal target tracking method through the computer program.

[0020] According to another aspect of the present invention, a computer program product is also provided, including computer instructions that, when executed by a processor, implement the steps of the above-described multimodal target tracking method.

[0021] In the embodiments described above, by acquiring visible light image frames and infrared light image frames of the target object simultaneously, a feature modal domain for target tracking is obtained. This feature modal domain includes a first modal domain and a second modal domain. A feature fusion domain is obtained by fusing features from a first search domain in the first modal domain and a second search domain in the second modal domain. Similarly, a template fusion domain is obtained by fusing features from a first template domain in the first modal domain and a second template domain in the second modal domain. This process leverages the complementarity of visible and infrared light to construct a high-quality joint representation. Then, the feature fusion domain is used to identify and perform differential analysis on the modal differences between visible and infrared light. By injecting information and utilizing the template fusion domain to perform cross-modal interaction between the first and second search domains, more robust first and second enhancement domains are generated. Based on the first and second enhancement domains, the target object is tracked in visible light and infrared light image frames according to the feature-enhanced first and second enhancement domains, respectively. This can effectively improve the tracking accuracy in complex scenes such as low light or occlusion, avoid tracking drift and accuracy degradation caused by single-modal failure, and achieve more stable and accurate target tracking. This achieves the technical effect of high tracking accuracy in multimodal target tracking, and solves the technical problem of low tracking accuracy in existing target tracking methods. Attached Figure Description

[0022] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0023] Figure 1 This is a flowchart of a multimodal target tracking method according to an embodiment of the present invention;

[0024] Figure 2 This is a schematic diagram of an improved TBSI model according to an embodiment of the present invention;

[0025] Figure 3 This is a schematic diagram of an adaptive feature fusion module according to an embodiment of the present invention;

[0026] Figure 4 This is a schematic diagram of a channel attention mechanism according to an embodiment of the present invention;

[0027] Figure 5 This is a schematic diagram of a spatial attention mechanism according to an embodiment of the present invention;

[0028] Figure 6 This is a schematic diagram of a difference information injection module according to an embodiment of the present invention;

[0029] Figure 7 This is a schematic diagram of a cross-modal interaction module according to an embodiment of the present invention;

[0030] Figure 8 This is a schematic diagram of a multimodal target tracking device according to an embodiment of the present invention;

[0031] Figure 9 This is a structural block diagram of a computer terminal according to an embodiment of the present invention. Detailed Implementation

[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0034] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0035] RGBT (RGB-Thermal): A multimodal image data format that integrates visible light (RGB) and thermal infrared (Thermal) image information. The RGB image reflects the appearance features of the target, such as texture and color, while the thermal infrared image reflects the temperature radiation characteristics of the target. It can achieve complementary target information in complex scenarios such as changes in lighting and occlusion.

[0036] Multi-Modal Object Tracking (MPT) is a technique that improves the robustness of target tracking by fusing image data from two or more modalities and leveraging the complementarity of each modality. Unlike single-modal tracking, it can effectively address scenarios where a single modality fails (such as RGB images failing at night or thermal infrared images failing under high-temperature interference).

[0037] According to an embodiment of the present invention, a method embodiment for multimodal target tracking is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0038] Figure 1 This is a flowchart of a multimodal target tracking method according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:

[0039] Step S102: In at least one image frame group of the video stream to be tracked, the feature modal domain of the image frame group to be tracked is obtained. Each image frame group includes: a visible light image frame and an infrared light image frame acquired at the same time for the same target object. The feature modal domain includes: a first modal domain corresponding to the visible light image frame and a second modal domain corresponding to the infrared light image frame. The first modal domain includes: a first search domain and a first template domain. The second modal domain includes: a second search domain and a second template domain. The first template domain and the second template domain include features for tracking the target object.

[0040] Step S104: Perform feature fusion on the first modal domain and the second modal domain to obtain a modal fusion domain, wherein the modal fusion domain includes: a feature fusion domain obtained based on the first search domain and the second search domain, and a template fusion domain obtained based on the first template domain and the second template domain;

[0041] Step S106: Based on the first search domain, the second search domain, the feature fusion domain, and the template fusion domain, determine the feature enhancement domain. The feature fusion domain is used to perform difference injection by combining the differences between the first search domain and the second search domain. The template fusion domain is used to perform cross-modal interaction between the first search domain and the second search domain. The feature enhancement domain includes at least: a first enhancement domain corresponding to the first search domain and a second enhancement domain corresponding to the second search domain.

[0042] Step S108: Track the target object in the visible light image frame based on the first enhancement domain, and track the target object in the infrared light image frame based on the second enhancement domain.

[0043] In the embodiments described above, by acquiring visible light image frames and infrared light image frames of the target object simultaneously, a feature modal domain for target tracking is obtained. This feature modal domain includes a first modal domain and a second modal domain. A feature fusion domain is obtained by fusing features from a first search domain in the first modal domain and a second search domain in the second modal domain. Similarly, a template fusion domain is obtained by fusing features from a first template domain in the first modal domain and a second template domain in the second modal domain. This process leverages the complementarity of visible and infrared light to construct a high-quality joint representation. Then, the feature fusion domain is used to identify and perform differential analysis on the modal differences between visible and infrared light. By injecting information and utilizing the template fusion domain to perform cross-modal interaction between the first and second search domains, more robust first and second enhancement domains are generated. Based on the first and second enhancement domains, the target object is tracked in visible light and infrared light image frames according to the feature-enhanced first and second enhancement domains, respectively. This can effectively improve the tracking accuracy in complex scenes such as low light or occlusion, avoid tracking drift and accuracy degradation caused by single-modal failure, and achieve more stable and accurate target tracking. This achieves the technical effect of high tracking accuracy in multimodal target tracking, and solves the technical problem of low tracking accuracy in existing target tracking methods.

[0044] The multimodal target tracking method described in this application uses a "dual-modal data acquisition - real-time calculation - result output / storage" as its core hardware architecture, which is suitable for typical scenarios such as nighttime security, vehicle-mounted assisted driving, and drone inspection. The hardware network elements and communication connections it relies on include: dual-modal image acquisition equipment, computing unit, and storage / output unit.

[0045] Optionally, the dual-modal image acquisition device can acquire the video stream to be tracked. The type of the dual-modal image acquisition device is: an RGB-TIR synchronous integrated camera (including a visible light CMOS sensor + an infrared thermal imaging sensor); the device form includes: fixed (such as a security monitoring camera, resolution 2K / 4K, frame rate 30fps), mobile (such as a vehicle-mounted forward-looking camera, resolution 1080P, frame rate 60fps), and portable (such as a drone-mounted camera, resolution 720P, frame rate 30fps); the specific functions include: synchronously acquiring RGB images (capturing texture and color) and TIR images (capturing temperature radiation, resisting low light / occlusion) of the target scene, and transmitting the raw data to the computing unit via Ethernet (fixed), PCIe 4.0 (vehicle-mounted), 4G / 5G (drone), with a transmission latency ≤50ms, ensuring dual-modal frame timing alignment (error ≤1ms); the communication relationship is: acquisition device → computing unit (real-time transmission of dual-modal image stream).

[0046] Optionally, the computing unit is used to perform target object tracking tasks based on visible light image frames and infrared light image frames in the video stream to be tracked. The type of computing unit can be: embedded processor (such as NVIDIA Jetson AGXXavier, adapted for vehicle / drone) or GPU server (such as NVIDIA A10, adapted for fixed security). The computing functions include: running core algorithms such as "dual-modal feature extraction, information fusion, channel-spatial attention optimization, and tracking decision", processing image data transmitted by the acquisition device, and outputting target coordinates and tracking confidence. The communication relationship is: computing unit → storage unit (asynchronously storing image stream + tracking log), computing unit → output unit (real-time transmission of tracking results).

[0047] Optionally, the storage / output unit can store the software for the aforementioned target tracking modality. Storage devices include: SSDs (mobile temporary scene storage, write speed 500MB / s) and NAS (fixed scene historical data storage, capacity 10TB+). Output devices include: monitoring displays (such as large screens in security centers, displaying target frames in real time), vehicle-mounted warning modules (such as voice prompts + dashboard pop-ups), and drone ground stations (for real-time transmission of tracking status). Communication: Both the storage / output unit and the computing unit communicate via TCP / IP protocol to ensure data reliability.

[0048] Optionally, the software architecture for performing modal target tracking is as follows: using a 12-layer ViT as the multimodal backbone network, and adopting the process of "input preprocessing - feature extraction - cross-modal interaction - tracking prediction".

[0049] In step S102 above, the video stream to be tracked can be an image acquired by a visible light image acquisition device and an infrared light image acquisition device. The visible light image frame and the infrared light image frame acquired at the same time for the same target object can be a group of image frames.

[0050] In step S102 above, the visible light image frames and infrared light image frames in the image frame group to be tracked can be image frames used to track the target object. By performing feature extraction on the image frame group to be tracked, a first search domain and a second search domain for tracking the target object can be obtained.

[0051] In step S102 above, the first template domain and the second template domain are pre-labeled features for tracking the target object. The first template domain and the second template domain can be based on the visible light image frame and infrared light image frame of the target object that have been pre-labeled.

[0052] Optionally, the first template domain and the second template domain can also be determined based on the tracking results of the previous image frame group of the image frame group to be tracked.

[0053] In step S102 above, the first modal domain may be a first feature matrix including: a plurality of visible light features, wherein the rows and columns of the first feature matrix represent the feature channels for acquiring visible light features and the feature positions of the acquired visible light features in the visible light image frame, respectively.

[0054] In step S102 above, the second modal domain may be a second feature matrix including: multiple infrared light features, wherein the rows and columns of the second feature matrix represent the feature channels for acquiring infrared light features and the feature positions of the acquired infrared light features in the infrared light image frame, respectively.

[0055] In step S102 above, the feature modality domain is obtained by feature extraction of image frames (such as visible light image frames and infrared light image frames). The specific process of feature extraction includes: dividing the image frame into fixed image blocks, then converting them into token features through modality-specific linear projection, and adding shared learnable position codes to achieve frame alignment of visible light image frames and infrared light image frames.

[0056] It should be noted that the segmented image blocks can be encoded, and image blocks with the same encoding correspond to the same position in visible light image frames and infrared light image frames.

[0057] The multimodal target tracking method described in this application can be implemented through an improved Template Bridging Search Region Interaction (TBSI).

[0058] Figure 2 This is a schematic diagram of an improved TBSI model according to an embodiment of the present invention, as shown below. Figure 2As shown, the above-mentioned multimodal target tracking method can be applied to the field of RGBT target tracking, implemented through an improved TBSI (Template-BridgedSearch Region Interaction) model. TBSI is based on the Vision Transformer (ViT) architecture and focuses on cross-modal interaction optimization. Its core objective is to solve the problems of coarse-grained modal fusion and insufficient context modeling in traditional methods. The specific technical path and implementation logic are as follows:

[0059] From an overall framework perspective, TBSI uses a 12-layer ViT multimodal backbone network and adopts a process of "input preprocessing - feature extraction - cross-modal interaction - tracking prediction". In the input stage, the model segments the search frames (256×256 resolution) and template frames (128×128 resolution) of the RGB and TIR modalities into fixed-size image patches. These patches are then transformed into token features through modality-specific linear projection and incorporating shared, learnable positional encodings (frame alignment has been implemented in the dataset) to form the RGB search domain. (such as the first search domain), RGB template domain (such as the first template domain), and the search domain corresponding to the TIR mode. (such as the second search domain), template domain (e.g., the second template domain). Subsequently, the search domains (e.g., the first and second search domains) and template domains (e.g., the first and second template domains) of each modality are concatenated and input into the ViT Transformer block with shared parameters. Cross-attention is used to complete the "search-template matching". At the same time, the core TBSI module is inserted into the 4th, 7th and 10th layers of ViT. Finally, the output RGB and TIR search region features are concatenated and input into the tracking head to achieve target bounding box prediction.

[0060] It's important to note that token features refer to a series of fixed-length vector representations transformed from input data. ViT-type models divide an image into a series of non-overlapping small image patches, each of which is treated as a token. Then, a linear projection layer (sometimes including additional layers such as normalization layers) compresses the pixel values ​​of each small image patch into a fixed-length vector, which is called the token feature. Token features also include positional encoding to preserve the relative position information of the image patches within the original image.

[0061] Optionally, in the ViT-based multimodal target tracking model, the input RGB and TIR image frames are first segmented into multiple image patches. Each image patch is linearly projected to generate corresponding token features. These token features, along with positional codes, are then fed into the Transformer encoder for cross-attention calculation, thereby achieving interaction and information fusion between features. Such token features are a key component for achieving advanced target detection and tracking tasks in multimodal tracking systems.

[0062] Optionally, the above-mentioned multimodal target tracking method can be implemented by an improved TBSI module, which is specifically divided into three parts: an adaptive feature fusion module, a differential information injection module, and a cross-modal interaction module. Specifically, step S104 is implemented by the adaptive feature fusion module, and steps S106 and S108 are implemented by the differential information injection module and the cross-modal interaction module.

[0063] In the embodiments described above, the adaptive feature fusion module, through the combined action of channel attention and spatial attention, achieves adaptive weight allocation for visible light and infrared modal information, constructs a joint representation for noise suppression in the feature space, and obtains high-quality fused feature information (such as a feature fusion domain). Through the difference information injection module, based on the fused feature information (such as a template fusion domain), the difference information between different modalities is extracted, and the specific information between modalities is supplemented to the visible light and infrared branches, thereby enhancing the model's ability to utilize multimodal features.

[0064] Optionally, the improved TBSI module, building upon the original function of "achieving precise cross-modal interaction through templates," incorporates the utilization of difference information. Before cross-modal interaction, an adaptive feature fusion module is designed first. This module concatenates the input RGB and TIR information and applies channel attention and spatial attention weights to obtain fused information. This achieves adaptive weight allocation for visible and infrared modal template information, constructing a joint representation for noise suppression in the feature space to obtain high-quality fused feature information. Second, a difference information injection module is designed. Based on the fused information, it extracts difference information between different modalities, supplementing the visible and infrared branches with modal-specific information, enhancing the model's ability to utilize multimodal features. Finally, the updated fused template domain information and the search domain information of the two branches are used in conjunction with the previous TBSI operation for bidirectional template bridging interaction and template updating.

[0065] The embodiments described above in this application, through the design of a cross-modal mutual enhancement network, fully explore and utilize the differential information between different modes, enhance the model's ability to utilize multimodal features, thereby improving the robustness and adaptability of the tracking system and solving the problem of target tracking performance degradation caused by single-modal failure.

[0066] As an optional implementation, in the case of multimodal target tracking, the loss function used to train the improved TBSI model is similar to the training method of TBSI. This loss function combines regression loss and classification loss, with bounding box regression loss using L1 loss and GIoU loss, denoted as follows: and The classification loss uses weighted focus loss, denoted as... .

[0067] Optionally, total loss It can be represented as: Among them, the regularization parameter and The values ​​are set to 5 and 2 respectively. For the GIoU loss, assuming A is the predicted bounding box, B is the ground truth bounding box, and G is the minimum closure area of ​​these two boxes (i.e., the area containing the minimum bounding box of both A and B), the GIoU loss can be obtained. Calculation formula: .

[0068] Optionally, for the weighted focus loss, for the center point of each ground truth box... A Gaussian kernel can be used to generate the target response map of the true bounding box. : ,in, The standard deviation of the object size is adaptive.

[0069] Alternatively, the weighted focus loss can be expressed as:

[0070] ;

[0071] Among them, hyperparameters and Set them to 2 and 4 respectively.

[0072] It should be noted that the data needs to be preprocessed before training. The size of the search region of the input image is set to 256×256, and the size of the template region is set to 128×128. This chapter uses ViT as the backbone network, connecting 12 Transformer modules, and inserts the designed DIMTBSI module into layers 4, 7, and 10 of the backbone network.

[0073] It should be noted that during the offline training phase, the LasHeR dataset was used as the training data. After training, the datasets were directly tested on the LasHeR test set, the RGBT210 dataset, and the RGBT234 dataset. The experiment was set to 18 training epochs with a batch size of 16. The learning rate of the ViT backbone network was set to 0.000005, and the learning rates of other parameters were also set to 0.00005. The learning rate decayed by a factor of 10 after 8 epochs. AdamW was chosen as the optimizer, with a weight decay of 0.0001.

[0074] Optionally, the adaptive feature fusion module can fuse visible light search region information (such as the first modal domain) with infrared search region information (such as the second modal domain) to obtain a fused modal domain.

[0075] As an optional embodiment, feature fusion of the first modal domain and the second modal domain to obtain the modal fusion domain includes: channel concatenation of the first modal domain and the second modal domain to obtain a spliced ​​feature domain. The first modal domain includes: a set of visible light features corresponding to multiple spatial locations in a visible light image frame, each visible light feature set including: visible light features extracted from the same spatial location using multiple first feature channels. The second modal domain includes: a set of infrared light features corresponding to the same multiple spatial locations in an infrared light image frame, each infrared light feature set including: infrared light features extracted from the same spatial location using multiple second feature channels. The first feature channels and the second feature channels... The number of channels in each feature field is the same. The splicing feature domain includes: multiple splicing feature sets corresponding to different spatial locations. Each splicing feature set includes: splicing features based on multiple fourth feature channels at the same spatial location. The fourth feature channel is either the first feature channel or the second feature channel. The splicing features are visible light features or infrared light features. The splicing feature domain is transformed in the channel dimension using the first linear layer to obtain the modal fusion domain. The modal fusion domain includes: multiple fusion feature sets corresponding to different spatial locations. Each fusion feature set includes: fusion features based on multiple third feature channels at the same spatial location. The number of third feature channels is the same as that of the first feature channels.

[0076] In the embodiments described above, the adaptive feature fusion module can intelligently integrate the feature domains of two modalities to construct a high-quality fused feature representation, providing a purer feature base for the subsequent differential information injection module, thereby improving the robustness and accuracy of the entire tracking system. This fusion process can effectively suppress noise and irrelevant features, strengthen key features related to the target, and lay a solid foundation for the extraction and utilization of differential information between different modalities.

[0077] Optionally, the first modal domain and the second modal domain are channel-concatenated to achieve the splicing of the first modal domain and the second modal domain, resulting in a spliced ​​feature domain. The number of feature channels in the spliced ​​feature domain is the sum of the number of feature channels in the first modal domain and the second modal domain. The first linear layer can perform feature transformation on the spliced ​​feature domain in the channel dimension to obtain a modal fusion domain with the same number of feature channels as the first modal domain.

[0078] Optionally, the first feature channel, the second feature channel, and the third feature channel have the same number of channels.

[0079] It should be noted that although multimodal approaches can provide rich information from different perspectives, blindly fusing information may cause modal interference in the tracking process, leading to poor results. Therefore, the input features are weighted and optimized by combining channel attention and spatial attention.

[0080] As an optional embodiment, after concatenating the channels of the first and second modal domains to obtain the spliced ​​feature domain, the method further includes: processing the spliced ​​feature domain using a channel attention mechanism to obtain a channel weight matrix, wherein the channel weight matrix includes: channel weights corresponding to multiple fourth feature channels, each channel weight representing the importance of each fourth feature channel; determining the channel feature domain based on the channel weight matrix and the spliced ​​feature domain, wherein the channel feature domain includes: multiple sets of channel features corresponding to multiple spatial locations, each set of channel features including: channel features corresponding to multiple fourth feature channels at the same spatial location, each channel feature being the product of the channel weight corresponding to the same fourth feature channel and the spliced ​​feature; and processing the channel feature domain using a spatial attention mechanism. A spatial weight matrix is ​​obtained, which includes position weights corresponding to multiple spatial locations, each position weight representing the importance of each spatial location. Based on the spatial weight matrix and the channel feature domain, a position feature domain is determined, which includes a set of position features corresponding to multiple spatial locations. Each set of position features includes position features corresponding to multiple fourth feature channels at the same spatial location, and each position feature is the product of the position weight and the channel feature corresponding to the same spatial location. The second linear layer is used to perform a weighted summation of the concatenated feature domain and the position feature domain to obtain the modal fusion domain. The fusion feature in the modal fusion domain is the result of weighted merging of the concatenated features in the concatenated feature domain and the position features in the position feature domain by channel dimension.

[0081] In the embodiments described above, after concatenating the first and second modal domains to obtain a spliced ​​feature domain, a channel attention mechanism can be used to process the spliced ​​feature domain to generate a channel weight matrix containing the channel weights of multiple fourth feature channels, thereby characterizing the relative importance of each channel. Through the synergistic effect of the channel weight matrix and the spliced ​​feature domain, a channel feature domain containing a set of channel features at multiple spatial locations is determined. Then, a spatial attention mechanism is used to process the channel feature domain to obtain a spatial weight matrix, which contains the position weights of multiple spatial locations, reflecting the importance of each position. Based on the spatial weight matrix and the channel feature domain, a position feature domain is determined. Then, a second linear layer is used to perform a feature weighted summation of the spliced ​​feature domain and the position feature domain to obtain the modal fusion domain. Thus, by adaptively allocating the channel weights and spatial weights of the channels, efficient fusion and complementarity between modalities are achieved, enhancing the robustness of the tracking system and ensuring that high tracking accuracy can be maintained even in the event of single-modal failure.

[0082] Figure 3 This is a schematic diagram of an adaptive feature fusion module according to an embodiment of the present invention, as shown below. Figure 3 As shown, first, the infrared search area information (such as the second search mode domain) is... Information related to the visible light search region (such as the first search mode domain). The splicing is performed to obtain the spliced ​​feature domain. Then, the channel weights (i.e., the channel weight matrix) are calculated through the channel attention mechanism, and the spatial weights (i.e., the spatial weight matrix) are calculated through the spatial attention mechanism. Finally, the key features are strengthened through weighted operations.

[0083] Optionally, the first and second modal domains are channel-concatenated to obtain a spliced ​​feature domain, the formula of which is expressed as: Where Concat is feature concatenation.

[0084] Optionally, the concatenated feature domain is processed using a channel attention mechanism to obtain the channel weight matrix; and the channel feature domain is determined based on the channel weight matrix and the concatenated feature domain, as expressed by the following formula: ,in, It is the channel weight matrix. It is global average pooling. It is a fully connected layer. It is the channel feature domain. It is an activation function.

[0085] Optionally, a spatial attention mechanism is used to process the channel feature domain to obtain a spatial weight matrix; and based on the spatial weight matrix and the channel feature domain, the position feature domain is determined, as expressed by the following formula: ,in, It is a spatial weight matrix. It is max pooling. It is a convolutional layer. It is the location feature domain. It is an activation function.

[0086] Optionally, the second linear layer is used to perform a weighted summation of the spliced ​​feature domain and the positional feature domain to obtain the modality fusion domain, the formula of which is expressed as: Wherein, Linear is a linear layer. It is a modal fusion domain. .

[0087] It's important to note that the attention mechanism mimics the selective focusing ability of the human visual system on key information, enabling neural networks to dynamically focus on critical regions of the input data and suppress interference from irrelevant information. Essentially, it uses learnable weight allocation strategies to assign different importance to different feature elements, thereby enhancing the model's ability to capture core information. In image processing tasks, the attention mechanism is widely used to enhance the discriminative power of features, with channel attention and spatial attention being particularly common.

[0088] Figure 4 This is a schematic diagram of a channel attention mechanism according to an embodiment of the present invention, such as... Figure 4 As shown, the core idea of ​​the channel attention mechanism is to assign different weights to each feature map channel, and enhance key information and suppress noise interference by dynamically evaluating the importance of different channels. Figure 1 shows the network structure diagram of the typical channel attention mechanism SENet

[65] . The method flow is as follows: First, the input features are subjected to global average pooling (GAP) to generate a one-dimensional channel description vector, where each element represents the global response intensity of the corresponding channel; then, the nonlinear dependency relationship between channels is learned through the fully connected layer to generate the channel weight vector; finally, the channel weight vector is multiplied element-wise with the original features after the dimension is expanded by vector broadcasting to strengthen key channels and suppress redundant channels.

[0089] In the embodiments described above, the channel attention mechanism primarily focuses on the information weight allocation of different channels in the feature map. In a standard CNN architecture, the feature map output by each layer consists of multiple channels, each typically corresponding to a different feature (e.g., edge, texture, color, etc.). The channel attention mechanism introduces an additional module to learn the relative importance of each channel, and then performs a weighted summation of the channels based on these importance weights, thereby enhancing the representation of key features and suppressing irrelevant or noisy features.

[0090] Figure 5This is a schematic diagram of a spatial attention mechanism according to an embodiment of the present invention, such as... Figure 5 As shown, spatial attention acts on the spatial dimension of the feature map, guiding the model to focus on key target regions by generating a two-dimensional weight mask. The figure illustrates the network structure of the spatial attention mechanism, and the process is as follows: First, the feature map's channel dimension is compressed using methods such as global average pooling and max pooling; then, spatial dependencies are learned through convolutional layers to generate a spatial attention matrix; finally, the attention matrix is ​​expanded in dimension through vector broadcasting and multiplied pixel-by-pixel with the original features to highlight the target region and weaken background interference.

[0091] In the embodiments described above, spatial attention operations focus on the weighting of information at different locations in the feature map. Compared to channel attention, which focuses on the global information of each channel, spatial attention emphasizes features at local or specific locations, dynamically adjusting the feature weights of these regions by identifying which regions are more important to the task (such as classification, detection, and segmentation).

[0092] For example, in object detection, the location of the target may be the most important, while background or irrelevant object regions may become noise, interfering with the model's judgment. Spatial attention mechanisms learn the importance of each location and assign a weight to each location (pixel or feature block) on the feature map, thereby enhancing the feature representation of key regions while suppressing the influence of unimportant regions. In this way, even if there are complex backgrounds or interference around the target, the model can focus more on the target and improve detection accuracy.

[0093] It should be noted that the "channel attention mechanism" and the "spatial attention mechanism" are important mechanisms used for feature learning and feature mapping in deep learning models, especially convolutional neural networks (CNNs) and Transformer models. They allow the model to dynamically adjust the weights of features based on the inherent characteristics of the input data, thereby improving the model's performance and adaptability.

[0094] In RGBT target tracking tasks, "channel attention mechanism" and "spatial attention mechanism" are often used in cross-modal fusion modules. The channel attention mechanism can dynamically suppress redundant noise, while the spatial attention mechanism can accurately eliminate background interference. Together, they improve the discriminativeness and anti-interference ability of the fused features, enabling the model to maintain stable tracking performance in complex scenes.

[0095] As an alternative example, in RGBT target tracking, channel attention can be used to adaptively adjust the weights of RGB and TIR feature channels, ensuring that the model utilizes the most effective target features under different lighting and environmental conditions. Spatial attention, on the other hand, can highlight the target's location, helping the model accurately locate and track the target even if it is partially occluded or unevenly lit in the image. By combining these two mechanisms, the tracking system can more intelligently process multimodal information, improving performance in various challenging scenarios.

[0096] It should be noted that the feature fusion domain is determined based on the first search domain and the second search domain. Therefore, the feature fusion domain includes all features of the first search domain and the second search domain, and can serve as a communication bridge between the first search domain and the second search domain. Thus, when using the feature fusion domain to combine the differences between the first search domain and the second search domain for difference injection, the process includes: determining the differences between the first search domain and the second search domain, and then using the feature fusion domain as a bridge to inject the differences between the first search domain and the second search domain into the first search domain, or to inject the differences between the second search domain and the first search domain into the second search domain.

[0097] As an optional embodiment, determining the feature enhancement domain based on the first search domain, the second search domain, the feature fusion domain, and the template fusion domain includes: processing the first search domain, the second search domain, and the feature fusion domain using a cross-attention mechanism to obtain a modality complementation domain, wherein the cross-attention mechanism is used to generate a cross-modal similarity domain based on the first search domain, the second search domain, and the feature fusion domain, the cross-modal similarity domain including: a first modal similarity domain representing the difference between the first search domain and the second search domain, and a second modal similarity domain representing the difference between the second search domain and the first search domain; the modality complementation domain including: a first complementation domain obtained by injecting a difference into the first search domain based on the first modal similarity domain, and a second complementation domain obtained by injecting a difference into the second search domain based on the second modal similarity domain; processing the first complementation domain, the second complementation domain, and the template fusion domain using the cross-attention mechanism to obtain the feature enhancement domain, wherein the cross-attention mechanism is used to perform bidirectional cross-modal interaction between the first complementation domain and the second complementation domain using the template fusion domain as a medium, the feature enhancement domain including at least: a first enhancement domain corresponding to the first complementation domain, and a second enhancement domain corresponding to the first complementation domain.

[0098] In the embodiments described above, based on the feature fusion domain, cross-modal similarity information and difference template information are calculated to supplement and enhance the original RGB (such as the first search domain) and TIR (such as the second search domain) branches, resulting in the first and second supplementary domains injected with difference information. This fully explores and utilizes the specific information between the two modalities, significantly enhancing the model's ability to utilize multimodal features. Then, using a bidirectional template bridging interaction mechanism, with the template fusion domain as the medium, a cross-attention mechanism is used to realize bidirectional cross-modal interaction between the updated RGB (such as the first supplementary domain) and TIR (such as the second supplementary domain) search regions, further improving the robustness and accuracy of target tracking. This effectively overcomes the tracking problem caused by single-modal failure, especially in complex environments such as insufficient lighting, excessive lighting, or thermal crossover, demonstrating strong anti-interference ability and tracking stability, meeting the needs of real-time tracking.

[0099] Optionally, the modality supplementation domain can be obtained by processing the first search domain, the second search domain, and the feature fusion domain using a cross-attention mechanism, which can be implemented through the difference information injection module; the feature enhancement domain can be obtained by processing the first supplementation domain, the second supplementation domain, and the template fusion domain using a cross-attention mechanism, which can be implemented through the cross-modal interaction module.

[0100] As an optional embodiment, the modality supplementation domain is obtained by processing the first search domain, the second search domain, and the feature fusion domain using a cross-attention mechanism, which includes at least the following steps: using the cross-attention mechanism, the first search domain is used as a query vector, the second search domain as a key vector, and the feature fusion domain as a value vector to obtain a first modality similarity domain; a first modality difference domain is determined based on the difference between the feature fusion domain and the first modality similarity domain; the first search domain and the first modality difference domain are superimposed and then subjected to layer normalization to obtain a first modality correction domain; the first modality correction domain is processed by a multilayer perceptron and then superimposed with the first modality correction domain, and then subjected to layer normalization after superposition to obtain a first supplementation domain.

[0101] The embodiments described above utilize a cross-attention mechanism to process the first search domain, the second search domain, and the feature fusion domain. This aims to capture and utilize intermodal difference information to enhance the robustness and accuracy of target tracking. Essentially, it searches for features in the cross-modal search domain that match those in the fusion domain, thereby extracting common information between the two modalities in the current scene. Based on the difference between the feature fusion domain and the first modal similarity domain, the first modal difference domain is determined, representing the difference between the two modal features. This difference information contains complementary characteristics, crucial for improving tracking performance in single-modal failure environments. Subsequently, the first search domain and the first modal difference domain are superimposed, and layer normalization is applied to obtain the first modal correction domain. This operation helps to correct intermodal features at the feature level, improving the quality of the search domain features. Then, the first modal correction domain is processed by a multilayer perceptron (MLP) and superimposed on itself. After superposition, layer normalization is applied to obtain the first supplementary domain, further enriching the feature representation of the search domain (such as the first search domain). By injecting difference information, the model's ability to utilize multimodal features is strengthened.

[0102] Figure 6 This is a schematic diagram of a difference information injection module according to an embodiment of the present invention, such as... Figure 6 As shown, the visible light search area information (such as the first search domain) is calculated first. In conjunction with infrared search area information (such as the second search domain) The similarity between them is then multiplied by information from the fusion search region (such as the feature fusion domain). This allows us to obtain cross-modal similarity information (such as the first-modal similarity domain). .

[0103] Optionally, using the cross-attention mechanism, the first search domain is used as the query vector, the second search domain as the key vector, and the feature fusion domain as the value vector to obtain the first modality similarity domain, which is expressed by the following formula: ,in, , , These represent the parameters for the query, key, and value projection layers, respectively (which can be determined based on the corresponding template fields). It is the number of feature channels. Used to scale attention scores to prevent them from becoming too large.

[0104] Optionally, the first modality difference domain is determined based on the difference between the feature fusion domain and the first modality similarity domain, and its formula is expressed as: ,in, It is the first modal difference domain.

[0105] Optionally, the first search domain and the first modal difference domain are superimposed and then subjected to layer normalization to obtain the first modal correction domain, the formula of which is expressed as: ,in, It is the first mode correction domain. Representative level normalization.

[0106] Optionally, the first modality correction domain is processed by a multilayer perceptron and then superimposed with the first modality correction domain. After superposition, layer normalization is performed to obtain the first supplementary domain, the formula of which is expressed as: ,in, It is the first supplementary domain. and These represent layer normalization and multilayer perceptron, respectively.

[0107] In the above embodiments of this application, visible light search region information (such as a first supplementary domain) injected with differential infrared information can be obtained. Similarly, in the opposite direction, visible light differential information is injected into the infrared search region to obtain a second supplementary domain.

[0108] It should be noted that the above process can also be applied to the second search domain and the feature fusion domain to achieve bidirectional information supplementation between modalities, further improve tracking performance, and make more effective use of the difference information between modalities. Especially in environments where single-modal tracking performance is poor, it can significantly improve the accuracy and stability of tracking.

[0109] As an optional embodiment, the modality supplementation domain is obtained by processing the first search domain, the second search domain, and the feature fusion domain using a cross-attention mechanism, which includes at least the following steps: using the cross-attention mechanism, the second search domain is used as the query vector, the first search domain is used as the key vector in the cross-attention mechanism, and the feature fusion domain is used as the value vector in the cross-attention mechanism to obtain a second modality similarity domain; based on the difference between the feature fusion domain and the second modality similarity domain, a second modality difference domain is determined; the second search domain and the second modality difference domain are superimposed and then subjected to layer normalization processing to obtain a second modality correction domain; the second modality correction domain is processed by a multilayer perceptron and then superimposed with the second modality correction domain, and then subjected to layer normalization processing after superposition to obtain a second supplementation domain.

[0110] Optionally, adaptive feature fusion can be performed first, followed by differential information injection. Adaptive fusion can be used to remove cluttered information and obtain high-quality feature information, which is convenient for subsequent functions such as extracting differential feature information. Alternatively, differential information can be extracted first, followed by feature fusion. Both methods improve the tracking performance compared to the baseline model. However, performing adaptive feature fusion first to extract high-quality information is more helpful for subsequent feature operations and improves the model's tracking ability more.

[0111] As an optional embodiment, the first supplementary domain, the second supplementary domain, and the template fusion domain are processed using a cross-attention mechanism to obtain a feature enhancement domain, which includes at least the following steps: using the cross-attention mechanism, the template fusion domain is used as a query vector, and the second supplementary domain is used as a key vector and a value vector to obtain second search context information; the second search context information is enriched into the template fusion domain to obtain a second template enhancement domain; and using the cross-attention mechanism, the second template enhancement domain is used as a key vector and a value vector, and the first supplementary domain is used as a query vector to obtain a first enhancement domain.

[0112] The embodiments described above utilize the cross-attention mechanism in the improved TBSI module to achieve bidirectional cross-modal interaction between the first supplementary domain (corresponding to the first modal domain) and the second supplementary domain (corresponding to the second modal domain) through the template fusion domain. This achieves the combined advantages of adaptive feature fusion and differential information injection, resulting in a high-quality first enhancement domain. This series of operations not only strengthens the representational ability of the template but also effectively complements the differential information between modalities into the search domains of their respective modalities, enhancing the tracking robustness and accuracy of the system in complex environments. Compared to relying solely on a single modality or static fusion strategy, dynamic interaction and intelligent reinforcement significantly improve the performance of the multimodal target tracking system, especially under extreme conditions such as insufficient illumination and heat source interference, maintaining high-precision target tracking.

[0113] Optionally, the cross-modal interaction module, which uses a template fusion domain as a medium for bidirectional cross-modal interaction, is a key technical aspect of a multimodal target tracking system. This process aims to enhance the robustness and accuracy of the system, especially in complex environments such as changes in illumination and occlusion, by fully utilizing information from RGB (visible light) and TIR (thermal infrared) modes to construct a more dynamic and complementary first enhancement domain.

[0114] Figure 7 This is a schematic diagram of a cross-modal interaction module according to an embodiment of the present invention, such as... Figure 7 As shown, the template fusion domain is fused using a cross-attention mechanism. As a query vector, the second supplementary field As key and value vectors, the second search context information is obtained; the second search context information is then enriched into the template fusion domain. The second template enhancement domain is obtained. Using a cross-attention mechanism, the second template is enhanced. As the key vector and value vector, the first supplementary field As the query vector, the first enhancement domain is obtained. .

[0115] It should be noted that the template fusion domain refers to the domain obtained by fusing features from the RGB template domain (i.e., the first template domain) and the TIR template domain (i.e., the second template domain), and can simultaneously carry target information from both modalities. This template fusion domain acts as a bridge, facilitating information exchange between search domains of different modalities.

[0116] It should be noted that cross-attention is a mechanism that allows models to compare and extract relevant information between different sequences. In the scenario of this application, cross-attention is used to achieve information exchange between RGB search region features (such as the first search domain, or the first complementary domain corresponding to the first search domain) and TIR search region features (such as the second search domain, or the second complementary domain corresponding to the second search domain). Specifically, it manifests as two-way operations:

[0117] TIR→Media→RGB: In this approach, the template fusion domain is first used as the query, and the second enhancement domain as the key and value. A cross-attention mechanism is used to extract the target context information (i.e., the second search context information) under the TIR modality. This information is then enriched into the template fusion domain (i.e., the second template enhancement domain). Next, the updated template fusion domain (i.e., the second template enhancement domain) is used as the key and value, and the first enhancement domain is used as the query. The information from the TIR modality is then passed to the RGB search region (i.e., the first enhancement domain) to enhance the RGB features (i.e., the first enhancement domain).

[0118] RGB → Medium → TIR: Conversely, information can also be transferred from the RGB search region (i.e., the first supplementary domain) to the TIR search region (i.e., the second supplementary domain). By using the template fusion domain as an intermediary, the target context information under the RGB modality (i.e., the first search context information) can be integrated into the TIR search region (i.e., the second supplementary domain), thereby improving the integrity of the TIR features (i.e., the second supplementary domain).

[0119] It should be noted that the above process can also be applied to determine the first enhancement domain to achieve bidirectional information supplementation between modes, further improving tracking performance. Especially in environments where single-mode tracking is ineffective, it can significantly improve the accuracy and stability of tracking.

[0120] As an optional embodiment, the first supplementary domain, the second supplementary domain, and the template fusion domain are processed using a cross-attention mechanism to obtain a feature enhancement domain, which includes at least the following steps: using the cross-attention mechanism, the template fusion domain is used as a query vector, and the first supplementary domain is used as a key vector and a value vector to obtain first search context information based on the first supplementary domain; the first search context information is enriched into the template fusion domain to obtain a first template enhancement domain; and using the cross-attention mechanism, the first template enhancement domain is used as a key vector and a value vector, and the second supplementary domain is used as a query vector to obtain a second enhancement domain.

[0121] It should be noted that the cross-modal interaction module can also update the first template domain and the second template domain.

[0122] As an optional embodiment, the feature enhancement domain further includes: a template update domain configured for the next image frame group of the image frame group to be tracked, wherein the template update domain includes: an updated first template domain and an updated second template domain; the feature enhancement domain is obtained by processing the first supplementary domain, the second supplementary domain, and the template fusion domain using a cross-attention mechanism, which includes at least: using the cross-attention mechanism, using the first template domain as a query vector and the template fusion domain as a key vector and a value vector to obtain first template context information, and enriching the first template context information into the first template domain to obtain an updated first template domain; using the cross-attention mechanism, using the second template domain as a query vector and the template fusion domain as a key vector and a value vector to obtain second template context information, and enriching the second template context information into the second template domain to obtain an updated second template domain.

[0123] The embodiments described above utilize a cross-attention mechanism to process the first supplementary domain, the second supplementary domain, and the template fusion domain, thereby achieving cross-modal information transfer and iterative updates of the template domain (such as the first template domain and the second template domain). This enables the dynamic capture and fusion of template context information (such as the first template context information and the second template context information) from different modalities, enhancing the representational capability of the template domain and improving the robustness and accuracy of target tracking. Even in complex and variable environments, such as insufficient or excessive lighting or occlusion, this scheme can effectively cope with the failure of single-modal information and maintain the stability of tracking performance by updating the template domain in real time.

[0124] Optionally, when updating the template domains (such as the first and second template domains), the cross-modal interaction module uses the feature enhancement domain as the key and value, and the first or second supplementary domain as the query, and again employs the cross-attention mechanism to feed back the feature information (such as the first and second search domains) from the latest search frame to the original template, generating updated RGB template domains (such as the first template domain) and TIR template domains (such as the second template domain). The purpose of this is to ensure that the first and second template domains always maintain the latest visual and temperature features of the target, so that they can still be effectively used for target object localization in the search of the next frame.

[0125] As an optional embodiment, the first enhancement domain and the second search domain are concatenated and then input into the tracking head to predict the target bounding box, which is a key step in multimodal target tracking tasks for finally determining the target location.

[0126] It should be noted that the tracking head is a component in a multimodal target tracking system used to process fused features and predict target bounding boxes. When the first and second augmentation domains are input into the tracking head, the algorithm within the tracking head calculates the precise position and size of the target in the current frame based on these fused features.

[0127] Within the tracking head, the fused features are converted into output that can be directly used for bounding box prediction. A bounding box is a rectangular region represented by coordinates, used to define the position and shape of a target in an image. The prediction process may involve regression analysis, classification decisions, or other machine learning algorithms, and the final output bounding box information will reflect the estimated position of the tracked target in the current frame.

[0128] The embodiments described above in this application concatenate RGB and TIR search region features and input them into the tracking head, realizing the transformation from raw image data to accurate target position estimation. This is an important step in the entire tracking system to transform multimodal information into actual target detection and tracking output.

[0129] As an optional implementation, for nighttime security monitoring scenarios, the triggering conditions are low nighttime illumination (severe noise in RGB images) and partial occlusion (target is obscured). Temperature information needs to be supplemented via TIR mode to avoid single-mode failure. The execution entities include: a fixed security camera (for data acquisition) and an embedded computing unit (NVIDIA Jetson AGX) (for computation). The specific processing procedure is as follows:

[0130] After each frame of image acquisition is completed, temporal synchronization is performed, i.e., the dual-modal frames are aligned according to the timestamp. After the target is selected in the first frame bounding box, it is sent to the software model, segmented into fixed-size image blocks, and converted into token features. Then, shared learnable positional encodings are added. Next, it enters the ViT architecture, where improved TBSI modules are inserted at layers 4, 7, and 10. In the TBSI module, the input infrared search region information and visible light search region information are first adaptively fused using spatial channel attention to obtain fused information. Then, based on the fused search region information, infrared search region information, and visible light search region information, the difference information is injected into different branches. Finally, cross-modal interaction is achieved using a template as a medium to update the template region and search region. Finally, the output RGB and TIR search region features are concatenated and input into the tracking head to achieve target bounding box prediction.

[0131] Optionally, streetlights of appropriate brightness can be added to the camera's shooting scene, and RGB cameras can be used for monitoring and target tracking. The streetlights can be used to ensure stable single-modal recognition and tracking.

[0132] The embodiments described above effectively alleviate the problems of inaccurate tracking and tracking drift in nighttime security monitoring scenarios under single-mode failure conditions. They also enable more accurate target identification in low-light occlusion scenarios and possess stronger anti-interference capabilities. Furthermore, the overall processing time of this method is short, meeting real-time requirements.

[0133] According to an embodiment of the present invention, a device embodiment for multimodal target tracking is also provided. It should be noted that the device for multimodal target tracking can be used to execute the multimodal target tracking method in the embodiments of the present invention, and the multimodal target tracking method in the embodiments of the present invention can be executed in the device for multimodal target tracking.

[0134] Figure 8 This is a schematic diagram of a multimodal target tracking device according to an embodiment of the present invention, as shown below. Figure 8As shown, the device may include: an acquisition module 62, configured to acquire the feature modal domain of at least one group of image frames in a video stream to be tracked, wherein each group of image frames includes: visible light image frames and infrared light image frames acquired at the same time for the same target object, the feature modal domain includes: a first modal domain corresponding to the visible light image frame and a second modal domain corresponding to the infrared light image frame, the first modal domain includes: a first search domain and a first template domain, the second modal domain includes: a second search domain and a second template domain, the first template domain and the second template domain include features for tracking the target object; and a feature fusion module 64, configured to perform feature fusion on the first modal domain and the second modal domain to obtain a modal fusion domain, wherein the modal fusion domain includes The system comprises: a feature fusion domain obtained based on a first search domain and a second search domain, and a template fusion domain obtained based on a first template domain and a second template domain; a feature enhancement module 66, used to determine a feature enhancement domain based on the first search domain, the second search domain, the feature fusion domain, and the template fusion domain, wherein the feature fusion domain is used to perform difference injection by combining the differences between the first search domain and the second search domain, and the template fusion domain is used to perform cross-modal interaction between the first search domain and the second search domain; the feature enhancement domain includes at least: a first enhancement domain corresponding to the first search domain and a second enhancement domain corresponding to the second search domain; and a tracking module 68, used to track target objects in visible light image frames based on the first enhancement domain and in infrared light image frames based on the second enhancement domain.

[0135] It should be noted that the acquisition module 62 in this embodiment can be used to execute step S102 in this application embodiment, the feature fusion module 64 in this embodiment can be used to execute step S104 in this application embodiment, the feature enhancement module 66 in this embodiment can be used to execute step S106 in this application embodiment, and the tracking module 68 in this embodiment can be used to execute step S108 in this application embodiment. The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments.

[0136] In the embodiments described above, by acquiring visible light image frames and infrared light image frames of the target object simultaneously, a feature modal domain for target tracking is obtained. This feature modal domain includes a first modal domain and a second modal domain. A feature fusion domain is obtained by fusing features from a first search domain in the first modal domain and a second search domain in the second modal domain. Similarly, a template fusion domain is obtained by fusing features from a first template domain in the first modal domain and a second template domain in the second modal domain. This process leverages the complementarity of visible and infrared light to construct a high-quality joint representation. Then, the feature fusion domain is used to identify and perform differential analysis on the modal differences between visible and infrared light. By injecting information and utilizing the template fusion domain to perform cross-modal interaction between the first and second search domains, more robust first and second enhancement domains are generated. Based on the first and second enhancement domains, the target object is tracked in visible light and infrared light image frames according to the feature-enhanced first and second enhancement domains, respectively. This can effectively improve the tracking accuracy in complex scenes such as low light or occlusion, avoid tracking drift and accuracy degradation caused by single-modal failure, and achieve more stable and accurate target tracking. This achieves the technical effect of high tracking accuracy in multimodal target tracking, and solves the technical problem of low tracking accuracy in existing target tracking methods.

[0137] As an optional embodiment, the feature fusion module includes: a feature stitching unit, used to concatenate the first modal domain and the second modal domain to obtain a stitched feature domain, wherein the first modal domain includes: a set of visible light features corresponding to multiple spatial locations in a visible light image frame, each set of visible light features including: visible light features extracted from the same spatial location using multiple first feature channels; the second modal domain includes: a set of infrared light features corresponding to the same multiple spatial locations in an infrared light image frame, each set of infrared light features including: infrared light features extracted from the same spatial location using multiple second feature channels, wherein the number of channels in the first feature channels and the number of channels in the second feature channels are the same. The splicing feature domain includes: multiple splicing feature sets corresponding to different spatial locations, each splicing feature set including: splicing features based on multiple fourth feature channels at the same spatial location, where the fourth feature channel is either the first feature channel or the second feature channel, and the splicing features are visible light features or infrared light features; a feature transformation unit is used to perform feature transformation on the splicing feature domain in the channel dimension using the first linear layer to obtain a modal fusion domain, wherein the modal fusion domain includes: multiple fusion feature sets corresponding to different spatial locations, each fusion feature set including: fusion features based on multiple third feature channels at the same spatial location, where the number of third feature channels is the same as the number of first feature channels.

[0138] As an optional embodiment, the device further includes: a first processing unit, configured to, after concatenating channels in the first modal domain and the second modal domain to obtain a spliced ​​feature domain, process the spliced ​​feature domain using a channel attention mechanism to obtain a channel weight matrix, wherein the channel weight matrix includes: channel weights corresponding to multiple fourth feature channels, each channel weight representing the importance of each fourth feature channel; a first determining unit, configured to determine a channel feature domain based on the channel weight matrix and the spliced ​​feature domain, wherein the channel feature domain includes: multiple sets of channel features corresponding to multiple spatial locations, each set of channel features including: channel features corresponding to multiple fourth feature channels at the same spatial location, each channel feature being the product of the channel weight corresponding to the same fourth feature channel and the spliced ​​feature; and a second processing unit, configured to process the channel feature domain using a spatial attention mechanism. The process involves several steps: First, a spatial weight matrix is ​​obtained, comprising position weights corresponding to multiple spatial locations, each representing the importance of a given spatial location. Second, a determining unit is used to determine a position feature domain based on the spatial weight matrix and the channel feature domain. This position feature domain comprises sets of position features corresponding to multiple spatial locations, each set including position features corresponding to multiple fourth feature channels at the same spatial location, where each feature is the product of the position weight and the channel feature corresponding to the same spatial location. Third, a feature processing unit is used to perform weighted summation of the concatenated feature domain and the position feature domain using a second linear layer to obtain a modal fusion domain. The fusion feature in the modal fusion domain is the result of weighted merging of the concatenated features in the concatenated feature domain and the position features in the position feature domain, performed according to the channel dimension.

[0139] As an optional embodiment, the feature enhancement module includes: a difference injection unit, used to process a first search domain, a second search domain, and a feature fusion domain using a cross-attention mechanism to obtain a modality complementation domain, wherein the cross-attention mechanism is used to generate a cross-modal similarity domain based on the first search domain, the second search domain, and the feature fusion domain, the cross-modal similarity domain including: a first modal similarity domain representing the difference between the first search domain and the second search domain, and a second modal similarity domain representing the difference between the second search domain and the first search domain, the modality complementation domain including: a first complementation domain obtained by difference injection of the first search domain based on the first modal similarity domain, and a second complementation domain obtained by difference injection of the second search domain based on the second modal similarity domain; and a cross-modal interaction unit, used to process the first complementation domain, the second complementation domain, and the template fusion domain using a cross-attention mechanism to obtain a feature enhancement domain, wherein the cross-attention mechanism is used to perform bidirectional cross-modal interaction between the first complementation domain and the second complementation domain using the template fusion domain as a medium, the feature enhancement domain including at least: a first enhancement domain corresponding to the first complementation domain, and a second enhancement domain corresponding to the first complementation domain.

[0140] As an optional embodiment, the difference injection unit includes at least: a first similarity domain determination subunit, used to obtain a first modality similarity domain by using a cross-attention mechanism, taking a first search domain as a query vector, a second search domain as a key vector, and a feature fusion domain as a value vector; a first determination subunit, used to determine a first modality difference domain based on the difference between the feature fusion domain and the first modality similarity domain; a first layer normalization subunit, used to perform layer normalization processing after superimposing the first search domain and the first modality difference domain, to obtain a first modality correction domain; and a second layer normalization subunit, used to perform layer normalization processing after superimposing the first modality correction domain using a multilayer perceptron, to obtain a first supplementary domain.

[0141] As an optional embodiment, the difference injection unit includes at least: a second similarity domain determination subunit, used to utilize a cross-attention mechanism, taking a second search domain as a query vector, a first search domain as a key vector in the cross-attention mechanism, and a feature fusion domain as a value vector in the cross-attention mechanism, to obtain a second modality similarity domain; a second determination subunit, used to determine a second modality difference domain based on the difference between the feature fusion domain and the second modality similarity domain; a third layer normalization subunit, used to superimpose the second search domain and the second modality difference domain and then perform layer normalization processing to obtain a second modality correction domain; and a fourth layer normalization subunit, used to process the second modality correction domain using a multilayer perceptron and then superimpose it with the second modality correction domain, and then perform layer normalization processing after superposition to obtain a second supplementary domain.

[0142] As an optional embodiment, the cross-modal interaction unit includes at least: a first information determination subunit, used to use a cross-attention mechanism to take the template fusion domain as a query vector and the first supplementary domain as a key vector and a value vector to obtain first search context information based on the first supplementary domain; a first enrichment subunit, used to enrich the first search context information into the template fusion domain to obtain a first template enhancement domain; and a first enhancement subunit, used to use a cross-attention mechanism to take the first template enhancement domain as a key vector and a value vector and the second supplementary domain as a query vector to obtain a second enhancement domain.

[0143] As an optional embodiment, the cross-modal interaction unit includes at least: a second information determination subunit, used to use a cross-attention mechanism to take the template fusion domain as a query vector and the second supplementary domain as a key vector and a value vector to obtain second search context information; a second enrichment subunit, used to enrich the second search context information into the template fusion domain to obtain a second template enhancement domain; and a second enhancement subunit, used to use a cross-attention mechanism to take the second template enhancement domain as a key vector and a value vector and the first supplementary domain as a query vector to obtain a first enhancement domain.

[0144] As an optional embodiment, the feature enhancement domain further includes: a template update domain configured for the next image frame group of the image frame group to be tracked, wherein the template update domain includes: an updated first template domain and an updated second template domain; the cross-modal interaction unit includes at least: a first template update subunit, used to use a cross-attention mechanism to take the first template domain as a query vector and the template fusion domain as a key vector and a value vector to obtain first template context information, and enrich the first template context information into the first template domain to obtain an updated first template domain; and a second template update subunit, used to use a cross-attention mechanism to take the second template domain as a query vector and the template fusion domain as a key vector and a value vector to obtain second template context information, and enrich the second template context information into the second template domain to obtain an updated second template domain.

[0145] Embodiments of the present invention can provide an electronic device, which can be a computer terminal, and the computer terminal can be any one of a group of computer terminal devices. Optionally, in this embodiment, the computer terminal can also be replaced by a mobile terminal or other terminal device.

[0146] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0147] In this embodiment, the computer terminal described above can execute the program code for the following steps in the multimodal target tracking method: In at least one group of image frames of the video stream to be tracked, the feature modal domain of the group of image frames to be tracked is obtained, wherein each group of image frames includes: visible light image frames and infrared light image frames acquired at the same time for the same target object; the feature modal domain includes: a first modal domain corresponding to the visible light image frame and a second modal domain corresponding to the infrared light image frame; the first modal domain includes: a first search domain and a first template domain; the second modal domain includes: a second search domain and a second template domain; the first template domain and the second template domain include features for tracking the target object; feature fusion is performed on the first modal domain and the second modal domain to obtain the modality. The fusion domain includes: a feature fusion domain obtained based on a first search domain and a second search domain, and a template fusion domain obtained based on a first template domain and a second template domain; a feature enhancement domain is determined based on the first search domain, the second search domain, the feature fusion domain, and the template fusion domain, wherein the feature fusion domain is used to perform difference injection by combining the differences between the first search domain and the second search domain, the template fusion domain is used to perform cross-modal interaction between the first search domain and the second search domain, and the feature enhancement domain includes at least: a first enhancement domain corresponding to the first search domain and a second enhancement domain corresponding to the second search domain; the first enhancement domain is used to track target objects in visible light image frames, and the second enhancement domain is used to track target objects in infrared light image frames.

[0148] Figure 9 This is a structural block diagram of a computer terminal according to an embodiment of the present invention, such as... Figure 9 As shown, the computer terminal 70 may include one or more (only one is shown in the figure) processors 72 and memory 74.

[0149] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the multimodal target tracking method and apparatus in this embodiment of the invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned multimodal target tracking method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal 70 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0150] The processor can invoke information and application programs stored in the memory via a transmission device to perform the following steps: In at least one group of image frames of the video stream to be tracked, acquire the feature modal domain of the group of image frames to be tracked, wherein each group of image frames includes: a visible light image frame and an infrared light image frame acquired at the same time for the same target object; the feature modal domain includes: a first modal domain corresponding to the visible light image frame and a second modal domain corresponding to the infrared light image frame; the first modal domain includes: a first search domain and a first template domain; the second modal domain includes: a second search domain and a second template domain; the first template domain and the second template domain include features for tracking the target object; perform feature fusion on the first modal domain and the second modal domain to obtain a modal fusion. The modality fusion domain includes: a feature fusion domain obtained based on a first search domain and a second search domain, and a template fusion domain obtained based on a first template domain and a second template domain; a feature enhancement domain is determined based on the first search domain, the second search domain, the feature fusion domain, and the template fusion domain, wherein the feature fusion domain is used to perform difference injection by combining the differences between the first search domain and the second search domain, the template fusion domain is used to perform cross-modal interaction between the first search domain and the second search domain, and the feature enhancement domain includes at least: a first enhancement domain corresponding to the first search domain, and a second enhancement domain corresponding to the second search domain; the first enhancement domain is used to track target objects in visible light image frames, and the second enhancement domain is used to track target objects in infrared light image frames.

[0151] Optionally, the processor may also execute program code for the following steps: concatenating channels of the first modal domain and the second modal domain to obtain a spliced ​​feature domain, wherein the first modal domain includes: visible light feature sets corresponding to multiple spatial locations in a visible light image frame, each visible light feature set including: visible light features extracted from the same spatial location using multiple first feature channels; the second modal domain includes: infrared light feature sets corresponding to the same multiple spatial locations in an infrared light image frame, each infrared light feature set including: infrared light features extracted from the same spatial location using multiple second feature channels; the number of channels of the first feature channels and the number of channels of the second feature channels are... Similarly, the splicing feature domain includes: multiple splicing feature sets corresponding to multiple spatial locations, each splicing feature set including: splicing features based on multiple fourth feature channels at the same spatial location, where the fourth feature channel is either the first feature channel or the second feature channel, and the splicing features are visible light features or infrared light features; the splicing feature domain is transformed in the channel dimension using the first linear layer to obtain the modal fusion domain, wherein the modal fusion domain includes: multiple fusion feature sets corresponding to multiple spatial locations, each fusion feature set including: fusion features based on multiple third feature channels at the same spatial location, where the number of third feature channels is the same as that of the first feature channels.

[0152] Optionally, the processor may also execute program code for the following steps: processing the concatenated feature domain using a channel attention mechanism to obtain a channel weight matrix, wherein the channel weight matrix includes: channel weights corresponding to multiple fourth feature channels, each channel weight representing the importance of each fourth feature channel; determining the channel feature domain based on the channel weight matrix and the concatenated feature domain, wherein the channel feature domain includes: multiple sets of channel features corresponding to multiple spatial locations, each set of channel features including: channel features corresponding to multiple fourth feature channels at the same spatial location, each channel feature being the product of the channel weight corresponding to the same fourth feature channel and the concatenated feature; processing the channel feature domain using a spatial attention mechanism to obtain a spatial weight matrix, wherein... The spatial weight matrix includes: position weights corresponding to multiple spatial locations, each position weight representing the importance of each spatial location; based on the spatial weight matrix and the channel feature domain, the position feature domain is determined, wherein the position feature domain includes: a set of position features corresponding to multiple spatial locations, each set of position features including: position features corresponding to multiple fourth feature channels at the same spatial location, each position feature being the product of the position weight and the channel feature corresponding to the same spatial location; the second linear layer is used to perform feature weighted summation on the concatenated feature domain and the position feature domain to obtain the modal fusion domain, wherein the fusion feature in the modal fusion domain is: the result of weighted merging of the concatenated features in the concatenated feature domain and the position features in the position feature domain by channel dimension.

[0153] Optionally, the processor may also execute program code for the following steps: processing the first search domain, the second search domain, and the feature fusion domain using a cross-attention mechanism to obtain a modality complementation domain, wherein the cross-attention mechanism is used to generate a cross-modal similarity domain based on the first search domain, the second search domain, and the feature fusion domain, the cross-modal similarity domain including: a first modal similarity domain representing the difference between the first search domain and the second search domain, and a second modal similarity domain representing the difference between the second search domain and the first search domain; the modality complementation domain including: a first complementation domain obtained by injecting a difference into the first search domain based on the first modal similarity domain, and a second complementation domain obtained by injecting a difference into the second search domain based on the second modal similarity domain; processing the first complementation domain, the second complementation domain, and the template fusion domain using a cross-attention mechanism to obtain a feature enhancement domain, wherein the cross-attention mechanism is used to perform bidirectional cross-modal interaction between the first complementation domain and the second complementation domain using the template fusion domain as a medium; the feature enhancement domain includes at least: a first enhancement domain corresponding to the first complementation domain, and a second enhancement domain corresponding to the first complementation domain.

[0154] Optionally, the processor may also execute program code with the following steps: using a cross-attention mechanism, the first search domain is used as the query vector, the second search domain as the key vector, and the feature fusion domain as the value vector to obtain a first modality similarity domain; based on the difference between the feature fusion domain and the first modality similarity domain, a first modality difference domain is determined; the first search domain and the first modality difference domain are superimposed and then subjected to layer normalization to obtain a first modality correction domain; the first modality correction domain is processed by a multilayer perceptron and then superimposed with the first modality correction domain, and then subjected to layer normalization after superposition to obtain a first supplementary domain.

[0155] Optionally, the processor may also execute program code with the following steps: using a cross-attention mechanism, the second search domain is used as the query vector, the first search domain is the key vector in the cross-attention mechanism, and the feature fusion domain is the value vector in the cross-attention mechanism to obtain a second modality similarity domain; based on the difference between the feature fusion domain and the second modality similarity domain, a second modality difference domain is determined; the second search domain and the second modality difference domain are superimposed and then subjected to layer normalization processing to obtain a second modality correction domain; the second modality correction domain is processed by a multilayer perceptron and then superimposed with the second modality correction domain, and then subjected to layer normalization processing after superposition to obtain a second supplementary domain.

[0156] Optionally, the processor may also execute program code with the following steps: using a cross-attention mechanism, using the template fusion domain as the query vector and the first supplementary domain as the key vector and value vector, to obtain first search context information based on the first supplementary domain; enriching the first search context information into the template fusion domain to obtain a first template enhancement domain; using a cross-attention mechanism, using the first template enhancement domain as the key vector and value vector and the second supplementary domain as the query vector, to obtain a second enhancement domain.

[0157] Optionally, the processor may also execute program code with the following steps: using a cross-attention mechanism, using the template fusion domain as the query vector and the second supplementary domain as the key vector and value vector to obtain second search context information; enriching the second search context information into the template fusion domain to obtain a second template enhancement domain; using a cross-attention mechanism, using the second template enhancement domain as the key vector and value vector and the first supplementary domain as the query vector to obtain a first enhancement domain.

[0158] Optionally, the feature enhancement domain further includes: a template update domain configured for the next image frame group of the image frame group to be tracked, wherein the template update domain includes: an updated first template domain and an updated second template domain. The processor may also execute program code that performs the following steps: using a cross-attention mechanism, the first template domain is used as a query vector, and the template fusion domain is used as a key vector and a value vector to obtain first template context information, and the first template context information is enriched into the first template domain to obtain an updated first template domain; using a cross-attention mechanism, the second template domain is used as a query vector, and the template fusion domain is used as a key vector and a value vector to obtain second template context information, and the second template context information is enriched into the second template domain to obtain an updated second template domain.

[0159] Those skilled in the art will understand that Figure 9 The structure shown is for illustrative purposes only. The computer terminal can also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a mobile internet device (MID), a PAD, and other terminal devices. Figure 9 This does not limit the structure of the aforementioned electronic device. For example, the computer terminal 70 may also include components that are more... Figure 9 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 9 The different configurations shown.

[0160] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a computer program instructing the hardware related to the terminal device. The computer program can be stored in a non-volatile medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0161] Embodiments of the present invention also provide a non-volatile storage medium. Optionally, in this embodiment, the aforementioned non-volatile storage medium can be used to store the program code executed by the multimodal target tracking method provided in the above embodiments.

[0162] Optionally, in this embodiment, the non-volatile storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0163] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: In at least one group of image frames of the video stream to be tracked, the feature modal domain of the group of image frames to be tracked is obtained, wherein each group of image frames includes: visible light image frames and infrared light image frames acquired at the same time for the same target object; the feature modal domain includes: a first modal domain corresponding to the visible light image frame and a second modal domain corresponding to the infrared light image frame; the first modal domain includes: a first search domain and a first template domain; the second modal domain includes: a second search domain and a second template domain; the first template domain and the second template domain include features for tracking the target object; feature fusion is performed on the first modal domain and the second modal domain to obtain a modality. The fusion domain includes: a feature fusion domain obtained based on a first search domain and a second search domain, and a template fusion domain obtained based on a first template domain and a second template domain; a feature enhancement domain is determined based on the first search domain, the second search domain, the feature fusion domain, and the template fusion domain, wherein the feature fusion domain is used to perform difference injection by combining the differences between the first search domain and the second search domain, the template fusion domain is used to perform cross-modal interaction between the first search domain and the second search domain, and the feature enhancement domain includes at least: a first enhancement domain corresponding to the first search domain and a second enhancement domain corresponding to the second search domain; the first enhancement domain is used to track target objects in visible light image frames, and the second enhancement domain is used to track target objects in infrared light image frames.

[0164] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: concatenating channels of the first modal domain and the second modal domain to obtain a spliced ​​feature domain, wherein the first modal domain includes: a set of visible light features corresponding to multiple spatial locations in a visible light image frame, each set of visible light features including: visible light features extracted from the same spatial location using multiple first feature channels; the second modal domain includes: a set of infrared light features corresponding to the same multiple spatial locations in an infrared light image frame, each set of infrared light features including: infrared light features extracted from the same spatial location using multiple second feature channels; the first feature channels and the second feature channels... The number of channels is the same. The splicing feature domain includes: multiple splicing feature sets corresponding to multiple spatial locations. Each splicing feature set includes: splicing features based on multiple fourth feature channels at the same spatial location. The fourth feature channel is either the first feature channel or the second feature channel. The splicing features are visible light features or infrared light features. The splicing feature domain is transformed in the channel dimension using the first linear layer to obtain the modal fusion domain. The modal fusion domain includes: multiple fusion feature sets corresponding to multiple spatial locations. Each fusion feature set includes: fusion features based on multiple third feature channels at the same spatial location. The number of third feature channels is the same as that of the first feature channels.

[0165] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: processing the concatenated feature domain using a channel attention mechanism to obtain a channel weight matrix, wherein the channel weight matrix includes: channel weights corresponding to multiple fourth feature channels, each channel weight representing the importance of each fourth feature channel; determining the channel feature domain based on the channel weight matrix and the concatenated feature domain, wherein the channel feature domain includes: multiple sets of channel features corresponding to multiple spatial locations, each set of channel features including: channel features corresponding to multiple fourth feature channels at the same spatial location, each channel feature being the product of the channel weight corresponding to the same fourth feature channel and the concatenated feature; processing the channel feature domain using a spatial attention mechanism to obtain a spatial feature domain. The weight matrix includes a spatial weight matrix, which contains position weights corresponding to multiple spatial locations, each position weight representing the importance of each spatial location. Based on the spatial weight matrix and the channel feature domain, a position feature domain is determined, which includes a set of position features corresponding to multiple spatial locations. Each set of position features includes position features corresponding to multiple fourth feature channels at the same spatial location, where each position feature is the product of the position weight and the channel feature corresponding to the same spatial location. A second linear layer is used to perform a weighted summation of the concatenated feature domain and the position feature domain to obtain the modal fusion domain. The fusion feature in the modal fusion domain is the result of weighted merging of the concatenated features in the concatenated feature domain and the position features in the position feature domain, performed according to the channel dimension.

[0166] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: processing the first search domain, the second search domain, and the feature fusion domain using a cross-attention mechanism to obtain a modality complementation domain, wherein the cross-attention mechanism is used to generate a cross-modal similarity domain based on the first search domain, the second search domain, and the feature fusion domain, the cross-modal similarity domain including: a first modal similarity domain representing the difference between the first search domain and the second search domain, and a second modal similarity domain representing the difference between the second search domain and the first search domain, the modality complementation domain including: a first complementation domain obtained by injecting a difference into the first search domain based on the first modal similarity domain, and a second complementation domain obtained by injecting a difference into the second search domain based on the second modal similarity domain; processing the first complementation domain, the second complementation domain, and the template fusion domain using a cross-attention mechanism to obtain a feature enhancement domain, wherein the cross-attention mechanism is used to perform bidirectional cross-modal interaction between the first complementation domain and the second complementation domain using the template fusion domain as a medium, the feature enhancement domain including at least: a first enhancement domain corresponding to the first complementation domain, and a second enhancement domain corresponding to the first complementation domain.

[0167] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: using a cross-attention mechanism, a first search domain is used as a query vector, a second search domain as a key vector, and a feature fusion domain as a value vector to obtain a first modality similarity domain; based on the difference between the feature fusion domain and the first modality similarity domain, a first modality difference domain is determined; the first search domain and the first modality difference domain are superimposed and then subjected to layer normalization processing to obtain a first modality correction domain; the first modality correction domain is processed by a multilayer perceptron and then superimposed with the first modality correction domain, and then subjected to layer normalization processing after superposition to obtain a first supplementary domain.

[0168] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: using a cross-attention mechanism, a second search domain is used as a query vector, a first search domain is used as a key vector in the cross-attention mechanism, and a feature fusion domain is used as a value vector in the cross-attention mechanism to obtain a second modality similarity domain; based on the difference between the feature fusion domain and the second modality similarity domain, a second modality difference domain is determined; the second search domain and the second modality difference domain are superimposed and then subjected to layer normalization processing to obtain a second modality correction domain; the second modality correction domain is processed by a multilayer perceptron and then superimposed with the second modality correction domain, and then subjected to layer normalization processing after superposition to obtain a second supplementary domain.

[0169] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: using a cross-attention mechanism, the template fusion domain is used as a query vector, and the first supplementary domain is used as a key vector and a value vector to obtain first search context information based on the first supplementary domain; the first search context information is enriched into the template fusion domain to obtain a first template enhancement domain; using a cross-attention mechanism, the first template enhancement domain is used as a key vector and a value vector, and the second supplementary domain is used as a query vector to obtain a second enhancement domain.

[0170] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: using a cross-attention mechanism, the template fusion domain is used as a query vector, and the second supplementary domain is used as a key vector and a value vector to obtain second search context information; the second search context information is enriched into the template fusion domain to obtain a second template enhancement domain; using a cross-attention mechanism, the second template enhancement domain is used as a key vector and a value vector, and the first supplementary domain is used as a query vector to obtain a first enhancement domain.

[0171] Optionally, in this embodiment, the feature enhancement domain further includes: a template update domain configured for the next image frame group of the image frame group to be tracked, wherein the template update domain includes: an updated first template domain and an updated second template domain, and the non-volatile storage medium is configured to store program code for performing the following steps: using a cross-attention mechanism, using the first template domain as a query vector and the template fusion domain as a key vector and value vector to obtain first template context information, and enriching the first template context information into the first template domain to obtain an updated first template domain; using a cross-attention mechanism, using the second template domain as a query vector and the template fusion domain as a key vector and value vector to obtain second template context information, and enriching the second template context information into the second template domain to obtain an updated second template domain.

[0172] Embodiments of the present invention also provide a computer program product, including a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the steps of the multimodal target tracking method provided in the above embodiments.

[0173] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0174] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0175] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0176] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0177] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0178] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a non-volatile storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned non-volatile storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0179] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for multimodal target tracking, characterized in that, include: In at least one group of image frames in a video stream to be tracked, the feature modal domain of the group of image frames to be tracked is obtained. Each group of image frames includes: a visible light image frame and an infrared light image frame acquired at the same time for the same target object. The feature modal domain includes: a first modal domain corresponding to the visible light image frame and a second modal domain corresponding to the infrared light image frame. The first modal domain includes: a first search domain and a first template domain. The second modal domain includes: a second search domain and a second template domain. The first template domain and the second template domain include features for tracking the target object. The first modal domain and the second modal domain are fused to obtain a modal fusion domain, wherein the modal fusion domain includes: a feature fusion domain obtained based on the first search domain and the second search domain, and a template fusion domain obtained based on the first template domain and the second template domain; Based on the first search domain, the second search domain, the feature fusion domain, and the template fusion domain, a feature enhancement domain is determined. The feature fusion domain is used to perform difference injection by combining the differences between the first search domain and the second search domain. The template fusion domain is used to perform cross-modal interaction between the first search domain and the second search domain. The feature enhancement domain includes at least: a first enhancement domain corresponding to the first search domain and a second enhancement domain corresponding to the second search domain. The target object is tracked in the visible light image frame based on the first enhancement domain, and the target object is tracked in the infrared light image frame based on the second enhancement domain.

2. The method according to claim 1, characterized in that, Feature fusion is performed on the first modal domain and the second modal domain to obtain a modal fusion domain including: The first modal domain and the second modal domain are channel-concatenated to obtain a spliced ​​feature domain. The first modal domain includes: a set of visible light features corresponding to multiple spatial locations in the visible light image frame, each set including visible light features extracted from multiple first feature channels at the same spatial location. The second modal domain includes: a set of infrared light features corresponding to multiple identical spatial locations in the infrared light image frame, each set including infrared light features extracted from multiple second feature channels at the same spatial location. The first and second feature channels have the same number of channels. The spliced ​​feature domain includes: a set of spliced ​​features corresponding to multiple spatial locations, each set including spliced ​​features based on multiple fourth feature channels at the same spatial location. The fourth feature channels are either the first or second feature channels, and the spliced ​​features are either the visible light features or the infrared light features. The first linear layer is used to perform feature transformation on the channel dimension of the spliced ​​feature domain to obtain the modal fusion domain. The modal fusion domain includes: multiple fusion feature sets corresponding to the spatial locations respectively. Each fusion feature set includes: fusion features based on multiple third feature channels corresponding to the same spatial location. The number of third feature channels is the same as that of the first feature channels.

3. The method according to claim 2, characterized in that, After concatenating the channels of the first modal domain and the second modal domain to obtain the spliced ​​feature domain, the method further includes: The concatenated feature domain is processed using a channel attention mechanism to obtain a channel weight matrix, wherein the channel weight matrix includes: channel weights corresponding to multiple fourth feature channels, and each channel weight is used to represent the importance of each fourth feature channel; Based on the channel weight matrix and the splicing feature domain, a channel feature domain is determined, wherein the channel feature domain includes: multiple channel feature sets corresponding to the spatial locations respectively, each channel feature set includes: channel features corresponding to multiple fourth feature channels at the same spatial location, and each channel feature is the product of the channel weight and the splicing feature corresponding to the same fourth feature channel; The channel feature domain is processed using a spatial attention mechanism to obtain a spatial weight matrix, wherein the spatial weight matrix includes: position weights corresponding to multiple spatial positions, and each position weight is used to represent the importance of each spatial position; Based on the spatial weight matrix and the channel feature domain, a position feature domain is determined, wherein the position feature domain includes: a plurality of position feature sets corresponding to the spatial positions respectively, each position feature set including: position features corresponding to the same spatial position based on the plurality of fourth feature channels respectively, each position feature being the product of the position weight and the channel feature corresponding to the same spatial position; The second linear layer is used to perform a weighted summation of the splicing feature domain and the position feature domain to obtain the modality fusion domain. The fusion feature in the modality fusion domain is the result of weighted merging of the splicing feature in the splicing feature domain and the position feature in the position feature domain by channel dimension.

4. The method according to claim 1, characterized in that, Based on the first search domain, the second search domain, the feature fusion domain, and the template fusion domain, the feature enhancement domain is determined to include: A cross-attention mechanism is used to process the first search domain, the second search domain, and the feature fusion domain to obtain a modality complementation domain. The cross-attention mechanism is used to generate a cross-modal similarity domain based on the first search domain, the second search domain, and the feature fusion domain. The cross-modal similarity domain includes a first modal similarity domain representing the difference between the first search domain and the second search domain, and a second modal similarity domain representing the difference between the second search domain and the first search domain. The modality complementation domain includes a first complementation domain obtained by injecting a difference into the first search domain based on the first modal similarity domain, and a second complementation domain obtained by injecting a difference into the second search domain based on the second modal similarity domain. The first supplementary domain, the second supplementary domain, and the template fusion domain are processed using a cross-attention mechanism to obtain a feature enhancement domain. The cross-attention mechanism is used to perform bidirectional cross-modal interaction between the first supplementary domain and the second supplementary domain using the template fusion domain as a medium. The feature enhancement domain includes at least a first enhancement domain corresponding to the first supplementary domain and a second enhancement domain corresponding to the first supplementary domain.

5. The method according to claim 4, characterized in that, By processing the first search domain, the second search domain, and the feature fusion domain using a cross-attention mechanism, the modality complementation domain is obtained, which includes at least: Using the cross-attention mechanism, the first search domain is used as the query vector, the second search domain as the key vector, and the feature fusion domain as the value vector to obtain the first modality similarity domain; The first modal difference domain is determined based on the difference between the feature fusion domain and the first modal similarity domain; The first search domain and the first modal difference domain are superimposed and then subjected to layer normalization to obtain the first modal correction domain. The first modality correction domain is processed by a multilayer perceptron and then superimposed with the first modality correction domain. After superposition, layer normalization is performed to obtain the first supplementary domain.

6. The method according to claim 4, characterized in that, The first supplementary domain, the second supplementary domain, and the template fusion domain are processed using a cross-attention mechanism to obtain a feature enhancement domain that includes at least: Using the cross-attention mechanism, the template fusion domain is used as the query vector, and the first supplementary domain is used as the key vector and value vector to obtain the first search context information based on the first supplementary domain; The first search context information is enriched into the template fusion domain to obtain the first template enhancement domain; Using the cross-attention mechanism, the first template enhancement domain is used as the key vector and value vector, and the second supplementary domain is used as the query vector to obtain the second enhancement domain.

7. The method according to claim 4, characterized in that, The feature enhancement domain further includes: a template update domain configured for the next image frame group of the image frame group to be tracked, wherein the template update domain includes: an updated first template domain and an updated second template domain; the first supplementary domain, the second supplementary domain, and the template fusion domain are processed using a cross-attention mechanism to obtain a feature enhancement domain that includes at least: Using the cross-attention mechanism, the first template domain is used as the query vector, and the template fusion domain is used as the key vector and value vector to obtain the first template context information. The first template context information is then enriched into the first template domain to obtain the updated first template domain. Using the cross-attention mechanism, the second template domain is used as the query vector, and the template fusion domain is used as the key vector and value vector to obtain the second template context information. The second template context information is then enriched into the second template domain to obtain the updated second template domain.

8. A device for multimodal target tracking, characterized in that, include: An acquisition module is configured to acquire the feature modality domain of at least one image frame group in a video stream to be tracked, wherein each image frame group includes: a visible light image frame and an infrared light image frame acquired at the same time for the same target object, and the feature modality domain includes: a first modality domain corresponding to the visible light image frame and a second modality domain corresponding to the infrared light image frame, the first modality domain including: a first search domain and a first template domain, the second modality domain including: a second search domain and a second template domain, and the first template domain and the second template domain including features for tracking the target object; The feature fusion module is used to perform feature fusion on the first modal domain and the second modal domain to obtain a modal fusion domain, wherein the modal fusion domain includes: a feature fusion domain obtained based on the first search domain and the second search domain, and a template fusion domain obtained based on the first template domain and the second template domain; The feature enhancement module is used to determine a feature enhancement domain based on the first search domain, the second search domain, the feature fusion domain, and the template fusion domain. The feature fusion domain is used to perform difference injection by combining the differences between the first search domain and the second search domain. The template fusion domain is used to perform cross-modal interaction between the first search domain and the second search domain. The feature enhancement domain includes at least a first enhancement domain corresponding to the first search domain and a second enhancement domain corresponding to the second search domain. The tracking module is configured to track the target object in the visible light image frame based on the first enhancement domain, and to track the target object in the infrared light image frame based on the second enhancement domain.

9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the multimodal target tracking method of any one of claims 1 to 7 through the computer program.

10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the multimodal target tracking method according to any one of claims 1 to 7.