RGB-T-based target tracking method and system
By adopting an RGB-T-based method in the target tracking technology, combining self-attention operation and Mask masking strategy to fusion multimodal features, dynamically adjusting the contribution ratio between modes, the problem of low information utilization caused by difficulty in ensuring tracking accuracy and fixed weights in the existing technology is solved, and higher tracking accuracy and model robustness are achieved.
Patent Information
- Application Number
- CN202510173049.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
In scenarios where the background is complex or the target is moving rapidly, the tracking accuracy is difficult to ensure, and the modal fusion method with fixed weights leads to low utilization of important information, affecting the robustness of the model.
The RGB-T-based target tracking method is adopted, and the search area and template area are determined by obtaining RGB images and TIR images, and feature extraction is used for feature extraction, and multimodal feature fusion is performed by combining self-attention operations and Mask masking strategies to dynamically adjust the contribution ratio between modes.
It significantly improves the tracking accuracy and the robustness of the model, enhances the mining of modal-specific information, and avoids the problem of low utilization of important information under fixed weight mode.
Smart Images

Figure CN120107312A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an RGB-T based target tracking method and system. Background Art
[0002] With the development of artificial intelligence, target tracking technology has been widely used in fields such as unmanned driving, security monitoring, and robot vision navigation. Among them, visible light (RGB) and thermal infrared (Thermal) are two common information media, each showing unique advantages in different scenarios with its imaging characteristics. Visible light images are rich in details, but are limited by external lighting conditions; thermal infrared images show strong robustness under low light conditions, but are difficult to distinguish when the temperature difference between the target and the background is small. Therefore, how to effectively fuse the information of the two modalities to overcome the limitations of a single modality is a research hotspot in the current field of target tracking.
[0003] Existing target tracking technologies mainly include modal fusion methods based on single-stream networks. The basic idea of single-stream networks is to fuse the information of visible light (RGB) and infrared (Thermal) modalities at the network input stage or the initial feature extraction stage, and input the fused unified feature stream into the subsequent feature extraction module for processing. Another type of modal fusion is based on dual-stream networks. The dual-stream network retains the independence of the modalities and improves the utilization of modality-specific information by performing multi-level extraction and fusion of features.
[0004] However, the existing single-stream network has insufficient mining of modal-specific information due to early fusion, and the two-stream network is limited to the internal fusion of the search area or template, and the modal synergy is weak, especially in scenes with complex backgrounds or fast-moving targets, the tracking accuracy is difficult to guarantee; at the same time, most of the existing methods use a fixed-weight modal fusion method, which is unable to adaptively adjust the contribution ratio between modalities according to the characteristics of dynamic scenes, resulting in low utilization of important information, which in turn affects the tracking accuracy and the robustness of the model. Summary of the invention
[0005] In order to solve the technical problems that the existing single-stream network has insufficient mining of modal-specific information due to early fusion, the dual-stream network is limited to the internal fusion of the search area or template, and the modal synergy is weak, especially in scenes with complex backgrounds or fast-moving targets, the tracking accuracy is difficult to guarantee; at the same time, most of the existing methods adopt a modal fusion method with fixed weights, which cannot adaptively adjust the contribution ratio between modalities according to the characteristics of dynamic scenes, resulting in low utilization of important information, thereby affecting the tracking accuracy and the robustness of the model. The present invention provides a target tracking method and system based on RGB-T.
[0006] The technical solution provided by the embodiment of the present invention is as follows:
[0007] First aspect:
[0008] An RGB-T based target tracking method provided by an embodiment of the present invention includes:
[0009] S1: Obtain RGB image and TIR image of the detection target;
[0010] S2: respectively determining a search area and a template area of the RGB image and the TIR image, wherein the template area includes a fixed template and an online template;
[0011] S3: using a feature extraction network, extracting features from the search area and the template area, and determining a search area feature sequence, a fixed template feature sequence, and an online template feature sequence of both the RGB modality and the TIR modality, respectively;
[0012] S4: performing a splicing operation on the search area feature sequence, the fixed template feature sequence and the online template feature sequence to obtain a total feature sequence;
[0013] S5: Combining the self-attention operation and the Mask strategy to perform multimodal feature fusion on the total feature sequence to generate preliminary modal fusion features;
[0014] S6: Determine the self-attention weight of the RGB modality and the self-attention weight of the TIR modality;
[0015] S7: generating a final modality fusion feature according to the preliminary modality fusion feature, the self-attention weight of the RGB modality, and the self-attention weight of the TIR modality;
[0016] S8: Input the final modal fusion feature into the target prediction head to determine the current position of the detected target.
[0017] Second aspect:
[0018] An RGB-T based target tracking system provided by an embodiment of the present invention includes:
[0019] processor;
[0020] A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the RGB-T-based target tracking method as described in the first aspect is implemented.
[0021] The third aspect:
[0022] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the RGB-T-based target tracking method as described in the first aspect is implemented.
[0023] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0024] (1) In the embodiment of the present invention, the search area feature sequence, the fixed template feature sequence and the online template feature sequence are spliced to generate a total feature sequence, and the self-attention operation and the Mask strategy are combined to perform multimodal feature fusion to generate preliminary modal fusion features. Not only the search area and template area features of the RGB and TIR modalities are completely preserved, but also the interference of irrelevant or noise information is effectively suppressed, further enhancing the mining of modality-specific information and greatly improving the tracking accuracy.
[0025] (2) In an embodiment of the present invention, by determining the self-attention weight of the RGB modality and the self-attention weight of the TIR modality, the final modality fusion feature is generated according to the preliminary modality fusion feature, the self-attention weight of the RGB modality and the self-attention weight of the TIR modality. The preliminary modality fusion feature is combined with the dynamic weight to generate a more scene-adaptive final modality fusion feature, which effectively avoids the problem of low utilization of important information caused by the fixed weight method and significantly improves the tracking accuracy and the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0027] Figure 1 A schematic diagram of a flow chart of an RGB-T based target tracking method provided by an embodiment of the present invention;
[0028] Figure 2 A schematic diagram of a target tracking structure provided by an embodiment of the present invention;
[0029] Figure 3 A schematic diagram of a process for updating an online template provided by an embodiment of the present invention;
[0030] Figure 4 A schematic diagram of the structure of an RGB-T based target tracking system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0031] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0032] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0033] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0034] In the embodiments of the present invention, sometimes the subscripts such as W 1 It may be written in non-subscript form such as W1. When the difference is not emphasized, the meaning is the same.
[0035] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0036] Reference Manual Attached Figure 1 , shows a flow chart of a target tracking method based on RGB-T provided in an embodiment of the present invention.
[0037] Reference Manual Attached Figure 2 , showing a schematic structural diagram of a target tracking provided by an embodiment of the present invention.
[0038] Figure 2 In the process, firstly, the RGB image and TIR image of the detection target are obtained, including Initial-Template, Search-Region and Online-Template, and the RGB image and TIR image are sequentially passed through PatchEmbedding, TransformerBlock, Regroup, self-Attention, Mask, and Softmax to generate preliminary modality fusion features; the RGB image and TIR image are sequentially passed through MLP, SUM, Regroup, and Softmax to determine the self-attention weight of the RGB modality and the self-attention weight of the TIR modality;
[0039] The final modality fusion feature is obtained by performing a weighted sum of the generated preliminary modality fusion features and the self-attention weights of the RGB modality and the self-attention weights of the TIR modality;
[0040] The final modal fusion features are input into the target prediction head to determine the current position of the detected target.
[0041] The embodiment of the present invention provides a target tracking method based on RGB-T, which can be implemented by a target tracking device based on RGB-T, and the target tracking device based on RGB-T can be a terminal or a server. The processing flow of the target tracking method based on RGB-T can include the following steps:
[0042] S1: Obtain RGB image and TIR image of the detection target.
[0043] Among them, an image representation method based on three color channels of red, green and blue is the most common image type in current digital image processing. RGB images can represent almost all colors visible to the human eye by mixing the different intensities of these three basic color channels.
[0044] Among them, TIR image, or thermal infrared image, is an imaging method based on the temperature distribution of the object surface. TIR image generates images by detecting the thermal radiation (infrared radiation) emitted by the object. Unlike visible light images (RGB images) that rely on light reflection, it can be imaged at night or in a dark environment. It is an important tool in the fields of target tracking, security monitoring and industrial detection.
[0045] In this invention, the multimodal fusion method of combining RGB images and TIR images is used to fully utilize the high expressiveness of RGB images for color and texture details, as well as the stability and robustness of TIR images in low light, at night and under complex backgrounds. Through the complementarity of RGB images and TIR images, not only the accuracy of target detection and tracking can be significantly improved, but also the adaptability of the model in different environments (such as daytime, nighttime, occlusion or bad weather) can be enhanced.
[0046] S2: Determine the search area and the template area of the RGB image and the TIR image respectively, wherein the template area includes a fixed template and an online template.
[0047] Among them, the search region is a key concept in target detection and tracking tasks. It is used to define a smaller range in the input image, limit the target search area, reduce computational complexity, and improve the tracking efficiency and accuracy of the model.
[0048] The template area is an area used to represent the appearance features of the target, and generally includes the fixed features of the target in the first frame (fixed template) and the features dynamically updated during the tracking process (online template).
[0049] Among them, the online template area is the target feature that is dynamically updated during the tracking process, reflecting the short-term changes of the target.
[0050] Among them, the fixed template area is the feature extracted from the target in the first frame, which usually remains unchanged during the tracking process.
[0051] In the present invention, the design of the search area and the template area (including fixed templates and online templates) can significantly improve the efficiency, robustness and adaptability of the target tracking task. The search area reduces the computational complexity by limiting the possible location range of the target, while reducing background interference and improving the tracking accuracy.
[0052] S3: Use the feature extraction network to extract features from the search area and the template area, and determine the search area feature sequence, fixed template feature sequence, and online template feature sequence of the RGB modality and the TIR modality, respectively.
[0053] Optionally, the feature extraction network is ViT.
[0054] Specifically, the search area and template area images of RGB and TIR modalities are segmented into patches of fixed size and flattened into token sequences;
[0055] The token sequence is input into the Transformer module with shared weights, and the intra-modal features are extracted through the self-attention mechanism. The interaction between the search area and the template area is realized, and the search area feature sequence, fixed template feature sequence and online template feature sequence of the RGB modality and TIR modality are obtained.
[0056] In the present invention, by dividing the RGB and TIR images into patches of fixed size, each patch can capture the local detail features of the image, and the self-attention mechanism of the Transformer can globally interact with all patches, thereby capturing the global context information of the image. At the same time, through the Transformer module with shared weights, the RGB and TIR modalities complete feature extraction respectively, and the interaction between the search area within the modality and the template area is realized through the self-attention mechanism. This design allows each modality to fully extract its own features, explore the texture details of the RGB modality and the thermal radiation characteristics of the TIR modality, and provide higher quality features for inter-modal interaction.
[0057] S4: performing a splicing operation on the search area feature sequence, the fixed template feature sequence and the online template feature sequence to obtain a total feature sequence.
[0058] In a possible implementation, the total feature sequence is specifically:
[0059]
[0060] Among them, X represents the total feature sequence, Z V Represents a fixed template of visible light, Z I Indicates the fixed template of thermal infrared, X V represents the search area of visible light, X I represents the search area of thermal infrared, An online template representing visible light, Online template representing thermal infrared.
[0061] In the present invention, the features of the RGB modality and the TIR modality are spliced into the same total feature sequence, so that the model can simultaneously utilize the color and texture information of the RGB modality and the thermal radiation information of the TIR modality. By integrating multimodal information, the complementarity of the two modalities is fully utilized, thereby enhancing the expressive power of the target features. At the same time, by splicing the features of the search area, fixed template, and online template into a unified sequence, the model can directly realize the interaction between modalities (RGB and TIR) and regions (search area and template area) using the self-attention mechanism. In this way, the complementary information between the modalities can be captured, and the relationship between the search area and the template area can be dynamically matched, thereby improving the accuracy of target tracking.
[0062] S5: Combine the self-attention operation and the Mask strategy to perform multimodal feature fusion on the total feature sequence to generate preliminary modal fusion features.
[0063] Among them, self-attention is a mechanism that can dynamically capture the relationship between elements in sequence data (such as text, images, time series, etc.). It highlights important features by calculating the correlation weight of each element (such as token, pixel block) in the sequence with other elements, thereby generating richer feature representations.
[0064] Among them, the Mask strategy is a technology widely used in deep learning models (especially models based on the self-attention mechanism). Its core idea is to introduce a mask in the input data or attention weight matrix to shield certain specific information or calculation areas, so as to achieve specific goals, such as optimizing model performance, avoiding interference from irrelevant information, or highlighting key features.
[0065] In a possible implementation, S5 specifically includes:
[0066] The interactive information between the search area and the template area of the same modality is extracted through self-attention operation, and the Mask strategy is used to mask the exchanged interactive information between the search area and the template area to generate preliminary modality fusion features.
[0067] Combined with the Mask strategy, the interactive information between the search area and the template area of the same modality is extracted through self-attention operation.
[0068] Specifically, the information between the search area and the template has been exchanged within the modalities of the two modalities, so we mask this part of the information. In this way, we mask out the calculation between the search area and the template of the same modality. In addition, since the information of the online template has changed compared to the fixed initial template after the update, the attention between different templates may introduce additional noise, so this part is also masked.
[0069] In a possible implementation, the calculation formula for mask processing is:
[0070] A m =A⊙M
[0071] Among them, A m represents the attention weight matrix after masking, A represents the attention weight matrix, M represents the mask matrix, and ⊙ represents the element-by-element product.
[0072] In one possible implementation, the calculation formula for the self-attention operation is:
[0073] Q,K,V=XW q ,XW k ,XW v
[0074]
[0075] Among them, Q represents the query vector, K represents the key vector, V represents the value vector, and W q It is used to map the total feature sequence into a query vector, W k It is used to map the total feature sequence into a key vector, W v It is used to map the total feature sequence into a value vector, A represents the attention weight matrix, QK T represents the similarity between the query vector and the key vector, and d represents the dimension of the query vector or the key vector.
[0076] In a possible implementation, the preliminary modality fusion features are specifically:
[0077]
[0078] in, represents the preliminary modality fusion feature of RGB modality, Z V Represents a fixed template of visible light, X V represents the search area for visible light, An online template representing visible light, represents the preliminary modal fusion features of the TIR modality, Z i Indicates the fixed template of thermal infrared, X i represents the search area of thermal infrared, Online template representing thermal infrared.
[0079] In the present invention, the Mask strategy avoids unnecessary calculations and information redundancy by masking the exchanged information between the same-modality search area and the template area during the calculation process, improves calculation efficiency, and makes the model more focused on those areas that have not yet interacted or are irrelevant. At the same time, the self-attention operation can capture the correlation between each element (such as the search area or the template area). This intra-modal interaction can effectively enhance the expressiveness of the modal features, allowing the model to fully understand the local and global features of the target in each modality, and improve the accuracy and robustness of tracking.
[0080] Furthermore, through the self-attention mechanism and masking strategy, it is possible to accurately control which areas' features should be paid attention to, thereby performing efficient feature matching between the template area and the search area. This precise feature matching improves the ability to distinguish between the target and the background, and can effectively avoid mismatching in complex backgrounds or when the similarity between targets is high.
[0081] S6: Determine the self-attention weight of the RGB modality and the self-attention weight of the TIR modality.
[0082] In a possible implementation, S6 specifically includes:
[0083] S601: Extract the interaction information between the search area and the template according to the attention weight matrix in the self-attention operation:
[0084]
[0085] Among them, A selecteed represents the part of the total attention map selected for feature interaction, B represents the batch size, H represents the number of attention heads, and L x Indicates the length of the search image area, L z Indicates the length of the template image area.
[0086] S602: By means of a deformation operation, the format of the interactive information is adjusted to a format that can be processed by the multi-layer perceptron.
[0087] S603: Use a multi-layer perceptron to transform the interactive information after the adjusted format to generate a transformed attention weight matrix A attn .
[0088] Among them, Multi-Layer Perceptron (MLP) is a feedforward neural network consisting of multiple linear layers (fully connected layers) and nonlinear activation functions. It is one of the most basic artificial neural network architectures and is widely used in tasks such as classification and regression.
[0089] In the present invention, the interactive information is formatted and input into subsequent calculations through MLP, which reduces the complexity of feature processing, optimizes the calculation process, and improves the calculation efficiency, which is particularly suitable for real-time tasks.
[0090] S604: transform the attention weight matrix A attn Add in the column direction to get the attention weight matrix after column summation
[0091] S605: Add the column-wise summed attention weight matrix Segmentation is performed according to the template features of RGB mode and TIR mode:
[0092] Z′ v , Z′ i =split(A sum , dim=3)
[0093] Among them, Z′ v Represents the attention matrix of the visible part after segmentation, Z′ i represents the attention matrix of the thermal infrared part after segmentation, split represents the segmentation operation, and A sum It represents the attention weight matrix after column summation, and dim represents the segmentation dimension.
[0094] S606: Rejoin the segmentation results to obtain a joining matrix:
[0095] A′=concat(Z r , Z t , dim=2)
[0096] Wherein, A′ represents the concatenation matrix, and concat represents the concatenation operation.
[0097] S607: Normalize the concatenated matrix to obtain a normalized weight matrix:
[0098]
[0099] Among them, Anorm represents the normalized weight matrix, i, j, k represent the dimension variables, It represents the exponential operation on the concatenated matrix in the i, j, k dimensions, B represents the batch size, H′ represents the height of the feature, and W represents the width of the feature.
[0100] S608: In the normalized weight matrix, insert a full 1 vector for the search area to obtain the self-attention weight of the RGB modality and the self-attention weight of the TIR modality:
[0101]
[0102] Among them, W RGB represents the self-attention weight of RGB modality, W TIR Represents the self-attention weight of the TIR modality.
[0103] In the present invention, through the weighted self-attention weights of the RGB modality and the self-attention weights of the TIR modality, the model can dynamically adjust the weights between different regions (such as the template region and the search region) and modalities (RGB and TIR) according to the requirements of the task, thereby enhancing the model's attention to important target features, so that the attention allocation can be adjusted according to changes in the environment during the target tracking process, thereby improving the accuracy of target positioning. At the same time, by inserting a full 1 vector, a constant weight is maintained for the search area part to ensure that it is not affected by the modal weights of other regions, avoiding redundant calculations, improving model efficiency, and making the model more focused on the target area rather than background information, thereby avoiding interference from irrelevant information and optimizing the target tracking task.
[0104] S7: Generate the final modality fusion features based on the preliminary modality fusion features, the self-attention weights of the RGB modality, and the self-attention weights of the TIR modality.
[0105] In a possible implementation, the final modality fusion feature is specifically:
[0106]
[0107] in, represents the visible light feature after weighted update, Represents the thermal infrared features after weighted update.
[0108] Specifically, by adding these interacted features to the original related features through weighted addition, the fusion of the features of this modality with the information of another modality is achieved.
[0109] In the present invention, the self-attention weights of the RGB modality are combined with the self-attention weights of the TIR modality, and the features are weighted and summed according to the weights, which effectively integrates the information from different modalities and improves the overall feature expression ability of the multimodal target tracking model, so that it can perform better in a variety of environments. At the same time, the weighted fused features can flexibly adjust the contribution of each modality according to the needs of target tracking, thereby enhancing the adaptability of the target tracking model.
[0110] S8: Input the final modal fusion features into the target prediction head to determine the current position of the detected target.
[0111] Among them, the target prediction head is the core module in the target detection and tracking task, which is responsible for generating the position, category and confidence of the target based on the input features. It combines classification and regression tasks to not only provide target positioning information, but also evaluate the tracking accuracy of the model through confidence.
[0112] In the present invention, the target prediction head combines the self-attention mechanism and multimodal weighted fusion to enhance the model's adaptability to target appearance changes and environmental dynamics, while reducing the risk of false detection and missed detection. In addition, the fused features improve computational efficiency and avoid redundant calculations, enabling the model to process large-scale data in real time and maintain efficient tracking performance.
[0113] Reference Manual Attached Figure 3 , showing a schematic diagram of a process for updating an online template provided by an embodiment of the present invention.
[0114] In a possible implementation manner, after S8, the step further includes: updating the online template.
[0115] Updates to online templates include:
[0116] S901: Flatten the score map generated during the tracking process and extract the maximum score value in the score map.
[0117] S902: Determine whether the maximum score is greater than a preset score. If so, cut out a template area that meets the template size requirement according to the coordinates of the current frame tracking frame, and store the template area as the target online template. Otherwise, continue to use the current online template.
[0118] S903: When the number of frames running downward is greater than the preset update interval, the current online template is replaced with the target online template.
[0119] It should be noted that the number of frames running downward means that after every fixed number of frames, the model will update the online template to ensure that the template reflects the latest appearance of the target.
[0120] Specifically, the score map generated during the tracking process is first flattened, and the maximum score value is extracted. This score value can reflect the confidence of the model in the tracking effect of the current frame to a certain extent. Based on this feature, a predefined threshold m is set. When the maximum score value exceeds the threshold, the tracking effect of the current frame can be considered to be relatively reliable. In this case, the area that meets the template size requirements is cropped according to the coordinates of the current frame tracking box, and it is stored as a new online template to enhance the model's adaptability to dynamically changing targets. To ensure the periodicity of template updates, a fixed update interval n is designed. Whenever the tracking passes through n frames, the model will replace the current online template with the latest template stored in the template library. It should be pointed out that only one frame of online templates is saved in the template library, and the initial online template is the same as the fixed template. Within an update interval, if multiple templates that meet the update conditions appear, only the last frame is retained.
[0121] In the present invention, by updating the online template in real time, it is ensured that the template can always reflect the latest state of the target, which can significantly improve the accuracy of target tracking, especially when the target appearance changes. At the same time, by determining whether to update the template according to the tracking score of the current frame in each update cycle. If the score is lower than the preset threshold, the model will keep the existing online template unchanged, thereby avoiding the template being replaced immediately when the tracking is inaccurate, resulting in insufficient tracking accuracy.
[0122] Furthermore, by setting an update interval, it is ensured that the online template is not updated frequently, thereby avoiding adjustments in each frame, which affects the stability of the template.
[0123] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0124] (1) In the embodiment of the present invention, the search area feature sequence, the fixed template feature sequence and the online template feature sequence are spliced to generate a total feature sequence, and the self-attention operation and the Mask strategy are combined to perform multimodal feature fusion to generate preliminary modal fusion features. Not only the search area and template area features of the RGB and TIR modalities are completely preserved, but also the interference of irrelevant or noise information is effectively suppressed, further enhancing the mining of modality-specific information and greatly improving the tracking accuracy.
[0125] (2) In an embodiment of the present invention, by determining the self-attention weight of the RGB modality and the self-attention weight of the TIR modality, the final modality fusion feature is generated according to the preliminary modality fusion feature, the self-attention weight of the RGB modality and the self-attention weight of the TIR modality. The preliminary modality fusion feature is combined with the dynamic weight to generate a more scene-adaptive final modality fusion feature, which effectively avoids the problem of low utilization of important information caused by the fixed weight method and significantly improves the tracking accuracy and the robustness of the model.
[0126] Reference Manual Attached Figure 4 , showing a structural schematic diagram of an RGB-T based target tracking system provided by the present invention.
[0127] The present invention further provides an RGB-T based target tracking system 20, which is applied to the above-mentioned RGB-T based target tracking method, comprising:
[0128] Processor 201.
[0129] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201, the RGB-T-based target tracking method of the method embodiment is implemented.
[0130] The RGB-T based target tracking system 20 provided by the present invention can execute the above-mentioned RGB-T based target tracking method and achieve the same or similar technical effects. To avoid repetition, the present invention will not go into details.
[0131] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0132] (1) In the embodiment of the present invention, the search area feature sequence, the fixed template feature sequence and the online template feature sequence are spliced to generate a total feature sequence, and the self-attention operation and the Mask strategy are combined to perform multimodal feature fusion to generate preliminary modal fusion features. Not only the search area and template area features of the RGB and TIR modalities are completely preserved, but also the interference of irrelevant or noise information is effectively suppressed, further enhancing the mining of modality-specific information and greatly improving the tracking accuracy.
[0133] (2) In an embodiment of the present invention, by determining the self-attention weight of the RGB modality and the self-attention weight of the TIR modality, the final modality fusion feature is generated according to the preliminary modality fusion feature, the self-attention weight of the RGB modality and the self-attention weight of the TIR modality. The preliminary modality fusion feature is combined with the dynamic weight to generate a more scene-adaptive final modality fusion feature, which effectively avoids the problem of low utilization of important information caused by the fixed weight method and significantly improves the tracking accuracy and the robustness of the model.
[0134] It should be understood that the processor in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0135] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0136] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0137] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0138] In the present invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0139] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0140] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0141] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0142] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0143] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0144] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0145] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0146] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the RGB-T based target tracking method as described in the method embodiment is implemented.
[0147] A computer-readable storage medium provided by the present invention can implement the steps and effects of the RGB-T-based target tracking method of the above method embodiment. To avoid repetition, the present invention will not go into details.
[0148] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0149] (1) In the embodiment of the present invention, the search area feature sequence, the fixed template feature sequence and the online template feature sequence are spliced to generate a total feature sequence, and the self-attention operation and the Mask strategy are combined to perform multimodal feature fusion to generate preliminary modal fusion features. Not only the search area and template area features of the RGB and TIR modalities are completely preserved, but also the interference of irrelevant or noise information is effectively suppressed, further enhancing the mining of modality-specific information and greatly improving the tracking accuracy.
[0150] (2) In an embodiment of the present invention, by determining the self-attention weight of the RGB modality and the self-attention weight of the TIR modality, the final modality fusion feature is generated according to the preliminary modality fusion feature, the self-attention weight of the RGB modality and the self-attention weight of the TIR modality. The preliminary modality fusion feature is combined with the dynamic weight to generate a more scene-adaptive final modality fusion feature, which effectively avoids the problem of low utilization of important information caused by the fixed weight method and significantly improves the tracking accuracy and the robustness of the model.
[0151] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
[0152] There are a few points to note:
[0153] (1) The drawings of the embodiments of the present invention only relate to the structures related to the embodiments of the present invention, and other structures may refer to the general design.
[0154] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present invention, the thickness of the layers or regions is exaggerated or reduced, that is, these drawings are not drawn according to the actual scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or there may be intermediate elements.
[0155] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to obtain new embodiments.
[0156] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A target tracking method based on RGB-T, characterized in that: include: S1: Obtain RGB image and TIR image of the detection target; S2: respectively determining a search area and a template area of the RGB image and the TIR image, wherein the template area includes a fixed template and an online template; S3: using a feature extraction network, extracting features from the search area and the template area, and determining a search area feature sequence, a fixed template feature sequence, and an online template feature sequence of both the RGB modality and the TIR modality, respectively; S4: performing a splicing operation on the search area feature sequence, the fixed template feature sequence and the online template feature sequence to obtain a total feature sequence; S5: Combining the self-attention operation and the Mask strategy to perform multimodal feature fusion on the total feature sequence to generate preliminary modal fusion features; S6: Determine the self-attention weight of the RGB modality and the self-attention weight of the TIR modality; S7: generating a final modality fusion feature according to the preliminary modality fusion feature, the self-attention weight of the RGB modality, and the self-attention weight of the TIR modality; S8: Input the final modal fusion feature into the target prediction head to determine the current position of the detected target.
2. The RGB-T based target tracking method according to claim 1, characterized in that: The total feature sequence is specifically: Among them, X represents the total feature sequence, Z V Represents a fixed template of visible light, Z I Indicates the fixed template of thermal infrared, X V represents the search area of visible light, X I represents the search area of thermal infrared, An online template representing visible light, Online template representing thermal infrared.
3. The RGB-T based target tracking method according to claim 1, characterized in that: The S5 specifically includes: The interactive information between the search area and the template area of the same modality is extracted through self-attention operation, and the Mask strategy is used to mask the exchanged interactive information between the search area and the template area to generate preliminary modality fusion features; Combined with the Mask strategy, the interactive information between the search area and the template area of the same modality is extracted through self-attention operation.
4. The RGB-T based target tracking method according to claim 3, characterized in that: The calculation formula for the mask processing is: A m =A⊙M Among them, A m represents the attention weight matrix after masking, A represents the attention weight matrix, M represents the mask matrix, and ⊙ represents the element-by-element product.
5. The RGB-T based target tracking method according to claim 1, characterized in that: The calculation formula of the self-attention operation is: Q,K,V=xW q ,XW k ,XW v Among them, Q represents the query vector, K represents the key vector, V represents the value vector, and W q It is used to map the total feature sequence into a query vector, W k It is used to map the total feature sequence into a key vector, W v It is used to map the total feature sequence into a value vector, A represents the attention weight matrix, represents the similarity between the query vector and the key vector, and d represents the dimension of the query vector or the key vector.
6. The RGB-T based target tracking method according to claim 1, characterized in that: The preliminary modality fusion features are specifically: in, represents the preliminary modality fusion feature of RGB modality, Z V Represents a fixed template of visible light, X V represents the search area for visible light, An online template representing visible light, represents the preliminary modal fusion features of the TIR modality, Z i Indicates the fixed template of thermal infrared, X i represents the search area of thermal infrared, Online template representing thermal infrared.
7. The target tracking method based on RGB-T according to claim 1, characterized in that: The S6 specifically includes: S601: Extracting the interaction information between the search area and the template according to the attention weight matrix in the self-attention operation: Among them, A selected represents the part of the total attention map selected for feature interaction, B represents the batch size, H represents the number of attention heads, and L x Indicates the length of the search image area, L z Indicates the length of the template image area; S602: adjusting the format of the interactive information to a format that can be processed by a multi-layer perceptron through a deformation operation; S603: Use a multi-layer perceptron to transform the interactive information after the adjusted format to generate a transformed attention weight matrix A attn ; S604: transform the attention weight matrix A attn Add in the column direction to get the attention weight matrix after column summation S605: Add the column-wise summed attention weight matrix Segmentation is performed according to the template features of RGB mode and TIR mode: Z′ v ,Z′ i =split(A sum ,dim=3) Among them, Z′ v Represents the attention matrix of the visible part after segmentation, Z′ i represents the attention matrix of the thermal infrared part after segmentation, split represents the segmentation operation, and A sum represents the attention weight matrix after column summation, and dim represents the segmentation dimension; S606: Rejoin the segmentation results to obtain a joining matrix: A′=concat(Z′ v ,Z′ i ,dim=2) Among them, A′ represents the concatenation matrix, and concat represents the concatenation operation; S607: Normalize the concatenated matrix to obtain a normalized weight matrix: Among them, A norm represents the normalized weight matrix, i, j, k represent the dimension variables, represents the exponential operation of the concatenated matrix in the i, j, k dimensions, B represents the batch size, H′ represents the height of the feature, and W represents the width of the feature; S608: In the normalized weight matrix, insert a full 1 vector for the search area to obtain the self-attention weight of the RGB modality and the self-attention weight of the TIR modality: Among them, W RGB represents the self-attention weight of RGB modality, W TIR Represents the self-attention weight of the TIR modality.
8. The RGB-T based target tracking method according to claim 5, characterized in that: The final modality fusion feature is specifically: in, represents the visible light feature after weighted update, Represents the thermal infrared features after weighted update.
9. The RGB-T based target tracking method according to claim 1, characterized in that: After S8, the method further includes: updating the online template; Updating the online template specifically includes: S901: Flatten the score map generated during the tracking process, and extract the maximum score value in the score map; S902: Determine whether the maximum score value is greater than a preset score value; if so, cut out a template area that meets the template size requirement according to the coordinates of the current frame tracking frame, and store the template area as the target online template; otherwise, continue to use the current online template; S903: When the number of frames running downward is greater than a preset update interval, the current online template is replaced with the target online template.
10. A target tracking system based on RGB-T, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the RGB-T-based target tracking method according to any one of claims 1 to 9 is implemented.
Citation Information
Cited By
RGB-T tracking method based on consistency information guide enhancement
CN121236118A
Consistency information based guidance enhanced rgb-t tracking method
CN121236118B
RGB-T target tracking method and device
CN121414788A
RGB-t object tracking method and device
CN121414788B