Cooperative dominant multi-mode RGB-T fusion target tracking method and device based on double-current symmetric adapter bridging
Through the collaborative dominant multimodal fusion method of dual-stream symmetric architecture and modal reciprocating adapter bridge, the problem of insufficient modal information interaction in extreme scenarios is solved, and the target tracking effect with high accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202510407106.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-08-12
AI Technical Summary
The existing RGB-T multimodal fusion tracking method fails to fully consider the effective interaction and dynamic fusion of modal information in different complex and variable scenarios, resulting in failed tracking or insufficient robustness in extreme scenarios.
The dual-stream symmetric architecture and modal reciprocity adapter bridge are adopted, and the parameter sharing and modal prompt information between the dual-streams are used, combined with the collaborative dominant multi-modal fusion module, the dominant advantages of the modal in different scenarios are dynamically perceived, and the weighted fusion of the main and auxiliary modals is achieved.
It improves tracking accuracy and robustness under extreme conditions, demonstrates multimodal complementarity, and achieves high-precision target tracking.
Smart Images

Figure CN120471952A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multimodal fusion target tracking, and in particular relates to a collaborative-dominant multimodal RGB-T fusion target tracking method and device based on dual-stream symmetric adapter bridging. Background Art
[0002] RGB-T multimodal fusion tracking has garnered widespread attention in recent years. Single RGB visible light cannot effectively track objects in low light conditions or extreme weather conditions. Therefore, the highly penetrating TIR thermal infrared modality, which is independent of visible light and possesses strong penetration, is introduced for complementary and enhanced performance. By leveraging the complementary nature of RGB and TIR modalities, RGB-T tracking can achieve greater competitiveness in challenging scenarios. However, as a multimodal target tracking task, robust real-time tracking remains a key challenge, effectively utilizing the limited modal information provided by the scene in complex and changing scenarios. Most existing RGB-T tracking methods have not fully considered this issue, and their design approaches can be broadly categorized into three directions: 1) Focusing on achieving a better unified representation of RGB and TIR modalities, i.e., modal fusion achieved through averaging between modalities or splicing operations with RGB as the primary modality. However, such methods overlook the fact that a particular modality may have a greater tracking advantage in specific scenarios. 2) Focusing on information interaction between modalities, leveraging the complementarity between modalities to achieve efficient tracking, these methods, however, do not fully consider that in extreme scenarios, both RGB and TIR modalities may lose their tracking effectiveness, leading to tracking failure. 3) Focusing on the real-time and robustness of RGB-T tracking, enabling robust tracking in challenging environments, these methods only employ fixed modal fusion and interaction strategies, which do not fully exploit the complementary advantages of different modalities. In summary, existing methods do not fully consider the effective interaction and dynamic fusion of information between modalities, nor do they fully demonstrate adaptability and robustness to complex and changing situations. Summary of the Invention
[0003] The present invention aims to overcome the above-mentioned shortcomings of the prior art and provide a collaborative-dominant multimodal RGB-T fusion target tracking method and device based on dual-stream symmetric adapter bridging.
[0004] The present invention includes a dual-stream symmetric architecture to fully expand and explore scenarios in which a certain scene sequence is in extremely challenging conditions; a modal reciprocal adapter is used to appropriately provide effective modal prompt information to each other during the training process of the dual-stream input; a collaborative leading multimodal fusion module is used to dynamically perceive the dominant mode that plays a more important role in tracking in different scenarios, realize dynamic weighted fusion of the main and auxiliary modes, and then make the optimal decision on the utilization of different modes in different scenarios, ultimately realizing multimodal fusion and tracking.
[0005] To achieve the above object, the technical solutions adopted by the present invention are as follows:
[0006] The first aspect of the present invention relates to a collaborative, leading, and multimodal RGB-T fusion target tracking method based on a dual-stream symmetric adapter bridge. Through a dual-stream symmetric architecture, this method fully expands and explores extremely challenging tracking scenarios. A modality reciprocity adapter, acting as a bridge between the two streams, provides effective modality cue interaction information for dual-stream training. A collaborative, leading, and multimodal fusion module fully adapts to and exploits the ever-changing relationship between visible light and thermal infrared primary and secondary modal information in different scenarios. The method includes the following steps:
[0007] 1) Obtain the multimodal video image sequence to be processed
[0008] 2) Using a dual-stream symmetric architecture, the input data is passed as the original data and the masked data into their respective Transformer encoders, and the two streams are jointly trained.
[0009] 3) Introduce several lightweight modal reciprocal adapters (MRAs) between the two streams to provide effective modal prompt information to each other during the training process of the two streams.
[0010] 4) The collaborative dominant multimodal fusion module (CDMF) is used to sense the advantages of the mode that plays a major role in tracking in different scenarios, dynamically use it as the dominant mode, and the other mode as the auxiliary mode. The two modes are weighted and fused, and work together to make the optimal decision on the utilization of different modes in different scenarios.
[0011] 5) The fused multimodal features are fed into the prediction head to obtain the target's position in the next frame, resulting in the final tracking result. A loss function, consisting of the weighted sum of three loss functions: weighted focal loss, L1 loss, and IoU loss, is used to compare the predicted target position with the ground truth for optimization training.
[0012] The step 1) of obtaining a multimodal video image sequence to be processed includes the following sub-steps:
[0013] (11) The multimodal data in the training set are preprocessed, that is, the search area is resized to 256×256, the template is resized to 128×128, and then the patch is embedded as the input of the Transformer encoder.
[0014] The dual-stream symmetrical architecture described in step 2) includes the following sub-steps:
[0015] (21) The two streams share parameters and conduct joint training, where the original data stream does not process the data in any way, and obtains X RGB 、X TIR ; The mask processing flow randomly masks 30% of the multimodal data and obtains
[0016] (22) For the original data stream, the overall operation of the multimodal features is as follows:
[0017] H RGB =[CDMF(X RGB ,MRA(T));CDMF(Z RGB ,MRA(T))] (1)
[0018] H TIR =[CDMf(X TIR ,MRA(R));CDMF(Z TIR ,MRA(R))] (2)
[0019] Among them, X RGB 、Z RGB 、X TIR 、Z TIR denotes the search area and template token of visible light and thermal infrared, respectively. R and T denote the visible light and thermal infrared modal information from the opposing streams. h RGB 、H TIR represents the multimodal feature information obtained by splicing, MRA represents the modal reciprocal adapter proposed in the present invention, and CDMF represents the collaborative-dominant multimodal fusion module proposed in the present invention.
[0020] (23) For the mask processing flow, the overall operation of the multimodal features is as follows:
[0021]
[0022] in, denotes the search area and template token of visible light and thermal infrared from the mask processing stream, R and T denote the visible light and thermal infrared modal information from the opposing stream, and H′ RGB , H′ TIR Represents the multimodal feature information obtained by splicing.
[0023] The modality reciprocity adapter described in step 3) provides modality information prompts between the two streams, including the following sub-steps:
[0024] (31) When the search area token is used as input, a template token from another modality is specifically introduced as a prompt token to provide target-related information about the search area. When the template token is used as input, no additional prompt token information is required. Taking the search area token input as an example, the prompt process is as follows:
[0025] T′=X RGB +Z TIR (5)
[0026] Among them, X RGB Represents the RGB search area, Z TIR represents the TIR template, and T′ represents the formal input of the modal reciprocal adapter.
[0027] (32) The input token first undergoes a dimensionality reduction linear transformation to compress it to a lower dimension, and then passes through the GELU activation function to introduce nonlinearity and enhance its feature representation. The process is as follows:
[0028] T down =GELU(Dw(T′)) (6)
[0029] Where Dw represents the dimensionality reduction linear transformation, T down Represents the modal token after dimensionality reduction transformation and GELU activation.
[0030] (33) The features are further subjected to a linear transformation without changing their dimensions, followed by a dimensionality-increasing linear transformation to restore them to their original dimensions, and a residual connection with an initial token is attached to obtain the output of the modality reciprocal adapter, which is formulated as follows:
[0031] T up =Up(Linear(T down ))+T′ (7)
[0032] Among them, Linear represents a simple linear transformation operation, Up represents a dimension-raising linear transformation operation, and T up represents the final output of the modal reciprocal adapter after the dimensionality-increasing linear transformation.
[0033] The collaborative leading multimodal fusion module described in step 4) includes the following sub-steps for the dynamic fusion process of multimodal information:
[0034] (41) The input multimodal tokens first undergo GELU activation to suppress the position where the response drops below zero. Then, global average pooling is used to estimate the relative magnitude of the two modalities, thereby outputting two weighting factors γ1 and γ2. The specific formula is as follows:
[0035] γ1=GAP(GELU(T R )) (8)
[0036] γ2=GAP(GELU(T T )) (9)
[0037] Among them, T R 、T T represents the input multimodal token, GAP represents the global average pooling operation, and γ1 and γ2 represent the two obtained weighting factors.
[0038] (42) To obtain unbiased multimodal information, each mode is multiplied by the corresponding weighting factor to obtain a coordinated and balanced modal result, which is as follows:
[0039] T w =T R *γ1+T T *γ2 (10)
[0040] Among them, T w represents the coordinated and balanced modal results obtained by weighted product.
[0041] (43) To prevent the previous operation from accidentally amplifying the original mode, T w Divide by the sum of γ1 and γ2 to make its amplitude close to the original data. At the same time, in order to further improve the obtained modal results, a trainable parameter β is introduced to normalize the relative influence of the two modes together with the layer normalization operation. The specific formula is as follows:
[0042]
[0043] Among them, LN is the layer normalization operation, T N Represents the final result of collaborative-dominant multimodal fusion.
[0044] Step 5) sends the fused multimodal features to the prediction head to obtain the position of the target in the next frame to obtain the final tracking result, which includes the following sub-steps:
[0045] (51) Two loss functions are used to simultaneously supervise the training of the original data stream and the mask data stream. The weighted sum of the three loss functions, weightfocal loss, L1 loss, and IoU loss, is used to compare the predicted target position with the groundtruth for optimization training. The loss function calculation process is as follows:
[0046]
[0047] L total =L raw +θL processed (14)
[0048] Among them, L cls represents the weighted focal loss of the training branch, L iou It is used to optimize the overlap between the predicted bounding box and the groundtruth. L1 is the measure of the mean absolute difference between the predicted value and the true value. iou =2 and are two trade-off parameters, and θ is set to 0.3 as a scaling factor to balance the two flows.
[0049] The second aspect of the present invention relates to a collaborative dominant multimodal RGB-T fusion target tracking device based on dual-stream symmetric adapter bridging, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the collaborative dominant multimodal RGB-T fusion target tracking method based on dual-stream symmetric adapter bridging of the present invention.
[0050] The third aspect of the present invention relates to a computer-readable storage medium having a computer program stored thereon, characterized in that when the program is executed by a processor, the collaborative dominant multimodal RGB-T fusion target tracking method based on dual-stream symmetric adapter bridging of the present invention is implemented.
[0051] The advantages of the present invention are:
[0052] 1) This paper proposes a novel dual-stream symmetric adapter bridging collaborative-dominant multimodal RGB-T fusion method, which comprehensively considers the shortcomings of existing methods in modal interaction, fusion and robustness. By utilizing dual-stream complementary modal data enhancement and adapter bridging for modal interaction and prompting, as well as a collaborative-dominant multimodal fusion strategy, it can significantly improve robustness and accuracy even under extreme conditions, demonstrate excellent multimodal complementarity, and achieve high tracking accuracy.
[0053] 2) We designed a two-stream symmetric architecture and jointly trained the two streams with shared parameters. The inputs are fed into separate Transformer encoders as the original data stream and the mask processing stream, respectively. This allows us to fully simulate and explore tracking scenarios in both routine and extremely challenging scenarios for the same scene.
[0054] 3) Design a modality reciprocal adapter as a bridge between the two streams, providing effective modality prompt interaction information for dual-stream training;
[0055] 4) Design a collaborative dominant multimodal fusion module to fully adapt to and exploit the ever-changing relationship between visible light and thermal infrared primary and secondary modal information in different scenarios, and make a fusion strategy more suitable for the current scenario for multimodal RGB-T information fusion. Compared with the method of only taking RGB as the dominant modality for fusion and tracking, it better utilizes the advantages of the more dominant modal information in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a flow chart of the method of the present invention;
[0057] Figure 2 Schematic diagram of the network architecture of the method of the present invention;
[0058] Figure 3 This is a conceptual diagram of the detailed process of feature interaction and fusion of the method of the present invention;
[0059] Figure 4 is a diagram of the modal reciprocal adapter architecture of the method of the present invention;
[0060] Figure 5 This is a diagram of the collaborative-dominant multimodal fusion architecture of the method of the present invention. Figure 6 It is a schematic diagram of the device of the present invention. DETAILED DESCRIPTION
[0061] In order to make the purpose, technical solutions and advantages of this application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application. The embodiments described are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0063] Example 1
[0064] like Figure 1 As shown, this embodiment proposes a collaborative-dominant multimodal RGB-T fusion target tracking method based on dual-stream symmetric adapter bridging, including the following steps:
[0065] Step S1: Obtain a multimodal video image sequence to be processed
[0066] The multimodal data in the training set are preprocessed by adjusting the search area to 256×256 and the template to 128×128, and then embedded as the input of the Transformer encoder through patch embedding.
[0067] Step S2: Using a dual-stream symmetric architecture, the input data is passed as the original data and the masked data into their respective Transformer encoders, and the two streams are jointly trained.
[0068] Among them, the original data stream does not do any processing on the data, and obtains X RGB 、X TIR 、Z RGB 、Z TIR ; The mask processing flow randomly masks 30% of the multimodal data and obtains
[0069] For the original data stream, the overall operation of the multimodal features is as follows:
[0070] H RGB =[CDMF(X RGB ,MRA(T));CDMF(Z RGB ,MRA(T))] (1)
[0071] H TIR =[CDMF(X TIR ,MRA(R));CDMF(Z TIR ,MRA(R))] (2)
[0072] Among them, X RGB 、Z RGB 、X TIR 、Z TIR denotes the search area and template token of visible light and thermal infrared, respectively. R and T denote the visible light and thermal infrared modal information from the opposing streams. H RGB 、H TIR represents the multimodal feature information obtained by splicing, MRA represents the modal reciprocal adapter proposed in the present invention, and CDMF represents the collaborative-dominant multimodal fusion module proposed in the present invention.
[0073] For the mask processing flow, the overall operation of the multimodal features is as follows:
[0074]
[0075] in, denotes the search area and template token of visible light and thermal infrared from the mask processing stream, R and T denote the visible light and thermal infrared modal information from the opposing stream, and H′ RGB , H′TIR Represents the multimodal feature information obtained by splicing.
[0076] Step S3: Introduce several lightweight modality reciprocal adapters (MRAs) between the two streams to provide effective modality prompt information to each other during the training process of the two streams.
[0077] When the search area token is used as input, a template token from another modality is specifically introduced as a hint token to provide target-related information about the search area. When the template token is used as input, no hint token information is required. Taking the search area token input as an example, the hint process is as follows:
[0078] T′=X RGB +Z TIR (5)
[0079] Among them, X RGB Represents the RGB search area, Z TIR represents the TIR template, and T′ represents the formal input of the modal reciprocal adapter.
[0080] The input token first undergoes a dimensionality reduction linear transformation to compress it to a lower dimension, and then passes through the GELU activation function to introduce nonlinearity and enhance its feature representation. The process is as follows:
[0081] T down =GELU(Dw(T′)) (6)
[0082] Where Dw represents the dimensionality reduction linear transformation, T down Represents the modal token after dimensionality reduction transformation and GELU activation.
[0083] The features are further linearly transformed without changing their dimensions, and then linearly transformed to restore the original dimensions, with a residual connection of the initial token, to obtain the output of the modality reciprocal adapter, which is formulated as follows:
[0084] T up =Up(Linear(T down ))+T′ (7)
[0085] Among them, Linear represents a simple linear transformation operation, Up represents a dimension-raising linear transformation operation, and T up represents the final output of the modal reciprocal adapter after the dimensionality-increasing linear transformation.
[0086] Step S4: Use the collaborative dominant multimodal fusion module (CDMF) to perceive the advantages of the mode that plays a major role in tracking in different scenarios, dynamically use it as the dominant mode, and the other mode as the auxiliary mode. The two modes are weighted and fused to work together to make the optimal decision on the utilization of different modes in different scenarios.
[0087] The input multimodal tokens first undergo GELU activation to suppress the position where the response drops below zero. Then, global average pooling is used to estimate the relative magnitude of the two modalities, thereby outputting two weighting factors γ1 and γ2. The specific formula is as follows:
[0088] γ1=GAP(GELU(T R )) (8)
[0089] γ2=GAP(GELU(T T )) (9)
[0090] Among them, T R 、T T represents the input multimodal token, GAP represents the global average pooling operation, and γ1 and γ2 represent the two obtained weighting factors.
[0091] In order to obtain unbiased multimodal information, each mode is multiplied by the corresponding weighting factor to obtain a coordinated and balanced modal result. The formula is as follows:
[0092] T w =T R *γ1+T T *γ2 (10)
[0093] Among them, T w represents the coordinated and balanced modal results obtained by weighted product.
[0094] To prevent the previous operation from accidentally amplifying the original mode, w Divide by the sum of γ1 and γ2 to make its amplitude close to the original data. At the same time, in order to further improve the obtained modal results, a trainable parameter β is introduced to normalize the relative influence of the two modes together with the layer normalization operation. The specific formula is as follows:
[0095]
[0096] Among them, LN is the layer normalization operation, T N Represents the final result of collaborative-dominant multimodal fusion.
[0097] In step S5, the fused multimodal features are fed into the prediction head to obtain the target's position in the next frame, thereby obtaining the final tracking result. The predicted target position is compared with the ground truth using a loss function consisting of the weighted sum of the three loss functions: weighted focal loss, L1 loss, and IoU loss, for optimization training.
[0098] Two Loss functions are used to simultaneously supervise the training of the original data stream and the mask data stream. The weighted sum of the three loss functions, weightfocal loss, L1 loss, and IoU loss, is used to compare the predicted target position with the Groundtruth for optimization training. The loss function calculation process is as follows:
[0099]
[0100] L total =L raw +θL processed (14)
[0101] Among them, L cls represents the weighted focal loss of the training branch, L iou It is used to optimize the overlap between the predicted bounding box and the groundtruth. L1 is the measure of the mean absolute difference between the predicted value and the true value. iou =2 and are two trade-off parameters, and θ is set to 0.3 as a scaling factor to balance the two flows.
[0102] Example 2
[0103] Reference Figure 3 This embodiment provides a collaborative dominant multimodal RGB-T fusion target tracking device based on dual-stream symmetric adapter bridging, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the collaborative dominant multimodal RGB-T fusion target tracking method based on dual-stream symmetric adapter bridging of Example 1.
[0104] Example 3
[0105] This embodiment provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the program is executed by a processor, the collaborative-dominant multimodal RGB-T fusion target tracking method based on dual-stream symmetric adapter bridging of Example 1 is implemented.
[0106] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0107] In the present invention, by obtaining a multimodal video image sequence to be processed and taking the input as the original data stream and the mask processing stream respectively, and passing them into their respective Transformer encoders, the two streams are jointly trained; a modal reciprocal adapter is introduced, which acts as a bridge connecting the interaction of the two streams and provides appropriate prompt information for the training process between the two streams; at the same time, a collaborative dominant multimodal fusion module is designed for multimodal fusion, which dynamically mines the more dominant modality in different tracking scenarios and weightedly fuses it with the auxiliary modality. Compared with the method of only taking RGB as the dominant modality for fusion and tracking, it better utilizes the advantages of the more dominant modal information in different scenarios. Ultimately, the multimodal features obtained through dual-stream adapter prompts and collaborative dominant fusion ensure the accuracy and robustness of tracking. This method provides reliable data support for various multimodal learning and visual target tracking tasks, thereby meeting the user's needs for high-quality visual tracking tasks and improving user satisfaction.
[0108] It should be understood that the processor in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0109] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0110] The above embodiments can be implemented in whole or in part through software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function according to the embodiments of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired method (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available media can be magnetic media (such as floppy disks, hard disks, tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0111] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.
[0112] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0113] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0114] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0115] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0116] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, and can be electrical, mechanical, or other forms.
[0117] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0118] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0119] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other media that can store program codes.
[0120] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
[0121] There are a few points to note:
[0122] (1) The drawings of the embodiments of the present invention only relate to the structures related to the embodiments of the present invention. Other structures may refer to conventional designs.
[0123] (2) For the sake of clarity, the thickness of layers or regions in the drawings used to describe the embodiments of the present invention are exaggerated or reduced, that is, these drawings are not drawn to scale. It is understood that when an element such as a layer, film, region, or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or intervening elements may be present.
[0124] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to form new embodiments.
[0125] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A collaborative-dominant multimodal RGB-T fusion target tracking method based on dual-stream symmetric adapter bridging, characterized by: The steps include: S1: Acquire a multimodal video image sequence. S2: A two-stream symmetric architecture is used, where the input data is passed as the original data and the masked data into their respective Transformer encoders, and the two streams are trained jointly. S3: Introduce several lightweight modality reciprocal adapters between the two streams to provide effective modality prompt information to each other during the training process of the two streams. S4: A collaborative dominant multimodal fusion module is used to perceive the advantages of the mode that plays a major role in tracking in different scenarios, dynamically use it as the dominant mode, and the other mode as the auxiliary mode. The two modes are weighted and fused, and work together to make the optimal decision on the utilization of different modes in different scenarios. S5: The fused multimodal features are fed into the prediction head to obtain the target's position in the next frame, resulting in the final tracking result. The predicted target position is compared with the groundtruth using a loss function consisting of the weighted sum of weightfocalloss, L1loss, and IoUloss for optimization training.
2. The collaborative-dominant multimodal RGB-T fusion target tracking method based on dual-stream symmetric adapter bridging according to claim 1 is characterized in that: The step S2 is specifically as follows: The two streams share parameters and conduct joint training, where the original data stream does not process the data in any way, and obtains X RGB 、X TIR ; The mask processing flow randomly masks 30% of the multimodal data and obtains The overall operation of multimodal features in the original data stream is as follows: H RGB =[CDMF(X RGB ,MRA(T));CDMF(Z RGB ,MRA(T))] (1) H TIR =[CDMF(X TIR ,MRA(R));CDMF(Z TIR ,MRA(R))] (2) Among them, X RGB 、Z RGB 、X TIR 、Z TIR denotes the search area and template token of visible light and thermal infrared, respectively. R and T denote the visible light and thermal infrared modal information from the opposing streams. H RGB 、H TIR represents the multimodal feature information obtained by splicing, MRA represents the modal reciprocal adapter proposed in the present invention, and CDMF represents the collaborative-dominant multimodal fusion module proposed in the present invention. Correspondingly, the overall operation of multimodal features in the mask processing flow is as follows: in, denotes the search area and template token of visible light and thermal infrared from the mask processing stream, R and T denote the visible light and thermal infrared modal information from the opposing stream, and H′ RGB , H′ TIR Represents the multimodal feature information obtained by splicing.
3. The collaborative-dominant multimodal RGB-T fusion target tracking method based on dual-stream symmetric adapter bridging according to claim 1 is characterized in that: The step S3 is specifically as follows: S301: When the search area token is used as input, a template token from another modality is specifically introduced as a hint token to provide target-related information about the search area. When the template token is used as input, no hint token information is required. Taking the search area token input as an example, the hint process is as follows: T′=X RGB +Z TIR (5) Among them, X RGB Represents the RGB search area, Z TIR represents the TIR template, and T′ represents the formal input of the modal reciprocal adapter. S302: The input token first undergoes a dimensionality reduction linear transformation to be compressed to a lower dimension, and then passes through the GELU activation function to introduce nonlinearity and enhance its feature representation. The process is as follows: T down =GELU(Dw(T′)) (6) Where Dw represents the dimensionality reduction linear transformation, T down Represents the modal token after dimensionality reduction transformation and GELU activation. S303: The features are further subjected to a linear transformation without changing their dimensions, followed by a dimensionality-increasing linear transformation to restore them to their original dimensions, and a residual connection with an initial token is attached to obtain the output of the modality reciprocal adapter, which is formulated as follows: T up =Up(Linear(T down ))+T′ (7) Among them, Linear represents a simple linear transformation operation, Up represents a dimension-raising linear transformation operation, and T ′ is the token information obtained in S301, T up represents the final output of the modal reciprocal adapter after the dimensionality-increasing linear transformation.
4. The collaborative-dominant multimodal RGB-T fusion target tracking method based on dual-stream symmetric adapter bridging according to claim 1 is characterized in that: The step S4 is specifically as follows: S401: The input multimodal token first undergoes GELU activation to suppress the position where the response drops below zero. Then, global average pooling is used to calculate and estimate the relative magnitude of the two modalities, thereby outputting two weighting factors γ1 and γ2. The specific formula is as follows: γ1=GAP(GELU(T R )) (8) γ2=GAP(GELU(T T )) (9) Among them, T R 、T T represents the input multimodal token, GAP represents the global average pooling operation, and γ1 and γ2 represent the two obtained weighting factors. S402: To obtain unbiased multimodal information, each mode is multiplied by the corresponding weighting factor to obtain a coordinated and balanced modal result. The formula is as follows: T w =T R *γ1+T T *γ2 (10) Among them, T w represents the coordinated and balanced modal results obtained by weighted product. S403: To prevent the S402 operation from accidentally amplifying the original mode, w Divide by the sum of γ1 and γ2 to make its amplitude close to the original data. At the same time, in order to further improve the obtained modal results, a trainable parameter β is introduced to normalize the relative influence of the two modes together with the layer normalization operation. The specific formula is as follows: Among them, LN is the layer normalization operation, T N Represents the final result of collaborative-dominant multimodal fusion.
5. The collaborative-dominant multimodal RGB-T fusion target tracking method based on dual-stream symmetric adapter bridging according to claim 1 is characterized in that: The step S5 is specifically as follows: Two loss functions are used to simultaneously supervise the training of the original data stream and the mask data stream. The weighted sum of the three loss functions, weight focalloss, L1 loss, and IoU loss, is used to compare the predicted target position with the ground truth for optimization training. The loss function calculation process is as follows: L total =L raw +θL processed (14) Among them, L cls represents the weighted focal loss of the training branch, L iou It is used to optimize the overlap between the predicted bounding box and the ground truth. L1 is the measure of the mean absolute difference between the predicted value and the true value. iou =2 and are two trade-off parameters, and θ is set to 0.3 as a scaling factor to balance the two flows.
6. A collaborative-dominant multimodal RGB-T fusion target tracking device based on dual-stream symmetric adapter bridging, characterized in that: The method comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the collaborative dominant multimodal RGB-T fusion target tracking method based on dual-stream symmetric adapter bridging according to any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that A computer program is stored thereon, characterized in that when the program is executed by a processor, it implements the collaborative dominant multimodal RGB-T fusion target tracking method based on dual-stream symmetric adapter bridging as described in any one of claims 1 to 5.