Parameter efficient multi-modal tracking method based on sparse-dense hybrid expert

By using the SDMoE module of sparse-density hybrid expert in the multimodal tracking method, the multimodal tracking performance limitation problem caused by the difference between modes is solved, and efficient and robust multimodal tracking effect is achieved.

CN120198461APending Publication Date: 2025-06-24ANHUI UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510270712.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Due to the differences between modalities, the existing high-efficiency fine-tuning tracking method is difficult to effectively process tracking data of multiple modalities with a unified model, which limits the multimodal tracking performance.

Method used

Using a parameter efficient multimodal tracking method based on sparse-dense mixing experts, the SDMoE module is embedded in the Transformer encoder branches of RGB mode and X mode, as a modal-specific adapter and multimodal adapter, effectively modeling the unique and shared information of different modes.

Benefits of technology

It realizes efficient multimodal tracking with parameters in an end-to-end training network, effectively reducing the demand for computing resources and parameter storage, and improving the performance and robustness of multimodal tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198461A_ABST
    Figure CN120198461A_ABST
Patent Text Reader

Abstract

The invention discloses a parameter efficient multi-modal tracking method based on sparse-dense hybrid experts, belongs to the technical field of deep learning, and solves the problem that the existing method is difficult to process multi-modal tracking data by using a unified model. According to the invention, a sparse-dense hybrid expert module is used as a modal specific adapter and a multi-modal adapter to be embedded into a frozen tracking network based on RGB; the SDMoE module can be used for effectively modeling specific and shared information of different specific modals or fusion modals, and a large number of computing resources and parameter storage are not needed; an SDMoE module is used as a modal specific adapter and a multi-modal adapter to be embedded into each modal specific branch and each fusion modal branch, and modeling is carried out on specific and shared information of different specific modals and fusion modals; according to the method, under the condition that the model capacity is remarkably expanded, the calculation workload is not greatly increased, the too high calculation and storage cost of multiple parallel experts is effectively reduced, and the challenge that a small number of shared experts are difficult to fully utilize the shared characteristics among modals is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning and relates to a parameter-efficient multi-modal tracking method based on sparse-dense hybrid experts. Background Art

[0002] Object tracking, as a fundamental task in the field of computer vision, plays an important role in application scenarios such as autonomous driving and visual surveillance. Although significant progress has been made in object tracking technology based on visible light (referred to as RGB here) images, it still faces many challenges in complex environments such as low light, fast motion, and occlusion. To address these problems, researchers have proposed various multi-modal fusion tracking methods, such as technical solutions that combine visible light with thermal infrared (RGB-T), events (RGB-E), and depth data (RGB-D). Given the small scale of existing multi-modal datasets, most methods use the RGB tracking model as a pre-trained model and fine-tune it using limited multi-modal data on this basis. However, these methods have several problems: firstly, the full model fine-tuning process is time-consuming; secondly, it has high requirements for GPU memory; thirdly, it increases the space cost required to store parameters; finally, it may also lead to the model having "catastrophic forgetting", that is, the newly learned knowledge covers the previously mastered information.

[0003] Furthermore, to address the above problems, researchers have explored some parameter-efficient fine-tuning strategies. For example, Cao et al., Bidirectional Adapter for Multimodal Tracking. 《the AAAI Conference on Artificial Intelligence》. 2024:927-935., and Zhu et al., Visual prompt multi-modal tracking. 《IEEE / CVF Conference on Computer Vision and Pattern Recognition》. 2023:9516-9526. adjusted the pre-trained RGB model to adapt to multi-modal data by using adapter or prompt techniques. However, these methods usually require training a separate model for each pair of modality combinations, which reduces the overall training efficiency. Wu et al., Single-model and Any-modality for Video Object Tracking. 《IEEE / CVF Conference on Computer Vision and Pattern Recognition》. 2024:19156-19166. proposed a unified method to attempt to solve this problem by sharing modality parameters. However, due to the significant differences between different modalities, it is difficult for this method to effectively process multi-modal data, thus limiting its performance. Summary of the Invention

[0004] The technical solution of the present invention is used to solve the problem that due to the significant inter-modal differences, the existing parameter-efficient fine-tuning tracking methods are difficult to effectively process the tracking data of multiple modalities with a unified model, which limits the multi-modal tracking performance.

[0005] The present invention solves the above technical problems through the following technical solutions:

[0006] The present invention provides a parameter-efficient multi-modal tracking method based on sparse-dense hybrid experts, including:

[0007] S1 Given the search regions and template images of the RGB modality and the X modality;

[0008] S2 Use the embedding layer to segment the input image into blocks and flatten them into one-dimensional tokens;

[0009] S3 Use a two-branch ViT-Base backbone network to extract the features of the RGB modality and the X modality;

[0010] S4 embeds the SDMoE module into each Transformer encoder branch of both the RGB modality and the X modality dual branches simultaneously;

[0011] After completing the feature extraction of the RGB modality and X modality branches, the obtained features are fused together and then sent to the prediction head to predict the position and bounding box of the target, thereby completing target tracking.

[0012] Further, the method of embedding the SDMoE module into each Transformer encoder branch of both the RGB modality and the X modality dual branches is as follows:

[0013] The SDMoE module is embedded after the multi-head attention of each Transformer encoder. Here, the embedded SDMoE module is named the modality-specific adapter; among them, the embedded MSA in the RGB modality branch aims to enhance the feature representation in different visible light image domains; the modality-specific adapter is also embedded in the X modality branch, aiming to simulate the shared and specific features of different specific X modalities;

[0014] The SDMoE module is also embedded after the fusion modality, called the multi-modal adapter; that is, first, the enhanced features are extracted by the multi-layer perceptron in each Transformer encoder of the two modality branches, then these features are added and fused together, and finally input into the multi-modal adapter.

[0015] Further, the SDMoE module consists of a sparse mixture of experts and a lightweight dense shared mixture of experts; the sparse mixture of experts is used to model the specific and shared information in the modality; the specific and shared information for modeling the fusion modality or specific modality.

[0016] Further, the sparse mixture of experts includes a router and multiple specific experts; the router consists of a fully connected layer and a Softmax function, and is used to select specific experts to participate in the modeling of modality-specific information; each specific expert consists of a feed-forward down-projection layer, a feed-forward up-projection layer, a gating layer, and an activation function. The feature dimension of the input to the feed-forward down-projection layer is [B, H, D]. After being processed by the feed-forward down-projection layer, the feature dimension of the token becomes [B, H, D / G]. The network structure of the gating layer is the same as that of the feed-forward down-projection layer, and the activation function uses SiLU. The final modality-shared feature is obtained through the feed-forward up-projection layer.

[0017] Further, the formula of the Softmax function is as follows:

[0018] s t,n = Softmax(FFN(Token t , N))

[0019] Among them, s t,n represents the score value of the t-th token corresponding to the n-th expert, Token t represents the r-th token, N is the total number of specific experts, and FFN represents the fully connected layer.

[0020] Furthermore, the balanced loss involved in the Softmax function is calculated as follows:

[0021]

[0022] Among them, T is the total number of tokens, K is the number of experts selected for each forward propagation of the network, 1(Token tselects EXpert n) represents the indicator function, which is a boolean judgment condition, that is, if the token t is assigned to the expert n, the value of the indicator function is 1, otherwise it is 0; L eb represents the balanced loss, f n represents the probability value of the n-th expert, P n represents the average score of the T tokens of the n-th expert, s t,n represents the score value of the t-th token (Token) corresponding to the n-th expert.

[0023] Furthermore, the lightweight dense shared mixture of experts consists of a feed-forward down-projection layer, a router, a parallel shared multi-expert network, and a feed-forward up-projection layer; single-modal or multi-modal tokens are sent to the feed-forward down-projection layer for feature dimensionality reduction; the dimension of the input features is reduced from [B, H, D] to [B, H, D / G]; the dimension-reduced features will be sent to the router and the parallel shared multi-expert network to extract the shared features of the modality; the multiple extracted shared features will be weighted and fused, and the final modality-shared features are obtained through the feed-forward up-projection layer.

[0024] Furthermore, each parallel shared expert in the parallel shared multi-expert network consists of only one fully connected layer.

[0025] The present invention also provides an electronic device, including a memory and a processor, the memory is used to store a program for supporting the processor to execute the above-mentioned parameter-efficient multi-modal tracking method based on sparse-dense mixture of experts, and the processor is configured to execute the program stored in the memory.

[0026] The present invention also provides a storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of the above-mentioned parameter-efficient multi-modal tracking method based on sparse-dense mixture of experts.

[0027] The advantages of the present invention are:

[0028] The present invention proposes a new hybrid expert fine-tuning framework that can achieve parameter-efficient multi-modal tracking in an end-to-end training network. Specifically, the present invention designs a Sparse-Dense Mixture of Experts (SDMoE) module and embeds it as a Modality-Specific Adapter (MSA) and a Multi-Modal Adapter (MMA) into an RGB-based frozen tracking network; the SDMoE module can effectively model specific and shared information of different specific modalities or fused modalities without requiring a large amount of computing resources and parameter storage. The SDMoE module is embedded as a modality-specific adapter (MSA) into each modality-specific branch to model specific and shared information of different specific modalities. In addition, the present invention also embeds the SDMoE module as a multi-modal adapter (MMA) into the fused modality branch to model specific and shared information of different fused modalities. The SDMoE consists of a sparse MoE and a lightweight dense shared MoE. The sparse MoE contains N experts and is responsible for modeling modality-specific information between different specific or fused modalities. The sparse MoE significantly expands the model capacity without significantly increasing the computational workload because only K (K < N) specific experts are selected for calculation during training and testing. The lightweight dense shared MoE is used to model shared information between modalities. Different from the previous mixture of experts method using a serial-parallel structure, the design of this sub-module adopts a parallel-serial structure. This architecture effectively reduces the excessive computational and storage costs of multiple parallel experts and overcomes the challenge that it is difficult for a small number of shared experts to fully utilize the shared features between modalities. All experts pass through a shared serial down-projection layer and up-projection layer and pass through their respective specific sub-networks in parallel between the two projection layers. In the implementation of the present invention, there are M = 4 parallel sub-networks. It should be noted that the input of the parallel sub-networks is the feature after dimensionality reduction of the down-projection layer, and each sub-network consists of only one fully connected layer, so its parameter quantity and computational quantity will not increase significantly. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 FIG. is an overall network framework diagram of the parameter-efficient multi-modal tracking method based on sparse-dense mixture of experts of the present invention;

[0030] Figure 2 FIG. is a framework diagram of the sparse-dense mixture of experts module of the parameter-efficient multi-modal tracking method based on sparse-dense mixture of experts of the present invention;

[0031] Figure 3 FIG. is a comparison of visual tracking results of the parameter-efficient multi-modal tracking method based on sparse-dense mixture of experts of the present invention. Detailed implementation manners

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings of the specification and specific embodiments:

[0034] Embodiment 1

[0035] As Figure 1 shown, the parameter-efficient multi-modal tracking method based on sparse-dense mixture of experts in the embodiment of the present invention includes:

[0036] Input: Given a search region and a template image in RGB modality and X modality (such as thermal, depth, and event).

[0037] Image patch partitioning: Use an embedding layer to segment the input image into patches and flatten them into one-dimensional tokens.

[0038] Feature extraction: Use a two-branch ViT-Base backbone network to extract features of the RGB modality and the X modality (including modalities such as thermal infrared, depth, and event). Here, the backbone network of each branch consists of 12 Transformer encoders, and each Transformer encoder consists of a normalization layer, a multi-head attention layer, a multi-layer perceptron, etc.

[0039] SDMoE module embedding: To effectively represent the features of multiple modalities, the embodiment of the present invention proposes an SDMoE module (sparse-dense mixture of experts module) and embeds it into the two-branch Transformer encoder. The specific operation steps are as follows:

[0040] (1) Embed the SDMoE module into each Transformer encoder branch of both the RGB modality and the X modality dual branches simultaneously. Specifically, this module is embedded after the multi-head attention of each Transformer encoder. Note that the embedded SDMoE module here is named the modality-specific adapter (MSA). Among them, the embedded MSA in the RGB modality branch aims to enhance the feature representation in different visible light image domains. It should be emphasized here that although ViT is trained on RGB datasets, there are significant domain differences in RGB data in cross-device collected datasets. Therefore, it is crucial to use MSA to model the shared and unique features of different RGB image domains. In addition, MSA is also embedded in the X modality branch to simulate the shared and specific features of different specific X modalities.

[0041] (2) Additionally, as Figure 1 shown, the SDMoE module is also embedded after the fusion modality, called the multi-modal adapter (MMA). Specifically, first use the multi-layer perceptron in each Transformer encoder of the two modality branches to extract the enhanced features, then add and fuse these features together, and finally input them into the MMA. This MMA can effectively enhance the shared and unique features of different fusion modalities.

[0042] As Figure 2 shown, the specific composition and operation steps of the SDMoE module are as follows:

[0043] The SDMoE module consists of a sparse mixture of experts and a lightweight dense shared mixture of experts.

[0044] 1) Sparse mixture of experts

[0045] As Figure 2 shown on the right side, the sparse mixture of experts contains a router and N specific experts for modeling specific information in the modality. The router consists of a fully connected layer and a Softmax function, which is used to select specific experts to participate in the modality-specific information modeling. The specific formula is as follows:

[0046] s t,n = Softmax(FFN(Token t , N))

[0047] where s t,n represents the score value of the t-th token corresponding to the n-th expert, and Token tDenote the t-th token, N is the total number of specific experts, and FFN represents the fully connected layer. For MSA and MMA, the values of N are 4 and 8 respectively. The routing strategy of adaptive learning may have the risk of routing collapse, that is, the model always only selects a few experts, resulting in other experts not being fully trained. To solve this problem, the present invention uses an expert-level balance loss. The calculation of the balance loss is as follows:

[0048]

[0049] where T is the total number of tokens, and K is the number of experts selected for each forward propagation of the network. In the present invention, for MMA and MSA, the values of K are 2 and 1 respectively. 1(Token t selects EXpert n) represents an indicator function, which is a boolean judgment condition, that is, if token t is assigned to expert n, the value of the indicator function is 1, otherwise it is 0; L eb denotes the balance loss, f n denotes the probability value of the n-th expert, P n denotes the average score of the n-th expert for T tokens, s t,n denotes the score value of the t-th token corresponding to the n-th expert.

[0050] As Figure 2 shown on the far right, each specific expert consists of a feed-forward down-projection layer, a feed-forward up-projection layer, a gating layer, and an activation function. This feed-forward bottom-up projection structure significantly reduces the computational overhead. The feature dimension input to the feed-forward down-projection layer here is [B, H, D]. After being processed by this feed-forward down-projection layer, the feature dimension of the token becomes [B, H, D / G], where G takes the value of 12. The network structure of the gating layer and the feed-forward down-projection layer is the same, and the activation function used here is SiLU. The final modality-shared feature is obtained through the feed-forward up-projection layer. It should be emphasized that sparse MoE significantly expands the model capacity, enabling it to well model the specific information of different modalities. However, its computational workload does not increase significantly because only some specific experts are selected to participate in the calculation during the training iteration and testing.

[0051] 2) Lightweight Dense Shared MoE

[0052] As Figure 2As shown on the left, the lightweight dense shared MoE consists of a feed-forward down-projection layer, a router, a parallel shared multi-expert network, and a feed-forward up-projection layer, which is used to model the shared information of the fusion modality or specific modalities. Specifically, single-modal or multi-modal tokens are sent to the feed-forward down-projection layer for feature dimensionality reduction. Similar to specific experts, the dimension of the input features will be reduced from [B, H, D] to [B, H, D / G]. The dimension-reduced features will be sent to the router and the parallel shared multi-expert network to extract the shared features of the modality. The multiple extracted shared features will be weighted and fused, and the final modality-shared features will be obtained through the up-projection layer. There are M parallel sub-shared expert networks, and in the setting of the present invention, M = 4. Each parallel shared expert consists of only one fully connected layer.

[0053] Different from the previous multi-shared experts using a parallel and serial structure, the lightweight dense shared MoE proposed in the present invention adopts a serial and parallel structure design. This design not only effectively solves the problems of excessive consumption of computing resources and excessive number of parameters caused by multiple parallel experts, but also overcomes the challenge that it is difficult for a small number of shared experts to fully utilize the shared information between modalities. Here, the parallel and serial structure of multiple shared experts means that multiple experts are parallel, and the network structure inside each expert is serial; while the serial and parallel structure of the present invention means that all experts pass through the shared serial feed-forward down-projection layer and feed-forward up-projection layer, and the parallel part is the only sub-network of each expert. It should be noted that due to its low-dimensional features and fewer parallel parameters, this serial-parallel structure will only require less computing and storage resources.

[0054] Target location and bounding box prediction: After completing the feature extraction of the RGB modality and X modality branches, the obtained features are fused together and then sent to the prediction head for predicting the location and bounding box of the target, thereby completing target tracking.

[0055] Comparison and verification

[0056] The method of the present invention is compared with advanced multi-modal methods on multiple RGB-D, RGB-T, and RGB-E tracking datasets to verify the effectiveness of the method of the present invention. The datasets evaluated include DepthTrack, VOT-RGBD2022, RGBT234, LasHeR, VTUAV, VisEvent, and COESOT.

[0057] Table 1 Comparison of tracking results on the DepthTrack dataset (%)

[0058]

[0059] Table 2 Comparison of tracking results on the VOT-RGBD2022 dataset (%)

[0060]

[0061] Table 3 Comparison of tracking results on RGB-T234 and LasHeR datasets (%)

[0062]

[0063] Table 4 Comparison of tracking results on VTUAV dataset (%)

[0064]

[0065] Table 5 Comparison of tracking results on VisEvent dataset (%)

[0066]

[0067] Table 6 Comparison of tracking results on COESOT dataset (%)

[0068]

[0069] As Figure 3 shown, in order to visually compare the tracking performance of the method of the present invention with that of other trackers, a qualitative comparison was made between the method of the present invention and two trackers with unified parameters on two representative multi-modal video sequences. Specifically, the present invention presented an RGB-T and an RGB-E video sequence, and these sequences contained four typical challenging attributes, namely fast motion, similar objects, low illumination, and background clutter. As Figure 3 shown, in complex challenging scenarios, the method of the present invention is more robust than other algorithms. For example, in the RGB-E scenario ( Figure 3 upper part), it has challenges such as fast motion and similar targets. Compared with other algorithms, the method of the present invention can still track the target well. Another example, as Figure 3 the lower half shows, the low illumination challenge in the RGB-T sequence causes the visible light modality to be unable to effectively capture the target object, and the video sequence is accompanied by background clutter and other challenges. In such a difficult challenging scenario, the method of the present invention can still track the target well, which indicates that the method of the present invention can truly exploit the complementary features between modalities to achieve more robust multi-modal tracking.

[0070] Embodiment 2

[0071] An electronic device, comprising a memory and a processor, wherein the memory is used to store a program for supporting the processor to execute the parameter-efficient multi-modal tracking method based on sparse-dense hybrid experts in Embodiment 1, and the processor is configured to execute the program stored in the memory.

[0072] Embodiment 3

[0073] A storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the parameter-efficient multi-modal tracking method based on sparse-dense hybrid experts in the first embodiment.

[0074] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A parameter-efficient multimodal tracking method based on sparse-dense hybrid experts, characterized in that: include: S1 Given the search area and template image in RGB mode and X mode; S2 uses an embedding layer to split the input image into blocks and flatten them into one-dimensional tokens; S3 uses a dual-branch ViT-Base backbone network to extract features of RGB modality and X modality; S4 embeds the SDMoE module into each Transformer encoder branch of the RGB modality and X modality dual branches simultaneously; After S5 completes the feature extraction of the RGB modality and X modality branches, the obtained features are fused together and then sent to the prediction head to predict the position and bounding box of the target, thereby completing the target tracking.

2. The parameter-efficient multimodal tracking method based on sparse-dense hybrid experts according to claim 1, characterized in that: The method of simultaneously embedding the SDMoE module into each Transformer encoder branch of the RGB modality and X modality dual branches is as follows: After embedding the SDMoE module into the multi-head attention of each Transformer encoder, the embedded SDMoE module here is named as modality-specific adapter; The embedded MSA in the RGB modality branch aims to enhance the feature representation in different visible light image domains; the modality-specific adapter is also embedded in the X-modality branch to simulate the shared and specific features of different X-modalities; The SDMoE module is also embedded after the fusion modality, called the multi-modal adapter; that is, the enhanced features are first extracted using the multi-layer perceptrons in each Transformer encoder of the two modal branches, and then these features are added and fused together and finally input into the multi-modal adapter.

3. The parameter-efficient multimodal tracking method based on sparse-dense hybrid experts according to claim 1, characterized in that: The SDMoE module consists of a sparse mixture expert and a lightweight dense shared mixture expert; the sparse mixture expert is used to model the specific and shared information in the modalities; the is used to model the specific and shared information of the fused modality or the specific modality.

4. The parameter-efficient multimodal tracking method based on sparse-dense hybrid experts according to claim 3, characterized in that: The sparse hybrid expert includes a router and multiple specific experts; the router is composed of a fully connected layer and a Softmax function, which is used to select specific experts to participate in the modeling of modal specific information; each of the specific experts is composed of a feedforward down-projection layer, a feedforward up-projection layer, a gating layer and an activation function. The feature dimension input to the feedforward down-projection layer is [B, H, D]. After being processed by the feedforward down-projection layer, the feature dimension of the token becomes [B, H, D / G]. The network structure of the gating layer is consistent with that of the feedforward down-projection layer. The activation function uses SiLU, and the final modal shared features are obtained through the feedforward up-projection layer.

5. The parameter-efficient multimodal tracking method based on sparse-dense hybrid experts according to claim 4, characterized in that: The formula of the Softmax function is as follows: s t,n =Softmax(FFN(Token t ,OF)) Among them, s t,n Indicates the score value of the tth token corresponding to the nth expert, Token t represents the t-th token, N is the total number of specific experts, and FFN stands for fully connected layer.

6. The parameter-efficient multimodal tracking method based on sparse-dense hybrid experts according to claim 5, characterized in that: The balance loss involved in the Softmax function is calculated as follows: Where T is the total number of tokens, K is the number of experts selected for each forward propagation of the network, 1 (Token t selects EXpert n) represents the indicator function, which is a Boolean judgment condition, that is, if token t is assigned to expert n, the value of the indicator function is 1, otherwise it is 0; L eb represents the balance loss, f n represents the probability value of the nth expert, P n represents the average score of T tokens of the nth expert, s t,n Indicates the score value of the tth token corresponding to the nth expert.

7. The parameter-efficient multimodal tracking method based on sparse-dense hybrid experts according to claim 3, characterized in that: The lightweight dense shared hybrid expert consists of a feed-forward down-projection layer, a router, a parallel shared multi-expert network, and a feed-forward up-projection layer; the unimodal or multimodal tokens are sent to the feed-forward down-projection layer for feature dimensionality reduction; The dimension of the input features is reduced from [B, H, D] to [B, H, D / G]; the reduced features will be sent to the router and the parallel shared multi-expert network to extract the shared features of the modalities; The extracted multiple shared features will be weighted fused, and the final modality shared features are obtained through the feedforward up-projection layer.

8. The parameter-efficient multimodal tracking method based on sparse-dense hybrid experts according to claim 7, characterized in that: Each parallel shared expert in the parallel shared multi-expert network consists of only one fully connected layer.

9. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the parameter-efficient multimodal tracking method based on sparse-dense hybrid experts as described in any one of claims 1 to 8, and the processor is configured to execute the program stored in the memory.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the parameter-efficient multimodal tracking method based on sparse-dense hybrid experts according to any one of claims 1 to 8 are performed.

Citation Information

Cited By

  • Training method and device of hybrid expert model, equipment and medium

    CN120911554A

  • An end-to-end automatic driving method and system based on multi-task expert fine-tuning

    CN122530974A