A Dense Prediction Multi-Task Learning Method Combining Mixture of Experts and Mamba Model
By introducing a hybrid expert mechanism and multi-directional scanning method into the parameter space of the Mamba model, the problem of insufficient task independence and global interaction capabilities in multi-task learning is solved, and multi-task collaborative optimization and single-task performance improvement are achieved, especially in the field of autonomous driving and embodied intelligence.
Patent Information
- Application Number
- CN202510645139.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The existing single-task learning methods are independent of each other in the field of computer vision and cannot be shared and represented, resulting in insufficient resource waste and generalization capabilities, and it is difficult to perform stably in complex environments. The multi-task learning methods have shortcomings in global interaction capabilities and task collaborative optimization.
The hybrid expert model is used to perform multi-task learning in the parameter space of the Mamba model. By introducing a hybrid expert mechanism and multi-directional scanning method, parameter interaction and global feature space perception between tasks are enhanced, and combined with task prior knowledge, the efficient fusion of dynamic modeling and task features is achieved.
It significantly improves the coordination ability and single-task performance between multitasking, and improves the model's adaptability and robustness to complex scenarios, especially in the fields of autonomous driving and embodied intelligence.
Smart Images

Figure CN120182791B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital image processing, and relates to a dense prediction multi-task learning method that combines a mixture of experts and a Mamba model. Background Art
[0002] In typical intelligent systems such as autonomous driving and embodied intelligence, environmental perception, as the first link in the perception-understanding-decision-making-execution loop, determines the modeling quality and response ability of the system to the external environment. Such systems need to identify key elements in the surrounding environment in real time and construct accurate and comprehensive semantic and geometric representations in a complex, changing, and even unknown real world, thus posing extremely high requirements on related computer vision algorithms.
[0003] As an important branch of artificial intelligence, computer vision has long adopted single-task learning (Single-Task Learning), that is, designing and training models separately for specific tasks such as image segmentation, depth estimation, and edge detection. This approach can achieve good performance in specific tasks, but there are obvious limitations: the models for each task are independent and cannot share representations, which not only causes waste of resources but also makes it difficult to fully explore the potential correlations between tasks, resulting in insufficient generalization ability and unstable performance when facing complex environments.
[0004] To overcome the above problems, multi-task learning (Multi-Task Learning, MTL) has gradually become a research hotspot in the field of computer vision. MTL optimizes multiple related tasks simultaneously in a unified framework, introduces a shared representation mechanism, enables different tasks to promote each other, and improves the overall performance and model robustness. This collaborative learning strategy not only improves computational efficiency but also enhances the model's adaptability to complex scenarios, especially suitable for application scenarios with extremely high requirements for environmental understanding such as autonomous driving and embodied intelligence. In these scenarios, dense prediction tasks such as object detection, semantic segmentation, and depth estimation are often highly correlated, and adopting the MTL framework helps to achieve more accurate, detailed, and robust environmental perception.
[0005] The advancements of Convolutional Neural Networks (CNNs) and Transformers have significantly enhanced the capabilities of multi-task learning. For example, Xu et al. introduced a new method in Pad-net: Multi-tasks guided prediction-and-distillation network for simultaneous depth estimation and scene parsing, which involves supervising a set of auxiliary tasks to generate predictions and then using these predictions as multi-modal inputs for the final task. Similarly, Simon et al.'s Mti-net: Multi-scale task interaction networks for multi-task learning explicitly models task interactions at multiple scales, effectively refining task information from low to high levels. However, due to its lack of global interaction capabilities, the performance of CNN-based methods is limited. Transformer-based methods have shown better results. For instance, Ye et al.'s Inverted pyramid multi-task transformer for dense scene understanding optimizes the utilization of multi-task information at a global scale through a multi-scale multi-task feature interaction module. However, due to the failure to decouple the shared information between tasks, there are still deficiencies in multi-task collaborative optimization. Additionally, recently, the Mixture of Experts (MoE) model in MTL captures different aspects of data from different perspectives by using a set of experts. Ye et al.'s Taskexpert: Dynamically assembling multi-task representations with memorial mixture-of-experts uses a set of multi-layer perceptron experts for dynamic task routing. This method uses a task-specific routing network to selectively fuse this modal data, promotes cross-task target information sharing, and can model complex task relationships. However, due to the lack of global modeling capabilities of the constructed experts, the MoE method still has limitations.
[0006] Therefore, in view of the deficiencies and challenges in the above research, the present invention innovatively introduces a mixture of experts into the parameter space of the Mamba model with global modeling capabilities, and utilizes the parameter interaction between multi-tasks to efficiently achieve dynamic modeling of tasks. At the same time, it can utilize the sequence modeling of Mamba to perceive the effective global feature space, effectively improving the performance of single tasks and the collaborative ability between multi-tasks. Summary of the Invention
[0007] The present invention proposes a strategy for constructing a dense prediction multi-task learning expert in the parameter space of Mamba using a mixture of experts model, while enhancing Mamba's ability to model 2D features. It can effectively improve the cooperation ability between multi-tasks and enhance the performance of single tasks at the same time.
[0008] The technical solution of the present invention:
[0009] A dense prediction multi-task learning method combining a mixture of experts and a Mamba model is as follows:
[0010] (1) Task primary feature extraction
[0011] First, use a pre-trained visual encoder to process the input image to obtain general multi-scale features :
[0012]
[0013] Then, in each scale, use a convolutional encoder for dense prediction tasks (such as semantic segmentation, edge detection, normal estimation, etc.) to obtain task primary features :
[0014]
[0015] where is the number of tasks, is the index of the task;
[0016] (2) Parameter-aware Mamba expert decoder
[0017] Regularize the task primary features to enhance the expression:
[0018]
[0019] where is the regularization operation;
[0020] Use depthwise separable convolution and activation layers to process local information:
[0021]
[0022] where is the depthwise separable convolution, is the activation layer;
[0023] After passing through a linear layer, map the channel dimension of the feature to the Mamba parameter space dimension to obtain the Mamba feature:
[0024]
[0025] Among them is a linear layer;
[0026] (2.1) Mixture Parameter Experts
[0027] During the calculation of the Mamba parameter space, the calculation processes of parameters B and C in the original Mamba model are both replaced with the calculation of a mixture task expert;
[0028] Use the convolutional network in the Mamba parameter space to obtain the multi-task expert prediction results:
[0029]
[0030] Among them, is the number of experts;
[0031] Then model the dependencies between different dense prediction tasks and different experts:
[0032]
[0033] Among them, is the global pooling operation, which compresses the spatial dimension to 1; The operation concatenates the features on the channels to obtain the feature ;
[0034] Subsequently, use the linear mapping layer to transform the channel dimension into the number of experts , and then use the activation layer and the function to normalize the channels to obtain the weights of the task-specific experts , and the process is as follows:
[0035]
[0036] Add noise to the expert weights during the training phase, and at the same time, by selecting the activated experts as the final activation weights, and the process is as follows:
[0037]
[0038] Among them, is the added noise, is the linear mapping weight, which maps the task features to a scalar as the weight of the noise;
[0039] Finally, use the weighted method to obtain the features of the mixture parameter experts:
[0040]
[0041] Among them, For the weight of the th expert in the task; for the parameters B and C in the Mamba model, obtain and ; the subscripts B and C represent two hybrid expert feature calculation paths;
[0042] (2.2) Parameter priors
[0043] Define a set of learnable spatially invariant prior parameters for different dense prediction tasks, denoted as , where is the dimension of the parameter space of the Mamba model; for the parameters and parameter add this prior parameter , and also expand the spatial dimension of to obtain :
[0044]
[0045] Among them, The operation changes the spatial dimension 1 to the length L of the input sequence of Mamba. For the parameters and parameter add the learnable prior parameters and , and obtain the final parameters and parameter through the following operations:
[0046]
[0047]
[0048] (2.3) Other parameter calculations
[0049] Calculate the parameters A, D, and parameter of Mamba using the parameter calculation method of the original Mamba model;
[0050] (2.4) Multi-directional Hilbert scan module
[0051] Assume that the scale size of the image feature is , and find the smallest positive integer such that and . At this time, is the order of the Hilbert space; the transformation of the two-dimensional feature space mapped to the one-dimensional Hilbert space is as follows:
[0052]
[0053] Among them, is the Hilbert space transformation formula of order, and by continuously rotating the serialized features 90 degrees around the feature center, three other transformation methods are obtained, denoted as:
[0054] ,
[0055] Among them, the operation is to rotate 90 degrees clockwise times of operations;
[0056] If or , then the transformation also needs to be cropped; the specific method is to retain the area in the lower left of the transformation that is the same size as the image;
[0057] By performing four Hilbert serialization transformations on the Mamba feature , the Mamba parameter simultaneously:
[0058]
[0059]
[0060]
[0061]
[0062] Then use the standard Mamba state space model to perform state space calculations:
[0063]
[0064] Among them, is the standard state model;
[0065] Finally, by summing the state space results in four directions and performing inverse Hilbert serialization, the operation result of the state space is obtained:
[0066]
[0067] Among them, is the operation of rotating counterclockwise degrees times, is the output of the task feature;
[0068] (2.5)Integrate Mamba output
[0069] Send to the input linear layer to restore the number of channels before entering the Mamba parameter space:
[0070]
[0071] Use the regularized task primary features to calculate the gating adjustment weights:
[0072]
[0073] Dynamically adjust the information flow using the gating adjustment weights, and at the same time use skip connections to transmit the task primary features to obtain the output of the parameter-aware Mamba expert decoder ; The process is as follows:
[0074]
[0075] For the outputs of the Mamba expert decoder at different scales , assign subscripts 1, 2, 3, 4 to represent the four scales from shallow to deep where the visual encoder output features are located, to obtain multi-task multi-scale features :
[0076]
[0077] (3) Task decoder
[0078] First, spatially align all scale features in the multi-task multi-scale features using the bilinear interpolation algorithm, and the interpolation resolution is of the resolution:
[0079]
[0080]
[0081]
[0082] Among them, is the bilinear interpolation algorithm for spatial interpolation alignment;
[0083] Then concatenate the aligned features on the channels to obtain :
[0084]
[0085] Use a linear layer to weight the scale features, use a convolutional decoder to decode and predict, and use the bilinear interpolation algorithm to restore the spatial scale to the size of the input image :
[0086]
[0087] For Obtain the final prediction result using the post - processing operation corresponding to the task:
[0088]
[0089] A conversion method from probability estimation result to label prediction result for task t;
[0090] (4)Multi - task joint optimization strategy
[0091] By combining task loss functions, jointly optimize the network structure using gradient backpropagation; for continuous regression tasks, namely depth estimation and surface normal estimation, use the L1 loss function; for discrete classification tasks, namely semantic segmentation, human parsing, saliency detection, and object boundary detection, use the cross - entropy loss function; the final loss is as follows:
[0092]
[0093] where, is the loss function for task t.
[0094] Advantages of the present invention:
[0095] (1)The present invention innovatively introduces an architecture combining Mamba and MoE mechanisms in the model decoder, specifically for enhancing the effect of multi - task intensive prediction.
[0096] (2)This module effectively uses MoE to balance the relationship between task - specific parameters and adopts the S4 model to capture long - range dependencies, thus significantly enhancing the selectivity and expressive ability of the hidden state.
[0097] (3)In addition, to strengthen the perception of different task characteristics, we incorporate task - specific prior knowledge. To solve the mismatch problem between the one - way modeling characteristic of Mamba and image data, we apply a multi - direction scanning method.
[0098] (4)Through extensive experiments on two challenging benchmark tests - NYUD - v2 and PASCAL - Context, the effectiveness of our design scheme is verified. Both quantitative analysis and qualitative evaluation results consistently show that the method we proposed is not only robust but also superior in performance.
[0099] (5) The present invention is of great significance in the fields of autonomous driving and embodied intelligence. In an autonomous driving system, the accurate perception and reasoning ability of the model for multi-tasks (such as road segmentation, obstacle detection, depth estimation, etc.) directly affect safety and reliability. This method effectively improves the generalization ability and real-time performance of the system in complex traffic scenarios by enhancing the model's selective expression and fusion ability for multi-task information. In embodied intelligence, the multi-modal perception and understanding of the environment by the intelligent agent are crucial, especially when performing tasks such as path planning and action decision-making. The multi-directional modeling strategy and task prior injection mechanism proposed by the present invention help to improve the modeling accuracy of the intelligent agent for the environmental state and task adaptability, providing strong support for the embodied intelligent agent to achieve more robust perception and operation capabilities in the real physical world. Description of the Drawings
[0100] Figure 1 It is the overall architecture diagram of the Mamba expert method for parameter perception related to this application.
[0101] Figure 2 It is the detailed diagram of the Mamba expert decoder for parameter perception related to this application.
[0102] Figure 3 It is the detailed diagram of the hybrid parameter expert related to this application.
[0103] Figure 4 It is the detailed diagram of the multi-directional Hilbert scan related to this application. Specific Embodiment
[0104] The following further illustrates the specific embodiment of the present invention in combination with the drawings and technical solutions.
[0105] Taking dense prediction tasks such as semantic segmentation, edge detection, and normal estimation as examples, the overall framework is as Figure 1 shown.
[0106] Step 1, Task Primary Feature Extraction, the specific steps are as follows:
[0107] The input image passes through a pre-trained visual encoder to extract multi-scale features common to the task. At different scales, convolutional modules for dense prediction tasks are respectively used for local feature decoding to obtain multi-task primary features.
[0108] Step 2, Parameter-Aware Mamba Expert Decoding, the specific steps are as follows:
[0109] As Figure 2 shown, the multi-task primary features are regularized to enhance the features. Depthwise separable convolutions and activation layers are used to process local information. After activation, the feature channel dimension is mapped to the Mamba parameter space dimension through a linear layer to obtain Mamba features. Subsequently, the following five steps are calculated in sequence:
[0110] (1)Mixed parameter expert calculation, the specific steps are as follows:
[0111] As Figure 3 shown, a multi-task representation of parameter B and parameter C is generated using a mixture-of-experts structure. Channel compression and global pooling are used to obtain the weighted weights of different experts. The final expert weights are obtained through Top-K selection and noise addition during training to fuse the expert outputs.
[0112] (2)Parameter prior calculation, the specific steps are as follows:
[0113] As Figure 1 shown, parameter B and parameter C also need to integrate the task parameter priors, define spatially invariant prior parameters for each task, align the feature spaces through dimensional expansion, and combine the outputs of the mixed parameter experts to obtain the final multi-task Mamba parameters B and C.
[0114] (3)Other parameter calculation, the specific steps are as follows:
[0115] Calculate the Mamba parameters A, D, and parameter .
[0116] (4)Perform multi-directional Hilbert scanning, the specific steps are as follows:
[0117] As Figure 4 shown, multi-directional Hilbert scanning is adopted to construct the Hilbert curve of the minimum covering feature map, generate 4 scanning directions through multiple 90-degree rotations, and crop the scans exceeding the feature size. Then generate the Mamba features, parameter B, parameter C, and parameter in these 4 directions. The state space calculation is performed using the SSM standard state equation of Mamba. Subsequently, the deserialization operations are respectively performed on the outputs of the 4 scanning directions and then added together to obtain the output of the multi-directional Hilbert scanning.
[0118] (5)Integrate the Mamba output, the specific steps are as follows:
[0119] As Figure 2 shown, the output after multi-directional Hilbert scanning is used with a linear layer to decouple from the Mamba parameter space and restore the channels of the task features before entering the Mamba parameter space. The gated adjustment weights are calculated using the regularized task primary features through the linear layer and the activation layer to dynamically adjust the information flow. The skip connections are used to transmit the task primary features. The outputs of different scales are collected to obtain the multi-scale features of the multi-task.
[0120] Step 3, task decoding, the specific steps are as follows:
[0121] As Figure 1As shown, the multi-scale features of multi-tasks need to go through a task decoder to obtain the final multi-task prediction. In the task decoder, bilinear interpolation is used to spatially align the multi-task multi-scale features, and then the features of different scales are concatenated in the channel dimension. The final task features are obtained by weighted summation through a linear layer. For prediction generation, a task-specific convolutional head is used to decode the features, and the interpolation outputs the probability estimation result. Finally, conventional post-processing operations are used to generate the multi-task prediction map.
[0122] Step 4, training optimization, the specific details are as follows:
[0123] In the training strategy, a combined task-specific loss function is used to jointly optimize the network structure through gradient backpropagation. For continuous regression tasks, namely depth estimation and surface normal estimation, the L1 loss function is adopted; while for discrete classification tasks, namely semantic segmentation, human parsing, saliency detection, and object boundary detection, the cross-entropy loss function is used. The final loss is as follows:
[0124]
[0125] where is the loss function for task t, is the final prediction map for task t.
[0126] Experimental verification:
[0127] (1) Verification details
[0128] For the experimental setup, two widely recognized benchmark datasets were selected: NYUD-v2 and PASCAL-Context. The NYUD-v2 dataset contains 795 training images and 654 test images, providing detailed annotations for various tasks such as semantic segmentation, monocular depth estimation, surface normals, and object boundary detection, offering a robust framework for evaluating the effectiveness of models. The PASCAL-Context dataset is larger, containing 4998 training images and 5105 test images, covering diverse annotations for various tasks from semantic segmentation to human parsing and saliency detection, and can be used to evaluate the scalability and robustness of MTL methods in complex and variable scenarios. In terms of implementation details, ViT (including ViT-B and ViT-L) was adopted as the backbone network, and ablation analysis was conducted based on ViT-B. The top-k value and the total number of experts in the parameter expert module were set to 9 and 15 respectively. The entire multi-task learning network was trained for 40,000 iterations on both datasets with a batch size of 6, and the same optimizer and loss function as specified in InvPT were used. The experiments were conducted using 2 Nvidia A100 80G GPUs. Additionally, for fair comparison, the task-averaged gain Δg was adopted to represent the overall performance of all tasks, and a unified multi-task loss was applied to supervise the learning process of the network.
[0129] (2)Quantitative Results
[0130] We conducted a fair comparison on two well-known public datasets using Vision Transformer as the backbone network, ensuring the same architecture and evaluation criteria as previous studies. Our method was benchmarked against methods such as InvPT, InvPT++, TaskPrompter, M3ViT, MQTransformer, and TaskExpert. As shown in Tables 1 and 2, our method achieved the highest average performance gain values on both datasets. Meanwhile, we also compared with methods using Swin Transformer, and our method outperformed in terms of average performance gain values. These findings demonstrate the effectiveness of our method and showcase the practical advantages of integrating state space models into multi-task learning.
[0131] Table 1: Comparison results of using two vision encoders, Swin-L, ViT-L, and ViT-B, with current state-of-the-art multi-task dense prediction methods on the PASCAL-Context dataset.
[0132]
[0133] As shown in Table 1, the method proposed in this study demonstrates comprehensive and remarkable performance advantages in the multi-task dense prediction task of the PASCAL-Context dataset. Through systematic comparative experiments, it can be found that when using ViT-L as the backbone network, our method achieves the optimal performance in four tasks: semantic segmentation (81.38 mIoU), human body part estimation (70.39 mIoU), normal estimation (13.36 mEu), and edge detection (75.30 odsF). Especially in the semantic segmentation and edge detection tasks, compared with the previous best method, the performance is improved by 1.16% and 1.10% respectively, and the overall performance gain reaches 3.60%, which is the most prominent among all comparison methods. It is worth noting that even on the ViT-B backbone network with lower computational resource requirements, our method still maintains an average performance gain of 1.86% and shows the best performance compared with other methods in two tasks: saliency detection (85.25 maxF) and normal estimation (13.40 mEu). This result indicates that our method can not only make full use of the advantages brought by large model capacity, but also its innovative architecture design can achieve efficient feature learning on smaller models. Compared with the latest methods such as InvRT++ and TaskPrompter, our improvement is more balanced across all tasks rather than only achieving improvement in specific tasks, which fully demonstrates that the method has better generalization ability.
[0134] Table 2: Comparison results of using two vision encoders, Swin-L, ViT-L and ViT-B respectively, with current advanced dense prediction methods on the NYUD-v2 dataset.
[0135]
[0136] As shown in Table 2, the method proposed in this study demonstrates excellent performance in the multi-task dense prediction task of the NYUD-v2 dataset. Through systematic comparative experiments, it can be found that when using VIT-L as the backbone network, our method achieves competitive results in four tasks: semantic segmentation (56.63 mIoU), depth estimation (0.5102 RMSE), normal estimation (18.52 mEu), and edge detection (78.10 odsE). Notably, in the semantic segmentation task, our method improves by 1.28 mIoU compared to the previous best TaskExpert method, and this significant progress fully reflects the advantage of the method in complex scene understanding. From the perspective of overall performance gain, our method achieves an average improvement of 5.23%, ranking first among all compared methods, which fully verifies the comprehensive performance advantage of the method. It is worth noting that even on the VIT-B backbone network with lower computational resource requirements, our method still maintains an average performance gain of 1.02% and reaches the optimal level in the normal estimation (18.90 mEu) and edge detection (78.00 odsE) tasks. This result indicates that our method can not only fully utilize the advantages brought by large model capacity, but also its innovative multi-task learning mechanism can achieve efficient feature representation on smaller models. Compared with the latest methods such as InvPT++ and TaskPromoter, our improvement is more comprehensive in all task dimensions, especially in the two challenging tasks of semantic segmentation and depth estimation, which fully proves that the method has better task adaptability and generalization ability.
Claims
1. A dense prediction multi-task learning method combining a mixture of experts and a Mamba model, characterized in that The details are as follows: (1) Task primary feature extraction; (2) Parameter-aware Mamba expert decoder; Regularize the task primary features to enhance the representation: N t = Norm(X t ) Among them, Norm is the regularization operation, and X t is the primary feature of the task; Use depthwise separable convolution and activation layers to process local information: M t = Act(DWConv(N t )) Among them, DWConv is the depthwise separable convolution, and Act is the activation layer; After passing the feature M through a linear layer t the channel dimension is mapped to the dimension of the Mamba parameter space to obtain the Mamba feature: x t = Linear(M t ) Among them, Linear is the linear layer; (2.1) Hybrid parameter expert During the calculation process of the Mamba parameter space, replace the calculation processes of parameter B and parameter C in the original Mamba model with the calculation of a hybrid task expert; Use the convolutional network in the Mamba parameter space to obtain the multi-task expert prediction results: Among them, I is the number of experts; Then model the dependencies between different dense prediction tasks and different experts: Among them, Pool is a global pooling operation that compresses the spatial dimension to 1; the Cat operation concatenates features on the channels to obtain features Subsequently, a linear mapping layer is used to transform the channel dimension into the number I of experts, and then an activation layer and the Softmax function are used to normalize the channels, obtaining the weights of the experts for a specific task. The process is as follows: Add noise to the expert weights during the training phase, and at the same time select the TopK activated experts as the final activation weights. The process is as follows: where N(0,1) is the added noise, and W noise is the weight of the linear mapping, which maps the task feature to a scalar as the weight of the noise; Finally, use a weighted method to obtain the features of the hybrid parameter expert: Among them, is the weight of the i-th expert for task t; for parameters B and C in the Mamba model, we obtain and The subscripts B and C represent two hybrid expert feature calculation paths; (2.2) Parameter prior Define a set of learnable spatially invariant prior parameters for different dense prediction tasks, denoted as where q is the dimension of the Mamba model parameter space; add this prior parameter P to the parameters B' and C' of the Mamba model t , and also expand the spatial dimension of P t to obtain P ′t = Expand(P t ) Among them, the Expand operation changes the spatial dimension 1 to the length L of the input sequence of Mamba, and adds learnable prior parameters to parameter B' and parameter C'. and Obtain the final parameter B″ and parameter C″ through the following operations: (2.3) Other parameter calculations Use the parameter calculation method of the original Mamba model to calculate the parameters A, D, and Δ of Mamba; (2.4) Multi-directional Hilbert scan module Let the scale size of the image feature be H×W, and find the smallest positive integer k such that H ≤ 2 k and W ≤ 2 k , where k is the order of the Hilbert space; the transformation of mapping the two-dimensional feature space to the one-dimensional Hilbert space is as follows: h1 = H k (x, y) Among them, H k is the Hilbert space transformation formula of order k. By continuously rotating the serialized features 90 degrees around the feature center, three other transformation methods are obtained, denoted as: h d+1 = rot d (h1), d ∈ {1, 2, 3} where rot d is an operation of rotating 90 degrees clockwise d times; If H < 2 k or W < 2 k , then the transformation also needs to be cropped; the specific method is to retain the area in the lower left of the transformation that is the same size as the image; By performing four Hilbert serialization transformations on the Mamba feature x t , the Mamba parameters B, C, and Δ are simultaneously transformed Then use the standard Mamba state space model for state space calculation: Among them, SSM is the standard state model; Finally, obtain the operation result of the state space by summing the state space results in four directions and performing inverse Hilbert serialization: Among them, is the operation of rotating counterclockwise by 90 degrees for d - 1 times, out t is the output of the task feature; (2.5) Integrate Mamba output Take out t Input into the linear layer to restore the number of channels before entering the Mamba parameter space: O t = Linear(out t ) Use the regularized task primary feature N t to calculate the gating adjustment weight: w t = Act(Linear(N t )) Dynamically adjust the information flow using gated adjustment weights, and at the same time use skip connections to transmit the primary features of the task to obtain the output E of the parameter-aware Mamba expert decoder t ; The process is as follows: E t = w t · O t + X t The output E of the Mamba expert decoder at different scales t , is assigned subscripts 1, 2, 3, 4 to represent the four scales from shallow to deep where the output features of the visual encoder are located, obtaining the multi-task multi-scale features H t : (3) Task decoder; First, spatially align all scale features in the multi-task multi-scale feature H t using the bilinear interpolation algorithm. The interpolation resolution is the resolution of: Among them, Bilinear is the bilinear interpolation algorithm for spatial interpolation alignment; Next, the aligned features are concatenated on the channels to obtain K t : Use the linear layer to perform weighted fusion on features at each scale, use the convolutional decoder for prediction decoding, and use the bilinear interpolation algorithm to restore the spatial scale to the size of the input image img: Result t = Bilinear(Conv(Linear(K t ))) For Result t Obtain the final prediction result using the post - processing operation corresponding to the task: prediction t = PostProcess(Result t ) Among them, PostProcess t is a conversion method from the probability estimation result to the label prediction result for task t; (4) Multi-task joint optimization strategy.
2. The dense prediction multi-task learning method combining a hybrid expert and a Mamba model according to claim 1, characterized in that The specific implementation process of step (1) task primary feature extraction is as follows: First, use the pre-trained visual encoder Encoder to process the input image img to obtain general multi-scale features s: s = Encoder(img) Next, a convolutional encoder for dense prediction tasks is used in each scale to obtain the task primary feature X t : X t = Conv t (s), where \(t\in\{1,2,\ldots,T\}\) Among them, T is the number of tasks, and t is the task index.
3. The dense prediction multi-task learning method combining a hybrid expert and a Mamba model according to claim 2, characterized in that The specific implementation process of step (4) multi-task joint optimization strategy is as follows: By combining task loss functions, jointly optimize the network structure using gradient backpropagation; for continuous regression tasks, namely depth estimation and surface normal estimation, use the L1 loss function; for discrete classification tasks, namely semantic segmentation, human parsing, saliency detection, and object boundary detection, use the cross-entropy loss function; The final loss is as follows: Among them, L t (·) is the loss function of task t.
Citation Information
Patent Citations
Prostate MRI image segmentation method based on Mamba-Unet
CN119579627A
Pathological image classification system and method based on Mama and hybrid expert model
CN119851017A