Intensive prediction multi-task learning method combining mixed expert and Mangbar model

By introducing a hybrid expert model into the parameter space of the Mamba model, the global interaction in multi-task learning and the decoupling of shared information between tasks is solved, and the problem of insufficient generalization ability in multi-task learning in the existing technology is significantly improved.

CN120182791AActive Publication Date: 2025-06-20DALIAN UNIV OF TECH

Patent Information

Application Number
CN202510645139.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-20
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

Existing computer vision models have limitations in multi-task learning, especially the lack of global interaction capabilities and the decoupling of shared information between tasks, resulting in insufficient generalization capabilities in complex environments.

Method used

By introducing a hybrid expert model to perform parameter interaction in the parameter space of the Mamba model, combining Mamba's sequence modeling to perform global feature space perception, dynamic modeling of tasks and multi-task collaborative optimization.

Benefits of technology

It significantly improves the coordination ability and single-task performance between multi-tasks, and improves the model's adaptability and robustness to complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182791A_ABST
    Figure CN120182791A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing methods, and discloses an intensive prediction multitask learning method combining a hybrid expert and a Mangbar model, which is a new framework based on a decoder, namely a parameter perception Mangbar model, and is used for designing intensive prediction in a multitask learning environment. The objective of the invention is to enhance connectivity between tasks by using rich and extensible parameters of a state space model. The method is provided with double-state space parameter experts, parameter priori specific to tasks can be integrated and set, and internal attributes of each task are captured. According to the method, accurate multi-task interaction is promoted, and task prior global integration is allowed to be achieved through a structured state space sequence model. In addition, a multi-direction Hilbert scanning method is adopted to construct a multi-angle feature sequence, so that the perception ability of a sequence model to 2D data is enhanced. Experiments on a common public reference verify the effectiveness of the proposed method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital image processing, and relates to a dense prediction multi-task learning method that combines a mixture of experts and a Mamba model. Background Art

[0002] In typical intelligent systems such as autonomous driving and embodied intelligence, environmental perception, as the first link in the perception-understanding-decision-making-execution loop, determines the modeling quality and response ability of the system to the external environment. Such systems need to identify key elements in the surrounding area in real time and construct accurate and comprehensive semantic and geometric representations in a complex, changing, and even unknown real world, thus posing extremely high requirements on related computer vision algorithms.

[0003] As an important branch of artificial intelligence, computer vision has long adopted single-task learning (Single-Task Learning), that is, designing and training models separately for specific tasks such as image segmentation, depth estimation, and edge detection. This approach can achieve good performance in specific tasks, but there are obvious limitations: the models for each task are independent and cannot share representations, which not only causes waste of resources but also makes it difficult to fully explore the potential correlations between tasks, resulting in insufficient generalization ability and unstable performance when facing complex environments.

[0004] To overcome the above problems, multi-task learning (Multi-Task Learning, MTL) has gradually become a research hotspot in the field of computer vision. MTL optimizes multiple related tasks simultaneously in a unified framework, introduces a shared representation mechanism, enables different tasks to promote each other, and improves the overall performance and model robustness. This collaborative learning strategy not only improves computational efficiency but also enhances the model's adaptability to complex scenarios, especially suitable for application scenarios with extremely high requirements for environmental understanding such as autonomous driving and embodied intelligence. In these scenarios, dense prediction tasks such as object detection, semantic segmentation, and depth estimation are often highly correlated, and adopting the MTL framework helps to achieve more accurate, detailed, and robust environmental perception.

[0005] The progress of Convolutional Neural Networks (CNNs) and Transformers has significantly enhanced the ability of multi-task learning. For example, Xu et al. introduced a new method in Pad-net: Multi-tasks guided prediction-and-distillation network for simultaneous depth estimation and scene parsing, which is to supervise a set of auxiliary tasks to generate predictions and then use these predictions as multi-modal inputs for the final task. Similarly, Simon et al.'s Mti-net: Multi-scale task interaction networks for multi-task learning explicitly models task interactions at multiple scales, effectively refining task information from low level to high level. However, due to its lack of global interaction ability, the performance of CNN-based methods is limited. Transformer-based methods have shown better results. For example, Ye et al.'s Inverted pyramid multi-task transformer for dense scene understanding optimizes the utilization of multi-task information at the global scale through a multi-scale multi-task feature interaction module. However, due to the failure to decouple the shared information between tasks, there are still deficiencies in multi-task collaborative optimization. In addition, recently, the Mixture of Experts (MoE) model in MTL captures different aspects of data from different perspectives by using a set of experts. Ye et al.'s Taskexpert: Dynamically assembling multi-task representations with memorial mixture-of-experts uses a set of multi-layer perceptron experts for dynamic task routing. This method uses a task-specific routing network to selectively fuse these modal data, promotes cross-task target information sharing, and can model complex task relationships. However, due to the lack of global modeling ability of the constructed experts, the MoE method still has limitations.

[0006] Therefore, in view of the deficiencies and challenges in the above research, the present invention innovatively introduces a mixture of experts into the parameter space of the Mamba model with global construction ability, and uses parameter interaction between multi-tasks to efficiently achieve dynamic modeling of tasks. At the same time, it can use the sequence modeling of Mamba to perceive the effective global feature space, effectively improving the performance of single tasks and the collaborative ability between multi-tasks. Summary of the Invention

[0007] The present invention proposes a strategy for constructing a dense prediction multi-task learning expert in the parameter space of Mamba using a mixture of experts model, while enhancing Mamba's ability to model 2D features. It can effectively improve the collaboration ability between multi-tasks and enhance the performance of single tasks at the same time.

[0008] The technical solution of the present invention:

[0009] A dense prediction multi-task learning method combining a mixture of experts and a Mamba model is as follows:

[0010] (1) Task primary feature extraction

[0011] First, use a pre-trained visual encoder to process the input image to obtain general multi-scale features :

[0012]

[0013] Then, in each scale, use a convolutional encoder for dense prediction tasks (such as semantic segmentation, edge detection, normal estimation, etc.) to obtain task primary features :

[0014]

[0015] Among them, is the number of tasks, is the index of the task;

[0016] (2) Parameter-aware Mamba expert decoder

[0017] Regularize the task primary features to enhance the expression:

[0018]

[0019] Among them is the regularization operation;

[0020] Use depthwise separable convolution and activation layers to process local information:

[0021]

[0022] Among them, is the depthwise separable convolution, is the activation layer;

[0023] After passing through a linear layer, map the channel dimension of the feature to the Mamba parameter space dimension to obtain Mamba features:

[0024]

[0025] Among them is a linear layer;

[0026] (2.1) Hybrid Parameter Expert

[0027] During the calculation of the Mamba parameter space, the calculation processes of parameters B and C in the original Mamba model are both replaced by the calculation of a hybrid task expert;

[0028] Use the convolutional network in the Mamba parameter space to obtain the multi-task expert prediction results:

[0029]

[0030] Among them, is the number of experts;

[0031] Then model the dependencies between different dense prediction tasks and different experts:

[0032]

[0033] Among them, is the global pooling operation, which compresses the spatial dimension to 1; The operation concatenates the features on the channels to obtain the feature ;

[0034] Subsequently, use the linear mapping layer to transform the channel dimension into the number of experts , and then use the activation layer and the function to normalize the channels to obtain the weights of the task-specific experts , and the process is as follows:

[0035]

[0036] Add noise to the expert weights during the training phase, and at the same time select the activated experts as the final activation weights, and the process is as follows:

[0037]

[0038] Among them, is the added noise, is the linear mapping weight, which maps the task features to a scalar as the weight of the noise;

[0039] Finally, use the weighted method to obtain the features of the hybrid parameter expert:

[0040]

[0041] Among them, For the weight of the th expert in the task; for the parameters B and C in the Mamba model, obtain and ; the subscripts B and C represent two calculation paths of the mixture of experts features;

[0042] (2.2) Parameter priors

[0043] Define a set of learnable spatially invariant prior parameters for different dense prediction tasks, denoted as , where is the dimension of the parameter space of the Mamba model; for the parameters and parameter of the Mamba model, add this prior parameter , and also expand the spatial dimension of to obtain :

[0044]

[0045] Among them, The operation changes the spatial dimension 1 to the length L of the input sequence of Mamba. For the parameters and parameter of the Mamba model, add the learnable prior parameters and , and obtain the final parameters and parameter through the following operations:

[0046]

[0047]

[0048] (2.3) Other parameter calculations

[0049] Calculate the parameters A, D, and parameter of Mamba using the parameter calculation method of the original Mamba model;

[0050] (2.4) Multi-directional Hilbert scan module

[0051] Assume that the scale size of the image features is , and find the smallest positive integer such that and . At this time, is the order of the Hilbert space; the transformation of the two-dimensional feature space mapped to the one-dimensional Hilbert space is as follows:

[0052]

[0053] Among them, is the Hilbert space transformation formula of order, and three other transformation methods are obtained by continuously rotating the serialized features 90 degrees around the feature center, denoted as:

[0054] ,

[0055] Among them, the operation is to rotate 90 degrees clockwise times of operations;

[0056] If or , then the transformation also needs to be cropped; the specific method is to retain the area in the lower left of the transformation that is the same size as the image;

[0057] By performing four Hilbert serialization transformations on the Mamba feature , Mamba parameter simultaneously:

[0058]

[0059]

[0060]

[0061]

[0062] Then use the standard Mamba state space model for state space calculation:

[0063]

[0064] Among them, is the standard state model;

[0065] Finally, the operation result of the state space is obtained by adding the state space results in four directions and performing inverse Hilbert serialization:

[0066]

[0067] Among them, is the operation of rotating counterclockwise degrees times, is the output of the task feature;

[0068] (2.5) Integrate Mamba output

[0069] Feed into the linear layer to restore the number of channels before entering the Mamba parameter space:

[0070]

[0071] Use the regularized task primary features to calculate the gating adjustment weights:

[0072]

[0073] Use the gating adjustment weights to dynamically adjust the information flow, and at the same time use skip connections to transmit the task primary features to obtain the output of the parameter-aware Mamba expert decoder ; The process is as follows:

[0074]

[0075] For the outputs of the Mamba expert decoder at different scales , assign subscripts 1, 2, 3, 4 to represent the four scales from shallow to deep where the output features of the visual encoder are located, to obtain multi-task multi-scale features :

[0076]

[0077] (3)Task decoder

[0078] First, spatially align all scale features in the multi-task multi-scale features using the bilinear interpolation algorithm, and the interpolation resolution is of the resolution:

[0079]

[0080]

[0081]

[0082] Among them, is the bilinear interpolation algorithm for spatial interpolation alignment;

[0083] Then concatenate the aligned features on the channels to obtain :

[0084]

[0085] Use a linear layer to weight the scale features, use a convolutional decoder to decode and predict, and use the bilinear interpolation algorithm to restore the spatial scale to the size of the input image :

[0086]

[0087] For Obtain the final prediction result using the post - processing operation corresponding to the task:

[0088]

[0089] A conversion method from the probability estimation result to the label prediction result for task t;

[0090] (4)Multi - task joint optimization strategy

[0091] By combining task loss functions, jointly optimize the network structure using gradient backpropagation; for continuous regression tasks, namely depth estimation and surface normal estimation, use the L1 loss function; for discrete classification tasks, namely semantic segmentation, human parsing, saliency detection, and object boundary detection, use the cross - entropy loss function; the final loss is as follows:

[0092]

[0093] where, is the loss function for task t.

[0094] Advantages of the present invention:

[0095] (1)The present invention innovatively introduces an architecture combining Mamba and MoE mechanisms in the model decoder, specifically for enhancing the effect of multi - task intensive prediction.

[0096] (2)This module effectively uses MoE to balance the relationship between task - specific parameters and adopts the S4 model to capture long - range dependencies, thus significantly enhancing the selectivity and expressive ability of the hidden state.

[0097] (3)In addition, to strengthen the perception of different task characteristics, we incorporate task - specific prior knowledge. To address the mismatch between the unidirectional modeling characteristic of Mamba and image data, we apply a multi - direction scanning method.

[0098] (4)Through extensive experiments on two challenging benchmark tests - NYUD - v2 and PASCAL - Context, the effectiveness of our design scheme is verified. Both quantitative analysis and qualitative evaluation results consistently show that the method we proposed is not only robust but also has superior performance.

[0099] (5) The present invention is of great significance in the fields of autonomous driving and embodied intelligence. In an autonomous driving system, the accurate perception and reasoning ability of the model for multi-tasks (such as road segmentation, obstacle detection, depth estimation, etc.) directly affects safety and reliability. This method effectively improves the generalization ability and real-time performance of the system in complex traffic scenarios by enhancing the model's selective expression and fusion ability for multi-task information. In embodied intelligence, the multi-modal perception and understanding of the environment by the intelligent agent are crucial, especially when performing tasks such as path planning and action decision-making. The multi-directional modeling strategy and task prior injection mechanism proposed by the present invention help to improve the modeling accuracy of the intelligent agent for the environmental state and task adaptability, providing strong support for the embodied intelligent agent to achieve more robust perception and operation capabilities in the real physical world. BRIEF DESCRIPTION OF THE DRAWINGS

[0100] Figure 1 It is the overall architecture diagram of the Mamba expert method for parameter perception related to this application.

[0101] Figure 2 It is the detailed diagram of the Mamba expert decoder for parameter perception related to this application.

[0102] Figure 3 It is the detailed diagram of the hybrid parameter expert related to this application.

[0103] Figure 4 It is the detailed diagram of the multi-directional Hilbert scan related to this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0104] The following further illustrates the specific embodiments of the present invention in conjunction with the drawings and technical solutions.

[0105] Taking dense prediction tasks such as semantic segmentation, edge detection, and normal estimation as examples, the overall framework is as Figure 1 shown.

[0106] Step 1, Task primary feature extraction, the specific steps are as follows:

[0107] The input image passes through a pre-trained visual encoder to extract multi-scale features common to the task. At different scales, convolutional modules for dense prediction tasks are used to decode local features respectively, obtaining multi-task primary features.

[0108] Step 2, Parameter-aware Mamba expert decoding, the specific steps are as follows:

[0109] As Figure 2 shown, the multi-task primary features are regularized to enhance the features. Depthwise separable convolutions and activation layers are used to process local information. After activation, the feature channel dimension is mapped to the Mamba parameter space dimension through a linear layer to obtain Mamba features. Subsequently, the following five steps are calculated in sequence:

[0110] (1)Mixed parameter expert calculation, the specific steps are as follows:

[0111] As Figure 3 shown, a multi-task representation of parameter B and parameter C is generated using a mixture of experts structure. Channel compression and global pooling are used to obtain the weighted weights of different experts. The final expert weights are obtained through Top-K selection and noise addition during training to fuse the expert outputs.

[0112] (2)Parameter prior calculation, the specific steps are as follows:

[0113] As Figure 1 shown, parameter B and parameter C also need to integrate the task parameter prior, define spatially invariant prior parameters for each task, align the feature space through dimension expansion, and combine the output of the mixed parameter expert to obtain the final multi-task Mamba parameters B and C.

[0114] (3)Other parameter calculations, the specific steps are as follows:

[0115] Use the parameter calculation method of the original Mamba model to calculate the Mamba parameters A, D, and parameter .

[0116] (4)Perform multi-directional Hilbert scanning, the specific steps are as follows:

[0117] As Figure 4 shown, adopt multi-directional Hilbert scanning, construct the Hilbert curve of the minimum covering feature map, generate 4 scanning directions through multiple 90-degree rotations, and crop the scans that exceed the feature size. Then generate Mamba features, parameter B, parameter C, and parameter in these 4 directions. Use the SSM standard state equation of Mamba for state space calculation. Subsequently, perform deserialization operations on the outputs of the 4 scanning directions respectively and then add them to obtain the output of the multi-directional Hilbert scanning.

[0118] (5)Integrate the Mamba output, the specific steps are as follows:

[0119] As Figure 2 shown, use a linear layer to decouple the output of the multi-directional Hilbert scanning from the Mamba parameter space and restore the channels of the task features before entering the Mamba parameter space. Calculate the gating adjustment weights using the regularized task primary features through linear and activation layers to dynamically adjust the information flow. Use skip connections to transfer the task primary features. Collect the outputs of different scales to obtain the multi-scale features of the multi-task.

[0120] Step 3, task decoding, the specific steps are as follows:

[0121] As Figure 1As shown, the multi-scale features of multi-tasks need to pass through the task decoder to obtain the final multi-task prediction. In the task decoder, bilinear interpolation is used to spatially align the multi-task multi-scale features, and then the features of different scales are concatenated in the channel dimension. The final task features are obtained by weighted summation through a linear layer. For prediction generation, a task-specific convolutional head is used to decode the features, and the interpolation outputs the probability estimation results. Finally, conventional post-processing operations are used to generate the multi-task prediction map.

[0122] Step 4, training optimization, the specific details are as follows:

[0123] In the training strategy, the task-specific loss functions are combined, and the network structure is jointly optimized through gradient backpropagation. For continuous regression tasks, namely depth estimation and surface normal estimation, the L1 loss function is adopted; while for discrete classification tasks, namely semantic segmentation, human parsing, saliency detection, and object boundary detection, the cross-entropy loss function is used. The final loss is as follows:

[0124]

[0125] where is the loss function of task t, is the final prediction map of task t.

[0126] Experimental verification:

[0127] (1) Verification details

[0128] For the experimental setup, two widely recognized benchmark datasets were selected: NYUD-v2 and PASCAL-Context. The NYUD-v2 dataset contains 795 training images and 654 test images, providing detailed annotations for various tasks such as semantic segmentation, monocular depth estimation, surface normals, and object boundary detection, offering a robust framework for evaluating the effectiveness of models. The PASCAL-Context dataset is larger, containing 4998 training images and 5105 test images, covering diverse annotations for various tasks from semantic segmentation to human parsing and saliency detection, and can be used to evaluate the scalability and robustness of MTL methods in complex and changing scenarios. In terms of implementation details, ViT (including ViT-B and ViT-L) was adopted as the backbone network, and ablation analysis was conducted based on ViT-B. The top-k value and the total number of experts in the parameter expert module were set to 9 and 15 respectively. The entire multi-task learning network was trained for 40,000 iterations on both datasets with a batch size of 6, and the same optimizer and loss function as specified in InvPT were used. The experiments were conducted using 2 Nvidia A100 80G GPUs. Additionally, for fair comparison, the task-average gain Δg was adopted to represent the overall performance of all tasks, and a unified multi-task loss was applied to supervise the learning process of the network.

[0129] (2)Quantitative Results

[0130] We conducted a fair comparison using Vision Transformer as the backbone network on two well-known public datasets, ensuring the same architecture and evaluation criteria as previous studies. Our method was benchmarked against methods such as InvPT, InvPT++, TaskPrompter, M3ViT, MQTransformer, and TaskExpert. As shown in Tables 1 and 2, our method achieved the highest average performance gain values on both datasets. At the same time, we also compared with methods using Swin Transformer, and our method performed better in terms of the average performance gain value. These findings demonstrate the effectiveness of our method and showcase the practical advantages of integrating state space models into multi-task learning.

[0131] Table 1: Comparison results of using two vision encoders, Swin-L, ViT-L, and ViT-B, with current state-of-the-art multi-task dense prediction methods on the PASCAL-Context dataset.

[0132]

[0133] As shown in Table 1, the method proposed in this study demonstrates comprehensive and remarkable performance advantages in the multi-task dense prediction task of the PASCAL-Context dataset. Through systematic comparative experiments, it can be found that when using ViT-L as the backbone network, our method achieves the optimal performance in four tasks: semantic segmentation (81.38 mIoU), human part estimation (70.39 mIoU), normal estimation (13.36 mEu), and edge detection (75.30 odsF). Especially in the semantic segmentation and edge detection tasks, our method improves by 1.16% and 1.10% respectively compared with the previous best method, and the overall performance gain reaches 3.60%, which is the most prominent among all comparison methods. It is worth noting that even on the ViT-B backbone network with lower computational resource requirements, our method still maintains an average performance gain of 1.86% and performs optimally compared with other methods in two tasks: saliency detection (85.25 maxF) and normal estimation (13.40 mEu). This result indicates that our method can not only fully utilize the advantages brought by large model capacity, but also its innovative architecture design can achieve efficient feature learning on smaller models. Compared with the latest methods such as InvRT++ and TaskPrompter, our improvement is more evenly reflected in all tasks rather than only achieving improvements in specific tasks, which fully demonstrates that the method has better generalization ability.

[0134] Table 2: Comparison results of using two vision encoders, Swin-L, ViT-L, and ViT-B, respectively, with the current advanced dense prediction methods on the NYUD-v2 dataset.

[0135]

[0136] As shown in Table 2, the method proposed in this study demonstrates excellent performance in the multi-task dense prediction task of the NYUD-v2 dataset. Through systematic comparative experiments, it can be found that when using VIT-L as the backbone network, our method achieves competitive results in four tasks: semantic segmentation (56.63 mIoU), depth estimation (0.5102 RMSE), normal estimation (18.52 mEu), and edge detection (78.10 odsE). Notably, in the semantic segmentation task, our method improves by 1.28 mIoU compared to the previous best TaskExpert method, which fully demonstrates the advantage of the method in complex scene understanding. From the perspective of overall performance gain, our method achieves an average improvement of 5.23%, ranking first among all compared methods, which fully verifies the comprehensive performance advantage of the method. It is worth noting that even on the VIT-B backbone network with lower computational resource requirements, our method still maintains an average performance gain of 1.02% and reaches the optimal level in the normal estimation (18.90 mEu) and edge detection (78.00 odsE) tasks. This result indicates that our method can not only fully utilize the advantages brought by large model capacity, but also its innovative multi-task learning mechanism can achieve efficient feature representation on smaller models. Compared with the latest methods such as InvPT++ and TaskPromoter, our improvement is more comprehensive in all task dimensions, especially in the two challenging tasks of semantic segmentation and depth estimation, which fully proves that the method has better task adaptability and generalization ability.

Claims

1. A dense prediction multi-task learning method combining mixed experts and Mamba model, characterized in that: The details are as follows: (1) Extraction of primary task features; (2) Parameter-aware Mamba expert decoder; Regularize the primary features of the task to enhance the expression: in is the regularization operation; Use depth-wise separable convolution and activation layers to process local information: in, is a depth-wise separable convolution, is the activation layer; After the linear layer, the features The channel dimension is mapped to the Mamba parameter space dimension to obtain the Mamba feature: in is a linear layer; (2.1) Hybrid Parameter Expert In the calculation process of Mamba parameter space, the calculation process of parameter B and parameter C in the original Mamba model is replaced by a calculation of a hybrid task expert; Use a convolutional network in the Mamba parameter space to obtain multi-task expert prediction results: in, is the number of experts; Then model the dependencies between different intensive prediction tasks and different experts: in, It is a global pooling operation, compressing the spatial dimension to 1; The operation connects the features on the channel to get the features ; A linear mapping layer is then used to transform the channel dimension into the number of experts , and then use the activation layer and function Normalize the channels to get the weights of task-specific experts , the process is as follows: Noise is added to the expert weights during the training phase, and by selecting The activated expert is used as the final activation weight, and the process is as follows: in, is the added noise, To linearly map weights, the task features are mapped to a scalar as the weight of the noise; Finally, a weighted approach is used to obtain the characteristics of the mixed parameter expert: in, For the task No. The weight of the experts; for the parameters B and C in the Mamba model, we get as well as ; The subscripts B and C represent two hybrid expert feature calculation paths; (2.2) Parameter priors Define a set of learnable spatially invariant prior parameters for different dense prediction tasks, denoted as ,in is the dimension of the Mamba model parameter space; is the parameter of the Mamba model and parameters Add this prior parameter , and also By expanding the spatial dimension of : in, The operation transforms the spatial dimension 1 into the length L of the input sequence of Mamba, which is the parameter and parameters Adding learnable prior parameters and , the final parameters are obtained by the following operations and parameters : (2.3) Calculation of other parameters Use the original Mamba model parameter calculation method to calculate Mamba's parameters A, D and ; (2.4) Multi-directional Hilber scanning module Assume the scale of image features is , find the smallest positive integer make and ,at this time is the order of the Hilbert space; the transformation of the two-dimensional feature space to the one-dimensional Hilbert space is as follows: in, for The Hilber space transformation formula of the order is obtained by rotating the serialized features continuously by 90 degrees around the feature center to obtain three other transformation methods, which are recorded as: , in, The operation is to rotate 90 degrees clockwise times of operation; like or , the transformation must also be cropped; the specific method is to retain the area at the lower left of the transformation that is the same size as the image; By using the Mamba feature , Mamba parameters Four Hilber serialization transformations are performed simultaneously: Then the state space calculation is performed using the standard Mamba state space model: in, It is the standard state model; Finally, the state space results in the four directions are summed and deserialized to obtain the state space operation results: in, For counterclockwise rotation Spend The operation of is the output of the task features; (2.5) Integrate Mamba output Will Enter a linear layer to recover the number of channels before entering the Mamba parameter space: Use regularized task-level features To calculate the gate adjustment weight: Use gated weights to dynamically adjust information flow and use skip connections to transfer primary features of the task to obtain the output of the parameter-aware Mamba expert decoder. ; The process is as follows: Output of Mamba expert decoder for different scales , assign subscripts 1, 2, 3, and 4 to represent the four scales from shallow to deep where the visual encoder outputs features, and obtain multi-task multi-scale features : ; (3) Task decoder; (4) Multi-task joint optimization strategy.

2. The dense prediction multi-task learning method of the joint hybrid expert and Mamba model according to claim 1 is characterized in that: The specific implementation process of step (1) task primary feature extraction is as follows: First, use the pre-trained visual encoder Processing input images Obtaining universal multi-scale features : Then, a convolutional encoder for dense prediction tasks is used at each scale to obtain the primary features of the task. : in, is the number of tasks, The index of the task.

3. The dense prediction multi-task learning method of the joint hybrid expert and Mamba model according to claim 2 is characterized in that: The specific implementation process of step (3) task decoder is as follows: First, the multi-task multi-scale features All scale features in are spatially aligned using a bilinear interpolation algorithm with an interpolation resolution of Resolution: in, It is a bilinear interpolation algorithm used for spatial interpolation alignment; Then the aligned features are spliced ​​on the channel to obtain : Use a linear layer to weight the scale features, use a convolutional decoder to decode the prediction, and use a bilinear interpolation algorithm to restore the spatial scale to the input image. Size: right Use the post-processing operations of the corresponding task to obtain the final prediction results: It is a method for converting the probability estimation result to the label prediction result of task t.

4. The dense prediction multi-task learning method of the joint hybrid expert and Mamba model according to claim 3 is characterized in that: The specific implementation process of step (4) multi-task joint optimization strategy is as follows: By combining the task loss functions, the network structure is jointly optimized using gradient back propagation; for continuous regression tasks, namely depth estimation and surface normal estimation, the L1 loss function is used; for discrete classification tasks, namely semantic segmentation, human body parsing, saliency detection, and object boundary detection, the cross entropy loss function is used; the final loss is as follows: in, is the loss function of task t.

Citation Information

Patent Citations

  • Hierarchical multi-task learning method for intensive prediction tasks

    CN118247635A

  • Prostate MRI image segmentation method based on Mamba-Unet

    CN119579627A

  • Pathological image classification system and method based on Mama and hybrid expert model

    CN119851017A

  • Copilot architecture: network of microservices including specialized machine learning tools

    US12242503B1

  • Vision-language model with an ensemble of experts

    US20240265690A1

Cited By

  • Multi-task prediction method and system based on hybrid expert network, electronic equipment and medium

    CN122132758A

  • Multi-task prediction method and system based on hybrid expert network, electronic device and medium

    CN122132758B