A method and system for optimizing model generalization ability based on mixed experts
By introducing hybrid expert modules and heterogeneous hybrid convolutional structures into multimodal large models, and combining them with low-rank adaptive methods, the inductive bias problem in the domain generalization process of multimodal large models is solved, thereby improving the model's cross-domain adaptability and semantic fusion effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XI AN JIAOTONG UNIV
- Filing Date
- 2025-11-19
- Publication Date
- 2026-07-14
AI Technical Summary
Existing multimodal large models suffer from insufficient utilization of inductive bias and low efficiency in information fusion between modalities during domain generalization, resulting in inaccurate semantic understanding and low adaptation efficiency when the model is applied to new domains, making it difficult to meet complex and ever-changing practical needs.
We adopt a text-visual bimodal Transformer architecture, introduce hybrid expert modules (MoE) and heterogeneous hybrid convolutional structures, and combine low-rank adaptive methods to embed a low-rank adaptive adjustment module in the visual encoder. We also introduce a channel-aware cross-modal adapter in the cross-modal layer to enhance modal interaction through dynamic channel recombination and conditional convolution, thereby improving the model's cross-domain adaptability.
By dynamically combining multi-scale convolutional kernels and channel-aware mechanisms, the model can extract more spatial information and inductive bias, solve the problem of local semantic confusion, enhance the semantic fusion of image features and text features, and improve the model's generalization and adaptation efficiency.
Smart Images

Figure CN121581146B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multimodal large model generalization technology, and specifically relates to a method and system for optimizing model generalization capability based on hybrid experts. Background Technology
[0002] With the rapid development of artificial intelligence technology, multimodal large models have become an important research direction in the field of artificial intelligence due to their ability to integrate information from multiple sources such as images, text, and audio. However, in practical applications, multimodal large models face the key challenge of domain generalization.
[0003] In terms of domain generalization, traditional full-parameter fine-tuning methods suffer from high computational costs, making it difficult to quickly adapt to diverse application domains. To address this issue, efficient parameter fine-tuning techniques have been widely researched and applied. For example, adapter-based techniques fine-tune for specific tasks without altering the original pre-trained model structure; some methods combine hybrid expert (MoE) and low-rank adaptation (LoRA) techniques to reduce computational costs while improving the model's cross-domain task adaptability; other methods achieve near-full-parameter fine-tuning effects with a small number of newly added trainable parameters through interactive prompts and deep parameter tuning, or enhance modal collaboration through dynamic gradient adjustment. However, these existing domain generalization methods generally suffer from a drawback: they fail to fully consider inductive biases in different domains, leading to inaccurate semantic understanding and low adaptation efficiency when the model is applied to new domains, making it difficult to meet complex and ever-changing practical needs.
[0004] Current multimodal large models suffer from technical bottlenecks in domain generalization, such as insufficient utilization of inductive bias and low efficiency in intermodal information fusion. Therefore, there is an urgent need for an innovative method that can integrate domain-specific knowledge, deeply explore modal interaction mechanisms, and achieve efficient domain adaptation and modal extension to overcome the shortcomings of existing technologies.
[0005] According to the applicant's research, the existing technologies related to this invention in the field of large-scale model generalization include: Patent CN119067236A provides a method and system for reducing the impact of large language model fine-tuning on generalization ability by combining system prompts. This patent constructs a training data template according to the professional domain problem to be solved; obtains several training data based on the training data template; then mixes the training data with the open-source data of the model to be fine-tuned, and adjusts the mixing ratio of the training data and the open-source data to obtain the optimal mixing ratio; finally, fine-tunes the model to be fine-tuned according to the optimal mixing ratio. However, its shortcomings are that adjusting the large model is done through full-parameter training, resulting in excessively long training time; it only processes the data and does not involve the model itself. Therefore, its generalization ability is insufficient.
[0006] The patent with publication number CN119863687A provides a diversity inductive bias learning method for generalizable large models. It proposes a text-level inductive bias by supplementing the prompt text with numerous LLM-generated descriptions to provide detailed information for each category. To enable the model to capture inductive bias effectively, a phrase adapter is designed for the text encoder to explicitly explore connections between adjacent words; a spatial adapter is designed for the image encoder to allow the model to see more local relationships and details. The patent reduces overfitting through optimized inductive bias, achieved through a dynamic training strategy that allows the model to learn different degrees of fit. However, this method tunes large models through full parameter tuning, resulting in excessively long training times. It lacks a channel-aware cross-modal expert adapter and fails to consider biases in modalities other than text, such as the inductive bias of the image modality, leading to poor preservation of cross-modal generalization ability. Summary of the Invention
[0007] To address the problems existing in the prior art, this invention provides a method and system for optimizing model generalization ability based on hybrid experts. This method is based on a text-visual bimodal Transformer architecture, embedding a Mixture-of-Experts (MoE) module in the visual encoder layer, introducing a heterogeneous hybrid convolutional structure, and employing a low-rank adaptive method for fine-tuning; a channel-aware cross-modal adapter is introduced in the cross-modal layer. This solves the problem of weak generalization ability of large models when facing new domains.
[0008] To achieve the above objectives, in a first aspect, the present invention provides a method for optimizing the generalization ability of a model based on hybrid experts, comprising the following steps:
[0009] Step 1: Using the text-visual bimodal Transformer architecture as the original model, inject low-rank parameterized increments into the self-attention projection matrix of the visual encoder and embed a low-rank adaptive adjustment module.
[0010] Step 2: Embed a hybrid expert module in the low-rank adaptive adjustment module of the self-attention module in the visual encoder, and introduce a heterogeneous hybrid convolutional structure; the hybrid expert module to be embedded includes multiple expert networks, a gating module, an amplification module, a heterogeneous convolutional module and a shrinking module;
[0011] Step 3: Integrate a channel-aware cross-modal expert adapter into the backbone network of the original model to obtain a model based on low-rank adaptive heterogeneous hybrid convolutional experts.
[0012] Furthermore, in step 1, based on the text-visual bimodal Transformer architecture, a low-rank adaptive method is used to adjust the projection matrix of the self-attention module in the visual encoder. A low-rank increment parameter is introduced for adjustment, wherein... The dimensions representing the input parameters are as follows:
[0013]
[0014] Let the input features of the i-th layer be... Output features Original weight matrix Keep the low-rank increment parameter constant during fine-tuning. .
[0015] Furthermore, in step 2, the gating network includes: image semantics. Routing weights among experts are assigned through a gating mechanism and generated by a lightweight gating network.
[0016]
[0017] In the formula, The semantic vector of the image modality. For image feature dimensions, These are learnable projection parameters.
[0018] Furthermore, in step 2, the specific process for amplifying the input features is as follows: For the input features... ,expert Press the input feature map horizontally Magnification, vertically Magnification:
[0019]
[0020] in, It indicates that after expert consultation The feature map after interpolation and magnification. The scaling factor controls the scaling ratio of height and width respectively. to times, The interpolation is represented by H and W, which are their original height and width.
[0021] The specific process of heterogeneous convolution is as follows: Expert The magnified feature map is convolved using convolution kernels of heterogeneous sizes, as follows:
[0022]
[0023] in, This represents the feature map after convolution. This indicates the expert's Convolution kernel, , which represents the size of the convolution kernel;
[0024] The specific process for narrowing down input features is as follows: Experts The convolutional feature maps are then sorted according to... Multiplier, vertically The magnification is reduced to restore the original size, using the following formula:
[0025]
[0026] in It indicates that after expert consultation Reduced feature map, This represents the feature map after convolution. The scaling factor controls the scaling ratio of height and width respectively. to The times, H, and W represent their original height and width.
[0027] Furthermore, the steps performed by the channel-aware cross-modal expert adapter include:
[0028] 1) Perform dynamic channel segmentation, assigning all channels of the input visual features to each expert without overlap;
[0029] 2) Each expert dynamically generates a convolution kernel and performs a convolution operation based on the cross-modal semantics of the text embedding vector;
[0030] 3) The features obtained from all expert convolutions are weighted and aggregated to obtain the aggregated high-dimensional features; the weighting criterion is based on the semantic relevance between the features obtained from all expert convolutions and the text embedding vector.
[0031] 4) The aggregated high-dimensional features are fused with the original image input features to achieve cross-modal semantic enhancement of image features.
[0032] Furthermore, in step 1), the input visual features are... and text features Divide the dataset into d subsets based on channel dimension, and compute the joint feature for each channel. With each expert semantic relevance score :
[0033]
[0034] in, Representing visual features Feature map of each channel Indicates the text feature number Feature map of each channel This represents the k-th output of the gating network.
[0035] Furthermore, in step 2), each expert The channel group assigned to it Dynamically generate conditional convolution kernels for text and perform convolution:
[0036]
[0037] in The kernel size is the convolution kernel size. For experts The number of channels allocated;
[0038] expert Channel characteristics assigned to it Perform convolution operations to extract single-channel features containing high-dimensional abstract semantics. , .
[0039] Furthermore, in step 3), higher weights are assigned to experts whose work is highly relevant to the text. The weights are calculated using the following formula:
[0040]
[0041]
[0042] in, As weight, For similarity, The temperature coefficient controls the sharpness of the weight distribution;
[0043] Aggregate the output features of all experts according to their weights:
[0044]
[0045] in, These are the high-dimensional features after aggregation. Single-channel characteristics.
[0046] Furthermore, in step 4), the weighted single-channel features are expanded to the original channel dimension and concatenated with the input feature residuals:
[0047]
[0048] in, This indicates that the copy will be performed along the channel dimension. Second-rate.
[0049] Secondly, this invention provides a model generalization capability optimization system based on hybrid experts. It uses a text-visual bimodal Transformer architecture as the original model, comprising an image encoder, a text encoder, and a large text model. The image encoder includes an image input module, which extracts features through N repetitions of the Transformer structure, and then obtains the final model using an MLPProjector. Simultaneously, the text input is obtained through tokenizing and embedding. Subsequently, the large text model receives... and By repeating Self Attention M times, the channel-aware cross-modal adapter Cross Attention and FFN complete cross-modal feature alignment and fusion; wherein, the self-attention projection matrix of the visual encoder is injected with low-rank parameterized increments and a low-rank adaptive adjustment module is embedded.
[0050] A hybrid expert module is embedded in the low-rank adaptive adjustment module of the self-attention module in the visual encoder, introducing a heterogeneous hybrid convolutional structure; the proposed hybrid expert module includes multiple expert networks, a gating module, an amplification module, a heterogeneous convolutional module, and a shrinking module;
[0051] By integrating a channel-aware cross-modal expert adapter into the backbone network of the original model, a system based on low-rank adaptive heterogeneous hybrid convolutional experts is obtained.
[0052] Compared with existing technologies, this invention has at least the following beneficial effects: This invention embeds a Mixture-of-Experts (MoE) module into the low-rank adaptive fine-tuning framework of the self-attention module in the visual encoder, introducing a heterogeneous hybrid convolutional structure. Innovatively, bilinear interpolation is used in the convolutional structure to achieve independent scaling in the horizontal and vertical directions, enabling convolutional operations to extract local priors that were previously unavailable due to spatial folding in a certain direction (such as when the original object and the lens direction are at a certain angle). Furthermore, it supports dynamic combination of multi-scale convolutional kernels through the MoE architecture, extracting richer multi-scale spatial prior knowledge and inductive biases from feature maps and injecting them into the model, thus giving the model better generalization ability. Simultaneously, a channel-aware cross-modal expert adapter is constructed. Through dynamic channel recombination, cross-modal conditional convolution, and weighted spatial feature aggregation, it not only effectively solves the problem of local semantic confusion but also enhances the semantic fusion of image and text features, injecting more effective spatial information and prior knowledge into the model, thereby improving the model's generalization ability. Attached Figure Description
[0053] Figure 1 This is a schematic diagram of the main structure of the present invention.
[0054] Figure 2 This is the flowchart of the algorithm of this invention. Detailed Implementation
[0055] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings and examples.
[0056] The present invention is as follows Figure 1 and Figure 2 As shown, a novel domain generalization method based on hybrid experts is proposed. It is grounded in a text-visual bimodal Transformer architecture, embedding a Mixture-of-Experts (MoE) module in the visual encoder layer, introducing a heterogeneous hybrid convolutional structure, and employing a low-rank adaptive method for fine-tuning. A channel-aware cross-modal adapter is introduced in the cross-modal layer. These two innovations primarily enhance the semantic fusion of image and text features, inject more spatial information, and generalize biases, thereby improving the model's generalization ability.
[0057] Step 1: Based on the text-visual bimodal Transformer architecture, the Low-Rank Adaptation (LoRA) method is used to project the self-attention module in the visual encoder onto the image. Adjustment is achieved by introducing a low-rank incremental parameter, where The dimension representing the input features.
[0058]
[0059] Let the input features of the i-th layer be... Output features Original weight matrix The low-rank incremental parameters are kept constant during adjustment to preserve general knowledge from the pre-training phase. .
[0060] Step 2: Embed a Mixture-of-Experts (MoE) module within the low-rank adaptive adjustment framework of the self-attention module in the visual encoder. This introduces a heterogeneous hybrid convolutional structure to enhance the spatial-semantic modeling capability of local features. Specifically, this includes:
[0061] Step 201: The proposed MoE module contains multiple expert networks and a gating module. During the forward pass, the MoE module dynamically selects which experts to activate. Image semantics. Routing weights among experts are assigned through a gating mechanism and generated by a lightweight gating network.
[0062]
[0063] In the formula, For routing weights among experts, The semantic vector of the image modality. For image feature dimensions, For learnable projection parameters, Sparsemax is the sparse maximization function.
[0064] Step 202: For input features ,expert Press the input feature map horizontally Multiplier, vertically Magnification is applied when Then isotropic amplification occurs, when Then anisotropic amplification occurs.
[0065]
[0066] in, It indicates that after expert consultation The feature map after interpolation and magnification. The scaling factor controls the scaling ratio of height and width respectively. to times, The interpolation is represented by H and W, which are the original height and width of the feature map.
[0067] Step 203: Expert It uses its proprietary heterogeneous-sized convolution kernels to convolve the magnified feature map, as shown in the formula:
[0068]
[0069] in, This represents the feature map after convolution. This indicates the expert's exclusive Convolution kernel, where This indicates the size of the convolution kernel.
[0070] Step 204: Expert The convolutional feature maps are then sorted according to... Multiplier, vertically The magnification is reduced to restore the original size, using the following formula:
[0071]
[0072] in, It indicates that after expert consultation Reduced feature map, This represents the feature map after convolution. The scaling factor controls the scaling ratio of height and width respectively. to H and W represent the original height and width of the feature map.
[0073] Step 3: Design a channel-aware cross-modal expert adapter and integrate it into a multimodal large-scale model backbone network. This adapter includes the following sub-steps:
[0074] Step 301: Input visual features and text features Divide the input into d subsets according to the channel dimension, and calculate the joint feature of the input visual features and input text features for each channel. With each expert semantic relevance score :
[0075]
[0076] in, Representing visual features Feature map of each channel Indicates the text feature number Characteristics of each channel The k-th output of the gated network in the channel-aware cross-modal expert adapter represents the channel-aware network.
[0077] Using sparsification functions , channel right The consecutive matching scores of each expert are converted into strict one-hot encoding, thus ensuring that each channel is assigned to only one expert:
[0078]
[0079] in, Indicates channel right An expert dimensional matching score vector, For binary values, when the channel Assigned to experts hour, Otherwise Ultimately, all experts' channels will be assigned to corresponding... Channel group And satisfy ;
[0080] Step 302: Each expert The channel group assigned to it Dynamically generate conditional convolution kernels for text and perform convolution.
[0081]
[0082] in, The kernel size is the convolution kernel size. For experts The number of channels allocated, experts Channel characteristics assigned to it Perform convolution operations to extract single-channel features containing high-dimensional abstract semantics. .
[0083]
[0084] Step 303: Assign higher weights to experts who are highly relevant to the text to ensure that the final aggregated features can highlight the key spatial semantics of the aggregated features. The weights are calculated using the following formula:
[0085]
[0086]
[0087] in, For routing weights among experts, For similarity, The temperature coefficient controls the sharpness of the weight distribution.
[0088] Step 304: Aggregate the output features of all experts according to the routing weights among the experts:
[0089]
[0090] The weighted single-channel features are then expanded to the original channel dimension and concatenated with the input image feature residuals.
[0091]
[0092] in, This indicates that the copy will be performed along the channel dimension. Second-rate.
[0093] On the other hand, reference Figure 1 This invention provides a model generalization optimization system based on hybrid experts, using a text-visual bimodal Transformer architecture as the original model. The overall structure of this model consists of three parts: an image encoder, a text encoder, and a large text model. The image encoder includes an image input module, which extracts features through N repetitions of the Transformer structure, and then obtains the final model using an MLP Projector. Simultaneously, the text input is obtained through tokenizing and embedding. The large text model then receives... and By repeating Self Attention M times, the channel-aware cross-modal adapter CrossAttention and FFN complete cross-modal feature alignment and fusion.
[0094] In this design, a hybrid expert module is embedded in the low-rank adaptive adjustment module of the self-attention module in the visual encoder, introducing a heterogeneous hybrid convolutional structure. The proposed hybrid expert module includes multiple expert networks, a gating module, an amplification module, a heterogeneous convolutional module, and a shrinking module, in order to extract from the feature map and inject richer multi-scale spatial prior knowledge and inductive bias into the model.
[0095] The original model backbone network integrates a channel-aware cross-modal expert adapter, which includes feature channel grouping and channel MOE. The channel MOE includes a gating network, multiple expert networks and convolutional modules to enhance image features through text features.
[0096] By combining all the above modules, a system based on low-rank adaptive heterogeneous hybrid convolutional expert and channel-aware cross-modal expert is obtained.
[0097] In summary, this invention discloses a method and system for optimizing model generalization ability based on hybrid experts. It embeds a Mixture-of-Experts (MoE) module into the low-rank adaptive fine-tuning framework of the self-attention module in a visual encoder, introducing a heterogeneous hybrid convolutional structure. Innovatively, bilinear interpolation is used within the convolutional structure to achieve independent horizontal and vertical scaling, and multi-scale convolutional kernels are supported. Dynamic combination is achieved through the MoE architecture. A channel-aware cross-modal expert adapter is constructed, employing methods such as dynamic channel recombination, cross-modal conditional convolution, and weighted spatial feature aggregation. This invention not only effectively solves the problem of local semantic confusion but also enhances the semantic fusion of image and text features, injecting more effective spatial information and prior knowledge into the model, thereby improving its generalization ability.
Claims
1. A method for optimizing the generalization ability of a model based on hybrid experts, characterized in that, Includes the following steps: Step 1: Using the text-visual bimodal Transformer architecture as the original model, inject low-rank parameterized increments into the self-attention projection matrix of the visual encoder and embed a low-rank adaptive adjustment module. Step 2: Embed a hybrid expert module into the low-rank adaptive adjustment module of the self-attention module in the visual encoder, introducing a heterogeneous hybrid convolutional structure; the proposed hybrid expert module includes multiple expert networks, a gating module, an amplification module, a heterogeneous convolutional module, and a reduction module; in the gating network: image semantics... Routing weights among experts are assigned through a gating mechanism and generated by a lightweight gating network. In the formula, The semantic vector of the image modality, For image feature dimensions, For learnable projection parameters; The specific process for amplifying input features is as follows: For input features ,expert Press the input feature map horizontally Magnification, vertically Magnification: in, It indicates that after expert consultation The feature map after interpolation and magnification. The scaling factor controls the scaling ratio of height and width respectively. to times, The interpolation is represented by H and W, which are their original height and width. The specific process of heterogeneous convolution is as follows: Expert The magnified feature map is convolved using convolution kernels of heterogeneous sizes, as follows: in, This represents the feature map after convolution. This indicates the expert's Convolution kernel, , which represents the size of the convolution kernel; The specific process for narrowing down input features is as follows: Experts The convolutional feature maps are then sorted according to... Multiplier, vertically The magnification is reduced to restore the original size. The formula is: in It indicates that after expert consultation Reduced feature map, This represents the feature map after convolution. The scaling factor controls the scaling ratio of height and width respectively. to The times, H, and W are their original height and width; Step 3: Integrate a channel-aware cross-modal expert adapter into the original model backbone network to obtain a model based on low-rank adaptive heterogeneous hybrid convolutional experts; the steps performed by the channel-aware cross-modal expert adapter include: 1) Perform dynamic channel segmentation, assigning all channels of the input visual features to each expert without overlap; 2) Each expert dynamically generates a convolution kernel and performs a convolution operation based on the cross-modal semantics of the text embedding vector; 3) The features obtained from all expert convolutions are weighted and aggregated to obtain the aggregated high-dimensional features; the weighting criterion is based on the semantic relevance between the features obtained from all expert convolutions and the text embedding vector. 4) The aggregated high-dimensional features are fused with the original image input features to achieve cross-modal semantic enhancement of image features.
2. The method for optimizing model generalization ability based on hybrid experts according to claim 1, characterized in that, In step 1, based on the text-visual bimodal Transformer architecture, a low-rank adaptive method is used to project the self-attention module's projection matrix in the visual encoder. A low-rank increment parameter is introduced for adjustment, wherein... The dimensions representing the input parameters are as follows: Let the input features of the i-th layer be... Output features Original weight matrix Keep the low-rank increment parameter constant during fine-tuning. .
3. The method for optimizing model generalization ability based on hybrid experts according to claim 1, characterized in that, In step 1), the input visual features and text features Divide the dataset into d subsets based on channel dimension, and compute the joint feature for each channel. With each expert semantic relevance score : in, Representing visual features Feature map of each channel Indicates the text feature number Feature map of each channel This represents the k-th output of the gating network.
4. The method for optimizing model generalization ability based on hybrid experts according to claim 1, characterized in that, In step 2), each expert The channel group assigned to it Dynamically generate conditional convolution kernels for text and perform convolution: in The kernel size is the convolution kernel size. For experts The number of channels allocated; expert Channel characteristics assigned to it Perform convolution operations to extract single-channel features containing high-dimensional abstract semantics. , .
5. The method for optimizing model generalization ability based on hybrid experts according to claim 3, characterized in that, In step 3), higher weights are assigned to experts whose work is highly relevant to the text. The formula for calculating the weights is as follows: in, As weight, For similarity, The temperature coefficient controls the sharpness of the weight distribution; Aggregate the output features of all experts according to their weights: in, These are the high-dimensional features after aggregation. Single-channel characteristics.
6. The method for optimizing model generalization ability based on hybrid experts according to claim 3, characterized in that, In step 4), the weighted single-channel features are expanded to the original channel dimension and concatenated with the input feature residuals: in, This indicates that the copy will be performed along the channel dimension. Second-rate.
7. A model generalization capability optimization system based on hybrid experts, characterized in that, This method, used to implement the hybrid expert-based model generalization optimization method as described in any one of claims 1-6, uses a text-visual bimodal Transformer architecture as the original model, comprising three parts: an image encoder, a text encoder, and a large text model. The image encoder includes an image input module, which extracts features through N repetitions of the Transformer structure, and then obtains the final model using an MLP Projector. Simultaneously, the text input is obtained through tokenizing and embedding. Subsequently, the large text model receives... and By repeating Self Attention M times, the channel-aware cross-modal adapter Cross Attention and FFN complete cross-modal feature alignment and fusion; wherein, the self-attention projection matrix of the visual encoder is injected with low-rank parameterized increments and a low-rank adaptive adjustment module is embedded. A hybrid expert module is embedded in the low-rank adaptive adjustment module of the self-attention module in the visual encoder, introducing a heterogeneous hybrid convolutional structure; the proposed hybrid expert module includes multiple expert networks, a gating module, an amplification module, a heterogeneous convolutional module, and a shrinking module; By integrating a channel-aware cross-modal expert adapter into the backbone network of the original model, a system based on low-rank adaptive heterogeneous hybrid convolutional experts is obtained.
Citation Information
Patent Citations
Method and system for reducing influence of large language model fine tuning on generalization ability in combination with system prompt
CN119067236A
Diversity induction bias learning method for generalizable large model
CN119863687A
Face forgery detection method and system based on reconstruction learning and hybrid expert mode
CN119625812A
Visual language model continuous learning method based on dynamic hybrid expert adapter
CN120766064A