The invention discloses a model generalization ability optimization method and
system based on
hybrid experts, and aims to solve the problems of weak generalization ability, difficulty in knowledge migration, low
semantic alignment efficiency and the like when a multi-
modal large model performs fine adjustment on downstream tasks in a new field. The method comprises the following steps: embedding a
hybrid expert module in a low-rank adaptive
fine tuning framework of a self-attention module in the visual
encoder, introducing a heterogeneous
hybrid convolution structure, innovatively realizing transverse and longitudinal independent scaling through bilinear interpolation in the
convolution structure, supporting a multi-scale
convolution kernel, and realizing dynamic combination through an MoE framework; constructing a cross-
modal expert adapter of channel
perception, and adopting
dynamic channel recombination, cross-
modal conditional convolution and weighted spatial
feature aggregation methods; according to the method, the problem of local semantic
confusion can be well solved, meanwhile, semantic fusion of image features and text features is enhanced, more effective spatial information and priori knowledge are injected into the model, and therefore the generalization of the model is improved.