The invention discloses a
singular value decomposition-based multi-
modal task
fine tuning method for enhancing a routing function, which comprises the following steps of: mapping
input language and visual features from a high-dimensional space to a low-rank space by using a PEFT method, performing
singular value decomposition on language features in the low-rank space, performing routing function alignment through a
tensor after efficient reconstruction, and performing multi-
modal task
fine tuning on a multi-
modal task based on a
singular value decomposition-enhanced routing function. And finally, after the low-rank space is recovered to the original dimension again, performing residual connection with the original language features, and outputting the features. According to the method,
singular value decomposition is applied to language features before a routing function, a low-rank dominant mode of the language features is extracted, the alignment precision of vision and language features is enhanced, interference of high-dimensional
noise is eliminated, and meanwhile calculation efficiency and model stability are kept. Routing calculation is carried out through the reconstructed
tensor, key information in the features can be better extracted and aligned, and therefore the precision and effect of feature alignment are improved. The method is suitable for VL tasks such as visual questioning and answering and
image description generation, and model performance can be obviously improved.