Multi-modal task fine tuning method based on singular value decomposition enhanced routing function
By optimizing multimodal tasks through singular value decomposition and routing functions, the alignment accuracy problem caused by high-dimensional noise in language features is solved, thereby improving the feature alignment accuracy and model performance of vision-language tasks.
Patent Information
- Application Number
- CN202511082633.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-25
AI Technical Summary
In existing multimodal tasks, the alignment effect of routing functions depends on the quality of input features. High-dimensional noise or redundant information in language features leads to information loss, affecting the accuracy of modality alignment and reducing computational efficiency and model stability.
Singular value decomposition is used to preprocess language features, which are then mapped to a low-dimensional space through a low-rank bottleneck to preserve core information. A routing function is used for cross-modal alignment, and a shared cache mechanism is combined to optimize computational efficiency and enhance feature alignment accuracy.
It improves feature alignment accuracy and model performance in vision-language tasks, reduces computational complexity, maintains model stability, and is suitable for tasks such as visual question answering and image description generation.
Smart Images

Figure CN121009338A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal task technology, specifically a fine-tuning method for multimodal tasks based on singular value decomposition to enhance routing functions. Background Technology
[0002] Existing parameter efficient fine-tuning (PEFT) techniques in multimodal tasks, such as Adapter and LoRA, primarily focus on cross-modal feature alignment of input visual and linguistic features to improve the model's adaptability to specific tasks. These methods introduce low-rank bottlenecks into the hidden layers of the pre-trained model, projecting high-dimensional features (e.g., visual and linguistic features) down to a low-dimensional space for processing, and then restoring the original dimensions using an up-projection matrix.
[0003] Before performing upprojection to restore the original dimension, a routing function operation is used to enhance the VL alignment in the low-rank bottleneck. Then, the low-dimensional features are mapped back to the high-dimensional space through the upprojection matrix, and residual connections are made with the original features to generate the final output. This method reduces the number of trainable parameters through the low-rank bottleneck, thereby reducing computation and storage costs, and is suitable for resource-constrained scenarios.
[0004] While introducing routing functions into existing multimodal tasks can effectively improve model performance and enable data from different modalities to interact through specific transformations, the effectiveness of routing functions depends on the quality and representation of the input features. Language features often contain high-dimensional noise or redundant information, and directly projecting them into a low-dimensional space may lead to information loss, affecting the alignment effect of the routing function. Routing functions directly manipulate the projected features and lack a preprocessing mechanism for language features, making it difficult to effectively extract dominant patterns and limiting the accuracy of modality alignment. When the batch size is large or the sequence length is long, the complexity of language features may lead to a decrease in computational efficiency and model stability. Summary of the Invention
[0005] In view of this, the present invention addresses the problem of insufficient language feature processing in existing PEFT methods in VL tasks, and provides a fine-tuning method for multimodal tasks based on singular value decomposition-enhanced routing functions, which aims to optimize the alignment accuracy of visual and language features in multimodal tasks, reduce computational complexity, and improve model performance.
[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is: a fine-tuning method for multimodal tasks based on singular value decomposition to enhance routing functions, comprising the following steps:
[0007] Input image sequences and text sequences, extract features from the input image sequences and text sequences respectively through pre-trained visual encoders and language models to obtain additional visual features and language hidden features that need to be integrated;
[0008] The language hidden features and additional visual features are mapped from the high-dimensional space to the low-rank space using the downprojection matrix of the PEFT method with a low-rank bottleneck, generating language features and visual features.
[0009] In the low-rank space, the downsampled language features are preprocessed by converting them into floating-point tensors. Singular value decomposition is then performed on the language features. To preserve core information and reduce dimensionality, the first R largest singular values are truncated. The energy proportion of the first R singular values is chosen to maximize their relative importance. 95%, compress the dimension to R, and use the truncated left singular matrix, singular value matrix and right singular matrix to perform efficient reconstruction to obtain low-rank language features, ensure that the shape is consistent with the shape of the input language features, and verify the output shape;
[0010] The low-rank language features and visual features are cross-modal aligned using a routing function;
[0011] The low-rank language features after cross-modal alignment are restored to their original dimensions by upprojection, and finally residually connected with the input language hidden state to obtain the output features.
[0012] Preferably, it also includes a shared caching mechanism for storing the singular value decomposition results of tensors with the same shape. After generating the language features and visual features, the shared caching mechanism checks whether there are valid results. If so, they can be reused directly. Otherwise, singular value decomposition is performed and the cache is updated. The cache is designed to be updated every N steps to balance efficiency and data freshness.
[0013] Preferably, additional visual features to be integrated are extracted from the pre-trained visual encoder and the Transformer block in the language model, respectively. and language hidden features The additional visual features Dimensions are ( ,d), where For the additional visual features The sequence length, the language hidden features Dimensions are ( ,d), where For the language hidden features The sequence length is d, and d is the feature dimension of the pre-trained model.
[0014] Using a PEFT method with a low-rank bottleneck through the downprojection matrix of LoRA or Adapter Hidden language features and the additional visual features Mapping from high-dimensional d to low-dimensional r, low-rank bottleneck dimension Generate language features and visual features The language features The shape is ( The visual features (r) The shape is ( ,r), where r represents the feature length and the low-rank dimension or level, and the low-rank dimension r can be adjusted according to task requirements;
[0015] The downsampled language features Preprocessing is performed to extract the language features. Convert to a floating-point tensor and apply the language features. Perform singular value decomposition ,in ( Let be a left singular matrix, and let the column vectors be an orthogonal basis. For singular value matrices, the diagonal elements Sort in descending order Let be a right singular matrix, and let the column vectors be an orthogonal basis of the feature space. To preserve the core information of the features and reduce the dimensionality, the first R largest singular values are truncated, where By choosing R, the energy proportion of the first R singular values is... 95% efficiency can be achieved by compressing the dimensionality to R while preserving the main feature information. Low-rank language features can be obtained through efficient reconstruction using the truncated left singular matrix, singular value matrix, and right singular matrix. ,in , , Let be the truncated left singular matrix, singular value matrix, and right singular matrix, respectively. According to the low-rank approximation error bound theory, the reconstruction error after truncation satisfies: Ensure that the shape matches the input language features. Consistent shape , ,r), and verify the output shape;
[0016] The routing function applies to the low-rank language features. and visual features After cross-modal alignment, the low-rank language features Projection matrix The original dimensions are restored, and finally compared with the input language hidden features. Perform residual connections to obtain output features. .
[0017] Furthermore, the routing function employs element-wise multiplication. ,in By broadcasting the visual features The resulting shape is ( , Tensors of r) are matched with the low-rank language features. The shape, the visual features After shape adaptation, it is combined with the low-rank language features. Element-wise multiplication, through Implement a broadcast mechanism to enhance the dimensions of visually relevant features. Each element is used as a scaling factor to adjust the low-rank language features. The size of the corresponding element, if If the value is larger, then the low-rank language features are enhanced. If the value is small, then the low-rank language features are suppressed. .
[0018] Furthermore, the routing function employs element-wise addition. ,in By broadcasting the visual features The resulting shape is ( , The tensor of r) will The features are directly added to the low-rank language features Above, making the low-rank language features Visual features in low-rank space By bringing language and visual features closer together, the spatial distance between them is reduced, thus promoting feature integration.
[0019] Furthermore, the routing function employs matrix multiplication. ,in The visual features are Remodeled into ( , A matrix of type r). As a transformation matrix, the low-rank language features are reshaped. To match the visual features The semantics of the low-rank language features Each line through Rescaling indirectly affects the visual features. Information integration, by reshaping the visual features for A matrix that adjusts the row scale of language features, where To achieve structured interaction between modalities, through the visual features Linear transformation to adjust the low-rank language features The semantic distribution.
[0020] The technical advantages of this invention are as follows: By applying singular value decomposition to language features before the routing function, its low-rank dominant patterns are extracted, enhancing the alignment accuracy of visual and linguistic features, eliminating the interference of high-dimensional noise, and maintaining computational efficiency and model stability. Routing calculations using the reconstructed tensors allow for better extraction and alignment of key information in the features, thereby improving the accuracy and effectiveness of feature alignment. It is applicable to Visual Language (VL) tasks such as visual question answering and image caption generation, and can significantly improve model performance. Attached Figure Description
[0021] Figure 1 A schematic diagram illustrating the technical path for introducing routing functionality in the efficient fine-tuning of visual-language parameters in existing technologies;
[0022] Figure 2 This is a micro-flow diagram illustrating the multimodal task based on singular value decomposition to enhance the routing function according to the present invention. Detailed Implementation
[0023] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings, so that the technical solution of the present invention can be more easily understood and mastered.
[0024] Example 1:
[0025] Definitions of some abbreviations and key terms:
[0026] PEFT, or Parameter-Efficient Fine-Tuning, is a fine-tuning method for pre-trained models that achieves efficient adaptation to downstream tasks by updating a small number of parameters.
[0027] VL, or Vision-Language, refers to multimodal tasks that combine visual and linguistic modalities, such as visual question answering and image description generation.
[0028] An adapter is a tool used to adapt a pre-trained model to a downstream task by inserting small neural network modules (adapter modules) into each layer of the model.
[0029] LoRA, or Low-Rand-Adaptation, is a PEFT method that achieves efficient fine-tuning by introducing a low-rank bottleneck into the model.
[0030] SVD: Singular Value Decomposition, a matrix factorization technique used to decompose a matrix into the product of three matrices. It is widely used in feature reduction and signal processing.
[0031] Routing functions are linear operations introduced in the efficient fine-tuning of vision-language parameters, designed to enhance feature alignment between vision and language modalities in low-rank bottlenecks.
[0032] like Figure 2 As shown, a fine-tuning method for multimodal tasks based on singular value decomposition to enhance routing functions includes the following steps:
[0033] Step 1: Input image and text sequences. Extract features from the input image and text sequences using a pre-trained visual encoder (e.g., ViT-B / 16, pre-trained on ImageNet-21k) and a Transformer block in a language model (e.g., GPT-2), respectively, to obtain the additional visual features that need to be integrated. (Typically, this refers to ViT's [CLS] features) and language hiding features The additional visual features Dimensions are ( ,d), where For the additional visual features sequence length (usually) Because [CLS] features are used or multiple token features are compressed into a single feature vector through average pooling; the language hidden features Dimensions are ( ,d), where d is the sequence length (number of tokens) of the hidden features of the language, and d is the feature dimension of the pre-trained model.
[0034] Step 2: Use the PEFT method with a low-rank bottleneck to project the lower projection matrix through LoRA or the Adapter. Hidden language features and the additional visual features Mapping from high-dimensional d to low-dimensional r, low-rank bottleneck dimension Generate language features and visual features The language features The shape is ( The visual features (r) The shape is ( ,r), where `r` represents the feature length (i.e., the number of tokens) and the low-rank dimension or level, where `r` is the low-rank dimension (e.g., `r=64 / 128`). The low-rank dimension `r` can be adjusted according to task requirements (e.g., `r=64` or `r=128`) to balance performance and efficiency. In practical applications, a single token / global feature vector is used to represent the visual content. In the case of multiple tokens, average pooling is used to simplify it to a single feature vector. This step maps the high-dimensional features to the low-rank space through linear projection, reducing the subsequent computational complexity from that of full-rank systems. Reduce to (Reducing by approximately 90%), while retaining key information about the features, laying the foundation for the low-rank approximation of SVD.
[0035] Step 3: In the low-rank space, the downsampled language features are processed... Preprocessing is performed to extract the language features. Convert to floating-point tensors (FP16 compatible) and perform low-dimensional processing of the language features. Perform singular value decomposition (SVD). ,in ( Let be a left singular matrix, and let the column vectors be an orthogonal basis; For singular value matrices, the diagonal elements Sort in descending order; Let be a right singular matrix, and let the column vectors be an orthogonal basis of the feature space. To preserve the core information of the features and reduce dimensionality, the first R largest singular values are truncated, where Low-rank language features are obtained by efficiently reconstructing the truncated left singular matrix, singular value matrix, and right singular matrix. ,in , (Only R singular values are retained) Let be the truncated left singular matrix, singular value matrix, and right singular matrix, respectively. According to the low-rank approximation error bound theory, the reconstruction error after truncation satisfies: By choosing R, the energy proportion of the first R singular values is... 95% (e.g.) This method can compress the dimension to R while preserving the main feature information, ensuring that the shape matches the input language features. Consistent shape , , r), and verify the output shape.
[0036] Step 3.5: To avoid redundant calculations, a shared cache mechanism (SharedSVDCacher) is also included for storing the singular value decomposition results of tensors with the same shape. The cache key is the tensor shape tuple (...). , r), with a value of ( , , The design updates the cache every 10 steps to balance efficiency and data freshness. Specifically, after the SVD method is called 10 times, the SVD decomposition results in the cache are updated once every 10 forward propagation iterations. After generating the language features and visual features, the shared cache mechanism checks for existing and valid results (not expired and with matching shapes). If a result exists, it can be reused directly; otherwise, SVD decomposition is performed and the cache is updated. Actual testing shows that the shared cache mechanism can reduce the time spent on repeated decomposition by approximately 70%, decreasing the average SVD processing time from 52ms to 15ms.
[0037] Step 4: Use routing functions to process the low-rank language features. and visual features Cross-modal alignment is performed, and the routing function is designed to optimize the performance of standard PEFT methods (such as LoRA or Adapter) on VL tasks without adding additional trainable parameters. Cross-modal alignment is improved by introducing guidance on visual features in the low-rank space. All operations are performed in the low-rank space after SVD to ensure that the energy of interactive features is concentrated in the subspace.
[0038] The routing function can choose to use one of the three linear routing functions:
[0039] (1) Element-by-element multiplication, ,in By broadcasting the visual features The resulting shape is ( , Tensors of r) are matched with the low-rank language features. The shape, the visual features After shape adaptation, it is combined with the low-rank language features. Element-wise multiplication, specifically through Implement a broadcast mechanism to enhance the dimensions of visually relevant features. Each element is used as a scaling factor to adjust the low-rank language features. The size of the corresponding element, if If the value is larger, then the low-rank language features are enhanced. If the value is small, then the low-rank language features are suppressed. .
[0040] By analyzing the visual features Selectively enhance or modify the low-rank language features Weakening, promoting the low-rank language features With visual features Semantic alignment of content. For example, if an image contains the feature of "red", It may amplify language tags associated with "red".
[0041] (2) Element addition, ,in By broadcasting the visual features The resulting shape is ( , The tensor of r) will The features are directly added to the low-rank language features Above, making the low-rank language features Visual features in low-rank space By bringing language and visual features closer together, the spatial distance between them is reduced, thus promoting feature integration.
[0042] By "translating" the low-rank language features To visual features Nearby, force the low-rank language features Incorporating visual features Information. For example, if The low-rank linguistic features represent the global semantics of an image (e.g., "dog"). Each tag will be adjusted to include more information related to "dog".
[0043] (3) Matrix multiplication, ,in The visual features are Remodeled into ( , A matrix of type r). As a transformation matrix, the low-rank language features are reshaped. To match the visual features The semantics of the low-rank language features Each line through Rescaling indirectly affects the visual features. Information integration, by reshaping the visual features for A matrix that adjusts the row scale of language features, where To achieve structured interaction between modalities, through the visual features Linear transformation to adjust the low-rank language features The semantic distribution.
[0044] The low-rank language features after the routing function operation described above By using nonlinear activation (such as ReLU) and upprojection sampling ( The low-rank language features are restored to their original dimension d. Central and visual features The relevant parts are enhanced and passed on to subsequent steps, further strengthening cross-modal interaction through attention mechanisms.
[0045] Step 5: The routing function applies the low-rank language features. and visual features After performing cross-modal alignment, the low-rank language features should be remembered. and visual features Both are achieved through the lower projection matrix. From the additional visual features and language hidden features Derivatives, you can learn about projection matrices. The parameters are used to select the original language hidden features. and additional visual features Specific dimensions. The aligned low-rank language features. Projection matrix The original dimensions are restored, and finally compared with the input language hidden features. Perform residual connections to obtain output features. The residual connectivity-preserving model enables efficient parameter fine-tuning (update). and While integrating visual and linguistic information (including routing parameters), it effectively enhances the performance of multimodal tasks.
[0046] The existing technical approach for introducing routing functionality in efficient fine-tuning of visual-language parameters is as follows: Figure 1 As shown, it includes steps 1, 2, 4 and 5 similar to those in this embodiment. In existing methods, the effectiveness of the routing function depends on the quality and representation of the input features. Language features usually contain high-dimensional noise or redundant information. Directly projecting them into a low-dimensional space may lead to information loss, affecting the alignment effect of the routing function. The routing function directly operates on the projected features and lacks a preprocessing mechanism for language features, which cannot effectively extract the dominant pattern and limits the accuracy of modality alignment. When the batch size is large or the sequence length is long, the complexity of language features may lead to a decrease in computational efficiency and model stability.
[0047] This application introduces singular value decomposition and optimized routing mechanisms into the efficient parameter fine-tuning of multimodal tasks, and reduces the time consumption of repeated SVD decomposition by 70% by using a shared cache mechanism; furthermore, it preserves the core feature information by truncating the first R largest singular values, and combines linear routing functions (element-wise multiplication, element-wise addition, matrix multiplication) to enhance the cross-modal alignment accuracy of visual and linguistic features, avoiding the alignment deficiency problem of existing traditional methods; and it uses residual connections to preserve the high-frequency information of the original features, preventing feature distortion caused by routing operations.
[0048] As shown in Tables 1 and 2, the experimental verification results demonstrate that the method of this application improves the BLEU-4, METEOR, and CIDEr metrics in COCO Captioning by more than 25%, while only increasing the inference time by 5%. This achieves comprehensive optimization of computational efficiency, feature quality, and multimodal task performance, providing key technical support for the efficient deployment of visual-language models.
[0049] Table 1 compares the experimental results of using SVD to optimize routing in the Adapter method with those of using only the routing function (+ indicates the improvement):
[0050]
[0051] Table 2 compares experimental results of using SVD to optimize routing in the LoRa method with using routing functions only (+ indicates the improvement):
[0052]
Claims
1. A fine-tuning method for multimodal tasks based on singular value decomposition to enhance routing functions, characterized in that: Includes the following steps: Input image sequences and text sequences, extract features from the input image sequences and text sequences respectively through pre-trained visual encoders and language models to obtain additional visual features and language hidden features that need to be integrated; The language hidden features and additional visual features are mapped from the high-dimensional space to the low-rank space using the downprojection matrix of the PEFT method with a low-rank bottleneck, generating language features and visual features. In the low-rank space, the downsampled language features are preprocessed by converting them into floating-point tensors. Singular value decomposition is then performed on the language features. To preserve core information and reduce dimensionality, the first R largest singular values are truncated. The energy proportion of the first R singular values is chosen to maximize their relative importance. 95%, compress the dimension to R, and use the truncated left singular matrix, singular value matrix and right singular matrix to perform efficient reconstruction to obtain low-rank language features, ensure that the shape is consistent with the shape of the input language features, and verify the output shape; The low-rank language features and visual features are cross-modal aligned using a routing function; The low-rank language features after cross-modal alignment are restored to their original dimensions by upprojection, and finally residually connected with the input language hidden state to obtain the output features.
2. The fine-tuning method for multimodal tasks based on singular value decomposition to enhance routing functions according to claim 1, characterized in that: It also includes a shared caching mechanism for storing singular value decomposition results of tensors with the same shape. After generating the language features and visual features, the shared caching mechanism checks whether there are valid results. If so, they can be reused directly. Otherwise, singular value decomposition is performed and the cache is updated. The cache is designed to be updated every N steps to balance efficiency and data freshness.
3. A fine-tuning method for multimodal tasks based on singular value decomposition to enhance routing functions, as described in claim 1 or 2, characterized in that: Additional visual features that need to be integrated are extracted from the pre-trained visual encoder and the Transformer block in the language model, respectively. and language hidden features The additional visual features Dimensions are ( ,d), where For the additional visual features The sequence length, the language hidden features Dimensions are ( ,d), where For the language hidden features The sequence length is d, and d is the feature dimension of the pre-trained model. Using a PEFT method with a low-rank bottleneck through the downprojection matrix of LoRA or Adapter Hidden language features and the additional visual features Mapping from high-dimensional d to low-dimensional r, low-rank bottleneck dimension Generate language features and visual features The language features The shape is ( The visual features (r) The shape is ( ,r), where r represents the feature length and the low-rank dimension or level, and the low-rank dimension r can be adjusted according to task requirements; The downsampled language features Preprocessing is performed to extract the language features. Convert to a floating-point tensor and apply the language features. Perform singular value decomposition ,in ( Let be a left singular matrix, and let the column vectors be an orthogonal basis. For singular value matrices, the diagonal elements Sort in descending order Let be a right singular matrix, and let the column vectors be an orthogonal basis of the feature space. To preserve the core information of the features and reduce the dimensionality, the first R largest singular values are truncated, where By choosing R, the energy proportion of the first R singular values is... 95% efficiency can be achieved by compressing the dimensionality to R while preserving the main feature information. Low-rank language features can be obtained through efficient reconstruction using the truncated left singular matrix, singular value matrix, and right singular matrix. ,in , , Let be the truncated left singular matrix, singular value matrix, and right singular matrix, respectively. According to the low-rank approximation error bound theory, the reconstruction error after truncation satisfies: Ensure that the shape matches the input language features. Consistent shape , ,r), and verify the output shape; The routing function applies to the low-rank language features. and visual features After cross-modal alignment, the low-rank language features Projection matrix The original dimensions are restored, and finally compared with the input language hidden features. Perform residual connections to obtain output features. .
4. The fine-tuning method for multimodal tasks based on singular value decomposition to enhance routing functions according to claim 3, characterized in that: The routing function uses element-wise multiplication. ,in By broadcasting the visual features The resulting shape is ( , Tensors of r) are matched with the low-rank language features. The shape, the visual features After shape adaptation, it is combined with the low-rank language features. Element-wise multiplication, through Implement a broadcast mechanism to enhance the dimensions of visually relevant features. Each element is used as a scaling factor to adjust the low-rank language features. The size of the corresponding element, if If the value is larger, then the low-rank language features are enhanced. If the value is small, then the low-rank language features are suppressed. .
5. The fine-tuning method for multimodal tasks based on singular value decomposition to enhance routing functions according to claim 3, characterized in that: The routing function uses element-wise addition. ,in By broadcasting the visual features The resulting shape is ( , The tensor of r) will The features are directly added to the low-rank language features Above, making the low-rank language features Visual features in low-rank space By bringing language and visual features closer together, the spatial distance between them is reduced, thus promoting feature integration.
6. The fine-tuning method for multimodal tasks based on singular value decomposition to enhance routing functions according to claim 3, characterized in that: The routing function uses matrix multiplication. ,in The visual features are Remodeled into ( , A matrix of type r). As a transformation matrix, the low-rank language features are reshaped. To match the visual features The semantics of the low-rank language features Each line through Rescaling indirectly affects the visual features. Information integration, by reshaping the visual features for A matrix that adjusts the row scale of language features, where To achieve structured interaction between modalities, through the visual features Linear transformation to adjust the low-rank language features The semantic distribution.
Citation Information
Cited By
Unmanned vehicle vision Transform model compression method and system
CN121708558A
Efficient video understanding method based on low-rank key tensor residual decomposition
CN122137972A