Visual Question Answering Optimization Method and System Based on Hierarchical Multimodal Fine Tuning
By adopting a hierarchical multimodal fine adjustment method in KI-VQA task, the CLIP model and multi-head attention mechanism are used for cross-modal interaction fusion, and fine-tuning is carried out through low-rank adaptation technology, the problem of semantic differences between images and text is solved, and the performance of VQA task is significantly improved.
Patent Information
- Application Number
- CN202510472099.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing KI-VQA methods have difficulties in processing semantic differences between images and texts and achieving effective knowledge transfer, resulting in difficulty in information fusion and low accuracy of question and answers.
The visual question-and-answer optimization method based on hierarchical multimodal fine adjustment is adopted to extract visual and text features by pre-training the CLIP model, and cross-modal interaction fusion is used to use adaptive weighting, linear projection and multi-head attention mechanisms to fine-tune the encoder, combined with low-rank adaptation technology.
It significantly improves the accuracy and effect of cross-modal learning, solves the problem of inconsistent matching of visual features and text semantics, and achieves more refined semantic alignment and modal fusion, thereby improving the performance of VQA tasks.
Smart Images

Figure CN120011547B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of visual question answering, and particularly relates to a visual question answering optimization method and system based on hierarchical multi-modal fine adjustment. Background Art
[0002] The Knowledge-Intensive Visual Question Answering (KI-VQA) task is a variant of visual question answering that combines image information with external knowledge. Its core lies in the integration of images and external knowledge bases to provide users with more accurate and rich answers. This task has important practical significance in fields such as e-commerce and education, especially in scenarios that require the combination of visual information and external knowledge bases. The challenge of the KI-VQA task is that the answers to questions usually cannot be obtained solely from the image itself but require external knowledge sources. For example, on an e-commerce website, users can take a photo of a product and ask relevant questions, and the system needs to combine the product image and detailed product information (such as brand, model, function, etc.) to answer the questions; in the field of education, students can ask questions about the illustrations in textbooks, and the system needs to combine the illustration content and relevant knowledge points to answer the questions.
[0003] Currently, most KI-VQA methods assume that external knowledge can be obtained from structured knowledge bases. First, features are extracted from images and text through image and text encoders; then, the image features and text features are fused, and multi-modal features are used for retrieval and question answering generation. These methods need to effectively fuse image and text information and retrieve content related to the question from the external knowledge base to improve the accuracy of question answering. However, existing dense retrieval methods usually face the problem of heterogeneity between multi-modal (image and text) inputs and single-modal (text) outputs, and the semantic differences between images and texts make it difficult to fuse information between modalities.
[0004] How to effectively capture the correlation between images and texts and retrieve information related to the question from external knowledge to improve the accuracy and richness of question answering is still an unsolved problem. To address these issues, existing solutions include using large-scale text corpora as external knowledge sources and retrieving questions and images through dense retrieval models. However, how to handle semantic differences between different modalities and achieve effective knowledge transfer remains a research difficulty in the KI-VQA field. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a visual question answering optimization method and system based on hierarchical multi-modal fine adjustment, which improves the performance of the visual question answering task by introducing refined visual and text feature interactions.
[0006] The technical solution provided by the present invention is specifically as follows:
[0007] An optimization method for visual question answering based on hierarchical multi-modal fine-tuning, comprising the steps of:
[0008] S1. Obtain text-image pairs for the visual question answering task, including images and corresponding texts;
[0009] S2. Use the visual encoder of the pre-trained CLIP model to extract the visual features of the image, and form a visual feature set from the hierarchical features extracted by the visual encoder;
[0010] S3. Use the text encoder of the pre-trained CLIP model to extract the text features of the text, and form a text feature set from the hierarchical features extracted by the text encoder;
[0011] S4. Perform adaptive weighted adjustment on the text features of each layer in the text feature set;
[0012] S5. Concatenate the weighted-adjusted text features of each layer into a vector and perform linear projection;
[0013] S6. Concatenate the visual features of each layer in the visual feature set into a vector, and perform cross-modal interaction fusion on it and the text features after linear projection through the multi-head attention mechanism;
[0014] S7. Fuse the feature vector obtained by cross-modal interaction fusion with the vector obtained by concatenating the visual features of each layer in the visual feature set to obtain the semantic visual perception fusion feature;
[0015] S8. Use the semantic visual perception fusion feature for the visual question answering task.
[0016] Further, perform multi-layer group-by-group stage fine-tuning on the visual encoder and text encoder through the low-rank adaptation technology, including the steps of:
[0017] Model layer group division: Divide the network levels of the visual encoder and text encoder into the front layer, middle layer, and rear layer in sequence according to the depth;
[0018] First-stage fine-tuning: Add LoRA adapters to the rear layers of the visual encoder and text encoder respectively, keep the pre-trained parameters of the visual encoder and text encoder unchanged, and only open the LoRA adapters inserted in the rear layers for training. Only update the parameters of the LoRA adapters during the training process;
[0019] Second-stage fine-tuning: LoRA adapters are added to the middle layers of the visual encoder and the text encoder respectively. The pre-trained parameters of the visual encoder and the text encoder are kept unchanged, and only the LoRA adapters inserted in the later layers and the middle layers are opened for training. During the training process, only the parameters of the LoRA adapters are updated.
[0020] Preferably, both the visual encoder and the text encoder of the pre-trained CLIP model include 12 Transformer layers. In the model layer group division, the 1st to 4th Transformer layers are regarded as the front layers, the 5th to 8th Transformer layers are regarded as the middle layers, and the 8th to 12th Transformer layers and the projection layer are regarded as the later layers.
[0021] Furthermore, it further includes third-stage fine-tuning: LoRA adapters are added to the front layers of the visual encoder and the text encoder respectively. The pre-trained parameters of the visual encoder and the text encoder are kept unchanged, and only the LoRA adapters inserted in the later layers, the middle layers and the front layers are opened for training. During the training process, only the parameters of the LoRA adapters are updated.
[0022] Preferably, during the third-stage fine-tuning process, first keep the parameters of the visual encoder, the text encoder and the middle and later layer LoRA adapters unchanged, and only open the LoRA adapters inserted in the front layer for low learning rate training; then open the middle and later layer LoRA adapters together for a few rounds of joint fine-tuning with an extremely small learning rate to make the parameters of all LoRA adapters reach the overall coordination optimum.
[0023] Furthermore, in step S4, the adaptive weighted adjustment is expressed as:
[0024] ,
[0025] where, represents the text feature of the i th layer after weighted adjustment, represents the corresponding element multiplication operation, represents the corresponding element addition operation, represents the adaptive weight W corresponding to the th i layer weight vector in represents the text feature extracted by the i th layer of the text encoder.
[0026] Furthermore, in step S5, the linear projection is expressed as:
[0027]
[0028] where, is the text feature after linear projection, concat represents vector concatenation, represents matrix multiplication operation, is the linear projection weight.
[0029] Further, in step S6, the text feature after linear projection is used as the key vector K and the value vector V, and the vector formed by concatenating the visual features of each layer in the visual feature set is used as the query vector Q, and the cross-modal interaction fusion feature is obtained through the multi-head attention mechanism.
[0030] Further, in step S7, residual connection is adopted, and the cross-modal interaction fusion feature output by the multi-head attention mechanism is added to the vector formed by concatenating the visual features of each layer in the visual feature set element by element to fuse and obtain the semantic visual perception fusion feature:
[0031] ,
[0032] wherein, f is the semantic visual perception fusion feature, is the cross-modal interaction fusion feature, represents the vector formed by concatenating the visual features of each layer in the visual feature set, represents the element-wise addition operation.
[0033] A visual question answering optimization system based on the above method, including a visual feature extraction module, a text feature extraction module, a weighted projection module, and a cross-modal bridging module;
[0034] The visual feature extraction module is used to extract the visual features of the image in the visual question answering task by using the visual encoder of the pre-trained CLIP model, and form a visual feature set with the hierarchical features extracted by the visual encoder;
[0035] The text feature extraction module is used to extract the text features of the text in the visual question answering task by using the text encoder of the pre-trained CLIP model, and form a text feature set with the hierarchical features extracted by the text encoder;
[0036] The weighted projection module is used to adaptively weight and adjust the text features of each layer in the text feature set, and concatenate the weighted and adjusted text features of each layer into a vector for linear projection;
[0037] The cross-modal bridging module is used to perform cross-modal fusion on the text feature after linear projection and the vector formed by concatenating the visual features of each layer in the visual feature set through the multi-head attention mechanism and residual connection to obtain the semantic visual perception fusion feature for the visual question answering task;
[0038] It also includes a LoRA fine-tuning module for performing multi-layer group-by-group stage fine-tuning on the visual encoder and the text encoder through low-rank adaptation technology.
[0039] Compared with the prior art, the present invention has at least the following beneficial effects:
[0040] Through the innovative adaptive cross-modal bridging and multi-layer group-by-group stage low-rank adaptation method, the present invention significantly improves the accuracy and effect of cross-modal learning. The method introduces text information with multi-level semantic perception into visual features, enabling the visual features to be more finely adjusted and optimized, thereby solving the problem of inconsistent matching between visual features and text semantics. Through multi-layer group-by-group stage low-rank adaptation, the accumulation of perceptual errors is effectively avoided, and a layer-by-layer adaptation mechanism is provided, enabling visual features and text features to be more accurately aligned in the multi-level learning process. The present invention provides new ideas and methods for visual-text feature fusion in cross-modal tasks, and can better capture high-level interactions in the text and visual fields, thereby improving the performance of VQA tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention.
[0042] Figure 1 It is a schematic diagram of the framework of feature adaptive cross-modal bridging provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] Embodiment 1
[0045] This embodiment provides a visual question answering optimization method based on hierarchical multi-modal fine adjustment, as Figure 1 shown, aiming to improve the performance of visual question answering tasks by introducing refined visual and text feature interactions.
[0046] The method mainly includes the following steps:
[0047] 1. Initial feature extraction
[0048] For the image given in the visual question answering task I and textT (Problem) Extract corresponding visual features and text features through the CLIP (Contrastive Language-Image Pre-training) model respectively.
[0049] The CLIP model is a multi-modal pre-training model designed to achieve matching and association between images and texts through contrastive learning. Its core idea is to encode images and texts into vectors and optimize the similarity of these vectors through a cosine similarity loss function, thereby realizing cross-modal understanding and interaction. The CLIP model consists of two main parts: a visual encoder and a text encoder. The visual encoder is responsible for converting the input image into a high-dimensional vector representation, while the text encoder converts the text information into a vector representation. Both are jointly optimized through contrastive learning to achieve semantic alignment between images and texts.
[0050] The visual encoder of CLIP adopts the Transformer (ViT) architecture. In the visual encoder, the input image is segmented into multiple patches, and features are extracted layer by layer through multiple Transformer layers. The output of each layer is a further abstraction and enhancement of the features of the previous layer. Finally, a high-dimensional feature representation of the image is generated through the stacking of multiple Transformer layers. Similarly, the text encoder of CLIP is based on the Transformer architecture. Inside the text encoder, the text is tokenized and converted into embedding vectors, and these embedding vectors are processed through multiple Transformer layers. Each layer performs linear and non-linear transformations on the input embedding vectors, thereby gradually extracting the semantic features of the text. Finally, the text encoder outputs a high-dimensional feature vector, which represents the semantic meaning of the input text.
[0051] In this embodiment, when extracting features from an image I different from the general method that only uses the output features of the final layer, instead, the features of each level extracted by the CLIP visual encoder are combined into a visual feature set , m represents the number of levels of the image encoder, that is, it represents the visual features extracted by the i th layer of the CLIP visual encoder. Similarly, for the text T feature extraction, the features of each level extracted by the CLIP text encoder are combined into a text feature set , n represents the number of levels of the text encoder, that is, it represents the text features extracted by the i th layer of the CLIP text encoder.
[0052] 2. Feature Adaptive Cross-Modal Bridging
[0053] (1)Adaptive Weighting of Text Features
[0054] For the text features of each layer in the text feature set weighted adjustment is performed through sample-independent adaptive weights W That is:
[0055] ,
[0056] where represents the text features of the i th layer after weighted adjustment, represents the element-wise multiplication operation, represents the element-wise addition operation, represents the adaptive weight W corresponding to the th layer weight vector in i , that is . The value of the adaptive weight W is randomly initialized at the beginning of training to ensure that it will not have too much impact on the model output in the initial stage of training, and is iteratively updated during the backpropagation of the training process.
[0057] (2)Text Feature Concatenation and Projection
[0058] The text features of each layer after weighted adjustment are concatenated into a vector and multiplied by the linear projection weight matrix to be mapped into the visual embedding space:
[0059]
[0060] where is the text feature after concatenation and projection, which can be effectively docked with visual features; concat represents vector concatenation, represents matrix multiplication operation; is the linear projection weight, which is a parameter matrix obtained through training and can map the concatenated text feature vector into the visual embedding space.
[0061] The iterative update during training can be expressed as:
[0062]
[0063] where is the learning rate, is the loss function defined by the downstream task (such as visual question answering).
[0064] (3)Visual Feature Concatenation
[0065] For the convenience of the interactive integration of visual and text features, similarly, the visual feature sets in each layer are concatenated into a vector: .
[0066] (4)Cross-modal multi-head attention mechanism
[0067] The cross-modal fusion of text features and visual features is achieved through the multi-head attention mechanism. The multi-head attention mechanism is a key component in the Transformer architecture and is widely used in natural language processing tasks as well as image processing tasks. Its core idea is to capture different features of the input through multiple different attention heads, thereby improving the performance of the model.
[0068] Specifically, the multi-head attention mechanism generally proceeds through the following steps:
[0069] 1) Linear mapping
[0070] The input query vector Q, key vector K, and value vector V are respectively projected into h different subspaces (i.e., different attention heads). For each attention head i, the following are calculated respectively:
[0071] ,
[0072] where , and are the linear projection matrices of attention head i.
[0073] 2) Parallel calculation of multiple attentions
[0074] For each head, assuming the dimension of each attention head is d, the scaled dot-product attention mechanism is used to calculate the attention output respectively:
[0075]
[0076] 3) Concatenate the outputs of the attention heads. The outputs of all h attention heads are concatenated to form a large vector:
[0077]
[0078] 4) Linear transformation
[0079] The concatenated result is projected through another linear transformation to obtain the final multi-head attention output:
[0080]
[0081] Projection matrices , and and after the output of the splicing head , are all randomly initialized at first and then gradually learned through training.
[0082] In this embodiment, the text features after splicing and projection are used as the key vector K and the value vector V, and the multi-layer visual features after splicing are used as the query vector Q, and the cross-modal interaction fusion features are obtained through the multi-head attention mechanism , that is . Finally, using residual connection, the cross-modal interaction fusion features output by the multi-head attention mechanism and the multi-layer visual features are added element by element again to further optimize the representation of the visual features, and the final semantic visual perception fusion features are obtained f , that is .
[0083] The number of attention heads is an important hyperparameter in the multi-head attention mechanism. Its setting directly affects the expressive power and computational complexity of the model, and is also restricted by the scale of the dataset and computing resources. In this embodiment, the number of heads is gradually adjusted through experiments, and the changes in the model performance are observed. It is found that setting the number of attention heads to 8 is more appropriate.
[0084] 3. Finally, the obtained semantic visual perception fusion features are applied to the downstream visual question answering task.
[0085] The semantic difference between images and texts is one of the core challenges in the VQA task. Traditional feature fusion methods (such as direct embedding, linear fusion, etc.) often cannot fully capture the complex interaction relationship between the two, resulting in limited model performance. In this embodiment, by introducing text information with multi-level semantic perception into the visual features, the visual features are more finely adjusted and optimized at multiple levels, realizing effective semantic alignment and modality fusion, thus solving the problem of inconsistent matching between visual features and text semantics, and being able to better capture the high-level interaction in the language and vision fields, thereby improving the performance of the VQA task.
[0086] In some embodiments, the pre-trained CLIP model is fine-tuned through the multi-layer group-by-stage low-rank adaptation technology (Low-Rank Adaptation, LoRA) to improve the adaptability or generalization of the pre-trained model to specific application fields, ensure the consistency of visual features and text features at multiple levels, and further effectively avoid the accumulation of perception errors, so that the visual features and text features can be more matched in subsequent tasks.
[0087] The multi-layer group-by-stage LoRA fine-tuning mainly includes the following links:
[0088] (1)Model layer group division
[0089] The visual encoder hierarchy usually consists of an image embedding layer (including patch partitioning and positional encoding) and several Transformer blocks. For example, for ViT-B / 16 with 12 Transformer layers, it can be divided as follows:
[0090] Front layers: Refer to several initial layers of the model (such as the 1st - 4th Transformer layers). These layers focus on extracting low-level visual features, such as local patterns like edges, textures, and colors, and they provide a general image representation.
[0091] Middle layers: Refer to several middle layers (such as the 5th - 8th Transformer layers). These layers gradually capture higher-level feature patterns (such as object parts, complex shapes), combining low-level features into intermediate representations.
[0092] Back layers: Refer to the last several layers (such as the 9th - 12th Transformer layers) and the final image feature projection layer. The back layers focus on extracting global semantic information, and the learned features are more discriminative and task-relevant.
[0093] Similarly, the hierarchy of the text encoder can be divided as follows:
[0094] Front layers: Such as the 1st - 4th Transformer layers in the Embedding hierarchy. These layers mainly capture word-level and local syntactic-level information (similar to the model learning basic lexical and phrasal features).
[0095] Middle layers: Such as the 5th - 8th Transformer layers. These layers gradually model longer-range dependencies and semantic combinations (representation from phrases to sentences).
[0096] Back layers: The 9th - 12th Transformer layers and the final text projection layer. The back layers output the semantic embedding of the entire sentence, understanding global semantics and context.
[0097] Dividing the visual encoder and text encoder into layer groups according to depth reflects the gradual abstraction from basic features to high-level semantics. Dividing layer groups helps to adopt different strategies during subsequent fine-tuning: preferentially adjusting the back layers to adapt to the semantics of new tasks, and gradually adjusting the middle and front layers, so as to retain the original general knowledge of the model to the greatest extent.
[0098] (2)Stage-by-stage optimization strategy
[0099] LoRA efficiently adapts to downstream tasks by adding parameters in the form of low-rank matrices to the existing model layers. During the fine-tuning process, the original weights of the model are not directly updated. Instead, trainable low-rank matrices (LoRA adapters) are injected into each specified layer. A and B , satisfying ( W is the original weight of the model, is the weight change), so that the pre-trained weights remain frozen and the model can be adjusted with very few new parameters.
[0100] To make the low-rank adapter fully effective, a stage-by-stage LoRA optimization strategy is designed:
[0101] Back-layer optimization: In the first stage, LoRA adapters are introduced into the back layers of the visual and text encoders, and the parameters of this part of the adapters are mainly optimized. The back layers directly determine the final embedding representations of images and texts, and their outputs are directly related to the contrast loss. Therefore, prior adjustment can quickly respond to the alignment requirements of new tasks. In this stage, a relatively high rank, such as 8 or 16, is set for the back-layer LoRA to ensure sufficient expressive power to capture new semantic mapping relationships. At the same time, an appropriate scaling factor is combined to stabilize the impact of LoRA updates on the original weights. The optimization goal is to correct the image-text embedding space for new data: for example, when new categories or concepts are introduced, the back-layer LoRA can learn the multimodal associations of these concepts, making the embedding distances of similar image-text pairs closer. Without changing the underlying feature extractor, these low-rank updates can focus on adjusting key parameters and efficiently adapt to the requirements of downstream tasks.
[0102] Middle-layer optimization: In the second stage, the LoRA adapters are extended to the middle layers. The middle layers control the transition from low-level features to high-level semantics, and their adjustment helps the model extract mid-scale patterns more relevant to the new task. In this stage, not only the back-layer LoRA from the previous stage is continued to be trained (the learning rate can be appropriately reduced to prevent over-adjustment), but also the LoRA adapters for the middle layers are newly trained. To prevent excessive perturbation of the pre-trained features, the rank of the middle-layer LoRA can be moderately medium (such as 4 or 8) to limit its parameter scale, thus encouraging it to learn detailed supplementary adjustments rather than significantly modifying the original features. By optimizing the middle-layer LoRA, the model can better represent the detailed patterns of new data (such as special textures, backgrounds, or industry terms in texts), and cooperate with the back layers to improve the multimodal alignment effect.
[0103] Front - layer Optimization: In the final stage, add LoRA adapters to the front - layer as needed. Since the front - layer extracts very general low - level features, generally only minor adjustments are made to the front - layer. If the data distribution of the new task is significantly different from the pre - training (such as the overall color style of images is different, or the words used in the text are very special), then the adjustment of the front - layer LoRA can help the model better perceive these underlying differences. When optimizing the front - layer LoRA, it is advisable to use a relatively small rank (such as 4) and a lower learning rate to ensure that the changes are controlled and avoid damaging the model's original general feature extraction ability. The front - layer LoRA is more about fine - tuning the basic features to better fit a specific domain (for example, slightly changing the first - layer convolution / projection to emphasize certain frequency components). At this stage, the LoRAs of each layer group work together: the front - layer LoRA fine - tunes the input representation, the middle - layer LoRA adjusts the way of feature combination, and the back - layer LoRA calibrates the final embedding, so as to fully adapt to the new task requirements.
[0104] (3) Progressive Fine - tuning Training
[0105] Based on the above - mentioned stage - by - stage optimization strategy, the entire fine - tuning is carried out in stages using a progressive unfreezing strategy, that is, starting from training only the back - layer and gradually increasing the layer groups involved in training. This way can ensure that only limited adjustments are made to the model each time, avoiding the instability caused by fine - tuning the entire model at once. The specific process is as follows:
[0106] Stage 1: Fine - tune the back - layer
[0107] In the initial stage, freeze all the weights of the CLIP model except the LoRA adapters in the back - layer, that is, the original parameters of the visual encoder and the text encoder remain unchanged, and only the LoRA adapters inserted in the back - layer are open for training. Use the AdamW optimizer to update only the parameters of the back - layer LoRA and continue contrastive learning on the model loaded from the pre - trained weights. During the training process, the image - text pairs pass through the encoder to obtain their respective embeddings, and the cosine similarity and contrastive loss are calculated to update the back - layer LoRA. Since the number of trainable parameters is very small, the learning rate can be set slightly higher appropriately to accelerate convergence, and at the same time, a short warmup (such as linearly increasing the learning rate in the first 0 - 0.1 epochs) is used to smooth the initial stage. The goal of this stage is to quickly align the semantics of the new data: the model adjusts its high - level representation so that the matching image - text pairs are closer in the embedding space and the non - matching ones are farther away, thus reducing the contrastive loss.
[0108] Continuously monitor the performance of the model on the validation set. When the metrics tend to be stable (for example, the evaluation accuracy no longer increases significantly after several epochs), the next stage can be entered. By focusing on high - level adaptation in Stage 1, the perturbation of the underlying general features is effectively avoided, reducing the risks of training instability and overfitting.
[0109] Stage 2: Fine-tune the middle layer
[0110] After ensuring that the high-level semantics are basically aligned, unfreeze the middle layer of the model (i.e., insert LoRA adapters into the middle layer and add them to the set of trainable parameters). At this time, the LoRA parameters of the middle layer and the later layer are involved in training, while the original weights remain frozen. During joint training, different learning rates can be used for different layer groups. For example, a slightly higher learning rate is used for the middle layer LoRA and a lower learning rate is used for the later layer LoRA to reflect that the middle layer is still far from the optimum and requires a larger step size for adjustment. As training progresses, the middle layer LoRA begins to adjust the intermediate representations of the model. For example, it makes the middle layer of the visual encoder pay more attention to the image region features related to the new task, and makes the middle layer of the text encoder better fuse the context of new domain terms. Through the cooperation of the middle layer and the later layer, the model's representation of new data will be more refined and accurate. The progressive unfreezing strategy ensures that the model adjustment is gradual: only when the previous higher layer group has adapted, the next lower layer group is unfrozen. Therefore, the newly unfrozen layer faces a sub-problem that is easier to optimize (the upper layer has been adjusted and the upper layer gradient change slows down), which can reduce the drastic update of parameters, slow down "catastrophic forgetting", and improve the stability of fine-tuning.
[0111] After training for several epochs in Stage 2, continue to evaluate the model performance through the validation set. If the metrics improve significantly and tend to be stable, the final stage can be entered.
[0112] Stage 3: Fine-tune the front layer
[0113] Usually, adjusting the front layer is significantly beneficial only when the new task is very different from the original pre-training task (for example, differences in image style or camera result in different low-level feature distributions, or the text contains a large number of new words not seen by the model). Otherwise, it can be chosen not to unfreeze the front layer to avoid interfering with the generalization ability.
[0114] If Stage 3 training is carried out, a lower learning rate (one order of magnitude lower than the previous two stages, such as from 1e-4 to 1e-5 or even lower) is used to train the front layer LoRA to slightly adjust the front layer weights. At the same time, the middle layer and the later layer LoRA can be continued to be frozen at the beginning, and only the front layer LoRA is updated to make it learn slowly; then jointly fine-tune with the middle and later layers for a few rounds at an extremely low learning rate to make the LoRA adapter parameters of the entire model reach the overall coordination optimum on the new data. In this stage, the front layer LoRA will make some changes to the input representation of the model. For example, it makes the first layer of the visual encoder more prominent in the common edge patterns or color distributions in the new domain, and makes the embedding layer of the text encoder better represent the newly introduced professional vocabulary. Since the adjustment amplitude is very small, most of the pre-trained features of the model are still retained, only the final bottom-level optimization correction is made for the new task. After Stage 3, all layers of the model from front to back have been fine-tuned for the new task (only through the LoRA adapters of each layer), achieving full adaptation.
[0115] The entire training process unfreezes layer by layer and trains progressively in a "three-step" manner, ensuring that the model only needs to adjust a limited number of parameters at each stage, greatly reducing the training instability and the risk of forgetting.
[0116] Example Two
[0117] Based on the above method, this example provides a visual question answering optimization system based on hierarchical multi-modal fine-tuning, which mainly includes:
[0118] A visual feature extraction module, which uses the visual encoder of the pre-trained CLIP model to extract the visual features of the image in the visual question answering task, and forms a visual feature set from the hierarchical features extracted by the visual encoder;
[0119] A text feature extraction module, which uses the text encoder of the pre-trained CLIP model to extract the text features of the text in the visual question answering task, and forms a text feature set from the hierarchical features extracted by the text encoder;
[0120] A weighted projection module, which adaptively weights and adjusts the text features of each layer in the text feature set, and splices the weighted text features of each layer into a vector for linear projection;
[0121] A cross-modal bridging module, which uses the multi-head attention mechanism and residual connection to perform cross-modal fusion on the linearly projected text features and the vector spliced from the visual features of each layer in the visual feature set to obtain the semantic visual perception fusion features for the visual question answering task.
[0122] It also includes a LoRA fine-tuning module, which performs multi-layer group-by-stage fine-tuning on the visual encoder and text encoder through low-rank adaptation technology.
[0123] The above system can execute the visual question answering optimization method based on hierarchical multi-modal fine-tuning described in Example One, and has the corresponding functional modules and beneficial effects of the method. For the technical details not described in detail in this example, reference can be made to the visual question answering optimization method based on hierarchical multi-modal fine-tuning provided in Example One of the present invention.
[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; under the idea of the present invention, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other changes in different aspects of the present invention as described above. For the sake of brevity, they are not provided in detail; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A visual question answering optimization method based on hierarchical multimodal fine-tuning, characterized in that: Includes steps: S1. Obtain text-image pairs for visual question answering tasks, including images and corresponding texts; S2. Use the visual encoder of the pre-trained CLIP model to extract the visual features of the image, and combine the features of each level extracted by the visual encoder into a visual feature set; S3, using the text encoder of the pre-trained CLIP model to extract text features of the text, and composing the features of each level extracted by the text encoder into a text feature set; S4. Adaptively weight the text features of each layer in the text feature set: , in, The weighted adjusted i Layer text features, represents the corresponding element multiplication operation, represents the addition operation of corresponding elements, Represents adaptive weight W Corresponding to No. i The layer weight vector, Represents the text encoder i The text features extracted by the layer; S5, concatenating the weighted text features of each layer into a vector and performing linear projection; S6, concatenate the visual features of each layer in the visual feature set into a vector, and perform cross-modal interactive fusion with the text features after linear projection through a multi-head attention mechanism; S7, fusing the feature vector obtained by cross-modal interactive fusion with the vector obtained by concatenating the visual features of each layer in the visual feature set to obtain a semantic visual perception fusion feature; S8, using semantic visual perception fusion features for visual question answering tasks; The visual encoder and the text encoder are fine-tuned in multiple layers and stages by low-rank adaptation technology, including the following steps: Model layer group division: The network layers of the visual encoder and text encoder are divided into front layer, middle layer and back layer according to the depth; First stage fine-tuning: Add LoRA adapters to the back layers of the visual encoder and text encoder respectively, keep the pre-trained parameters of the visual encoder and text encoder unchanged, and only open the LoRA adapter inserted in the back layer for training. During the training process, only the parameters of the LoRA adapter are updated; The second stage of fine-tuning: add LoRA adapters to the middle layer of the visual encoder and text encoder respectively, keep the pre-trained parameters of the visual encoder and text encoder unchanged, and only open the LoRA adapters inserted in the back layer and middle layer for training. During the training process, only the parameters of the LoRA adapter are updated.
2. The visual question answering optimization method according to claim 1, characterized in that: The visual encoder and text encoder of the pre-trained CLIP model both contain 12 Transformer layers. In the model layer group division, the 1st to 4th Transformer layers are used as the front layers, the 5th to 8th Transformer layers are used as the middle layers, and the 8th to 12th Transformer layers and the projection layer are used as the back layers.
3. The visual question answering optimization method according to claim 1, characterized in that: It also includes the third stage of fine-tuning: adding LoRA adapters to the front layers of the visual encoder and text encoder respectively, keeping the pre-trained parameters of the visual encoder and text encoder unchanged, and only opening the LoRA adapters inserted in the back, middle and front layers for training. During the training process, only the parameters of the LoRA adapters are updated.
4. The visual question answering optimization method according to claim 3, characterized in that: In the third stage of fine-tuning, the parameters of the visual encoder, text encoder, and middle and back layer LoRA adapters are first kept unchanged, and only the LoRA adapter inserted in the front layer is opened for low learning rate training; then the middle and back layer LoRA adapters are opened and fine-tuned together for several rounds at a very small learning rate to make all LoRA adapter parameters reach the overall coordinated optimal.
5. The visual question answering optimization method according to claim 1, characterized in that: In step S5, the linear projection is expressed as: in, is the text feature after linear projection, concat represents vector concatenation, represents the matrix multiplication operation, is the linear projection weight.
6. The visual question answering optimization method according to claim 1, characterized in that: In step S6, the text features after linear projection are used as the key vector K and the value vector V, and the vector concatenated from the visual features of each layer in the visual feature set is used as the query vector Q, and the cross-modal interactive fusion features are obtained through the multi-head attention mechanism.
7. The visual question answering optimization method according to claim 1, characterized in that: In step S7, a residual connection is used to add the corresponding elements of the cross-modal interaction fusion features output by the multi-head attention mechanism and the vector concatenated by the visual features of each layer in the visual feature set to obtain the semantic visual perception fusion features: , in, f is the semantic visual perception fusion feature, For cross-modal interactive fusion features, Represents the vector concatenated from each layer of visual features in the visual feature set. Represents the addition operation of corresponding elements.
8. A visual question answering optimization system based on the method according to any one of claims 1 to 7, characterized in that: It includes visual feature extraction module, text feature extraction module, weighted projection module and cross-modal bridging module; The visual feature extraction module is used to extract the visual features of the image in the visual question answering task using the visual encoder of the pre-trained CLIP model, and to form a visual feature set from the features of each level extracted by the visual encoder; The text feature extraction module is used to extract text features of text in the visual question answering task using the text encoder of the pre-trained CLIP model, and to form a text feature set from the features of each level extracted by the text encoder; The weighted projection module is used to perform adaptive weighted adjustment on the text features of each layer in the text feature set, and splice the weighted adjusted text features of each layer into a vector for linear projection; The cross-modal bridging module is used to perform cross-modal fusion of the linearly projected text features with the vectors formed by splicing the visual features of each layer in the visual feature set through a multi-head attention mechanism and residual connection, so as to obtain semantic visual perception fusion features for visual question answering tasks; It also includes a LoRA fine-tuning module for performing multi-layer group-by-stage fine-tuning on the visual encoder and the text encoder through a low-rank adaptation technique.
Citation Information
Patent Citations
Multi-round visual dialogue method based on multi-level sorting learning
CN113435399A
Universal multi-modal learning method based on deep interactive adaptation network model
CN116882477A