Visual question and answer optimization method and system based on hierarchical multi-modal fine adjustment
By adopting a hierarchical multimodal fine adjustment method in KI-VQA task, the multi-layer features of images and text are extracted and fused with the CLIP model, the problem of inconsistent cross-modal feature matching is solved, and the performance of VQA task is significantly improved.
Patent Information
- Application Number
- CN202510472099.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing KI-VQA method encounters difficulties in dealing with semantic differences and information fusion between images and text, resulting in inconsistent cross-modal feature matching, affecting the accuracy and richness of question-and-answer.
The visual question-and-answer optimization method based on hierarchical multimodal fine adjustment is adopted. Multi-layer features of images and text are extracted by pre-training the CLIP model, adaptive weighting adjustment and linear projection are performed, and cross-modal interaction fusion is used to generate semantic visual perception fusion features.
It significantly improves the accuracy and effect of cross-modal learning, solves the problem of inconsistent matching of visual features and text semantics, and can better capture high-level text and visual interactions, thereby improving the performance of VQA tasks.
Smart Images

Figure CN120011547A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of visual question answering, and in particular, relates to a visual question answering optimization method and system based on hierarchical multimodal fine adjustment. Background Art
[0002] The knowledge-intensive visual question answering (KI-VQA) task is a variant of visual question answering that combines image information with external knowledge. Its core is to provide users with more accurate and rich answers by integrating images with external knowledge bases. This task has important practical significance in e-commerce, education and other fields, especially in scenarios where visual information needs to be combined with external knowledge bases. The challenge of the KI-VQA task is that the answer to the question cannot usually be obtained only from the image itself, but requires the help of external knowledge sources. For example, on an e-commerce website, users can take photos of products and ask related questions. The system needs to combine product images and detailed information about the product (such as brand, model, function, etc.) to answer questions; in the field of education, students can ask questions about illustrations in textbooks, and the system needs to combine the content of the illustrations and related knowledge points to answer questions.
[0003] At present, most KI-VQA methods assume that external knowledge can be obtained from a structured knowledge base. First, the image and text encoders are used to extract their respective features. Then, the image features and text features are fused, and multimodal features are used for retrieval and question-answer generation. These methods need to effectively fuse image and text information and retrieve question-related content from external knowledge bases to improve the accuracy of question-answering. However, existing dense retrieval methods usually face the heterogeneity problem between multimodal (image and text) input and single modality (text) output. The semantic difference between images and text makes it difficult to fuse information between modalities.
[0004] How to effectively capture the correlation between images and texts, and retrieve information related to questions from external knowledge to improve the accuracy and richness of question answering is still an unsolved problem. Existing solutions to these problems include using large-scale text corpora as external knowledge sources and retrieving questions and images through dense retrieval models, but how to deal with semantic differences between different modalities and achieve effective knowledge transfer is still a research difficulty in the field of KI-VQA. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a visual question answering optimization method and system based on hierarchical multimodal fine-tuning, which improves the performance of visual question answering tasks by introducing refined visual and text feature interaction.
[0006] The technical solution provided by the present invention is as follows: A visual question answering optimization method based on hierarchical multimodal fine-tuning includes the following steps: S1. Obtain text-image pairs for visual question answering tasks, including images and corresponding texts; S2. Use the visual encoder of the pre-trained CLIP model to extract the visual features of the image, and combine the features of each level extracted by the visual encoder into a visual feature set; S3, using the text encoder of the pre-trained CLIP model to extract text features of the text, and composing the features of each level extracted by the text encoder into a text feature set; S4, adaptively weighting and adjusting the text features of each layer in the text feature set; S5, concatenating the weighted text features of each layer into a vector and performing linear projection; S6, concatenate the visual features of each layer in the visual feature set into a vector, and perform cross-modal interactive fusion with the text features after linear projection through a multi-head attention mechanism; S7, fusing the feature vector obtained by cross-modal interactive fusion with the vector obtained by concatenating the visual features of each layer in the visual feature set to obtain a semantic visual perception fusion feature; S8. Use semantic visual perception fusion features for visual question answering tasks.
[0007] Furthermore, the visual encoder and the text encoder are fine-tuned in multiple layers and stages by using a low-rank adaptation technique, including the steps of: Model layer group division: The network layers of the visual encoder and text encoder are divided into front layer, middle layer and back layer according to the depth; First stage fine-tuning: Add LoRA adapters to the back layers of the visual encoder and text encoder respectively, keep the pre-trained parameters of the visual encoder and text encoder unchanged, and only open the LoRA adapter inserted in the back layer for training. During the training process, only the parameters of the LoRA adapter are updated; The second stage of fine-tuning: add LoRA adapters to the middle layer of the visual encoder and text encoder respectively, keep the pre-trained parameters of the visual encoder and text encoder unchanged, and only open the LoRA adapters inserted in the back layer and middle layer for training. During the training process, only the parameters of the LoRA adapter are updated.
[0008] Preferably, the visual encoder and text encoder of the pre-trained CLIP model each include 12 Transformer layers. In the model layer group division, the 1st to 4th Transformer layers are used as the front layers, the 5th to 8th Transformer layers are used as the middle layers, and the 8th to 12th Transformer layers and the projection layer are used as the back layers.
[0009] Furthermore, it also includes the third stage of fine-tuning: adding LoRA adapters to the front layers of the visual encoder and text encoder respectively, keeping the pre-trained parameters of the visual encoder and text encoder unchanged, and only opening the LoRA adapters inserted in the back layer, middle layer and front layer for training. Only the parameters of the LoRA adapters are updated during the training process.
[0010] Preferably, in the third stage fine-tuning process, the parameters of the visual encoder, text encoder and the middle and rear layer LoRA adapters are first kept unchanged, and only the LoRA adapter inserted in the front layer is opened for low learning rate training; then the middle and rear layer LoRA adapters are opened and jointly fine-tuned for several rounds at a very small learning rate, so that all LoRA adapter parameters reach the overall coordinated optimal.
[0011] Further, in step S4, the adaptive weighted adjustment is expressed as: , in, The weighted adjusted i Layer text features, represents the corresponding element multiplication operation, represents the addition operation of corresponding elements, Represents adaptive weight W Corresponding to No. i The layer weight vector, Represents the text encoder i The text features extracted by the layer.
[0012] Further, in step S5, the linear projection is expressed as:
[0013] in, is the text feature after linear projection, concat represents vector concatenation, represents the matrix multiplication operation, is the linear projection weight.
[0014] Furthermore, in step S6, the text features after linear projection are used as the key vector K and the value vector V, and the vector concatenated from the visual features of each layer in the visual feature set is used as the query vector Q, and the cross-modal interactive fusion features are obtained through the multi-head attention mechanism.
[0015] Furthermore, in step S7, a residual connection is used to add the corresponding elements of the cross-modal interaction fusion features output by the multi-head attention mechanism and the vector concatenated by the visual features of each layer in the visual feature set to obtain the semantic visual perception fusion features: , in, f is the semantic visual perception fusion feature, For cross-modal interactive fusion features, Represents the vector concatenated from each layer of visual features in the visual feature set. Represents the addition operation of corresponding elements.
[0016] A visual question answering optimization system based on the above method, comprising a visual feature extraction module, a text feature extraction module, a weighted projection module and a cross-modal bridging module; The visual feature extraction module is used to extract the visual features of the image in the visual question answering task using the visual encoder of the pre-trained CLIP model, and to form a visual feature set from the features of each level extracted by the visual encoder; The text feature extraction module is used to extract text features of text in the visual question answering task using the text encoder of the pre-trained CLIP model, and to form a text feature set from the features of each level extracted by the text encoder; The weighted projection module is used to perform adaptive weighted adjustment on the text features of each layer in the text feature set, and splice the weighted adjusted text features of each layer into a vector for linear projection; The cross-modal bridging module is used to perform cross-modal fusion of the linearly projected text features with the vectors formed by splicing the visual features of each layer in the visual feature set through a multi-head attention mechanism and residual connection, so as to obtain semantic visual perception fusion features for visual question answering tasks; It also includes a LoRA fine-tuning module for performing multi-layer group-by-stage fine-tuning on the visual encoder and the text encoder through a low-rank adaptation technique.
[0017] Compared with the prior art, the present invention has at least the following beneficial effects: The present invention significantly improves the accuracy and effect of cross-modal learning through innovative adaptive cross-modal bridging and multi-layer group stage-by-stage low-rank adaptation methods. The method introduces multi-level semantically perceived text information into the visual features, so that the visual features are more finely adjusted and optimized, thereby solving the inconsistency problem between the visual features and the text semantics. Through multi-layer group stage-by-stage low-rank adaptation, the accumulation of perceptual errors is effectively avoided, and a layer-by-layer adaptation mechanism is provided, so that the visual features and text features can be more accurately aligned in the multi-level learning process. The present invention provides new ideas and methods for the fusion of visual-text features in cross-modal tasks, which can better capture the high-level interaction between text and visual fields, thereby improving the performance of VQA tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0019] Figure 1 It is a schematic diagram of a framework of feature adaptive cross-modal bridging provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, other embodiments obtained by ordinary technicians in this field without making creative work are all within the scope of protection of the present invention.
[0021] Embodiment 1 This embodiment provides a visual question answering optimization method based on hierarchical multimodal fine-tuning, such as Figure 1 As shown in the figure, it aims to improve the performance of visual question answering tasks by introducing refined interactions between visual and textual features.
[0022] The method mainly includes the following steps: 1. Initial feature extraction For a given image in a visual question answering task I and text T (Question), the corresponding visual features and text features are extracted respectively through the CLIP (Contrastive Language-Image Pre-training) model.
[0023] The CLIP model is a multimodal pre-training model that aims to achieve matching and association between images and texts through contrastive learning. The core idea is to encode images and texts into vectors and optimize the similarity of these vectors through the cosine similarity loss function to achieve cross-modal understanding and interaction. The CLIP model consists of two main parts: the visual encoder and the text encoder. The visual encoder is responsible for converting the input image into a high-dimensional vector representation, while the text encoder converts the text information into a vector representation. The two are jointly optimized through contrastive learning to achieve semantic alignment between images and texts.
[0024] CLIP's visual encoder uses the Transformer (ViT) architecture. In the visual encoder, the input image is segmented into multiple patches, and features are extracted layer by layer through multiple Transformer layers. The output of each layer is a further abstraction and enhancement of the features of the previous layer. Finally, a high-dimensional feature representation of the image is generated by stacking multiple layers of Transformers. Similarly, CLIP's text encoder is based on the Transformer architecture. Inside the text encoder, the text is segmented and converted into embedding vectors. These embedding vectors are processed through multiple Transformer layers. Each layer performs linear and nonlinear transformations on the input embedding vectors, thereby gradually extracting the semantic features of the text. Finally, the text encoder outputs a high-dimensional feature vector that represents the semantic meaning of the input text.
[0025] In this embodiment, when the image I When extracting features, unlike the general method of using only the output features of the final layer, the features of each level extracted by the CLIP visual encoder are combined into a visual feature set. , m Indicates the number of levels of the image encoder, This means that the CLIP visual encoder i The visual features extracted by the layer. Similarly, for text T The features of each level extracted by CLIP text encoder are combined into a text feature set , n Indicates the number of levels of the text encoder, This means that the CLIP text encoder i The text features extracted by the layer.
[0026] 2. Feature Adaptive Cross-Modal Bridging (1) Adaptive weighting of text features Text feature set The text features of each layer in the W Make a weighted adjustment, namely: , in, The weighted adjusted i Layer text features, represents the corresponding element multiplication operation, represents the addition operation of corresponding elements, Represents adaptive weight W Corresponding to No. i The layer weight vector is . Adaptive weights W The value of is randomly initialized at the beginning of training to ensure that it does not have a significant impact on the model output in the early stages of training, and is subsequently iteratively updated during the back propagation of the training process.
[0027] (2) Text feature splicing and projection The weighted text features of each layer are concatenated into a vector and mapped to the visual embedding space by matrix multiplication with the linear projection weights:
[0028] in, To stitch the projected text features, it can be effectively connected with the visual features; concat represents vector concatenation, Represents a matrix multiplication operation; is the linear projection weight, which is a parameter matrix learned through training, and it can map the concatenated text feature vector to the visual embedding space.
[0029] The iterative update during the training process can be expressed as:
[0030] in, is the learning rate, is the loss function defined by downstream tasks such as visual question answering.
[0031] (3) Visual feature stitching To facilitate the interactive fusion of visual and text features, similarly, the visual feature set The visual features of each layer in are concatenated into a vector: .
[0032] (4) Cross-modal multi-head attention mechanism The multi-head attention mechanism is used to achieve cross-modal fusion of text features and visual features. The multi-head attention mechanism is a key component in the Transformer architecture and is widely used in natural language processing tasks and image processing tasks. Its core idea is to capture different features of the input through multiple different attention heads, thereby improving the expressiveness of the model.
[0033] Specifically, the multi-head attention mechanism generally proceeds through the following steps: 1) Linear mapping The input query vector Q, key vector K and value vector V are projected into h different subspaces (i.e. different attention heads). For each attention head i, calculate: , in, , and is the linear projection matrix of attention head i.
[0034] 2) Calculate multiple attentions in parallel For each head, assuming the dimension of each attention head is d, the scaled dot product attention mechanism is used to calculate the attention output:
[0035] 3) Concatenate the outputs of the attention heads and concatenate the outputs of all h attention heads to form a large vector:
[0036] 4) Linear transformation The concatenated result is projected through another linear transformation to obtain the final multi-head attention output:
[0037] Projection Matrix , and And for the output of the splicing , are initially randomly initialized and then gradually learned through training.
[0038] In this embodiment, the text features after splicing and projection are As the key vector K and value vector V, the concatenated multi-layer visual features As the query vector Q, the cross-modal interactive fusion features are obtained through the multi-head attention mechanism ,Right now Finally, residual connection is used to fuse the cross-modal interactive features output by the multi-head attention mechanism. With multi-layer visual features Then add the corresponding elements to further optimize the representation of visual features and obtain the final semantic visual perception fusion feature f ,Right now .
[0039] The number of attention heads is an important hyperparameter in the multi-head attention mechanism. Its setting directly affects the expressiveness and computational complexity of the model, and is also limited by the size of the data set and computing resources. This embodiment gradually adjusts the number of heads through experiments and observes the changes in model performance. It is found that setting the number of attention heads to 8 is more appropriate.
[0040] 3. Finally, the obtained semantic visual perception fusion features are applied to the downstream visual question answering task.
[0041] The semantic difference between images and text is one of the core challenges of the VQA task. Traditional feature fusion methods (such as direct embedding, linear fusion, etc.) often cannot fully capture the complex interactive relationship between the two, resulting in limited model performance. This embodiment introduces multi-level semantically-aware text information into the visual features, so that the visual features can be more finely adjusted and optimized at multiple levels, achieving effective semantic alignment and modal fusion, thereby solving the inconsistency problem between visual features and text semantic matching, and can better capture high-level interactions in the language and visual fields, thereby improving the performance of the VQA task.
[0042] In some embodiments, the pre-trained CLIP model is fine-tuned through multi-layer group stage-by-stage low-rank adaptation technology (LoRA) to improve the adaptability or generalization of the pre-trained model to specific application fields, ensure that the visual features and text features can remain consistent in multiple layers, and thus effectively avoid the accumulation of perceptual errors, so that the visual features and text features can be more matched in subsequent tasks.
[0043] The LoRA fine-tuning of the multi-layer group stage by stage mainly includes the following steps: (1) Model layer group division The visual encoder layer usually consists of an image embedding layer (including patch partitioning and position encoding) and several Transformer blocks. For example, for ViT-B / 16 with 12 Transformer layers, it can be divided into: Front layers: refers to the initial layers of the model (e.g. the 1st to 4th Transformer layers). These layers focus on extracting low-level visual features, such as local patterns such as edges, textures, and colors. They provide general image representations.
[0044] Middle layer: refers to the middle layers (such as the 5th to 8th Transformer layers), which gradually capture higher-level feature patterns (such as object parts, complex shapes) and combine low-level features into mid-level representations.
[0045] Back layers: refers to the last few layers (such as the 9th to 12th Transformer layers) and the final image feature projection layer. The back layers focus on extracting global semantic information, and the learned features are more discriminative and task-relevant.
[0046] Similarly, the levels of the text encoder can be divided as follows: Front layer: such as the 1st to 4th Transformer layers in the Embedding layer. These layers mainly capture information at the word level and local grammatical level (similar to the basic lexical and phrase features learned by the model).
[0047] Middle layers: such as the 5th to 8th Transformer layers, these layers gradually model longer-range dependencies and semantic combinations (phrase to sentence representations).
[0048] Back layers: 9th to 12th Transformer layers and the final text projection layer. The back layers output the semantic embedding of the entire sentence to understand the global semantics and context.
[0049] Dividing the visual encoder and text encoder into layers according to depth reflects the step-by-step abstraction from basic features to high-level semantics. Dividing the layers into groups helps to adopt different strategies in subsequent fine-tuning: first adjust the back layers to adapt to the semantics of the new task, and gradually adjust the middle and front layers, so as to retain the original general knowledge of the model to the greatest extent.
[0050] (2) Optimize strategy step by step LoRA efficiently adapts to downstream tasks by adding parameters in the form of low-rank matrices to existing model layers. During fine-tuning, the original weights of the model are not directly updated, but a trainable low-rank matrix (LoRA adapter) is injected into each specified layer. A and B ,satisfy ( W is the original weight of the model, This not only keeps the pre-trained weights frozen, but also enables model adjustment with very few new parameters.
[0051] In order to make the low-rank adapter play its full role, a phase-by-phase LoRA optimization strategy is designed: Post-layer optimization: In the first stage, the LoRA adapter is introduced in the post-layer of the visual and text encoders, and the parameters of this part of the adapter are mainly optimized. The post-layer directly determines the final embedding representation of the image and text, and its output is directly related to the contrast loss, so priority adjustment can quickly respond to the alignment requirements of new tasks. At this stage, a relatively high rank is set for the post-layer LoRA, such as 8 or 16, to ensure sufficient expressive power to capture new semantic mapping relationships. At the same time, an appropriate scaling factor is combined to stabilize the impact of LoRA updates on the original weights. The optimization goal is to correct the image-text embedding space for new data: for example, when new categories or concepts are introduced, the post-layer LoRA can learn the multimodal associations of these concepts, making the embedding distance of similar image-text pairs closer. Without changing the underlying feature extractor, these low-rank updates can focus on adjusting key parameters and efficiently adapt to downstream task requirements.
[0052] Middle-layer optimization: The second stage extends the LoRA adapter to the middle layer. The middle layer controls the transition from low-level features to high-level semantics, and its adjustment helps the model extract mesoscale patterns that are more relevant to new tasks. In this stage, not only the rear-layer LoRA of the previous stage continues to be trained (its learning rate can be appropriately reduced to prevent over-adjustment), but the middle-layer LoRA adapter is also newly trained. To prevent excessive perturbations to pre-trained features, the rank of the middle-layer LoRA can be appropriately medium (for example, 4 or 8) to limit its parameter scale, thereby encouraging it to learn detailed supplementary adjustments rather than drastically modifying the original features. By optimizing the middle-layer LoRA, the model can better represent the detailed patterns of new data (such as special textures, backgrounds, or industry terms in text, etc.), and work with the back layers to improve the multimodal alignment effect.
[0053] Front layer optimization: In the final stage, LoRA adapters are added to the front layer as needed. Since the front layer extracts very general low-level features, generally only minor adjustments are required in the front layer. If the data distribution of the new task is significantly different from the pre-training (for example, the overall color style of the image is different, or the text words are very special), the adjustment of the front layer LoRA can help the model better perceive these underlying differences. When optimizing the front layer LoRA, a smaller rank (such as 4) and a lower learning rate are preferred to ensure that the changes are controlled and avoid destroying the original general feature extraction ability of the model. The front layer LoRA is more about fine-tuning the basic features to make them more suitable for a specific domain (for example, slightly changing the first layer of convolution / projection to emphasize certain frequency components). At this stage, the LoRAs of each layer group work together: the front layer LoRA fine-tunes the input representation, the middle layer LoRA adjusts the feature combination method, and the back layer LoRA calibrates the final embedding to fully adapt to the new task requirements.
[0054] (3) Progressive fine-tuning training Based on the above-mentioned phase-by-phase optimization strategy, the entire fine-tuning is carried out in stages using a progressive unfreezing strategy, that is, starting from training only the back layers, and gradually increasing the number of layers involved in the training. This approach ensures that only limited adjustments are made to the model each time, avoiding instability caused by fine-tuning the entire model at once. The specific process is as follows: Stage 1: Fine-tuning the back layer In the initial stage, all weights of the CLIP model except the rear-layer LoRA adapter are frozen, that is, the original parameters of the visual encoder and text encoder remain unchanged, and only the LoRA adapter inserted in the rear layer is open for training. The AdamW optimizer is used to update only the rear-layer LoRA parameters, and contrastive learning continues on the model loaded with pre-trained weights. During the training process, the image-text pairs are passed through the encoder to obtain their own embeddings, and the rear-layer LoRA is updated by calculating cosine similarity and contrast loss. Since there are very few trainable parameters, the learning rate can be set slightly higher to accelerate convergence, and a short warmup (for example, linearly increase the learning rate for the first 0 to 0.1 epochs) is used to smooth the initial stage. The goal of this stage is to quickly align the semantics of new data: the model adjusts its high-level representation so that matching image-text pairs are closer in the embedding space and mismatched ones are farther away, thereby reducing contrast loss.
[0055] Continuously monitor the performance of the model on the validation set. When the indicators tend to be stable (for example, the evaluation accuracy no longer increases significantly after several epochs), you can enter the next stage. By focusing on high-level adaptation in stage 1, the disturbance of the underlying common features is effectively avoided, reducing the risk of training instability and overfitting.
[0056] Phase 2: Fine-tuning the middle layer After ensuring that the high-level semantics are basically aligned, the middle layer of the model is unfrozen (that is, the LoRA adapter is inserted into the middle layer and added to the set of trainable parameters). At this time, the LoRA parameters of the middle and back layers are involved in the training, while the original weights remain frozen. During joint training, group learning rates can be used for different layer groups, such as a slightly higher learning rate for the middle layer LoRA and a lower learning rate for the back layer LoRA, to reflect that the middle layer is still far from the optimal and needs a larger step adjustment. As the training proceeds, the middle layer LoRA begins to adjust the intermediate representation of the model, such as making the middle layer of the visual encoder pay more attention to the image area features related to the new task, and making the middle layer of the text encoder better integrate the context of new domain terms. Through the cooperation of the middle and back layers, the model's representation of new data will be more refined and accurate. The progressive unfreezing strategy ensures that the model adjustment is gradual: only when the previous set of higher layers has been adapted, the next set of lower layers is unfrozen, so the newly unfrozen layer faces a sub-problem that is easier to optimize (the upper layer has been adjusted, and the upper layer gradient changes slowly), which can reduce the drastic update of parameters, slow down "catastrophic forgetting", and improve the stability of fine-tuning.
[0057] After training for several epochs in stage 2, continue to evaluate the model performance through the validation set. If the indicators are significantly improved and tend to be stable, you can enter the final stage.
[0058] Stage 3: Fine-tuning the front layer Generally, adjusting the front layers will only have significant benefits when the new task is very different from the original pre-training task (for example, the image style or camera differences lead to different low-level feature distributions, or the text contains a large number of new words that the model has not seen). Otherwise, you can choose not to unfreeze the front layers to avoid interfering with generalization ability.
[0059] If the third stage of training is carried out, a lower learning rate (one order of magnitude lower than the first two stages, such as from 1e-4 to 1e-5 or even lower) is used to train the front layer LoRA to slightly adjust the weights of the front layer. At the same time, the middle and rear layers LoRA can continue to be frozen at the beginning, and only the front layer LoRA is updated to make it learn slowly; then it is fine-tuned together with the middle and rear layers at a very small learning rate for several rounds, so that the LoRA adapter parameters of the entire model can reach the overall coordinated optimality on the new data. At this stage, the front layer LoRA will make some changes to the input representation of the model, such as making the first layer of the visual encoder more prominent in the new field Common edge patterns or color distributions, and allowing the embedding layer of the text encoder to better represent the newly introduced professional vocabulary. Since the adjustment is small, most of the pre-trained features of the model are still retained, and only the final underlying optimization correction is made for the new task. After the third stage, all layers of the model from front to back have been fine-tuned for the new task (only through the LoRA adapters of each layer), achieving full adaptation.
[0060] The entire training process is unfrozen and trained layer by layer in a "three-step" manner, ensuring that only a limited number of model parameters need to be adjusted at each stage, greatly reducing training instability and the risk of disasters.
[0061] Embodiment 2 Based on the above method, this embodiment provides a visual question answering optimization system based on hierarchical multimodal fine adjustment, which mainly includes: The visual feature extraction module is used to extract the visual features of the image in the visual question answering task using the visual encoder of the pre-trained CLIP model, and to form a visual feature set from the features of each level extracted by the visual encoder; The text feature extraction module is used to extract the text features of the text in the visual question answering task using the text encoder of the pre-trained CLIP model, and to form a text feature set from the features of each level extracted by the text encoder; The weighted projection module is used to perform adaptive weighted adjustment on the text features of each layer in the text feature set, and concatenate the weighted adjusted text features of each layer into a vector for linear projection; The cross-modal bridging module is used to cross-modally fuse the linearly projected text features with the vector concatenated from the visual features of each layer in the visual feature set through a multi-head attention mechanism and residual connection, so as to obtain the semantic visual perception fusion features for visual question answering tasks.
[0062] It also includes a LoRA fine-tuning module for performing multi-layer group-by-stage fine-tuning on the visual encoder and the text encoder through a low-rank adaptation technique.
[0063] The above-mentioned system can execute the visual question answering optimization method based on hierarchical multimodal fine-tuning described in Example 1, and has the corresponding functional modules and beneficial effects of the method. For technical details not described in detail in this embodiment, please refer to the visual question answering optimization method based on hierarchical multimodal fine-tuning provided in Example 1 of the present invention.
[0064] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution, in essence or in other words, the part that contributes to the relevant technology, can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.
[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Under the concept of the present invention, the technical features in the above embodiments or different embodiments may also be combined, the steps may be implemented in any order, and there are many other changes in different aspects of the present invention as described above, which are not provided in detail for the sake of simplicity. Although the present invention has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A visual question answering optimization method based on hierarchical multimodal fine-tuning, characterized in that: Includes steps: S1. Obtain text-image pairs for visual question answering tasks, including images and corresponding texts; S2. Use the visual encoder of the pre-trained CLIP model to extract the visual features of the image, and combine the features of each level extracted by the visual encoder into a visual feature set; S3, using the text encoder of the pre-trained CLIP model to extract text features of the text, and composing the features of each level extracted by the text encoder into a text feature set; S4, adaptively weighting and adjusting the text features of each layer in the text feature set; S5, concatenating the weighted text features of each layer into a vector and performing linear projection; S6, concatenate the visual features of each layer in the visual feature set into a vector, and perform cross-modal interactive fusion with the text features after linear projection through a multi-head attention mechanism; S7, fusing the feature vector obtained by cross-modal interactive fusion with the vector obtained by concatenating the visual features of each layer in the visual feature set to obtain a semantic visual perception fusion feature; S8. Use semantic visual perception fusion features for visual question answering tasks.
2. The visual question answering optimization method according to claim 1, characterized in that: The visual encoder and the text encoder are fine-tuned in multiple layers and stages by low-rank adaptation technology, including the following steps: Model layer group division: The network layers of the visual encoder and text encoder are divided into front layer, middle layer and back layer according to the depth; First stage fine-tuning: Add LoRA adapters to the back layers of the visual encoder and text encoder respectively, keep the pre-trained parameters of the visual encoder and text encoder unchanged, and only open the LoRA adapter inserted in the back layer for training. During the training process, only the parameters of the LoRA adapter are updated; The second stage of fine-tuning: add LoRA adapters to the middle layer of the visual encoder and text encoder respectively, keep the pre-trained parameters of the visual encoder and text encoder unchanged, and only open the LoRA adapters inserted in the back layer and middle layer for training. During the training process, only the parameters of the LoRA adapter are updated.
3. The visual question answering optimization method according to claim 2, characterized in that: The visual encoder and text encoder of the pre-trained CLIP model both contain 12 Transformer layers. In the model layer group division, the 1st to 4th Transformer layers are used as the front layers, the 5th to 8th Transformer layers are used as the middle layers, and the 8th to 12th Transformer layers and the projection layer are used as the back layers.
4. The visual question answering optimization method according to claim 2, characterized in that: It also includes the third stage of fine-tuning: adding LoRA adapters to the front layers of the visual encoder and text encoder respectively, keeping the pre-trained parameters of the visual encoder and text encoder unchanged, and only opening the LoRA adapters inserted in the back, middle and front layers for training. During the training process, only the parameters of the LoRA adapters are updated.
5. The visual question answering optimization method according to claim 4, characterized in that: In the third stage of fine-tuning, the parameters of the visual encoder, text encoder, and middle and back layer LoRA adapters are first kept unchanged, and only the LoRA adapter inserted in the front layer is opened for low learning rate training; then the middle and back layer LoRA adapters are opened and fine-tuned together for several rounds at a very small learning rate to make all LoRA adapter parameters reach the overall coordinated optimal.
6. The visual question answering optimization method according to claim 1, characterized in that: In step S4, the adaptive weighted adjustment is expressed as: , in, The weighted adjusted i Layer text features, Represents the corresponding element multiplication operation, represents the addition operation of corresponding elements, Represents adaptive weight W Corresponding to No. i The layer weight vector, Represents the text encoder i The text features extracted by the layer.
7. The visual question answering optimization method according to claim 6, characterized in that: In step S5, the linear projection is expressed as: in, is the text feature after linear projection, concat represents vector concatenation, represents the matrix multiplication operation, is the linear projection weight.
8. The visual question answering optimization method according to claim 1, characterized in that: In step S6, the text features after linear projection are used as the key vector K and the value vector V, and the vector concatenated from the visual features of each layer in the visual feature set is used as the query vector Q, and the cross-modal interactive fusion features are obtained through the multi-head attention mechanism.
9. The visual question answering optimization method according to claim 1, characterized in that: In step S7, a residual connection is used to add the corresponding elements of the cross-modal interaction fusion features output by the multi-head attention mechanism and the vector concatenated by the visual features of each layer in the visual feature set to obtain the semantic visual perception fusion features: , in, f is the semantic visual perception fusion feature, For cross-modal interactive fusion features, Represents the vector concatenated from each layer of visual features in the visual feature set. Represents the addition operation of corresponding elements.
10. A visual question answering optimization system based on the method according to any one of claims 1 to 9, characterized in that: It includes visual feature extraction module, text feature extraction module, weighted projection module and cross-modal bridging module; The visual feature extraction module is used to extract the visual features of the image in the visual question answering task using the visual encoder of the pre-trained CLIP model, and to form a visual feature set from the features of each level extracted by the visual encoder; The text feature extraction module is used to extract text features of text in the visual question answering task using the text encoder of the pre-trained CLIP model, and to form a text feature set from the features of each level extracted by the text encoder; The weighted projection module is used to perform adaptive weighted adjustment on the text features of each layer in the text feature set, and splice the weighted adjusted text features of each layer into a vector for linear projection; The cross-modal bridging module is used to perform cross-modal fusion of the linearly projected text features with the vectors formed by splicing the visual features of each layer in the visual feature set through a multi-head attention mechanism and residual connection, so as to obtain semantic visual perception fusion features for visual question answering tasks; It also includes a LoRA fine-tuning module for performing multi-layer group-by-stage fine-tuning on the visual encoder and the text encoder through a low-rank adaptation technique.
Citation Information
Patent Citations
Multi-round visual dialogue method based on multi-level sorting learning
CN113435399A
Universal multi-modal learning method based on deep interactive adaptation network model
CN116882477A
Multi-modal emotion recognition method based on multi-level interaction
CN118152919A
Remote sensing image semantic segmentation method based on Mama and Transform architecture fusion
CN119206227A
Multi-modal data associative learning model training method and device
JP2022137145A
Cited By
Intelligent question answering method and device based on domain knowledge lightweight fine tuning and cross-domain dynamic knowledge base and readable storage medium thereof
CN120296110A
Chart visual question-answering method based on dynamic routing and low-rank mixing and related device
CN120723951A
Alignment method based on natural language and machine vision
CN121117953A
Nerve symbol learning-based natural language problem programmed analysis method
CN121859889A
Multi-modal large model continuous learning method oriented to visual question and answer task
CN122133755A