Visual language multi-modal fusion method based on parameter-free cross attention

Through parameter-free cross-attention calculation and multi-scale visual feature generation mechanism, the problems of high computational overhead and parameter redundancy in visual-language multimodal fusion are solved, and efficient visual-language task performance improvement is achieved, especially in picture-text question answering and image generation tasks.

CN120654176APending Publication Date: 2025-09-16BEIJING INST OF TECH

Patent Information

Application Number
CN202510704740.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing vision-language multimodal fusion methods suffer from high computational overhead, redundant parameters in cross-modal interaction modules, and uneven contribution of visual features to language understanding under limited computing resources, resulting in unbalanced and insufficient modal fusion.

Method used

A parameter-free cross-attention calculation method is adopted, combined with a multi-scale visual feature generation mechanism and a dynamic feature selection strategy, and a language modality-guided cross-modal attention design is used to reduce the model parameter scale and improve the semantic alignment effect.

Benefits of technology

It achieves efficient multimodal fusion of visual language under limited computing resources, improves the performance of multimodal tasks, especially in tasks such as image-text question answering, image generation, and multimodal instruction understanding, improving model accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654176A_ABST
    Figure CN120654176A_ABST
Patent Text Reader

Abstract

The invention discloses a visual language multi-modal fusion method based on parameter-free cross attention, and belongs to the field of computer vision. The implementation method comprises the following steps: using a fixed pre-training language model as a trunk, using a visual encoder to extract image features, and calculating a cross attention weight between language query and visual features through a parameter-free activation function, replacing a plurality of groups of learnable projection matrixes introduced by a traditional cross attention module, and significantly reducing the model parameter scale. A multi-scale visual feature generation mechanism based on pooling operation is introduced, and rich visual semantic prompt information is provided for a language model. A dynamic feature selection module is designed in combination with cross attention, visual areas corresponding to all text tokens are screened, low-correlation areas are discarded, only visual content more contributing to the current language context is reserved, accurate information matching and efficient fusion between modals are achieved, and the accuracy and efficiency of information fusion are improved. And the performance of the visual language model in tasks such as image-text question answering, image generation and multi-modal instruction understanding is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multimodal fusion method, in particular to a visual language multimodal fusion method based on parameter-free cross-attention, and belongs to the field of computer vision. Background Art

[0002] With the rapid development of artificial intelligence technology, the Vision-Language Model (VLM) has demonstrated outstanding performance in joint image-text modeling tasks and has received widespread attention. As an important branch of multimodal learning, VLM has made breakthrough progress in core tasks such as image description, visual question answering, and command following by integrating visual and language modal information, greatly expanding the application boundaries of large-scale language models in perception and reasoning. Compared with traditional single-modality processing methods, the Vision-Language Model can more comprehensively understand the semantic information contained in the image content, and make more accurate reasoning and judgments in combination with the natural language context. However, despite the excellent results currently achieved by VLM in multiple tasks, its huge computational overhead and resource consumption still restrict its ability to be implemented in actual scenarios.

[0003] Existing mainstream methods typically use cross-attention mechanisms to achieve deep fusion of visual and linguistic modalities, or use input splicing to directly feed image features into a language model. Cross-attention mechanisms primarily insert additional fusion modules into the intermediate layers of a pre-trained language model, using the linguistic modality as a query and the visual modality as key / value pairs for information exchange. These methods enable more fine-grained modality fusion, facilitating the model's learning of complex image-text correspondences. However, the additional attention modules required at each layer introduce a large number of trainable parameters, significantly increasing the model's parameter size and memory usage. On the other hand, input splicing transforms image features into pseudo "language tokens" and directly appends them to the text input sequence, where they are uniformly processed by a language model. This approach offers a simpler structure, is easier to implement, and does not rely on complex module insertion. However, to achieve image-text alignment, it typically requires feature mapping using a high-dimensional projector, which is trained throughout the fine-tuning phase. Furthermore, the projection process loses crucial visual topological information (such as the relative position of pixels), leading to erroneous understanding of local image regions during the fusion process.

[0004] Therefore, how to design a visual language modeling mechanism that is both fusion-expressive and efficient under limited computing resources has become a key issue in current research. Summary of the Invention

[0005] In order to address the problems in existing visual-language multimodal fusion methods, such as high computational overhead caused by input sequence length extension, redundant trainable parameters in cross-modal interaction modules, and uneven contribution of visual features to language understanding, which lead to unbalanced and insufficient fusion of different modalities, the present invention aims to provide a visual-language multimodal fusion method based on parameter-free cross-attention. This method is based on an efficient fusion mechanism of cross-modal information screening, a language modality-guided cross-modal attention design and a multi-scale visual feature prompt mechanism, and a parameter-free cross-attention calculation method and a dynamic visual feature selection strategy. While maintaining the original language model structure, it reduces model complexity, improves semantic alignment effect and model accuracy in visual-language tasks, and achieves efficient visual-language multimodal fusion.

[0006] The main purpose of the present invention is achieved through the following technical solutions:

[0007] The present invention discloses a method for visual-language multimodal fusion based on parameter-free cross-attention, which adopts a fixed pre-trained language model as the backbone, uses a visual encoder to extract image features, and calculates the cross-attention weights between language queries and visual features through a parameter-free activation function, replacing the multiple sets of learnable projection matrices introduced in the traditional cross-attention module, thereby significantly reducing the scale of model parameters. At the same time, a multi-scale visual feature generation mechanism based on pooling operation is introduced to provide rich visual semantic prompt information for the language model without increasing the visual forward computation cost. A dynamic feature selection module is designed in combination with cross-attention to screen the visual area corresponding to each text token, discard low-correlation areas, and retain only visual content that contributes more to the current language context, thereby achieving accurate information matching and efficient fusion between modalities, and improving the performance of visual language models in tasks such as picture-text question answering, image generation, and multimodal instruction understanding.

[0008] The present invention discloses a method for multimodal fusion of vision and language based on parameter-free cross-attention, comprising the following steps:

[0009] Step 1: Extract multi-scale visual features based on the image encoder, convert the image block features into a two-dimensional grid through progressive pooling operations, and generate downsampled features of different scales. After splicing, a visual representation containing local to global semantics is formed. At the same time, the classification label [cls] is used as a global semantic guide and spliced ​​with the text sequence to realize the introduction of multi-scale information in cross-attention fusion and improve the accuracy of the multimodal model.

[0010] Step 1.1: Input an image with a resolution of N×N into the visual encoder. The input image is divided into C non-overlapping image blocks. After encoding, the visual encoder outputs C visual feature vectors; each visual feature vector corresponds to the semantic representation of a local image region; adjacent visual feature vectors are merged through pooling operations to provide multi-scale visual features for the language model. That is, C one-dimensional visual feature vectors are rearranged into an M×M two-dimensional feature grid according to their original spatial positions. Pooling operations with different kernel sizes are used to obtain downsampled feature maps of different scales; the downsampled feature maps of each scale are flattened and spliced ​​through the channel dimension to form the final multi-scale visual features;

[0011] Step 1.2: The classification tag [cls] output by the encoder is used as the global semantic summary of the input image. The feature vector of [cls] contains high-level semantic information of the entire image. Following the input space hint method, it is connected to the starting position of the text tag as a prefix tag to aggregate global information.

[0012] Step 2: Construct a parameter-free cross-attention mechanism, introduce the nonlinear transformation function SiLU as the feature mapping function, decompose the projection matrix into a parameter-free mapping function, fuse the visual features into the language feature space, and obtain the similarity score matrix. The parameter-free cross-attention-based fusion is achieved through weighted method, which reduces the computational complexity and improves the parameter efficiency of the model.

[0013] Step 2.1: Given a natural language sequence input, it is encoded to obtain text tokens As the query object q, L is the length of the text sequence, d is the feature dimension; the multi-scale visual features obtained in step 1 As key k and value v, N is the number of visual features, d' is the dimension of the visual feature. The query object and key are projected through the matrix to map the features to dimension d k , the cross attention weight is obtained by scaling and softmax function, and the cross attention weight is used to weight the value vector V to obtain the fused feature. The fused feature expression is shown in formula (1):

[0014]

[0015] Where, are the learned projection matrices for query, key, and value, respectively. Q, K, and V are the query vector for text features, the key vector for visual features, and the value vector for visual features, respectively. is the output projection matrix, used to adjust the output dimensions.

[0016] Step 2.2: Introduce the nonlinear transformation function SiLU as the feature mapping function, decompose the projection matrix into a parameter-free mapping function, fuse the visual features into the language feature space, and obtain the similarity score matrix. Through a weighted method, parameter-free cross-attention-based fusion is achieved to reduce the computational complexity and improve the parameter efficiency of the model.

[0017] Perform dot product on each query vector Q in step 2.1 with all key vectors K, and obtain the cross-attention weight through the softmax function. The cross-attention weight can be regarded as the similarity between the query vector Q and the key vector K. Use the cross-attention weight to perform weighted summation on the value vector to obtain the fusion result. The fusion result expression is shown in formula (2):

[0018]

[0019] in, Calculate the i-th query vector Q i With the j-th key vector K j The similarity between them is determined by the constraint that sin(Q i ,K j ) must be non-negative, so sin(Q i ,K j ) can be any kernel function

[0020] Define the kernel function as k(x,y)=φ(x)φ(y) T , φ is a projection function, and the cross attention is calculated as Select the SiLU activation function as the parameter-free projection function and transform the projection matrix W Q and W K The low-rank decomposition is a parameter-free projection function. The cross attention is obtained through the parameter-free projection function, that is, the similarity score matrix. The cross attention is used for weighted summation to obtain the fusion feature. The expression is shown in formula (3):

[0021] XAttn(X l ,X v )=φ(X l )φ(X v ) T X v , (3)

[0022] Among them, φ(X l )φ(X v ) T Represents the similarity score matrix between each text tag and each visual feature; φ(·)=SiLU(·); Output XAttn(X l ,Xv ) is the fusion result; that is, the visual features are fused into the language feature space to achieve parameter-free cross-attention-based fusion.

[0023] Step 3: Sort the similarity score matrix obtained in step 2, i.e., the cross-attention matrix, and mask some low-scoring visual features to form a dynamic binary mask. This allows the text tags to dynamically extract visual information in the parameter-free cross-attention module, focusing more on information-rich visual features, reducing the interference of redundant features, and improving the fusion effect of text and image.

[0024] In formula (3) represents the similarity score between each text token and each visual feature, where L represents the length of the text sequence and N represents the number of visual features. A mask is generated for each row of the similarity score matrix, creating a binary mask vector M. This is initialized with an all-one vector and the lowest value is masked using the masking rate γ. This means that for all elements ranked after the unmasked rate in the sorted sequence, the mask values ​​at their corresponding positions are set to 0. The original cross-attention score is element-wise multiplied by the mask, discarding redundant visual features. Since the sum of the attention scores for each text token is not 1, the remaining attention scores after deleting some features retain their original numerical values, reducing the number of visual features involved in the weighted summation. During backpropagation, the gradient propagates through the unmasked feature positions.

[0025] Visual feature selection based on parameter-free cross-attention can adapt to different text tags and input images, making each text tag focus more on useful information and achieving better fusion results.

[0026] Beneficial effects:

[0027] 1. The present invention discloses a method for multimodal visual language fusion based on parameter-free cross-attention. It converts image block features into two-dimensional grids through progressive pooling, and generates downsampled features of different scales. After splicing, it forms a visual representation containing local to global semantics. At the same time, the classification label [cls] is used as a global semantic guide and spliced ​​with the text sequence to realize the introduction of multi-scale information in cross-attention fusion. Compared with traditional multi-scale visual features, it realizes the combination of global and local information of visual images, improves the representation effect of visual features, and improves the accuracy of multimodal visual language models.

[0028] 2. This invention discloses a parameter-free cross-attention-based visual-linguistic multimodal fusion method. This method constructs a parameter-free attention mechanism, introduces the nonlinear transformation function SiLU as a feature mapping function, performs a low-rank decomposition of the projection matrix, fuses visual features into the language feature space, and obtains a similarity score matrix. This method then achieves parameter-free attention-based fusion through a weighted approach. Compared with traditional cross-attention mechanisms, this method reduces computational complexity and improves the model's parameter efficiency.

[0029] 3. The present invention discloses a visual language multimodal fusion method based on parameter-free cross-attention. For the cross-attention score matrix generated for each text tag, it sorts and shields some visual features with low scores to form a dynamic binary mask, so that the text tag can dynamically extract visual information in the parameter-free cross-attention module. Compared with the traditional cross-attention strategy, it focuses more on information-rich visual features, reduces the interference of redundant features, and effectively improves the fusion effect of text and image. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 This is a schematic diagram of the framework of a visual-language multimodal fusion method based on parameter-free cross-attention disclosed in this example.

[0031] Figure 2 This is a flowchart of a visual-language multimodal fusion method based on parameter-free cross-attention disclosed in this example. DETAILED DESCRIPTION

[0032] The present invention will be described in detail below with reference to the accompanying drawings and embodiments, and the technical problems solved by the technical solution of the present invention and the beneficial effects thereof will be discussed. It should be noted that the described embodiments are intended to facilitate understanding of the present invention and do not have any limiting effect on the present invention.

[0033] Example 1

[0034] like Figure 1As shown, this embodiment discloses a parameter-free cross-attention-based vision-language multimodal fusion method that fine-tunes a pre-trained language model for vision-language tasks. LLaMA (a family of base language models with parameters ranging from 7 billion to 65 billion) is used as the language model. The experiments primarily use the 7- and 13-billion-parameter versions. To extract visual information, the CLIP model with ViT-L / 14 as the backbone network is used as the image encoder. Training is performed on the ScienceQA dataset (a multimodal dataset for scientific question answering. The dataset includes plain text and image-text samples covering three subjects, 26 topics, and 379 skills, divided into training, validation, and test sets. Average precision on the test set is reported) for 20 epochs with a global batch size of 32. The learning rate is initially set to 9e-3 and decayed using a cosine annealing strategy. For visual cues, dual-scale features are used: original visual features and features downsampled by average pooling as multi-scale visual cues.

[0035] like Figure 2 As shown, this embodiment discloses a method for multimodal fusion of vision and language based on parameter-free cross attention, and the specific implementation steps are as follows:

[0036] Step 1: Extract multi-scale visual features based on the CLIP image encoder, convert image block features into a two-dimensional grid through progressive pooling operations, and generate downsampled features of different scales. After splicing, a visual representation containing local to global semantics is formed. At the same time, the classification label [cls] is used as a global semantic guide and spliced ​​with the text sequence to realize the introduction of multi-scale information in cross-attention fusion and improve the accuracy of the multimodal model.

[0037] Step 1.1: Input an image with a resolution of N×N into the visual encoder CLIP. The input image is divided into C non-overlapping image blocks. After encoding, the visual encoder CLIP outputs C visual feature vectors; each visual feature vector corresponds to the semantic representation of a local image region; adjacent visual feature vectors are merged through pooling operations to provide multi-scale visual features to the language model. That is, C one-dimensional visual feature vectors are rearranged into an M×M two-dimensional feature grid according to their original spatial positions. Pooling operations with different kernel sizes are used to obtain downsampled feature maps of different scales; the downsampled feature maps of each scale are flattened and spliced ​​through the channel dimension to form the final multi-scale visual features;

[0038] Table 1 shows the experimental results comparing the effects of different scales and methods on obtaining multi-scale features. The effect of downsampling scale was investigated using a two-dimensional average pooling method. Downsampling operations with kernel sizes of 2 and 4 were used, resulting in 64 and 16 downsampled visual features, respectively. When using only visual features at a single scale, performance was lower than using the full 256 visual features, regardless of whether the number of features was downsampled to 64 or 16 features. The performance degradation became more pronounced as the number of features decreased.

[0039] When constructing visual cues by concatenating multi-scale features, downsampling with a kernel size of 2 effectively improves model performance. However, increasing the kernel size to 4 often results in a decrease in performance. Excessive downsampling can lead to overly coarse features, hindering the model from learning fine-grained information. Therefore, when designing multi-scale visual cues, it is important to avoid excessive scale differences. Furthermore, the effects of maximum pooling and average pooling were compared under optimal scale configurations. Experimental results show that average pooling is the superior downsampling method.

[0040] Table 1 Comparison of different downsampling methods and scales for generating multimodal visual cues using LLaMA-7B as the language model

[0041]

[0042]

[0043] Step 1.2: The classification tag [cls] output by the encoder CLIP is used as the global semantic summary of the input image. The feature vector of [cls] contains high-level semantic information of the entire image. Following the input space hint method, it is connected to the starting position of the text tag as a prefix tag to aggregate global information.

[0044] To verify the compatibility of intermediate feature fusion and input-level fusion methods, we compared the effects of introducing visual features of different scales at the input stage. As shown in Table 2, the model achieves optimal performance when only the [cls] tag is used as an additional input in addition to the text markup. However, introducing other visual features at the input stage significantly degrades performance, regardless of the scale used. This indicates that this information has been fully integrated during the intermediate-level fusion process, and repeated introduction at the input stage interferes with the feature fusion mechanism, resulting in performance degradation.

[0045] Table 2 Comparison of different input-level fusion schemes using LLaMA-7B as the language model

[0046] Visual Marker Dimension Whether to connect the classification mark [cls] Average accuracy (%) 0 no 92.97 0 yes 93.85 64 no 92.47 64 yes 92.86 256 no 89.86 256 yes 90.17

[0047] Step 2: Construct a parameter-free cross-attention mechanism, introduce the nonlinear transformation function SiLU as the feature mapping function, decompose the projection matrix into a parameter-free mapping function, fuse the visual features into the language feature space, and obtain the similarity score matrix. The parameter-free cross-attention-based fusion is achieved through weighted method, which reduces the computational complexity and improves the parameter efficiency of the model.

[0048] Step 2.1: Given a natural language sequence input, it is encoded to obtain text tokens As the query object q, L is the length of the text sequence, d is the feature dimension; the multi-scale visual features obtained in step 1 As key k and value v, N is the number of visual features, d' is the dimension of the visual feature. The query object and key are projected through the matrix to map the features to dimension d k , the cross attention weight is obtained by scaling and softmax function, and the cross attention weight is used to weight the value vector V to obtain the fused feature. The fused feature expression is shown in formula (1):

[0049]

[0050] Where, are the learned projection matrices for query, key, and value, respectively. Q, K, and V are the query vector for text features, the key vector for visual features, and the value vector for visual features, respectively. is the output projection matrix, used to adjust the output dimensions.

[0051] Step 2.2: Perform dot product on each query vector Q in step 2.1 and all key vectors K, and obtain the cross attention weight through the softmax function. The cross attention weight can be regarded as the similarity between the query vector Q and the key vector K. Use the cross attention weight to perform weighted summation on the value vector to obtain the fusion result. The fusion result expression is shown in formula (2):

[0052]

[0053] in, Calculate the i-th query vector Q i With the j-th key vector K j The similarity between them is determined by the constraint that sin(Q i ,K j ) must be non-negative, so sin(Q i ,K j ) can be any kernel function

[0054] Define the kernel function as k(x,y)=φ(x)φ(y) T, φ is a projection function, and the cross attention is calculated as Select the SiLU activation function as the parameter-free projection function and transform the projection matrix W Q and W K The low-rank decomposition is a parameter-free projection function. The cross attention is obtained through the parameter-free projection function, that is, the similarity score matrix. The cross attention is used for weighted summation to obtain the fusion feature. The expression is shown in formula (3):

[0055] XAttn(X l ,X v )=φ(X l )φ(X v ) T X v , (3)

[0056] Among them, φ(X l )φ(X v ) T Represents the similarity score matrix between each text tag and each visual feature; φ(·)=SiLU(·); Output XAttn(X l ,X v ) is the fusion result; that is, the visual features are fused into the language feature space to achieve parameter-free cross-attention-based fusion.

[0057] To explore the impact of different projection functions, we compare various schemes in Table 3. The experimental results show that identity projection achieves an average accuracy of 92.16%, while using more complex projections (such as softmax or ReLU) does not guarantee performance improvement. In contrast, using SiLU as the projection function achieves an accuracy of 93.85%.

[0058] Table 3 Comparison of different non-parametric linear projections using LLaMA-7B as the language model

[0059]

[0060]

[0061] The parameter-free cross-attention-based visual-language multimodal fusion method was evaluated on the ScienceQA dataset, and the results are shown in Table 4. Experiments show that the parameter-free cross-attention-based visual-language multimodal fusion method achieved the best average precision among all compared methods. The parameter-efficient fine-tuning (PEFT) method outperformed the zero-shot / small-shot methods and the traditional full-parameter fine-tuning method. Among the PEFT methods, when LLaMA-7B was used as the large language model, the parameter-free cross-attention-based visual-language multimodal fusion method surpassed the second place by 0.78% with an average accuracy of 93.85%. When using LLaMA-13B, the accuracy was further improved to 94.55%, 0.77% higher than the second best result.

[0062] At the same time, experiments were conducted using the more advanced LLaMA2-7B as a pre-trained model (maintaining the same visual-linguistic tuning scheme as LLaMA-7B). The results showed that the performance of MemVP based on LLaMA2-7B was lower than its LLaMA-7B version, while the visual-linguistic multimodal fusion method based on parameter-free cross-attention showed stronger compatibility and outperformed the original LLaMA-7B version. This shows that when equipped with a more powerful pre-trained natural language model (LLM), the visual-linguistic multimodal fusion method based on parameter-free cross-attention has the potential to achieve better performance.

[0063] Table 4 Comparison of test results of different VL methods on the ScienceQA dataset

[0064]

[0065] Step 3: Sort the similarity score matrix obtained in step 2, i.e., the cross-attention matrix, and mask some low-scoring visual features to form a dynamic binary mask. This allows the text tags to dynamically extract visual information in the parameter-free cross-attention module, focusing more on information-rich visual features, reducing the interference of redundant features, and effectively improving the fusion effect of text and image.

[0066] In formula (3) Represents the similarity score between each text tag and each visual feature, where L represents the length of the text sequence and N represents the number of visual features. A mask is generated separately for each row of the similarity score matrix, a binary mask vector M is created, and a full 1 vector is initialized. The lowest value is masked using the masking rate γ, that is, for all elements ranked after the non-masking rate in the sorted sequence, the mask value of the corresponding position is set to 0, and the original cross-attention score is multiplied element-by-element by the mask to discard redundant visual features. Since the sum of the attention scores of each text tag is not 1, the remaining attention scores after deleting some features maintain the original numerical value, reducing the number of visual features involved in the weighted summation; during the backpropagation process, the gradient will propagate through the unmasked feature positions. The overall process is as follows Figure 2 shown.

[0067] We conducted ablation experiments on the ScienceQA dataset using LLaMA-7B to investigate the effectiveness of the dynamic feature selection strategy. The results are shown in Table 5. After removing the least important visual features using dynamic feature selection, the model achieved an average accuracy of 93.85%. This demonstrates that the dynamic selection strategy helps improve the performance of the fine-tuned VL model.

[0068] Table 5 Impact of dynamic feature selection on model performance

[0069] Whether to use dynamic feature selection Average accuracy (%) yes 93.85 no 93.47

[0070] Regarding the selection of the shielding rate γ, an ablation experiment was designed to compare different configurations of the shielding rate. The results are shown in Table 6. As γ increases, the performance gradually improves and reaches a peak when γ = 0.2. At the same time, further increasing γ has a marginal benefit on adaptive fusion.

[0071] Table 6 Effect of different shielding rates on model performance

[0072] Shielding rate γ Average accuracy (%) 0.0 93.48 0.05 93.67 0.1 93.74 0.2 93.85 0.3 93.51

[0073] Therefore, the present invention discloses a visual language multimodal fusion method based on parameter-free cross-attention, an efficient fusion mechanism based on cross-modal information screening, a cross-modal attention design guided by language modality and a multi-scale visual feature prompt mechanism, and a parameter-free cross-attention calculation method and a dynamic visual feature selection strategy. While maintaining the original language model structure, the model complexity is reduced, the semantic alignment effect and model accuracy in visual language tasks are improved, and efficient visual language multimodal fusion is achieved, providing theoretical support for visual language multimodal information fusion processing.

[0074] The above specific description further illustrates the purpose, technical solutions and beneficial effects of the invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A parameter-free cross-attention-based vision-language multimodal fusion method, characterized by: The steps include: Step 1: Extract multi-scale visual features based on the image encoder. Convert image block features into a two-dimensional grid through progressive pooling operations, and generate downsampled features of different scales. After splicing, form a visual representation that contains local to global semantics. At the same time, the classification label [cls] is used as a global semantic guide and spliced ​​with the text sequence to realize the introduction of multi-scale information in cross-attention fusion and improve the accuracy of the multimodal model. Step 2: Construct a parameter-free cross-attention mechanism, introduce the nonlinear transformation function SiLU as the feature mapping function, decompose the projection matrix into a parameter-free mapping function, fuse the visual features into the language feature space, and obtain the similarity score matrix. The parameter-free cross-attention-based fusion is achieved through weighted method, which reduces the computational complexity and improves the parameter efficiency of the model. Step 3: Sort the similarity score matrix obtained in step 2, i.e., the cross-attention matrix, and mask some low-scoring visual features to form a dynamic binary mask. This allows the text tags to dynamically extract visual information in the parameter-free cross-attention module, focusing more on information-rich visual features, reducing the interference of redundant features, and improving the fusion effect of text and image.

2. The method for multimodal visual-linguistic fusion based on parameter-free cross-attention according to claim 1, characterized in that: Step 1 is implemented as follows: Step 1.

1. Input an image with a resolution of N×N into the visual encoder, and divide the input image into C non-overlapping image blocks. After encoding, the visual encoder outputs C visual feature vectors; each visual feature vector corresponds to the semantic representation of a local image area; adjacent visual feature vectors are merged through pooling operations to provide multi-scale visual features for the language model, that is, C one-dimensional visual feature vectors are rearranged into an M×M two-dimensional feature grid according to their original spatial positions, and pooling operations with different kernel sizes are used to obtain downsampled feature maps of different scales; after flattening, the downsampled feature maps of each scale are spliced ​​through the channel dimension to form the final multi-scale visual features. Step 1.2: The classification tag [cls] output by the encoder is used as the global semantic summary of the input image. The feature vector of [cls] contains high-level semantic information of the entire image. Following the input space hint method, it is connected to the starting position of the text tag as a prefix tag to aggregate global information.

3. The method for visual-linguistic multimodal fusion based on parameter-free cross-attention according to claim 2, characterized in that: Step 2 is implemented as follows: Step 2.1: Given a natural language sequence input, get a text token after encoding As the query object q, L is the length of the text sequence, and d is the feature dimension; Multi-scale visual features obtained in step 1 As key k and value v, N is the number of visual features, d′ is the dimension of the visual feature; the query object and key are projected through the matrix to map the feature to dimension d k , cross attention weights are obtained by scaling and softmax function, and the cross attention weights are used to weight the sum of the value vector V to obtain the fused features; Step 2.2: Introduce the nonlinear transformation function SiLU as the feature mapping function, decompose the projection matrix into a parameter-free mapping function, fuse the visual features into the language feature space, and obtain the similarity score matrix. The parameter-free cross-attention-based fusion is achieved through weighted method, which reduces the computational complexity and improves the parameter efficiency of the model.

4. The method for visual-linguistic multimodal fusion based on parameter-free cross-attention according to claim 3, characterized in that: In step 2.1, The fusion feature expression is shown in formula (1): Where, are the learned projection matrices for query, key, and value, respectively. Q, K, and V are the query vector for text features, the key vector for visual features, and the value vector for visual features, respectively. is the output projection matrix, used to adjust the output dimensions.

5. The method for visual-linguistic multimodal fusion based on parameter-free cross-attention according to claim 4, characterized in that: Step 2.2 is implemented as follows: Perform dot product on each query vector Q in step 2.1 with all key vectors K, and obtain the cross-attention weight through the softmax function. The cross-attention weight is equivalent to the similarity between the query vector Q and the key vector K. Use the cross-attention weight to perform weighted summation on the value vector to obtain the fusion result. The fusion result expression is shown in formula (2): in, Calculate the i-th query vector Q i With the j-th key vector K j The similarity between them is determined by the constraint that sin(Q i ,K j ) must be non-negative, so sin(Q i ,K j ) is an arbitrary kernel function Define the kernel function as k(x,y)=φ(x)φ(y) T , φ is a projection function, and the cross attention is calculated as Select the SiLU activation function as the parameter-free projection function and transform the projection matrix W Q and W K The low-rank decomposition is converted into a parameter-free projection function, and the cross attention is obtained through the parameter-free projection function, that is, the similarity score matrix, and the cross attention is used for weighted summation to obtain the fusion feature.

6. The method for visual-linguistic multimodal fusion based on parameter-free cross-attention according to claim 5, characterized in that: In step 2.2, Use cross attention to perform weighted summation to obtain the fusion feature, as shown in formula (3): XAttn(X l ,X v )=φ(X l )φ(X v ) T X v , (3) Among them, φ(X l )φ(X v ) T Represents the similarity score matrix between each text tag and each visual feature; φ(·)=SiLU(·); Output XAttn(X l ,X v ) is the fusion result; that is, the visual features are fused into the language feature space to achieve parameter-free cross-attention-based fusion.

7. The method for visual-linguistic multimodal fusion based on parameter-free cross-attention according to claim 6, characterized in that: In step 3, In formula (3) Represents the similarity score between each text token and each visual feature, where L represents the length of the text sequence and N represents the number of visual features. A mask is generated separately for each row of the similarity score matrix, a binary mask vector M is created, and an all-1 vector is initialized. The lowest value is masked using the masking rate γ, that is, for all elements ranked after the non-masking rate in the sorted sequence, the mask value of the corresponding position is set to 0, and the original cross-attention score is multiplied element-by-element by the mask to discard redundant visual features. Since the sum of the attention scores of each text token is not 1, the remaining attention scores after deleting some features maintain the original numerical value, reducing the number of visual features involved in the weighted summation. During the backpropagation process, the gradient will be propagated through the unmasked feature positions.

8. The method for visual-linguistic multimodal fusion based on parameter-free cross-attention according to claim 7, characterized in that: In step 3, the visual feature selection based on parameter-free cross-attention is adapted to different text tags and input images, so that each text tag pays more attention to useful information and achieves better fusion results.

Citation Information

Patent Citations

  • Target counting method and system based on multi-modal multi-scale cross attention

    CN119785057A

Cited By

  • Self-adaptive three-dimensional large language model system based on query guidance

    CN120849595A

  • Adaptive text-guided fiber bundle feature fusion method and system based on large language model

    CN121030694A

  • Power robot control method and system based on visual voice action model

    CN121043156A

  • Multi-modal large language model reasoning method based on parallel visual token scheduling

    CN121189504A

  • A multi-modal large language model inference method based on parallel visual token scheduling

    CN121189504B