A Visual Question Answering Method, Device and Medium Based on Multimodal Large Model

By meshing and interpolation of low-resolution images, high-resolution visual features are generated and combined with text features are fused, the problems of missing details and rising hallucinations under low-resolution image input are solved, and the accuracy and reliability of visual questions and answers are improved, and suitable for mobile devices and edge computing nodes.

CN119938872BActive Publication Date: 2025-07-18INSPUR GENERSOFT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510429240.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-18
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The existing visual question-and-answer technology has problems of missing details and rising hallucinations in low-resolution image input scenarios, which leads to the improvement of the accuracy of question-and-answer results.

Method used

By obtaining the original Q&A image data input by the user, grid processing and interpolation operations are performed to generate high-resolution visual feature data, and using high-resolution visual feature data for feature enhancement, combining Q&A text features for feature fusion to generate answers.

Benefits of technology

It effectively compensates for the high-frequency details lost by low-resolution images due to downsampling compression, reduces visual hallucinations, reduces computing volume and memory usage, and is suitable for mobile devices and edge computing nodes, improving the accuracy and reliability of Q&A systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938872B_ABST
    Figure CN119938872B_ABST
Patent Text Reader

Abstract

An embodiment of this specification discloses a visual question answering method, device, and medium based on a multimodal large model, which relates to the technical field of data processing. The method includes: obtaining the original question answering image data and original question answering text data input by the user, converting the original question answering image data to determine the corresponding high-resolution visual feature data; using the high-resolution visual feature data to enhance the features of the original visual features corresponding to the pre-obtained original question answering image data to determine the enhanced visual token features; extracting the question answering text features of the original question answering text data, determining a comprehensive feature vector based on the enhanced visual token features and the question answering text features, and generating an answer through the multimodal large model and the comprehensive feature vector. Through the targeted processing and feature enhancement of the original image data, while ensuring the acquisition of key details, relatively low computational complexity is maintained, meeting the resource limitations in practical applications and broadening the model application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the technical field of data processing, and particularly to a visual question answering method, device, and medium based on a multimodal large model. Background Art

[0002] Visual Question Answering (VQA) technology is an interdisciplinary field of computer vision and natural language processing. The VQA task can be specifically described as follows: Given a picture and a text question, the model extracts image features from the given image, extracts text features from the text question, performs feature fusion processing on the two types of modal features, and finally generates the best answer based on the fused features.

[0003] In the Visual Question Answering (VQA) task, the input resolution of the visual modality directly affects the model's ability to capture image details and the reliability of context reasoning. When the existing technology uses a low-resolution image (such as 224×224 pixels) as the input, the low-resolution image loses high-frequency details (such as text, texture, small objects) due to downsampling compression, resulting in the model being unable to accurately identify key visual cues, and there are problems of detail loss and visual hallucination. In addition, the object boundaries in the low-resolution image are blurred and the spatial relationship is weakened, making it difficult for the model to parse complex scenes and resulting in answer deviation. If the input resolution is increased to 448×448 or higher, although the details can be improved, the computational cost and memory occupancy increase, making it difficult to deploy to mobile devices or edge computing nodes; if more image patches are generated through dense cropping, although the local perception is enhanced, the computational complexity of self-attention surges and the training cost rises. That is to say, the current visual question answering scheme has a vicious cycle of detail loss and increasing hallucination under low-resolution input, and the improvement methods that rely on high-resolution input or brute-force expansion of visual tokens face rigid constraints on computational efficiency-resource cost. Therefore, in the traditional visual question answering process, for the low-resolution image input scenario, there is a risk of detail loss and increasing hallucination, resulting in the need to improve the accuracy of the question answering results. Summary of the Invention

[0004] One or more embodiments of this specification provide a visual question answering method, device, and medium based on a multimodal large model, which are used to solve the following technical problems: In the traditional visual question answering process, for the low-resolution image input scenario, there is a risk of detail loss and increasing hallucination, resulting in the need to improve the accuracy of the question answering results.

[0005] One or more embodiments of this specification adopt the following technical solutions:

[0006] One or more embodiments of this specification provide a visual question answering method based on a multimodal large model. The method includes: obtaining the original question answering image data and original question answering text data input by a user, converting the original question answering image data to determine corresponding high-resolution visual feature data; using the high-resolution visual feature data to enhance the features of the original visual features corresponding to the pre-obtained original question answering image data to determine enhanced visual token features; extracting the question answering text features of the original question answering text data, performing feature fusion based on the enhanced visual token features and the question answering text features to determine a comprehensive feature vector, and generating an answer through the multimodal large model and the comprehensive feature vector.

[0007] One or more embodiments of this specification provide a visual question answering device based on a multimodal large model, including:

[0008] At least one processor; and,

[0009] A memory communicatively connected to the at least one processor; wherein,

[0010] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above method.

[0011] A non-volatile computer storage medium provided by one or more embodiments of this specification stores computer-executable instructions, and the computer-executable instructions are configured to: execute the above method.

[0012] The above at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects: Through the above technical solution, the original question-and-answer image data is converted to obtain high-resolution visual feature data, effectively compensating for the high-frequency details lost due to downsampling compression in low-resolution images; Using the high-resolution visual feature data to perform feature enhancement on the original visual features, generating enhanced visual token features, more rich detail information can be extracted from the high-resolution features and incorporated into the original features, enabling the model to capture more subtle details when processing images; Since high-resolution visual feature data can be obtained and effective feature enhancement can be performed, the visual information on which the model is based is more accurate and complete, reducing the phenomenon of visual hallucinations caused by insufficient or inaccurate information; The enhanced visual token features and the question-and-answer text features are fused, and through text-image semantic alignment processing, the model can comprehensively consider the information of both the image and the text when generating answers. The mutual verification and supplementation of multi-modal information avoid the one-sided understanding caused by the model relying only on single-modal information, further reducing the risk of visual hallucinations; Different from the method of directly increasing the resolution of the input image, while enhancing the image details, the computational complexity and memory occupancy are not significantly increased. Through targeted processing and feature enhancement of the original image data, while ensuring the acquisition of key details, a relatively low computational complexity is maintained, enabling the model to be more easily deployed to mobile devices or edge computing nodes, meeting the resource limitation requirements in practical applications and broadening the application scenarios of the model; Compared with the method of generating more image patches through dense cropping, through reasonable feature processing and fusion methods, the computational cost is effectively controlled, reducing the time and resources required for training; From the acquisition of the original question-and-answer image data and the original question-and-answer text data, to the generation, feature enhancement, and feature fusion of high-resolution visual feature data, comprehensive, accurate, and interrelated information is provided for the multi-modal large model, enabling it to more accurately understand the question intention and image content, and thus generate more accurate answers; Optimized for low-resolution image input scenarios, while avoiding the problems brought by high-resolution input and brute-force expansion of visual tokens, it has strong robustness, enabling the model to stably perform in various practical application scenarios, improving the overall performance and reliability of the question-and-answer system. Brief Description of the Drawings

[0013] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:

[0014] Figure 1Schematic flowchart of a visual question answering method based on a multimodal large model provided by an embodiment of this specification;

[0015] Figure 2 Schematic flowchart of visual feature extraction for low-resolution images provided by an embodiment of this specification;

[0016] Figure 3 Schematic flowchart of feature enhancement for visual tokens of low-resolution images provided by an embodiment of this specification;

[0017] Figure 4 Schematic flowchart of an answer generation process based on feature splicing provided by an embodiment of this specification;

[0018] Figure 5 Schematic structural diagram of a visual question answering device based on a multimodal large model provided by an embodiment of this specification. Detailed implementation manners

[0019] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.

[0020] An embodiment of this specification provides a visual question answering method based on a multimodal large model. It should be noted that the execution subject in the embodiments of this specification can be a server or any device with data processing capabilities. Figure 1 Schematic flowchart of a visual question answering method based on a multimodal large model provided by an embodiment of this specification, as Figure 1 shown, mainly includes the following steps:

[0021] Step S101, obtain the original question-and-answer image data and original question-and-answer text data input by the user, and convert the original question-and-answer image data to determine the corresponding high-resolution visual feature data.

[0022] In an embodiment of this specification, the original question-and-answer image data and original question-and-answer text data input by the user are obtained through a multimodal large model. It should be noted that the original question-and-answer image data here is low-resolution data, such as 224×224 pixels, etc.

[0023] Convert the original Q&A image data to determine the corresponding high-resolution visual feature data, specifically including: performing grid processing on the original Q&A image data according to a preset grid division method to determine a plurality of pixel grids, where each pixel grid includes at least one original pixel point; performing an interpolation operation on the original pixel points in each pixel grid to determine the interpolation pixel color attributes; obtaining an interpolation image matrix through the combination of a plurality of the interpolation pixel color attributes, determining the corresponding high-resolution image data, and extracting the high-resolution visual feature data corresponding to the high-resolution image data.

[0024] In one embodiment of this specification, the original Q&A image data is divided into a plurality of pixel grids according to a preset grid division rule. The image is disassembled into small regional units, and each pixel grid contains at least one original pixel point, which is convenient for performing specific operations on different regions. It should be noted that the original low-resolution image can be divided into regular grids, and each grid contains a fixed number of original pixel points (such as a 2×2 grid). Setting regular grids maintains computational efficiency and avoids the overhead of complex dynamic division.

[0025] Perform an interpolation operation on the original pixel points in each pixel grid. During the interpolation process, the neighborhood range can be dynamically expanded according to the local content. Taking the original pixel point as the center, determine multiple neighboring pixel points around it according to its position in the image, and obtain the positions and color attributes of these neighboring pixel points. Then, according to the relationship between the neighboring pixel positions and the original pixel position, assign neighborhood assignment weights to each neighboring pixel point. Finally, combine the weights and color attributes of the neighboring pixel points to calculate the interpolation pixel color attribute corresponding to the original pixel point. The interpolation method based on neighboring pixel information and weight assignment can reasonably infer the color of the interpolation pixel according to the surrounding pixels, effectively supplementing the image details. Combine a plurality of interpolation pixel color attributes to form an interpolation image matrix, and thus obtain high-resolution image data. After obtaining the high-resolution image data, use relevant algorithms or models to extract the high-resolution visual feature data from it.

[0026] Through the above technical solution, through gridification and interpolation operations, the number of pixels in the image is effectively increased, and the image resolution is improved. The blurred parts or areas with missing details in the original low-resolution image can present more details after processing. During the interpolation process, by considering the position of neighboring pixels to assign weights, the local information of the image is fully utilized, making the high-resolution image closer to the real scene in terms of details and avoiding the distortion and blurring phenomena caused by simply magnifying the image. The high-resolution visual feature data corresponding to the high-resolution image data contains richer semantic and detailed information. In the visual question answering task of the multimodal large model, when the high-quality visual feature data is fused with the text feature data, it can provide a more accurate basis for the model to understand the image, improve the model's ability to understand and answer questions, and thus improve the accuracy and reliability of visual question answering.

[0027] For each original pixel point in the pixel grid, an interpolation operation is performed to determine the color attribute of the interpolated pixel, specifically including: taking the original pixel point as the center, and according to the original pixel position of the original pixel point, determining a plurality of neighboring pixel points in the original question-and-answer image data to obtain the neighboring pixel position and neighboring pixel color attribute of each neighboring pixel point; through the neighboring pixel position and the original pixel position, assigning a neighboring assignment weight to each neighboring pixel point; and determining the color attribute of the interpolated pixel corresponding to the original pixel point according to the neighboring assignment weight corresponding to each neighboring pixel point and the neighboring pixel color attribute.

[0028] In an embodiment of this specification, when processing the original question-and-answer image data, for a specific original pixel point, based on its position, a plurality of neighboring pixel points adjacent to it are determined within the range of the original image data, and at the same time, the position information and color attribute of each neighboring pixel point are obtained. Here, the color attribute can be the values corresponding to the R channel, G channel, and B channel. Based on the position relationship between the neighboring pixel points and the original pixel point, weights are assigned to each neighboring pixel point. Generally speaking, neighboring pixel points closer to the original pixel point will be assigned higher weights, while those farther away will have lower weights, which reflects the difference in the influence degree of neighboring pixel points on the interpolation of the original pixel point. Using the assigned neighboring assignment weights and the corresponding neighboring pixel color attributes, the color attribute of the interpolated pixel corresponding to the original pixel point is determined by means of weighted calculation. Multiply the color attribute of each neighboring pixel point by its corresponding weight, and then accumulate these products. The final result obtained is the color attribute of the interpolated pixel of the original pixel point, thereby realizing the supplement and optimization of the color information of the original pixel point.

[0029] It should be noted that based on the positional relationship between the neighborhood pixel points and the original pixel points, the weight assignment for each neighborhood pixel point can be achieved through the following method: For each original pixel point, find its four nearest pixels (upper left, upper right, lower left, lower right) in the original image. According to the difference in the positions of the old and new pixels, calculate the contribution weights of these four pixels to the pixel value of the original pixel point. The weights are determined based on the distance, and the closer the pixel, the greater the weight. For each color channel (red, green, blue), apply the bilinear interpolation formula, combine the color values of the four neighboring pixels and their corresponding weights, and calculate the color value of the new grid point, taking into account the interpolation in both the horizontal and vertical directions. Repeat the above steps for each grid point in the target image until all new pixels are assigned the color values obtained by bilinear interpolation. Combine all the calculated new pixel values into a complete image matrix to generate the final high-resolution image.

[0030] In addition to the above method, through the neighborhood pixel position and the original pixel position, assign neighborhood assignment weights to each neighborhood pixel point, specifically including: determining the distance weight parameter corresponding to the neighborhood pixel point according to the neighborhood pixel position and the original pixel position, where the distance weight parameter has a negative correlation with the distance; pre-identifying local features of the original Q&A image data to determine the regional local feature parameters; dynamically correcting the distance weight parameter through the regional local feature parameters to determine the neighborhood assignment weight corresponding to the neighborhood pixel point.

[0031] In one embodiment of this specification, based on the position information of neighboring pixel points and the original pixel point, the distance weight parameter corresponding to the neighboring pixel point is calculated. The distance weight parameter and the distance between the neighboring pixel point and the original pixel point show a negative correlation, that is, the closer the neighboring pixel point is to the original pixel point, the larger the corresponding distance weight parameter; conversely, the farther the distance, the smaller the distance weight parameter. This negative correlation setting is based on the principle that the closer the distance in image interpolation, the stronger the correlation. Neighboring pixel points closer to the original pixel point have a greater impact on its interpolation result. When processing the original Q&A image data, local feature recognition operations are performed on it in advance. Through image analysis algorithms, local features of different regions in the image are extracted, and these features cover edge intensity and texture complexity, and finally form regional local feature parameters. The edge intensity parameter can reflect the obviousness and direction of the edge in a certain region of the image, and the texture complexity parameter describes the complex situation of the texture in this region. Using the identified regional local feature parameters, the previously calculated distance weight parameter is dynamically adjusted. For example, if the edge intensity is high in the region where the original pixel point is located, the weight of the neighboring pixel points along the edge direction will be appropriately increased to better maintain the coherence and accuracy of the edge; if the texture complexity is high, it means that the region is rich in details, and more refined adjustments need to be made to the weights of the surrounding neighboring pixel points to ensure that the interpolated image can retain these details. Through this dynamic correction method, the neighborhood allocation weight corresponding to each neighboring pixel point is finally determined.

[0032] Through the regional local feature parameters, the distance weight parameter is dynamically corrected to determine the neighborhood allocation weight corresponding to the neighboring pixel point, which specifically includes: determining the edge intensity parameter and texture complexity parameter corresponding to each original pixel point in the regional local feature parameters; obtaining the interpolation path between the neighboring pixel point and the original pixel point, so as to determine the edge correction factor corresponding to the neighboring pixel point through the interpolation path and the edge intensity parameter, where the edge intensity parameter is the gradient direction parameter corresponding to each original pixel point; centering on the neighboring pixel position, determining the local variance parameter within the preset neighborhood, so as to determine the texture correction factor corresponding to the neighboring pixel point according to the local variance parameter; dynamically correcting the distance weight parameter through the edge correction factor and the texture correction factor to determine the neighborhood allocation weight corresponding to the neighboring pixel point.

[0033] In one embodiment of this specification, when determining the edge strength parameter, the gradient magnitude and the pixel gradient direction are calculated through the Sobel operator. The gradient magnitude is used to determine the edge region, and through the pixel gradient direction and the interpolation direction, it is possible to effectively judge whether the interpolation path crosses the edge. Assume that bilinear interpolation is to be performed on the interpolation point (x, y) in the low-resolution image, and its 4 surrounding neighboring pixels are (i, j), with corresponding coordinates: upper left (x1, y1), upper right (x2, y1), lower left (x1, y2), lower right (x2, y2). If the interpolation point is located in the edge region, the weight is enhanced along the edge direction to suppress the contribution of pixels that cross the edge. If the edge is vertical (such as a tree trunk), the weights of the left and right pixels should be higher than those of the upper and lower pixels. First, detect the edge direction and calculate the gradient direction θ of the interpolation point (x, y) through the Sobel operator, , where Gx and Gy are the horizontal and vertical gradients respectively, from the Sobel filter. Then, judge whether the neighboring pixels cross the edge. For each neighboring pixel (i, j), calculate the connection direction θ(i, j) between it and the interpolation point, . The direction difference . If Δθ > 45°, it is considered that the interpolation path crosses the edge. Apply a penalty to the neighboring pixels that cross the edge: , where λ is the penalty coefficient (usually taken as 0.5 - 1.0), and the larger the value, the stronger the weight attenuation of the pixels that cross the edge.

[0034] In complex texture regions (such as grass, hair), reduce the interpolation neighborhood to retain details; in flat regions (such as the sky, wall), expand the neighborhood to suppress noise. First, calculate the local texture complexity. Taking the interpolation point (x, y) as the center, take a 3×3 neighborhood and calculate the local variance σ 2 . The higher the texture complexity (the larger σ 2 ), the stronger the weight attenuation of the neighboring pixels: w2(i, j) = 1 / (1 + γσ 2 ), where γ is the suppression coefficient (usually taken as 0.1 - 0.3), which controls the amplitude of the weight attenuation in the texture region. Fuse the basic weight (distance weight parameter), edge correction, and texture correction in a multiplicative manner and perform normalization processing to ensure that the sum of all weights is 1, obtaining the neighborhood allocation weights corresponding to the neighboring pixel points. Independently calculate the interpolation color for each color channel (R / G / B) to obtain the interpolation pixel color attribute. Through edge-sensitive correction and texture-adaptive correction, the dynamic weight mechanism realizes edge protection and detail retention, enhances the contribution of effective pixels along the edge direction to suppress blurring; focuses on core pixels in complex texture regions to avoid over-smoothing, and expands the effective neighborhood in flat regions to improve the anti-noise ability.

[0035] Through the above technical solution, the neighborhood allocation weight is determined by comprehensively considering the distance and local image features, making the interpolation result more accurately reflect the real situation of the image; compared with the method of determining the weight only based on the distance, it can better handle the complex structures and detailed parts in the image, reducing the information loss or distortion problems caused by simple distance weighting, and making the interpolated image closer to the original high-resolution image in terms of visual effect; the dynamic correction of the distance weight parameter based on local features can focus on protecting the edges and details of the image during the interpolation process; in the edge area, reasonably adjusting the weights of neighboring pixel points can avoid the appearance of edge blurring or jagged phenomena, making the edges smoother and more natural; for areas with rich textures, it can more carefully retain the texture features, improving the clarity and realism of the image. In visual question answering tasks, it helps the model to more accurately understand the image content and improve the accuracy of answers; in addition, it has the ability to automatically adjust the weight according to the characteristics of different regions of the image, and has good adaptability to various different types of images. Whether it is a natural scene image containing a large amount of details or an engineering drawing with obvious edges and regular textures, a better interpolation effect can be obtained through dynamic weight correction, expanding its application scope in different fields.

[0036] Step S102: Feature enhancement is performed on the original visual features corresponding to the pre-acquired original question-and-answer image data through high-resolution visual feature data to determine enhanced visual token features.

[0037] In an embodiment of this specification, first, the original visual features corresponding to the original question-and-answer image data are extracted, that is, multi-grid visual embeddings are extracted from the low-resolution image, and the image is divided into blocks according to the grid. Each image patch is subjected to feature extraction through the Visual Transformer in the pre-trained CLIP (Contrastive Language-Image Pretraining) model, and each block corresponds to a feature vector. Figure 2 It is a schematic flowchart of the extraction of visual features of a low-resolution image provided by an embodiment of this specification, as Figure 2 shown. The specific steps are as follows: First, the low-resolution image is represented as , and the low-resolution image is segmented into multiple image patches of uniform size, ready to be input into the Visual Transformer. It should be noted that if the patch size is , then image patches can be created, denoted as , where , It is the length of the sequence, similar to the number of words in a sentence. The patch size is set to 16, and the shape of a patch is (3, 16, 16), with a dimension of 3. After flattening, the size is 3 x 16 x 16 = 768. That is, a patch becomes a vector of length 768. The vector of 768 is not consistent with the set visual feature dimension. Therefore, a linear transformation is applied to each image patch to convert it into a vector of a fixed dimension, which serves as the input sequence for the Visual Transformer. Since the model does not know the position information of each patch in the sequence data, the patches must first append a position information. Different position encoding embeddings have little impact on the final result. In this embodiment, the fixed position embedding vector used by ViT is added to the corresponding output patch embeddings. The commonly used position embedding formula in ViT can be expressed as:

[0038]

[0039]

[0040] Here, represents the position embedding matrix, is the position index (starting from 0), is the dimension index, and is the total dimension of the position embedding. The above formula actually generates position encoding based on sine and cosine functions, ensuring the difference between different positions and encoding position information at different scales (through the exponential term in the denominator), enabling the model to learn local and global position relationships. Concatenate the vector of each block and the position information encoding, and input it into the Transformer encoder to obtain the visual feature of each block. The visual features of all blocks are combined into the visual feature of the low-resolution image .

[0041] Through this high-resolution visual feature data, the original visual feature corresponding to the pre-obtained original Q&A image data is feature-enhanced to determine the enhanced visual token feature, specifically including: using the low-resolution feature token corresponding to the original visual feature as the query, and using this high-resolution visual feature as the key and value. Through the cross-resolution attention mechanism, the original visual feature is feature-enhanced to determine the enhanced visual token feature.

[0042] In an embodiment of this specification, in order to align with the low-resolution image visual feature, this step uses a ConvNeXt model pre-trained on the LAION dataset as the visual encoder for the high-resolution image. By upsampling and concatenating the features of different convolutional stages of ConvNeXt at a 1 / 4 input ratio, the feature map of the high-resolution image is obtained , where represents the number of high - resolution image features, where the high - resolution image reflects the pixel - level feature count in each high - resolution image segmentation block. The visual encoding of each stage in the ConvNeXt model mainly includes downsampling layers, ConvNeXt blocks, and key steps of normalization. At the beginning of some stages, depth convolution with stride - 2 is used for downsampling to reduce the spatial dimension and increase the number of channels; depth convolution with a larger kernel size (such as 7x7) is used to capture a larger range of context information, which can cover more spatial regions. Layer Normalization is used, and Pointwise Convolution (1x1 Conv) is used to adjust the number of channels, helping to reduce or increase the deep activation function of the feature map. ReLU is used as the activation function. The result of the above operations is added to the input to maintain information flow and promote deeper learning. The formula expression of the ConvNeXt block is as follows:

[0043]

[0044] where is the input feature map, and represent depth convolution and point convolution operations respectively, is layer normalization, is the activation function.

[0045] In addition, the visual token features of the low - resolution image are enhanced based on the attention mechanism, Figure 3 is a schematic flow diagram of feature enhancement for visual tokens of low - resolution images provided by an embodiment of this specification. As Figure 3 shown, the above - generated low - resolution image features are and the high - resolution image features are , using to enhance the visual token features in to improve the visual question - answering ability of the multi - modal large model. To maintain the number of final visual tokens to improve the efficiency of the language large model for visual and text semantic understanding and reasoning, the low - resolution visual feature is used as the query , aiming to match relevant visual cues from the feature map. At the same time, the high - resolution image feature map is used as the key and the value . The low - resolution blocks in and contain Associate with the corresponding high-resolution sub-regions of the pixel point features. The patch information mining process can be described as follows:

[0046] ;

[0047] Among them, and respectively represent the projection layer and the multi-layer perceptron, generating enhanced visual tokens for subsequent processing by the language large model, ensuring that for each query only mine the corresponding sub-regions with features, thus maintaining the overall computational efficiency. The method for enhancing visual token features of low-resolution images based on the attention mechanism extracts details in high-resolution images without expanding the number of visual tokens, thus maintaining a balance between rich details and saving computational resources.

[0048] Step S103: Extract the question-and-answer text features of the original question-and-answer text data, perform feature fusion based on the enhanced visual token features and the question-and-answer text features, and determine the comprehensive feature vector to generate an answer through the multi-modal large model and the comprehensive feature vector.

[0049] In an embodiment of this specification, extract the question-and-answer text features of the original question-and-answer text data. When the multi-modal large model processes visual questions, it usually fuses data of two modalities, text and image, to obtain a more comprehensive and in-depth understanding. When processing the text part, the question text undergoes a series of preprocessing and transformation steps to be input into the Transformer structure of the model to obtain text features. First, tokenization splits the text into words or sub-words. For Chinese, it involves splitting the sentence into individual characters or words; for English, it is usually split into words. This process does not directly correspond to a formula and depends on specific tokenization algorithms or toolkits (such as NLTK, jieba, etc.). Then, tokenization maps the tokenized words to digital IDs that the model can understand, and this process is called tokenization. Each word corresponds to a unique ID in the model's vocabulary.

[0050] ,

[0051] Among them, is a word, is its corresponding ID. Further, Positional Encoding enables the model to understand the relative order of Tokens at different positions in the sequence. Common positional encoding methods assign a fixed-length vector to each position, and these vectors are added to the embeddings of the Tokens. The Embedding Layer converts the ID of a Token into a high-dimensional dense vector, which can capture the semantic relationships between words. If the vocabulary size is , and the embedding dimension is , then the embedding matrix is . The embedding representation of a Token is:

[0052]

[0053] Finally, the Transformer encoder deeply extracts and encodes the text features through the Self-Attention mechanism and multiple layers of fully connected feed-forward networks. The self-attention calculation process can be summarized as:

[0054] ;

[0055] ,

[0056] where, is the input embedding matrix, is the weight matrix for queries, keys, and values, is the dimension of the key vector, The output is the text feature representation .

[0057] Based on the enhanced visual token features and the Q&A text features, feature fusion is performed to determine a comprehensive feature vector, specifically including: pre-acquiring a set of matching text-image pair data, and optimizing the projection parameters of the projection layer through the matching text-image pairs in the set of matching text-image pair data; through the projection layer, embedding the enhanced visual token features into the text feature space corresponding to the Q&A text features, performing text-image semantic alignment processing for feature splicing, and determining the comprehensive feature vector.

[0058] By matching the matching image-text pairs in the data set, the projection parameters of the projection layer are optimized, specifically including: obtaining the high-resolution image features, low-resolution image features, and matching text features in the matching image-text pairs; determining the high-resolution matching similarity based on the high-resolution image features and the matching text features, and determining the low-resolution matching similarity based on the low-resolution image features and the matching text features; determining the similarity difference loss based on the difference between the high-resolution matching similarity and the low-resolution matching similarity, using the similarity difference loss as the loss function for contrastive learning, and using backpropagation to optimize the projection parameters of the projection layer.

[0059] The projection layer is a key module in the multimodal model. Its core goal is to map visual features and text features to the same semantic space to solve the problem of inconsistent cross-modal data distributions. In one embodiment of this specification, the projection layer can be implemented through a multi-layer perceptron structure. A data set of matching image-text pairs is collected in advance, and the matching image-text pairs contain image and corresponding text information. The projection parameters of the projection layer are optimized using the image-text pairs in the data set of matching image-text pairs. The optimized projection layer embeds enhanced visual token features into the text feature space corresponding to the question-and-answer text features. In this process, image-text semantic alignment is achieved, enabling the features of images and texts to match and understand each other at the semantic level. Finally, through feature concatenation, the aligned image and text features are combined together to form a comprehensive feature vector, providing richer and more relevant information for the multimodal large model to generate accurate answers.

[0060] Extract high-resolution image features, low-resolution image features, and matching text features from the matching image-text pairs. Calculate the high-resolution matching similarity between the high-resolution image features and the matching text features, and the low-resolution matching similarity between the low-resolution image features and the matching text features respectively. These similarity metrics reflect the semantic association degree between image features and text features at different resolutions. Use the difference between the high-resolution matching similarity and the low-resolution matching similarity as the similarity difference loss. Apply this loss function to contrastive learning. Contrastive learning maximizes the similarity between matching pairs and minimizes the similarity between non-matching pairs, enabling the model to learn more effective feature representations. Use the backpropagation algorithm to adjust and optimize the projection parameters of the projection layer according to the loss value, enabling the projection layer to better embed image features into the text feature space and achieve more accurate image-text semantic alignment.

[0061] By optimizing the parameters of the projection layer to achieve graphic and text semantic alignment and feature splicing, it is possible to better fuse image and text features in the same space, reduce information loss caused by modal differences, and the comprehensive feature vector contains richer and more matching graphic and text information. When the multi-modal large model understands and processes visual question answering tasks, it can more accurately grasp the relationship between the question and the image, thereby improving the accuracy of the answer; considering the matching similarity between high-resolution and low-resolution image features and text features, and optimizing the projection layer based on this, can fully explore the connection between the information of different resolutions of the image and the text; based on the matching graphic and text pairs for contrastive learning to optimize the projection layer parameters, the model can learn more general graphic and text matching patterns, enabling the model to more effectively perform feature fusion and semantic understanding when facing unseen combinations of images and texts; using the similarity difference loss as the optimization basis, compared with other complex training methods, it can more specifically adjust the projection layer parameters, accelerate the convergence speed of the model, and reduce unnecessary training time and computational resource consumption.

[0062] Figure 4 FIG. is a schematic flow chart of an answer generation process based on feature splicing provided by an embodiment of the present specification, as Figure 4 shown, the processed visual feature vector and text feature vector are directly spliced together by dimension to form a comprehensive feature vector. After feature splicing, it will first pass through a multi-modal fusion layer (a neural network layer with parameter sharing or specific design) to further integrate the information of the two modalities and enhance the interaction between them. The comprehensive feature vector is input into the language large model, which usually contains multiple layers of Transformer structures, can handle long-distance dependencies, and understand the fused features in a rich context. Inside the language large model, the attention mechanism is used to dynamically focus on the parts of the visual and text features that are most relevant to answering the question, which helps the model focus on key information and improve the accuracy of reasoning. After multiple layers of transformation and calculation, the model finally generates an answer. The embodiment of the present specification generates the answer in an autoregressive manner, which is implemented through a decoder network, predicting the next most likely word based on the previous feature vector until the end marker is reached or the maximum length limit is reached.

[0063] Through the above technical solutions, the original question-and-answer image data is transformed to obtain high-resolution visual feature data, effectively compensating for the high-frequency details lost due to downsampling compression in low-resolution images; the original visual features are enhanced using the high-resolution visual feature data to generate enhanced visual token features, which can extract richer detailed information from the high-resolution features and integrate it into the original features, enabling the model to capture more nuances when processing images; due to the ability to obtain high-resolution visual feature data and perform effective feature enhancement, the visual information on which the model is based is more accurate and complete, reducing the visual hallucination phenomenon caused by insufficient or inaccurate information; the enhanced visual token features and the question-and-answer text features are fused, and through text-image semantic alignment processing, the model can comprehensively consider the information from both the image and the text when generating answers. The mutual verification and complementation of multi-modal information avoid the one-sided understanding caused by the model relying only on single-modal information, further reducing the risk of visual hallucinations; different from the method of directly increasing the resolution of the input image, while enhancing the image details, the computational complexity and memory occupancy are not significantly increased. Through targeted processing and feature enhancement of the original image data, while ensuring the acquisition of key details, a relatively low computational complexity is maintained, enabling the model to be more easily deployed to mobile devices or edge computing nodes, meeting the resource limitation requirements in practical applications and broadening the application scenarios of the model; compared with the method of generating more image patches through dense cropping, through reasonable feature processing and fusion methods, the computational cost is effectively controlled, reducing the time and resources required for training; from the acquisition of the original question-and-answer image data and the original question-and-answer text data to the generation, feature enhancement, and feature fusion of the high-resolution visual feature data, comprehensive, accurate, and interrelated information is provided for the multi-modal large model, enabling it to more accurately understand the question intention and image content, and thus generate more accurate answers; it is optimized for low-resolution image input scenarios, while avoiding the problems brought by high-resolution input and brute-force expansion of visual tokens, and has strong robustness, enabling the model to stably perform in various practical application scenarios and improving the overall performance and reliability of the question-and-answer system.

[0064] This embodiment of the specification also provides a visual question-and-answer device based on a multi-modal large model, as Figure 5 shown. The device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.

[0065] This embodiment of the specification also provides a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set to: execute the above method.

[0066] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the embodiments of devices, equipment, and non-volatile computer storage media, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0067] The above description is of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0068] The devices and media provided in the embodiments of this specification correspond one-to-one with the methods. Therefore, the devices and media also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.

[0069] Those skilled in the art should understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0070] This specification is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0071] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction means that implements the function specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0072] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0073] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0074] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0075] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0076] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising said element.

[0077] The above description is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and variations can be made to one or more embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.

Claims

1. A visual question answering method based on a multimodal large model, characterized in that, The method includes: Obtain the original Q&A image data and the original Q&A text data input by the user, and convert the original Q&A image data to determine the corresponding high-resolution visual feature data; Through the high-resolution visual feature data, perform feature enhancement on the original visual features corresponding to the pre-obtained original Q&A image data to determine enhanced visual token features; Extract the Q&A text features of the original Q&A text data, perform feature fusion based on the enhanced visual token features and the Q&A text features, determine a comprehensive feature vector, and generate an answer through a multimodal large model and the comprehensive feature vector; Performing feature fusion based on the enhanced visual token features and the Q&A text features to determine a comprehensive feature vector specifically includes: Pre-obtain a set of matching text-image pair data, and optimize the projection parameters of the projection layer through the matching text-image pairs in the set of matching text-image pair data; Through the projection layer, embed the enhanced visual token features into the text feature space corresponding to the Q&A text features, perform text-image semantic alignment processing, and perform feature splicing to determine the comprehensive feature vector; Optimizing the projection parameters of the projection layer through the matching text-image pairs in the set of matching text-image pair data specifically includes: Obtain the high-resolution image features, low-resolution image features, and matching text features in the matching text-image pairs; Determine the high-resolution matching similarity based on the high-resolution image features and the matching text features, and determine the low-resolution matching similarity based on the low-resolution image features and the matching text features; Use the difference between the high-resolution matching similarity and the low-resolution matching similarity to determine the similarity difference loss, use the similarity difference loss as a loss function to perform contrastive learning, and use backpropagation to optimize the projection parameters of the projection layer.

2. The visual question answering method based on a multimodal large model according to claim 1, wherein Converting the original Q&A image data to determine the corresponding high-resolution visual feature data specifically includes: Perform grid processing on the original Q&A image data according to a preset grid division method to determine a plurality of pixel grids, where each pixel grid includes at least one original pixel point; Perform an interpolation operation on the original pixel points in each pixel grid to determine the interpolation pixel color attributes; Obtain an interpolation image matrix through the combination of a plurality of the interpolation pixel color attributes, determine the corresponding high-resolution image data, and extract the high-resolution visual feature data corresponding to the high-resolution image data.

3. A visual question answering method based on a multimodal large model according to claim 2, characterized in that, Performing an interpolation operation on the original pixel points in each pixel grid to determine the interpolation pixel color attributes specifically includes: Taking the original pixel point as the center, determine a plurality of neighboring pixel points in the original Q&A image data according to the original pixel position of the original pixel point, so as to obtain the neighboring pixel positions and neighboring pixel color attributes of each neighboring pixel point; Assign neighboring assignment weights to each neighboring pixel point through the neighboring pixel positions and the original pixel position; Determine the interpolated pixel color attribute corresponding to the original pixel point according to the neighborhood allocation weight corresponding to each neighborhood pixel point and the neighborhood pixel color attribute.

4. A visual question answering method based on a multimodal large model according to claim 1, characterized in that Perform feature enhancement on the original visual features corresponding to the pre-acquired original Q&A image data through the high-resolution visual feature data to determine enhanced visual token features, specifically including: Use the low-resolution feature token corresponding to the original visual feature as a query, use the high-resolution visual feature as the key and value, and perform feature enhancement on the original visual feature through a cross-resolution attention mechanism to determine the enhanced visual token feature.

5. A visual question answering method based on a multimodal large model according to claim 3, characterized in that, Allocate neighborhood allocation weights to each neighborhood pixel point according to the neighborhood pixel position and the original pixel position, specifically including: Determine the distance weight parameter corresponding to the neighborhood pixel point according to the neighborhood pixel position and the original pixel position, where the distance weight parameter has a negative correlation with the distance; Perform local feature recognition on the original Q&A image data in advance to determine the regional local feature parameters; Dynamically correct the distance weight parameter through the regional local feature parameters to determine the neighborhood allocation weight corresponding to the neighborhood pixel point.

6. The visual question answering method based on a multimodal large model according to claim 5, characterized in that Dynamically correct the distance weight parameter through the regional local feature parameters to determine the neighborhood allocation weight corresponding to the neighborhood pixel point, specifically including: Determine the edge intensity parameter and texture complexity parameter corresponding to each original pixel point in the regional local feature parameters; Obtain the interpolation path between the neighborhood pixel point and the original pixel point, and determine the edge correction factor corresponding to the neighborhood pixel point through the interpolation path and the edge intensity parameter, where the edge intensity parameter is the gradient direction parameter corresponding to each original pixel point; Determine the local variance parameter within a preset neighborhood centered on the neighborhood pixel position, and determine the texture correction factor corresponding to the neighborhood pixel point according to the local variance parameter; Dynamically correct the distance weight parameter through the edge correction factor and the texture correction factor to determine the neighborhood allocation weight corresponding to the neighborhood pixel point.

7. A visual question answering device based on a multimodal large model, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-6.

8. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are set to: execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Multi-modal language generation method and system guided by multi-granularity visual information

    CN118708071A

  • High-efficiency high-resolution image visual mark generation method for multi-modal large model

    CN118735932A

  • Response method and device based on multi-modal tooth problem consultation

    CN118762821A