Visual question and answer method and device based on multi-modal large model and medium

By converting and feature enhancement of low-resolution images and combining feature fusion with Q&A text features, the problems of missing details and rising hallucinations in low-resolution image input scenes are solved, the accuracy and reliability of visual Q&A are improved, and the application scenarios are adapted to resource-limited application scenarios.

CN119938872AActive Publication Date: 2025-05-06INSPUR GENERSOFT CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510429240.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

In the traditional visual question-and-answer process, for low-resolution image input scenes, there is a risk of missing details and rising hallucinations, which leads to the accuracy of question-and-answer results.

Method used

By obtaining the original Q&A image data and text data input by the user, the image data is converted to determine high-resolution visual feature data, and feature enhancement of the original visual feature to generate enhanced visual token features. Then, the enhanced visual token features are featured and the Q&A text features are featured to generate answers through the multimodal big model.

Benefits of technology

It effectively compensates for the high-frequency details lost by low-resolution images due to downsampling compression, reduces visual hallucinations, improves the accuracy and reliability of Q&A results, and avoids a significant increase in computing volume and memory usage. It is suitable for mobile devices or edge computing nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938872A_ABST
    Figure CN119938872A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a visual question and answer method and device based on a multi-modal large model and a medium, and relates to the technical field of data processing, the method comprises the following steps: obtaining original question and answer image data and original question and answer text data input by a user, and converting the original question and answer image data to determine corresponding high-resolution visual feature data; performing feature enhancement on original visual features corresponding to pre-acquired original question and answer image data through the high-resolution visual feature data to determine enhanced visual token features; and extracting question and answer text features of the original question and answer text data, performing feature fusion based on the enhanced visual token features and the question and answer text features to determine a comprehensive feature vector, and generating answers through the multi-modal large model and the comprehensive feature vector. Through targeted processing and feature enhancement of original image data, relatively low calculation complexity is maintained on the premise of ensuring acquisition of key details, resource limitation in practical application is met, and a model application scene is widened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present specification relates to the field of data processing technology, and in particular to a visual question answering method, device and medium based on a multimodal large model. Background Art

[0002] Visual Question Answering (VQA) technology is an intersection of computer vision and natural language processing. The VQA task can be specifically described as: given a picture and a text question, the model extracts image features from the given image, extracts text features from the text question, fuses the two types of modal features, and finally generates the best answer based on the fused features.

[0003] In the task of visual question answering (VQA), the input resolution of the visual modality directly affects the model's ability to capture image details and the reliability of contextual reasoning. When the existing technology uses low-resolution images (such as 224×224 pixels) as input, the low-resolution images lose high-frequency details (such as text, textures, and small objects) due to downsampling compression, resulting in the model's inability to accurately identify key visual clues, and there are problems of detail loss and visual hallucinations. In addition, the boundaries of objects in low-resolution images are blurred and the spatial relationships are weakened, making it difficult for the model to parse complex scenes, resulting in answer deviations. If the input resolution is increased to 448×448 or higher, although the details can be improved, the amount of calculation and memory usage increase, making it difficult to deploy to mobile devices or edge computing nodes; if more image blocks are generated by dense cropping, although local perception is enhanced, the complexity of self-attention calculation increases sharply, and the training cost increases. In other words, the current visual question answering scheme has a vicious cycle of missing details and increasing hallucinations under low-resolution input, and the improvement methods that rely on high-resolution input or violent expansion of visual tokens face the rigid constraints of computational efficiency-resource cost. Therefore, in the traditional visual question answering process, there is a risk of missing details and increased hallucinations for low-resolution image input scenarios, resulting in the accuracy of the question answering results needing to be improved. Summary of the invention

[0004] One or more embodiments of the present specification provide a visual question answering method, device and medium based on a multimodal large model, which are used to solve the following technical problems: In the traditional visual question answering process, there is a risk of missing details and increased hallucinations for low-resolution image input scenarios, resulting in the accuracy of the question answering results needing to be improved.

[0005] One or more embodiments of this specification adopt the following technical solutions: One or more embodiments of the present specification provide a visual question-answering method based on a multimodal large model, the method comprising: obtaining original question-answering image data and original question-answering text data input by a user, converting the original question-answering image data to determine corresponding high-resolution visual feature data; performing feature enhancement on the original visual features corresponding to the original question-answering image data acquired in advance through the high-resolution visual feature data to determine enhanced visual token features; extracting question-answering text features of the original question-answering text data, performing feature fusion based on the enhanced visual token features and the question-answering text features, and determining a comprehensive feature vector to generate an answer through the multimodal large model and the comprehensive feature vector.

[0006] One or more embodiments of this specification provide a visual question answering device based on a multimodal large model, including: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above method.

[0007] One or more embodiments of the present specification provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the above method.

[0008] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: through the above technical solution, the original question and answer image data is converted to obtain high-resolution visual feature data, which effectively compensates for the high-frequency details lost in the low-resolution image due to downsampling compression; the original visual features are enhanced by using the high-resolution visual feature data to generate enhanced visual token features, which can extract richer detail information from the high-resolution features and integrate them into the original features, so that the model can capture more subtleties when processing images; since high-resolution visual feature data can be obtained and effective feature enhancement can be performed, the visual information based on the model is more accurate and complete, reducing the visual illusion phenomenon caused by insufficient or inaccurate information; the enhanced visual token features and question and answer text features are integrated, and through the semantic alignment of images and texts, the model can comprehensively consider the information of both images and texts when generating answers, and the mutual confirmation and supplementation of multimodal information avoids the one-sided understanding of the model based on only a single modal information, further reducing the risk of visual illusions; different from the method of directly improving the resolution of the input image, While improving image details, the amount of calculation and memory usage have not increased significantly. Through targeted processing and feature enhancement of raw image data, while ensuring the acquisition of key details, the computational complexity is maintained at a relatively low level, making it easier for the model to be deployed on mobile devices or edge computing nodes, meeting the resource constraints in practical applications and broadening the application scenarios of the model. Compared with the method of generating more image blocks through dense cropping, reasonable feature processing and fusion methods are used to effectively control the computational cost and reduce the time and resources required for training. From the acquisition of raw question-and-answer image data and raw question-and-answer text data to the generation, feature enhancement and feature fusion of high-resolution visual feature data, comprehensive, accurate and interrelated information is provided for the multimodal large model, which can more accurately understand the question intent and image content, thereby generating more accurate answers. It is optimized for low-resolution image input scenarios, while avoiding the problems caused by high-resolution input and violent expansion of visual tokens, and has strong robustness, so that the model can perform stably in various practical application scenarios, improving the overall performance and reliability of the question-and-answer system. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art description. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. In the drawings: Figure 1 A flowchart of a visual question answering method based on a multimodal large model provided in an embodiment of this specification; Figure 2 A schematic diagram of a process for extracting visual features from a low-resolution image provided in an embodiment of this specification; Figure 3 A schematic diagram of a process for enhancing features of visual tokens of low-resolution images provided in an embodiment of this specification; Figure 4 A flowchart of a feature-based answer generation process provided in an embodiment of this specification; Figure 5 A schematic diagram of the structure of a visual question answering device based on a multimodal large model provided in an embodiment of this specification. DETAILED DESCRIPTION

[0010] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0011] The embodiments of this specification provide a visual question answering method based on a multimodal large model. It should be noted that the execution subject in the embodiments of this specification can be a server or any device with data processing capabilities. Figure 1 A flowchart of a visual question answering method based on a multimodal large model provided in an embodiment of this specification is shown in FIG. Figure 1 As shown, it mainly includes the following steps: Step S101, obtaining original question-answer image data and original question-answer text data input by the user, and converting the original question-answer image data to determine corresponding high-resolution visual feature data.

[0012] In one embodiment of the present specification, the original question and answer image data and the original question and answer text data input by the user are obtained through a multimodal large model. It should be noted that the original question and answer image data here is low-resolution data, such as 224×224 pixels.

[0013] The original question-and-answer image data is converted to determine the corresponding high-resolution visual feature data, specifically comprising: gridding the original question-and-answer image data according to a preset grid division method to determine a plurality of pixel grids, wherein each of the pixel grids includes at least one original pixel point; interpolating the original pixel points in each of the pixel grids to determine an interpolated pixel color attribute; obtaining an interpolated image matrix through a combination of a plurality of the interpolated pixel color attributes to determine the corresponding high-resolution image data, so as to extract the high-resolution visual feature data corresponding to the high-resolution image data.

[0014] In one embodiment of the present specification, the original question-and-answer image data is divided into multiple pixel grids according to a pre-set grid division rule. The image is broken down into small area units, each pixel grid contains at least one original pixel point, which is convenient for performing specific operations on different areas. It should be noted that the original low-resolution image can be divided into regular grids, each grid containing a fixed number of original pixels (such as a 2×2 grid). Setting a regular grid maintains computational efficiency and avoids the overhead of complex dynamic division.

[0015] Interpolation operations are performed on the original pixels in each pixel grid. During the interpolation process, the neighborhood range can be dynamically expanded according to the local content. With the original pixel as the center, multiple neighboring pixels around it are determined according to its position in the image, and the positions and color attributes of these neighboring pixels are obtained. Then, according to the relationship between the neighborhood pixel position and the original pixel position, a neighborhood allocation weight is assigned to each neighborhood pixel. Finally, the interpolation pixel color attribute corresponding to the original pixel is calculated by combining the weight and color attribute of the neighborhood pixel. The interpolation method based on neighborhood pixel information and weight allocation can reasonably infer the color of the interpolated pixel according to the situation of the surrounding pixels, and effectively supplement the image details. Multiple interpolation pixel color attributes are combined to form an interpolation image matrix, thereby obtaining high-resolution image data. After obtaining the high-resolution image data, high-resolution visual feature data is extracted from it using relevant algorithms or models.

[0016] Through the above technical scheme, the number of pixels of the image is effectively increased through gridding and interpolation operations, and the image resolution is improved. The blurred parts or detail-missing areas in the original low-resolution image can present more details after processing; in the interpolation process, by considering the position of the neighborhood pixels to allocate weights, the local information of the image is fully utilized, so that the high-resolution image is closer to the real scene in details, avoiding the distortion and blurring caused by simply enlarging the image; the high-resolution visual feature data corresponding to the high-resolution image data contains richer semantic and detail information. In the visual question-answering task of the multimodal large model, when the high-quality visual feature data is fused with the text feature data, it can provide the model with a more accurate basis for image understanding, enhance the model's understanding and answering capabilities of questions, and thus improve the accuracy and reliability of visual question-answering.

[0017] An interpolation operation is performed on each original pixel point in the pixel grid to determine the interpolated pixel color attribute, specifically comprising: taking the original pixel point as the center, and according to the original pixel position of the original pixel point, determining a plurality of neighborhood pixel points in the original question-and-answer image data to obtain the neighborhood pixel position and the neighborhood pixel color attribute of each of the neighborhood pixel points; assigning a neighborhood allocation weight to each of the neighborhood pixel points through the neighborhood pixel position and the original pixel position; and determining the interpolated pixel color attribute corresponding to the original pixel point according to the neighborhood allocation weight corresponding to each of the neighborhood pixel points and the neighborhood pixel color attribute.

[0018] In one embodiment of the present specification, when processing the original question-answer image data, for a specific original pixel, based on its position, multiple neighboring pixel points adjacent to it are determined within the original image data range, and the position information and color attributes of each neighboring pixel point are obtained at the same time. The color attributes here can be the values ​​corresponding to the R channel, the G channel and the B channel. Based on the positional relationship between the neighborhood pixel point and the original pixel point, a weight is assigned to each neighborhood pixel point. Generally speaking, the neighborhood pixel point closer to the original pixel point will be assigned a higher weight, and the neighborhood pixel point farther away will be assigned a lower weight, which reflects the difference in the degree of influence of the neighborhood pixel point on the interpolation of the original pixel point. Using the assigned neighborhood allocation weight and the corresponding neighborhood pixel color attribute, the interpolation pixel color attribute corresponding to the original pixel point is determined by weighted calculation. The color attribute of each neighborhood pixel point is multiplied by its corresponding weight, and then these products are accumulated. The final result is the interpolation pixel color attribute of the original pixel point, thereby realizing the supplementation and optimization of the color information of the original pixel point.

[0019] It should be noted that, based on the positional relationship between the neighborhood pixels and the original pixels, the weights assigned to each neighborhood pixel can be achieved in the following way: For each original pixel, find the four nearest pixels in the original image (upper left, upper right, lower left, lower right). According to the position difference between the new and old pixels, calculate the contribution weights of these four pixels to the pixel value of the original pixel. The weight is determined by the distance, and the closer the pixel, the greater the weight. For each color channel (red, green, and blue), apply the bilinear interpolation formula, combine the color values ​​of the four neighboring pixels and their corresponding weights, calculate the color value of the new grid point, and consider the interpolation in the horizontal and vertical directions. Repeat the above steps for each grid point in the target image until all new pixels are assigned the color values ​​obtained according to the bilinear interpolation. All the calculated new pixel values ​​are combined into a complete image matrix to generate the final high-resolution image.

[0020] In addition to the above method, a neighborhood allocation weight is assigned to each neighborhood pixel point through the neighborhood pixel position and the original pixel position, specifically including: determining a distance weight parameter corresponding to the neighborhood pixel point according to the neighborhood pixel position and the original pixel position, wherein the distance weight parameter is negatively correlated with the distance; performing local feature recognition on the original question and answer image data in advance to determine the regional local feature parameters; and dynamically correcting the distance weight parameter through the regional local feature parameters to determine the neighborhood allocation weight corresponding to the neighborhood pixel point.

[0021] In one embodiment of the present specification, the distance weight parameter corresponding to the neighborhood pixel point is calculated based on the position information of the neighborhood pixel point and the original pixel point. The distance weight parameter and the distance between the neighborhood pixel point and the original pixel point show a negative correlation, that is, the closer the neighborhood pixel point is to the original pixel point, the larger the corresponding distance weight parameter; conversely, the farther the distance is, the smaller the distance weight parameter is. This negative correlation is set based on the principle that the closer the distance is, the stronger the correlation is in image interpolation, and the neighborhood pixel point close to the original pixel point has a greater impact on its interpolation result. When processing the original question-and-answer image data, a local feature recognition operation is performed on it in advance. Through the image analysis algorithm, the local features of different regions in the image are extracted. These features cover edge strength and texture complexity, and finally form regional local feature parameters. The edge strength parameter can reflect the obvious degree and direction of the edge of a certain area in the image, and the texture complexity parameter describes the complexity of the texture of the area. Using the identified regional local feature parameters, the distance weight parameter calculated previously is dynamically adjusted. For example, if the edge strength is high in the area where the original pixel is located, the weight of the neighboring pixels along the edge direction will be appropriately increased to better maintain the continuity and accuracy of the edge; if the texture complexity is high, it means that the area is rich in details, and the weight of the surrounding neighboring pixels needs to be adjusted more finely to ensure that the interpolated image can retain these details. Through this dynamic correction method, the neighborhood allocation weight corresponding to each neighboring pixel is finally determined.

[0022] The distance weight parameter is dynamically corrected through the local feature parameters of the region to determine the neighborhood allocation weight corresponding to the neighborhood pixel point, specifically including: determining the edge strength parameter and texture complexity parameter corresponding to each original pixel point in the local feature parameters of the region; obtaining the interpolation path between the neighborhood pixel point and the original pixel point, so as to determine the edge correction factor corresponding to the neighborhood pixel point through the interpolation path and the edge strength parameter, wherein the edge strength parameter is a gradient direction parameter corresponding to each original pixel point; taking the neighborhood pixel position as the center, determining the local variance parameter within a preset neighborhood, so as to determine the texture correction factor corresponding to the neighborhood pixel point according to the local variance parameter; and dynamically correcting the distance weight parameter through the edge correction factor and the texture correction factor to determine the neighborhood allocation weight corresponding to the neighborhood pixel point.

[0023] In one embodiment of the present specification, when determining the edge strength parameters, the gradient amplitude and pixel gradient direction are calculated by the Sobel operator. The gradient amplitude is used to determine the edge area. The pixel gradient direction and interpolation direction can be used to effectively determine whether the interpolation path crosses the edge. Assume that bilinear interpolation is required for the interpolation point (x, y) in the low-resolution image, and the four neighboring pixels around it are (i, j), with corresponding coordinates: upper left (x1, y1), upper right (x2, y1), lower left (x1, y2), and lower right (x2, y2). If the interpolation point is located in the edge area, the weight is enhanced along the edge direction to suppress the contribution of pixels crossing the edge. If the edge is in the vertical direction (such as a tree trunk), the weights of the left and right pixels should be higher than those of the upper and lower pixels. First, the edge direction is detected by calculating the gradient direction θ of the interpolation point (x, y) through the Sobel operator. , Gx, Gy are horizontal and vertical gradients respectively, which come from Sobel filtering. After that, determine whether the neighboring pixels cross the edge. For each neighboring pixel (i, j), calculate the direction of the line connecting it and the interpolation point θ(i, j). Directional Differences , if Δθ>45°, the interpolation path is considered to cross the edge. Penalize the neighboring pixels that cross the edge: , λ is the penalty coefficient (usually 0.5~1.0). The larger the value, the stronger the weight attenuation across the edge.

[0024] In complex texture areas (such as grass and hair), the interpolation neighborhood is reduced to preserve details; in flat areas (such as the sky and walls), the neighborhood is expanded to suppress noise. First, the local texture complexity is calculated, and the local variance σ is calculated by taking a 3×3 neighborhood with the interpolation point (x, y) as the center. 2 , the higher the texture complexity (σ 2 The larger the value, the stronger the weight decay of the neighboring pixels: w2 (i,j)=1 / (1+γσ 2 ), where γ is the suppression coefficient (usually 0.1~0.3), which controls the weight attenuation of the texture area. The basic weight (distance weight parameter), edge correction, and texture correction are fused by product, and normalized to ensure that the sum of all weights is 1, and the neighborhood distribution weight corresponding to the neighborhood pixel point is obtained. The interpolated color is calculated independently for each color channel (R / G / B) , and get the color attribute of the interpolated pixel. Through edge-sensitive correction and texture adaptive correction, the dynamic weight mechanism realizes edge protection and detail retention, enhances the contribution of effective pixels along the edge direction, and suppresses blur; focuses on core pixels in complex texture areas to avoid over-smoothing, expands the effective neighborhood in flat areas, and improves noise resistance.

[0025] Through the above technical scheme, the neighborhood allocation weights are determined by comprehensively considering the distance and local features of the image, so that the interpolation result can more accurately reflect the real situation of the image; compared with the method of determining the weight only based on the distance, it can better handle the complex structures and details in the image, reduce the information loss or distortion caused by simple distance weighting, and make the interpolated image closer to the original high-resolution image in visual effect; the distance weight parameter is dynamically corrected based on local features, which can focus on protecting the edges and details of the image during the interpolation process; in the edge area, the weights of the neighborhood pixels are reasonably adjusted to avoid the appearance of edge blur or jagged phenomena, making the edge smoother and more natural; for areas with rich textures, the texture features can be more carefully preserved, the clarity and realism of the image can be improved, and in the visual question-answering task, it helps the model to understand the image content more accurately and improve the accuracy of the answer; in addition, it has the ability to automatically adjust the weights according to the characteristics of different areas of the image, and has good adaptability to various types of images. Whether it is a natural scene image with a lot of details or an engineering drawing with obvious edges and regular textures, it can obtain a better interpolation effect by dynamically correcting the weights, expanding its application range in different fields.

[0026] Step S102, using high-resolution visual feature data, feature enhancement is performed on the original visual features corresponding to the pre-acquired original question-and-answer image data to determine enhanced visual token features.

[0027] In one embodiment of the present specification, the original visual features corresponding to the original question-answering image data are first extracted, that is, multi-grid visual embeddings are extracted from the low-resolution image, and the image is divided into blocks according to the grid. Each image block (patch) is subjected to feature extraction by the Visual Transformer in the pre-trained CLIP (Contrastive Language -Image Pretraining) model, and each block corresponds to a feature vector. Figure 2 A schematic diagram of a process for extracting visual features from a low-resolution image provided in an embodiment of this specification is shown in FIG. Figure 2 The specific steps are as follows: First, the low-resolution image is represented as , split the low-resolution image into multiple uniformly sized image patches (patches) and prepare them for input to the Visual Transformer. It should be noted that the patch size is , then you can create image patches, represented as ,in , It is the length of the sequence, similar to the number of words in a sentence. The patches size is set to 16, the shape of a patch is (3,16,16), the dimension is 3, and the size after flattening is 3x16x16=768. That is, a patch becomes a vector of length 768. The vector of 768 is inconsistent with the set visual feature dimension. Therefore, a linear transformation is applied to each image block to convert it into a vector of fixed dimension as the input sequence of the Visual Transformer. Since the model does not know the position information of each patch in the sequence data, the patches must first be appended with a position information. Different position encoding embeddings have little effect on the final result. In this embodiment, the fixed position embedding vector used by ViT is added to the corresponding output patchembeddings. The commonly used position embedding formula in ViT can be expressed as:

[0028]

[0029] here, represents the position embedding matrix, is the position index (starting from 0), is the dimension index, and is the total dimension of position embedding. The above formula actually generates position encoding based on sine and cosine functions, ensuring the difference between different positions, and encoding position information at different scales (through the exponential term in the denominator), so that the model can learn local and global position relationships. The vector and position information encoding of each block are concatenated and input into the Transformer encoder to obtain the visual features of each block. The visual features of all blocks are combined into low-resolution image visual features. .

[0030] The high-resolution visual feature data is used to perform feature enhancement on the original visual features corresponding to the original question-and-answer image data acquired in advance to determine the enhanced visual token features, specifically including: taking the low-resolution feature token corresponding to the original visual feature as a query, and the high-resolution visual feature as a key and value, and performing feature enhancement on the original visual feature through a cross-resolution attention mechanism to determine the enhanced visual token features.

[0031] In one embodiment of the present specification, in order to align with the visual features of the low-resolution image, this step uses the ConvNeXt model pre-trained with the LAION dataset as the visual encoder of the high-resolution image. The feature map of the high-resolution image is obtained by upsampling and concatenating the features of different convolution stages of ConvNeXt at a ratio of 1 / 4 of the input. ,in Represents the number of high-resolution image features, where the high-resolution image reflects the number of pixel-level features in each high-resolution image segmentation block. The visual encoding of each stage of the ConvNeXt model mainly includes the downsampling layer ConvNeXt block, and the key steps of normalization. At the beginning of some stages, a deep convolution with stride-2 is used for downsampling to reduce the spatial dimension and increase the number of channels; a larger kernelsize (such as 7x7) is used for deep convolution to capture a wider range of contextual information and cover more spatial areas. Using Layer Normalization, Pointwise Convolution (1x1 Conv) is used to adjust the number of channels to help reduce or increase the deep activation function of the feature map using ReLU. The results of the above operations are added to the input to keep the information flowing and promote deeper learning. The formula of the ConvNeXt block is expressed as follows:

[0032] in is the input feature map, and Represent the depth convolution and point convolution operations respectively, is the layer normalization, is the activation function.

[0033] In addition, the visual token features of low-resolution images are enhanced based on the attention mechanism. Figure 3 The present invention provides a flow chart of a method for enhancing the features of low-resolution image visual tokens. Figure 3 As shown, the low-resolution image features generated above are and high-resolution image features are ,use Enhancement In order to maintain the number of final visual tokens and improve the efficiency of the language large model for visual and textual semantic understanding and reasoning, the low-resolution visual features are used to improve the ability of multimodal large models to answer visual questions. As a query , the purpose is to Match relevant visual cues in the feature map. At the same time, the high-resolution image feature map As a key Sum . The low-resolution blocks in and Included The corresponding high-resolution sub-regions of the pixel features are associated, and the patch information mining process can be expressed as: ; in, and Represent the projection layer and multi-layer perceptron, respectively, to generate enhanced visual tokens Used for subsequent language model processing to ensure that each query Mining only In The corresponding sub-region of the feature is selected, thereby maintaining the overall computational efficiency. The low-resolution image visual token feature enhancement method based on the attention mechanism does not enlarge The proposed method extracts details from high-resolution images without reducing the number of visual tokens, thus maintaining a balance between enriching details and saving computing resources.

[0034] Step S103, extracting question and answer text features of the original question and answer text data, performing feature fusion based on enhanced visual token features and question and answer text features, and determining a comprehensive feature vector to generate an answer through a multimodal large model and the comprehensive feature vector.

[0035] In one embodiment of the present specification, question-answering text features of the original question-answering text data are extracted. When processing visual questions, multimodal large models usually fuse data from both text and image modes to obtain a more comprehensive and in-depth understanding. When processing the text portion, the question text undergoes a series of preprocessing and conversion steps to be input into the Transformer structure of the model to obtain text features. First, tokenization divides the text into words or subwords. For Chinese, it involves dividing sentences into individual characters or words; for English, it is usually divided into words. This process does not directly correspond to a formula, but depends on a specific word segmentation algorithm or tool library (such as NLTK, jieba, etc.). Then, Tokenization maps the segmented words to a digital ID that the model can understand. This process is called Tokenization. Each word corresponds to a unique ID in the model vocabulary.

[0036] , in, It's vocabulary. is its corresponding ID. Furthermore, positional encoding is used to allow the model to understand the relative order of tokens at different positions in the sequence. Common positional encoding methods assign a fixed-length vector to each position, which is added to the embedding of the token. The embedding layer converts the token ID into a high-dimensional dense vector that can capture the semantic relationship between words. If the vocabulary size is , the embedding dimension is , then the embedding matrix The embedding representation of Token is:

[0037] Finally, the Transformer encoder uses the self-attention mechanism and a multi-layer fully connected feedforward network to deeply extract and encode text features. The self-attention calculation process can be summarized as: ; , in, is the embedding matrix of the input, is the weight matrix of query, key and value, is the dimension of the key vector, Output is text feature representation .

[0038] Based on the enhanced visual token feature and the question and answer text feature, feature fusion is performed to determine a comprehensive feature vector, which specifically includes: pre-acquiring a matching image-text pair data set, and optimizing the projection parameters of the projection layer through the matching image-text pairs in the matching image-text pair data set; through the projection layer, embedding the enhanced visual token feature into the text feature space corresponding to the question and answer text feature, performing image-text semantic alignment processing to perform feature splicing, and determining the comprehensive feature vector.

[0039] The projection parameters of the projection layer are optimized through the matching image-text pairs in the matching image-text pair data set, specifically including: obtaining high-resolution image features, low-resolution image features and matching text features in the matching image-text pairs; determining high-resolution matching similarity based on the high-resolution image features and the matching text features, and determining low-resolution matching similarity based on the low-resolution image features and the matching text features; determining the similarity difference loss based on the difference between the high-resolution matching similarity and the low-resolution matching similarity, performing comparative learning based on the similarity difference loss as a loss function, and optimizing the projection parameters of the projection layer by using back propagation.

[0040] The projection layer is a key module in the multimodal model, and its core goal is to map visual features and text features to the same semantic space to solve the problem of inconsistent cross-modal data distribution. In one embodiment of the present specification, the projection layer can be implemented by a multi-layer perceptron structure. A matching image-text pair data set is collected in advance, and the matching image-text pair contains the image and the corresponding text information. The projection parameters of the projection layer are optimized using the image-text pairs in the matching image-text pair data set. The optimized projection layer embeds the enhanced visual token features into the text feature space corresponding to the question and answer text features. In this process, image-text semantic alignment is achieved, so that the features of the image and text can match and understand each other at the semantic level. Finally, the aligned image and text features are combined together through feature splicing to form a comprehensive feature vector, which provides richer and more relevant information for the multimodal large model to generate accurate answers.

[0041] Extract high-resolution image features, low-resolution image features, and matching text features from matching image-text pairs. Calculate the high-resolution matching similarity between high-resolution image features and matching text features, and the low-resolution matching similarity between low-resolution image features and matching text features. These similarity metrics reflect the degree of semantic association between image features and text features at different resolutions. The difference between high-resolution matching similarity and low-resolution matching similarity is used as the similarity difference loss. This loss function is applied to contrastive learning, which maximizes the similarity between matching pairs and minimizes the similarity between mismatching pairs, allowing the model to learn more effective feature representations. Using the back-propagation algorithm, the projection parameters of the projection layer are adjusted and optimized according to the loss value, so that the projection layer can better embed image features into the text feature space and achieve more accurate semantic alignment of images and texts.

[0042] By optimizing the projection layer parameters to achieve semantic alignment and feature splicing of images and texts, the image and text features can be better integrated in the same space, reducing the information loss caused by modal differences. The comprehensive feature vector contains richer and more matching image and text information, so that the multimodal large model can more accurately grasp the relationship between questions and images when understanding and processing visual question-answering tasks, thereby improving the accuracy of answers; considering the matching similarity of high-resolution and low-resolution image features and text features, and optimizing the projection layer based on this, the connection between information and text at different image resolutions can be fully explored; based on the comparative learning and optimization of the projection layer parameters for matching image and text pairs, the model can learn a more general image-text matching mode, so that the model can more effectively perform feature fusion and semantic understanding when facing an unprecedented combination of images and texts; using the similarity difference loss as the optimization basis, compared with other complex training methods, the projection layer parameters can be adjusted more specifically, which can speed up the convergence of the model and reduce unnecessary training time and computing resource consumption.

[0043] Figure 4 A flowchart of a feature-based answer generation process provided in an embodiment of this specification is shown in FIG. Figure 4 As shown, the processed visual feature vector and the text feature vector are directly spliced ​​together by dimension to form a comprehensive feature vector. After the features are spliced, they will first pass through a multimodal fusion layer (a neural network layer with parameter sharing or a specific design) to further integrate the information of the two modes and enhance the interactivity between them. The comprehensive feature vector is input into the language model, which usually contains a multi-layer Transformer structure that can handle long-distance dependencies and understand the fused features in a rich context. Inside the language model, an attention mechanism is used to dynamically focus on the parts of the visual and text features that are most relevant to answering questions, which helps the model focus on key information and improves the accuracy of reasoning. After multiple layers of transformation and calculation, the model finally generates an answer. The embodiment of this specification uses an autoregressive method to generate an answer, which is implemented through a decoder network to predict the next most likely word based on the previous feature vector until the end mark is reached or the maximum length limit is reached.

[0044] Through the above technical scheme, the original question and answer image data is converted to obtain high-resolution visual feature data, which effectively compensates for the high-frequency details lost in the low-resolution image due to downsampling compression; the original visual features are enhanced by using the high-resolution visual feature data to generate enhanced visual token features, which can extract richer detail information from the high-resolution features and integrate them into the original features, so that the model can capture more subtleties when processing images; since high-resolution visual feature data can be obtained and effective feature enhancement can be performed, the visual information based on the model is more accurate and complete, reducing the visual illusion phenomenon caused by insufficient or inaccurate information; the enhanced visual token features and question and answer text features are fused, and through the semantic alignment of pictures and texts, the model can comprehensively consider the information of both images and texts when generating answers. The mutual confirmation and supplementation of multimodal information avoids the one-sided understanding of the model based on only a single modal information, further reducing the risk of visual illusions; unlike the method of directly increasing the resolution of the input image, while improving the image details, there is no significant The amount of computation and memory usage are increased, but through targeted processing and feature enhancement of the original image data, while ensuring the acquisition of key details, the computational complexity is maintained at a relatively low level, making it easier for the model to be deployed on mobile devices or edge computing nodes, meeting the resource constraints in practical applications and broadening the application scenarios of the model. Compared with the method of generating more image blocks through dense cropping, reasonable feature processing and fusion methods are used to effectively control the computational cost and reduce the time and resources required for training. From the acquisition of original question and answer image data and original question and answer text data to the generation, feature enhancement and feature fusion of high-resolution visual feature data, comprehensive, accurate and interrelated information is provided for the multimodal large model, which can more accurately understand the question intent and image content, thereby generating more accurate answers. It is optimized for low-resolution image input scenarios, while avoiding the problems caused by high-resolution input and violent expansion of visual tokens, and has strong robustness, so that the model can perform stably in various practical application scenarios, improving the overall performance and reliability of the question and answer system.

[0045] The embodiment of this specification also provides a visual question answering device based on a multimodal large model, such as Figure 5 As shown, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.

[0046] The embodiments of the present specification also provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the above method.

[0047] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0048] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0049] The devices and media provided in the embodiments of this specification correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0050] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0051] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0052] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0053] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0054] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0055] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0056] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0057] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0058] The above description is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included in the scope of the claims of this specification.

Claims

1. A visual question answering method based on a multimodal large model, characterized in that: The method comprises: Acquire original question-answer image data and original question-answer text data input by a user, and convert the original question-answer image data to determine corresponding high-resolution visual feature data; By using the high-resolution visual feature data, feature enhancement is performed on the original visual features corresponding to the original question-answer image data acquired in advance to determine enhanced visual token features; Extract the question and answer text features of the original question and answer text data, perform feature fusion based on the enhanced visual token features and the question and answer text features, determine a comprehensive feature vector, and generate an answer through a multimodal large model and the comprehensive feature vector.

2. The visual question answering method based on a multimodal large model according to claim 1, characterized in that: The original question-answering image data is converted to determine corresponding high-resolution visual feature data, specifically including: According to a preset grid division method, the original question-answer image data is gridded to determine a plurality of pixel grids, wherein each of the pixel grids includes at least one original pixel point; Performing an interpolation operation on each original pixel point in the pixel grid to determine an interpolated pixel color attribute; An interpolated image matrix is ​​obtained by combining a plurality of the interpolated pixel color attributes, and the corresponding high-resolution image data is determined to extract high-resolution visual feature data corresponding to the high-resolution image data.

3. The visual question answering method based on a multimodal large model according to claim 2, characterized in that: Performing an interpolation operation on each original pixel point in the pixel grid to determine an interpolated pixel color attribute, specifically comprising: Taking the original pixel point as the center, and according to the original pixel position of the original pixel point, determining a plurality of neighborhood pixel points in the original question-and-answer image data, so as to obtain a neighborhood pixel position and a neighborhood pixel color attribute of each of the neighborhood pixel points; Allocating a neighborhood allocation weight to each of the neighborhood pixel points according to the neighborhood pixel position and the original pixel position; The interpolation pixel color attribute corresponding to the original pixel point is determined according to the neighborhood allocation weight corresponding to each of the neighborhood pixel points and the neighborhood pixel color attribute.

4. The visual question answering method based on a multimodal large model according to claim 1, characterized in that: By using the high-resolution visual feature data, feature enhancement is performed on the original visual features corresponding to the original question-answer image data acquired in advance to determine enhanced visual token features, specifically including: The low-resolution feature token corresponding to the original visual feature is used as a query, and the high-resolution visual feature is used as a key and a value. The original visual feature is enhanced through a cross-resolution attention mechanism to determine the enhanced visual token feature.

5. The visual question answering method based on a multimodal large model according to claim 1, characterized in that: Performing feature fusion based on the enhanced visual token feature and the question-answer text feature to determine a comprehensive feature vector specifically includes: Pre-acquire a matching image-text pair data set, and optimize the projection parameters of the projection layer through the matching image-text pairs in the matching image-text pair data set; Through the projection layer, the enhanced visual token feature is embedded into the text feature space corresponding to the question and answer text feature, and image-text semantic alignment processing is performed to perform feature splicing and determine the comprehensive feature vector.

6. The visual question answering method based on a multimodal large model according to claim 5, characterized in that: Optimizing the projection parameters of the projection layer by using the matching image-text pairs in the matching image-text pair data set, specifically comprising: Acquire high-resolution image features, low-resolution image features, and matching text features in the matching image-text pair; Determining high-resolution matching similarity based on the high-resolution image features and the matching text features, and determining low-resolution matching similarity based on the low-resolution image features and the matching text features; A similarity difference loss is determined based on the difference between the high-resolution matching similarity and the low-resolution matching similarity. Comparative learning is performed based on the similarity difference loss as a loss function, and back propagation is used to optimize the projection parameters of the projection layer.

7. The visual question answering method based on a multimodal large model according to claim 3, characterized in that: Allocating a neighborhood allocation weight to each of the neighborhood pixels according to the neighborhood pixel position and the original pixel position specifically includes: Determine a distance weight parameter corresponding to the neighborhood pixel point according to the neighborhood pixel position and the original pixel position, wherein the distance weight parameter is negatively correlated with the distance; Preliminarily performing local feature recognition on the original question-answering image data to determine regional local feature parameters; The distance weight parameter is dynamically modified by using the local feature parameter of the region to determine the neighborhood allocation weight corresponding to the neighborhood pixel point.

8. The visual question answering method based on a multimodal large model according to claim 7, characterized in that: The distance weight parameter is dynamically modified by the local feature parameter of the region to determine the neighborhood allocation weight corresponding to the neighborhood pixel point, specifically including: Determine the edge intensity parameter and texture complexity parameter corresponding to each original pixel point in the local feature parameters of the region; Obtaining an interpolation path between the neighborhood pixel point and the original pixel point, so as to determine an edge correction factor corresponding to the neighborhood pixel point through the interpolation path and the edge strength parameter, wherein the edge strength parameter is a gradient direction parameter corresponding to each original pixel point; Taking the neighborhood pixel position as the center, determining a local variance parameter in a preset neighborhood, and determining a texture correction factor corresponding to the neighborhood pixel point according to the local variance parameter; The distance weight parameter is dynamically corrected by the edge correction factor and the texture correction factor to determine the neighborhood allocation weight corresponding to the neighborhood pixel point.

9. A visual question answering device based on a multimodal large model, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-modal language generation method and system guided by multi-granularity visual information

    CN118708071A

  • High-efficiency high-resolution image visual mark generation method for multi-modal large model

    CN118735932A

  • Response method and device based on multi-modal tooth problem consultation

    CN118762821A

  • Multi-modal question and answer method and device based on multi-angle image, and electronic equipment

    CN119739814A

  • Image processing device, image processing method, and program

    US20130216130A1