Method for enhancing fine-grained sensing capability of multi-modal large model, and image processing method and device based on multi-modal large model
By integrating a target detection branch during training, the method enhances the visual perception of multiple-modal large models, addressing their fine-grained visual perception limitations and improving visual question answering accuracy.
Patent Information
- Application Number
- CN202510804247.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-16
AI Technical Summary
Current multiple-modal large models lack sufficient fine-grained visual perception capabilities, leading to suboptimal performance in tasks requiring detailed visual understanding.
Enhance the visual perception of multiple-modal large models by incorporating a target detection branch during training, utilizing a visual encoder and an enhancement coding module to refine visual features through target detection tasks, thereby improving the model's ability to accurately predict object classes and locations.
The proposed method significantly enhances the fine-grained visual perception of multiple-modal large models, enabling more accurate responses to visual question answering tasks by refining visual features through target detection tasks.
Smart Images

Figure CN120318606A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of artificial intelligence, and more specifically, to a method for enhancing the fine-grained perception ability of a multimodal large model, an image processing method and apparatus based on the multimodal large model. Background Art
[0002] The multimodal large model is an important technical foundation for realizing visual question answering and also an important step towards general artificial intelligence. The visual question answering task requires the multimodal large model to simultaneously process and understand data in visual (image, video) and language (question, text description) modalities and generate reasonable answers. Structurally, the multimodal large model usually consists of a visual encoder, a connection layer, and a large language model, and its capabilities are mainly reflected in the following aspects: 1) Modal fusion: By means of cross-modal attention mechanisms and multimodal alignment techniques, visual features and language features are effectively fused to understand the question and locate relevant regions in the image; 2) Language generation ability: The multimodal large model usually adopts an autoregressive generation method and utilizes the powerful generation ability of the pre-trained large language model to provide smooth and context-consistent answers when answering questions.
[0003] However, the currently used pre-trained visual encoder of the multimodal large model has weak fine-grained visual perception ability, resulting in poor performance of the multimodal large model in some visual question answering tasks that require fine-grained visual perception. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method for enhancing the fine-grained perception ability of a multimodal large model, an image processing method and apparatus based on the multimodal large model, which can effectively solve the problem of insufficient fine-grained perception ability of the multimodal large model in the prior art.
[0005] In one general aspect, a method for enhancing the fine-grained perception ability of a multimodal large model is provided. The multimodal large model includes a visual encoder, an enhancement encoding module, and a large language model. The method includes: obtaining a target detection dataset and a text-image pair dataset, where the target detection dataset includes a plurality of first images and the first ground truth labels of each first image, and the text-image pair dataset includes a plurality of text-image pairs and the second ground truth labels of each text-image pair. The first ground truth labels include the true categories and true location information of the objects included in the corresponding first images, and the second ground truth labels include the true answers to the text questions in the corresponding text-image pairs; for each image in the target detection dataset and the text-image pair dataset, perform the following processing: input the current image into the visual encoder to obtain a first visual feature; input the first visual feature into the enhancement encoding module to obtain a second visual feature; in response to the current image belonging to the target detection dataset, perform target detection processing on the current image based on the second visual feature to obtain a prediction result corresponding to the current image; adjust the parameters of the visual encoder and the enhancement encoding module based on the detection loss determined by the prediction result and the first ground truth label of the current image; in response to the current image belonging to the text-image pair dataset, input the second visual feature and the text question in the text-image pair where the current image is located into the large language model to obtain a predicted answer to the text question; adjust the parameters of the visual encoder, the enhancement encoding module, and the large language model based on the first generation loss determined by the predicted answer and the second ground truth label of the text-image pair where the current image is located.
[0006] Optionally, performing target detection processing on the current image based on the second visual feature to obtain a prediction result corresponding to the current image includes: performing cross-attention processing on the second visual feature and a preset query vector to obtain a cross-processed vector, where the query vector indicates querying the category and location information of the objects included in the current image from the second visual feature; inputting the cross-processed vector into a decoder to obtain a decoded feature; inputting the decoded feature into a feed-forward neural network to obtain a prediction result corresponding to the current image.
[0007] Optionally, when the target detection dataset further includes the text question corresponding to the current image and the first ground truth label further includes the true answer to the text question, before adjusting the parameters of the visual encoder and the enhancement encoding module based on the detection loss determined by the prediction result and the first ground truth label of the current image, it further includes: inputting the second visual feature and the text question corresponding to the current image into the large language model to obtain a predicted answer to the text question corresponding to the current image; determining a second generation loss based on the predicted answer and the true answer to the text question corresponding to the current image; where adjusting the parameters of the visual encoder and the enhancement encoding module based on the detection loss determined by the prediction result and the first ground truth label of the current image includes: adjusting the parameters of the visual encoder, the enhancement encoding module, and the large language model based on the detection loss and the second generation loss.
[0008] Optionally, before adjusting the parameters in the visual encoder and the enhancement encoding module based on the detection loss determined from the prediction result and the first ground truth label of the current image, it further includes: for each object included in the current image, determining the matching relationship between the prediction result of the current object and the ground truth category and ground truth location information of the current object in the first ground truth label of the current image; and determining the detection loss based on the matching relationship of each object and the first ground truth label.
[0009] Optionally, the multimodal large model further includes an image projection layer and a text projection layer. Among them, inputting the second visual feature and the text question in the text-image pair where the current image is located into the large language model to obtain the predicted answer to the text question includes: inputting the second visual feature into the image projection layer to obtain a third visual feature; inputting the text question in the text-image pair where the current image is located into the text projection layer to obtain a first text feature; concatenating the third visual feature and the first text feature to obtain a first concatenated result; and inputting the first concatenated result into the large language model to obtain the predicted answer to the text question in the text-image pair where the current image is located. In another general aspect, there is provided an image processing method based on a multimodal large model. The multimodal large model includes a visual encoder, an enhancement encoding module, and a large language model. The image processing method includes: inputting the image to be processed into the visual encoder to obtain a first visual feature; inputting the first visual feature into the enhancement encoding module to obtain a second visual feature; and inputting the second visual feature and the text question in the text-image pair where the image to be processed is located into the large language model to obtain the answer to the text question, where the parameters of the visual encoder, the enhancement encoding module, and the large language model in the multimodal large model are obtained by the method as above.
[0010] In another general aspect, an apparatus for enhancing the fine-grained perception ability of a multi-modal large model is provided. The multi-modal large model includes a visual encoder, an enhanced encoding module, and a large language model. The apparatus includes: an acquisition unit configured to acquire a target detection data set and a text-image pair data set. The target detection data set includes a plurality of first images and the first ground truth labels of each first image. The text-image pair data set includes a plurality of text-image pairs and the second ground truth labels of each text-image pair. The first ground truth label includes the ground truth category and the ground truth location information of the objects included in the corresponding first image. The second ground truth label includes the ground truth answer to the text question in the corresponding text-image pair. An execution unit configured to perform the following processing for each image in the target detection data set and the text-image pair data set: input the current image into the visual encoder to obtain a first visual feature; input the first visual feature into the enhanced encoding module to obtain a second visual feature; in response to the current image belonging to the target detection data set, perform target detection processing on the current image based on the second visual feature to obtain a prediction result corresponding to the current image; adjust the parameters of the visual encoder and the enhanced encoding module based on the detection loss determined based on the prediction result and the first ground truth label of the current image; in response to the current image belonging to the text-image pair data set, input the second visual feature and the text question in the text-image pair where the current image is located into the large language model to obtain a predicted answer to the text question; adjust the parameters of the visual encoder, the enhanced encoding module, and the large language model based on the first generation loss determined based on the predicted answer and the second ground truth label of the text-image pair where the current image is located.
[0011] Optionally, the execution unit is further configured to perform cross-attention processing on the second visual feature and a preset query vector to obtain a cross-processed vector, where the query vector is used to indicate querying the category and location information of the objects included in the current image from the second visual feature; input the cross-processed vector into a decoder to obtain a decoded feature; input the decoded feature into a feed-forward neural network to obtain a prediction result corresponding to the current image.
[0012] Optionally, when the target detection data set further includes the text question corresponding to the current image and the first ground truth label further includes the ground truth answer to the text question, before adjusting the parameters of the visual encoder and the enhanced encoding module based on the detection loss determined based on the prediction result and the first ground truth label of the current image, the execution unit is further configured to input the second visual feature and the text question corresponding to the current image into the large language model to obtain a predicted answer to the text question corresponding to the current image; determine a second generation loss based on the predicted answer and the ground truth answer to the text question corresponding to the current image; adjust the parameters of the visual encoder, the enhanced encoding module, and the large language model based on the detection loss and the second generation loss.
[0013] Optionally, the execution unit is further configured to, before adjusting the parameters of the visual encoder and the enhancement encoding module based on the detection loss determined by the prediction result and the first ground truth label of the current image, for each object included in the current image, determine the matching relationship between the prediction result of the current object and the ground truth category and ground truth location information of the current object in the first ground truth label of the current image; and determine the detection loss based on the matching relationship of each object and the first ground truth label.
[0014] Optionally, the multi-modal large model further includes an image projection layer and a text projection layer. Among them, the execution unit is further configured to input the second visual feature into the image projection layer to obtain a third visual feature; input the text question in the text-image pair where the current image is located into the text projection layer to obtain a first text feature; splice the third visual feature and the first text feature to obtain a first splicing result; and input the first splicing result into the large language model to obtain a predicted answer to the text question in the text-image pair where the current image is located. In another general aspect, there is provided an image processing apparatus based on a multi-modal large model. The multi-modal large model includes a visual encoder, an enhancement encoding module, and a large language model. The apparatus includes: a first acquisition unit configured to input an image to be processed into the visual encoder to obtain a first visual feature; a second acquisition unit configured to input the first visual feature into the enhancement encoding module to obtain a second visual feature; and a third acquisition unit configured to input the second visual feature and the text question in the text-image pair where the image to be processed is located into the large language model to obtain an answer to the text question, where the parameters of the visual encoder, the enhancement encoding module, and the large language model in the multi-modal large model are obtained by any of the above methods.
[0015] In another general aspect, there is provided a computer-readable storage medium storing instructions, where when the instructions are run by at least one computing device, at least one computing device is prompted to execute any of the above methods for enhancing the fine-grained perception ability of the multi-modal large model or the image processing method based on the multi-modal large model.
[0016] In another general aspect, there is provided a system including at least one computing device and at least one storage device storing instructions, where when the instructions are run by at least one computing device, at least one computing device is prompted to execute any of the above methods for enhancing the fine-grained perception ability of the multi-modal large model or the image processing method based on the multi-modal large model.
[0017] In another general aspect, there is provided a computer program product including computer instructions, where when the computer instructions are executed by a processor, the methods for enhancing the fine-grained perception ability of the multi-modal large model or the image processing method based on the multi-modal large model as described above are implemented.
[0018] A method for enhancing the fine-grained perception ability of an enhanced multi-modal large model according to an embodiment of the present disclosure, an image processing method and apparatus based on a multi-modal large model. On the basis of the original generation task training, a target detection data set is introduced, and the target detection process of an image is added, so that the category and position information of the objects included in the image can be predicted, and based on the detection loss determined by this part of the prediction results and their true categories and true position information, the parameters of the visual encoder and the enhanced encoding module are updated, enhancing the fine-grained visual perception ability of the visual features during the encoding process. That is, the present disclosure additionally optimizes the target detection task on the basis of the generation task, and enhances the fine-grained perception ability of the visual features encoded by the multi-modal large model through this visual downstream task of the target detection task, so as to obtain a multi-modal large model with better generation effect. Therefore, through the present disclosure, the problem of insufficient fine-grained perception ability of the multi-modal large model in the prior art can be effectively solved.
[0019] Some other aspects and / or advantages of the general concept of the present disclosure will be described in part in the following description, some of which will be clear from the description, or can be learned through the implementation of the general concept of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Through the following description with reference to the drawings showing embodiments, the above and other objects and features of the embodiments of the present disclosure will become clearer, where: Figure 1 is a flowchart showing a method for enhancing the fine-grained perception ability of an enhanced multi-modal large model according to an embodiment of the present disclosure; Figure 2 is an architecture diagram showing an enhanced multi-modal large model with enhanced fine-grained perception ability according to an embodiment of the present disclosure; Figure 3 is a general flowchart showing an enhanced multi-modal large model with enhanced fine-grained perception ability according to an embodiment of the present disclosure; Figure 4 is a flowchart showing an image processing method based on a multi-modal large model according to an embodiment of the present disclosure; Figure 5 is a block diagram showing an apparatus for enhancing the fine-grained perception ability of an enhanced multi-modal large model according to an embodiment of the present disclosure; Figure 6 is a block diagram showing an image processing apparatus based on a multi-modal large model according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The following specific embodiments are provided to assist the reader in obtaining a comprehensive understanding of the methods, apparatuses, and / or systems described herein. However, after understanding the disclosure of the present application, various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will be apparent. For example, the order of operations described herein is merely illustrative and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of the present application, except for operations that must occur in a specific order. Additionally, descriptions of features known in the art may be omitted for greater clarity and conciseness.
[0022] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. On the contrary, the examples described herein are provided only to illustrate some of the many possible ways of implementing the methods, apparatuses, and / or systems described herein, which will be apparent after understanding the disclosure of the present application.
[0023] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more thereof.
[0024] Although terms such as "first," "second," and "third" may be used herein to describe various components, elements, regions, layers, or parts, these components, elements, regions, layers, or parts should not be limited by these terms. On the contrary, these terms are only used to distinguish one component, element, region, layer, or part from another. Thus, a first component, first element, first region, first layer, or first part as referred to in the examples described herein may also be referred to as a second component, second element, second region, second layer, or second part without departing from the teachings of the examples.
[0025] In the specification, when an element (such as a layer, region, or substrate) is described as "on," "connected to," or "coupled to" another element, the element may be directly "on," "connected to," or "coupled to" the other element, or there may be one or more other elements therebetween. In contrast, when an element is described as "directly on," "directly connected to," or "directly coupled to" another element, there may be no other elements therebetween.
[0026] The terms used herein are for the purpose of describing various examples only and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. The terms "comprising," "including," and "having" specify the presence of the recited features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.
[0027] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains after understanding this disclosure. Unless explicitly defined herein, terms (such as those defined in a general dictionary) shall be construed to have a meaning consistent with their meaning in the context of the relevant art and this disclosure, and shall not be interpreted in an idealized or overly formal manner.
[0028] In addition, in the description of the examples, when it is considered that a detailed description of a relevant structure or function that is well-known will cause an ambiguous interpretation of this disclosure, such detailed description will be omitted.
[0029] In the visual instruction fine-tuning stage of a multimodal large model, the model is generally set to causal language modeling. The input sequence (visual features + question text) is concatenated into a continuous sequence. Through an autoregressive generation task, the next token is gradually predicted, and finally a complete answer is generated. This way, the fine-grained perception ability of visual features is limited by the visual encoder of the multimodal large model. Recent research shows that common visual encoders (such as clip) used in multimodal large models have inherent deficiencies in fine-grained perception ability due to the training paradigm, resulting in weak fine-grained visual perception ability of multimodal large models.
[0030] This disclosure provides a method for enhancing the fine-grained perception ability of visual features of a multimodal large model, which can improve the fine-grained perception ability of the visual encoder in the visual instruction fine-tuning stage, so as to answer more precisely questions that require fine-grained visual perception in visual question answering tasks. Specifically, an enhanced encoding module is added after the visual encoder and a target detection branch is added when training the multimodal large model. That is, in the visual instruction fine-tuning stage, on the basis of the autoregressive generation task, a target detection task is additionally introduced, and the fine-grained visual perception ability of the multimodal large model during encoding is strengthened through this visual downstream task of the target detection task.
[0031] The following will describe in detail the method for enhancing the fine-grained perception ability of a multimodal large model, the image processing method and device based on a multimodal large model with reference to the accompanying drawings.
[0032] This disclosure proposes a method for enhancing the fine-grained perception ability of a multimodal large model, Figure 1 is a flowchart showing the method for enhancing the fine-grained perception ability of an embodiment of this disclosure. Referring to Figure 1 , the multimodal large model includes a visual encoder, an enhanced encoding module, and a large language model. The method for enhancing the fine-grained perception ability of the multimodal large model includes the following steps: In step S101, a target detection dataset and an image-text pair dataset are obtained. The target detection dataset includes multiple first images and first ground-truth labels for each first image. The image-text pair dataset includes multiple image-text pairs and second ground-truth labels for each image-text pair. The first ground-truth labels include the ground-truth categories and ground-truth location information of the objects included in the corresponding first images. The second ground-truth labels include the ground-truth answers to the text questions in the corresponding image-text pairs.
[0033] As an example, the first images included in the above target detection dataset only need to contain objects and have the ground-truth categories and ground-truth location information of the corresponding objects, and it is not required that the first images must correspond to text questions. Suppose the image is an image containing two playing puppies. Then the ground-truth category of the objects included in the image is dog, and the ground-truth location is the actual locations of the two dogs in the image. For example, the two dogs can be boxed separately, and the location of each box is the ground-truth location of each dog. It should be noted that the image can contain both dogs and cows at the same time, indicating that the image includes two objects, and the ground-truth categories of each object are dog and cow respectively.
[0034] As an example, the image-text pairs included in the above image-text pair dataset include an image and the text question corresponding to the image for training in the generation task. It should be noted that the images in the image-text pairs included in the image-text pair dataset are generally images that cannot be used for target detection to avoid the appearance of samples for target detection in the image-text pair dataset, resulting in the same samples in the two datasets and redundant calculations.
[0035] In step S102, for each image in the target detection dataset and the image-text pair dataset, the following processing is performed: input the current image into a visual encoder to obtain a first visual feature; input the first visual feature into an enhancement encoding module to obtain a second visual feature; in response to the current image belonging to the target detection dataset, perform target detection processing on the current image based on the second visual feature to obtain a prediction result corresponding to the current image; adjust the parameters of the visual encoder and the enhancement encoding module based on the detection loss determined by the prediction result and the first ground-truth label of the current image; in response to the current image belonging to the image-text pair dataset, input the second visual feature and the text question in the image-text pair where the current image is located into a large language model to obtain a predicted answer to the text question; adjust the parameters of the visual encoder, the enhancement encoding module, and the large language model based on the first generation loss determined by the predicted answer and the second ground-truth label of the image-text pair where the current image is located.
[0036] As an example, the above processing can be first performed on the images in the object detection dataset. After the image processing in the object detection dataset is completed, the above processing can be then performed on the images in the image-text pair dataset; alternatively, the above processing can be first performed on the images in the image-text pair dataset. After the image processing in the image-text pair dataset is completed, the above processing can be then performed on the images in the object detection dataset; or, after the processing of some images in the object detection dataset is completed, such as after the processing of a batch of images in the object detection dataset, the above processing can be performed on a batch of images in the image-text pair dataset. Then, the processing of the next batch of images in the object detection dataset can be carried out, and the processing of the next batch of images in the image-text pair dataset can be performed until all the data in the object detection dataset and the image-text pair dataset are processed.
[0037] As an example, for each image in the object detection dataset and the image-text pair dataset, it can be first input into the vision encoder to obtain the first visual feature (i.e., visual token), and then the first visual feature can be input into the enhancement encoding module to re-encode the first visual feature, further enhancing the visual fine-grained perception to obtain the second visual feature. After obtaining the second visual feature, the subsequent processing needs to determine which dataset the current image belongs to. Specifically: If the current image belongs to the object detection dataset, object detection processing is performed on the current image to obtain the prediction result corresponding to the current image. The prediction result includes the predicted category and predicted location information of the object in the current image. Then, based on the predicted category, predicted location information and the true category and true location information in the first true label of the current image, the detection loss is determined, and through this detection loss, the parameters of the vision encoder and the enhancement encoding module are adjusted; If the current image belongs to the image-text pair dataset, the second visual feature and the text question in the image-text pair where the current image is located are input into the large language model to obtain the estimated answer to the text question. Then, based on the estimated answer and the true answer in the second true label of the image-text pair where the current image is located, the first generation loss is determined, and through the first generation loss, the parameters of the vision encoder, the enhancement encoding module and the large language model are adjusted; After the above processing is performed on all the images in the object detection dataset and the image-text pair dataset, a trained multi-modal large model is obtained.
[0038] It should be noted that in order to adapt to the output of the visual encoder of the multimodal large model, the object detection branch of the present disclosure may adopt an end-to-end object detection framework composed of pure transformers, but the present disclosure does not limit this. The enhanced encoding module may be the encoding layer that is more than the original visual encoding in the multimodal large model in the transformer encoder. Here, only the transformer encoder may be connected, but only the encoding layer that is more than the original visual encoding in the multimodal large model in the transformer encoder is used. It is also possible to replace the visual encoder and the enhanced encoding module with a transformer encoder, but the present disclosure does not limit this.
[0039] The object detection branch of the present disclosure is to use the visual encoder of the multimodal large model as the backbone network, and the connected transformer encoder as the parameter to be additionally trained, and re-encode the visual tokens transmitted by the visual encoder to inject fine-grained perception information. The following takes the end-to-end object detection framework composed of pure transformers as an example for illustration.
[0040] According to an embodiment of the present disclosure, the object detection processing of the current image based on the second visual feature can be implemented in the following manner to obtain the prediction result corresponding to the current image: perform cross-attention processing on the second visual feature and a preset query vector to obtain a cross-processed vector, where the query vector indicates querying the category and location information of the objects included in the current image from the second visual feature; input the cross-processed vector into a decoder to obtain a decoded feature; and input the decoded feature into a feed-forward neural network to obtain the prediction result corresponding to the current image. Through this embodiment, object detection processing can be performed quickly and accurately.
[0041] As an example, the above query vector can be preset and initialized as needed. Through this query vector, the feature information of the object (i.e., the above cross-processed vector) can be obtained from the second visual feature. For example, the preset query vector can indicate querying the category and location information of the objects included in the current image from the second visual feature, but the present disclosure does not limit this.
[0042] As an example, object detection can be implemented through a decoder, a feed-forward neural network, and the previous visual encoder and enhanced encoding model.
[0043] Specifically, assuming that the current image belongs to the object detection data set, perform cross-attention processing on the second visual feature and a preset query vector to obtain a cross-processed vector. Then, input the cross-processed vector into a decoder to obtain a decoded feature, and then input the decoded feature into a feed-forward neural network, and the prediction result corresponding to the current image can be obtained, that is, the predicted category and predicted location information of the objects included in the current image (such as the bounding box used to frame the objects).
[0044] It should be noted that if the image does not contain the required object, the feedforward neural network can input information without an object, such as the object category being the "no object" class, and the present disclosure does not limit this.
[0045] According to an embodiment of the present disclosure, before adjusting the parameters of the visual encoder and the enhancement coding module based on the detection loss determined from the prediction result and the first ground truth label of the current image, the detection loss can be determined. Specifically, for each object included in the current image, determine the matching relationship between the prediction result of the current object and the true category and true location information of the current object in the first ground truth label of the current image; based on the matching relationship of each object and the first ground truth label, determine the detection loss.
[0046] Through this embodiment, by using the matching relationship between the prediction result and the true category and true location information in the ground truth label, a detection loss that is more consistent with the actual situation can be determined, thereby enabling better training of the multi-modal large model.
[0047] As an example, the Hungarian matching cost function can be used to calculate the matching cost between the prediction result and the ground truth label, and the matching cost for each object in the current image can be calculated by the following formula: ) where, ( ) represents the ground truth label of the i th object in the current image, including the category of the object and the bounding box (i.e., the location information); ( ) represents the prediction result of the i th object in the current image, including the predicted category and the predicted bounding box ; represents the classification loss, which is used to evaluate the difference between the predicted category and the true category; represents the L1 loss, which is used to evaluate the coordinate difference between the predicted bounding box and the true bounding box; represents the GIoU loss, which is used to evaluate the overlap quality between the predicted bounding box and the true bounding box; and represent the loss weights, which are used to balance the contributions of the L1 and GIoU losses.
[0048] For each object in the current image, by minimizing the corresponding Hungarian matching cost function, the best matching relationship between the prediction result and the ground truth label can be found. Then, based on the matching relationship of each object, the detection loss is determined, and this detection loss is obtained based on the Hungarian matching result and can be as follows:
[0049] Among them, represents the best matching relationship between the prediction result found using the Hungarian matching cost function and the true label, and N represents the number of object categories included in the current image.
[0050] According to an embodiment of the present disclosure, the multimodal large model further includes an image projection layer and a text projection layer. Among them, inputting the second visual feature and the text question in the text-image pair where the current image is located into the large language model to obtain an estimated answer for the text question includes: inputting the second visual feature into the image projection layer to obtain a third visual feature; inputting the text question in the text-image pair where the current image is located into the text projection layer to obtain a first text feature; splicing the third visual feature and the first text feature to obtain a first splicing result; inputting the first splicing result into the large language model to obtain an estimated answer for the text question in the text-image pair where the current image is located.
[0051] Through this embodiment, the original training path of the multimodal large model can be completed conveniently and quickly.
[0052] As an example, this embodiment is the original training path of the multimodal large model, that is, training on the text-image pair dataset, aiming to align visual features and text features and perform text generation. Specifically, after obtaining the second visual feature of the image in the text-image pair, input the text question in the text-image pair into a tokenizer to obtain a tokenization result. Then, map the second visual feature and the tokenization result to the same dimension through the image projection layer and the text projection layer respectively, that is, obtain the third visual feature and the first text feature of the same dimension. After splicing the third visual feature and the first text feature, input them into the large language model to obtain the corresponding estimated answer, and perform autoregressive training based on the estimated answer and the true answer.
[0053] As an example, the generation loss in autoregressive training is as follows:
[0054] Among them, T represents the total length of the generated sequence (i.e., the estimated answer); represents the true token at time step t; represents the generated token sequence before time step t; x represents the result after splicing the third visual feature and the first text feature; represents the multimodal large model under the condition x and the generated token sequence to predict probability.
[0055] After obtaining the first generation loss through the above-mentioned generation loss function, the parameters of the visual encoder, the enhancement encoding module, and the large language model can be adjusted based on the first generation loss to achieve autoregressive training.
[0056] According to an embodiment of the present disclosure, when the target detection dataset further includes a text question corresponding to the current image and the first ground truth label further includes the ground truth answer to the text question, before adjusting the parameters of the visual encoder and the enhancement encoding module based on the detection loss determined based on the prediction result and the first ground truth label of the current image, the second visual feature and the text question corresponding to the current image are input into the large language model to obtain a predicted answer to the text question corresponding to the current image; a second generation loss is determined based on the predicted answer and the ground truth answer to the text question corresponding to the current image; wherein, adjusting the parameters in the visual encoder and the enhancement encoding module based on the detection loss determined based on the prediction result and the first ground truth label of the current image includes: adjusting the parameters of the visual encoder, the enhancement encoding module, and the large language model based on the detection loss and the second generation loss.
[0057] Through this embodiment, when some of the first images in the target detection dataset correspond to text questions, these first images can also go through a generation task, thereby enhancing the training of the large language model.
[0058] As an example, when some of the first images in the target detection dataset correspond to text questions, the second visual feature and the text question corresponding to the first image can also be input into the large language model to obtain a predicted answer to the text question corresponding to the first image, and then, a second generation loss is determined based on the predicted answer and the ground truth answer corresponding to the first image, and then a total loss is obtained based on the second generation loss and the detection loss of the first image, and further, the parameters of the visual encoder, the enhancement encoding module, and the large language model are adjusted through the total loss.
[0059] It should be noted that the above total loss can be the sum of the second generation loss and the detection loss, and the present disclosure does not limit this.
[0060] For a better understanding of the present disclosure, the following is combined with Figure 2 and Figure 3 for a systematic description.
[0061] Figure 2 shows an architecture diagram for enhancing the fine-grained perception ability of the enhanced multi-modal large model, Figure 3 shows a flowchart for enhancing the fine-grained perception ability of the enhanced multi-modal large model, as shown in Figure 2 and Figure 3 shown, this process includes the following parts: First, obtain the object detection dataset and the image-text pair dataset, and obtain the current image and the corresponding ground truth label (the first ground truth label or the second ground truth label) from the object detection dataset or the image-text pair dataset; Secondly, input the current image ( Figure 2 a picture containing two puppies) into the visual encoder to obtain the first visual feature; Thirdly, input the first visual feature into the enhancement encoding module to obtain the second visual feature; Once again, if the current image is from the object detection dataset, perform cross-attention processing on the second visual feature and a preset query vector to obtain the cross-processing result (i.e., the above-mentioned cross-processing vector). This cross-processing result is the feature information of the puppies in the image obtained from the second visual feature based on the query vector. Then, input the cross-processing vector into the Transformer decoder to obtain the decoded feature; input the decoded feature into a feed-forward neural network (Feedforward Neural Network, abbreviated as FFN) to obtain the prediction result corresponding to the current image. This prediction result includes the class (CLASS) - dog and the bounding boxes (BOX) of the two dogs; match the prediction result with the ground truth label one by one, use the Hungarian matching cost function to calculate the matching cost between the prediction result and the ground truth label, and determine the detection loss of the current image based on the matching cost; it should be noted that considering that an image may contain multiple objects, generally multiple FFNs are set, one FFN corresponds to the detection of one object, and if there is no corresponding object, the corresponding FFN outputs no object (NO OBEJECT); Once again, if the current image is from the image-text pair dataset, input the second visual feature into the image projection layer to obtain the third visual feature, input the text question in the image-text pair where the current image is located into the tokenizer to obtain the tokenization result, input the tokenization result into the text projection layer to obtain the text feature. Then, concatenate the third visual feature and the text feature and input them into the large language model to obtain the estimated answer to the text question. The estimated answer is that there are two dogs in the picture, and their positions in the picture are <x1, y1, x2, y2>, <x3, y3, x4, y4>; determine the first generation loss based on the estimated answer and the ground truth answer; Finally, if the current image is an image in the object detection dataset and there is no corresponding text question for the current image, then adjust the parameters of the visual encoder and the enhancement encoding module based on the above detection loss; if the current image is an image in the object detection dataset and there is a corresponding text question for the current image, then input the second visual feature into the image projection layer to obtain a third visual feature, input the text question in the text-image pair where the current image is located into the tokenizer to obtain a tokenization result, input the tokenization result into the text projection layer to obtain a text feature, and then, concatenate the third visual feature and the text feature and input them into the large language model to obtain a predicted answer to the text question; determine the second generation loss based on the predicted answer and the ground truth answer, and adjust the parameters of the visual encoder, the enhancement encoding module, and the large language model through the detection loss and the second generation loss; if the current image is an image in the text-image pair dataset, then adjust the parameters of the visual encoder, the enhancement encoding module, and the large language model through the first generation loss; until all samples in the object detection dataset and the text-image pair dataset are processed.
[0062] The present disclosure also provides an image processing method based on a multimodal large model. Figure 4 FIG. is a flowchart showing an image processing method based on a multimodal large model according to an embodiment of the present disclosure. Referring to Figure 4 FIG., the multimodal large model includes a visual encoder, an enhancement encoding module, and a large language model. The image processing method based on the multimodal large model includes the following steps: In step S401, input the image to be processed into the visual encoder to obtain a first visual feature; In step S402, input the first visual feature into the enhancement encoding module to obtain a second visual feature; In step S403, input the second visual feature and the text question in the text-image pair where the image to be processed is located into the large language model to obtain an answer to the text question.
[0063] The parameters of the visual encoder, the enhancement encoding module, and the large language model in the multimodal large model as described above are obtained by the method as described above.
[0064] In summary, the present disclosure belongs to the field of multimodal artificial intelligence. Based on the visual question answering task, especially for the part that requires fine-grained perception in the visual question answering task, the specific purpose is to enhance the visual fine-grained perception ability of the multimodal large model to solve the problem that the multimodal large model performs poorly in the visual question answering task that requires visual fine-grained perception. Specifically, an enhanced encoding module is added after the visual encoder, and a target detection branch is added during the training of the multimodal large model. That is, during the visual instruction fine-tuning stage, an additional target detection task is introduced on the basis of the autoregressive generation task. Through this visual downstream task of the target detection task, the fine-grained perception ability during the visual encoding of the multimodal large model is enhanced. During the inference settlement after the training is completed, the target detection branch can be omitted, and only the enhanced encoding module for injecting visual fine-grained perception ability and the original multimodal large model structure are retained.
[0065] Figure 5 is a block diagram of a device for enhancing the fine-grained perception ability of the multimodal large model according to an embodiment of the present disclosure. As Figure 5 shown, the multimodal large model includes a visual encoder, an enhanced encoding module, and a large language model. The device includes: an acquisition unit 50 and an execution unit 52.
[0066] The acquisition unit 50 is configured to acquire a target detection data set and a text-image pair data set. Among them, the target detection data set includes a plurality of first images and the first ground truth labels of each first image. The text-image pair data set includes a plurality of text-image pairs and the second ground truth labels of each text-image pair. The first ground truth label includes the true category and true location information of the objects included in the corresponding first image. The second ground truth label includes the true answer to the text question in the corresponding text-image pair. The execution unit 52 is configured to perform the following processing for each image in the target detection data set and the text-image pair data set: input the current image into the visual encoder to obtain a first visual feature; input the first visual feature into the enhanced encoding module to obtain a second visual feature; in response to the current image belonging to the target detection data set, perform target detection processing on the current image based on the second visual feature to obtain a prediction result corresponding to the current image; adjust the parameters of the visual encoder and the enhanced encoding module based on the detection loss determined by the prediction result and the first ground truth label of the current image; in response to the current image belonging to the text-image pair data set, input the second visual feature and the text question in the text-image pair where the current image is located into the large language model to obtain a predicted answer to the text question; adjust the parameters of the visual encoder, the enhanced encoding module, and the large language model based on the first generation loss determined by the predicted answer and the second ground truth label of the text-image pair where the current image is located.
[0067] According to an embodiment of the present disclosure, the execution unit 52 is further configured to perform cross-attention processing on the second visual feature and a preset query vector to obtain a cross-processed vector, where the query vector indicates querying the category and location information of the objects included in the current image from the second visual feature; input the cross-processed vector into a decoder to obtain a decoded feature; and input the decoded feature into a feed-forward neural network to obtain a prediction result corresponding to the current image.
[0068] According to an embodiment of the present disclosure, when the target detection dataset further includes a text question corresponding to the current image and the first ground truth label further includes the ground truth answer to the text question, before adjusting the parameters of the visual encoder and the enhancement encoding module based on the detection loss determined by the prediction result and the first ground truth label of the current image, the execution unit 52 is further configured to input the second visual feature and the text question corresponding to the current image into a large language model to obtain an estimated answer to the text question corresponding to the current image; determine a second generation loss based on the estimated answer and the ground truth answer to the text question corresponding to the current image; and adjust the parameters of the visual encoder, the enhancement encoding module, and the large language model based on the detection loss and the second generation loss.
[0069] According to an embodiment of the present disclosure, before adjusting the parameters in the visual encoder and the enhancement encoding module based on the detection loss determined by the prediction result and the first ground truth label of the current image, for each object included in the current image, the execution unit 52 is further configured to determine the matching relationship between the prediction result of the current object and the ground truth category and ground truth location information of the current object in the first ground truth label of the current image; and determine the detection loss based on the matching relationship of each object and the first ground truth label.
[0070] According to an embodiment of the present disclosure, the multimodal large model further includes an image projection layer and a text projection layer. The execution unit 52 is further configured to input the second visual feature into the image projection layer to obtain a third visual feature; input the text question in the text-image pair where the current image is located into the text projection layer to obtain a first text feature; splice the third visual feature and the first text feature to obtain a first splicing result; and input the first splicing result into the large language model to obtain an estimated answer to the text question in the text-image pair where the current image is located. Figure 6 is a block diagram of an image processing device based on a multimodal large model showing an embodiment of the present disclosure. As Figure 6 shown, the multimodal large model includes a visual encoder, an enhancement encoding module, and a large language model. The device includes a first acquisition unit 60, a second acquisition unit 62, and a third acquisition unit 64.
[0071] A first acquisition unit 60 is configured to input an image to be processed into a vision encoder to obtain first vision features; a second acquisition unit 62 is configured to input the first vision features into an enhancement encoding module to obtain second vision features; a third acquisition unit 64 is configured to input the second vision features and a text question in a text-image pair where the image to be processed is located into a large language model to obtain an answer to the text question, wherein the parameters of the vision encoder, the enhancement encoding module, and the large language model in the multimodal large model are obtained by the method described above.
[0072] According to an embodiment of the present disclosure, there is provided a computer-readable storage medium storing instructions, wherein when the instructions are run by at least one computing device, the at least one computing device is caused to execute the method for enhancing the fine-grained perception ability of the multimodal large model or the image processing method based on the multimodal large model in any of the above embodiments.
[0073] According to an embodiment of the present disclosure, there is provided a system including at least one computing device and at least one storage device storing instructions, wherein when the instructions are run by at least one computing device, the at least one computing device is caused to execute the method for enhancing the fine-grained perception ability of the multimodal large model or the image processing method based on the multimodal large model in any of the above embodiments.
[0074] In another general aspect, there is provided a computer program product including computer instructions that, when executed by a processor, implement the method for enhancing the fine-grained perception ability of the multimodal large model or the image processing method based on the multimodal large model in any of the above embodiments.
[0075] Although some embodiments of the present disclosure have been shown and described, those skilled in the art should understand that these embodiments can be modified without departing from the principles and spirit of the present disclosure defined by the claims and their equivalents.
Claims
1. A method for enhancing the fine-grained perception ability of a multi-modal large model, characterized in that The multimodal large model includes a visual encoder, an enhancement encoding module, and a large language model. The method includes: Obtain a target detection dataset and a text-image pair dataset. Among them, the target detection dataset includes multiple first images and the first ground truth labels of each first image. The text-image pair dataset includes multiple text-image pairs and the second ground truth labels of each text-image pair. The first ground truth labels include the true categories and true location information of the objects contained in the corresponding first images. The second ground truth labels include the true answers to the text questions in the corresponding text-image pairs. For each image in the target detection dataset and the text-image pair dataset, perform the following processing: Input the current image into the visual encoder to obtain a first visual feature. Input the first visual feature into the enhancement encoding module to obtain a second visual feature. In response to the current image belonging to the target detection dataset, perform target detection processing on the current image based on the second visual feature to obtain a prediction result corresponding to the current image. Adjust the parameters of the visual encoder and the enhancement encoding module based on the detection loss determined by the prediction result and the first ground truth label of the current image. In response to the current image belonging to the text-image pair dataset, input the second visual feature and the text question in the text-image pair where the current image is located into the large language model to obtain an estimated answer to the text question. Adjust the parameters of the visual encoder, the enhancement encoding module, and the large language model based on the first generation loss determined by the estimated answer and the second ground truth label of the text-image pair where the current image is located.
2. The method according to claim 1, wherein The performing target detection processing on the current image based on the second visual feature to obtain a prediction result corresponding to the current image includes: Perform cross-attention processing on the second visual feature and a preset query vector to obtain a cross-processed vector. The query vector indicates querying the category and location information of the objects contained in the current image from the second visual feature. Input the cross-processed vector into a decoder to obtain a decoded feature. Input the decoded feature into a feed-forward neural network to obtain a prediction result corresponding to the current image.
3. The method according to claim 1, characterized in that When the target detection dataset further includes the text question corresponding to the current image and the first ground truth label further includes the true answer to the text question, before adjusting the parameters of the visual encoder and the enhancement encoding module based on the detection loss determined by the prediction result and the first ground truth label of the current image, it further includes: Input the second visual feature and the text question corresponding to the current image into the large language model to obtain an estimated answer to the text question corresponding to the current image. Determine a second generation loss based on the estimated answer and the true answer to the text question corresponding to the current image. Among them, adjusting the parameters of the visual encoder and the enhancement encoding module based on the detection loss determined by the prediction result and the first ground truth label of the current image includes: Adjust the parameters of the visual encoder, the enhancement encoding module, and the large language model based on the detection loss and the second generation loss.
4. The method according to claim 1, wherein Before adjusting the parameters of the visual encoder and the enhancement encoding module based on the detection loss determined from the prediction result and the first ground truth label of the current image, it further includes: For each object included in the current image, determine the matching relationship between the prediction result of the current object and the true class and true location information of the current object in the first ground truth label of the current image; Based on the matching relationship of each object and the first ground truth label, determine the detection loss.
5. The method according to claim 1, characterized in that, The multimodal large model further includes an image projection layer and a text projection layer. Among them, the step of inputting the second visual feature and the text question in the text-image pair where the current image is located into the large language model to obtain the estimated answer to the text question includes: Input the second visual feature into the image projection layer to obtain a third visual feature; Input the text question in the text-image pair where the current image is located into the text projection layer to obtain a first text feature; Concatenate the third visual feature and the first text feature to obtain a first concatenated result; Input the first concatenated result into the large language model to obtain the estimated answer to the text question in the text-image pair where the current image is located.
6. An image processing method based on a multimodal large model, characterized in that, The multimodal large model includes a visual encoder, an enhancement encoding module, and a large language model. The method includes: Input the image to be processed into the visual encoder to obtain a first visual feature; Input the first visual feature into the enhancement encoding module to obtain a second visual feature; Input the second visual feature and the text question in the text-image pair where the image to be processed is located into the large language model to obtain the answer to the text question. Among them, the parameters of the visual encoder, the enhancement encoding module, and the large language model in the multimodal large model are obtained by the method described in any one of claims 1 to 5.
7. An apparatus for enhancing the fine-grained perception ability of a multi-modal large model, characterized in that, The multimodal large model includes a visual encoder, an enhancement encoding module, and a large language model. The device includes: An acquisition unit configured to acquire a target detection data set and a text-image pair data set. Among them, the target detection data set includes a plurality of first images and the first ground truth label of each first image, the text-image pair data set includes a plurality of text-image pairs and the second ground truth label of each text-image pair, the first ground truth label includes the true class and true location information of the objects included in the corresponding first image, and the second ground truth label includes the true answer to the text question in the corresponding text-image pair; An execution unit configured to perform the following processing for each image in the target detection data set and the text-image pair data set: Input the current image into the visual encoder to obtain a first visual feature; Input the first visual feature into the enhancement encoding module to obtain a second visual feature; In response to the current image belonging to the target detection data set, perform target detection processing on the current image based on the second visual feature to obtain a prediction result corresponding to the current image; adjust the parameters of the visual encoder and the enhancement coding module based on the detection loss determined by the prediction result and the first ground truth label of the current image; In response to the current image belonging to the text-image pair data set, input the second visual feature and the text question in the text-image pair where the current image is located into a large language model to obtain a predicted answer to the text question; adjust the parameters of the visual encoder, the enhancement coding module, and the large language model based on the first generation loss determined by the predicted answer and the second ground truth label of the text-image pair where the current image is located.
8. An image processing device based on a multimodal large model, characterized in that, The multimodal large model includes a visual encoder, an enhancement coding module, and a large language model, and the device includes: A first acquisition unit configured to input an image to be processed into the visual encoder to obtain a first visual feature; A second acquisition unit configured to input the first visual feature into the enhancement coding module to obtain a second visual feature; A third acquisition unit configured to input the second visual feature and the text question in the text-image pair where the image to be processed is located into a large language model to obtain an answer to the text question, wherein the parameters of the visual encoder, the enhancement coding module, and the large language model in the multimodal large model are obtained by the method according to any one of claims 1 to 5.
9. A computer-readable storage medium for storing instructions, characterized in that, When the instruction is run by at least one computing device, cause the at least one computing device to execute the method for enhancing the fine-grained perception ability of the multimodal large model according to any one of claims 1 to 5 or the image processing method based on the multimodal large model according to claim 6.
10. A system comprising at least one computing device and at least one storage device storing instructions, characterized in that, When the instruction is run by the at least one computing device, cause the at least one computing device to execute the method for enhancing the fine-grained perception ability of the multimodal large model according to any one of claims 1 to 5 or the image processing method based on the multimodal large model according to claim 6.
11. A computer program product comprising computer instructions, characterized in that, When the computer instruction is executed by a processor, implement the method for enhancing the fine-grained perception ability of the multimodal large model according to any one of claims 1 to 5 or the image processing method based on the multimodal large model according to claim 6.
Citation Information
Patent Citations
Image report generation method and model training method
CN118072898A
Conversational target positioning method and device based on arbitrary granularity text input
CN118350464A
Multi-modal model visual perception ability enhancement method and device, and medium
CN119809925A
KR20250070982A
Cited By
Weak supervision scene understanding method, system and equipment for multi-modal information interaction
CN121305053A