A multimodal large language model training and visual task processing method, device and medium

CN119202157BActive Publication Date: 2026-09-25NINGBO DIGITAL TWIN (EASTERN UNIV OF TECH) RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411198395.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2026-09-25
Estimated Expiration
2044-08-29

AI Technical Summary

Technical Problem

然而,现有的多模态大模型往往在图像理解的基础上往往只能实现一种下游视觉任务(目标检测或图像分割),或是只能通过大模型本身的Agent能力调用下游各种现有模型,缺乏端到端的训练和实现能力,并且缺乏一种显式、统一的视觉表征,难以实现多视觉任务的整合

Benefits of technology

[0038](1)本发明可以通过一种大语言模型实现多种常见视觉任务的整合,如检测、分割、关键点检测、深度估计等视觉任务,并且支持低成本的任务扩展。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119202157B_ABST
    Figure CN119202157B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of multimodal large language model training and visual task processing method, equipment, medium, wherein model training method includes the following steps: constructing the visual task instruction fine-tuning scheme of multimodal large language model;For different visual tasks, using existing dataset based on visual task instruction fine-tuning scheme constructs multimodal question and answer dataset;Build the multimodal large language model of supporting multiple user visual interaction mode understanding dialogue and visual feature decoding, using multimodal question and answer dataset trains the multimodal large language model.Compared with prior art, the present application has the advantages of being able to process multiple visual tasks simultaneously, fast inference speed, low training cost, supporting task expansion, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large language model technology, and in particular to a method, device, and medium for training multimodal large language models and processing visual tasks. Background Technology

[0002] Large-scale multimodal models have brought breakthroughs to various advanced visual understanding tasks by compressing image and language information into a single autoregressive model based on the Transformer architecture. Existing work includes multimodal large language models similar to the LLAVA architecture, which construct linear layers to realize the mapping relationship between images and text, enabling large language models to perform image understanding tasks; multimodal large models such as Shikra, Ferret, and Kosmos-2 utilize techniques similar to Pix2Seq to convert image location information into sequence information by adding new words from an existing vocabulary or directly inputting text, thereby achieving object detection; and multimodal large models such as GLaMM, PSALM, and GroundHog utilize multiple visual encoders and SAM models or Mask2Former architectures as decoders to achieve image segmentation based on image understanding. However, existing multimodal large models often only achieve one downstream visual task (object detection or image segmentation) based on image understanding, or can only call various existing downstream models through the agent capabilities of the large model itself. They lack end-to-end training and implementation capabilities, and lack an explicit and unified visual representation, making it difficult to integrate multiple visual tasks. Summary of the Invention

[0003] The purpose of this invention is to provide a method, device, and medium for training a multimodal large language model and processing visual tasks. By constructing a new data logic, an explicit visual representation is formed at the model level, enabling the multimodal large language model to simultaneously achieve functions such as image captioning, visual question answering (VQA), object detection, referring expression generation (REG), referring expression comprehension (REC), referring expression segmentation (RES), instance segmentation, grounded conversation generation detection (GCG detection), grounded conversation generation segmentation (GCG segmentation), keypoint detection, and monocular depth estimation.

[0004] The objective of this invention can be achieved through the following technical solutions:

[0005] A method for training a multimodal large language model for visual task processing includes the following steps:

[0006] A scheme for fine-tuning visual task instructions in constructing a multimodal large language model;

[0007] For different visual tasks, a multimodal question-answering dataset is constructed by using existing datasets and fine-tuning the scheme based on visual task instructions;

[0008] A multimodal large language model supporting understanding dialogue and visual feature decoding for multiple user visual interaction methods is constructed, and the multimodal large language model is trained using a multimodal question-answering dataset.

[0009] The visual task instruction fine-tuning scheme for constructing a multimodal large language model includes the following steps:

[0010] Different visual tasks are constructed, and a plurality of question templates with different expression modes are designed for each task based on GPT-4. In the construction process of corresponding task data, questions are randomly extracted from the templates to construct dialogues. The visual tasks include image description, visual question answering, referring expression comprehension, referring expression generation, referring expression segmentation, and dialogue and localization;

[0011] Design special tokens for referring to the visual prompts input by the user, and configure the text encoder not to tokenize the special tokens;

[0012] Construct a visual decoding chain-of-thought, which is used to generate chain-of-thought content to further generate decoded text based on the chain-of-thought content, and the generated chain-of-thought content includes at least one or more of phrases, units and quantities;

[0013] A general and extensible model response scheme is constructed for visual decoding tasks. Specifically, in the model response scheme, "<phrase>" and "< / phrase>" are used to wrap the categories predicted by the model, "<unit>" and "< / unit>" are used to wrap the decoding types of the model, with " <ref>"Refers to the target of the model decoding and is used for subsequent decoding."

[0014] The construction of the multimodal question-answering dataset is specifically as follows:

[0015] In the image description task, questions corresponding to the task are selected from the constructed question template, and the annotations of image content in the existing image description dataset are used as answers to form an image description question-and-answer dialogue.

[0016] In visual question answering tasks, question-answer pairs from existing visual question answering datasets are used to construct image-based visual question answering dialogues.

[0017] In the pointing expression and understanding task, questions corresponding to the task are selected from the constructed question template, special words designed are used to encode the user's visual input, and the annotations of image content in the existing pointing expression and understanding dataset are used as answers to form an image description question-and-answer dialogue.

[0018] In the index expression segmentation task, questions corresponding to the task are selected from the constructed question template, user visual input is encoded using specially designed words, and answers are generated based on the existing index expression segmentation dataset using the constructed model answer scheme, thus forming an image segmentation content question-and-answer dialogue.

[0019] In the index expression generation task, the corresponding task question is selected from the constructed question template, the user's visual input is encoded using specially designed words, and the constructed model answering scheme is used to generate the answer based on the existing index expression generation dataset, thus forming an image object detection content question-and-answer dialogue.

[0020] In the dialogue and localization task, questions corresponding to the task are selected from the constructed question template, and the constructed model answering scheme is used to generate answers based on the existing dialogue and localization dataset, thus forming image dialogue and localization question answering.

[0021] The multimodal large language model includes a visual encoder, a visual cue encoder, a visual feature mapper, a basic language model, a visual representation optimizer, and a visual content decoder.

[0022] The visual encoder adopts the CLIP-VIT model, which encodes global visual features by uniformly scaling the input image to a preset size.

[0023] The visual cue encoder converts different types of visual cue inputs into masks of different shapes, pools the masks to the same size as the global visual features, and performs point-by-point multiplication with the global visual features.

[0024] The visual feature mapper maps the features output by the visual encoder and the visual cue encoder to the same dimension of the text features of the multimodal large language model, thereby aligning multimodal information.

[0025] The underlying language model adopts the Vicuna-7B model;

[0026] The visual representation optimizer optimizes the global visual features obtained from the visual encoder and reconstructs the decoding region of the current task by introducing a learnable query token at the end of the problem, thereby restoring large-scale details.

[0027] The visual content decoder adopts different designs for different visual tasks, and combines global visual features and visual representations optimized by the visual representation optimizer to decode the decoding target in the model decoding scheme.

[0028] The training of the multimodal large language model is divided into two stages. The first stage uses text and image data to train the basic language model and the visual feature mapper to align visual and text features. The second stage trains all models in the multimodal large language model except for the visual encoder.

[0029] A visual task processing method based on a multimodal large language model includes the following steps:

[0030] The trained multimodal large language model is obtained based on the model training method described above;

[0031] Obtain visual task information input by the user;

[0032] The trained multimodal large language model is used for visual task processing, and the processing results are output.

[0033] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the model training method as described above.

[0034] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the visual task processing method as described above.

[0035] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the model training method as described above.

[0036] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the visual task processing method as described above.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] (1) This invention can integrate a variety of common visual tasks, such as detection, segmentation, key point detection, depth estimation and other visual tasks, through a large language model, and supports low-cost task expansion.

[0039] (2) The present invention constructs a large number of directly usable multimodal question-answering datasets, and the data can be rapidly and continuously expanded according to the construction scheme.

[0040] (3) The present invention optimizes the performance of visual representation, encoding and other processes, making the model inference faster and the training cost lower. Attached Figure Description

[0041] Figure 1 This is a flowchart of the method of the present invention;

[0042] Figure 2 This is a schematic diagram of the multimodal large language model structure of the present invention;

[0043] Figure 3 This is a schematic diagram of the visual cue encoder structure of the present invention;

[0044] Figure 4 This is a schematic diagram of the visual representation optimizer structure of the present invention. Detailed Implementation

[0045] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0046] This embodiment provides a method for training a multimodal large language model for visual task processing, such as... Figure 1 As shown, it includes the following steps:

[0047] S1, a scheme for fine-tuning visual task instructions for constructing a multimodal large language model.

[0048] This embodiment constructs an instruction tuning scheme that ensures a multimodal large language model can perform common visual tasks such as detection, segmentation, keypoint detection, and depth map estimation during dialogue, and supports extension to other vision-language multimodal tasks. Specifically, it includes the following steps:

[0049] S101 constructs various visual tasks, including image description, visual question answering, referring expression comprehension (REC), referring expression generation (REG), referring expression segmentation (RES), and grounded conversation generation (GCG). Based on GPT-4, 100-150 question templates with different expressions are designed for each task. During the data construction process for the corresponding task, questions are randomly selected from the templates to construct the dialogue.

[0050] S102 provides a special token "[VPT]" to represent visual cues input by the user, such as points, boxes, masks, and scribbles. This token is then set as a separate token in the text encoder and is no longer segmented.

[0051] S103, Construct a visual decoding thought chain. Based on the current problem, the model needs to first generate a thought chain according to its content format, and then use this content to generate specific decoded text. The thought chain content generated by the model needs to include at least one or more of the following: phrase, unit, and quantity.

[0052] In one embodiment, the thought process for the instance segmentation task is illustrated below:

[0053] <task>

[0054] Unit decode(True).Class name,target unit and number:

[0055] -Name:a small stone building Unit: <unit> mask< / unit> Num:1

[0056] -Name:a serene japanese garden with a rock garden Unit: <unit> mask< / unit> Num:1

[0057] < / task>

[0058] S104, for visual decoding tasks such as Box, Mask, Keypoint, and Depth Map, constructs a general and scalable model response solution. Specifically, it designs... <phrase> "and"< / phrase> "The categories used for prediction by the front and back package model," <unit> "and"< / unit> "Used to wrap the model decoding type," <ref>"Used to refer to the target of the model decoding and used for subsequent decoding. In this case, an example of the response to the instance segmentation task can be represented as: <phrase> a small stone building< / phrase> ( <unit> mask< / unit> [0] <ref> ), <phrase> a serene japanese garden with arock garden< / phrase> ( <unit> mask< / unit> [0] <ref>).

[0059] S2, for different visual tasks, constructs a multimodal question-answering dataset by using existing datasets and fine-tuning schemes based on visual task instructions.

[0060] Specifically:

[0061] In S201, in the image description task, datasets such as COCO2014, llava-pretrain, llava-instruct, Grand caption, and GRIT caption are used as existing image description datasets. Questions corresponding to the task are selected from the question template constructed in step S101, and the annotations of image content in the existing image description datasets are used as answers to form an image description question-and-answer dialogue.

[0062] S202, in the visual question answering task, uses the PointQA, VQAv2, VQAE, and VQAX datasets as existing visual question answering datasets, and uses question-answer pairs from existing visual question answering datasets to construct image-based visual question answering dialogues.

[0063] S203, in the pointing expression comprehension task, the existing pointing expression comprehension datasets such as RefCOCO, RefCOCO+, RefCOCOg, GRIT REC, Grand REC, and COCO Interact Box are used. Questions corresponding to the task are selected from the question template constructed in step S101. The user's visual input is encoded using special words designed in step S102. The answers are based on the annotations of image content in the existing pointing expression comprehension dataset, thus forming an image description question-and-answer dialogue.

[0064] S204. In the indexing expression segmentation task, datasets such as Grand RES and COCO Interact Mask are used as existing indexing expression segmentation datasets. Questions corresponding to the task are selected from the question template constructed in step S101. The user's visual input is encoded using special words designed in step S102. The model answering scheme constructed in step S104 is used to generate answers based on the existing indexing expression segmentation dataset, thus forming an image segmentation content question-and-answer dialogue.

[0065] S205, in the index expression generation task, GRIT REG, Grand REG, LLAVAg REG, COCOInteract REG, RefCOCO, RefCOCO+, RefCOCOg and other datasets are used as existing index expression generation datasets. The corresponding task question is selected from the question template constructed in step S101. The user's visual input is encoded using special words designed in step S102. The answer is generated based on the existing index expression generation dataset using the model answering scheme constructed in step S104, thus forming an image object detection content question-and-answer dialogue.

[0066] S206. In the dialogue and localization task, LLAVAg GCG, PNG, GRIT GCG, Grand GCG and other datasets are used as existing dialogue and localization datasets. Questions corresponding to the task are selected from the question template constructed in step S101. The model answering scheme constructed in step S104 is used to generate answers based on the existing dialogue and localization datasets, thus forming image dialogue and localization question answering.

[0067] S3. Construct a multimodal large language model that can support understanding dialogue and decoding visual features such as bounding boxes, masks, keypoints, and depth maps, and train the multimodal large language model using a multimodal question-answering dataset.

[0068] like Figure 2 As shown, the multimodal large language model includes a visual encoder, a visual cue encoder, a visual feature mapper, a basic language model, a visual representation optimizer, and a visual content decoder.

[0069] 1. Visual Encoder: In this embodiment, the CLIP-VIT model is used as the visual encoder. By uniformly scaling the input image to a size of 336*336, it is encoded to obtain a global visual feature with a shape of [576, 1024], where 576 represents the feature length, which is equivalent to flattening the 24*24 image features.

[0070] 2. Visual Cue Encoder: For different types of visual cue inputs, such as points, boxes, masks, and scribbles, the visual cue encoder converts them into masks of different shapes. Its structure is as follows: Figure 3 As shown in the diagram. Specifically, for a Point, it is converted into a circular mask with a radius of 10 centered at the point; for a Scribble, it is converted into a bar mask with a width of 10 and unchanged shape; and for a Box, it is converted into a rectangular mask of the same area. A mask patch pooling scheme is used, pooling the mask to the same size as the global visual feature, i.e., 24*24, and then multiplying it point-by-point with the global visual feature.

[0071] 3. Visual Feature Mapper: The visual feature mapper maps the features output by the visual encoder and the visual cue encoder to the same dimension of the text features of the multimodal large language model, thus aligning the multimodal information.

[0072] 4. Basic Language Model: In this embodiment, the basic language model adopts the Vicuna-7B model.

[0073] 5. Visual Representation Optimizer: The global features acquired in the visual encoder suffer from a loss of detail for fine-grained visual decoding tasks (such as Mask, Depth Map, etc.). This embodiment employs a visual representation optimizer to optimize the global visual features acquired in the visual encoder. By introducing a learnable query token at the end of the problem, the decoding region of the current task is reconstructed, restoring large-scale details. Its structure is as follows: Figure 4 As shown.

[0074] 6. Visual Content Decoder: For different visual tasks, the visual content decoder employs different designs, but uniformly uses the model-generated query tokens as learnable tokens for subsequent decoding. The visual content decoder combines global visual features and visual representations optimized by the visual representation optimizer to decode the decoding targets in the model's decoding scheme. For Box and Keypoint decoding tasks, the visual content decoder uses a Detr-Decoder, taking CLIP-VIT features as visual features and fusing them with the visual representations optimized by the visual representation optimizer to decode the... <ref>"Decoding is performed; for Mask and DepthMap decoding tasks, the visual content decoder uses MaskFormer-Decoder, taking CLIP-VIT features as visual features, and after fusing the visual representation optimized by the visual representation optimizer, it performs..." <ref>"Decode."

[0075] The training of the multimodal large language model is divided into two stages. The first stage uses text and image data to train the basic language model and visual feature mapper to align visual and text features. The second stage trains all models in the multimodal large language model except for the visual encoder. The dataset constructed in step S2 ensures that the large language model has multimodal understanding capabilities and responds to visual decoding tasks in a predefined paradigm.

[0076] This embodiment also provides a visual task processing method based on a multimodal large language model, including the following steps:

[0077] The trained multimodal large language model is obtained based on the model training method described above;

[0078] Obtain visual task information input by the user;

[0079] The trained multimodal large language model is used for visual task processing, and the processing results are output.

[0080] In one embodiment, the electronic device includes a computing unit that can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit into a random access memory (RAM). The RAM may also store various programs and data required for device operation. The computing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0081] Multiple components in an electronic device are connected to an I / O interface, including: input units such as keyboards and mice; output units such as various types of displays and speakers; storage units such as disks and optical discs; and communication units such as network interface cards (NICs), modems, and wireless transceivers. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0082] The computing unit can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing units include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit performs the various methods and processes described above, such as the model training methods / visual task processing methods described above. For example, in some embodiments, the model training methods / visual task processing methods described above can be implemented as computer software programs tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed on the device via ROM and / or a communication unit. When the computer program is loaded into RAM and executed by the computing unit, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit can be configured to perform the methods described above by any other suitable means (e.g., by means of firmware).

[0083] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0084] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0085] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0086] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.< / ref> < / ref> < / ref> < / ref> < / ref> < / ref>

Claims

1. A method for training a multimodal large language model for visual task processing, characterized in that, Includes the following steps: A scheme for fine-tuning visual task instructions in constructing a multimodal large language model; For different visual tasks, a multimodal question-answering dataset is constructed by using existing datasets and fine-tuning the scheme based on visual task instructions; Construct a multimodal large language model that supports understanding dialogue and visual feature decoding for multiple user visual interaction methods, and train the multimodal large language model using a multimodal question-answering dataset; The visual task instruction fine-tuning scheme for constructing a multimodal large language model includes the following steps: Different visual tasks are constructed, and multiple question templates with different expressions are designed for each task based on GPT-4. During the construction of corresponding task data, questions are randomly extracted from the templates to construct dialogues. The visual tasks include image description, visual question answering, pointer expression understanding, pointer expression generation, pointer expression segmentation, and dialogue and localization. Design special words to refer to the visual cues input by the user, and set the text encoder not to segment the special words; Construct a visual decoding mind chain to generate mind chain content for further generating decoded text based on the mind chain content. The generated mind chain content includes at least one or more of phrases, units, and quantities. A general and scalable model answering scheme is constructed for visual decoding tasks. Specifically, in the model answering scheme, the category predicted by the model is wrapped before and after by "<phrase>" and "< / phrase>", the decoding type of the model is wrapped by "<unit>" and "< / unit>", <ref> "Refers to the target of the model decoding and is used for subsequent decoding;< / ref> The multimodal large language model includes a visual encoder, a visual cue encoder, a visual feature mapper, a basic language model, a visual representation optimizer, and a visual content decoder. The visual encoder adopts the CLIP-VIT model, which encodes global visual features by uniformly scaling the input image to a preset size. The visual cue encoder converts different types of visual cue inputs into masks of different shapes, pools the masks to the same size as the global visual features, and performs point-by-point multiplication with the global visual features. The visual feature mapper maps the features output by the visual encoder and the visual cue encoder to the same dimension of the text features of the multimodal large language model, thereby aligning multimodal information. The underlying language model adopts the Vicuna-7B model; The visual representation optimizer optimizes the global visual features obtained from the visual encoder and reconstructs the decoding region of the current task by introducing a learnable query token at the end of the problem, thereby restoring large-scale details. The visual content decoder adopts different designs for different visual tasks, and combines global visual features and visual representations optimized by the visual representation optimizer to decode the decoding target in the model decoding scheme.

2. The scalable visual task processing method based on a multimodal large language model according to claim 1, characterized in that, The construction of the multimodal question-answering dataset is specifically as follows: In the image description task, questions corresponding to the task are selected from the constructed question template, and the annotations of image content in the existing image description dataset are used as answers to form an image description question-and-answer dialogue. In visual question answering tasks, question-answer pairs from existing visual question answering datasets are used to construct image-based visual question answering dialogues. In the pointing expression and understanding task, questions corresponding to the task are selected from the constructed question template, special words designed are used to encode the user's visual input, and the annotations of image content in the existing pointing expression and understanding dataset are used as answers to form an image description question-and-answer dialogue. In the index expression segmentation task, questions corresponding to the task are selected from the constructed question template, user visual input is encoded using specially designed words, and answers are generated based on the existing index expression segmentation dataset using the constructed model answer scheme, thus forming an image segmentation content question-and-answer dialogue. In the index expression generation task, the corresponding task question is selected from the constructed question template, the user's visual input is encoded using specially designed words, and the constructed model answering scheme is used to generate the answer based on the existing index expression generation dataset, thus forming an image object detection content question-and-answer dialogue. In the dialogue and localization task, questions corresponding to the task are selected from the constructed question template, and the constructed model answering scheme is used to generate answers based on the existing dialogue and localization dataset, thus forming image dialogue and localization question answering.

3. The scalable visual task processing method based on a multimodal large language model according to claim 1, characterized in that, The training of the multimodal large language model is divided into two stages. The first stage uses text and image data to train the basic language model and the visual feature mapper to align visual and text features. The second stage trains all models in the multimodal large language model except for the visual encoder.

4. A visual task processing method based on a multimodal large language model, characterized in that, Includes the following steps: The trained multimodal large language model is obtained based on the method described in any one of claims 1-3; Obtain visual task information input by the user; The trained multimodal large language model is used for visual task processing, and the processing results are output.

5. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 3.

6. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the program, it implements the method as described in claim 4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 3.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in claim 4.